Build an n8n Market Research Pipeline That Finds Trends Early
n8n market research automation scrapes Hacker News and Reddit, uses AI to identify emerging niches and gaps, then emails a weekly briefing. Browse templates.
By the time a trend appears in a newsletter, two or three funded teams have been building for 90 days. n8n market research automation closes that gap by running the same research loop on a schedule — no subscription fees, no two-hour forum sessions on Sunday night.
The setup is a single workflow: fetch top posts from Hacker News and Reddit, score them by engagement, aggregate the signal, send it through an OpenAI analysis prompt, and deliver a structured briefing before Monday's standup. All of it runs inside n8n. No external scraping service required.
There are edge cases that trip up first-time builds. Here's what the production version looks like.
What n8n market research automation can track
An automated research pipeline handles more than weekly trends. With small changes to the HTTP Request nodes, the same pattern covers:
- High-signal Hacker News stories filtered by upvote threshold and creation date
- Weekly top posts from any public subreddit — r/SaaS, r/startups, r/entrepreneur, or niche communities in a specific product category
- Pain-point threads across multiple subreddits in parallel, merged into one ranked list
- Competitor mentions: posts and comments containing a specific product name
- Keyword frequency trends — how often "I wish X did Y" appears in recent discussions
- ProductHunt launches tagged with a relevant category, via their public RSS feed
It's not a replacement for customer interviews. But it's a reliable early-warning layer that most small teams skip because the manual version doesn't scale.
The weekly market research pipeline
Schedule Trigger (Mon 08:00, cron: 0 8 * * 1)
→ Code (calculate Unix timestamp for 7 days ago)
→ HTTP Request (HN Algolia — stories, points>50, created_at > last_week)
→ HTTP Request (Reddit — top.json, t=week, limit=25)
→ Code (merge, normalize, sort by aggregate score, slice to 60)
→ OpenAI (generate trend briefing from aggregated titles)
→ Gmail (send formatted briefing to inbox)
→ Google Sheets (append timestamp + summary to historical log)
The Code node before the HTTP Requests calculates Math.floor(Date.now() / 1000) - 604800 and sets it as the created_at_i filter value for Algolia. Hard-coding a timestamp breaks the workflow within a week, so this step isn't optional.
Step-by-step breakdown
1. Collect — Pulling high-signal posts
The HN Algolia endpoint:
https://hn.algolia.com/api/v1/search?tags=story&numericFilters=created_at_i>LAST_WEEK,points>50&hitsPerPage=100
The hitsPerPage default is 20. Bump it to 100 — at 50+ points, the result set is clean, and 100 stories per week gives the AI more material to identify real patterns.
Reddit's endpoint (no authentication required for public subreddits):
https://www.reddit.com/r/SaaS/top.json?limit=25&t=week
For multiple subreddits, add separate HTTP Request nodes and merge in the aggregation step. Reddit's unauthenticated rate limit is 60 requests per minute per IP — not an issue for a weekly run, but add a Wait node (2 seconds between requests) if you expand to 10 or more sources.
2. Normalize — Merging and scoring
The Code node (v2, jsCode parameter — not the deprecated Function node from pre-1.0 n8n) merges both sources and sorts by engagement:
const hn = $('HTTP Request - HN').first().json.hits || [];
const reddit = $('HTTP Request - Reddit').first().json.data?.children || [];
const hnItems = hn.map(h => ({ src: 'HN', title: h.title, score: h.points }));
const redItems = reddit.map(c => ({ src: 'Reddit', title: c.data.title, score: c.data.score }));
const merged = [...hnItems, ...redItems].sort((a, b) => b.score - a.score);
const lines = merged.slice(0, 60).map(t => `[${t.src}:${t.score}] ${t.title}`).join('\n');
return [{ json: { titles: lines } }];
Slicing to 60 entries keeps the prompt under 2,000 tokens for GPT-4o-mini, which is enough for a solid briefing without hitting the model's context ceiling mid-run.
3. Analyze — What GPT makes of the signal
The @n8n/n8n-nodes-langchain.openAi v2.1 node takes the aggregated titles and returns output at $json.output[0].content[0].text. A prompt structure that produces consistently useful results:
Analyze these Hacker News and Reddit posts from the past week.
Return exactly:
1. Five emerging themes (name + one sentence each)
2. Three market gaps (problems discussed without clear solutions)
3. Two product opportunity descriptions (specific and buildable)
Be specific. Reference actual post titles where relevant. No generic observations.
Posts:
[paste the aggregated titles string here]
The "reference actual post titles" instruction changes the output quality dramatically. Without it, the model returns abstractions like "AI tooling is growing." With it, you get observations like "Developer friction with LLM token costs: multiple posts from indie builders describe abandoning GPT-4o after a single workflow run cost $0.40, with no cost-control tooling in their stack." That's actionable.
4. Deliver — Email and log
The Gmail node sends the briefing as plain text or light HTML. Keep it readable — one section per theme, no markdown headers in the email body (they don't render in plain-text clients).
The Google Sheets append step logs timestamp, total posts analyzed, and briefing summary. After 12 weeks, that log is a trend index. You can grep it for when a specific topic first appeared, or spot which themes keep recurring without ever converting into funded products.
5. Extend — Filters and domain targeting
Two low-effort extensions that change the output significantly:
Topic filtering in the Code node: before passing titles to the AI, filter merged to only entries whose titles contain strings like "pricing", "alternative", or "too expensive". Pain-point language. The AI briefing shifts from general trends to specific switching signals in your product category.
Domain-specific subreddits: r/webdev, r/devops, r/datascience, and r/marketing each surface different demand signals than r/SaaS. Add them as separate HTTP Request nodes and merge into the same aggregation step. The weekly briefing becomes a multi-community picture rather than a single-community snapshot.
Implementation patterns
Pattern 1: Weekly general trend digest
The full pipeline above, running every Monday at 08:00. No filters applied before the AI step. The model categorizes everything from infra tooling complaints to product launch traction to pricing debates. Works well when you're exploring rather than validating a specific hypothesis.
Pattern 2: Pain-point targeting for a specific product category
Apply a keyword filter in the Code node before the AI step. If the product is a scheduling tool, filter merged to only titles containing "calendar", "booking", "scheduling", or "appointment" before the prompt. Smaller input, more focused output — and the AI can go deeper because it's not trying to cover the entire product landscape in 500 tokens.
The most common cause of low-quality analysis is a noisy input set. Setting points>50 on HN and t=week on Reddit already filters out most noise, but if briefings still feel vague, try sorting by comment count instead of upvote score. A post with 40 upvotes and 180 comments signals active debate, not passive agreement. Debates reveal fractures in the market — that's where the interesting signal lives.
n8n nodes you'll use in every market research workflow
| Node | Version | Purpose |
|---|---|---|
| Schedule Trigger | v1.2 | Weekly cadence with 0 8 * * 1 cron expression |
| HTTP Request | v4.2 | Fetch HN Algolia and Reddit JSON endpoints |
| Code | v2 | Timestamp calculation; merge, sort, and slice post lists |
@n8n/n8n-nodes-langchain.openAi | v2.1 | Generate briefing; output at $json.output[0].content[0].text |
| Gmail | v2 | Send formatted briefing to inbox |
| Google Sheets | v4.4 | Append timestamped row to historical trend log |
Running your first n8n market research pipeline
You can have a working briefing in under an hour:
- Create a new workflow and add a Schedule Trigger with cron
0 8 * * 1. - Add a Code node with:
return [{ json: { lastWeek: Math.floor(Date.now() / 1000) - 604800 } }]. This is the timestamp the Algolia filter will use. - Add an HTTP Request node pointing at the HN Algolia endpoint above. Use
{{ $node["Code"].json.lastWeek }}as thecreated_at_ivalue. Run manually and confirm 50–100 posts are coming back. - Add a second HTTP Request for one Reddit subreddit. Verify the response structure looks right before adding more.
- Add the aggregation Code node. Merge, sort by score, slice to 60, join as a string.
- Add the OpenAI node. Paste the analysis prompt and run manually. If the output is too generic, tighten the prompt's "be specific" instruction or reduce the input slice further.
- Wire Gmail and Google Sheets at the end. Activate and wait for Monday.
If starting from a working template is faster than building from scratch, the Market Trend Analyzer template ships this pipeline pre-configured: HN collection with the right filters, Reddit aggregation, the OpenAI prompt, Gmail delivery, and Sheets logging. Setup is about 10 minutes.
Deploy the Market Trend Analyzer →The Market Trend Analyzer ships this end-to-end — a Schedule trigger fires each week, pulls HN and Reddit posts filtered by engagement score, passes them through a GPT-4o-mini analysis prompt, and emails a structured niche-and-gap briefing to your inbox. It's part of The Complete n8n Templates Bundle, a one-time lifetime license to the whole catalog (plus every template added later) if you run more than one of these automations.
The Competitor Price Intelligence template handles a related job: tracking competitor pricing changes on a schedule rather than open trend signals. Running both in parallel gives a fuller picture — what's bubbling up in communities and what competitors are quietly adjusting.
For building on the data layer, the Sheets log from weekly runs fits cleanly into the approach covered in n8n data pipeline automation: the historical log becomes a source table you can query for trend recurrence over time. And if competitor mentions are part of your research scope, n8n competitor monitoring automation covers the alert layer that sits on top of the same HTTP Request pattern used here.
Browse all n8n automation templates →Common questions
How does n8n scrape Hacker News for market research?
Can n8n analyze Reddit posts for trend monitoring?
What does the AI step output in an n8n market research workflow?
Get the workflow templates this guide is built on
Import-ready n8n JSON, step-by-step setup, and tested end-to-end. One-time payment, own it forever.
Get 3 tested n8n templates, free
The full customer package for three real catalog templates — workflow JSON, step-by-step setup guide, credential checklist. Built through the same live-instance release process as everything we sell. Plus new templates and automation guides in your inbox. No spam, unsubscribe anytime.
- 01Smart To-Do List ManagerPre-built n8n workflow template that automates productivity with OpenAI. Live in about 10 minutes.$14
- 02Email Follow-Up AutomatorPre-built n8n workflow template that automates crm with OpenAI. Live in about 15 minutes.$12
- 03Market Trend AnalyzerPre-built n8n workflow template that automates data processing with OpenAI. Live in about 10 minutes.$14
More automation guides

How to Automate Meta Ads Campaigns with n8n
The Manual Campaign Launch Problem Writing three Facebook ad variations from scratch is creative work, but barely. It's mostly rewriting the same benefit statement with slightly different hooks. Then…

How to Automate Appointment Reminders with n8n
A hair salon with 20 appointments on a Tuesday loses $240 in revenue from a single no-show if the average ticket is $80. That's one client, one empty hour, and no way to fill the slot on short notice.…

Automate Ad Performance Reporting with n8n Workflows
Most marketing teams lose the first two hours of Monday to the same ritual: open Google Ads, export yesterday's CSV, switch to Facebook Ads Manager, export another one, paste both into a spreadsheet,…