# /hk01/articles/search

`POST /api/hk01/articles/search`

Price: 20 credits

Search the full HK01 (香港01) archive — 1.2M articles and videos published since December 2015. Filter by keyword, channel, topic tag, video-only, sponsored-only, sponsored content type and publication date range. An empty keyword lists the archive newest-first. Returns headline, lead, URL, cover image, channel, categories, tags, authors and timestamps for every match.

## How to use it

Full-text search across the HK01 archive back to 2015-12. Every result is a live, published article whose id can be passed straight to hk01/articles under the same locale. Combine keyword with category (channel name such as 突發 or 財經快訊), tag (topic name such as 特朗普), is_video, is_sponsored, sponsor_type and published_after/published_before (unix seconds or ISO date) to slice the archive. Each item has id, article_title, description (lead), web_url, image, main_category/main_category_id, categories, tags, authors (name and newsroom email), is_video/is_sponsored/is_featured, sponsor_type (特約內容 advertorial or 導購文章 shopping guide — the reliable source of that classification), video_duration in seconds for video items and published_at/created_at/updated_at in unix seconds. With locale=zh-CN only the simplified edition (published since August 2025) is searched.

## Parameters

- `access-token` (string, required)

## Request body

- `timeout` (integer) — Max scrapping execution timeout (in seconds) (default: 300; min: 20; max: 1500)
- `keyword` (string) — Free-text query over headline, lead, tags, categories and author names. Leave empty to list the archive newest-first (default: ""; examples: "香港", "特朗普")
- `count` (integer, required) — Max number of articles to return (min: 1)
- `category` (string, nullable) — Restrict to a channel by its HK01 name (examples: "突發", "財經快訊", "即時娛樂")
- `tag` (string, nullable) — Restrict to a topic tag by name (examples: "特朗普", "香港樓市")
- `is_video` (boolean, nullable) — Keep only video (true) or only text (false) items
- `is_sponsored` (boolean, nullable) — Keep only sponsored (true) or only editorial (false) items
- `sponsor_type` (string, nullable) — Restrict to a sponsored content type (one of: "特約內容", "導購文章")
- `published_after` (integer, nullable) — Keep articles published at or after this moment (unix seconds or ISO-8601 date)
- `published_before` (integer, nullable) — Keep articles published before this moment (unix seconds or ISO-8601 date)
- `locale` (string) — Edition the results belong to. The simplified-Chinese edition only carries articles published since August 2025, so older matches are not part of it (default: "zh-HK"; one of: "zh-HK", "zh-CN")

## Response

### 200 — Successful Response

- `@type` (string) (default: "Hk01ArticleSearchResult")
- `id` (integer, required)
- `article_title` (string, required)
- `description` (string, nullable)
- `web_url` (string, nullable)
- `image` (string, nullable)
- `main_category_id` (integer, nullable)
- `main_category` (string, nullable)
- `main_category_url` (string, nullable)
- `categories` (array) (default: [])
- `tags` (array) (default: [])
- `authors` (array) (default: [])
  - `@type` (string) (default: "Hk01SearchAuthor")
  - `name` (string, required)
  - `email` (string, nullable)
- `is_video` (boolean) (default: false)
- `is_sponsored` (boolean) (default: false)
- `is_featured` (boolean) (default: false)
- `sponsor_type` (string, nullable)
- `video_duration` (integer, nullable)
- `published_at` (integer, nullable)
- `created_at` (integer, nullable)
- `updated_at` (integer, nullable)

## Errors

### 422 — Validation Error

The request body did not validate

What to do: Check the fields against this schema. A URN with the wrong prefix is the most common cause.

- `detail` (array)
  - `loc` (array, required)
  - `msg` (string, required)
  - `type` (string, required)
  - `input` (any)
  - `ctx` (object)

### 408

The request ran past its time limit

What to do: Raise `timeout` in the request body, up to the maximum this endpoint documents. Lowering `count` or turning off the `with_*` flags also helps, because less work finishes sooner.

### 412

The entity was not found, or a precondition failed

What to do: Retrying will not help: either the entity does not exist, or the input points at a different one.

### 429

Too many requests: a rate limit or a usage window is exhausted

What to do: When the response carries an X-Retry-After header, wait that many seconds and retry: the same number is in the body as `detail.retry_after`, and the limit clears once that window passes. The message in the body names the limit that was hit.

### 500

Something broke on our side

What to do: Retrying will not help. If it keeps happening, send us the X-Request-ID from the response headers.

### 529

Rate limit reached, or the endpoint is overloaded

What to do: Wait at least 30 seconds, then retry.

## Response envelope

Success: Array of objects (may be empty if no results)

Error: Error may coexist with partial results if it occurs mid-execution. Check X-Error header and status code.

Every response carries these headers:

- `X-Error` — Error message text (present only on error)
- `X-Request-ID` — Unique request identifier
- `X-Execution-Time` — Execution time in seconds
- `X-Result-Count` — How many records the body carries. 0 means an empty result, which is a normal answer and not by itself an error. A non-zero count can come back together with X-Error when the failure happened partway through — read this header and X-Error independently.
- `X-Total-Available-Results` — How many records exist for this query, when the endpoint can say. On a `dry_run` request this is the answer and the body is empty. It saturates: the endpoint's documented maximum means 'at least that many', any smaller number is exact.
- `X-Warning` — Present when the request body carried keys this endpoint does not document. They were ignored, so any filter you meant to apply through them did not apply. Check the spelling against this schema and retry.
- `X-Retry-After` — Seconds to wait before retrying. Present only on 429.

