news-watch: Indonesia's top news websites scraper

news-watch is a Python package that scrapes structured news data from Indonesia's top news websites, offering keyword and date filtering queries for targeted research

⚠️ Ethical Considerations & Disclaimer ⚠️

Purpose: For educational and research purposes only. Not designed for commercial use that could be detrimental to news source providers.

User Responsibility: Users must comply with each website's Terms of Service and robots.txt. Aggressive scraping may lead to IP blocking. Scrape responsibly and respect server limitations.

Installation

Using pip (standard)

pip install news-watch
playwright install chromium

Development setup: see https://okky.dev/news-watch/getting-started/

Performance Notes

⚠️ Works best locally. Cloud environments (Google Colab, servers) may experience degraded performance or blocking due to anti-bot measures.

Usage

To run the scraper from the command line:

newswatch -k <keywords> -sd <start_date> -s [<scrapers>] -of <output_format> -v

Command-Line Arguments

Argument	Description
`-k, --keywords`	Required. Comma-separated keywords to scrape (e.g., `"ojk,bank,npl"`)
`-sd, --start_date`	Required. Start date in YYYY-MM-DD format (e.g., `2025-01-01`)
`-s, --scrapers`	Scrapers to use: specific names (e.g., `"kompas,viva"`), `"auto"` (default, platform-appropriate), or `"all"` (force all, may fail)
`-of, --output_format`	Output format: `csv`, `xlsx`, or `json` (default: csv)
`-o, --output_path`	Custom output file path (optional)
`-v, --verbose`	Show detailed logging output (default: silent)
`--list_scrapers`	List all supported scrapers and exit

Examples

# Basic usage
newswatch --keywords ihsg --start_date 2025-01-01

# Multiple keywords with specific scraper
newswatch -k "ihsg,bank" -s "detik" --output_format xlsx -v

# List available scrapers
newswatch --list_scrapers

Python API Usage

import newswatch as nw

# Basic scraping - returns list of article dictionaries
articles = nw.scrape("ekonomi,politik", "2025-01-01")
print(f"Found {len(articles)} articles")

# Get results as pandas DataFrame for analysis
df = nw.scrape_to_dataframe("teknologi,startup", "2025-01-01")
print(df['source'].value_counts())

# Save directly to file
nw.scrape_to_file(
    keywords="bank,ihsg", 
    start_date="2025-01-01",
    output_path="financial_news.xlsx"
)

# Quick recent news
recent_news = nw.quick_scrape("politik", days_back=3)

# Get available news sources
sources = nw.list_scrapers()
print("Available sources:", sources)

See the comprehensive guide for detailed usage examples and advanced patterns. For interactive examples, see the API reference notebook.

Run on Google Colab

You can run news-watch on Google Colab

Output

The scraped articles are saved as a CSV, XLSX, or JSON file in the current working directory with the format news-watch-{keywords}-YYYYMMDD_HH.

The output file contains the following columns:

title
publish_date
author
content
keyword
category
source
link

Supported Websites

Note:

On Linux platforms / cloud: Katadata relies on bearer-token capture (Playwright) and may fail in restricted environments, so it is automatically excluded in auto mode on Linux.

Use -s all to force-run all scrapers (may cause errors/timeouts).

Limitation: Kontan scraper maximum 50 pages.

Contributing

Contributions are welcome! If you'd like to add support for more websites or improve the existing code, please open an issue or submit a pull request.

License

This project is licensed under the MIT License - see the LICENSE file for details. The authors assume no liability for misuse of this software.

Citation

@software{mabruri_newswatch,
  author = {Okky Mabruri},
  title = {news-watch},
  year = {2025},
  doi = {10.5281/zenodo.14908389}
}

Name		Name	Last commit message	Last commit date
Latest commit History 105 Commits
.github/workflows		.github/workflows
docs		docs
notebook		notebook
scripts		scripts
src/newswatch		src/newswatch
tests		tests
.gitignore		.gitignore
CITATION.cff		CITATION.cff
LICENSE		LICENSE
Makefile		Makefile
README.md		README.md
mkdocs.yml		mkdocs.yml
pyproject.toml		pyproject.toml
uv.lock		uv.lock

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

news-watch: Indonesia's top news websites scraper

⚠️ Ethical Considerations & Disclaimer ⚠️

Installation

Using pip (standard)

Performance Notes

Usage

Examples

Python API Usage

Run on Google Colab

Output

Supported Websites

Contributing

License

Citation

Related Work

About

Uh oh!

Releases 11

Packages

Uh oh!

Contributors 2

Uh oh!

Languages

License

okkymabruri/news-watch

Folders and files

Latest commit

History

Repository files navigation

news-watch: Indonesia's top news websites scraper

⚠️ Ethical Considerations & Disclaimer ⚠️

Installation

Using pip (standard)

Performance Notes

Usage

Examples

Python API Usage

Run on Google Colab

Output

Supported Websites

Contributing

License

Citation

Related Work

About

Topics

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases 11

Packages 0

Uh oh!

Contributors 2

Uh oh!

Languages

Packages