See the full portfolio

Rental finder

A system that consolidates rental listings from three sites every day, deduplicates and geocodes them, and publishes them in a viewer with filters, metrics and history.

Year
2026
Status
Live, updated daily
Stack
Python, Playwright, Leaflet, OpenStreetMap, launchd, GitHub Pages
Rental finder

The problem

It started with a very non-technical conversation. My flatmate and I decided to move and we knew exactly what we wanted: three bedrooms (two to sleep in, one as an office), near Parque Bustamante, Lastarria, Barrio Italia or Manuel Montt, and a cap of CLP 800,000 a month including building fees.

Sounds reasonable. Searching for it by hand, less so:

  • Listings are spread across several sites and the same apartment shows up two or three times.
  • Almost nobody publishes building fees, so a CLP 700,000 rent can end up costing 850,000.
  • Some prices come in UF (an inflation-indexed unit) and others in pesos.
  • No search engine understands "close to the metro" or "a big two-bedroom that works as a three".

After a week of manual searching I turned the requirement into a data problem and built the pipeline that solves it.

What it found this morning

These numbers are read from the live finder every time this site is published.

861listings checked
438new since the previous run
30days of history kept
4perfect matches

Updated 2026-10-01 11:34

It's the real thing. Try the Perfect match button and the deals tab.

An end-to-end pipeline

The design is deliberately simple. Each piece does one thing and can run on its own.

  1. 01
    ScrapersOne file per site, all with the same output. scrapers/*.py
  2. 02
    ConsolidateMerges sources, removes duplicates and computes the total cost with fees. consolidate.py
  3. 03
    GeocodeAddresses to coordinates with OpenStreetMap and a cache. geocode.py
  4. 04
    EnrichPlaywright opens key listings to extract real fees, building age and pet policy. enrich_pi.py
  5. 05
    Static viewerHTML, CSS and JavaScript with Leaflet. No server, no database. viewer/
  6. 06
    Every daylaunchd runs it at 10:00, saves the daily snapshot and publishes to GitHub Pages.

The most important decision was defining a shared schema before writing the first scraper. Every site has its own HTML, but all of them return the same structure. Adding a new site means writing one file that returns a list of listings, nothing else.

Scraping: every site is its own world

SiteTechniqueResult
Portal Inmobiliariorequests + HTMLMain source, by neighborhood
Chilepropiedadesrequests + HTMLFast and complete
YapoScrapling / CamoufoxBlocks bots, about 40 s per page
Facebook MarketplaceManual CSVScraping it breaks their terms

A Chilean classic: prices arrive as "$ 750.000", "UF 13,5" or "13,5 UF". The conversion uses that day’s UF value from mindicador.cl, with a fallback if the API is down.

def parse_precio(texto: str) -> int:
    t = texto.upper().replace("\xa0", " ")
    num = re.sub(r"[^\d.,]", "", t)   # solo dígitos, comas y puntos
    if "UF" in t:                      # la UF usa coma decimal
        return round(float(num.replace(".", "").replace(",", ".")) * uf_hoy())
    return int(num.replace(".", ""))

Sale listings kept sneaking in among the rentals. The rule to catch them was simple: no monthly rent costs 158 million pesos.

Duplicates and geocoding

With over a thousand listings from several sources, duplicates are inevitable. Deduplication happens in two passes: first a stable id (a hash of the URL), then a key made of the normalized address plus price, which catches the same listing posted by two agencies.

The map needed coordinates. OpenStreetMap's Nominatim is free but allows one request per second, so everything goes through a cache that only looks up new addresses. If an address can't be found, it falls back to the neighborhood center and then the municipality center, with a small offset so pins don't stack. Those listings are flagged as approximate.

And "close to the metro" needed no scraping at all. With each listing's coordinates and those of 23 stations, the haversine formula gives the distance. Every card says, for example, "Santa Isabel, 257 m".

Encoding what we actually wanted

It wasn’t just "three bedrooms under 800,000". A two-bedroom of 68 m² or more also worked. That human criterion became a rule:

# un 2D amplio sirve como oficina y también califica
sirve = dorm >= 3 or (dorm >= 2 and m2 >= 68)
match_perfecto = sirve and total_con_gastos <= 800_000 and en_barrio_objetivo

The result is a funnel. This is how it looked on September 15, 2026:

Each listing also gets a 0–100 score weighing budget, bedrooms, neighborhood and metro distance. The most promising ones always float to the top.

Is it expensive or a deal?

A price is only expensive compared with something similar. The pipeline compares each rent with the median of its group, and if the group has fewer than five listings, it moves up to a broader one:

  1. Neighborhood+ bedrooms + bathrooms + m²
  2. Neighborhood+ bedrooms + m²
  3. Municipality+ bedrooms + m²
  4. Municipality+ bedrooms

A listing 10% below its group median is flagged as a deal and gets its own tab. That’s when the finder stopped showing what exists and started showing what’s worth it.

Automatic updates and history

A macOS LaunchAgent runs everything each morning and pushes to GitHub Pages. Every run stores a full snapshot of the market, and with 30 days of snapshots the finder knows which listings are new, which dropped in price and which disappeared (probably rented). A date picker shows the market on any day of the past month.

From over a thousand daily listings to a couple dozen concrete candidates every morning.

Technical lessons

  • Shared schema first, scrapers second. Adding sources becomes trivial.
  • Cache everything expensive. Reruns are cheap and polite to external services.
  • Cascading fallbacks. The pipeline never breaks over one missing value.
  • Static output. Free hosting, zero maintenance.
  • Automate from day one. The data stays fresh without anyone remembering to run it.

Want it for your own search? All the criteria live in config.py: change the municipalities, budget and neighborhoods, run python run_all.py and you have your own finder.