
Job Data Converter
Our Job Data Feeds are delivered as gzip-compressed JSON Lines files on AWS S3 and AWS Data Exchange. This page shows how to turn them into CSV, XML, RSS 2.0, Atom, Parquet, Excel or NDJSON - with complete, tested source code in JavaScript, Python, Java and Bash that produces the same structure as the formats of our Job Postings API.
Table of Contents
1. From JSON Lines to Your Format
Every delivery file is named techmap_jobs_{countryCode}_{YYYY-MM-DD}.jsonl.gz - for example techmap_jobs_us_2026-09-06.jsonl.gz - and contains one job posting as a JSON object per line. JSON Lines is ideal for streaming and for data warehouses, but many tools expect a table, a feed or a columnar file. The converters below:
- stream the gzip file line by line, so even files with hundreds of thousands of postings never have to fit into memory,
- map each posting to the job structure of the Job Postings API (same field names as in the API result, see Section 2),
- write one or more output formats in a single pass - with the same columns, element names and types as the API's
format=parameter.
Because the output matches the API, you can mix both sources: e.g. backfill a job board from the daily files and keep it current with the API, using the same importer.
| Format | Output file | Same as API | Typical use |
|---|---|---|---|
| CSV | .csv | format=csv | Excel, Google Sheets, pandas, SQL imports |
| XML | .xml | format=xml | Enterprise integrations, ATS and HR systems |
| RSS 2.0 | .rss.xml | format=rss | Job boards, WordPress job plugins, feed readers |
| Atom | .atom.xml | format=atom | Feed consumers that prefer Atom 1.0 |
| Parquet | .parquet | format=parquet | Snowflake, BigQuery, Redshift, Athena, Spark, DuckDB |
| Excel | .xlsx | (CSV columns) | Business users, small countries |
| NDJSON | .ndjson | format=json (result array) | Code that already reads the API |
2. Field Mapping
The files contain the full export schema documented in the data dictionary - raw text and html, the original page JSON, company.*, location.*, salary.* and more. The converters derive the compact job structure of the API from it. Arrows (→) are fallbacks used when the first value is missing or empty.
2.1 File Fields to API Fields
| API field / CSV column | Field in the .jsonl.gz file |
|---|---|
title | name |
company | company.name → company.nameOrg → json.jsonLD.hiringOrganization.name |
jsonLD | json.jsonLD (schema.org JobPosting incl. description) → json.schemaOrg |
occupation | json.jsonLD.relevantOccupation → position.name → "N/A" |
industry | json.jsonLD.industry → json.inferredTags.INDUSTRIES[0] → "N/A" |
department | json.jsonLD.employmentUnit → json.inferredTags.DEPARTMENTS[0] → "N/A" |
city / state / postCode | json.jsonLD.jobLocation.address.addressLocality / addressRegion / postalCode → location.orgAddress.city / state / postCode |
geoPoint | {lat, lon} from json.jsonLD.jobLocation.latitude/longitude → location.orgAddress.geoPoint (lat/lng) |
countryCode | sourceCC |
language / locale | first two letters of locale / locale |
timezone / timezoneOffset | json.jsonLD.applicantLocationRequirements ("CEST Timezone" → "CEST") / location.orgAddress.timezoneOffset |
workType / workPlace / careerLevel / contractType | json.inferredTags.WORK_TYPES / WORK_PLACES / CAREER_LEVELS / CONTRACT_TYPES - unique, sorted, ["N/A"] if empty |
skills | json.jsonLD.skills → json.inferredTags.SKILLS |
hasSalary | true if salary.minValue, salary.maxValue or json.jsonLD.baseSalary is set |
dateCreated / dateExpired / dateActive | dateCreated / json.jsonLD.validThrough / dateExpired, else dateCreated + 1 month |
source / portal / isDuplicate / isRecruiter / isDirect | same field names |
description (CSV and Excel only) | json.jsonLD.description |
Need more columns, such as salary.minValue, company.info.companySize or the plain text? Add them in the mapping function (toApiJob / to_api_job / to_api) - every writer picks up new fields automatically except the fixed CSV and Parquet column lists, which you extend the same way.
2.2 RSS and Atom Elements
RSS items and Atom entries use the same elements as the API feeds. Values of N/A become empty elements, arrays are joined with commas.
| RSS 2.0 <item> | Atom <entry> | Source (API job field) |
|---|---|---|
title | title | title |
description | content | jsonLD.description |
pubDate (RFC 822) | updated, published (ISO 8601) | dateCreated |
link | link | jsonLD.url |
guid | id | jsonLD.url |
category | category term="..." | occupation |
location | location | jsonLD.jobLocation.name |
city, state | city, state | city, state |
country | country | jsonLD.jobLocation.address.addressCountry → countryCode |
company | company | company |
company_url, company_logo | company_url, company_logo | jsonLD.hiringOrganization.url, .logo |
apply_link | apply_link | jsonLD.sameAs |
salary | salary | jsonLD.baseSalary.name |
workType, contractType | workType, contractType | workType joined with ", " → jsonLD.employmentType |
industry, department, occupation | industry, department, occupation | industry, department, occupation |
careerLevel, workPlace | careerLevel, workPlace | careerLevel, workPlace joined with ", " |
skills | skills | jsonLD.skills joined with ", " |
3. Output Formats
Pick a format to see what the converters produce and how to create only this format in each language.
CSV (.csv - API: format=csv)
One row per job posting and one column per top-level field of the API result. Nested objects and arrays (jsonLD, geoPoint, skills, workType, ...) are written as JSON strings, and jsonLD.description is copied into an extra description column at the end - exactly like format=csv of the API (json2csv). Strings are quoted, numbers and booleans are not. Use it for Excel, Google Sheets, pandas or database imports.
node convert-jobs.js techmap_jobs_us_2026-09-06.jsonl.gz csv
python convert_jobs.py techmap_jobs_us_2026-09-06.jsonl.gz csv
java -cp "lib/*" ConvertJobs.java techmap_jobs_us_2026-09-06.jsonl.gz csv
./convert-jobs.sh techmap_jobs_us_2026-09-06.jsonl.gz csv"occupation","dateActive","city","timezone","contractType","language","industry","jsonLD","source","locale","geoPoint","title","skills","dateCreated","timezoneOffset","countryCode","company","state","isDuplicate","portal","department","workPlace","isRecruiter","hasSalary","careerLevel","workType","postCode","isDirect","dateExpired","description"
"Engineer","2026-10-06T00:00:00.000Z","Austin","CDT","[""N/A""]","en","IT","{""@context"":""https://schema.org"",""@type"":""JobPosting"",""title"":""Data Engineer"", ...}","linkedin_us","en_US","{""lat"":30.2672,""lon"":-97.7431}","Data Engineer","[""Python"",""SQL"",""AWS""]","2026-09-06T00:00:00+0000",,"us","Example Corp","Texas",false,"linkedin","IT","[""Hybrid""]",false,false,"[""N/A""]","[""FullTime""]","78701",false,"2026-10-06T00:00:00.000Z","We are looking for a Data Engineer ..."4. Converter Source Code
Each converter is a single file you can copy, run and adapt. All four produce identical CSV, XML, RSS, Atom and Parquet output for the same input file (except for number formatting such as 0.0 vs. 0 inside JSON strings). Pass the formats you need as comma-separated second argument - the default is all of them.
Dependencies: npm: xml-js (XML, RSS, Atom) and parquetjs-lite (Parquet). Gzip, line reading and CSV use the Node.js standard library.
# Node.js 18+ - xml-js and parquetjs-lite are the same libraries the Job Postings API uses
npm install xml-js@^1.6.11 parquetjs-lite@^0.8.7
node convert-jobs.js techmap_jobs_us_2026-09-06.jsonl.gz # all formats
node convert-jobs.js techmap_jobs_us_2026-09-06.jsonl.gz csv,rss # only CSV and RSS#!/usr/bin/env node
// convert-jobs.js - convert Techmap job postings (.jsonl.gz) to CSV, XML, RSS 2.0, Atom, Parquet and NDJSON
// Usage: node convert-jobs.js techmap_jobs_us_2026-09-06.jsonl.gz [csv,xml,rss,atom,parquet,ndjson]
// Install: npm install xml-js parquetjs-lite (the same libraries the Techmap Job Postings API uses)
const fs = require('fs');
const zlib = require('zlib');
const readline = require('readline');
const { once } = require('events');
const { json2xml } = require('xml-js');
const { ParquetSchema, ParquetWriter } = require('parquetjs-lite');
const input = process.argv[2];
if (!input) {
console.error('Usage: node convert-jobs.js <techmap_jobs_cc_YYYY-MM-DD.jsonl.gz> [csv,xml,rss,atom,parquet,ndjson]');
process.exit(1);
}
const formats = (process.argv[3] || 'csv,xml,rss,atom,parquet,ndjson').toLowerCase().split(',');
const base = input.replace(/\.jsonl?(\.gz)?$/, '');
// Fields of a job in the API result (api.techmap.io /api/v2/jobs/search)
const FIELDS = ['occupation', 'dateActive', 'city', 'timezone', 'contractType', 'language', 'industry', 'jsonLD',
'source', 'locale', 'geoPoint', 'title', 'skills', 'dateCreated', 'timezoneOffset', 'countryCode', 'company',
'state', 'isDuplicate', 'portal', 'department', 'workPlace', 'isRecruiter', 'hasSalary', 'careerLevel',
'workType', 'postCode', 'isDirect', 'dateExpired'];
const BOOLEAN_FIELDS = ['isDuplicate', 'isRecruiter', 'isDirect', 'hasSalary'];
const NUMBER_FIELDS = ['timezoneOffset'];
// ---------- helpers ----------
const asList = (v) => (Array.isArray(v) ? v : v ? [v] : []);
const str = (v) => (typeof v === 'string' ? v : '');
const sanitize = (v) => (v === undefined || v === null || v === 'N/A' || v === '' ? '' : v);
const toDate = (s) => new Date(String(s).replace(/([+-]\d\d)(\d\d)$/, '$1:$2')); // "+0000" -> "+00:00"
const cleanXml = (s) => s.replace(/[\u0000-\u0008\u000B\u000C\u000E-\u001F]/g, ''); // chars not allowed in XML
const escAttr = (s) => String(s).replace(/&/g, '&').replace(/&/g, '&').replace(/</g, '<'); // xml-js only escapes quotes in attributes
const toXml = (obj) => cleanXml(json2xml(JSON.stringify(obj), { compact: true, spaces: 2 }))
.replace(/<@/g, '<AT').replace(/<\/@/g, '</AT'); // "@type" -> <ATtype> like the API
function tagList(raw, key) {
const values = ((raw.json && raw.json.inferredTags && raw.json.inferredTags[key]) || [])
.filter((v) => v && v !== 'N/A');
return values.length ? [...new Set(values)].sort() : ['N/A'];
}
// Map one line of the S3/ADX file to the job structure returned by the API
function toApiJob(raw) {
const json = raw.json || {};
const jsonLD = json.jsonLD || json.schemaOrg || {};
const tags = json.inferredTags || {};
const place = jsonLD.jobLocation || {};
const address = place.address || {};
const org = (raw.location && raw.location.orgAddress) || {};
const salary = raw.salary || {};
let geoPoint = null;
if (typeof place.latitude === 'number' && typeof place.longitude === 'number') geoPoint = { lat: place.latitude, lon: place.longitude };
else if (org.geoPoint && org.geoPoint.lat != null) geoPoint = { lat: org.geoPoint.lat, lon: org.geoPoint.lng };
const dateExpired = jsonLD.validThrough || null;
let dateActive = dateExpired;
if (!dateActive && raw.dateCreated) {
const d = toDate(raw.dateCreated);
d.setUTCMonth(d.getUTCMonth() + 1);
dateActive = d.toISOString();
}
return {
occupation: jsonLD.relevantOccupation || (raw.position && raw.position.name) || 'N/A',
dateActive,
city: address.addressLocality || org.city || '',
timezone: (str(jsonLD.applicantLocationRequirements).match(/^(\S+) Timezone$/) || [])[1] || str(org.timezone),
contractType: tagList(raw, 'CONTRACT_TYPES'),
language: (raw.locale || '').slice(0, 2),
industry: jsonLD.industry || (tags.INDUSTRIES || [])[0] || 'N/A',
jsonLD,
source: raw.source || '',
locale: raw.locale || '',
geoPoint,
title: raw.name || jsonLD.title || '',
skills: asList(jsonLD.skills).length ? asList(jsonLD.skills) : asList(tags.SKILLS),
dateCreated: raw.dateCreated,
timezoneOffset: typeof org.timezoneOffset === 'number' ? org.timezoneOffset : null,
countryCode: raw.sourceCC || '',
company: (raw.company && (raw.company.name || raw.company.nameOrg)) || (jsonLD.hiringOrganization || {}).name || '',
state: address.addressRegion || org.state || '',
isDuplicate: !!raw.isDuplicate,
portal: raw.portal || '',
department: jsonLD.employmentUnit || (tags.DEPARTMENTS || [])[0] || 'N/A',
workPlace: tagList(raw, 'WORK_PLACES'),
isRecruiter: !!raw.isRecruiter,
hasSalary: !!(salary.minValue || salary.maxValue || jsonLD.baseSalary),
careerLevel: tagList(raw, 'CAREER_LEVELS'),
workType: tagList(raw, 'WORK_TYPES'),
postCode: address.postalCode || org.postCode || '',
isDirect: !!raw.isDirect,
dateExpired,
};
}
// ---------- format writers (same mapping as format=csv|xml|rss|atom|parquet in the API) ----------
// CSV: one column per top-level field, nested objects/arrays as JSON, jsonLD.description as extra "description" column
const csvCell = (v) => {
if (v === undefined || v === null) return '';
if (typeof v === 'number' || typeof v === 'boolean') return String(v);
const s = typeof v === 'object' ? JSON.stringify(v) : String(v);
return '"' + s.replace(/"/g, '""') + '"';
};
const csv = {
ext: 'csv',
header: () => [...FIELDS, 'description'].map(csvCell).join(',') + '\n',
row: (job) => [...FIELDS.map((f) => job[f]), job.jsonLD.description || ''].map(csvCell).join(',') + '\n',
footer: () => '',
};
const xml = {
ext: 'xml',
header: () => '<?xml version="1.0" encoding="UTF-8"?>\n<response>\n<api>Techmap.io Job Posting API</api>\n<result>\n',
row: (job) => toXml({ job }) + '\n',
footer: () => '</result>\n</response>\n',
};
// RSS item and Atom entry use the same element names as the API
function feedFields(job) {
const ld = job.jsonLD;
return {
location: (ld.jobLocation || {}).name,
city: job.city,
state: job.state,
country: ((ld.jobLocation || {}).address || {}).addressCountry || job.countryCode,
company: job.company,
company_url: (ld.hiringOrganization || {}).url,
company_logo: sanitize((ld.hiringOrganization || {}).logo),
apply_link: ld.sameAs,
salary: (ld.baseSalary || {}).name,
workType: sanitize(job.workType.join(', ') || ld.employmentType),
contractType: sanitize(job.workType.join(', ') || ld.employmentType),
industry: sanitize(job.industry),
careerLevel: sanitize(job.careerLevel.join(', ')),
workPlace: sanitize(job.workPlace.join(', ')),
skills: sanitize(asList(ld.skills).join(', ')),
department: sanitize(job.department || ld.employmentUnit),
occupation: sanitize(job.occupation),
};
}
const rss = {
ext: 'rss.xml',
header: () => '<?xml version="1.0" encoding="UTF-8"?>\n<rss version="2.0">\n<channel>\n' + toXml({
title: 'Techmap.io Job Postings',
link: 'https://api.techmap.io',
description: 'Techmap job postings from ' + input + ' in RSS 2.0 Feed format.',
pubDate: new Date().toUTCString(),
docs: 'https://api.techmap.io',
ttl: '60',
}) + '\n',
row: (job) => toXml({ item: {
title: job.title,
description: job.jsonLD.description || '',
pubDate: toDate(job.dateCreated).toUTCString(),
link: job.jsonLD.url,
guid: job.jsonLD.url,
category: sanitize(job.occupation),
...feedFields(job),
} }) + '\n',
footer: () => '</channel>\n</rss>\n',
};
const atom = {
ext: 'atom.xml',
header: () => '<?xml version="1.0" encoding="UTF-8"?>\n<feed version="1.0" xmlns="http://www.w3.org/2005/Atom">\n' + toXml({
id: 'https://api.techmap.io',
title: 'Techmap.io Job Postings',
updated: new Date().toISOString(),
subtitle: 'Techmap job postings from ' + input + ' in Atom 1.0 Feed format.',
category: { _attributes: { term: 'jobs' } },
docs: 'https://api.techmap.io',
}) + '\n',
row: (job) => toXml({ entry: {
id: job.jsonLD.url,
title: job.title,
updated: toDate(job.dateCreated).toISOString(),
link: job.jsonLD.url,
content: job.jsonLD.description || '',
published: toDate(job.dateCreated).toISOString(),
category: { _attributes: { term: escAttr(sanitize(job.occupation)) } },
...feedFields(job),
} }) + '\n',
footer: () => '</feed>\n',
};
const ndjson = { ext: 'ndjson', header: () => '', row: (job) => JSON.stringify(job) + '\n', footer: () => '' };
// Parquet: booleans as BOOLEAN, numbers as DOUBLE, objects/arrays as JSON strings (UTF8) - like the API
function parquetWriter() {
const fields = {};
for (const f of FIELDS) {
const type = BOOLEAN_FIELDS.includes(f) ? 'BOOLEAN' : NUMBER_FIELDS.includes(f) ? 'DOUBLE' : 'UTF8';
fields[f] = { type, optional: true };
}
let writer;
return {
open: async () => { writer = await ParquetWriter.openFile(new ParquetSchema(fields), base + '.parquet'); },
append: async (job) => {
const row = {};
for (const f of FIELDS) {
const v = job[f];
if (v === undefined || v === null) row[f] = null;
else if (typeof v === 'object') row[f] = JSON.stringify(v);
else row[f] = v;
}
await writer.appendRow(row);
},
close: () => writer.close(),
};
}
// ---------- main: stream the gzip file line by line ----------
async function write(stream, text) {
if (text && !stream.write(text)) await once(stream, 'drain'); // respect backpressure
}
async function main() {
const textFormats = { csv, xml, rss, atom, ndjson };
const outputs = formats.filter((f) => textFormats[f]).map((f) => ({
fmt: textFormats[f],
stream: fs.createWriteStream(base + '.' + textFormats[f].ext, 'utf8'),
}));
const parquet = formats.includes('parquet') ? parquetWriter() : null;
if (parquet) await parquet.open();
for (const o of outputs) await write(o.stream, o.fmt.header());
const lines = readline.createInterface({
input: fs.createReadStream(input).pipe(zlib.createGunzip()),
crlfDelay: Infinity,
});
let count = 0;
for await (const line of lines) {
if (!line.trim()) continue;
const job = toApiJob(JSON.parse(line));
for (const o of outputs) await write(o.stream, o.fmt.row(job));
if (parquet) await parquet.append(job);
count++;
}
for (const o of outputs) {
await write(o.stream, o.fmt.footer());
o.stream.end();
await once(o.stream, 'finish');
}
if (parquet) await parquet.close();
console.log('Converted ' + count + ' job postings from ' + input + ' to: ' + formats.join(', '));
}
main().catch((err) => { console.error(err); process.exit(1); });
Converting many files
Download a month of files with the AWS CLI and loop over them - each file is converted independently:
aws s3 cp --recursive s3://YOUR_BUCKET_ALIAS/ . --exclude "*" --include "techmap_jobs_us_2026-09-*.jsonl.gz"
for f in techmap_jobs_us_2026-09-*.jsonl.gz; do
python convert_jobs.py "$f" csv,parquet
done5. Shortcuts: DuckDB and the API
Raw files to Parquet or CSV with one DuckDB command
If you want all original fields instead of the API structure, DuckDB reads the gzip files directly and keeps nested objects as STRUCT columns (Parquet) or JSON (CSV):
duckdb -c "COPY (SELECT * FROM read_json_auto('techmap_jobs_us_2026-09-*.jsonl.gz', union_by_name = true))
TO 'techmap_jobs_us_2026-09.parquet' (FORMAT parquet)"
duckdb -c "COPY (SELECT sourceCC, dateCreated, name, url, company.name AS company, location.orgAddress.city AS city, text
FROM read_json_auto('techmap_jobs_us_2026-09-06.jsonl.gz')) TO 'jobs.csv' (HEADER)"Get CSV, XML, RSS, Atom or Parquet directly from the API
The Job Postings API returns the same formats without any conversion - add format=csv, xml, rss, atom or parquet to a search request (default is json). Sample responses in every format are available on the API sample data page.
curl -G "https://daily-international-job-postings.p.rapidapi.com/api/v2/jobs/search" \
--data-urlencode "countryCode=us" \
--data-urlencode "dateCreated=2026-09-06" \
--data-urlencode "format=csv" \
-H "X-RapidAPI-Key: YOUR_API_KEY" \
-o techmap_jobs_us_2026-09-06_page1.csvUse the API for filtered, near real-time queries and the data feeds for complete daily volumes per country - the converters on this page make both look the same.
6. References
- Job Data Overview - file format, export schema and data dictionary of the data feeds
- Developer Resources - AWS CLI commands, API quick start and code examples
- Job API Overview - endpoints, parameters (incl.
format) and response schema - API sample data - JSON, CSV, RSS, Atom, XML and Parquet samples for download
- API field reference - all fields of the API job structure
- JSON Lines specification - the format of all deliveries
- RSS 2.0 and Atom (RFC 4287) specifications
- Techmap on AWS Marketplace - subscribe to Data Feeds or purchase Historical Datasets
Need another format or a custom field mapping? Contact us - we are happy to help.