{"service":"speech","display_name":"Speech","version":"0.1.0","description":"Text-to-speech behind x402 micropayments. A single POST /speech/synthesize endpoint accepts up to 4,000 characters of text plus a voice and returns synthesized audio. Powered by the qwen-audio-3.0-tts-plus model (premium speech tier of the Qwen Audio family) with two flagship voices: longanlufeng (male, bright and cheerful) and longanlingxin (female, warm and empathetic), each fluent in Chinese (Mandarin) and English. Output formats: mp3 (default), wav, pcm, opus, with volume, speech-rate and pitch controls. Clips are stored for 2 hours after generation — poll GET /speech/status/{audio_id} (free) and download via GET /speech/audio/{audio_id} (free).","archetype":"simple-inference","maturity":"stable","tags":["speech","text-to-speech","tts","audio","voice","ai","x402"],"networks":[{"name":"Base","chain_id":"eip155:8453","asset":"USDC/EURC"},{"name":"Arbitrum","chain_id":"eip155:42161","asset":"USDC"},{"name":"Polygon","chain_id":"eip155:137","asset":"USDC"},{"name":"Avalanche","chain_id":"eip155:43114","asset":"USDC"}],"accepted_assets":["USDC","EURC"],"facilitators":["https://facilitator.x402.press"],"payment_scheme":"exact","persona":"","operations":1}