🔊 Teach the Computer to Speak · สอนคอมพิวเตอร์ให้พูด · 教电脑说话

Step by step, in Python. English, ไทย and 中文 voices, free, with no key and no bill. Then let the program choose the voice by itself. · ทีละขั้นตอน ด้วย Python เสียงภาษาอังกฤษ ไทย และจีน ฟรี ไม่ต้องใช้คีย์ ไม่มีค่าใช้จ่าย แล้วให้โปรแกรมเลือกเสียงเองได้ด้วย · 用 Python 一步一步来。英文、泰文、中文的声音,免费,不用 key,也不用付钱。最后让程序自己选声音。

English voice
en-GB-LibbyNeural
Good morning. Today we will teach a computer to speak.

Real audio, made by the program in this lesson. Press a language button above — English, ไทย or 中文 — and the voice here changes with it. That swap is exactly what you are about to build. เสียงจริง สร้างด้วยโปรแกรมในบทเรียนนี้ ลองกดปุ่มภาษาด้านบน ไม่ว่าจะเป็น English ไทย หรือ 中文 แล้วเสียงตรงนี้จะเปลี่ยนตาม การสลับแบบนี้แหละคือสิ่งที่คุณกำลังจะสร้าง 真实的音频,就是这一课的程序做出来的。点一下上面的语言按钮 —— English、ไทย 或 中文 —— 这里的声音就会跟着换。你接下来要做的,正是这个切换。

🇬🇧 English

A computer that can speak is useful for learning a language. You can hear a word as many times as you like, and it never gets tired.

We use edge-tts. It is free. There is no key, no account and no bill.

We start in English. Then Thai. Then Chinese. At the end, the program looks at your letters and picks the voice by itself.

🇹🇭 ไทย

คอมพิวเตอร์ที่พูดได้มีประโยชน์มากสำหรับการเรียนภาษา คุณฟังคำเดิมกี่ครั้งก็ได้ และมันไม่มีวันเหนื่อย

เราจะใช้ edge-tts มันฟรี ไม่ต้องมีคีย์ ไม่ต้องมีบัญชี และไม่มีค่าใช้จ่าย

เราจะเริ่มที่ภาษาอังกฤษ ต่อด้วยภาษาไทย แล้วก็ภาษาจีน สุดท้ายโปรแกรมจะดูตัวอักษรของคุณแล้วเลือกเสียงเอง

🇨🇳 中文

会说话的电脑对学语言很有用。同一个词你想听多少遍都行,它永远不会累。

我们用 edge-tts。它是免费的,不用 key,不用账号,也不用付钱。

我们先做英文,再做泰文,然后是中文。最后,程序会看你的文字,自己挑声音。

0 Get ready · เตรียมตัว · 做好准备

🇬🇧 English

One library. The voices are made on a server, so you need wifi — but you never need a password.

🇹🇭 ไทย

ไลบรารีเดียว เสียงถูกสร้างบนเซิร์ฟเวอร์ คุณจึงต้องมีไวไฟ แต่ไม่ต้องใช้รหัสผ่านเลย

🇨🇳 中文

只要一个库。声音是在服务器上生成的,所以需要 wifi —— 但完全不需要密码。

pip3 install edge-tts

mkdir ~/Documents/Code/speak
cd ~/Documents/Code/speak
vim say.py

1 Say one thing in English · พูดหนึ่งประโยคเป็นภาษาอังกฤษ · 先用英文说一句

🇬🇧 English

Six lines and your computer speaks. You get an .mp3 file you can play, send, or put on a phone.

async and await look strange. They mean: this part waits for the internet, so let other things happen while it waits. You do not need to understand them fully today. Copy the shape.

🇹🇭 ไทย

หกบรรทัด คอมพิวเตอร์ของคุณก็พูดได้ คุณจะได้ไฟล์ .mp3 ที่เปิดฟังได้ ส่งต่อได้ หรือเอาใส่มือถือก็ได้

async กับ await ดูแปลก ๆ มันแปลว่า ส่วนนี้ต้องรออินเทอร์เน็ต ระหว่างรอก็ให้อย่างอื่นทำงานไปได้ วันนี้ยังไม่ต้องเข้าใจทั้งหมด ลอกรูปแบบไปก่อน

🇨🇳 中文

六行,你的电脑就会说话了。你会得到一个 .mp3 文件,可以播放、发送,也可以放进手机。

asyncawait 看起来很怪。它们的意思是:这一段要等网络,等的时候让别的事先做。今天不用完全弄懂,照着形状抄就行。

say.py
import asyncio
import edge_tts

async def speak(text, voice, out_path):
    talker = edge_tts.Communicate(text, voice)
    await talker.save(out_path)

asyncio.run(speak("Good morning, teacher.",
                  "en-GB-LibbyNeural",
                  "hello.mp3"))
▶ RUN IT python3 say.py, then open hello.mp3. That is your computer talking. python3 say.py แล้วเปิดไฟล์ hello.mp3 นั่นคือเสียงคอมพิวเตอร์ของคุณ python3 say.py,然后打开 hello.mp3。那就是你的电脑在说话。

2 Now Thai · ทีนี้ภาษาไทย · 现在换泰文

🇬🇧 English

Change one thing: the voice name. There are two Thai voices, one woman and one man.

Voice names have a pattern: language - country - name - Neural. Once you see it, you can read any voice name in the list.

An English voice cannot say Thai properly, and a Thai voice cannot say English properly. The words and the voice must match. That is the whole problem this lesson ends up solving.

🇹🇭 ไทย

เปลี่ยนแค่อย่างเดียว คือชื่อเสียง ภาษาไทยมีสองเสียง เป็นผู้หญิงหนึ่ง ผู้ชายหนึ่ง

ชื่อเสียงมีรูปแบบว่า ภาษา - ประเทศ - ชื่อ - Neural พอเห็นรูปแบบแล้ว คุณจะอ่านชื่อเสียงตัวไหนในรายการก็ได้

เสียงภาษาอังกฤษพูดไทยไม่ชัด และเสียงภาษาไทยก็พูดอังกฤษไม่ชัด คำกับเสียงต้องเข้าคู่กัน นี่คือปัญหาที่บทเรียนนี้จะไปแก้ในตอนท้าย

🇨🇳 中文

只改一样东西:声音的名字。泰文有两个声音,一女一男。

声音名字有固定格式:语言 - 国家 - 名字 - Neural。看懂这个格式,列表里任何一个名字你都能读懂。

英文的声音说不好泰文,泰文的声音也说不好英文。文字和声音必须配对。这正是这一课最后要解决的问题。

asyncio.run(speak("สวัสดีตอนเช้าค่ะ",
                  "th-TH-PremwadeeNeural",     # woman
                  "thai.mp3"))

# th-TH-NiwatNeural is the man's voice.
▶ TRY IT Now put the Thai words with the English voice and listen. That is why matching matters. ลองเอาคำภาษาไทยไปใส่กับเสียงภาษาอังกฤษแล้วฟังดู นี่แหละคือเหตุผลว่าทำไมต้องจับคู่ให้ถูก 现在把泰文配上英文的声音听听看。你就明白为什么必须配对了。

3 Now Chinese · ทีนี้ภาษาจีน · 现在换中文

🇬🇧 English

Same change again. Mandarin has eight voices, so you have a real choice.

You now have three files and three voices. Notice what your code did not need: no new library, no new function, no new idea. One string changed.

Ask the library what else it has. Reading the list is a normal part of using any library.

🇹🇭 ไทย

เปลี่ยนแบบเดิมอีกครั้ง ภาษาจีนกลางมีแปดเสียง คุณจึงมีตัวเลือกจริง ๆ

ตอนนี้คุณมีสามไฟล์ สามเสียง ลองสังเกตว่าโค้ดของคุณไม่ต้องใช้อะไรเพิ่มเลย ไม่มีไลบรารีใหม่ ไม่มีฟังก์ชันใหม่ ไม่มีแนวคิดใหม่ เปลี่ยนแค่ข้อความเดียว

ลองถามไลบรารีว่ามีอะไรอีกบ้าง การอ่านรายการแบบนี้เป็นเรื่องปกติของการใช้ไลบรารีทุกตัว

🇨🇳 中文

还是同样的改法。普通话有八个声音,所以你真的可以挑一挑。

现在你有三个文件、三个声音。注意你的代码需要什么:没有新库,没有新函数,没有新概念。只改了一个字符串。

问问这个库还有什么。看列表是使用任何库的正常动作。

asyncio.run(speak("老师早上好",
                  "zh-CN-XiaoxiaoNeural",
                  "chinese.mp3"))
python3 say.py --voices          # see them all

en-GB  (5)
   en-GB-LibbyNeural              Female
   en-GB-RyanNeural               Male
th-    (2)
   th-TH-NiwatNeural              Male
   th-TH-PremwadeeNeural          Female
zh-CN  (8)
   zh-CN-XiaoxiaoNeural           Female
   zh-CN-YunxiNeural              Male

4 Let the program choose · ให้โปรแกรมเลือกเอง · 让程序自己选

🇬🇧 English

Choosing the voice by hand every time is work. The program can see which language you typed.

Every letter has a number behind it. Thai letters sit in one range of numbers. Chinese characters sit in another. So we just count them.

Count the Thai letters. Count the Chinese ones. Whichever wins picks the voice. Everything else is English.

This is the same swap the buttons at the top of this page do. You press ไทย, the voice becomes Thai.

🇹🇭 ไทย

การเลือกเสียงเองทุกครั้งเป็นงานที่น่าเบื่อ โปรแกรมสามารถดูออกได้ว่าคุณพิมพ์ภาษาอะไร

ตัวอักษรทุกตัวมีตัวเลขอยู่เบื้องหลัง ตัวอักษรไทยอยู่ในช่วงตัวเลขหนึ่ง ตัวอักษรจีนอยู่อีกช่วงหนึ่ง เราก็แค่นับมัน

นับตัวอักษรไทย นับตัวอักษรจีน ฝ่ายไหนชนะก็เลือกเสียงนั้น นอกนั้นเป็นภาษาอังกฤษ

นี่คือการสลับแบบเดียวกับปุ่มด้านบนของหน้านี้ พอคุณกด ไทย เสียงก็เปลี่ยนเป็นภาษาไทย

🇨🇳 中文

每次都手动选声音很麻烦。程序其实能看出你打的是哪种文字。

每个字符背后都有一个数字。泰文字母在某一段数字范围里,汉字在另一段。所以我们数一数就行了。

数泰文字符,数汉字,谁多就用谁的声音。其余的都算英文。

这跟这一页顶上那几个按钮做的切换是同一回事。你按 ไทย,声音就变成泰文。

def detect_language(text):
    thai = sum(1 for c in text if "฀" <= c <= "๿")
    han  = sum(1 for c in text if "一" <= c <= "鿿")
    if thai > han and thai > 0:
        return "th"
    if han > 0:
        return "zh"
    return "en"

VOICES = {
    "en": "en-GB-LibbyNeural",
    "th": "th-TH-PremwadeeNeural",
    "zh": "zh-CN-XiaoxiaoNeural",
}

lang = detect_language(text)
voice = VOICES[lang]
▶ RUN IT Try "Hello สวัสดี". It picks Thai, because Thai letters win the count. A guess is not always right — that is why --lang exists. ลองพิมพ์ "Hello สวัสดี" มันจะเลือกภาษาไทย เพราะตัวอักษรไทยชนะการนับ การเดาไม่ได้ถูกเสมอ นี่คือเหตุผลที่ต้องมี --lang 试试 "Hello สวัสดี"。它会选泰文,因为泰文字符数量占优。猜测不总是对的 —— 所以才要有 --lang

5 Slow it down for a beginner · พูดช้าลงสำหรับผู้เริ่มเรียน · 为初学者放慢速度

🇬🇧 English

Normal speed is too fast when you are learning. rate fixes that, and it is the single most useful option here for a classroom.

Write it as a percent, with a plus or a minus sign. -25% is a good speed for a beginner.

🇹🇭 ไทย

ความเร็วปกติเร็วเกินไปสำหรับคนที่เพิ่งเริ่มเรียน rate ช่วยแก้เรื่องนี้ได้ และเป็นตัวเลือกที่มีประโยชน์ที่สุดสำหรับห้องเรียน

เขียนเป็นเปอร์เซ็นต์ พร้อมเครื่องหมายบวกหรือลบ ค่า -25% เป็นความเร็วที่กำลังดีสำหรับผู้เริ่มเรียน

🇨🇳 中文

刚开始学的时候,正常语速太快了。rate 能解决这个问题,而且是这里对课堂最有用的一个选项。

写成百分比,前面加正号或负号。-25% 对初学者来说速度正好。

talker = edge_tts.Communicate(text, voice, rate="-25%")
python3 say.py --rate=-30% "Where is the railway station?"

6 When no sound comes back · ตอนที่เสียงไม่กลับมา · 当声音没回来的时候

🇬🇧 English

Here is a real thing that happened while we made this page. The Thai line failed with NoAudioReceived. Nothing was wrong with it. Two seconds later, the same line worked perfectly.

Free services do this. They are busy, or something hiccups. The answer is not to give up. The answer is to try again, and wait a little longer each time.

Check the file size too. A file can exist and still be empty, and an empty file is not a failure your program will notice unless you look.

🇹🇭 ไทย

นี่คือเรื่องที่เกิดขึ้นจริงตอนเราทำหน้านี้ บรรทัดภาษาไทยล้มเหลวด้วย NoAudioReceived ทั้งที่ไม่มีอะไรผิดเลย สองวินาทีต่อมา บรรทัดเดิมทำงานได้ปกติ

บริการฟรีเป็นแบบนี้ บางทีมันยุ่ง หรือมีอะไรสะดุด ทางออกไม่ใช่การยอมแพ้ แต่คือลองใหม่ และรอนานขึ้นอีกนิดในแต่ละครั้ง

ตรวจขนาดไฟล์ด้วย ไฟล์อาจมีอยู่แต่ว่างเปล่า และไฟล์ว่างจะไม่ถูกนับว่าล้มเหลว ถ้าคุณไม่ไปดูเอง

🇨🇳 中文

这是我们做这一页时真实发生的事。泰文那一行失败了,报 NoAudioReceived,但它本身一点问题都没有。两秒钟后,同一行完美地跑通了。

免费服务就是这样。可能它忙,也可能只是打了个嗝。答案不是放弃,而是再试一次,每次多等一会儿。

也要检查文件大小。文件可能存在但是空的,而空文件不会被当成失败 —— 除非你去看。

async def speak(text, voice, out_path, rate, tries=4):
    for attempt in range(1, tries + 1):
        try:
            talker = edge_tts.Communicate(text, voice, rate=rate)
            await talker.save(str(out_path))
            if out_path.stat().st_size > 0:      # a file can exist and be empty
                return attempt
        except edge_tts.exceptions.NoAudioReceived:
            pass
        if attempt < tries:
            print("  (no sound came back, trying again)")
            await asyncio.sleep(2 * attempt)     # wait longer each time
    raise RuntimeError("The voice service gave nothing back.")

7 Play it out loud · เล่นเสียงออกลำโพง · 把声音放出来

🇬🇧 English

Making the file is the hard part. Playing it is easy, but every computer plays sound differently.

So we try each one until something works, and if nothing does we say so kindly. The file is saved either way, so nothing is lost.

🇹🇭 ไทย

การสร้างไฟล์คือส่วนที่ยาก ส่วนการเล่นนั้นง่าย แต่คอมพิวเตอร์แต่ละเครื่องเล่นเสียงคนละแบบ

เราจึงลองทีละวิธีจนกว่าจะมีอันที่ใช้ได้ ถ้าไม่มีเลยก็บอกผู้ใช้ดี ๆ ยังไงไฟล์ก็ถูกบันทึกไว้แล้ว ไม่มีอะไรหาย

🇨🇳 中文

做出文件是难的部分。播放很简单,但每台电脑放声音的方式都不一样。

所以我们一个一个试,直到有一个能用;都不行就好好地说一声。反正文件已经存下来了,什么都不会丢。

if sys.platform == "darwin":                     # Mac
    commands = [["afplay", str(path)]]
elif sys.platform.startswith("win"):             # Windows
    commands = [["powershell", "-NoProfile", "-c", ...]]
else:                                            # Linux
    commands = [["mpg123", "-q", str(path)],
                ["ffplay", "-nodisp", "-autoexit", str(path)]]

8 The whole program · โปรแกรมทั้งหมด · 完整的程序

🇬🇧 English

Here is the finished file. The three sound clips at the top of this page were made with it.

🇹🇭 ไทย

นี่คือไฟล์ที่เสร็จแล้ว คลิปเสียงสามอันด้านบนของหน้านี้ก็สร้างด้วยไฟล์นี้

🇨🇳 中文

这是做完的文件。这一页顶上那三段声音,就是用它做出来的。

python3 say.py "Good morning, teacher."
python3 say.py "สวัสดีตอนเช้าค่ะ"
python3 say.py "老师早上好"
python3 say.py --rate=-25% --lang th "ขอบคุณครับ"

9 Now change it · ลองแก้ดู · 现在动手改它

🇬🇧 English

A speaking computer joins up with almost everything else you have built.

🇹🇭 ไทย

คอมพิวเตอร์ที่พูดได้ต่อยอดเข้ากับเกือบทุกอย่างที่คุณสร้างมาแล้ว

🇨🇳 中文

一台会说话的电脑,几乎能接到你做过的所有东西上。

Try thisHowLevel
Say the man's Thai voice--voice th-TH-NiwatNeural. One word changes.easy
Read a whole file out loudOpen a .txt file and pass the text in.easy
Speak your flashcardsAdd a speaker button to the flashcard app.harder
Say the air reportSend the words from the air program to say.py.harder
Give your AI teacher a voiceSpeak the answer from the AI tutor.harder
Make a listening testSay ten words, save ten files, shuffle, play, and mark.harder