Methods, apparatus, and storage media for converting text data into audiovisual data.

By embedding 13×14 dot matrix fonts containing Chinese character voice information into handheld electronic audio-visual devices, the problems of complicated operation and inability of users to freely select content in existing technologies have been solved. This has enabled automatic conversion and synchronous playback of text data to audio-visual data, reducing labor costs.

CN122489786APending Publication Date: 2026-07-31王保君
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
王保君
Filing Date
2023-09-18
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing handheld electronic audio-visual devices suffer from problems such as cumbersome operation, high labor costs, and the inability of users to freely select content when playing audio-visual files.

Method used

By embedding Chinese character phonetic information into 13×14 dot matrix Chinese character fonts, the phonetic and graphic data of Chinese characters are automatically obtained, realizing the conversion of text data into audiovisual data and synchronous playback.

Benefits of technology

It enables the integrated use of the sound and form of Chinese characters, reducing the manual recording process, saving costs, and allowing users to freely choose the content to listen to.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489786A_ABST
    Figure CN122489786A_ABST
Patent Text Reader

Abstract

This specification provides a method, apparatus, and storage medium for converting text data into audiovisual data. The method embeds Chinese character speech information into innovative 13x14 dot matrix Chinese character font data, thereby achieving integrated sound-image calling of Chinese characters. The method includes receiving text data; obtaining a character font display based on the text data and a preset font library; obtaining the pronunciation of the character based on the speech information embedded in the dot matrix font and a preset pronunciation library; thus automatically converting the text data into audiovisual data. This integrated sound-image calling technology for Chinese characters provided in this application overcomes the shortcomings of current handheld electronic audiovisual products, such as Chinese text documents not being recorded with audio, having text but no sound, or having sound but no text; and the defect that users cannot freely select text for listening and reading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of electronic audiovisual technology, and in particular to a method, apparatus and storage medium for converting text data into audiovisual data. Background Technology

[0002] With the development of science and technology, handheld electronic audio-visual devices have been widely accepted and loved by people due to the emergence of large-capacity storage and their portability.

[0003] Handheld electronic audio-visual devices are used to play audio-visual files. In the field of audio-visual reading, there are basically two ways to obtain audio-visual files. One common method is to pre-dub the Chinese text document and record it manually. This method is cumbersome, requires significant labor costs, and has limitations. The other method is to download Chinese character pronunciation software from the internet to obtain the pronunciation of Chinese characters. However, these software programs cannot run due to the very limited software and hardware resources of handheld audio-visual products. In addition to the above, the audio-visual content of handheld electronic audio-visual products is currently mostly specified by the manufacturer, and users cannot freely choose what they want. Summary of the Invention

[0004] In view of the above problems, this application aims to propose and implement a method, apparatus and storage medium for converting text data into audiovisual data in a handheld electronic audiovisual product, so as to solve the above-mentioned technical problems.

[0005] Firstly, embodiments of this specification provide a method for automatically converting text data into audiovisual data. The method includes: embedding Chinese character phonetic information into innovative 13x14 dot matrix Chinese character font data. This achieves integrated phonetic and graphical representation of Chinese characters. The program automatically acquires the Chinese character font data and phonetic data, enabling the automatic conversion of text data into audiovisual data and synchronized playback.

[0006] Receive input text data;

[0007] Based on the text data, determine the dot matrix font from the preset font library;

[0008] Based on the dot matrix font, the text data is converted into corresponding characters;

[0009] Based on the speech information embedded in the dot matrix font, the pronunciation of the text is obtained from a preset pronunciation library;

[0010] Audiovisual data is obtained based on the text and its pronunciation.

[0011] Furthermore, the dot matrix font is a 13×14 dot matrix font;

[0012] The step of converting the text data into corresponding characters based on the dot matrix font includes:

[0013] Based on the 13×14 dot matrix font, the Chinese characters to be read in the text data are converted into dot matrix fonts.

[0014] Further, the step of converting the Chinese characters to be read in the text data into dot matrix character data based on the 13×14 dot matrix font includes:

[0015] Based on the Chinese characters to be read in the text data, determine the corresponding dot matrix font in the preset 13×14 dot matrix font library;

[0016] Detect whether the dot matrix font encoding contains a correction identifier;

[0017] When the correction identifier exists in the font encoding, the corresponding dot matrix font is corrected.

[0018] Furthermore, the dot matrix font has a correction area, which is located between two characters;

[0019] When the correction identifier exists in the font encoding, the corresponding dot matrix font is corrected, including:

[0020] Identify the strokes in the dot matrix font that need to be corrected;

[0021] For the strokes that need correction, a dot matrix is ​​added to the correction area.

[0022] Furthermore, the step of adding a dot matrix to the correction area for the stroke that needs correction includes:

[0023] Determine the end point of the stroke that needs to be corrected;

[0024] Within the correction region, a target region adjacent to the endpoint is determined;

[0025] Fill in the new points in the target area

[0026] Further

[0027] Based on the speech information embedded in the dot matrix font, the pronunciation of the text is obtained from a preset pronunciation library, including:

[0028] Detect whether the speech information embedded in the dot matrix font contains polyphonic character identifiers;

[0029] When the polyphonic character identifier is not present in the dot matrix font, the starting address of the speech data for the pronunciation is determined based on the speech information;

[0030] Based on the starting address of the voice data, determine the pronunciation of the text;

[0031] When the polyphonic character identifier exists in the dot matrix font, the pronunciation of the corresponding polyphonic character is determined from the pronunciation library.

[0032] When the polyphonic character identifier exists in the dot matrix font, the pronunciation of the corresponding polyphonic character is determined from the pronunciation library.

[0033] Furthermore, the pronunciation database includes: a set of words for each pronunciation of each polyphonic character;

[0034] When the polyphonic character identifier exists in the character encoding, determining the pronunciation of the corresponding polyphonic character from a preset pronunciation database includes:

[0035] Perform semantic analysis on the text data to determine the characters adjacent to the polyphonic characters;

[0036] Based on the polyphonic character identifier and the phonetic information, a target word set is determined in the pronunciation database;

[0037] Determine whether the target word set includes the adjacent characters;

[0038] When the target word set includes the adjacent characters, the pronunciation corresponding to the target word set is determined to be the pronunciation of the polyphonic character.

[0039] Secondly, embodiments of this application provide an apparatus for converting text data into audiovisual data, including: a receiving module and a data processing and conversion module;

[0040] The receiving module is used to receive text data input from outside;

[0041] The data processing and conversion module is used to determine a dot matrix font from a preset font library based on the text data; convert the text data into corresponding characters based on the dot matrix font; obtain the pronunciation of the characters from a preset pronunciation library based on the speech information embedded in the dot matrix font; and obtain audiovisual data based on the obtained characters and their pronunciations.

[0042] Thirdly, embodiments of this application provide an electronic device for converting text data into audiovisual data, comprising: a memory and a processor; the memory stores a computer program, and the processor executes the computer program to implement the text reading and automatic playback method described in any one of the first aspects.

[0043] Fourthly, embodiments of this application provide a storage medium, including:

[0044] Used to store a character template library, a pronunciation library, and computer-executable instructions, wherein the computer-executable instructions, when executed, implement the method described in any one of the first aspects;

[0045] or,

[0046] Used to store computer-executable instructions, which, when executed, implement the method described in any one of the first aspects.

[0047] Compared with the prior art, this application can achieve at least the following technical effects:

[0048] By embedding the phonetic information of Chinese characters into their fonts, a phonetic-graphic integrated playback technology for Chinese characters is achieved. After importing the audiovisual file as text data, 13x14 dot matrix data with embedded phonetic information is obtained based on the text data and a pre-defined dot matrix font library. Simultaneously, the pronunciation of the character is obtained from a pronunciation library using the phonetic information embedded in the dot matrix font. This converts the text data into audiovisual data and enables synchronized playback. This phonetic-graphic integrated playback technology for Chinese characters eliminates the need for manual recording, saving labor and costs. It also overcomes the limitation of users not being able to freely select text for listening and reading. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in one or more embodiments or prior art of this specification, the accompanying drawings used in the description of the embodiments or prior art will be briefly introduced below. The accompanying drawings described below are only some embodiments recorded in this specification. For those skilled in the art, more technical understanding can be obtained from these drawings without creative effort.

[0050] Figure 1 A flowchart illustrating a method for automatically converting Chinese character documents into audio-visual files and playing them, provided for one or more embodiments of this specification;

[0051] Figure 2 A schematic diagram of the dot matrix structure of a 13×14 dot matrix font with embedded Chinese character speech information provided for one or more embodiments of this specification;

[0052] Figure 3 A flowchart illustrating another method for automatically converting Chinese character documents into audio-visual files and playing them, provided for one or more embodiments of this specification. Detailed Implementation

[0053] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.

[0054] For existing handheld electronic audio-visual products, the audio-visual materials not only need to be pre-recorded, but are mostly determined by the equipment manufacturer, and users cannot freely choose them completely.

[0055] To solve the above technical problems, embodiments of this specification provide a method for converting text data into audio-visual data, as Figure 1 shown, including the following steps:

[0056] Step 1: Receive externally input text data.

[0057] In embodiments of this application, the externally input text data is selected text that the user is interested in or likes. The text content corresponding to the text data is used for playback.

[0058] Step 2: Determine a dot matrix font with embedded voice information from a preset font library according to the text data.

[0059] Step 3: Convert the text data into corresponding characters according to the dot matrix font.

[0060] In embodiments of this application, based on a 13×14 dot matrix font, the Chinese characters to be audio-visualized are converted into dot matrix data. Among them, the dot matrix structure of the 13×14 dot matrix font is as Figure 2 shown.

[0061] In Chinese calligraphy, the ending stroke of a character makes the character appear more artistic. However, the existing dot matrix fonts cannot reflect this effect of the ending stroke. Therefore, in order to make the displayed Chinese characters more beautiful in this application, a correction area is added to the font, and in the form of dots, this effect of the ending stroke is expressed in the correction area.

[0062] For example, the Chinese characters "ai" and "ai" are both represented by a 13×14 dot matrix font. There is a blank column gap between the two Chinese characters, and it is also a preset font correction area (in the actual design of the font, two bytes of this column are embedded with the voice information of the Chinese character); there are two blank lines below the 13×14 dot matrix font, which are used to mark whether the Chinese character is a polyphonic character and whether the font needs to be corrected. Among them, the polyphonic character identifier is set in the row where "4" is located in the bottom two rows in the figure in the form of dots, and the font correction identifier is set in the row where "8" is located in the bottom two rows in the figure in the form of dots.

[0063] Specifically, the correction process is as follows:

[0064] According to the Chinese characters to be listened to and read in the text data, determine the corresponding dot matrix font in the preset 13×14 dot matrix font library. Detect whether the font code of the dot matrix font contains a correction flag, and determine the strokes that need to be corrected in the dot matrix font; for the strokes that need to be corrected, determine the end points of the strokes that need to be corrected; in the correction area, determine the target area adjacent to the end point; fill in new points in the target area. Among them, there is a correction area in the dot matrix font, and the correction area is located between two characters. Determine the strokes that need to be corrected in the dot matrix font; for the strokes that need to be corrected, add dots in the correction area.

[0065] It should be noted that the stroke that needs to be corrected in this application is the right-falling stroke. According to the writing rules, the ending of the right-falling stroke is relatively important. For example, Figure 2 in, the last stroke of "ai" is the right-falling stroke, so the right-falling stroke of "ai" is the stroke that needs to be corrected. And the end point of the right-falling stroke is Figure 2 the points corresponding to the last three rows of "2" in. After that, determine the target area adjacent to the dot matrix in the modification area, and fill in new points in the target area.

[0066] Step 4: According to the voice information of the Chinese characters embedded in the dot matrix font, obtain the pronunciation of the text from the preset pronunciation library.

[0067] In the embodiments of this application, there are often polyphonic characters in the text, so polyphonic character flags are set in the voice information. The process of determining whether it is a polyphonic character based on the polyphonic character flag is: detect whether the dot matrix font contains a polyphonic character flag; when the polyphonic character flag exists in the dot matrix font, determine the pronunciation of the corresponding polyphonic character from the pronunciation library.

[0068] Specifically, the device automatically extracts the corresponding 13×14 dot matrix font code of the document Chinese characters in the preset font library; the last two bytes, 27 and 28, of the font code embed the voice information of the Chinese character (see Figure 2 ); combined with the high two-bit flags of the 26th byte of the font (the highest bit of the "8" row is the font correction flag, and the second highest bit of the "4" row is the polyphonic character flag). Determine whether the character is a non-polyphonic character or a polyphonic character. If it is a non-polyphonic character flag, such as Figure 2 in the character "ai", the stored data in bytes 27 and 28 is: 08 06, which is the voice phonetic sequence value data of the Chinese character. Based on this, the starting address of the voice data can be directly calculated by the program, and then the voice data can be read for reading and playing.

[0069] Based on the database of polyphonic word sets, the specific process of determining the pronunciation of the corresponding text from the pronunciation library is:

[0070] Perform semantic analysis on the audio-visual data to determine the adjacent characters; according to the polyphonic character identifier and the voice information, determine the target word set in the pronunciation library; determine whether the adjacent characters are included in the target word set; when the adjacent characters are included in the target word set, determine that the pronunciation corresponding to the target word set is the pronunciation of the polyphonic character.

[0071] If it is detected that there is a polyphonic character identifier, such as Figure 2 In the "挨" character in, the data stored in bytes 27 and 28 is: 10 30, which is the initial absolute address for reading the set of polyphonic words for this character; the set of words with each pronunciation of the polyphonic character can be extracted from this address. For example, the Chinese character "长" has two pronunciations, "chang" and "zhang". When the pronunciation of "长" is "chang", it is paired with "很、太、剪、江、河、城……", and the paired words form a set of words. When the pronunciation of "长" is "zhang", it is paired with "军、师、团、营、者、辈、子……", and the paired words form another set of words. During the user's listening and reading process, based on each set of words, determine which set of words the words before and after the polyphonic character belong to, and confirm the pronunciation of the corresponding polyphonic character according to the corresponding set of words.

[0072] Detect whether the voice information embedded in the dot matrix font contains a polyphonic character identifier; when there is no polyphonic character identifier in the dot matrix font, determine the starting address of the voice data of the pronunciation according to the voice information. Such as Figure 2 In the "皑" character in, the data stored in bytes 27 and 28 is: 08 06, which is the voice sequence value data of this Chinese character. Based on this, the starting address for reading the voice data can be directly calculated by the program, and then the voice data can be read for reading and playing.

[0073] Step 5, obtain the audio-visual data according to the characters and their pronunciations.

[0074] In the embodiment of the present application, after step 5, the user generates a playlist for the audio-visual data. Then, the user can play the corresponding text data in the order of the playlist.

[0075] In the embodiment of the present application, the user can select multiple Chinese character document materials to be listened to and read. Then, use the identifiers of these Chinese character document materials to make a playlist. Among them, the identifier of the Chinese character document material is the name of the corresponding Chinese character document.

[0076] To illustrate the feasibility of the above technical solution, for each Chinese character in the Chinese character document material to be listened to and read, the present application gives the following example, such as Figure 3 as shown:

[0077] Step S1, after identifying the Chinese character, determine the font extraction address according to the GB code of the Chinese character.

[0078] Step S2: Based on the determined template address, determine the corresponding dot matrix font from the preset 13×14 dot matrix font library with voice information.

[0079] Step S3: Detect whether there is a correction mark in the dot matrix font. If yes, proceed to step S4; otherwise, proceed to step S5.

[0080] Step S4: Determine the correction area and set new points within the correction area.

[0081] Step S5: Based on the speech information embedded in the character pattern data, detect whether there are polyphonic characters in the dot matrix character pattern. If so, proceed to steps S6 and S7; otherwise, proceed to step S8.

[0082] Step S6: Determine the set of words for each pronunciation of a polyphonic character based on the polyphonic character identifier.

[0083] Step S7: Determine the phonetic address of the polyphonic character based on the set of words containing each pronunciation of the polyphonic character.

[0084] Step S8: Obtain the pronunciation of the Chinese characters.

[0085] Step S9: Obtain the audiovisual data of the Chinese characters based on their pronunciation and character model.

[0086] In this embodiment of the application, the audiovisual data of the Chinese characters can be combined to obtain the audiovisual data of the entire article, so as to achieve synchronized playback.

[0087] As can be seen from the above embodiments, this application embeds voice information into dot matrix fonts, establishes the relationship between dot matrix fonts and Chinese character pronunciations, and thus establishes the relationship between text and voice, thereby achieving the technical effect of "sound and shape synchronization" playback.

[0088] This specification provides an apparatus for converting text data into audiovisual data, comprising: a receiving module and a data processing and conversion module;

[0089] The receiving module is used to receive input text data;

[0090] The data processing and conversion module is used to determine a dot matrix font from a preset font library based on the text data; convert the text data into corresponding characters based on the dot matrix font; obtain the pronunciation of the characters from a preset pronunciation library based on the speech information embedded in the dot matrix font; and obtain audiovisual data based on the obtained characters and their pronunciations.

[0091] This specification provides an electronic device for converting text data into audiovisual data, including: a memory and a processor;

[0092] The memory stores a character template database, a speech database, a polyphonic word set database, and Chinese character document files in the local memory. The processor resides a conversion program for audio-visual files. By running the program, it extracts, processes, and calculates the source data, ultimately converting it into corresponding character template data and speech data. The actuator then plays the character template data and speech data on the corresponding display and broadcasting devices.

[0093] The processor is the CPU. This device uses the AU7860 chip with audio decoding from Shanghai Integrated Circuit Co., Ltd. Other microcontroller chips with audio decoding can also be used in this technical solution.

[0094] The memory is a high-capacity erasable NAND flash memory chip. This device uses a Samsung K9F8G08UOM memory chip and a Micro SD card as both fixed and removable storage. Other similar memory chips can also be used. Furthermore, this technology and similar chips can be used to manufacture other audio-visual devices and products.

[0095] This application provides a storage medium for storing computer-executable instructions, which, when executed, implement the methods described in any of the above embodiments. Furthermore, the storage medium may also store a character set library and a pronunciation library.

[0096] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0097] In the 1930s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many improvements to the methodology today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that an improvement to the methodology cannot be implemented using a hardware physical module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0098] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0099] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0100] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.

Claims

1. A method of converting text data into audiovisual data, characterized by, include: Receive input text data; Based on the text data, determine the dot matrix fonts embedded with speech information from the preset font library; Based on the dot matrix font, the text data is converted into corresponding characters; Based on the speech information embedded in the dot matrix font, the pronunciation of the text is obtained from a preset pronunciation library; Based on the obtained text and its pronunciation, audiovisual data is obtained.

2. The method according to claim 1, characterized in that, The dot matrix font is a 13×14 dot matrix font; The step of converting the text data into corresponding characters based on the dot matrix font includes: Based on the 13×14 dot matrix font, the Chinese characters to be read in the text data are converted into dot matrix fonts.

3. The method according to claim 2, characterized in that, The process of converting the Chinese characters to be read from the text data into dot matrix character data based on the 13×14 dot matrix font includes: Based on the Chinese characters to be read in the text data, determine the corresponding dot matrix font in the preset 13×14 dot matrix font library; Detect whether the dot matrix font encoding contains a correction identifier; When the correction identifier exists in the font encoding, the corresponding dot matrix font is corrected.

4. The method according to claim 3, characterized in that, The dot matrix font has a correction area, which is located between two characters; When the correction identifier exists in the font encoding, the corresponding dot matrix font is corrected, including: Identify the strokes in the dot matrix font that need to be corrected; For the strokes that need correction, a dot matrix is ​​added to the correction area.

5. The method according to claim 4, characterized in that, The step of adding a dot matrix to the correction area for the strokes that need correction includes: Determine the end point of the stroke that needs to be corrected; Within the correction region, a target region adjacent to the endpoint is determined; Fill in the new point in the target area.

6. The method according to claim 1, characterized in that, The step of obtaining the pronunciation of the text from a preset pronunciation library based on the speech information embedded in the dot matrix font includes: Detect whether the speech information embedded in the dot matrix font contains polyphonic character identifiers; When the polyphonic character identifier is not present in the dot matrix font, the starting address of the speech data for the pronunciation is determined based on the speech information; Based on the starting address of the voice data, determine the pronunciation of the text; When the polyphonic character identifier exists in the dot matrix font, the pronunciation of the corresponding polyphonic character is determined from the pronunciation library.

7. The method according to claim 6, characterized in that, The pronunciation database includes: a set of words for each pronunciation of each polyphonic character; When the polyphonic character identifier exists in the character template, determining the pronunciation of the corresponding polyphonic character from a preset pronunciation database includes: Perform semantic analysis on the text data to determine the characters adjacent to the polyphonic characters; Based on the polyphonic character identifier and the phonetic information, a target word set is determined in the pronunciation database; Determine whether the target word set includes the adjacent characters; When the target word set includes the adjacent characters, the pronunciation corresponding to the target word set is determined to be the pronunciation of the polyphonic character.

8. An apparatus for converting text data into audiovisual data, characterized by include: Receiving module and data processing and conversion module; The receiving module is used to receive input text data; The data processing and conversion module is used to determine a dot matrix font from a preset font library based on the text data; convert the text data into corresponding characters based on the dot matrix font; and obtain the pronunciation of the characters from a preset pronunciation library based on the speech information embedded in the dot matrix font. Based on the obtained text and its pronunciation, audiovisual data is obtained.

9. An electronic device that converts text data into audiovisual data, comprising: A memory and a processor; the memory stores a computer program, characterized in that the processor, when executing the computer program, implements the text reading and automatic playback method according to any one of claims 1 to 7.

10. A storage medium, characterized by include: Used to store a character template library, a pronunciation library, and computer-executable instructions, wherein the computer-executable instructions, when executed, implement the method described in any one of claims 1-7; or, Used to store computer-executable instructions, which, when executed, implement the method according to any one of claims 1-7.