English pronunciation calibration device and system
By monitoring the learner's mouth shape when pronouncing a word and providing the difference in mouth shape between the English text and the standard pronunciation, the problem of insufficient information acquisition in existing English pronunciation learning methods is solved, and learning efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202110049538.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-14
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2041-01-14
AI Technical Summary
Existing English pronunciation learning methods mainly rely on shadowing, and learners obtain little information with low diversity, resulting in insufficient pronunciation learning efficiency and accuracy.
By monitoring the learner's mouth shape when pronouncing, the difference between the English text and the standard pronunciation of the English mouth shape is provided, and more relevant information is obtained by using the memory, word playback component, image acquisition component and pronunciation correction processor to correct pronunciation defects.
It improves the learning efficiency and accuracy of English pronunciation, and helps learners correct pronunciation errors by displaying difference images and syllable missing information on the display screen.
Smart Images

Figure CN112699275B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of English multimedia teaching, and more particularly to an English pronunciation calibration device and an English pronunciation calibration system. Background Art
[0002] Learning English pronunciation is crucial for learning English. Existing English pronunciation learning methods often rely on shadowing, where learners read aloud to determine if their pronunciation is correct. If a pronunciation error occurs, an error message is returned, prompting the learner to repeat until their pronunciation is correct. During the reading process, pronunciation learning software or systems only provide English characters and their corresponding English pronunciation audio. Learners rely solely on this information, resulting in limited and less diverse information. Summary of the Invention
[0003] The purpose of the present invention is to provide an English pronunciation calibration device, which corrects pronunciation defects by monitoring the learner's mouth shape during pronunciation, provides the learner with English text and standard English pronunciation, and provides the English mouth shape differences corresponding to the standard pronunciation. By enabling the learner to obtain more relevant information, the learner's learning efficiency and English pronunciation accuracy are improved.
[0004] The purpose of the present invention is to provide an English pronunciation calibration system, which provides learners with English text, English standard pronunciation, and English mouth shape differences corresponding to the standard pronunciation, thereby enabling learners to obtain more relevant information, thereby improving their learning efficiency and English pronunciation accuracy.
[0005] A first aspect of the present invention provides an English pronunciation calibration device, which includes: a memory, a word playing component, an image acquisition component and a pronunciation correction processor.
[0006] The memory stores a pre-stored word library. Each word in the word library has a corresponding pronunciation video file and a pronunciation waveform graphic file corresponding to the pronunciation audio file. The pronunciation video file is an image of the mouth area changes when each word is correctly pronounced.
[0007] The word playing component has a display screen connected to a memory. The word playing component obtains a current word from a pre-stored word library and displays the current word through the display screen.
[0008] The image acquisition component includes: a fixing plate, two fixing clamping rods, a camera, an audio acquisition module and a wireless transmission module.
[0009] The fixing plate has a fixing surface. The fixing plate has an arcuate edge. One end of each of the two fixing clamping rods is movably connected to the fixing plate, and the fixing clamping rods extend perpendicular to the fixing surface. The other ends of the two fixing clamping rods each have a clamping member; the clamping member has two clamping ends, and the two clamping ends are arranged opposite each other. The opposing ends of the two clamping members each have a clamping surface, and the two clamping members continuously apply a clamping force to the opposing ends. The fixing surface of the fixing plate is arranged in the direction of the two fixing clamping rods.
[0010] The camera is mounted on a fixed surface and is capable of capturing a current video image facing away from the fixed surface. The current video image is a video image of the user's mouth when the user is speaking the current word. The camera is capable of outputting the current video image. The audio acquisition module is mounted on the fixed surface and is capable of capturing the current audio. The audio acquisition module is capable of outputting the current audio. The wireless transmission module is connected to the camera and is capable of transmitting the received current video image to a remote location.
[0011] The pronunciation correction processor is connected to the memory, the word playback component and the image acquisition component. The pronunciation correction processor is configured as follows:
[0012] The system can receive the current video image sent by the wireless transmission module and the current audio from the audio acquisition module. It can retrieve the pronunciation video file and pronunciation waveform graphic file corresponding to the current word from a pre-stored word library. It can also retrieve the pronunciation waveform corresponding to each pronunciation syllable in the pronunciation waveform graphic file and the corresponding image in the pronunciation video file corresponding to the pronunciation waveform.
[0013] Based on the correspondence between the current video image playback frame and the current audio playback frame, a current image of each currently pronounced syllable is obtained. The current image is compared with the corresponding image in the pronunciation video file corresponding to the pronunciation waveform in the word library. A difference image of each currently pronounced syllable is obtained. The difference image is sent to the word playback component. The word playback component displays the difference image on a display screen.
[0014] In another embodiment of the English pronunciation calibration device of the present invention, the memory further stores a syllable waveform graphic library, in which each English syllable and its corresponding syllable waveform are pre-stored.
[0015] In another embodiment of the English pronunciation calibration device of the present invention, the pronunciation correction processor is further configured to obtain the pronunciation waveform corresponding to each pronunciation syllable in the pronunciation waveform graphic file and the corresponding image in the pronunciation video file corresponding to the pronunciation waveform, including the following specific steps:
[0016] Obtain the playback frame coordinates from the pronunciation waveform image file. Obtain the syllable information corresponding to the current word based on the current word. Search a pre-stored syllable waveform image library based on the syllable information corresponding to the current word to obtain the syllable waveform corresponding to each syllable information. Identify the syllable frame information in the pronunciation waveform image with the playback frame coordinates based on the syllable waveform corresponding to each syllable information.
[0017] The playback frame coordinates of the pronunciation waveform graphic file correspond to the playback frame coordinates of the pronunciation video file, and the corresponding image in the pronunciation video file corresponding to the pronunciation waveform is extracted from the pronunciation video file according to the identification syllable frame information.
[0018] In another embodiment of the English pronunciation calibration device of the present invention, the pronunciation correction processor is further configured to obtain a current pronunciation waveform corresponding to the current audio according to the playback frame of the current audio. The current pronunciation waveform is compared with the syllable waveform corresponding to each syllable information to obtain the current pronunciation syllable;
[0019] Compare the currently pronounced syllable with the syllable information corresponding to the current word to determine whether a syllable is missing. If so, output syllable missing information. Send the syllable missing information to the word playback component. The word playback component displays the syllable missing information on the display screen.
[0020] In another embodiment of the English pronunciation calibration device of the present invention, the step of obtaining the current image of each currently pronounced syllable according to the correspondence between the current video image playback frame and the current audio playback frame specifically includes:
[0021] The current pronunciation waveform is obtained based on the current audio playback frame. The playback frames corresponding to each current syllable are obtained based on the waveform segments of the waveform change in the pronunciation waveform. The current image of each current pronunciation syllable is obtained by comparing the playback frames corresponding to each current syllable with the playback frames of the current video image.
[0022] In another embodiment of the English pronunciation calibration device of the present invention, the difference image is an image of a non-superimposed portion of the mouth image of the current reader superimposed with a corresponding mouth pronunciation image in a pre-stored word library.
[0023] In another embodiment of the English pronunciation calibration device of the present invention, the transparency of the difference image is greater than 90%, and the edge of the difference image has a dotted border.
[0024] In another embodiment of the English pronunciation calibration device of the present invention, the pronunciation correction processor is further configured to send the pronunciation video file and the pronunciation audio file corresponding to the current word in the pre-stored word library to the word playback component. The word playback component plays the pronunciation video file and the pronunciation audio file through the display screen.
[0025] Meanwhile, the present invention also provides an English pronunciation calibration system, which includes: a local terminal and a remote terminal. The remote terminal and the local terminal are connected via wireless communication.
[0026] The local end includes: a memory, a word playing component, an image collecting component and a pronunciation correction processor.
[0027] The memory stores a pre-stored word library. Each word in the word library has a corresponding pronunciation video file and a pronunciation waveform graphic file corresponding to the pronunciation audio file. The pronunciation video file is an image of the mouth area changes when each word is correctly pronounced.
[0028] The word playing component has a display screen connected to a memory. The word playing component obtains a current word from a pre-stored word library and displays the current word through the display screen.
[0029] The image acquisition component includes a fixing plate, two fixing clamps, a camera, an audio acquisition module and a wireless transmission module:
[0030] The fixing plate has a fixing surface. The fixing plate has an arcuate edge. One end of each of the two fixing clamping rods is movably connected to the fixing plate, and the fixing clamping rods extend perpendicular to the fixing surface. The other ends of the two fixing clamping rods each have a clamping member; the clamping member has two clamping ends, and the two clamping ends are arranged opposite each other. The opposing ends of the two clamping members each have a clamping surface, and the two clamping members continuously apply a clamping force to the opposing ends. The fixing surface of the fixing plate is arranged in the direction of the two fixing clamping rods.
[0031] The camera is mounted on a fixed surface and is capable of capturing a current video image facing away from the fixed surface. The current video image is a video image of the user's mouth when the user is speaking the current word. The camera is capable of outputting the current video image.
[0032] The audio acquisition module is set on a fixed surface and can collect the current audio. The audio acquisition module can output the current audio. The wireless transmission module is connected to the camera and can transmit the received current video image to a remote location.
[0033] The pronunciation correction processor is connected to the memory, the word playback component and the image acquisition component. The pronunciation correction processor is configured as follows:
[0034] The system can receive the current video image sent by the wireless transmission module and the current audio from the audio acquisition module. It can retrieve the pronunciation video file and pronunciation waveform graphic file corresponding to the current word from a pre-stored word library. It can also retrieve the pronunciation waveform corresponding to each pronunciation syllable in the pronunciation waveform graphic file and the corresponding image in the pronunciation video file corresponding to the pronunciation waveform.
[0035] Based on the correspondence between the current video image playback frame and the current audio playback frame, a current image of each currently pronounced syllable is obtained. The current image is compared with the corresponding image in the pronunciation video file corresponding to the pronunciation waveform in the word library. A difference image of each currently pronounced syllable is obtained. The difference image is sent to the word playback component. The word playback component displays the difference image on a display screen.
[0036] In another embodiment of the English pronunciation calibration system of the present invention, the memory further stores a syllable waveform graphic library, in which each English syllable and its corresponding syllable waveform are pre-stored.
[0037] The characteristics, technical features, advantages and implementation methods of the English pronunciation calibration device and the English pronunciation calibration system will be further explained in a clear and understandable manner with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic diagram for illustrating the composition of an English pronunciation calibration device in one embodiment of the present invention.
[0039] Figure 2 It is a schematic structural diagram for illustrating an image acquisition component in one embodiment of the present invention.
[0040] Figure 3 It is a schematic diagram for illustrating the structure of the image acquisition component when worn in one embodiment of the present invention.
[0041] Figure 4 It is a schematic diagram for illustrating the composition of an English pronunciation calibration system in one embodiment of the present invention. DETAILED DESCRIPTION
[0042] In order to have a clearer understanding of the technical features, purposes and effects of the invention, the specific embodiments of the present invention are now described with reference to the accompanying drawings. The same reference numerals in the drawings represent components with the same structure or similar structures but the same functions.
[0043] As used herein, "schematic" means "serving as an example, instance, or illustration." Any illustration or implementation described herein as "schematic" should not be construed as a preferred or advantageous technical solution. For simplicity, each figure schematically depicts only the portions relevant to the exemplary embodiment and does not represent the actual structure or true proportions of the product.
[0044] like Figure 1 As shown, the first aspect of the present invention provides an English pronunciation calibration device, which includes: a memory 10, a word playing component 20, an image acquisition component 30 and a pronunciation correction processor 40.
[0045] like Figure 1 As shown, the memory 10 stores a pre-stored word library. Each word in the word library has a corresponding pronunciation video file and a pronunciation waveform graphic file corresponding to the pronunciation audio file. The pronunciation video file is an image of the mouth area changes when each word is correctly pronounced, that is, an image of the mouth shape that continuously changes during pronunciation.
[0046] like Figure 1 As shown, the word playing component 20 has a display screen 21 which is connected to the memory 10. The word playing component 20 obtains the current word from the pre-stored word library and displays the current word through the display screen 21.
[0047] like Figure 1 、 2 As shown in FIG. 3 , the image acquisition component 30 includes: a fixing plate 31 , two fixing clamping rods 32 , 33 , a camera 34 , an audio acquisition module 35 and a wireless transmission module 36 .
[0048] like Figure 1 、 2 As shown in Figures 3 and 4, the fixed plate 31 has a fixed surface 51. The fixed plate 31 has an arc-shaped edge. One end of the two fixed clamping rods 32 and 33 is movably connected to the fixed plate 31 respectively, and the extension direction of the fixed clamping rod is perpendicular to the fixed surface 51. The other end of the two fixed clamping rods 32 and 33 has a clamping member 52 and 53 respectively. The clamping members 52 and 53 have two clamping ends and the two clamping ends are arranged opposite to each other. The opposite ends of the two clamping members respectively have a clamping surface, and the two clamping members continuously apply a clamping force to their opposite ends. Among them, the fixed surface 51 of the fixed plate 31 is arranged in the direction of the two fixed clamping rods 32 and 33.
[0049] like Figure 1 、 2 As shown in FIG3 , camera 34 is mounted on fixed surface 51 and is capable of capturing a current video image facing away from fixed surface 51. The current video image is a video image of the user's mouth when the user is speaking the current word. Camera 34 is capable of outputting the current video image. Audio acquisition module 35 is mounted on fixed surface 51 and is capable of capturing the current audio. Audio acquisition module 35 is capable of outputting the current audio. Wireless transmission module 36 is connected to camera 34 and is capable of transmitting the received current video image to a remote location.
[0050] like Figure 1 As shown, the pronunciation correction processor 40 is connected to the memory 10, the word playback component 20 and the image acquisition component 30. The pronunciation correction processor 40 is configured as follows:
[0051] The system can receive the current video image sent by the wireless transmission module 36 and the current audio from the audio acquisition module 35. It can retrieve the pronunciation video file and pronunciation waveform graphic file corresponding to the current word from the pre-stored word library. It can also retrieve the pronunciation waveform corresponding to each syllable in the pronunciation waveform graphic file and the corresponding image in the pronunciation video file corresponding to the pronunciation waveform.
[0052] Based on the correspondence between the playback frame of the current video image and the playback frame of the current audio, a current image of each currently pronounced syllable is obtained. The current image is compared with the corresponding image in the pronunciation video file corresponding to the pronunciation waveform in the word library. A difference image of each currently pronounced syllable is obtained. The difference image is sent to the word playback component 20. The word playback component 20 displays the difference image on the display screen 21.
[0053] In another embodiment of the English pronunciation calibration device of the present invention, the memory 10 further stores a syllable waveform graphic library, in which each English syllable and its corresponding syllable waveform are pre-stored.
[0054] In another embodiment of the English pronunciation calibration device of the present invention, the pronunciation correction processor 40 is further configured to obtain the pronunciation waveform corresponding to each pronunciation syllable in the pronunciation waveform graphic file and the corresponding image in the pronunciation video file corresponding to the pronunciation waveform. The specific steps include:
[0055] Obtain the playback frame coordinates from the pronunciation waveform image file. Obtain the syllable information corresponding to the current word based on the current word. Search a pre-stored syllable waveform image library based on the syllable information corresponding to the current word to obtain the syllable waveform corresponding to each syllable information. Identify the syllable frame information in the pronunciation waveform image with the playback frame coordinates based on the syllable waveform corresponding to each syllable information.
[0056] The playback frame coordinates of the pronunciation waveform graphic file correspond to the playback frame coordinates of the pronunciation video file, and the corresponding image in the pronunciation video file corresponding to the pronunciation waveform is extracted from the pronunciation video file according to the identification syllable frame information.
[0057] In another embodiment of the English pronunciation calibration device of the present invention, the pronunciation correction processor 40 is further configured to obtain a current pronunciation waveform corresponding to the current audio according to the playback frame of the current audio, and obtain the current pronunciation syllable according to the current pronunciation waveform and the syllable waveform corresponding to each syllable information.
[0058] Compare the current pronunciation syllable and the syllable information corresponding to the current word to determine whether a syllable is missing. If so, output syllable missing information. Send the syllable missing information to the word playback component 20. The word playback component 20 displays the syllable missing information through the display screen 21.
[0059] In another embodiment of the English pronunciation calibration device of the present invention, the step of obtaining the current image of each currently pronounced syllable according to the correspondence between the current video image playback frame and the current audio playback frame specifically includes:
[0060] The current pronunciation waveform is obtained based on the current audio playback frame. The playback frames corresponding to each current syllable are obtained based on the waveform segments of the waveform change in the pronunciation waveform. The current image of each current pronunciation syllable is obtained by comparing the playback frames corresponding to each current syllable with the playback frames of the current video image.
[0061] In another embodiment of the English pronunciation calibration device of the present invention, the difference image is an image of a non-superimposed portion of the mouth image of the current reader superimposed with a corresponding mouth pronunciation image in a pre-stored word library.
[0062] In another embodiment of the English pronunciation calibration device of the present invention, the transparency of the difference image is greater than 90%, and the edge of the difference image has a dotted border.
[0063] In another embodiment of the English pronunciation calibration device of the present invention, the pronunciation correction processor 40 is further configured to send the pronunciation video file and the pronunciation audio file corresponding to the current word in the pre-stored word library to the word playback component 20. The word playback component 20 plays the pronunciation video file and the pronunciation audio file through the display screen 21.
[0064] like Figure 4 As shown, the present invention also provides an English pronunciation calibration system, which includes: a local terminal 101 and a remote terminal 201. The remote terminal 201 is connected to the local terminal 101 via wireless communication.
[0065] The local terminal 101 includes: a memory 10, a word playing component 20, an image acquisition component 30 and a pronunciation correction processor 40.
[0066] The memory 10 stores a pre-stored word library. Each word in the word library has a corresponding pronunciation video file and a pronunciation waveform graphic file corresponding to the pronunciation audio file. The pronunciation video file is an image of the mouth area changes when each word is correctly pronounced.
[0067] The word playing component 20 has a display screen 21 connected to the memory 10. The word playing component 20 obtains the current word from the pre-stored word library and displays the current word through the display screen 21.
[0068] The image acquisition component 30 includes a fixing plate 31, two fixing clamping rods 32 and 33, a camera 34, an audio acquisition module 35 and a wireless transmission module 36:
[0069] The fixed plate 31 has a fixed surface 51. The fixed plate 31 has an arcuate edge. One end of each of the two fixed clamping rods is movably connected to the fixed plate 31, and the fixed clamping rods extend perpendicular to the fixed surface 51. The other end of each of the two fixed clamping rods has a clamping member, and the two clamping ends are arranged opposite each other. The opposing ends of the two clamping members each have a clamping surface, and the two clamping members continuously apply a clamping force to their opposing ends. The fixed surface 51 of the fixed plate 31 is arranged toward the two fixed clamping rods 32 and 33.
[0070] The camera 34 is disposed on the fixed surface 51 and can capture the current video image facing away from the fixed surface 51. The current video image is a video image of the user's mouth when the user is reading the current word. The camera 34 can output the current video image.
[0071] The audio collection module 35 is disposed on the fixed surface 51 and can collect the current audio. The audio collection module 35 can output the current audio. The wireless transmission module 36 is connected to the camera 34 and can transmit the received current video image to a remote location.
[0072] The pronunciation correction processor 40 is connected to the memory 10, the word playback component 20 and the image acquisition component 30. The pronunciation correction processor 40 is configured as follows:
[0073] The system can receive the current video image sent by the wireless transmission module 36 and the current audio from the audio acquisition module 35. It can retrieve the pronunciation video file and pronunciation waveform graphic file corresponding to the current word from the pre-stored word library. It can also retrieve the pronunciation waveform corresponding to each syllable in the pronunciation waveform graphic file and the corresponding image in the pronunciation video file corresponding to the pronunciation waveform.
[0074] Based on the correspondence between the playback frame of the current video image and the playback frame of the current audio, a current image of each currently pronounced syllable is obtained. The current image is compared with the corresponding image in the pronunciation video file corresponding to the pronunciation waveform in the word library. A difference image of each currently pronounced syllable is obtained. The difference image is sent to the word playback component 20. The word playback component 20 displays the difference image on the display screen 21.
[0075] In another embodiment of the English pronunciation calibration system of the present invention, the memory 10 further stores a syllable waveform graphic library, in which each English syllable and its corresponding syllable waveform are pre-stored.
[0076] A non-volatile computer-readable storage medium can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as the program instructions / modules corresponding to the speech signal processing method in the embodiments of the present invention. One or more program instructions stored in the non-volatile computer-readable storage medium, when executed by a processor, perform the speech signal processing method in any of the aforementioned method embodiments.
[0077] The non-volatile computer-readable storage medium may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the voice signal processing device, etc. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the voice signal processing device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0078] An embodiment of the present invention further provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any of the above-mentioned speech signal processing methods.
[0079] The above-mentioned product can execute the method provided by the embodiment of the present invention, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided by the embodiment of the present invention.
[0080] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to:
[0081] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and are primarily designed to provide voice and data communications. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.
[0082] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, have computing and processing capabilities, and generally also have mobile Internet access. These terminals include PDAs, MIDs, and UMPCs.
[0083] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.
[0084] (4) Server: A device that provides computing services. The server consists of a processor, hard disk, memory, system bus, etc. The server is similar to a general computer architecture, but because it needs to provide highly reliable services, it has higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.
[0085] (5) Other electronic devices with data interaction functions.
[0086] The device embodiments described above are merely illustrative. Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of these modules may be selected based on actual needs to achieve the objectives of this embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0087] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or certain parts of the embodiment.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An English pronunciation calibration device, characterized in that: It includes: a memory storing a pre-stored word library; Each word in the word library has a corresponding pronunciation video file and a pronunciation waveform graphic file corresponding to the pronunciation video file; the pronunciation video file is an image of the mouth area changes when each word is correctly pronounced; a word playing component having a display screen connected to the memory; The word playing component obtains the current word from the pre-stored word library; and displaying the current word via the display screen; An image acquisition component, comprising: A fixing plate having a fixing surface; the fixing plate having an arc-shaped edge; Two fixing clamping rods, one end of each fixing clamping rod is movably connected to the fixing plate and the extension direction of each fixing clamping rod is perpendicular to the fixing surface, and the other end of each fixing clamping rod has a clamping member; the clamping member has two clamping ends and the two clamping ends are arranged opposite to each other; the opposite ends of each clamping member have a clamping surface, and the two clamping members continuously apply a clamping force to the opposite ends thereof; Wherein, the fixing surface of the fixing plate is arranged toward the direction of the two fixing clamping rods; a camera, which is disposed on the fixed surface and is capable of capturing a current video image facing away from the fixed surface; the current video image is a video image of the user's mouth when the user is speaking the current word; the camera is capable of outputting the current video image; an audio acquisition module, which is disposed on the fixed surface and capable of acquiring current audio; the audio acquisition module is capable of outputting the current audio; and a wireless transmission module, which is connected to the camera and capable of transmitting the received current video image to a remote location; A pronunciation correction processor is connected to the memory, the word playback component and the image acquisition component; the pronunciation correction processor is configured as follows: The system can receive the current video image sent by the wireless transmission module and the current audio of the audio acquisition module; obtain the pronunciation video file and pronunciation waveform graphic file corresponding to the current word from the pre-stored word library; obtain the pronunciation waveform corresponding to each pronunciation syllable in the pronunciation waveform graphic file and the corresponding image in the pronunciation video file corresponding to the pronunciation waveform; According to the correspondence between the playback frame of the current video image and the playback frame of the current audio, a current image of each currently pronounced syllable is obtained; according to the current image, a corresponding image in the pronunciation video file corresponding to the pronunciation waveform in the word library is compared; a difference image of each currently pronounced syllable is obtained; and the difference image is sent to the word playback component; The word playing component displays the difference image through a display screen.
2. The English pronunciation calibration device according to claim 1, characterized in that: The memory also stores a syllable waveform graphic library; the syllable waveform graphic library pre-stores each English syllable and its corresponding syllable waveform.
3. The English pronunciation calibration device according to claim 1, characterized in that: The pronunciation correction processor is further configured to obtain the pronunciation waveform corresponding to each pronunciation syllable in the pronunciation waveform graphic file and the corresponding image in the pronunciation video file corresponding to the pronunciation waveform, including the following specific steps: Get the playback frame coordinates in the pronunciation waveform graphic file; Acquire the syllable information corresponding to the current word according to the current word; retrieve the pre-stored syllable waveform graphic library according to the syllable information corresponding to the current word, and acquire the syllable waveform corresponding to each syllable information; Identifying syllable frame information in the pronunciation waveform graph having the playback frame coordinates according to the syllable waveform corresponding to each syllable information; The playback frame coordinates of the pronunciation waveform graphic file correspond to the playback frame coordinates of the pronunciation video file, and the corresponding image in the pronunciation video file corresponding to the pronunciation waveform is extracted from the pronunciation video file according to the identification syllable frame information.
4. The English pronunciation calibration device according to claim 1, characterized in that: The pronunciation correction processor is further configured to obtain a current pronunciation waveform corresponding to the current audio according to the playback frame of the current audio; and obtain a current pronunciation syllable according to the current pronunciation waveform and the syllable waveform corresponding to each syllable information; Comparing the currently pronounced syllable with the syllable information corresponding to the current word to determine whether a syllable is missing, and if so, outputting syllable missing information; Sending the syllable missing information to the word playing component; The word playing component displays the syllable missing information through a display screen.
5. The English pronunciation calibration device according to claim 4, characterized in that: The step of obtaining the current image of each currently pronounced syllable according to the correspondence between the current video image playback frame and the current audio playback frame specifically includes: Obtaining a current pronunciation waveform according to the playback frame of the current audio; obtaining a playback frame corresponding to each current syllable according to each waveform segment of the waveform change in the pronunciation waveform; The current image of each currently pronounced syllable is obtained by comparing the playback frame corresponding to each current syllable with the playback frame of the current video image.
6. The English pronunciation calibration device according to claim 1, characterized in that: The difference image is an image of the non-superimposed portion of the image after the mouth image of the current reader is superimposed with the corresponding mouth pronunciation image in the pre-stored word library.
7. The English pronunciation calibration device according to claim 1, characterized in that: The transparency of the difference image is greater than 90%, and the edge of the difference image has a dotted border.
8. The English pronunciation calibration device according to claim 1, characterized in that: The pronunciation correction processor is further configured to: Sending the pronunciation video file corresponding to the current word in the pre-stored word library to the word playing component; The word playing component plays the pronunciation video file through a display screen.
9. An English pronunciation calibration system, characterized in that: It includes: A local end and a remote end; The remote end is connected to the local end via wireless communication; The local end includes: A memory storing a pre-stored word library; each word in the word library has a corresponding pronunciation video file and a pronunciation waveform graphic file corresponding to the pronunciation video file; the pronunciation video file is an image of the mouth area changes when each word is correctly pronounced; A word playing component having a display screen connected to the memory; the word playing component obtains a current word from the pre-stored word library; and displays the current word through the display screen; An image acquisition component, comprising: A fixing plate having a fixing surface; the fixing plate having an arc-shaped edge; Two fixing clamping rods, one end of each fixing clamping rod is movably connected to the fixing plate and the extension direction of each fixing clamping rod is perpendicular to the fixing surface, and the other end of each fixing clamping rod has a clamping member; the clamping member has two clamping ends and the two clamping ends are arranged opposite to each other; the opposite ends of each clamping member have a clamping surface, and the two clamping members continuously apply a clamping force to the opposite ends thereof; Wherein, the fixing surface of the fixing plate is arranged toward the direction of the two fixing clamping rods; a camera, which is disposed on the fixed surface and is capable of capturing a current video image facing away from the fixed surface; the current video image is a video image of the user's mouth when the user is speaking the current word; the camera is capable of outputting the current video image; an audio acquisition module, which is disposed on the fixed surface and capable of acquiring current audio; the audio acquisition module is capable of outputting the current audio; and a wireless transmission module, which is connected to the camera and capable of transmitting the received current video image to a remote location; A pronunciation correction processor is connected to the memory, the word playback component and the image acquisition component; the pronunciation correction processor is configured as follows: The system can receive the current video image sent by the wireless transmission module and the current audio of the audio acquisition module; obtain the pronunciation video file and pronunciation waveform graphic file corresponding to the current word from the pre-stored word library; obtain the pronunciation waveform corresponding to each pronunciation syllable in the pronunciation waveform graphic file and the corresponding image in the pronunciation video file corresponding to the pronunciation waveform; According to the correspondence between the playback frame of the current video image and the playback frame of the current audio, a current image of each currently pronounced syllable is obtained; according to the current image, a corresponding image in the pronunciation video file corresponding to the pronunciation waveform in the word library is compared; a difference image of each currently pronounced syllable is obtained; and the difference image is sent to the word playback component; The word playing component displays the difference image through a display screen.
10. The English pronunciation calibration system according to claim 9, characterized in that: The memory also stores a syllable waveform graphic library; the syllable waveform graphic library pre-stores each English syllable and its corresponding syllable waveform.
Citation Information
Patent Citations
System and method for pronunciation correction
CN107424450A
Pronunciation correction of text-to-speech systems between different spoken languages
US20090006097A1