Artificial intelligence (AI) voice control system and experience device

By establishing a database containing audio data of different regions and ethnic keywords in the voice control system, using geographical location matching and keyword matching, the recognition failure caused by dialect differences is solved, and more efficient voice control is achieved.

CN120544553AInactive Publication Date: 2025-08-26SHANDONG AEROSPACE ARTIFICIAL INTELLIGENCE SECURITY CHIP RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510460962.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-08-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Differences in dialects in different regions and ethnic groups have caused speech recognition systems to fail to recognize when controlling smart homes, which are of low practicality and functionality.

Method used

Establish a keyword database, including keyword audio data samples from different regions and ethnic groups, and improve recognition efficiency through geographical location matching and keyword matching after semantic recognition failure.

Benefits of technology

Through the combination of geographical location matching and keyword database, the recognition accuracy and practicality of the voice control system are improved, ensuring that users can quickly and accurately control smart devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544553A_ABST
    Figure CN120544553A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent AI voice control, and particularly relates to an artificial intelligence AI voice control system. The system comprises a voice acquisition module, a voice recognition module, an equipment control module and a storage unit. Voice module data is stored on the storage unit, the voice module data comprises an instruction statement database and a keyword database, instruction statements are used for controlling the intelligent equipment by using voice input of a user, and keywords are used for generating voice prompt statements containing the keywords; the voice acquisition module comprises a microphone and is used for acquiring a voice instruction of a user and transmitting the voice instruction into the voice recognition module; according to the method, the keyword database is established, after semantic recognition fails, the voice instruction is matched with the keywords in the keyword database, whether the voice instruction contains the keywords in the keyword database or not is judged, if matching succeeds, instruction statements related to the successfully matched keywords are prompted through voice, a user is prompted to continue operation, and the user experience is improved. Therefore, the practicability of the artificial intelligence AI voice control system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent AI voice control technology, and specifically relates to an artificial intelligence AI voice control system and experience device. Background Art

[0002] AI, also known as artificial intelligence, extracts "artificial intelligence" as the core feature of the current technological development of the industry, and fully integrates it with industries such as industry, commerce, and finance to promote the continuous evolution of economic forms, thereby driving the vitality of social and economic entities.

[0003] Voice control technology is an artificial intelligence technology that uses voice input to control devices or systems. It converts speech into text through voice recognition, and uses natural language processing to process and understand the text generated by voice recognition, parsing the user's intentions and commands. "AI" control of the home uses the home as a platform to achieve greater efficiency, safety, energy conservation, intelligence, convenience, and comfort. This is "AI" achieved through the automation and intelligence of home products and anthropomorphic requirements through the network. "AI" frees users' hands by controlling the home, providing a better user experience. "AI" in the home is primarily achieved through voice control, which is mainly achieved through voice recognition and speech synthesis technologies. Typical furniture voice control devices recognize and analyze voice commands, retrieve the corresponding connected modules based on the recognized signals, and control the operation of the corresponding modules.

[0004] However, dialects vary from place to place in my country, and people in different regions have different speaking habits. It is easy for non-standard Mandarin speech recognition to fail and dialects to be unrecognizable. It is impossible to quickly and accurately control smart homes through the voice system, resulting in problems of low practicality and functionality. Summary of the Invention

[0005] In response to the above-mentioned deficiencies in the prior art, the present invention provides an artificial intelligence (AI) voice control system and experience device to solve the problems in the above-mentioned background technology.

[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0007] An artificial intelligence (AI) voice control system, comprising a voice acquisition module, a voice recognition module, a device control module, and a storage unit;

[0008] The storage unit stores voice module data, which includes an instruction statement database and a keyword database. The instruction statement is used to control the smart device using the user's voice input, and the keyword is used to generate a voice prompt statement containing the keyword.

[0009] The voice acquisition module includes a microphone for collecting the user's voice commands and transmitting them to the voice recognition module;

[0010] The voice recognition module includes a speaker, which performs semantic recognition on the user's voice based on the user's voice command collected by the voice collection module. When the semantic recognition is successful, the control command is transmitted to the device control module. When the semantic recognition fails, the voice command is matched with the keywords in the keyword database to determine whether the voice command contains the keywords in the keyword database. If the match is successful, a voice prompt is given with the instruction sentence of the successfully matched keyword, prompting the user to continue the operation;

[0011] The device control module transmits control instructions to the corresponding smart device via Wi-Fi and Bluetooth to control the smart device.

[0012] Furthermore, the keyword database contains keyword voice files containing multiple audio data and label data of keywords, wherein each audio data corresponds to one label data, and the audio data is audio data samples of pronunciation of keywords in different regions and / or different nationalities.

[0013] Furthermore, it also includes a positioning unit that records the geographic location of the current user, matches and compares the voice instructions that failed semantic recognition with all audio data whose label data is the geographic location of the user, and if a keyword is matched, issues a voice prompt statement about the keyword that successfully matched.

[0014] Furthermore, the audio data includes an original file of speech, and MFCC features and GFCC features are extracted from the original file, and feature fusion is performed on the MFCC features and GFCC features to obtain a comprehensive feature vector of the keyword audio data.

[0015] Furthermore, the tag data corresponds to the region and / or ethnicity where the pronunciation audio data sample is collected.

[0016] Furthermore, the voice acquisition module collects the user's voice instructions through a microphone, performs feature extraction and text conversion on the received voice through a voice recognition algorithm, recognizes the user's voice intention based on natural language understanding, and generates control instructions.

[0017] Furthermore, the speech recognition module performs semantic recognition on the user's speech based on the user's speech instructions collected by the microphone. When the semantic recognition fails, the speech instructions are matched with the keywords in the keyword database to determine whether the speech instructions contain the keywords in the keyword database. When the keyword matching is unsuccessful, the ununderstood instructions are output and transmitted to the speaker.

[0018] Furthermore, the audio data of the keyword is stored in the keyword database in a wav format.

[0019] An experience device equipped with the above-mentioned artificial intelligence AI voice control system.

[0020] Compared with the prior art, the present invention has the following beneficial effects:

[0021] 1. The present invention establishes a keyword database. After semantic recognition fails, the voice command is matched with the keywords in the keyword database to determine whether the voice command contains the keywords in the keyword database. If a match is successful, a voice prompt is given for the instruction sentence related to the successfully matched keyword, prompting the user to continue the operation, thereby improving the practicality of the artificial intelligence (AI) voice control system.

[0022] 2. The present invention establishes a keyword voice file to record multiple audio data and label data of the same keyword, wherein each audio data corresponds to a label data, and the audio data is an audio data sample of the pronunciation of the keyword in different regions and / or different nationalities. Through the positioning unit, the geographical location of the current user is recorded, and the voice instructions that fail semantic recognition are matched and compared with all audio data whose label data is the geographical location of the user. If a keyword is matched, a voice prompt sentence related to the successfully matched keyword is issued, thereby narrowing the search by geographical location, quickly performing matching, and improving matching efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a block diagram of an artificial intelligence (AI) voice control system of the present invention;

[0024] Figure 2 This is a flow chart of an artificial intelligence (AI) voice control system of the present invention; DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the present invention, the technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0026] Among them, the drawings are only used for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting this patent; in order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0027] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "inside", "outside" and the like indicate an orientation or position relationship based on the orientation or position relationship shown in the drawings, it is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting this patent. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0028] In the description of the present invention, unless otherwise expressly specified or limited, when the term "connection" or the like appears to indicate a connection relationship between components, such term should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be internal communication between two components or an interaction between two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood in specific circumstances.

[0029] Example 1:

[0030] like Figure 1-2 As shown, an artificial intelligence (AI) voice control system includes a voice acquisition module, a voice recognition module, a device control module, and a storage unit;

[0031] The storage unit stores voice module data, which includes an instruction statement database and a keyword database. The instruction statement is used to control the smart device using the user's voice input, and the keyword is used to generate a voice prompt statement containing the keyword.

[0032] The voice acquisition module includes a microphone for collecting the user's voice commands and transmitting them to the voice recognition module;

[0033] The voice recognition module includes a speaker, which performs semantic recognition on the user's voice based on the user's voice command collected by the voice collection module. When the semantic recognition is successful, the control command is transmitted to the device control module. When the semantic recognition fails, the voice command is matched with the keywords in the keyword database to determine whether the voice command contains the keywords in the keyword database. If the match is successful, a voice prompt is given with a command sentence related to the keyword, prompting the user to continue the operation;

[0034] The device control module transmits control instructions to the corresponding smart device via Wi-Fi and Bluetooth to control the smart device.

[0035] The keyword database contains keyword voice files containing multiple audio data and label data containing keywords, where each audio data item corresponds to a label data item. The audio data items are audio data samples of the pronunciation of keywords from different regions and / or different ethnic groups. MFCC features and GFCC features are extracted from the pronunciation audio data samples and then feature fusion is performed to obtain a comprehensive feature vector of the keyword pronunciation audio data sample. The label data corresponds to the region and / or ethnic group from which the pronunciation audio data sample was collected. Different weight coefficients are assigned to the comprehensive feature vector and the label data, and keywords are matched by calculating similarity during semantic recognition.

[0036] The AI ​​voice control system also includes a positioning unit that records the current user's geographic location and traverses a keyword database based on the geographic location. It first traverses the tag data and compares the voice commands that failed semantic recognition with all audio data with the tag data corresponding to the user's geographic location. If a keyword is matched, a voice prompt related to the keyword is issued. The specific matching method is to extract the MFCC features and GFCC features of the voice command audio and compare them with the MFCC features and GFCC features of the pronunciation audio data samples in the keyword database. The weight coefficient of the combined feature vector of the MFCC features and GFCC features and the tag data is determined based on the similarity between different regions or ethnic languages ​​and Mandarin. For example, in large cities such as Beijing, Shanghai, Guangzhou, and Shenzhen, the regional weight coefficient is lower because not only are there more people speaking Mandarin in these areas, but they also speak more standard Mandarin. In regions such as Tibet and Xinjiang, where the number of people speaking Mandarin is small and non-standard, the regional weight coefficient is larger, facilitating more direct matching and comparison in the keyword database. The specific weight coefficient value can be obtained through extensive experimentation or manually set. The preset keywords include smart home device names including but not limited to TV, air conditioner, light, open, close, and temperature.

[0037] The voice acquisition module collects the user's voice commands through a microphone, extracts features and converts the received voice into text using a voice recognition algorithm, identifies the user's voice intent based on natural language understanding, and generates control commands. Based on a microphone device, the module captures the user's voice input, converts the user's voice into an electrical signal, and converts the electrical signal captured by the microphone into a digital signal using an analog-to-digital converter. The module extracts acoustic features from the digital signal, including Mel-frequency cepstral coefficients, and uses automatic speech recognition (ASR) to convert the user's voice into text based on the extracted acoustic features. The module performs semantic analysis on the converted text using natural language understanding to extract entity information and generate control commands. The generated control commands are then transmitted to the corresponding smart device via wireless communication methods, including Wi-Fi and Bluetooth.

[0038] The speech recognition module is based on the user voice instructions collected by the microphone. When the keyword matching is unsuccessful, it outputs the ununderstood instructions and transmits the ununderstood instructions into the speaker.

[0039] The audio data of the keyword is stored in the keyword database in wav format.

[0040] The present application also provides an artificial intelligence (AI) voice control experience device, which is configured with the above-mentioned artificial intelligence (AI) voice control system.

[0041] The above are only embodiments of the present invention, and the circuits, electronic components and modules involved are all prior art, which can be fully implemented by those skilled in the art. It is needless to say that the content protected by this application does not involve improvements to software and methods. Common knowledge such as the specific structures and characteristics known in the scheme are not described in detail here. Ordinary technicians in the field are aware of all common technical knowledge in the technical field of the invention before the application date or priority date, can obtain all prior art in the field, and have the ability to apply conventional experimental means before that date. Ordinary technicians in the field can improve and implement this scheme in combination with their own abilities under the inspiration given by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the field to implement this application. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several variations and improvements can be made, which should also be regarded as the scope of protection of the present invention. These will not affect the effect of the implementation of the present invention and the practicality of the patent.

Claims

1. An artificial intelligence (AI) voice control system, characterized by: It includes a voice acquisition module, a voice recognition module, a device control module and a storage unit; The storage unit stores voice module data, which includes an instruction statement database and a keyword database. The instruction statement is used to control the smart device using the user's voice input, and the keyword is used to generate a voice prompt statement containing the keyword. The voice acquisition module includes a microphone for collecting the user's voice commands and transmitting them to the voice recognition module; The voice recognition module includes a speaker, which performs semantic recognition on the user's voice based on the user's voice command collected by the voice collection module. When the semantic recognition is successful, the control command is transmitted to the device control module. When the semantic recognition fails, the voice command is matched with the keywords in the keyword database to determine whether the voice command contains the keywords in the keyword database. If the match is successful, a voice prompt is given with the instruction sentence of the successfully matched keyword, prompting the user to continue the operation; The device control module transmits control instructions to the corresponding smart device via Wi-Fi and Bluetooth to control the smart device.

2. The artificial intelligence (AI) voice control system according to claim 1, characterized in that: The keyword database contains keyword voice files containing multiple audio data and label data containing keywords, wherein each audio data corresponds to one label data, and the audio data is audio data samples of the pronunciation of keywords in different regions and / or different nationalities.

3. The artificial intelligence (AI) voice control system according to claim 2, characterized in that: It also includes a positioning unit that records the geographic location of the current user, matches and compares the voice instructions that failed semantic recognition with all audio data with label data that is the geographic location of the user, and if a keyword is matched, issues a voice prompt statement about the keyword that was successfully matched.

4. The artificial intelligence (AI) voice control system according to claim 2, characterized in that: The audio data includes an original file of speech, and MFCC features and GFCC features are extracted from the original file. Feature fusion is performed on the MFCC features and the GFCC features to obtain a comprehensive feature vector of the keyword audio data.

5. The artificial intelligence (AI) voice control system according to claim 2, characterized in that: The tag data corresponds to the region and / or ethnicity from which the pronunciation audio data sample was collected.

6. The artificial intelligence (AI) voice control system according to claim 1, characterized in that: The voice acquisition module collects the user's voice instructions through a microphone, performs feature extraction and text conversion on the received voice through a voice recognition algorithm, recognizes the user's voice intention based on natural language understanding, and generates control instructions.

7. The artificial intelligence (AI) voice control system according to claim 1, characterized in that: The speech recognition module performs semantic recognition on the user's speech based on the user's speech instructions collected by the microphone. When the semantic recognition fails, the speech instruction is matched with the keywords in the keyword database to determine whether the speech instruction contains the keywords in the keyword database. When the keyword matching is unsuccessful, the ununderstood instruction is output and transmitted to the speaker.

8. The artificial intelligence (AI) voice control system according to claim 7, characterized in that: The audio data of the keyword is stored in the keyword database in wav format.

9. An experience device, characterized in that: The experience device is equipped with an artificial intelligence (AI) voice control system as described in any one of claims 1 to 8.