A vehicle-mounted voice control system and method

By designing an on-board voice control system, including voice reception, recognition, service adaptation and execution devices, the problem of long response time for voice recognition in the prior art is solved, and a fast response and high-accuracy voice control experience is achieved.

CN115691492BActive Publication Date: 2025-06-20CHONGQING CHANGAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211340123.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-29
Publication Date
2025-06-20
Estimated Expiration
2042-10-29

AI Technical Summary

Technical Problem

After the existing vehicle voice recognition technology obtains voice intentions at the terminal, it needs to pass protocol analysis and calls from multiple execution modules, resulting in long reaction time and slow feedback, which affects the user experience.

Method used

A vehicle-mounted voice control system is designed, including a voice receiving device, a voice recognition device, a service adapter device and an execution device. The recording data is obtained through the voice receiving device, the voice recognition device recognizes and processes the voice data, generates pre-start instructions and execution instructions, the service adapter matches the semantic information and generates execution instructions, and the execution device starts or performs corresponding operations according to the instructions.

Benefits of technology

It improves the response experience of voice operations, shortens the time from voice analysis to execution, improves the efficiency and accuracy of semantic analysis, and improves the user's voice control experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115691492B_ABST
    Figure CN115691492B_ABST
Patent Text Reader

Abstract

The present invention provides an in-vehicle voice control system and method. The system includes: a voice receiving device, which is used to acquire recording data, extract the valid content of the recording data, and obtain voice data; a voice recognition device, connected to the voice receiving device, the voice recognition device receives and processes the voice data, acquires the execution object and the optimal semantic information of the voice data, and the voice recognition device generates a pre-start instruction according to the execution object; a service adaptation device, connected to the voice recognition device, the service adaptation device receives the optimal semantic information, matches the service content corresponding to the optimal semantic information, and generates an execution instruction; and an execution device, connected to the service adaptation device and the voice recognition device, when receiving the pre-start instruction, the execution device starts, and when receiving the execution instruction, the execution device performs the operation corresponding to the optimal semantic information. The in-vehicle voice control system and method provided by the present invention can accurately recognize voice content and make a quick response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice interaction technology, and particularly relates to an in-vehicle voice control system and method. Background Art

[0002] Speech recognition technology, also known as Automatic Speech Recognition (ASR), aims to convert the lexical content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.

[0003] The key to speech recognition technology lies in the recognition accuracy of speech, and the research focus of speech recognition technology is mainly on the optimization of recognition models, etc. And the user's experience of using speech is more reflected in the interactive response speed of speech. In current in-vehicle speech recognition technology, after the terminal obtains the speech intention, it needs to parse the protocol and call multiple execution modules to achieve speech recognition and interactive response, with a long reaction time and slow feedback, which greatly affects the user experience. Summary of the Invention

[0004] In view of the above-mentioned disadvantages of the prior art, the present invention provides one to solve the above technical problems.

[0005] An in-vehicle voice control system provided by the present invention includes:

[0006] A voice receiving device, configured to obtain recording data and extract the valid content of the recording data to obtain voice data;

[0007] A voice recognition device, connected to the voice receiving device, the voice recognition device receives and processes the voice data, obtains the execution object and the optimal semantic information of the voice data, and the voice recognition device generates a pre-start instruction according to the execution object;

[0008] A service adaptation device, connected to the voice recognition device, the service adaptation device receives the optimal semantic information and matches the service content corresponding to the optimal semantic information to generate an execution instruction; and

[0009] An execution device, connected to the service adaptation device and the voice recognition device, when receiving the pre-start instruction, the execution device starts, and when receiving the execution instruction, the execution device executes the operation corresponding to the optimal semantic information.

[0010] In an embodiment of the present invention, the voice recognition device includes:

[0011] An audio recognition module, connected to the voice receiving device, which receives the voice data and converts the voice data into text data;

[0012] A semantic recognition module, connected to the audio recognition module, which receives the text data and offline analyzes the semantics of the text data to obtain offline semantic data;

[0013] An intention pre-judgment module, connected to the audio recognition module and the execution device, which receives the text data, obtains the execution object of the text data, and generates the pre-start instruction according to the execution object; and

[0014] A voice cloud module, connected to the audio recognition module, which online analyzes the semantics of the text data to obtain online semantic data.

[0015] In an embodiment of the present invention, the voice recognition device includes a semantic arbitration module, which is connected to the semantic recognition module and the voice cloud module. The semantic arbitration module receives the online semantic data and the offline semantic data, and screens out the optimal semantic information from the online semantic data and the offline semantic data.

[0016] In an embodiment of the present invention, the audio recognition module includes:

[0017] A wake-up word recognition unit, connected to the voice receiving device, which receives the voice data and recognizes the wake-up word in the voice data according to the audio characteristics;

[0018] An instruction segment extraction unit, connected to the wake-up word recognition unit. When the wake-up word is included in the voice data, the voice data after the wake-up word is intercepted in the recording order to obtain instruction segment data; and

[0019] An audio conversion unit, connected to the instruction segment extraction unit, which receives the instruction segment data and converts the instruction segment data into text data.

[0020] In an embodiment of the present invention, the semantic recognition module includes:

[0021] An execution instruction confirmation unit, which receives the text data and determines whether the execution instruction is included in the text data; and

[0022] A semantic analysis unit, connected to the execution instruction confirmation unit. When the instruction is included in the text data, the semantic analysis unit receives the text data and analyzes the semantic information of the text data to obtain the offline semantic data.

[0023] In one embodiment of the present invention, the voice recognition device includes a custom intent prediction library, and the custom corpus stores a plurality of activation words.

[0024] In one embodiment of the present invention, the intent pre-judgment module includes:

[0025] A long sentence splitting unit that receives the text data and splits the text data into multiple segments word by word or according to phrases;

[0026] An intent object matching unit connected to the long sentence splitting unit and the custom corpus. While the long sentence splitting unit splits the text data, the intent object matching unit receives the split text data and compares the text data with the words in the custom corpus until the execution object of the text data is obtained; and

[0027] A pre-start unit connected to the intent object matching unit, and the pre-start unit matches the execution device corresponding to the execution object and generates the pre-start instruction.

[0028] In one embodiment of the present invention, the voice receiving device includes:

[0029] A recording unit for recording the voice of the user to obtain recording data; and

[0030] A data processing unit that receives the recording data and removes the interference frequency band of the recording data to obtain voice data.

[0031] The present invention provides a vehicle-mounted voice control method, including the following steps:

[0032] Obtain recording data through a voice receiving device, extract the valid content of the recording data to obtain voice data;

[0033] Receive and process the voice data through a voice recognition device, obtain the execution object and the optimal semantic information of the voice data, and generate a pre-start instruction through the voice recognition device according to the execution object;

[0034] Receive the optimal semantic information through a service adaptation device, match the service content corresponding to the optimal semantic information, and generate an execution instruction; and

[0035] Receive the pre-start instruction and the execution instruction through an execution device. When the execution device receives the pre-start instruction, the execution device starts. When the execution device receives the execution instruction, the execution device executes the operation corresponding to the optimal semantic information.

[0036] In an embodiment of the present invention, the step of obtaining the pre-start instruction includes:

[0037] Converting the voice data into text data through an audio recognition module;

[0038] Receiving the text data through a long sentence splitting unit, and splitting the text data into multiple segments word by word or according to phrases;

[0039] While the long sentence splitting unit splits the text data, receiving the multiple segments of the text data split out through an intent object matching unit, and comparing each segment of the text data with the words in a custom corpus until the execution object of the text data is obtained; and

[0040] Matching the execution device corresponding to the execution object through a pre-start unit, and generating the pre-start instruction.

[0041] Advantages of the present invention: The in-vehicle voice control system in the present invention can pre-start the corresponding execution device according to the intent of the voice content while parsing the semantic content of the voice, thereby improving the response experience of the user's voice operation, and being beneficial to quickly execute the operation after semantic parsing. The voice control system provided by the present invention is also applicable to online and offline semantic parsing, with high parsing efficiency and high parsing accuracy, and being beneficial to sample training to continuously improve the accuracy of semantic parsing. The voice control system provided by the present invention has the characteristics of accurate recognition and fast response, greatly improving the user experience of voice control.

[0042] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Brief Description of the Drawings

[0043] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can obtain other drawings according to these drawings without creative efforts. In the drawings:

[0044] Figure 1 is a schematic diagram of in-vehicle voice control shown in an exemplary embodiment of this application.

[0045] Figure 2 is an implementation schematic diagram of a voice control device shown in an exemplary embodiment of this application.

[0046] Figure 3 is a flowchart of a voice control method shown in an exemplary embodiment of this application.

[0047] Figure 4 Schematic diagram of a voice receiving device shown in an exemplary embodiment of the present application.

[0048] Figure 5 Schematic diagram of an audio recognition module shown in an exemplary embodiment of the present application.

[0049] Figure 6 Schematic diagram of an audio recognition module shown in an exemplary embodiment of the present application.

[0050] Figure 7 Schematic diagram of a semantic recognition module shown in an exemplary embodiment of the present application.

[0051] Figure 8 Schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. Detailed implementation manners

[0052] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for explaining the present invention, rather than for limiting the protection scope of the present invention.

[0053] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0054] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0055] First of all, it should be noted that vehicle-mounted voice control is based on voice recognition technology and can convert the input of the user's voice into the actions of the in-vehicle execution components. Figure 1 Schematic diagram of vehicle-mounted voice control shown in an exemplary embodiment of the present application. As Figure 1As shown, the user issues a voice command. The recognition device 110 receives the user's voice, recognizes the voice content, and converts the voice content into data that can be received by the computer 120. After analyzing the recognized data, the computer 120 generates computer instructions to control the execution component 130 to perform corresponding operations. For example, when the user issues a command to turn on the air conditioner, after the recognition device 110 recognizes the voice content, it converts the voice into text or a code sequence for the computer. The computer 120 then recognizes the code sequence, determines that the operation to be performed is to turn on the air conditioner, generates computer instructions, and transmits the computer instructions to the execution component 130, such as an air conditioner unit. Thus, the in-vehicle air conditioner can be controlled by voice in the vehicle.

[0056] As Figure 1 shown, after the user issues a voice command, the voice first undergoes information recognition and conversion by the recognition device 110, then is processed by the computer 120 to determine whether operation instructions can be generated, and finally the corresponding operation is performed by the execution component 130. The process reaction chain for the recognition, conversion, and instruction generation to be performed is relatively long. Moreover, during the crucial recognition process, multiple operations are required to achieve voice recognition, and for long and complex sentences or sentences with an accent, the recognition becomes even more complex. Therefore, the elongation of the entire control chain causes a relatively obvious delay in the response of the in-vehicle voice control device. For users, the accuracy and timeliness of voice control are directly related to the user experience.

[0057] Figure 2 is a schematic diagram of a voice control device shown in an exemplary embodiment of the present application. As Figure 2 shown, the voice control device 200 includes a voice receiving device 210, a voice recognition device 220, a service adaptation device 230, and an execution device 240. Among them, the voice receiving device 210 can receive the user's voice data. Specifically, the voice receiving device 210 can be a recorder. In the present invention, a recorder refers to an audio storage medium that can be used to store voice content, such as a magnetic tape or other audio storage medium made of hard magnetic materials. The voice receiving device 210 can record and convert the voice into a sound control signal. The voice recognition device 220 can recognize the sound control signal and convert the sound control signal into a digital signal. The service adaptation device 230 can recognize the digital signal and adapt the corresponding execution device 240 according to the digital signal. Under the coordination and control of the service adaptation device 230, the corresponding execution device 240 performs the operation corresponding to the voice.

[0058] As Figure 2As shown, in an embodiment of the present invention, the speech recognition device 220 includes an audio recognition module 221, a semantic recognition module 222, and a semantic arbitration module 223. Among them, the audio recognition module 221 is electrically connected to the speech receiving device 210 to receive the speech data of the speech receiving device 210. Specifically, the speech receiving device 210 receives a sound control signal, and the audio recognition module 221 can convert the sound control signal into a digital signal. The audio recognition module 221 can convert the speech data into, for example, text data. The speech recognition module 222 is electrically connected to the audio recognition module 221 and can receive the text data transmitted by the audio recognition module 221. The semantic recognition module 222 performs semantic recognition on the text data to determine whether the user's speech is a command. If it is a command message, it is executed; if it is not a command message, it is not executed. Among them, the semantic recognition module 222 can recognize the semantics of the text data in one or more ways, and different recognition models can recognize different key semantic data. For example, by analyzing the text data through keywords, the semantic data can include the characters, actions, and execution objects of the text, etc. For example, by analyzing the similarity of the model of the text data, the semantic data can include the execution intention and the corresponding execution result, etc. Different semantic recognition algorithms can obtain different semantic data. In this embodiment, through the semantic recognition module 222, for example, 1 offline semantic data can be obtained.

[0059] As Figure 2 shown, in an embodiment of the present invention, the speech recognition device 220 includes an audio recognition module 221, a speech cloud module 224, and a semantic arbitration module 223. Among them, the speech cloud module 224 is a natural language processing engine and data aggregation platform that allows users to enjoy semantic recognition technology and provides a large amount of knowledge data for users for semantic recognition. The speech cloud module 224 is communicatively connected to the audio recognition module 221. After the speech data is converted into text data by the audio recognition module 221, the speech cloud module 224 receives the text data and converts the text data into online semantic data. Multiple semantic data are received by the semantic arbitration module 223, and the best semantic data is selected as the semantic information. Therefore, through the semantic arbitration module 223, the online semantic data and the offline semantic data are arbitrated, and the best semantic data is selected as the optimal semantic information.

[0060] As Figure 2As shown, in an embodiment of the present invention, the voice cloud module 224 is an online recognition technology, and the semantic recognition module 222 is an offline recognition technology. The present invention is not limited to implementing online recognition technology or offline recognition technology. When connected to the network online, after receiving the text data processed by the audio recognition module 221, the voice cloud module 224 and the semantic recognition module 222 can synchronously perform semantic recognition to obtain semantic data, and then the semantic arbitration module 223 filters out the optimal semantic data. When in the network online state, it is also possible to obtain multiple semantic data only by recognizing the text data through the voice cloud module 224. And when connected to the network online, it is also possible to train the semantic recognition module 222 with the semantic data of the voice cloud module 224.

[0061] As Figure 2 shown, in an embodiment of the present invention, the voice recognition device 220 is electrically connected to the service adaptation device 230. Specifically, the semantic arbitration module 223 is electrically connected to the service adaptation device 230. Among them, the service adaptation device 230 selects the appropriate service content according to the semantic information and performs protocol parsing to drive the corresponding execution device 240 to work.

[0062] As Figure 2 shown, in an embodiment of the present invention, the voice recognition device 220 includes a custom intent corpus 225 and an intent judgment module 226. Among them, the output end of the audio recognition module 221 is electrically connected to the intent pre-judgment module 226. After receiving the text data, the intent pre-judgment module 226 compares the text data with the custom intent corpus, and can pre-judge the execution object and execution intent of the text data and form pre-execution information. The output end of the intent pre-judgment module 226 is electrically connected to the execution device 240. The execution device 240 receives the pre-execution information and pre-starts the corresponding execution device 240.

[0063] Based on the above voice control device 200, the present invention also provides a vehicle-mounted voice control method, which can quickly and accurately recognize the voice content of the user to improve the user experience of voice control. Figure 3 is a flowchart of the voice control method shown in an exemplary embodiment of the present application. As Figure 3 shown, the steps of the vehicle-mounted voice control method of the present invention include step S310.

[0064] Step S310: Receive the recording data through the voice receiving device, and extract the valid content of the recording data to obtain voice data.

[0065] As Figure 3As shown, in an embodiment of the present invention, in step S310, after the user wakes up the voice control system, the voice receiving device 210 starts to record the user's voice content to obtain recording data. Among them, the voice receiving device 210 performs pulse code modulation (PCM) on the voice data, and performs noise reduction and echo cancellation processing on the voice data to remove the interference content in the voice data, and obtains voice data that can be recognized by the audio recognition module 221.

[0066] Figure 4 Schematic diagram of the voice receiving device shown in an exemplary embodiment of the present application. As Figure 4 shown, in an embodiment of the present invention, the voice receiving device 210 includes a recording unit 401 and a data processing unit 402. In step S310, the recording unit 401 records the user's voice data to obtain recording data and sends the recording data to the data processing unit 402. The data processing unit 402 can remove the background noise and echo of the recording data, and perform modulation processing on the recording data to obtain voice data.

[0067] As Figure 3 shown, the steps of the vehicle-mounted voice control method of the present invention include step S320.

[0068] Step S320: Determine whether the voice data is valid through the audio recognition module. If the voice data is valid, obtain the instruction segment data in the voice data and convert the instruction segment data into text data.

[0069] As Figure 3 shown, in an embodiment of the present invention, in an in-vehicle environment for example, the user's voice data and environmental data are complex. Therefore, setting a wake-up word can determine whether the user issues an instruction. The wake-up word can be pre-set by the operator and stored in the audio recognition module 221, or the user can be given the permission to modify it, and the user can set the wake-up word by himself.

[0070] Figure 5 Schematic diagram of the audio recognition module shown in an exemplary embodiment of the present application. As Figure 3 and Figure 5As shown, in an embodiment of the present invention, the audio recognition module 221 includes a wake word recognition unit 501, an instruction segment extraction unit 502, and an audio conversion unit 503. In step S320, the wake word recognition unit 501 can perform acoustic feature recognition on the audio band of the wake word. Specifically, the wake word recognition unit 501 can receive multiple segments of voice data and number and sort the multiple segments of voice data according to the reception time. In step S320, in accordance with the sorting of the multiple segments of voice data, the wake word recognition unit 501 sequentially determines whether there is a wake word in each segment of voice data. If there is no wake word in the current segment of voice data, the current segment of voice data is discarded, and the wake word recognition unit 501 obtains the next segment of voice data. Among them, discarding the voice data can be to permanently delete the voice data to protect user privacy. If there is a wake word in the current segment of voice data, the wake word recognition unit 501 transmits the voice data to the instruction segment extraction unit 502. Among them, the voice data transmitted by the wake word recognition unit 501 includes the voice data of the segment where the wake word is located and each segment of voice data after the segment where the wake word is located.

[0071] As Figure 3 and Figure 5 shown, in an embodiment of the present invention, in step S320, the voice data intercepted and processed by the wake word recognition unit 501 includes an instruction segment and an interference segment. Among them, the instruction segment is the voice data after the wake word. Therefore, in this embodiment, the instruction segment extraction unit 502 intercepts the voice data. Specifically, taking the wake word as a node, the voice data in the current segment that is in the time sequence after the wake word is retained to obtain the instruction segment data. In step S320, the instruction segment data is processed by the audio conversion unit 503 to convert the instruction segment data into text data. Among them, the instruction segment data is PCM-encoded data obtained by transcoding an audio-type file, and the audio conversion unit 503 can parse the PCM-encoded data and convert the corresponding instruction segment data into text data. When converting the text data, the voice is converted word by word to ensure the integrity of the voice data, so as to improve the accuracy of semantic analysis. Among them, the audio conversion unit 503 can also correct errors in the instruction segment data to improve the accuracy of the text data output. For example, recognition errors may occur during the conversion of dialects. By comparing with the word library, the misrecognized phrases in the text data can be corrected.

[0072] As Figure 3 shown, the steps of the in-vehicle voice control method of the present invention include step S330.

[0073] Step S330: According to the text data, the intention pre-judgment module obtains the execution object pre-called by the user and pre-starts the corresponding execution device.

[0074] As Figure 2 andFigure 3 As shown, in an embodiment of the present invention, the audio recognition module 221 transmits text data to the intention pre-judgment module 226. In step S330, the intention pre-judgment module 226 obtains the execution object in the text data, and pre-starts the execution device before determining the specific execution content. Taking the start of the air conditioner in a vehicle environment as an example, if the user issues an instruction "Xiaobai, turn the air conditioner to gear 5". Among them, Xiaobai can be a wake-up word, and turning the air conditioner to gear 5 is the instruction segment. Through the intention pre-judgment module 226, it is pre-known that the execution object is the air conditioner. And the execution content is to turn on and set it to gear 5, which can only be known after subsequent semantic analysis. And in step S330, when it is already known that the execution object is the air conditioner, the air conditioner is pre-started to give the user a better response experience. During the process of starting the air conditioner, semantic analysis and protocol parsing are synchronized to determine the specific operation content of the air conditioner.

[0075] Figure 6 Schematic diagram of the audio recognition module shown in an exemplary embodiment of the present application. As Figure 2 、 Figure 3 and Figure 6 As shown, in an embodiment of the present invention, in step S330, the intention pre-judgment module 226 includes a long sentence splitting unit 601, an intention object matching unit 602, and a pre-start unit 603. Among them, the long sentence splitting unit 601 can split the text data into segments of multiple parts of speech according to the part of speech, including the subject, predicate, attributive, object, adverbial, and so on. Among them, the segments of parts of speech that do not target the execution object, such as attributives, adverbials, and predicates, are filtered by the intention object matching unit 602. Then, the intention object matching unit 602 analyzes the execution object corresponding to the subject and object. Among them, the intention object matching unit 602 matches the subject and object with the existing words, phrases, or short sentences in the custom intention corpus 225 to obtain the execution object in the text data. In step S330, after confirming the execution object, the intention object matching unit 602 obtains the execution device 240 corresponding to the execution object, and starts the corresponding execution device 240 through the pre-start unit 603. Among them, the long sentence splitting unit 601 can also split the text data word by word. In this embodiment, while the long sentence splitting unit 601 splits the text data, the intention object matching unit 602 matches the split words or characters until the corresponding information is matched in the custom intention corpus 225, and the execution object is obtained.

[0076] As Figure 2 、 Figure 3 and Figure 6 ​As shown, in an embodiment of the present invention, when the execution device 240 is started by the pre-start unit 603, the execution objects include in-vehicle hardware and software installed in the vehicle, such as the air conditioner in the vehicle environment, the lights of various parts of the vehicle body, the windows and seats, etc. Also, for example, the navigation information on the central control screen, the in-vehicle radio, and application software such as music and video. The present invention does not limit this. Among them, the pre-start does not include specific operation content. Among them, the pre-start method is pre-set by developers and can be to switch the corresponding execution device 240, such as turning on the air conditioner and lights, turning off the lights, etc. Among them, for the pre-start items corresponding to switching the execution device 240, when matching the intent object, corresponding entries for turning on and off the execution device 240 can be set in the custom intent corpus 225. For example, turning on the air conditioner, turning off the lights, and turning on the turn signal, etc. Among them, the pre-start method can also be to open and close the software interface, etc. For example, opening the music software interface, closing the navigation interface, etc.

[0077] As Figure 3 shown, the steps of the in-vehicle voice control method of the present invention include step S340.

[0078] Step S340: Match multiple semantic data corresponding to the text data through the semantic recognition module and / or the voice cloud module, and obtain the optimal semantic information among the multiple semantic data through the semantic arbitration module.

[0079] As Figure 2 and Figure 3 shown, in an embodiment of the present invention, in step S340, the audio recognition module 221 transmits the text data to the semantic recognition module 222 and the voice cloud module 224. In this embodiment, in the in-vehicle online environment, semantic data can be obtained through the voice cloud module 224 and the semantic recognition module 222, and the semantic data of both paths are sent to the semantic arbitration module 223. In the offline environment, semantic data can also be obtained through the semantic recognition module 222. Among them, Figure 7 is a schematic diagram of the semantic recognition module shown in an exemplary embodiment of the present application. As Figure 7 shown, the semantic recognition module 222 includes an execution instruction confirmation unit 701 and a semantic analysis unit 702.

[0080] As Figure 2 and Figure 3 and Figure 6 and Figure 7As shown, in an embodiment of the present invention, in step S340, before obtaining semantic data, it is confirmed whether there is an execution instruction in the text data through the execution instruction confirmation unit 701 to prevent the mis-start of the execution device 240 caused by the user's conversation content. If there is no execution instruction in the text data, for example, the user is just chatting with the intelligent software and does not issue an instruction, the work of the pre-start unit 603 is stopped, and the user is only communicated with through the chat word library in the voice. Among them, in another embodiment of the present invention, if no instruction is issued, it can also be defaulted that the execution device is a speaker, and the control system retrieves the content of the word library to communicate with the user. If there is an execution instruction in the text data, the text data is continuously semantically analyzed.

[0081] As Figure 2 and Figure 3 and Figure 6 and Figure 7 As shown, in an embodiment of the present invention, it can be confirmed whether there is an execution instruction by whether there is a start word of the execution device in the text data. Among them, the start word can be preset in advance. Taking the air conditioner as an example, the start words of the execution device 240 of the air conditioner can be words such as air conditioner, fan, temperature, heat, and cold. Taking the light as an example, the start words of the execution device 240 of the car light can be words such as bright, dark, and light. If the text data contains the preset start word, the text data is transmitted to the semantic analysis unit 702. The semantic analysis unit 702 obtains the semantic data corresponding to the text data through the semantic analysis algorithm model. Among them, the semantic analysis algorithm module includes models such as Probabilistic Latent Semantic Analysis (PLSA), Non-negative Matrix Factorization (NMF), and Linear Discriminant Analysis (LDA). The present invention is not limited to the above algorithms, and the present invention includes one or more semantic analysis algorithm models.

[0082] As Figure 2 and Figure 3 and Figure 6 and Figure 7As shown, in an embodiment of the present invention, online semantic data and / or offline semantic data can be obtained through the voice cloud module 224 and / or the semantic recognition module 222. Therefore, the optimal semantic information is screened out from the online semantic data and the offline semantic data by the semantic arbitration module 223. In this embodiment, the semantic arbitration module 223 can screen out semantic data with ambiguous content through ambiguity analysis, such as turning on / off the air conditioner. For multiple semantic data that pass the ambiguity check, semantic priorities can be preset in the semantic arbitration module 223. Among them, the method of setting priorities can be to set the priorities of multiple semantic data according to the integrity of the parts of speech of the semantic data, or according to the matching degree between the execution object and the execution operation, etc. The arbitration strategy can be customized based on the product strategy, and the present invention does not limit this. Among them, the integrity of the parts of speech of the semantic data refers to the integrity of the subject, predicate, attributive, object, and adverbial. Semantic data with richer parts of speech can be set with higher priorities. The matching degree between the execution object and the execution operation can be the rationality of the execution operation in the semantics. For example, the operations corresponding to the air conditioner can be turning on, turning off, adjusting the gear to, for example, gears 1-5, etc. If the operation corresponding to the air conditioner in the semantic data is putting down the air conditioner or adjusting the gear to, for example, gear 10, then the execution content and the execution object do not correspond, and the priority of the corresponding semantic data is reduced. After sorting the semantic data according to the priorities, the semantic arbitration module 223 can output the semantic data ranked first as the optimal semantic information.

[0083] As Figure 3 shown, the steps of the in-vehicle voice control method of the present invention include step S350.

[0084] Step S350: According to the optimal semantic information, the service adaptation device 230 calls the execution device corresponding to the optimal semantic information.

[0085] As Figure 2 、 Figure 3 and Figure 6 shown, in an embodiment of the present invention, in step S350, the service adaptation module 230 can obtain the execution object and the execution instruction corresponding to the optimal semantic information from the optimal semantic information, and thus send the execution instruction to the corresponding execution device 240. On the basis that the pre-start unit 603 has been pre-started, the corresponding execution device 240 completes the execution content in the optimal semantic information. Among them, if the intention pre-judgment unit 226 fails to obtain an accurate execution object, the pre-start unit 603 may not work. After the execution device 240 completes the execution content, the start word can be set for the execution object according to the optimal semantic information, and the new start word is updated to the custom intention corpus 225.

[0086] It should be noted that the voice control method provided in the above embodiments and the voice control system provided in the above embodiments belong to the same concept. The specific ways of each module and the operations performed by the module have been described in detail in the method embodiments, and will not be repeated here. In actual applications, the voice control system provided in the above embodiments can allocate the above functions to different functional modules as needed, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This is not limited here either.

[0087] An embodiment of the present application further provides an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the voice control method provided in each of the above embodiments.

[0088] Figure 8 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. It should be noted that Figure 8 The computer system 800 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0089] As Figure 8 shown, the computer system 800 includes a central processing module (Central Processing Unit, CPU) 801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (Read-Only Memory, ROM) 802 or the program loaded from the storage part 808 into the random access memory (Random Access Memory, RAM) 803, such as executing the method described in the above embodiments. In the RAM 803, various programs and data required for system operation are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other through a bus 804. The input / output (Input / Output, I / O) interface 805 is also connected to the bus 804.

[0090] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as required. A removable medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 810 as required so that a computer program read therefrom is installed into the storage section 808 as required.

[0091] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, the computer program including a computer program for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by a central processing module (CPU) 801, various functions defined in the system of the present application are executed.

[0092] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. A computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0094] The modules involved in the embodiments described in this application can be implemented in software or in hardware, and the described modules can also be provided in a processor. Among them, the names of these modules do not constitute a limitation to the modules themselves in some cases.

[0095] Another aspect of this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of the computer, the computer is caused to execute the voice control method as described above. The computer-readable storage medium can be included in the electronic device described in the above embodiments, or can exist alone without being assembled into the electronic device.

[0096] Another aspect of this application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the voice control methods provided in the above various embodiments.

[0097] The above embodiments are only used to exemplarily illustrate the principles and effects of the present invention, rather than to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A vehicle-mounted voice control system, characterized in that, Comprising: A voice receiving device, configured to obtain recording data, extract the valid content of the recording data, and obtain voice data; A voice recognition device, connected to the voice receiving device, the voice recognition device receives and processes the voice data, obtains the execution object and the optimal semantic information of the voice data, and the voice recognition device generates a pre-start instruction according to the execution object; A service adaptation device, connected to the voice recognition device, the service adaptation device receives the optimal semantic information, matches the service content corresponding to the optimal semantic information, and generates an execution instruction; And An execution device, connected to the service adaptation device and the voice recognition device, when receiving the pre-start instruction, the execution device starts, and when receiving the execution instruction, the execution device executes the operation corresponding to the optimal semantic information; Wherein the voice recognition device includes: An audio recognition module, connected to the voice receiving device, the audio recognition module receives the voice data and converts the voice data into text data; A semantic recognition module, connected to the audio recognition module, the semantic recognition module receives the text data and offline-parses the semantics of the text data to obtain offline semantic data; An intention pre-judgment module, connected to the audio recognition module and the execution device, the intention pre-judgment module receives the text data and obtains the execution object of the text data, and the intention pre-judgment module generates the pre-start instruction according to the execution object; and A voice cloud module, connected to the audio recognition module, the voice cloud module online-parses the semantics of the text data to obtain online semantic data; Wherein the audio recognition module includes: A wake-up word recognition unit, connected to the voice receiving device, the wake-up word recognition unit receives the voice data and recognizes the wake-up word in the voice data according to audio features; An instruction segment extraction unit, connected to the wake-up word recognition unit, when the wake-up word is included in the voice data, intercepts the voice data after the wake-up word in the recording order to obtain instruction segment data; and An audio conversion unit, connected to the instruction segment extraction unit, the audio conversion unit receives the instruction segment data and converts the instruction segment data into text data.

2. The vehicle-mounted voice control system according to claim 1, characterized in that, The voice recognition device includes a semantic arbitration module, the semantic arbitration module is connected to the semantic recognition module and the voice cloud module, the semantic arbitration module receives the online semantic data and the offline semantic data, and filters out the optimal semantic information from the online semantic data and the offline semantic data.

3. The vehicle-mounted voice control system according to claim 1, characterized in that, The semantic recognition module includes: An execution instruction confirmation unit, which receives the text data and determines whether the text data contains an execution instruction; and A semantic analysis unit, connected to the execution instruction confirmation unit, when the instruction is included in the text data, the semantic analysis unit receives the text data and parses the semantic information of the text data to obtain the offline semantic data.

4. The vehicle-mounted voice control system according to claim 1, characterized in that, The voice recognition device includes a custom intent corpus, and the custom intent corpus stores a plurality of activation words.

5. The vehicle-mounted voice control system according to claim 4, characterized in that, The intent pre-judgment module includes: A long sentence splitting unit that receives the text data and splits the text data into multiple segments word by word or according to phrases; An intent object matching unit connected to the long sentence splitting unit and the custom intent corpus. While the long sentence splitting unit splits the text data, the intent object matching unit receives the segmented text data and compares the text data with the words in the custom intent corpus until the execution object of the text data is obtained; and A pre-start unit connected to the intent object matching unit. The pre-start unit matches the execution device corresponding to the execution object and generates the pre-start instruction.

6. The vehicle-mounted voice control system according to claim 1, characterized in that, The voice receiving device includes: A recording unit for recording the user's voice to obtain recording data; and A data processing unit that receives the recording data and removes the interference frequency band of the recording data to obtain voice data.

7. A vehicle-mounted voice control method, characterized in that, It includes the following steps: Obtain recording data through the voice receiving device, and extract the effective content of the recording data to obtain voice data; Receive and process the voice data through the voice recognition device, obtain the execution object and the optimal semantic information of the voice data, and generate a pre-start instruction through the voice recognition device according to the execution object; Receive the optimal semantic information through the service adaptation device, match the service content corresponding to the optimal semantic information, and generate an execution instruction; And Receive the pre-start instruction and the execution instruction through the execution device. When the execution device receives the pre-start instruction, the execution device starts. When the execution device receives the execution instruction, the execution device executes the operation corresponding to the optimal semantic information; Among them, the processing steps of the voice recognition device include: Receive the voice data through the audio recognition module and convert the voice data into text data; Receive the text data through the semantic recognition module and offline parse the semantics of the text data to obtain offline semantic data; Receive the text data through the intent pre-judgment module and obtain the execution object of the text data, and the intent pre-judgment module generates the pre-start instruction according to the execution object; and Online parse the semantics of the text data through the voice cloud module to obtain online semantic data; Among them, the processing steps of the audio recognition module include: Receive the voice data through the wake-up word recognition unit and recognize the wake-up word in the voice data according to the audio characteristics; When the wake-up word is included in the voice data, through the instruction segment extraction unit, intercept the voice data after the wake-up word in the recording order to obtain instruction segment data; and Receive the instruction segment data through the audio conversion unit and convert the instruction segment data into text data.

8. A vehicle-mounted voice control method according to claim 7, wherein, The steps of obtaining the pre-start instruction include: Convert the voice data into text data through the audio recognition module; Receive the text data through the long sentence splitting unit and split the text data into multiple segments word by word or according to phrases; While the long sentence splitting unit splits the text data, the intention object matching unit receives the split text data and compares each piece of the text data with the words in the custom intention corpus until the execution object of the text data is obtained; and The pre-start unit matches the execution device corresponding to the execution object and generates the pre-start instruction.

Citation Information

Patent Citations

  • A control method, system, device and apparatus of an electrical apparatus and a medium

    CN109192208A

  • Semantic parsing with multiple parsers

    US9026431B1