Vehicle-mounted scene hot word identification method and related equipment

By constructing a hot word language model and a general language model in a car-mounted scenario, combining pronunciation and text features, the problem of low accuracy in hot word recognition is solved, and a higher recognition accuracy and stronger detection ability of hot words is achieved.

CN119993135AActive Publication Date: 2025-05-13VOYAH AUTOMOBILE TECH CO LTD

Patent Information

Application Number
CN202510028193.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-13
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

In car-based scenarios, the recognition of hot word words such as person names, place names and naming entities with low word frequency or strongly related to users cannot achieve the accuracy of general recognition.

Method used

By obtaining the voice data entered by the user, the hot word text is extracted and the fusion characteristics of the voice data and the hot word text are obtained. Then, a hot word language model is constructed based on multiple mapping paths, the fusion features are decoded, and the hot word decoding results are generated. At the same time, the same features are decoded using a common language model to generate a general decoding result. Finally, the target recognition results are determined based on hot words and general decoding results.

Benefits of technology

It improves the recognition accuracy of hot words in car-mounted scenarios, reduces interference with general recognition, and enhances the overall detection ability of hot words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993135A_ABST
    Figure CN119993135A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle-mounted scene hot word recognition method and related equipment, and relates to the technical field of Internet, and the method comprises the steps: obtaining voice data inputted by a user, extracting a hot word text from the voice data, and obtaining fusion features of the voice data and the hot word text; analyzing the hot word text to generate a plurality of mapping paths, and constructing a hot word language model based on the plurality of mapping paths; decoding the fusion feature based on the hot word language model to obtain a hot word decoding result; decoding the fusion feature based on the universal language model to obtain a universal decoding result; and determining a target recognition result based on the hot word decoding result and the general decoding result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a method for identifying hot words in vehicle scenes and related equipment. Background Art

[0002] At present, with the development of speech recognition technology, the recognition effect in general fields has been significantly improved. Under ideal conditions, the recognition accuracy can usually reach 95% or even higher. At the same time, speech recognition technology has been widely used in the automotive field to realize functions including in-vehicle infotainment systems, navigation, voice assistants, etc.

[0003] However, the recognition of hot words with low frequency or strong user relevance, such as names of people, places, and named entities, cannot achieve the same accuracy as general recognition. Therefore, it is necessary to propose a method for identifying hot words in vehicle scenarios to at least solve some of the above problems. Summary of the invention

[0004] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further described in detail in the Detailed Description of the Invention section. The Summary of the Invention section of this application does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the scope of protection of the claimed technical solution.

[0005] In a first aspect, an embodiment of the present application provides a method for identifying hot words in a vehicle scene, the method comprising:

[0006] Acquire voice data input by a user, extract hot word text from the voice data, and acquire fusion features of the voice data and the hot word text;

[0007] Parsing the hot word text to generate multiple mapping paths, and building a hot word language model based on the multiple mapping paths;

[0008] Decoding the fused features based on the hot word language model to obtain a hot word decoding result;

[0009] Decoding the fused features based on the universal language model to obtain a universal decoding result;

[0010] A target recognition result is determined based on the hot word decoding result and the general decoding result.

[0011] In one embodiment of the present invention, the step of obtaining the fusion features of the voice data and the hot word text includes:

[0012] Encoding the speech data to obtain a speech embedding vector;

[0013] Encode the hot word text to obtain a text embedding vector;

[0014] Calculate the correlation between the speech embedding vector and the text embedding vector to obtain a correlation weight;

[0015] The speech embedding vector and the text embedding vector are fused based on the correlation weight to obtain a fusion feature.

[0016] In one embodiment of the present invention, the step of fusing the speech embedding vector and the text embedding vector based on the correlation weight to obtain a fusion feature includes:

[0017] In the absence of the hot word text, the speech embedding vector is used as the fusion feature.

[0018] In one embodiment of the present invention, the step of constructing a hot word language model based on the multiple mapping paths includes:

[0019] Parsing the hot word text to obtain multiple hot word mapping relationships;

[0020] Constructing multiple hot word mapping paths based on multiple hot word mapping relationships;

[0021] Connecting multiple hot word mapping paths in parallel to obtain an initial hot word language model;

[0022] Obtaining sentence text associated with the hot word text;

[0023] Parsing the sentence text to obtain multiple sentence mapping relationships;

[0024] Constructing multiple sentence mapping paths based on multiple sentence mapping relationships;

[0025] Connecting a plurality of the sentence mapping paths in parallel to obtain a sentence language model;

[0026] The sentence language model and the initial hot word language model are connected end to end to form a hot word language model.

[0027] In one embodiment of the present invention, the step of determining the target recognition result based on the hot word decoding result and the general decoding result includes:

[0028] Calculating the hot word decoding result to obtain a hot word path score and a hot word pronunciation sequence distance, wherein the hot word path score includes a hot word acoustic score and a hot word language score;

[0029] Calculating the universal decoding result to obtain a universal path score and a universal pronunciation sequence distance, wherein the universal path score includes a universal acoustic score and a universal language score;

[0030] When the difference between the hot word acoustic score and the universal acoustic score is greater than a first preset threshold, and the difference between the hot word language score and the universal language score is greater than a second preset threshold, determining the similarity between the hot word pronunciation sequence distance and the universal pronunciation sequence distance;

[0031] If the similarity is greater than or equal to a third preset threshold, the hot word decoding result is used as the recognition result.

[0032] In one embodiment of the present invention, the steps after determining the similarity between the hot word pronunciation sequence distance and the common pronunciation sequence distance include:

[0033] In the case where the similarity is less than the third preset threshold, determining whether the difference between the hot word acoustic score and the universal acoustic score is greater than the first preset threshold, and whether the difference between the hot word language score and the universal language score is greater than the second preset threshold;

[0034] When the difference between the hot word acoustic score and the universal acoustic score is greater than a first preset threshold, and the difference between the hot word language score and the universal language score is greater than a second preset threshold, taking the hot word decoding result as the recognition result;

[0035] When the difference between the hot word acoustic score and the universal acoustic score is less than a first preset threshold, or the difference between the hot word language score and the universal language score is less than a second preset threshold, the universal decoding result is used as the recognition result.

[0036] In one embodiment of the present invention, the step of determining the target recognition result based on the hot word decoding result and the general decoding result further includes:

[0037] In the absence of the hot word text, the universal decoding result is used as the recognition result.

[0038] In a second aspect, the present application proposes a vehicle scene hot word recognition system, the system comprising: a data acquisition module, a parsing module and a recognition module;

[0039] The data acquisition module is configured to: acquire voice data input by a user, extract hot word text from the voice data, and acquire fusion features of the voice data and the hot word text;

[0040] The parsing module is configured to: parse the hot word text, generate multiple mapping paths, and construct a hot word language model based on the multiple mapping paths;

[0041] The recognition module is configured to: decode the fused features based on the hot word language model to obtain a hot word decoding result; decode the fused features based on the universal language model to obtain a universal decoding result; and determine a target recognition result based on the hot word decoding result and the universal decoding result.

[0042] In a third aspect, an electronic device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the steps of a method for recognizing hot words in a vehicle scene as described in any one of the first aspects above when executing the computer program stored in the memory.

[0043] In a fourth aspect, the present application further proposes a computer-readable storage medium having a computer program stored thereon, and when the above-mentioned computer program is executed by a processor, the steps of a method for identifying hot words in a vehicle scene according to any one of the first aspects are implemented.

[0044] In summary, a method for identifying hot words in vehicle scenarios in an embodiment of the present application generates fixed context dependencies through a hot word language model constructed based on a mapping path, improves the overall detection of hot words, and determines the target recognition results based on the hot word decoding results and the general decoding results, thereby reducing interference with general recognition and improving the accuracy of hot word recognition.

[0045] The method for identifying hot words in vehicle scenes proposed in this application, and other advantages, objectives and features of this application will be reflected in part through the following description, and in part will also be understood by technical personnel in this field through research and practice of this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present specification. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0047] Figure 1 A schematic diagram of the process of identifying hot words in a vehicle scene provided in an embodiment of the present application;

[0048] Figure 2 A schematic diagram of the structure of a vehicle scene hot word recognition system provided in an embodiment of the present application;

[0049] Figure 3 A schematic diagram of the structure of an electronic device for identifying and controlling hot words in a vehicle scene provided in an embodiment of the present application. DETAILED DESCRIPTION

[0050] In order to better understand the technical solutions provided by the embodiments of this specification, the technical solutions of the embodiments of this specification are described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of this specification and the specific features in the embodiments are detailed descriptions of the technical solutions of the embodiments of this specification, rather than limitations on the technical solutions of this specification. In the absence of conflict, the embodiments of this specification and the technical features in the embodiments can be combined with each other.

[0051] In this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of more restrictions, the elements limited by the statement "comprise one..." do not exclude the existence of other identical elements in the process, method, article or equipment including the elements. The term "more than two" includes two or more than two situations.

[0052] See also Figure 1 , which is a schematic diagram of a process of identifying hot words in a vehicle scene provided in an embodiment of the present application, which may specifically include:

[0053] S110, obtaining voice data input by a user, extracting hot word text from the voice data, and obtaining fusion features of the voice data and the hot word text.

[0054] For example, when a user speaks, the user's voice data is collected using the microphone of the device, and the voice data is analyzed and processed, specifically: the voice is converted into text using voice recognition technology, and then hot words are identified from the converted text using a natural language processing algorithm (such as a keyword extraction algorithm), and the extracted hot words are combined into a hot word text. Then, the fusion features of the voice data and the hot word text are obtained.

[0055] S120: Parse the hot word text to generate multiple mapping paths, and build a hot word language model based on the multiple mapping paths.

[0056] Exemplarily, after obtaining the hot word text, it is necessary to parse it and generate multiple mapping paths, where the mapping path can be understood as the connection method from the hot word text to different concepts or entities. For the hot word text "theme park", the mapping paths that may be generated are: pointing to the name of a specific well-known theme park (such as Disney theme park), pointing to the type of theme park (such as water theme park, adventure theme park, etc.), and pointing to activities related to theme parks (such as roller coasters, carousels, etc.). These mapping paths can help the system understand the meaning of the hot word text more comprehensively. Use the generated multiple mapping paths to build a hot word language model.

[0057] S130: Decode the fused features based on the hot word language model to obtain a hot word decoding result.

[0058] Exemplarily, the fusion feature is decoded by the hot word language model to obtain the hot word decoding result. The decoding process is to use the hot word language model to analyze and understand the fusion feature. For example, if the hot word text corresponding to the fusion feature is "park", the hot word decoding result may be a specific park name (such as "Central Park"), a description of the park type (such as "large city park"), or a recommendation of activities related to the park (such as "a park suitable for walking"), etc.

[0059] S140: Decode the fused features based on the universal language model to obtain a universal decoding result.

[0060] Exemplarily, the fused feature is decoded by a universal language model to obtain a universal decoding result. When the fused feature is decoded using a universal language model, the universal language model will try to analyze and interpret the fused feature based on the language knowledge and patterns it has learned. The universal language model will take a broader language perspective, rather than being limited to specific hot words. For example, if the voice data corresponding to the fused feature is "I want to go to that place to play", the universal decoding result is "The user expressed the desire to go to a specific place to play."

[0061] S150, determining a target recognition result based on the hot word decoding result and the general decoding result.

[0062] Exemplarily, the target recognition result is determined based on the hot word decoding result and the general decoding result, wherein the target recognition result is the final judgment obtained after comprehensively considering the hot word decoding result and the general decoding result, so as to obtain a more accurate and targeted result.

[0063] In summary, the method for identifying hot words in vehicle scenarios proposed in the embodiment of the present application generates fixed context dependencies through a hot word language model constructed based on a mapping path, improves the overall detection of hot words, and determines the target recognition results based on the hot word decoding results and the general decoding results, thereby reducing interference with general recognition and improving the accuracy of hot word recognition.

[0064] In some examples, the step of obtaining the fusion features of the voice data and the hot word text includes:

[0065] Encoding the speech data to obtain a speech embedding vector;

[0066] Encode the hot word text to obtain a text embedding vector;

[0067] Calculate the correlation between the speech embedding vector and the text embedding vector to obtain a correlation weight;

[0068] The speech embedding vector and the text embedding vector are fused based on the correlation weight to obtain a fusion feature.

[0069] Exemplarily, the speech data and the hot word text are encoded respectively by the text injection method to obtain the speech embedding vector and the text embedding vector. Text injection is a technical means to introduce text information in the process of encoding speech data. Specifically, when encoding the speech data and the hot word text, the text related to the speech (such as the transcribed text of the speech content, the related hot word text, etc.) will be used as an additional input or feature and input into the encoding model together with the original speech data and the hot word text.

[0070] The encoding process usually uses deep learning models such as convolutional neural networks, recurrent neural networks, or Transformer architectures. These models convert speech data into a fixed-length speech embedding vector by learning the correspondence between a large amount of speech data and related text. This vector can capture various features of speech, including semantics, timbre, intonation, and other information, while also incorporating additional information from text injection. At the same time, the hot word text also needs to be converted into a text embedding vector that can represent its semantic and grammatical features.

[0071] The purpose of calculating the correlation between the speech embedding vector and the text embedding vector is to determine the degree of correlation between the speech data and the hot word text. There are many ways to calculate the correlation, such as calculating the cosine similarity between vectors, the Pearson correlation coefficient, etc. These methods can measure the similarity of the direction of two vectors or the strength of the linear relationship. The obtained correlation weight is a numerical value that reflects the closeness between the speech data and the hot word text. If the correlation is high, the weight value will be large; otherwise, the weight value will be small.

[0072] The speech embedding vector and the text embedding vector are fused based on the relevance weight to obtain the fused feature. The fusion process is to integrate the information of the speech data and the hot word text to obtain a more comprehensive feature representation. The relevance weight determines the relative importance of the speech embedding vector and the text embedding vector in the fusion process. If the relevance weight is large, it means that the correlation between the speech and the hot word text is high. In this case, a larger weight can be given to the text embedding vector during fusion, so that the fused feature is more inclined to the characteristics of the text; conversely, if the relevance weight is small, a larger weight can be given to the speech embedding vector. The fusion method can be a simple weighted summation, or a more complex neural network structure can be used for fusion. The obtained fused feature is equivalent to a bias vector related to a hot word, which enhances the acoustic bias information of the hot word.

[0073] In some examples, the step of fusing the speech embedding vector and the text embedding vector based on the relevance weight to obtain a fused feature includes:

[0074] In the absence of the hot word text, the speech embedding vector is used as the fusion feature.

[0075] Exemplarily, if there is no hot word text and the relevance weight is 0, then the speech embedding vector obtained by encoding the speech data is directly used as a fusion feature. Normally, the speech data and the hot word text are encoded separately, and then the relevance weight is calculated and fused to obtain the fusion feature. However, when there is no hot word text, the fusion operation cannot be performed according to the normal process. At this time, in order to ensure that the system still has a certain feature representation available, the only speech embedding vector is directly used as a fusion feature. In this way, in the absence of hot word text, the encoding results of the speech data can still be used to perform subsequent processing tasks, avoiding the situation where effective features cannot be generated due to the lack of hot word text.

[0076] In some examples, the step of building a hot word language model based on the plurality of mapping paths includes:

[0077] Parsing the hot word text to obtain multiple hot word mapping relationships;

[0078] Constructing multiple hot word mapping paths based on multiple hot word mapping relationships;

[0079] Connecting multiple hot word mapping paths in parallel to obtain an initial hot word language model;

[0080] Obtaining sentence text associated with the hot word text;

[0081] Parsing the sentence text to obtain multiple sentence mapping relationships;

[0082] Constructing multiple sentence mapping paths based on multiple sentence mapping relationships;

[0083] Connecting a plurality of the sentence mapping paths in parallel to obtain a sentence language model;

[0084] The sentence language model and the initial hot word language model are connected end to end to form a hot word language model.

[0085] For example, for speech recognition based on the hybrid Encoder-Decoder structure, the language model can use Decoder, FST (Finite State Transducer), etc. Considering the determinism of the FST decoding path, this solution uses FST as the language model. For the FST in the general field, it is based on the probability statistics of the general corpus; for the hot words in the vehicle field, it is necessary to manually construct an effective path mapping to generate the hot word FST.

[0086] The hot word text is parsed to obtain multiple hot word mapping relationships. The process of parsing the hot word text involves performing semantic analysis and part-of-speech tagging on each hot word to determine the relationship between the hot words. For example, some hot words may have semantically similar, opposite, and inclusive relationships. Finally, multiple hot word mapping relationships are obtained, and these relationships can be expressed in the form of "hot word A is semantically similar to hot word B".

[0087] Use the hot word mapping relationship obtained above to build a path. For example, if hot word A and hot word B are semantically similar, then you can build a mapping path from hot word A to hot word B. Since there are multiple hot word mapping relationships, you can build multiple such mapping paths.

[0088] Put multiple hot word mapping paths together to form a network-like structure, where there is no order between the paths, just like the parallel connection in a circuit. This network structure is the initial hot word language model, which can be used to represent the relationship between hot words and possible combinations.

[0089] At the same time, the sentence texts of the hot word text related context are sorted out, similar to the analysis of the hot word text, and the relationship between different parts of the sentence is analyzed to determine the relationship between the different parts of the sentence to obtain multiple sentence mapping relationships. The sentence mapping relationship is used to build a sentence mapping path similar to the hot word mapping path. Similar to the construction of the initial hot word language model, multiple sentence mapping paths are connected in parallel to form a sentence language model. This sentence language model can represent the relationship and combination of different sentence parts. The sentence language model and the initial hot word language model are connected end to end to form a complete hot word language model with sentences.

[0090] In some examples, the step of determining a target recognition result based on the hot word decoding result and the general decoding result includes:

[0091] Calculating the hot word decoding result to obtain a hot word path score and a hot word pronunciation sequence distance, wherein the hot word path score includes a hot word acoustic score and a hot word language score;

[0092] Calculating the universal decoding result to obtain a universal path score and a universal pronunciation sequence distance, wherein the universal path score includes a universal acoustic score and a universal language score;

[0093] When the difference between the hot word acoustic score and the universal acoustic score is greater than a first preset threshold, and the difference between the hot word language score and the universal language score is greater than a second preset threshold, determining the similarity between the hot word pronunciation sequence distance and the universal pronunciation sequence distance;

[0094] If the similarity is greater than or equal to a third preset threshold, the hot word decoding result is used as the recognition result.

[0095] Exemplarily, the hot word decoding result is calculated to obtain the hot word path score and the hot word pronunciation sequence distance, wherein the hot word path score includes the hot word acoustic score and the hot word language score, and the higher the hot word acoustic score, the better the match between the acoustic features of the hot word and the model. The higher the hot word language score, the more the hot word meets expectations at the language level. The smaller the hot word pronunciation sequence distance, the closer the pronunciation of the hot word is to the standard pronunciation. The universal decoding result is calculated to obtain the universal path score and the universal pronunciation sequence distance, wherein the universal path score includes the universal acoustic score and the universal language score.

[0096] First, determine whether the difference between the hot word acoustic score and the universal acoustic score is greater than the first preset threshold, and whether the difference between the hot word language score and the universal language score is greater than the second preset threshold. If both conditions are met, it means that the hot word is significantly different from the universal data in terms of acoustics and language. When the above conditions are met, further determine the similarity between the hot word pronunciation sequence distance and the universal pronunciation sequence distance. If the similarity between the hot word pronunciation sequence distance and the universal pronunciation sequence distance is greater than or equal to the third preset threshold, it means that although there is a difference between the pronunciation of the hot word and the pronunciation of the universal data, the difference is within an acceptable range. In this case, the hot word decoding result is output as the final recognition result, that is, the hot word decoding result is considered to be a more reliable result.

[0097] In some examples, the steps after determining the similarity between the hot word pronunciation sequence distance and the common pronunciation sequence distance include:

[0098] In the case where the similarity is less than the third preset threshold, determining whether the difference between the hot word acoustic score and the universal acoustic score is greater than the first preset threshold, and whether the difference between the hot word language score and the universal language score is greater than the second preset threshold;

[0099] When the difference between the hot word acoustic score and the universal acoustic score is greater than a first preset threshold, and the difference between the hot word language score and the universal language score is greater than a second preset threshold, taking the hot word decoding result as the recognition result;

[0100] When the difference between the hot word acoustic score and the universal acoustic score is less than a first preset threshold, or the difference between the hot word language score and the universal language score is less than a second preset threshold, the universal decoding result is used as the recognition result.

[0101] Exemplarily, when the similarity between the hot word pronunciation sequence distance and the common pronunciation sequence distance is less than the third preset threshold, it means that the pronunciation difference between the hot word and the common data exceeds the acceptable range. At this time, it is necessary to further judge the difference between the hot word and the common data in acoustics and language. Specifically, it is to judge whether the difference between the hot word acoustic score and the common acoustic score is greater than the first preset threshold, and whether the difference between the hot word language score and the common language score is greater than the second preset threshold.

[0102] If the difference between the hot word acoustic score and the general acoustic score is greater than the first preset threshold, and the difference between the hot word language score and the general language score is greater than the second preset threshold, it means that the hot word has a large difference from the general data in terms of acoustics and language, and this difference is significant. In this case, the hot word decoding result is used as the recognition result. This is because hot words have special acoustic and language characteristics, which are significantly different from general data, so the hot word decoding result is more reliable.

[0103] If the difference between the hot word acoustic score and the general acoustic score is less than the first preset threshold, or the difference between the hot word language score and the general language score is less than the second preset threshold, it means that the difference between the hot word and the general data in acoustics or language is not significant enough. In this case, the general decoding result is used as the recognition result. This is because the difference between the hot word and the general data is not obvious, and the general decoding result is more universal and reliable.

[0104] In some examples, the step of determining the target recognition result based on the hot word decoding result and the general decoding result further includes:

[0105] In the absence of the hot word text, the universal decoding result is used as the recognition result.

[0106] For example, if there is no hot word text at present, the general decoding result is directly output as the final recognition result. If there is no hot word text, there is no need to perform special processing on the hot word, and the more universal general decoding result can be directly used, because there is no special hot word text that needs to be given priority as the recognition result.

[0107] like Figure 2 As shown, the present application proposes a vehicle scene hot word recognition system, the system comprising: a data acquisition module 21, a parsing module 22 and a recognition module 23;

[0108] The data acquisition module 21 is configured to: acquire voice data input by a user, extract hot word text from the voice data, and acquire fusion features of the voice data and the hot word text;

[0109] The parsing module 22 is configured to: parse the hot word text, generate multiple mapping paths, and construct a hot word language model based on the multiple mapping paths;

[0110] The recognition module 23 is configured to: decode the fused features based on the hot word language model to obtain a hot word decoding result; decode the fused features based on the universal language model to obtain a universal decoding result; and determine a target recognition result based on the hot word decoding result and the universal decoding result.

[0111] The effects of the above-mentioned system when applying the above-mentioned method can be found in the description of the above-mentioned method embodiment, which will not be repeated here.

[0112] like Figure 3 As shown, an embodiment of the present application also provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, the steps of any method for recognizing hot words in the above-mentioned vehicle scene are implemented.

[0113] Since the electronic device introduced in this embodiment is a device used to implement a vehicle-mounted scene hot word recognition device in the embodiment of the present application, based on the method introduced in the embodiment of the present application, technical personnel in this field can understand the specific implementation mode of the electronic device of this embodiment and its various variations. Therefore, how the electronic device implements the method in the embodiment of the present application is not introduced in detail here. As long as the equipment adopted by technical personnel in this field to implement the method in the embodiment of the present application is within the scope of protection of this application.

[0114] In the specific implementation process, when the computer program 311 is executed by the processor, it can achieve Figure 1 Any implementation manner in the corresponding embodiments.

[0115] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0116] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-readable program code.

[0117] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0118] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0120] An embodiment of the present application also provides a computer program product, which includes computer software instructions. When the computer software instructions are executed on a processing device, the processing device executes the process of the LDPC decoding method of the solid-state hard disk controller.

[0121] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integration. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)), etc.

[0122] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0123] In the several embodiments provided in the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0124] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0125] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0126] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.

[0127] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

[0128] Although the preferred embodiments of this specification have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of this specification.

[0129] Obviously, those skilled in the art can make various changes and modifications to this specification without departing from the spirit and scope of this specification. Thus, if these modifications and variations of this specification fall within the scope of the claims of this specification and their equivalents, this specification is also intended to include these modifications and variations.

Claims

1. A method for identifying hot words in vehicle scenes, characterized in that: The method comprises: Acquire voice data input by a user, extract hot word text from the voice data, and acquire fusion features of the voice data and the hot word text; Parsing the hot word text to generate multiple mapping paths, and building a hot word language model based on the multiple mapping paths; Decoding the fused features based on the hot word language model to obtain a hot word decoding result; Decoding the fused features based on the universal language model to obtain a universal decoding result; A target recognition result is determined based on the hot word decoding result and the general decoding result.

2. The method for identifying hot words in vehicle scenes according to claim 1, characterized in that: The step of obtaining the fusion features of the voice data and the hot word text includes: Encoding the speech data to obtain a speech embedding vector; Encode the hot word text to obtain a text embedding vector; Calculate the correlation between the speech embedding vector and the text embedding vector to obtain a correlation weight; The speech embedding vector and the text embedding vector are fused based on the correlation weight to obtain a fusion feature.

3. The method for identifying hot words in vehicle scenes according to claim 2, characterized in that: The step of fusing the speech embedding vector and the text embedding vector based on the correlation weight to obtain a fusion feature comprises: In the absence of the hot word text, the speech embedding vector is used as the fusion feature.

4. The method for identifying hot words in vehicle scenes according to claim 1, characterized in that: The step of constructing a hot word language model based on the multiple mapping paths comprises: Parsing the hot word text to obtain multiple hot word mapping relationships; Constructing multiple hot word mapping paths based on multiple hot word mapping relationships; Connecting multiple hot word mapping paths in parallel to obtain an initial hot word language model; Obtaining sentence text associated with the hot word text; Parsing the sentence text to obtain multiple sentence mapping relationships; Constructing multiple sentence mapping paths based on multiple sentence mapping relationships; Connecting a plurality of the sentence mapping paths in parallel to obtain a sentence language model; The sentence language model and the initial hot word language model are connected end to end to form a hot word language model.

5. The method for identifying hot words in vehicle scenes according to claim 1, characterized in that: The step of determining the target recognition result based on the hot word decoding result and the general decoding result comprises: Calculating the hot word decoding result to obtain a hot word path score and a hot word pronunciation sequence distance, wherein the hot word path score includes a hot word acoustic score and a hot word language score; Calculating the universal decoding result to obtain a universal path score and a universal pronunciation sequence distance, wherein the universal path score includes a universal acoustic score and a universal language score; When the difference between the hot word acoustic score and the universal acoustic score is greater than a first preset threshold, and the difference between the hot word language score and the universal language score is greater than a second preset threshold, determining the similarity between the hot word pronunciation sequence distance and the universal pronunciation sequence distance; If the similarity is greater than or equal to a third preset threshold, the hot word decoding result is used as the recognition result.

6. The method for identifying hot words in vehicle scenes according to claim 5, characterized in that: The steps after determining the similarity between the hot word pronunciation sequence distance and the common pronunciation sequence distance include: In the case where the similarity is less than the third preset threshold, determining whether the difference between the hot word acoustic score and the universal acoustic score is greater than the first preset threshold, and whether the difference between the hot word language score and the universal language score is greater than the second preset threshold; When the difference between the hot word acoustic score and the universal acoustic score is greater than a first preset threshold, and the difference between the hot word language score and the universal language score is greater than a second preset threshold, taking the hot word decoding result as the recognition result; When the difference between the hot word acoustic score and the universal acoustic score is less than a first preset threshold, or the difference between the hot word language score and the universal language score is less than a second preset threshold, the universal decoding result is used as the recognition result.

7. The method for identifying hot words in vehicle scenes according to claim 1, characterized in that: The step of determining the target recognition result based on the hot word decoding result and the general decoding result also includes: In the absence of the hot word text, the universal decoding result is used as the recognition result.

8. A vehicle scene hot word recognition system, characterized in that: The system comprises: a data acquisition module, a parsing module and a recognition module; The data acquisition module is configured to: acquire voice data input by a user, extract hot word text from the voice data, and acquire fusion features of the voice data and the hot word text; The parsing module is configured to: parse the hot word text, generate multiple mapping paths, and construct a hot word language model based on the multiple mapping paths; The recognition module is configured to: decode the fused features based on the hot word language model to obtain a hot word decoding result; decode the fused features based on the universal language model to obtain a universal decoding result; and determine a target recognition result based on the hot word decoding result and the universal decoding result.

9. An electronic device, comprising: A memory and a processor, characterized in that the processor is used to implement the steps of a method for identifying hot words in a vehicle scene as described in any one of claims 1-7 when executing a computer program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a method for identifying hot words in a vehicle scene as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Interactive system for vehicle-mounted voice

    CN101281745A

  • Speech recognition method, device and equipment and storage medium

    CN110164416A

  • Voice recognition method and device, electronic equipment and storage medium

    CN111508497A

  • Speech recognition method and device, equipment and storage medium

    CN114360499A

  • Speech recognition method and device based on hot word coding and storage medium

    CN115881104A

Cited By

  • Speech recognition illusion suppression method and system based on double-decoding decision network, and medium

    CN121354569A