Text recognition method and device, equipment and storage medium

By combining the current identification text and historical interaction information, using multiple semantic extraction models for feature extraction and classification, the problem of speech recognition system mistakenly or missed recognition in complex scenarios is solved, and higher recognition accuracy and user experience are achieved.

CN119988629APending Publication Date: 2025-05-13CHONGQING CHANGAN TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510100677.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In complex scenarios, it is difficult for the speech recognition system to accurately understand the user's intentions, especially when the user's expression is concise or noise, resulting in the problem of false recognition or missed recognition.

Method used

By combining the current identification text and historical interaction information, multiple semantic extraction models are used for feature extraction and classification, improving the ability to understand user intentions and reducing the situation of false rejection and missed recognition.

Benefits of technology

It improves the accuracy and efficiency of text recognition, reduces the situation of false rejection and omission of intentional text, and improves the robustness and user experience of the voice interaction system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988629A_ABST
    Figure CN119988629A_ABST
Patent Text Reader

Abstract

The invention relates to a text recognition method and device, equipment and a storage medium, and the method comprises the steps: carrying out the feature extraction of a current recognition text and historical interaction information before the current recognition text based on at least one semantic extraction model, and obtaining a feature vector outputted by each semantic extraction model; wherein the at least one semantic extraction model is determined in a plurality of feature extraction models with different feature extraction types on the basis of the contribution degree of feature extraction on the current recognition text by each feature extraction model; based on the feature vector output by each semantic extraction model, classifying the current recognition text to obtain a classification result; in response to the fact that the classification result meets the interaction condition, the current recognition text is recognized, and a recognition result is obtained. Therefore, based on the semantics of the current recognition text and the historical semantics provided by the historical interaction information, the understanding ability of the current recognition text can be improved, and thus the accuracy of text recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of language processing technology, and relate to but are not limited to a text recognition method, apparatus, device and storage medium. Background Art

[0002] At present, voice human-computer interaction technology is developing rapidly. This technology realizes natural communication between people and computers through sound signals. It uses core technologies such as speech recognition, speech synthesis and natural language processing to enable computers to understand and generate natural language to perform tasks. Voice human-computer interaction can free the user's hands and eyes, making it more convenient for users to perform other tasks. In the car scene, users can order music by voice.

[0003] In complex scenarios, speech recognition systems often collect a lot of noise, which inevitably introduces invalid intentions or problems. At this time, it is necessary to reject the recognition of such text intentions. However, sometimes when users are interacting with the computer, the speech recognition system cannot accurately understand the user's intentions due to the simplicity of expression, resulting in the user's intentions being rejected by the speech recognition system. Therefore, how to reject noisy text while retaining the user's intentions is an urgent problem to be solved. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a text recognition method, apparatus, device and storage medium to solve the problem in the related art that when rejecting noisy text, the user's intention cannot be understood due to the concise expression of the user, thus rejecting the user's intention.

[0005] In order to achieve the above object, the present application provides a text recognition method, which includes:

[0006] Based on at least one semantic extraction model, feature extraction is performed on the current recognized text and the historical interaction information before the current recognized text to obtain a feature vector output by each semantic extraction model; wherein the at least one semantic extraction model is determined among a plurality of feature extraction models with different feature extraction types based on the contribution degree of each feature extraction model to the feature extraction of the current recognized text; based on the feature vector output by each semantic extraction model, the current recognized text is classified to obtain a classification result; in response to the classification result satisfying an interaction condition, the current recognized text is recognized to obtain a recognition result.

[0007] According to the above-mentioned technical means, by combining the current recognized text and the historical interaction information, the semantics and contextual information of the current recognized text can be understood more comprehensively, and the ability to understand the current recognized text is improved, thereby improving the accuracy of classification; at the same time, at least one semantic extraction model performs feature extraction on the current recognized text and the historical interaction information in parallel, which can improve efficiency on the basis of extracting a variety of different semantic features; and through different types of feature extraction, the semantic information and structural features of the current recognized text can be more comprehensively captured, the accuracy of classification is improved, and the ability to judge whether to recognize or reject text is improved, thereby reducing the situation of false rejection and omission of intended text.

[0008] In some embodiments, the text recognition method also includes: performing word segmentation processing on the current recognized text and the historical interaction information respectively to obtain at least one first word corresponding to the current recognized text and at least one second word corresponding to the historical interaction information; performing vector mapping on the at least one first word and the at least one second word respectively to obtain at least one first vector and at least one second vector corresponding to the current recognized text; based on the timing of the at least one first word in the current recognized text and the timing of the at least one second word in the historical interaction information, position encoding each first vector and each second vector to obtain a timing coding sequence corresponding to the current recognized text and the historical interaction information; correspondingly, based on at least one semantic extraction model, feature extraction is performed on the current recognized text and the historical interaction information before the current recognized text to obtain a feature vector output by each semantic extraction model, including: based on the at least one semantic extraction model, feature extraction is performed on the timing coding sequence respectively to obtain a feature vector output by each semantic extraction model.

[0009] According to the above technical means, words are converted into vectors so that similar words are closer in the vector space, which can capture the relationship between words and is conducive to extracting context information during subsequent feature extraction. Position encoding is performed so that the relative position of each word in the sentence can be understood during subsequent feature extraction, thereby improving the accuracy of text recognition or rejection.

[0010] In some embodiments, the text recognition method also includes: obtaining multiple feature extraction models with different task types and contribution weights of each feature extraction model, wherein the contribution weights are obtained by training based on sample interaction data with classification labels; calculating the contribution degree of each feature extraction model based on the contribution weights of each feature extraction model and the temporal coding sequence; and among multiple feature extraction models, determining at least one feature extraction model whose contribution degree meets the contribution condition as the at least one semantic extraction model.

[0011] According to the above technical means, the contribution degree of each feature extraction model is used to evaluate which models are more suitable for processing the current recognition text, and at least one semantic extraction model is selected to extract features of different aspects of the text to improve the accuracy of text recognition or rejection.

[0012] In some embodiments, the currently recognized text is classified based on the feature vector output by each semantic extraction model to obtain a classification result, including: based on the contribution degree of each semantic extraction model, weighted calculation is performed on the feature vector output by each semantic extraction model to obtain a weighted feature vector corresponding to each semantic extraction model; the weighted feature vectors corresponding to each semantic extraction model are spliced ​​to obtain a spliced ​​vector; based on the self-attention mechanism, feature extraction is performed on the spliced ​​vector to obtain a target feature vector; classification prediction is performed on the target feature vector to obtain the classification result.

[0013] In some embodiments, the text recognition method also includes: receiving voice data; performing voice recognition on the voice data to obtain voice recognition text; performing semantic recognition on the voice recognition text to obtain a semantic recognition result; in response to the semantic recognition result characterizing that the voice recognition text is non-noise text, determining the voice recognition text as the current recognition text.

[0014] According to the above technical means, noisy text and meaningless text input are rejected through semantic recognition, so that the currently recognized text is non-noise text, avoiding the problem of multiple semantic extraction models extracting noisy text, reducing meaningless calculations, and improving interaction efficiency.

[0015] In some embodiments, performing semantic recognition on the speech recognition text to obtain a semantic recognition result includes: performing vector mapping on the speech recognition text to obtain a speech recognition vector sequence corresponding to the speech recognition text; performing similarity comparison between the speech recognition vectors in the speech recognition vector sequence and the word vectors in a preset vocabulary to obtain a similarity result; in response to the similarity result not satisfying a similarity condition, performing feature extraction on the speech recognition vector sequence to obtain a speech recognition feature; and based on the speech recognition feature, performing noise prediction on the speech recognition text to obtain the semantic recognition result.

[0016] According to the above technical means, text similarity matching can be used to quickly screen out common vehicle control problems, thereby improving the response speed of interaction; when the speech recognition text is a more general statement or is outside the candidate intention or question, feature extraction is used to understand and predict whether the input text is noise text, thereby avoiding rejection of text with user intention.

[0017] In some embodiments, the method further includes: in response to the similarity result satisfying a similarity condition, determining the speech recognition text as the current recognition text.

[0018] In some embodiments, the vector mapping of the speech recognition text to obtain a speech recognition vector sequence corresponding to the speech recognition text includes: performing word segmentation processing on the speech recognition text to obtain at least one third word corresponding to the speech recognition text; vector mapping the at least one third word to obtain a speech recognition vector corresponding to each third word; and sorting the speech recognition vectors corresponding to each third word based on the time sequence of each third word in the speech recognition text to obtain the speech recognition vector sequence.

[0019] According to the above technical means, the speech recognition text is converted into a vector, which is conducive to extracting context information during feature extraction to improve the accuracy of noise recognition.

[0020] In some embodiments, the feature extraction of the speech recognition vector sequence to obtain the speech recognition feature includes: performing forward feature extraction and backward feature extraction on the speech recognition vector sequence respectively to obtain a first feature vector sequence and a second feature vector sequence; splicing the first feature vector sequence and the second feature vector sequence to obtain a spliced ​​vector sequence; and based on a self-attention mechanism, performing feature extraction on the spliced ​​vector sequence to obtain the speech recognition feature.

[0021] An embodiment of the present application further provides a text recognition device, which includes: a feature extraction module, which is used to perform feature extraction on a current recognized text and historical interaction information before the current recognized text based on at least one semantic extraction model, to obtain a feature vector output by each semantic extraction model; wherein the at least one semantic extraction model is determined based on the contribution degree of each feature extraction model to the feature extraction of the current recognized text among multiple feature extraction models of different feature extraction types; a classification module, which is used to classify the current recognized text based on the feature vector output by each semantic extraction model, to obtain a classification result; and a recognition module, which is used to recognize the current recognized text in response to the classification result satisfying an interaction condition, to obtain a recognition result.

[0022] According to the above-mentioned technical means, by combining the current recognized text and the historical interaction information, the semantics and contextual information of the current recognized text can be understood more comprehensively, thereby improving the ability to understand the current recognized text and thus improving the accuracy of classification; at the same time, at least one semantic extraction model performs feature extraction on the current recognized text and the historical interaction information in parallel, thereby improving efficiency on the basis of extracting multiple features; and through different types of feature extraction, the semantic information and structural features of the current recognized text can be captured more comprehensively, thereby improving the accuracy of classification and the ability to judge rejection, thereby reducing the cases of false rejection and missed rejection of intended text.

[0023] An embodiment of the present application further provides a text recognition device, which includes a memory and a processor, wherein the memory stores a computer program that can be executed on the processor, and the processor implements the above-mentioned text recognition method when executing the program.

[0024] An embodiment of the present application further provides a computer-readable storage medium having executable instructions stored thereon, which is used to cause a processor to execute the executable instructions to implement the above-mentioned text recognition method. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is an optional architectural diagram of a text recognition system provided in an embodiment of the present application;

[0026] Figure 2 is a structural diagram of a text recognition device provided in an embodiment of the present application;

[0027] Figure 3 This is an optional flowchart of the text recognition method provided in the embodiment of the present application;

[0028] Figure 4 This is an optional flowchart of the text recognition method provided in the embodiment of the present application;

[0029] Figure 5 It is a flowchart of a method for rejecting text recognition of voice interaction in a car cockpit proposed in an embodiment of the present application;

[0030] Figure 6 is a schematic diagram of the noise rejection process proposed in the embodiment of the present application;

[0031] Figure 7 is a schematic diagram of the structure of the language model used in noise rejection provided in an embodiment of the present application;

[0032] Figure 8 It is a schematic diagram of the structure of the adaptive multi-round rejection model provided in the embodiment of the present application;

[0033] Fig. 9A schematic diagram of the structure of an expert model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.

[0035] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as those commonly understood by those skilled in the art of the technical field of the embodiments of the present application. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0036] Currently, the most common human-computer interaction method in car cockpits is for users to control the vehicle and ask questions through voice, such as opening the window, turning on the air conditioner, navigating to a certain place, calling someone, and asking where there are gas stations, restaurants, fun places nearby, etc. The recognition of user intentions and questions is often achieved by parsing or understanding the text corresponding to the voice. Therefore, text understanding is currently a key research direction in voice vehicle control and voice question and answer.

[0037] There are still many challenges in understanding text in the car cockpit scenario. On the one hand, in complex scenarios, the in-vehicle voice recognition system often collects a large amount of noisy text, which will inevitably introduce invalid intentions or questions. At this time, it is necessary to directly refuse to recognize the intention of such text. For example, at this time, the user is chatting with other people in the car, making phone calls, and other behaviors that are not interacting with the car computer. The car computer needs to refuse to recognize instead of recognizing all of them. On the other hand, as users have higher and higher requirements for voice car control and voice Q&A of car computers, users hope that car computers will become smarter and smarter. In many scenarios, it is no longer limited to one conversation to clarify the user's intention or question. In the process of multiple rounds of conversations, the user's expression may be brief, semantically incomplete, or even meaningless. In this case, the car computer may refuse to recognize the user's question, or it may not be able to refuse to recognize it.

[0038] Therefore, it is particularly important for the car computer to reject the semantics of the input voice through a text recognition module before executing the voice interaction business process. It can reject the user's voice when it is unclear, Mandarin is not standard, there are people chatting around, and the intention is unclear. It can reject or not reject the user in the case of single-round or multi-round interaction. By rejecting meaningless text input in advance, the pressure of subsequent business can be reduced and back-end computing resources can be saved. Multiple rounds of understanding can give users a better interactive experience, thereby improving the robustness and experience of the entire smart cockpit interaction.

[0039] The text recognition method of the related art uses rule matching to determine whether the user input text contains intent. For example, first preset a limited number of common fields and common intent candidate sets, then make the input text match the candidate set based on the rules, and finally execute the corresponding intent if it matches, otherwise refuse to respond or voice broadcast "Sorry, I don't understand what you mean".

[0040] The disadvantage of the rule-based rejection method in related technologies is that it has poor generalization ability. As long as the user's statement is slightly changed, it will cause matching failure and be rejected. With the rapid development of deep learning, deep learning models are applied to text understanding, and the corresponding text recognition method in voice interaction is changed to a rule plus model approach. The use of deep learning models can generalize more similar descriptions, and then use rule matching to intercept common intentions and problem expressions or to solve examples that the model cannot cover.

[0041] The voice interaction in the car cockpit can be "User: What is Sentry Mode? Assistant: Sentry Mode is a safety guard function when the vehicle is parked. It monitors the situation around the vehicle, records accidents and promptly notifies the owner. User: How to turn it on? Assistant: Usually in the vehicle's central control system or driving mode menu, find the safety settings or vehicle monitoring options and follow the prompts to enable Sentry Mode. User: Turn it on for me. Assistant: Sentry Mode has been turned on for you."

[0042] In the process of human-computer interaction in an actual environment, when the user asks "What is the Sentinel Mode?" and "How to turn it on?", due to defects in Automatic Speech Recognition (ASR) or severe environmental noise, the recognized text may become "Ten yards is the sentinel magic, loving beauty" and "How to turn it on". For input text without semantics like this, it is called noise text, and such text should be rejected. In the second and third rounds of conversation, if it is judged whether to reject based only on the questions in the second and third rounds, the expected result should be rejection, because from the semantic understanding, for "How to turn it on" and "Help me turn it on", it is not clear what the user intends to turn on, and the semantics of these two input texts are incomplete. However, if the history of the first round can be combined to infer that it is to turn on the "Sentinel Mode", then the rejection results in the second and third rounds should be non-rejection for a more reasonable result.

[0043] However, in the solutions for text rejection using deep learning in related technologies, due to the poor semantic understanding ability of shallow convolutional networks, it is difficult to reject noise text introduced by speech recognition defects, and historical interaction data is not used to determine the current input content, resulting in the speech interaction system being unable to understand the meaning of the input content and rejecting input text with actual intentions of the user, causing the speech interaction system to be unable to give timely feedback to the user and affecting the user experience.

[0044] Based on the problems existing in related technologies, the embodiments of this application provide a text recognition method, which can perform different types of feature extraction on the current recognized text and previous historical interaction information based on multiple different semantic extraction models, and then perform classification prediction on the current recognized text based on the feature vectors output by each model, and judge whether to recognize or reject the current recognized text based on the classification result.

[0045] In this way, by combining the current recognized text and historical interaction information, the embodiments of this application can more comprehensively understand the semantics and context information of the current recognized text, improve the understanding ability of the current recognized text, and thus improve the accuracy of classification; and through different types of feature extraction, the semantic information and structural features of the current recognized text can be more comprehensively captured, improving the accuracy of classification, improving the judgment ability of recognizing or rejecting text, and thus reducing the situations of mis-rejection and missed rejection of text with intentions.

[0046] The following describes an exemplary application of the text recognition device of the embodiment of the present application. The text recognition device provided by the embodiment of the present application can be implemented as a terminal or a server. In one implementation, the text recognition device provided by the embodiment of the present application can be implemented as a laptop, a tablet computer, a desktop computer, a mobile device (e.g., a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device), an intelligent robot, an intelligent home appliance, and an intelligent vehicle-mounted device, etc., any terminal with a data processing function; in another implementation, the text recognition device provided by the embodiment of the present application can also be implemented as a server, wherein the server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected by wired or wireless communication, which is not limited in the embodiment of the present application. Below, the exemplary application of the text recognition device when it is implemented as a server will be described.

[0047] This application embodiment provides a text recognition system, see Figure 1 , Figure 1 It is an optional architecture diagram of the text recognition system provided by the embodiment of the present application. The text recognition system 10 includes at least a vehicle 100, a network 200 and a server 300, wherein the server 300 can constitute the text recognition device of the embodiment of the present application. The vehicle 100 is connected to the server 300 through the network 200, and the network 200 can be a wide area network or a local area network, or a combination of the two. When performing text rejection, the server 300 obtains the historical interaction information before the current recognition text and the current recognition text, and extracts features of the historical interaction information before the current recognition text and the current recognition text based on at least one semantic extraction model to obtain the feature vector output by each semantic extraction model, and classifies the current recognition text based on the feature vector output by each semantic extraction model to obtain the classification result. In response to the classification result satisfying the interaction condition, the current recognition text is recognized to obtain the recognition result. The server 300 can send the recognition result to the vehicle 100 based on the network 200, and the vehicle 100 makes feedback based on the recognition result. When the classification result does not meet the interaction condition, that is, when the current recognition text needs to be rejected, the server 300 does not recognize the current recognition text.

[0048] In some embodiments, if the classification result meets the interaction conditions, it means that the currently recognized text has the user's intention. At this time, the vehicle 100 provides feedback to the currently recognized text, such as executing the action of the currently recognized text (such as opening the sunroof), or answering the question raised by the currently recognized text.

[0049] The following describes an exemplary application of the text recognition method of the embodiment of the present application. The text recognition device provided in the embodiment of the present application can be located on a vehicle, or it can be implemented as a cloud server. In one implementation, the server can be an independent physical server, or it can be a server cluster or distributed system composed of multiple physical servers, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiment of the present application. Below, an exemplary application when the text recognition device is a server will be described.

[0050] It should be noted here that cloud technology refers to a hosting technology that unifies hardware, software, network and other resources in a wide area network or local area network to achieve data computing, storage, processing and sharing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool that can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The backend services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark, and all need to be transmitted to the backend system for logical processing. Data of different levels will be processed separately. All kinds of industry data require strong system backing support, which can only be achieved through cloud computing.

[0051] Figure 2 is a structural diagram of a text recognition device provided in an embodiment of the present application, Figure 2 The text recognition device shown includes: at least one processor 210, a memory 250, at least one network interface 220 and a user interface 230. The various components in the text recognition device are coupled together via a bus system 240. It is understood that the bus system 240 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 240 is not described in detail. Figure 2 Various buses are labeled as bus system 240 .

[0052] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0053] The user interface 230 includes one or more output devices 231 that enable presentation of media content, and one or more input devices 232. Here, the input device 232 may be a microphone in the vehicle, and the output device may be a display device (eg, a central control screen) or a voice output device of the vehicle.

[0054] The memory 250 may be removable, non-removable or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drive, optical disk drive, etc. The memory 250 may optionally include one or more storage devices physically away from the processor 210. The memory 250 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 250 described in the embodiment of the present application is intended to include any suitable type of memory. In some embodiments, the memory 250 can store data to support various operations, and examples of these data include programs, modules, and data structures or subsets or supersets thereof, as exemplarily described below.

[0055] The operating system 251 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic businesses and process hardware-based tasks.

[0056] The network communication module 252 is used to reach other computing devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, Wireless Fidelity (Wi-Fi), and Universal Serial Bus (USB).

[0057] The input processing module 253 is used to detect one or more user inputs or interactions from one of the one or more input devices 232 and translate the detected inputs or interactions.

[0058] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 2 A text recognition device 254 stored in the memory 250 is shown. The text recognition device 254 may be a text recognition device in a text recognition device, which may be software in the form of a program or a plug-in, and includes the following software modules: a feature extraction module 2541, a classification module 2542, and a recognition module 2543. These modules are logical, and thus may be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.

[0059] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the text recognition method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components.

[0060] The text recognition method provided in each embodiment of the present application can be executed by a text recognition device, wherein the text recognition device can be any vehicle-mounted device with a text rejection processing function, or it can also be a server, that is, the text recognition method provided in each embodiment of the present application can be executed by a vehicle-mounted device, or it can be executed by a server, or it can also be executed by interaction between the vehicle-mounted device and the server.

[0061] See also Figure 3 , Figure 3 This is an optional flow chart of the text recognition method provided in the embodiment of the present application. Figure 3 The steps shown are explained, and it should be noted that Figure 3 The text recognition method in the embodiment of the present application is described by taking the server as the execution subject as an example. The text recognition method provided in the embodiment of the present application can be implemented by steps S301 to S303:

[0062] Step S301: Based on at least one semantic extraction model, feature extraction is performed on the current recognized text and the historical interaction information before the current recognized text to obtain a feature vector output by each semantic extraction model; wherein the at least one semantic extraction model is determined among multiple feature extraction models of different feature extraction types based on the contribution degree of each feature extraction model to the feature extraction of the current recognized text.

[0063] In some embodiments, when the text recognition method is applied to a scenario where a user interacts with a vehicle computer, the currently recognized text refers to the current text of the user's interaction with the vehicle computer, which may be text converted from the user's current voice. For example, the currently recognized text may be "help me drive".

[0064] In an embodiment of the present application, the currently recognized text is the text that has undergone noise recognition, that is, the acquired user voice is converted into text, the text is judged for noise, the chat or other non-instructive conversations between passengers in the car are removed, and the non-noise text is determined as the currently recognized text.

[0065] In some embodiments, historical interaction information refers to multiple rounds of interactions between the user and the vehicle computer before the current text recognition, for example, user: What is sentry mode? Assistant: Sentry mode is a safety guard function when the vehicle is parked, which monitors the situation around the vehicle, records accidents and notifies the owner in time; User: How to turn it on? Assistant: Usually find the safety settings or vehicle monitoring options in the vehicle's central control system or driving mode menu, and follow the prompts to enable sentry mode.

[0066] Here, since the current recognized text "Help me open it" does not clearly understand what the user intends to open from the perspective of semantic understanding, that is, the semantics of the current recognized text are incomplete. If only feature extraction is performed on the current recognized text, the user's intention cannot be obtained and it will be rejected. Therefore, the embodiment of the present application combines historical interaction information to understand the current recognized text and avoid rejecting text with user intention, so as to get closer to the user's actual expectations.

[0067] In an embodiment of the present application, different semantic extraction models have different feature extraction types, and each semantic extraction model is used to extract features of one aspect of the current recognized text and historical interaction information. For example, the semantic extraction model can be a semantic strength extraction model, which is used to extract features of whether the user question is a vehicle control question; it can also be a semantic completeness extraction model, which is used to extract features of whether the user question is semantically complete; it can also be a multi-round relationship extraction model, which is used to extract features associated with the current recognized text and historical interaction information; it can also be a keyword extraction model, which is used to extract features of the keywords of the current recognized text; it can also be an emotion extraction model, which is used to extract features of the emotional tendency of the current recognized text; it can also be a topic extraction model, which is used to extract features of the topic of the current recognized text.

[0068] Here, at least one semantic extraction model extracts features of at least one aspect of the current recognition text and the historical interaction information respectively to obtain at least one feature vector.

[0069] Each semantic extraction model can extract corresponding semantic features. The embodiment of the present application integrates the features extracted by all semantic extraction models to learn more comprehensive information about the current recognized text.

[0070] Here, at least one semantic extraction model can be selected from multiple feature extraction models with different feature extraction types, and the contribution of each feature extraction model to the current recognized text is evaluated, that is, the response degree and adaptability of each feature extraction model to the current recognized text, that is, the ability of each feature extraction model to extract features from the current recognized text, to determine which feature extraction models are more suitable for processing the current recognized text, calculate the probability of each feature extraction model, and select which feature extraction models to use through probability distribution to obtain at least one semantic extraction model.

[0071] Step S302: classify the current recognized text based on the feature vector output by each semantic extraction model to obtain a classification result.

[0072] In the embodiment of the present application, the feature vectors output by each semantic extraction model may be fused, and then a 2-classification prediction may be performed through softmax to calculate the classification result of the current recognized text.

[0073] Here, the classification results can be divided into two categories: recognizable and unrecognizable. Recognizable means that after feature extraction of the current recognized text and historical interaction information, the user's clear intention can be obtained, and the car computer can give feedback to the user based on the current recognized text, such as opening the sunroof on the car; unrecognizable means that after feature extraction of the current recognized text and historical interaction information, the user's intention cannot be determined.

[0074] Step S303: In response to the classification result satisfying the interaction condition, the current recognition text is recognized to obtain a recognition result.

[0075] Here, the classification result satisfying the interaction condition may mean that the classification result is recognizable, and the user intention can be determined based on historical interaction information. Therefore, the interaction condition is met, the current recognized text is recognized, and the intention of the current recognized text is obtained, and the vehicle computer makes feedback based on the intention; the classification result not satisfying the interaction condition may mean that the classification result is unrecognizable, and the user intention cannot be determined. Therefore, the interaction condition is not met, the current recognized text is rejected, and no response is made to the current recognized text, thereby avoiding misoperation or interference with the user's normal communication.

[0076] The embodiments of the present application can more comprehensively understand the semantics and contextual information of the currently recognized text by combining the currently recognized text and the historical interaction information, thereby improving the ability to understand the currently recognized text and thus improving the accuracy of classification; at the same time, at least one semantic extraction model performs feature extraction on the currently recognized text and the historical interaction information in parallel, thereby improving efficiency on the basis of extracting multiple features; and through different types of feature extraction, the semantic information and structural features of the currently recognized text can be more comprehensively captured, thereby improving the accuracy of classification and the ability to judge whether to recognize or reject text, thereby reducing the situation where intentional text is mistakenly rejected or missed.

[0077] In some embodiments, before extracting features from the current recognized text and the historical interaction information, the current recognized text and the historical interaction information need to be preprocessed. Therefore, the text recognition method provided by the present application may further include steps S1 to S3:

[0078] Step S1: perform word segmentation processing on the current recognition text and the historical interaction information respectively to obtain at least one first word corresponding to the current recognition text and at least one second word corresponding to the historical interaction information.

[0079] In an embodiment of the present application, before performing word segmentation processing, the text can be preprocessed, including removing noise such as punctuation marks, and then using word segmentation tools such as Jieba word segmentation (Jieba), Simplified Chinese Text Processing (SnowNLP), Language Technology Platform (LTP) and Han Language Processing (HanNLP) to perform word segmentation processing on the current recognized text and historical interaction information, respectively, to obtain at least one first word corresponding to the current recognized text and at least one second word corresponding to the historical interaction information.

[0080] Step S2: Perform vector mapping on the at least one first word and the at least one second word respectively to obtain corresponding at least one first vector and at least one second vector.

[0081] In some embodiments, at least one first word and at least one second word can be vector mapped through word embedding, the semantic relationship between the words can be learned, and the words can be mapped into the vector space so that the vector can capture the semantic information of the words, thereby obtaining at least one first vector corresponding to at least one first word and at least one second vector corresponding to at least one second word.

[0082] Step S3: Based on the timing of the at least one first word in the current recognition text and the timing of the at least one second word in the historical interaction information, position encode each first vector and each second vector to obtain a timing coding sequence corresponding to the current recognition text and the historical interaction information.

[0083] Here, after representing the text as a word vector, it is necessary to add positional encoding to each word vector to introduce the temporal position information of the word in the current recognized text and historical interaction information. The vector of each word is added to its corresponding position encoding vector to obtain the final vector of each word, and the temporal encoding sequence corresponding to the current recognized text and historical interaction information is obtained.

[0084] Correspondingly, step S301 can be implemented by step S3011:

[0085] Step S3011: Based on the at least one semantic extraction model, feature extraction is performed on the temporal coding sequence respectively to obtain a feature vector output by each semantic extraction model.

[0086] In the present application, after the current recognition text and historical interaction information are converted into a temporal coding sequence, feature extraction can be performed on the temporal coding sequence based on at least one semantic extraction model to obtain a feature vector output by each semantic extraction model.

[0087] The embodiment of the present application converts words into vectors so that similar words are closer in the vector space, which can capture the relationship between words, facilitate the extraction of context information during subsequent feature extraction, and perform position encoding so that the relative position of each word in the sentence can be understood during subsequent feature extraction, thereby improving the accuracy of text recognition or rejection.

[0088] In some embodiments, at least one semantic extraction model may be at least one model with the greatest contribution selected by a routing network based on the current recognized text among multiple feature extraction networks. Therefore, the text recognition method provided by the present application may also include steps S4 to S6:

[0089] Step S4: obtaining a plurality of feature extraction models with different task types and a contribution weight of each feature extraction model, wherein the contribution weight is obtained by training based on sample interaction data with classification labels.

[0090] In the embodiment of the present application, the multiple feature extraction models with different task types can be a semantic strength extraction model, a semantic complete extraction model, a multi-round relationship extraction model, a keyword extraction model, a sentiment extraction model, and a topic extraction model, etc. The feature extraction model is a pre-trained model that can realize different types of feature extraction.

[0091] The contribution weight of each feature extraction model can be obtained by training based on sample interaction data with classification labels. For example, feature extraction is performed on the sample interaction data through multiple feature extraction models, and then the feature vectors obtained based on each feature extraction model and the corresponding contribution weights are fused to obtain a fusion result. Classification is performed based on the fusion result to obtain a classification result. The classification result is compared with the classification label of the sample interaction data. When the classification result is different from the classification label, the contribution weight is adjusted until the classification result is the same as the classification label, thereby obtaining the contribution weight of each feature extraction model.

[0092] Step S5: Calculate the contribution degree of each feature extraction model based on the contribution weight of each feature extraction model and the temporal coding sequence.

[0093] In the embodiment of the present application, the contribution weight of each feature extraction model can be W r , the contribution weight can be a weight matrix, and the temporal coding sequence is x = [x1, ..., x n ], the contribution degree p(x) of each feature extraction model can be calculated by formula (1):

[0094] p(x)=softmax(W r ·x) (1);

[0095] Here, based on the result of multiplying the weight matrix and the time series coding sequence, the SoftMax function is applied to generate the contribution degree of each feature extraction model to the time series coding sequence, that is, the contribution probability.

[0096] In some embodiments, the contribution weight of each feature extraction model may be determined by a gating mechanism.

[0097] Step S6: Among the multiple feature extraction models, determine at least one feature extraction model whose contribution degree meets the contribution condition as the at least one semantic extraction model.

[0098] In the embodiment of the present application, satisfying the contribution condition may mean that the three feature extraction models with the highest contribution, or the feature extraction models with a contribution greater than a threshold (which may be 20%), are determined as at least one semantic extraction model.

[0099] In an embodiment of the present application, the contribution of each feature extraction model is used to evaluate which models are more suitable for processing the current recognition text, and at least one semantic extraction model is selected to extract features of different aspects of the text to improve the accuracy of text recognition or rejection.

[0100] In some embodiments, step S302 may be implemented by steps S3021 to S3024:

[0101] Step S3021: Based on the contribution degree of each semantic extraction model, perform weighted calculation on the feature vector output by each semantic extraction model to obtain a weighted feature vector corresponding to each semantic extraction model.

[0102] In the embodiment of the present application, the feature vector output by each semantic extraction model can be E i (x), the contribution of each semantic extraction model is p i (x), the weighted feature vector after weighted calculation can be p1(x)·E1(x),…,p k (x)·E k (x), where k is the number of semantic extraction models.

[0103] Step S3022: concatenate the weighted feature vectors corresponding to each semantic extraction model to obtain a concatenated vector.

[0104] In the embodiment of the present application, the feature vector is obtained based on the time sequence of the words, and the concatenation can be cascaded along the time sequence dimension to obtain the concatenation vector concat(p1(x)·E1(x),…,p k (x)·E k (x)).

[0105] Step S3023: Based on the self-attention mechanism, feature extraction is performed on the concatenated vector to obtain a target feature vector.

[0106] In an embodiment of the present application, the spliced ​​splicing vector enters the self-attention layer, and the self-attention mechanism of the self-attention layer converts the splicing vector into a query (Q) vector, a key (K) vector, and a value (V) vector through a linear transformation. Then, the dot product between the query (Q) vector and the key (K) vector representation is calculated to obtain the attention score, and the softmax function is applied to obtain the weight matrix. The obtained weight matrix is ​​used to perform weighted aggregation on the value (V) vector to obtain the target feature vector.

[0107] The self-attention mechanism can capture the correlation features between the feature vectors output by different semantic extraction models.

[0108] Step S3024: perform classification prediction on the target feature vector to obtain the classification result.

[0109] Here, the classification prediction can be performed through the softmax algorithm to perform 2-classification prediction and calculate the classification result of the current recognized text.

[0110] In some embodiments, before classification, the dimension of the target feature vector may be compressed by a multilayer perceptron (MLP) to further extract higher-level feature information.

[0111] In an embodiment of the present application, noise rejection can be performed on the currently recognized text first, and feature extraction can be performed when the currently recognized text is not noise, so as to avoid feature extraction of a large amount of noise text, which increases the amount of calculation. Figure 4 is an optional flow chart of the text recognition method provided in the embodiment of the present application. As shown in the figure, the text recognition method provided in the embodiment of the present application may also include steps S401 to S404:

[0112] Step S401: Receive voice data.

[0113] Here, the car computer can receive voice data sent by the user based on the vehicle's microphone array.

[0114] Step S402: Perform speech recognition on the speech data to obtain speech recognition text.

[0115] In the embodiment of the present application, the speech can be converted into text by automatic speech recognition (ASR) to obtain speech recognition text. Here, the speech conversion can be performed by any feasible speech-to-text method.

[0116] Step S403: Perform semantic recognition on the speech recognition text to obtain a semantic recognition result.

[0117] In the embodiment of the present application, semantic recognition is used to reject noisy text and meaningless text input, and is also used to ensure that the semantics of multiple rounds of historical interaction data (i.e., historical interaction information) are clear, which also makes it easier to understand the historical interaction text later.

[0118] Here, semantic recognition is used to understand the meaning of the input text (i.e., speech recognition text) to determine whether the speech recognition text is noise text or meaningless text. Semantic recognition can use the method of text similarity matching plus language understanding model to identify whether the current round of text is noise text. For example, the speech recognition text is first matched with the words in the preset vocabulary for text similarity. If the similarity is greater than the threshold, it is considered non-noise. When the similarity is lower than the threshold, the language understanding model is used to encode and understand the text information, and predict whether it is noise text to obtain the semantic recognition result. The semantic recognition result is noise text or non-noise text.

[0119] Step S404: In response to the semantic recognition result indicating that the speech recognition text is a non-noise text, the speech recognition text is determined as the current recognition text.

[0120] In some embodiments, when the semantic recognition result indicates that the speech recognition text is non-noise text, the speech recognition text is determined as the current recognition text, and feature extraction is performed to determine whether to reject the recognition.

[0121] Here, if it is determined to be noise text, the user can be prompted to speak again to avoid missing the user's intention.

[0122] The embodiment of the present application can reject noisy text and meaningless text input through semantic recognition, so that the currently recognized text is non-noise text, avoiding the problem of multiple semantic extraction models extracting noisy text, reducing meaningless calculations, and improving interaction efficiency.

[0123] In some embodiments, step S403 may be implemented by steps S4031 to S4034:

[0124] Step S4031, perform vector mapping on the speech recognition text to obtain a speech recognition vector sequence corresponding to the speech recognition text.

[0125] In some embodiments, step S4031 may be implemented by steps S10 to S12:

[0126] Step S10: performing word segmentation processing on the speech recognition text to obtain at least one third word corresponding to the speech recognition text.

[0127] In some embodiments, the speech recognition text may be segmented using segmentation tools such as Jieba, SnowNLP, LTP, and HanNLP to obtain at least one third word corresponding to the speech recognition text.

[0128] Step S11: perform vector mapping on the at least one third word to obtain a speech recognition vector corresponding to each third word.

[0129] In the embodiment of the present application, vector mapping can be performed through word embedding to obtain a speech recognition vector corresponding to each third word.

[0130] Step S12: sorting the speech recognition vectors corresponding to each third word based on the time sequence of each third word in the speech recognition text to obtain the speech recognition vector sequence.

[0131] The speech recognition vectors corresponding to each third word may be sorted based on the time sequence of each third word in the speech recognition text to obtain the speech recognition vector sequence, that is, the speech recognition vectors are position-encoded.

[0132] Step S4032: perform a similarity comparison between the speech recognition vectors in the speech recognition vector sequence and the word vectors in a preset vocabulary to obtain a similarity result.

[0133] In some embodiments, the preset vocabulary includes word vectors of multiple words in common vehicle control fields. The words can be multiple intentions and question candidate sets related to fields such as air conditioning, help, map, music, navigation, telephone, weather, etc.

[0134] The similarity comparison can be calculated by cosine similarity, as shown in formula (2):

[0135]

[0136] Among them, X is the speech recognition vector, and Y is the word vector in the preset vocabulary.

[0137] Here, if the similarity between the speech recognition vector and any word vector in the preset word library is greater than a threshold (e.g. 95%), the speech recognition text can be considered non-noise. If the similarity between the speech recognition vector and any word vector in the preset word library is less than the threshold, feature extraction is required to understand the text intent,

[0138] Step S4033: In response to the similarity result not satisfying the similarity condition, feature extraction is performed on the speech recognition vector sequence to obtain speech recognition features.

[0139] Here, the similarity result not meeting the similarity condition may mean that the similarity between the speech recognition vector and each word vector in the preset vocabulary is less than a threshold. In this case, a speech understanding model is needed to understand the text intent.

[0140] In some embodiments, to determine whether the speech recognition text is noisy text, a model with a recurrent neural network (RNN), a long short-term memory network (LSTM), a Transformer structure or a convolutional structure can be used to extract temporal feature information in the speech recognition text.

[0141] In some embodiments, step S4033 may be implemented by steps S13 to S15:

[0142] Step S13: Perform forward feature extraction and backward feature extraction on the speech recognition vector sequence to obtain a first feature vector sequence and a second feature vector sequence.

[0143] In an embodiment of the present application, two layers of LSTM can be used to capture forward features and backward features in a speech recognition vector sequence to obtain a first feature vector sequence and a second feature vector sequence.

[0144] Step S14: concatenate the first feature vector sequence and the second feature vector sequence to obtain a concatenated vector sequence.

[0145] Here, based on the time sequence, the forward first feature vector sequence and the backward second feature vector sequence can be spliced ​​to obtain a spliced ​​vector sequence, thereby obtaining a richer feature representation and enhancing the subsequent understanding of the sequence data.

[0146] Step S15: Based on the self-attention mechanism, feature extraction is performed on the concatenated vector sequence to obtain the speech recognition feature.

[0147] In an embodiment of the present application, the hidden states at all times in the splicing vector sequence can be weighted based on the self-attention mechanism to evaluate the correlation between the time series of the splicing vector sequence, extract the more important hidden information in the splicing vector sequence, and obtain the speech recognition feature.

[0148] Here, a MLP layer can also be used to extract higher-level feature information from the speech recognition features.

[0149] Step S4034: Based on the speech recognition feature, noise prediction is performed on the speech recognition text to obtain the semantic recognition result.

[0150] In an embodiment of the present application, the probability value of the speech recognition text being a noise text can be predicted by softmax. When the predicted probability is greater than a threshold (0.5 can be selected as an example), the input text is determined to be noise or meaningless text, otherwise it is not noise text.

[0151] The embodiment of the present application uses text similarity matching to quickly screen out common vehicle control problems, thereby improving the response speed of the interaction; when the speech recognition text is a more general statement or is outside the candidate intention or question, feature extraction is used to understand and predict whether the input text is noise text, thereby avoiding rejection of text with user intention.

[0152] In some embodiments, the text recognition method further includes step S20:

[0153] Step S20: In response to the similarity result satisfying a similarity condition, determining the speech recognition text as the current recognition text.

[0154] In some embodiments, if the similarity between the speech recognition vector and any word vector in the preset vocabulary is greater than a threshold (eg, 95%), the speech recognition text can be considered to be non-noise and the speech recognition text is determined as the current recognition text.

[0155] The embodiment of the present application further provides an application of a text recognition method in a practical scenario.

[0156] Based on the problems existing in the related technology, the embodiment of the present application proposes a text rejection method based on car cockpit voice interaction, which can effectively reject noisy text and at the same time use historical multi-round interaction information to make more reasonable rejection decisions.

[0157] The text rejection method provided in the embodiment of the present application is a two-level rejection method. The first is noise rejection, which is used to reject noise input caused by environmental reasons and remove noise text to ensure the semantic clarity of historical interactive texts. Then, an adaptive multi-round rejection model is used to more reasonably comprehensively predict the rejection result of the current input. At the level of the multi-round rejection method, the MOE structure is used to design expert models in three professional domains to extract semantic strength features, semantic completeness features, and multi-round relationship features. The features extracted by the expert model are adaptively weighed through parameter learning. In the final feature fusion, the weighted average fusion is changed to a cascade plus self-attention fusion method, so that the correlation features between experts can be learned in addition, which helps to improve the rejection effect.

[0158] Figure 5 is a flow chart of a method for rejecting text recognition in a car cockpit voice interaction proposed in an embodiment of the present application, such as Figure 5As shown, the voice interaction text rejection method is implemented through steps S501 to S509:

[0159] Step S501: Convert speech into text through ASR.

[0160] In some embodiments, when a user inputs voice, the voice can be recognized into text through ASR voice recognition, and then the recognized text is input into the first level noise rejection to determine whether the current round of text is noise text. If it is noise text, the car computer voice broadcasts a reminder to say it again, and then inputs the text without noise as the current round of text. If it is not noise text, it enters the next level of multi-round rejection.

[0161] The first level of noise rejection is used to reject noisy text and meaningless text input. It is also used to ensure the clarity of multi-round historical semantics, which makes subsequent multi-round rejection easier when understanding historical text.

[0162] Step S502: Understand the text by using text similarity matching and language model.

[0163] The core problem to be solved by noise rejection is to understand the meaning of the input text, so as to judge whether the input text is noise text or meaningless text. In this embodiment, noise rejection uses the method of text similarity matching plus language understanding model to identify whether the current round of text is noise text.

[0164] Figure 6 is a schematic diagram of the noise rejection process proposed in the embodiment of the present application, such as Figure 6 As shown, determining whether a text is noise by text similarity matching plus a language model can be implemented through steps S601 to S605:

[0165] Step S601: perform word segmentation and word embedding on the text to obtain word vectors.

[0166] In an embodiment of the present application, when user text enters noise rejection, the text will be segmented, that is, the continuous text will be decomposed into words or phrases, and then word embedding will be performed, that is, each word or phrase will be mapped to the vector space to obtain a word vector.

[0167] In an embodiment of the present application, the word segmentation operation can segment the text through the Chinese jieba word segmentation, and the word embedding can be performed through the Embedding model or any feasible method.

[0168] Step S602: Calculate the similarity between the word vector and the preset domain set.

[0169] In an embodiment of the present application, more than 20 commonly used word sets for common vehicle control domains can be preset in advance, such as air conditioning, help, map, music, navigation, telephone, weather, etc., as well as more than 300 domain-related intent and question candidate sets to obtain a preset domain set.

[0170] Here, the similarity between the input text word vector and the text word vectors of all preset candidate intent sets in the preset domain set can be calculated. When the similarity is greater than a threshold (which can be set to 95% for example), it is predicted to be non-noise text. When the similarity is less than the threshold, it enters the language model to encode and understand the text information through the language model, and predicts whether it is noise text.

[0171] Here, the text similarity measurement index can be exemplarily selected to be calculated by cosine similarity, as shown in formula (3):

[0172]

[0173] Among them, X is the word vector of the text, and Y represents the word vector in the preset field set.

[0174] Step S603: determine whether the similarity is greater than a threshold.

[0175] In the embodiment of the present application, when the similarity is greater than the threshold, the text is considered to be non-noise text, and step S604 is executed, and the non-noise text enters multiple rounds of rejection; if the similarity is less than the threshold, the text is considered to be noise text, and step S605 is executed.

[0176] Step S604: non-noise text.

[0177] Here, when the similarity is greater than the threshold, the text is considered to be non-noise text and enters multiple rounds of rejection. The text is understood using historical multiple rounds of interaction information to determine again whether rejection is required.

[0178] Step S605: Understand the text information through the language model and predict whether it is noise text.

[0179] In some embodiments, the most commonly used language model for determining whether the current round of input text is noise text is a time series model with an RNN, LSTM or Transformer structure. A convolutional structure may also be used to extract time series feature information from the text.

[0180] Figure 7 is a schematic diagram of the structure of the language model used in noise rejection provided in the embodiment of the present application, such as Figure 7 As shown, the model first uses a two-layer long short-term memory network (LSTM) 701 to capture the text X = (x1, ..., x n), and obtain the vector H′=(h1′,…,h n ′), and then based on the Attention mechanism in the Attention layer 702, the vector H′=(h1′,…,h n ′) The hidden states of all time are weighted to evaluate the correlation between time series, which can extract the more important hidden information in the input text and obtain the weighted summed attention vector. Then, the MLP layer 703 extracts higher-level feature information from the attention vector, and finally the softmax layer predicts the probability value of the input text being a noise text. When the predicted probability is greater than a threshold (0.5 can be selected as an example), the input text is judged to be noise or meaningless text, otherwise it is not a noise text.

[0181] In the embodiment of the present application, in noise rejection, text similarity matching is first used to quickly screen out common vehicle control problems, thereby improving the response speed of the interaction; when the user input is a more generalized statement or is outside the candidate intention or question, the language model is used to understand and predict whether the input text is noise text, thereby improving the accuracy of rejection.

[0182] Step S503: determine whether the text is noise text.

[0183] In the embodiment of the present application, if the text is determined to be noise, then step S504 is executed, and the vehicle computer can remind the user to speak again; if the text is determined not to be noise, then step S505 is executed.

[0184] Step S504: reject the text.

[0185] Step S505: Use historical multi-round interaction information to understand the text.

[0186] The embodiment of the present application uses multiple rounds of historical interaction information to comprehensively determine whether the current input should be rejected. If multiple rounds of rejection ultimately determine that the current round of input is rejected, the vehicle computer will not respond. If it is determined that it is not rejected, the input text will flow to subsequent services.

[0187] The second-level multi-round rejection can predict a more reasonable rejection result for the current round of input text by utilizing historical multi-round interaction information. For example, input that is semantically incomplete but can find default information in historical interaction information should not be rejected. If the semantics is incomplete but no default information can be found in history, it should be rejected.

[0188] In some embodiments, the core problem to be solved by the multi-round rejection recognition step is how to use historical multi-round interaction information to comprehensively decide whether the input text of the current round should be rejected, so as to approach the actual expectations of users. For example, in the case where the semantics of the input of the current round is incomplete, for vehicle control interactions, it could be like "User: What is Sentry Mode; Assistant: Sentry Mode is a security protection function when the vehicle is parked, monitoring the situation around the vehicle, recording accidents and notifying the owner in a timely manner; User: How do I turn it on; Assistant: Usually in the vehicle's central control system or driving mode menu, find the security settings or vehicle monitoring options and follow the prompts to enable Sentry Mode; User: Turn it on for me; Assistant: (Sentry Mode has been enabled for you)", and for Q&A interactions, it could be like "User: What's the weather like tomorrow? Assistant: Which city's weather do you want to query? User: Beijing; Assistant: Beijing will be sunny turning to cloudy tomorrow, with temperatures ranging from 22 degrees Celsius (°C) to 32°C, and southeast wind level 3-4. The weather is relatively hot, it is recommended to drink more water and take sun protection measures. Do you need me to set up a commuting reminder for you?".

[0189] Here, the user's input in each round may include noisy text due to the influence of ASR or the environment. Most of these noisy texts can be rejected and removed in the noise rejection recognition step. Therefore, noisy texts like "zen me da ka" (the actual intention is how to open) will not appear in the historical interaction information of the multi-round rejection recognition step.

[0190] When deciding whether to reject the non-noisy problem of the current round, it is often comprehensively decided after considering many factors. For example, it is necessary to consider whether the problem of this round has strong semantics, that is, whether it is asking about a vehicle, whether the semantics of this round and historical rounds are complete, the relationship between this round and historical questions, etc. Therefore, the embodiments of the present application propose an adaptive multi-round rejection recognition model based on the MOE structure to solve the multi-round rejection recognition problem. This model is designed to use 3 expert models (i.e., semantic extraction models) plus comprehensive decision-making to determine whether to reject the text of the current round. This model can involve 1 semantic strength expert to learn the characteristics of whether the user's question is a vehicle control question, 1 semantic integrity expert to learn the characteristics of whether the user's question is semantically complete, and 1 multi-round relationship expert to learn the characteristics of the relationship between the current question and historical questions. Obviously, the experts can be increased, decreased or modified according to actual business needs.

[0191] Compared with the traditional rejection recognition method that uses one model for end-to-end learning or a series of models with rules, in the embodiments of the present application, the task of each expert model is simpler and can capture effective features more easily. Each expert model does what it is good at; compared with the hard fusion method of using empirical rules to string models, in MOE, multiple expert models are fused by learning weight parameters, which shows more adaptive ability; in addition, the embodiments of the present application adopt multiple models in parallel or data parallelism, and the inference speed is faster than the serial multi-round rejection recognition speed.

[0192] Figure 8 is a schematic diagram of the structure of the adaptive multi-round rejection model provided in the embodiment of the present application, such as Figure 8 As shown in FIG. 8 , the input of the model is the current round question and the historical multi-round interaction text. The current round question and the multi-round interaction text are first segmented in the preprocessing, and then enter the Embedding model 801 to first perform word embedding, represent the text as a word vector, and then add the position vector to each word vector through position encoding 802 to introduce the position information between words, and obtain the time sequence word vector S = (s1, ..., s n ).

[0193] Then, the time-series word vector is input into the routing network 803 (also called the gating network), which selects which expert models to use (e.g., Figure 8 The semantic strength expert E1, semantic completeness expert E2 and multi-round relationship expert E3 in the text are selected, and then the temporal word vectors will be encoded by each selected expert model to obtain their respective encoding features. Then the output features of each expert will be multiplied by their respective assigned pre-trained contribution weights, and cascaded along the temporal dimension in the Concat layer 804, and then pass through a self-attention layer 805 to learn the correlation features between each expert, and then pass through an MLP layer 806 to compress the dimension and further extract higher-level feature information. Finally, a two-class prediction is made through the softmax layer 807 to determine whether the current round of input text should be rejected. When the prediction probability is greater than the threshold (0.5 can be selected as an example), it is judged to be rejected, otherwise it is not rejected.

[0194] Here, the routing network consists of a linear transformation layer and a softmax layer. Its function is to evaluate the contribution of each expert model to the input and select which expert models to use through probability distribution. That is to say, if the probability of some expert models tends to 0, then this expert model will not actually be executed. It can also be considered that the input does not need to utilize the representation ability of this expert. This probability weight will also directly affect the feature fusion of each expert.

[0195] In some embodiments, assume that the input of the routing network is x=[x1, ..., x n ], then the probability distribution of routing network generation is calculated according to formula (4):

[0196] p(x)=softmax(W r ·x) (4);

[0197] Among them, W r Refers to the weight of the routing network.

[0198] Here, the structure of the expert model is not limited to the sequence model, and a transformer structure can be selected. Because the expert model should be designed to be as small as possible, the embodiment of the present application can adopt the following structure: Fig. 9 The model structure shown, Fig. 9 A schematic diagram of the structure of an expert model provided in an embodiment of the present application is shown in FIG. Fig. 9 As shown, the model takes the text word vector as input and uses a 6-layer transformer encoder to encode the input text features. The transformer encoder structure includes a Norm layer 901, a multi-head attention layer (Multi-Head Attention) 902, a Norm layer 903 and an MLP layer 904.

[0199] In some embodiments, it is assumed that the output of the expert model is E i (x), then the fusion of the expert coding information can be expressed by formula (5):

[0200] y=selfatt(concat(p1(x)·E1(x),…,p k (x)·E k (x)), k = 3 (5);

[0201] Here, the training method for the adaptive multi-round rejection model is to first select data to train each expert model separately, and then fix the parameters of the expert model and then train other network parts including the routing network.

[0202] Compared with the standard MOE, which is based on the granularity of the model layer, the larger and heavier layers are split into multiple small expert models, and then weighted and combined. For example, in Switch Transformers, the FFN (feedforward network) layer in the original Transformer model is replaced by the MoE layer. The MoE layer contains multiple independent expert networks, and each expert is responsible for processing a specific part of the input sequence. The outputs of these experts are selectively weighted and combined through a gating network or a routing mechanism, and finally fused before entering the next layer. The structure of the multi-round rejection model provided in the embodiment of the present application is designed based on the granularity of the task level. When the output features of each expert are fused, the standard MOE is fused by weighted summation, while the present application is fused by cascading first and then self-attention. This fusion method can learn the correlation features between each expert.

[0203] Step S506: Determine whether recognition is rejected.

[0204] Here, the softmax layer will predict whether it will be rejected. If it is rejected, step S507 will be executed; if it is not rejected, step S508 will be executed.

[0205] Step S507: reject the text.

[0206] Step S508: Do not reject the text.

[0207] In the embodiment of the present application, not rejecting the text indicates that the text can be subsequently processed.

[0208] Step S509: Perform subsequent business flow processing based on the text.

[0209] Here, performing subsequent process processing may refer to performing vehicle control based on text, or answering based on a large model.

[0210] The application scenarios of the embodiments of the present application include various text understanding and rejection scenarios under human-computer voice interaction, especially human-computer voice interaction text rejection in the scenario of a car smart cockpit. The method of the embodiments of the present application can improve the accuracy of text recognition or rejection in multi-round human-computer question and answer, and remove text that has no semantic meaning after voice recognition. Therefore, it saves computing resources for back-end processing, improves the interactive robustness of the entire dialogue system, improves the user experience, and avoids erroneous semantic understanding from causing erroneous feedback to the user end.

[0211] Based on the above embodiments, Figure 2 As shown, the text recognition device 80 includes a feature extraction module 2541 , a classification module 2542 and a recognition module 2543 .

[0212] Among them, the feature extraction module 2541 is used to extract features from the current recognized text and the historical interaction information before the current recognized text based on at least one semantic extraction model, and obtain the feature vector output by each semantic extraction model; wherein the at least one semantic extraction model is determined among multiple feature extraction models with different feature extraction types based on the contribution degree of each feature extraction model to the feature extraction of the current recognized text; the classification module 2542 is used to classify the current recognized text based on the feature vector output by each semantic extraction model, and obtain the classification result; the recognition module 2543 is used to recognize the current recognized text in response to the classification result satisfying the interaction condition, and obtain the recognition result.

[0213] In some embodiments, the text recognition device further includes: a word segmentation processing module, used to perform word segmentation processing on the current recognition text and the historical interaction information respectively, to obtain at least one first word corresponding to the current recognition text and at least one second word corresponding to the historical interaction information; a vector mapping module, used to perform vector mapping on the at least one first word and the at least one second word respectively, to obtain at least one first vector and at least one second vector corresponding thereto; a position encoding module, used to perform position encoding on each first vector and each second vector based on the time sequence of the at least one first word in the current recognition text and the time sequence of the at least one second word in the historical interaction information, to obtain a time sequence encoding sequence corresponding to the current recognition text and the historical interaction information;

[0214] Correspondingly, the feature extraction module 2541 is further used to perform feature extraction on the temporal coding sequence based on the at least one semantic extraction model to obtain a feature vector output by each semantic extraction model.

[0215] In some embodiments, the text recognition device also includes: an acquisition module for acquiring multiple feature extraction models with different task types and contribution weights of each feature extraction model, wherein the contribution weights are obtained by training based on sample interaction data with classification labels; a calculation module for calculating the contribution degree of each feature extraction model based on the contribution weight of each feature extraction model and the temporal coding sequence; and a determination module for determining, among multiple feature extraction models, at least one feature extraction model whose contribution degree meets the contribution condition as the at least one semantic extraction model.

[0216] In some embodiments, the classification module 2542 is also used to perform weighted calculation on the feature vector output by each semantic extraction model based on the contribution weight of each semantic extraction model to obtain a weighted feature vector corresponding to each semantic extraction model; to splice the weighted feature vectors corresponding to each semantic extraction model to obtain a spliced ​​vector; to perform feature extraction on the spliced ​​vector based on the self-attention mechanism to obtain a target feature vector; and to perform classification prediction on the target feature vector to obtain the classification result.

[0217] In some embodiments, the text recognition device also includes: a receiving module for receiving voice data; a voice recognition module for performing voice recognition on the voice data to obtain voice recognition text; a semantic recognition module for performing semantic recognition on the voice recognition text to obtain a semantic recognition result; and a first determination module for determining the voice recognition text as the current recognition text in response to the semantic recognition result characterizing the voice recognition text as non-noise text.

[0218] In some embodiments, the semantic recognition module is also used to perform vector mapping on the speech recognition text to obtain a speech recognition vector sequence corresponding to the speech recognition text; perform similarity comparison between the speech recognition vectors in the speech recognition vector sequence and the word vectors in a preset vocabulary to obtain a similarity result; in response to the similarity result not meeting the similarity condition, perform feature extraction on the speech recognition vector sequence to obtain speech recognition features; based on the speech recognition features, perform noise prediction on the speech recognition text to obtain the semantic recognition result.

[0219] In some embodiments, the apparatus further comprises: a second determination module configured to determine the speech recognition text as the current recognition text in response to the similarity result satisfying a similarity condition.

[0220] In some embodiments, the semantic recognition module is also used to perform word segmentation processing on the speech recognition text to obtain at least one third word corresponding to the speech recognition text; perform vector mapping on the at least one third word to obtain a speech recognition vector corresponding to each third word; and sort the speech recognition vectors corresponding to each third word based on the time sequence of each third word in the speech recognition text to obtain the speech recognition vector sequence.

[0221] In some embodiments, the semantic recognition module is also used to perform forward feature extraction and backward feature extraction on the speech recognition vector sequence respectively to obtain a first feature vector sequence and a second feature vector sequence; to splice the first feature vector sequence and the second feature vector sequence to obtain a spliced ​​vector sequence; and to perform feature extraction on the spliced ​​vector sequence based on a self-attention mechanism to obtain the speech recognition feature.

[0222] It should be noted that the description of the device of the embodiment of the present application is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment, so it is not repeated. For technical details not disclosed in the embodiment of the device, please refer to the description of the method embodiment of the present application for understanding.

[0223] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0224] The embodiment of the present application provides a computer program product or a computer program, which includes an executable instruction, which is a computer instruction; the executable instruction is stored in a computer-readable storage medium. When the processor of the text recognition device reads the executable instruction from the computer-readable storage medium and the processor executes the executable instruction, the text recognition device executes the method provided in the embodiment of the present application.

[0225] An embodiment of the present application provides a storage medium storing executable instructions, wherein executable instructions are stored. When the executable instructions are executed by a processor, the processor will be caused to execute the method provided by the embodiment of the present application.

[0226] In some embodiments, executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.

[0227] As an example, executable instructions may, but need not necessarily, correspond to a file in a file system, may be stored as part of a file storing other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions). As an example, executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected by a communication network.

[0228] It should be noted here that the description of the above storage medium is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiment of this application for understanding.

[0229] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0230] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0231] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0232] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be separately configured as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0233] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A text recognition method, characterized in that: The text recognition method comprises: Based on at least one semantic extraction model, feature extraction is performed on the current recognized text and the historical interaction information before the current recognized text to obtain a feature vector output by each semantic extraction model; wherein the at least one semantic extraction model is determined based on the contribution degree of each feature extraction model to the feature extraction of the current recognized text among a plurality of feature extraction models with different feature extraction types; Classifying the current recognized text based on the feature vector output by each semantic extraction model to obtain a classification result; In response to the classification result satisfying the interaction condition, the current recognition text is recognized to obtain a recognition result.

2. The text recognition method according to claim 1, characterized in that: The text recognition method further comprises: Performing word segmentation processing on the current recognition text and the historical interaction information respectively to obtain at least one first word corresponding to the current recognition text and at least one second word corresponding to the historical interaction information; Performing vector mapping on the at least one first word and the at least one second word respectively to obtain corresponding at least one first vector and at least one second vector; Based on the time sequence of the at least one first word in the current recognition text and the time sequence of the at least one second word in the historical interaction information, position encoding is performed on each first vector and each second vector to obtain a time sequence encoding sequence corresponding to the current recognition text and the historical interaction information; Correspondingly, the feature extraction of the current recognized text and the historical interaction information before the current recognized text is performed based on at least one semantic extraction model to obtain the feature vector output by each semantic extraction model, including: Based on the at least one semantic extraction model, feature extraction is performed on the temporal coding sequence to obtain a feature vector output by each semantic extraction model.

3. The text recognition method according to claim 1, characterized in that: The text recognition method further comprises: Acquire multiple feature extraction models with different task types and contribution weights of each feature extraction model, wherein the contribution weights are obtained by training based on sample interaction data with classification labels; Calculating the contribution degree of each feature extraction model based on the contribution weight of each feature extraction model and the temporal coding sequence; Among the multiple feature extraction models, at least one feature extraction model whose contribution degree satisfies the contribution condition is determined as the at least one semantic extraction model.

4. The text recognition method according to claim 3, characterized in that: The step of classifying the current recognized text based on the feature vector output by each semantic extraction model to obtain a classification result includes: Based on the contribution weight of each semantic extraction model, weighted calculation is performed on the feature vector output by each semantic extraction model to obtain a weighted feature vector corresponding to each semantic extraction model; The weighted feature vectors corresponding to each semantic extraction model are concatenated to obtain a concatenated vector; Based on the self-attention mechanism, feature extraction is performed on the concatenated vector to obtain a target feature vector; Perform classification prediction on the target feature vector to obtain the classification result.

5. The text recognition method according to any one of claims 1 to 4, characterized in that: The text recognition method further comprises: receiving voice data; Performing speech recognition on the speech data to obtain speech recognition text; Performing semantic recognition on the speech recognition text to obtain a semantic recognition result; In response to the semantic recognition result characterizing that the speech recognition text is a non-noise text, the speech recognition text is determined as the current recognition text.

6. The text recognition method according to claim 5, characterized in that: The performing semantic recognition on the speech recognition text to obtain a semantic recognition result includes: Performing vector mapping on the speech recognition text to obtain a speech recognition vector sequence corresponding to the speech recognition text; Comparing the speech recognition vectors in the speech recognition vector sequence with the word vectors in a preset vocabulary to obtain a similarity result; In response to the similarity result not satisfying the similarity condition, performing feature extraction on the speech recognition vector sequence to obtain speech recognition features; Based on the speech recognition feature, noise prediction is performed on the speech recognition text to obtain the semantic recognition result.

7. The text recognition method according to claim 6, characterized in that: The text recognition method further comprises: In response to the similarity result satisfying a similarity condition, the speech recognition text is determined as the current recognition text.

8. The text recognition method according to claim 6, characterized in that: The performing vector mapping on the speech recognition text to obtain a speech recognition vector sequence corresponding to the speech recognition text includes: Performing word segmentation processing on the speech recognition text to obtain at least one third word corresponding to the speech recognition text; Performing vector mapping on the at least one third word to obtain a speech recognition vector corresponding to each third word; Based on the time sequence of each third word in the speech recognition text, the speech recognition vectors corresponding to each third word are sorted to obtain the speech recognition vector sequence.

9. The text recognition method according to claim 6, characterized in that: The step of extracting features from the speech recognition vector sequence to obtain speech recognition features includes: Performing forward feature extraction and backward feature extraction on the speech recognition vector sequence respectively to obtain a first feature vector sequence and a second feature vector sequence; splicing the first feature vector sequence and the second feature vector sequence to obtain a spliced ​​vector sequence; Based on the self-attention mechanism, feature extraction is performed on the concatenated vector sequence to obtain the speech recognition feature.

10. A text recognition device, characterized in that: The text recognition device comprises: A feature extraction module, configured to extract features from a current recognized text and historical interaction information before the current recognized text based on at least one semantic extraction model, and obtain a feature vector output by each semantic extraction model; wherein the at least one semantic extraction model is determined based on the contribution of each feature extraction model to feature extraction of the current recognized text among a plurality of feature extraction models with different feature extraction types; A classification module, used for classifying the current recognized text based on the feature vector output by each semantic extraction model to obtain a classification result; The recognition module is used to recognize the current recognition text in response to the classification result satisfying the interaction condition to obtain a recognition result.

11. A text recognition device, characterized in that: The text recognition device includes a processor and a memory storing instructions executable by the processor; when the instructions are executed by the processor, the method according to any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Automatic speech recognition method and device, computer equipment and medium

    CN120673762A