A data processing method, device, and readable storage medium
By generating an enhanced audio recognition model and a general audio recognition model in the smart terminal device, and dynamically adjusting the probability reference ratio, the problem of poor recognition results caused by frequent updates of third-party applications is solved, and the accuracy and accuracy of audio recognition are improved.
Patent Information
- Application Number
- CN202111158784.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-09-30
AI Technical Summary
Due to frequent updates of contents of third-party applications, the basic voice recognition system of smart terminal devices cannot follow up in time, resulting in poor recognition results and action errors.
By obtaining the interface text of the current display interface, the enhanced audio recognition model is generated and fused with the general audio recognition model, the probability reference ratio is dynamically adjusted to improve the recognition accuracy.
Improve the audio recognition accuracy of the current display interface, ensure that the recognition results match the interface text, and reduce errors.
Smart Images

Figure CN114333832B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a data processing method, device, and readable storage medium. Background Art
[0002] Currently, with the increasing maturity of voice interaction technology, there are more and more products based on voice interaction. For example, intelligent terminal devices (such as smart phones) usually include many voice interaction products (such as voice interaction products like music player applications, video applications, navigation applications, mini-program applications, etc.). These voice interaction products installed in intelligent terminal devices are usually called third-party applications. The content, functions, etc. in these third-party applications are mainly customized by the creators of the third-party applications. Users can issue voice commands based on the content in the application interface of the third-party application, and the intelligent terminal device can perform corresponding actions based on the voice commands.
[0003] However, due to the diversity and rapid update of the content of third-party applications, the manufacturers of intelligent terminal devices cannot keep up in a timely manner. The intelligent terminal device cannot optimize the speech recognition system in a timely manner according to the updated content in the third-party application, resulting in poor recognition effects of the unupdated basic speech recognition system on the content in the third-party application, and there is an error between the actually executed action and the action expected by the user. Summary of the Invention
[0004] Embodiments of this application provide a data processing method, device, equipment, and readable storage medium, which can improve the accuracy of audio recognition.
[0005] On the one hand, an embodiment of this application provides a data processing method, including:
[0006] Obtain the interface text included in the current display interface, and obtain an enhanced audio recognition model generated based on the interface text;
[0007] When receiving the target audio data of the target object, perform audio recognition on the target audio data according to the probability reference ratio corresponding to the enhanced audio recognition model, the enhanced audio recognition model, and the general audio recognition model generated based on the general corpus, to obtain the initial selection probabilities respectively corresponding to one or more initial candidate texts, and determine the first target text from the one or more initial candidate texts according to the one or more initial selection probabilities;
[0008] If the first target text does not match the interface text, the probability reference ratio is adjusted. Based on the adjusted probability reference ratio, the enhanced audio recognition model, and the general audio recognition model, audio recognition is performed on the target audio data to obtain the target selection probabilities corresponding to one or more target candidate texts. The second target text is determined from the one or more target candidate texts according to the target selection probabilities.
[0009] If the second target text matches the interface text, the second target text is determined as the audio recognition result for the target audio data.
[0010] One aspect of the embodiments of the present application provides a data processing device, including:
[0011] A model acquisition module, configured to acquire the interface text included in the current display interface;
[0012] The model acquisition module is further configured to acquire an enhanced audio recognition model generated based on the interface text;
[0013] A selection probability determination module, configured to, when receiving the target audio data of the target object, perform audio recognition on the target audio data according to the probability reference ratio corresponding to the enhanced audio recognition model, the enhanced audio recognition model, and the general audio recognition model generated based on the general corpus, to obtain the initial selection probabilities corresponding to one or more initial candidate texts;
[0014] A target text determination module, configured to determine a first target text from the one or more initial candidate texts according to the one or more initial selection probabilities;
[0015] A ratio adjustment module, configured to adjust the probability reference ratio if the first target text does not match the interface text;
[0016] The target text determination module is further configured to perform audio recognition on the target audio data according to the adjusted probability reference ratio, the enhanced audio recognition model, and the general audio recognition model, to obtain the target selection probabilities corresponding to one or more target candidate texts, and determine a second target text from the one or more target candidate texts according to the target selection probabilities;
[0017] A result determination module, configured to determine the second target text as the audio recognition result for the target audio data if the second target text matches the interface text.
[0018] In one embodiment, the model acquisition module includes:
[0019] A text extraction unit, configured to extract the valid text information included in the interface text;
[0020] A normalization processing unit for normalizing the valid text information to obtain standard text information;
[0021] A text tokenization unit for tokenizing the standard text information to obtain text tokens, training a language processing model based on the text tokens, and determining the trained language processing model as an enhanced audio recognition model.
[0022] In one embodiment, the normalization processing unit includes:
[0023] A text filtering subunit for obtaining the auxiliary text identifiers included in the valid text information, deleting the auxiliary text identifiers, and obtaining filtered text information;
[0024] A character replacement subunit for obtaining the character information included in the filtered text information and the corresponding text expression information of the character information, and replacing the character information with the text expression information to obtain standard text information.
[0025] In one embodiment, the selection probability determination module includes:
[0026] A candidate word acquisition unit for obtaining the phoneme features corresponding to the sub-audio data at time T in the target audio data, and obtaining the target candidate text words that match the phoneme features at time T in the general corpus; i is a positive integer; i At time T in the target audio data, obtain the phoneme features corresponding to the sub-audio data, and in the general corpus, obtain the target candidate text words that match the phoneme features at time T; i At time T, i is a positive integer;
[0027] A first word probability determination unit for obtaining the historical candidate text words corresponding to the sub-audio data at time T in the target audio data, and determining the first word selection probability corresponding to the target candidate text words according to the historical candidate text words in the enhanced audio recognition model; T i-1 At time T in the target audio data, obtain the historical candidate text words corresponding to the sub-audio data, and in the enhanced audio recognition model, determine the first word selection probability corresponding to the target candidate text words according to the historical candidate text words; T i-1 Time T is the previous time of time T i At time;
[0028] A second word probability determination unit for determining the second word selection probability corresponding to the target candidate text words according to the historical candidate text words in the general audio recognition model;
[0029] A word combination unit for combining the historical candidate text words and the target candidate text words in chronological order to obtain one or more initial candidate texts;
[0030] A selection probability determination unit for determining the initial selection probabilities corresponding to one or more initial candidate texts according to the first word selection probability, the second word selection probability, and the probability reference ratio.
[0031] In one embodiment, the first word probability determination unit includes:
[0032] A corpus acquisition subunit for acquiring an interface corpus corresponding to the interface text; the interface corpus includes K text sample sentences associated with the interface text;
[0033] A statement statistics subunit for obtaining, from the K text sample sentences, the text sample sentences including the historical candidate text words as the set of text sample sentences to be statistically analyzed;
[0034] The statement statistics subunit is further configured to, in the set of text sample sentences to be statistically analyzed, use the text sample sentences in which the next sample text word of the historical candidate text word is the target candidate text word as the target text sample sentences;
[0035] The statement statistics subunit is further configured to obtain the total number of statements corresponding to the set of text sample sentences to be statistically analyzed and the target number corresponding to the target text sample sentences;
[0036] A word probability determination subunit for determining a first word selection probability according to the ratio between the target number and the total number of statements.
[0037] In one embodiment, the historical candidate text words include the historical candidate text word C j , and one or more initial candidate texts include the initial candidate text M constructed from the historical candidate text word C j and the target candidate text word; j is a positive integer; j ; j is a positive integer;
[0038] The selection probability determination unit includes:
[0039] A probability fusion subunit for fusing the first word selection probability and the second word selection probability according to a probability reference ratio to obtain a target fusion word selection probability corresponding to the target candidate text word;
[0040] A probability operation subunit for obtaining the historical fusion word selection probability corresponding to the historical candidate text word C j , performing an operation process on the historical fusion word selection probability and the target fusion word selection probability to obtain an initial selection probability corresponding to the initial candidate text M j ;
[0041] In one embodiment, the probability fusion subunit is further specifically configured to perform an operation process on the first word selection probability and the probability reference ratio to obtain a first operation word selection probability;
[0042] The probability fusion subunit is further specifically configured to obtain the ratio difference between the model fusion coefficient and the probability reference ratio, and perform an operation process on the ratio difference and the second word selection probability to obtain a second operation word selection probability;
[0043] The probability fusion subunit is further specifically configured to obtain the maximum operator selection probability from the first operator selection probability and the second operator selection probability, and determine the maximum operator selection probability as the target fusion word selection probability corresponding to the target candidate text word.
[0044] In one embodiment, the target text determination module includes:
[0045] The maximum probability acquisition unit is configured to obtain the maximum initial selection probability from one or more initial selection probabilities;
[0046] The target text determination unit is configured to determine the initial candidate text corresponding to the maximum initial selection probability among one or more initial candidate texts as the first target text.
[0047] In one embodiment, the ratio adjustment module includes:
[0048] The adjustment times acquisition unit is configured to obtain the current adjustment times corresponding to the probability reference ratio if the first target text does not match the interface text;
[0049] The growth step length determination unit is configured to obtain the unit ratio growth step length, and determine the current ratio growth step length corresponding to the probability reference ratio according to the current adjustment times and the unit ratio growth step length;
[0050] The ratio adjustment unit is configured to perform arithmetic processing on the probability reference ratio and the ratio growth step length to obtain the target probability reference ratio, and use the target probability reference ratio as the adjusted probability reference ratio.
[0051] In one embodiment, the data processing device further includes:
[0052] The statement set acquisition module is configured to acquire the set of valid text statements corresponding to the interface text;
[0053] The text matching module is configured to match the first target text with the set of valid text statements;
[0054] The text matching module is further configured to determine that the first target text matches the interface text if there is a valid text statement in the set of valid text statements that matches the first target text;
[0055] The text matching module is further configured to determine that the first target text does not match the interface text if there is no valid text statement in the set of valid text statements that matches the first target text.
[0056] In one embodiment, the data processing device further includes:
[0057] An audio result determination module, configured to determine the first target text as the audio recognition result for the target audio data if the first target text matches the interface text.
[0058] In one embodiment, the data processing device further includes:
[0059] A ratio matching module, configured to match the adjusted probability reference ratio with a ratio threshold;
[0060] A step execution module, configured to, if the adjusted probability reference ratio is less than the ratio threshold, execute the step of performing audio recognition on the target audio data according to the adjusted probability reference ratio, the enhanced audio recognition model, and the general audio recognition model to obtain target selection probabilities corresponding to one or more target candidate texts, and determining a second target text from the one or more target candidate texts according to the target selection probabilities;
[0061] A model cancellation module, configured to, if the adjusted probability reference ratio is greater than the ratio threshold, cancel the enhanced audio recognition model, and perform audio recognition on the target audio data according to the general audio recognition model to obtain the audio recognition result for the target audio data.
[0062] In one embodiment, the data processing device further includes:
[0063] A result sending module, configured to send the audio recognition result for the target audio data to the target terminal device, so that the target terminal device generates an action instruction based on the audio recognition result and executes the action instruction in the target display interface; the target terminal device is the terminal device corresponding to the current display interface; the target display interface includes the current display interface.
[0064] On the one hand, an embodiment of the present application provides a computer device, including: a processor and a memory;
[0065] The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the method in the embodiment of the present application.
[0066] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program includes program instructions, and when the program instructions are executed by the processor, the method in the embodiment of the present application is executed.
[0067] On the one hand, the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in one aspect of the embodiment of the present application.
[0068] In the embodiment of the present application, an enhanced audio recognition model is generated by obtaining the interface text in the current display interface. When recognizing the target audio data of the target object, the enhanced audio recognition model and the general audio recognition model (generated based on the general corpus) can be dynamically fused through the probability reference ratio corresponding to the enhanced audio recognition model, that is, the target audio data is jointly recognized by the enhanced audio recognition model and the general audio recognition model to obtain the first target text. If the first target text does not match the interface text in the current display interface, the target audio data can be recognized again by adjusting the probability reference ratio to obtain the second target text. By adjusting the probability reference ratio, the fusion ratio between the enhanced audio recognition model and the general audio recognition model can be adjusted to make it more reasonable, that is, it is more likely that the second target text can match the interface text in the current display interface with a greater probability. When the second target text matches the interface text of the current display interface, the second target text can be used as the audio recognition result. It should be understood that usually, the target audio data of the target object is generally emitted based on the application interface, and the enhanced audio recognition model is generated in real time by the interface text of the current display interface. Then the enhanced audio recognition model is highly targeted at the current display interface, and can more effectively and flexibly and accurately identify the recognition result of the target audio data according to the interface text of the current display interface and the general audio recognition model. Dynamically adjusting the probability reference ratio corresponding to the enhanced audio recognition model according to the matching result between the recognition result and the interface text can improve the unreasonable problem of the fusion ratio between the enhanced audio recognition model and the general audio recognition model, thereby improving the problem of low audio recognition accuracy caused by the unreasonable fusion ratio, and further improving the audio recognition accuracy of the current display interface. In summary, the present application can improve the audio recognition accuracy of the current display interface. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0070] Figure 1 is a network architecture diagram provided by an embodiment of the present application;
[0071] Figure 2 is an audio recognition scenario diagram provided by an embodiment of the present application;
[0072] Figure 3 It is a schematic flowchart of a data processing method provided by an embodiment of the present application;
[0073] Figure 4 It is a system flowchart provided by an embodiment of the present application;
[0074] Figure 5 It is a schematic flowchart of an audio recognition process provided by an embodiment of the present application;
[0075] Figure 6 It is a schematic diagram of an audio recognition scenario provided by an embodiment of the present application;
[0076] Figure 7 It is a system interaction diagram provided by an embodiment of the present application;
[0077] Figure 8 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0078] Figure 9 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0079] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0080] The embodiments of the present application relate to artificial intelligence technology (Artificial Intelligence, AI). For ease of understanding, the following will elaborate on artificial intelligence and its related concepts.
[0081] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to enable the machine to have the functions of perception, reasoning, and decision-making.
[0082] Artificial intelligence technology is an interdisciplinary subject that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0083] The solution provided by the embodiments of this application belongs to Natural Language Processing (NLP), Speech Technology, and Machine Learning (ML) under the field of artificial intelligence.
[0084] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph, and other technologies. For example, this application can perform natural language processing on text data.
[0085] The key technologies of Speech Technology include automatic speech recognition technology, speech synthesis technology, and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future. For example, this application can use automatic speech recognition technology to recognize speech data.
[0086] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. For example, in this application, machine learning can be used to train a speech recognition model so that it can be used to recognize speech data.
[0087] Please refer to Figure 1 , Figure 1 which is a network architecture diagram provided by an embodiment of the present application. As Figure 1 shown, the network architecture may include a service server 1000 and a cluster of terminal devices. The cluster of terminal devices may include one or more terminal devices, and the number of terminal devices will not be limited here. As Figure 1 shown, the multiple terminal devices may include terminal device 100a, terminal device 100b, terminal device 100c, …, terminal device 100n; as Figure 1 shown, terminal device 100a, terminal device 100b, terminal device 100c, …, terminal device 100n may be respectively network-connected to the service server 1000, so that each terminal device can perform data interaction with the service server 1000 through this network connection.
[0088] Each terminal device may be integrally installed with a target application. When the target application runs on each terminal device, it can perform data interaction with the service server 1000 as described above Figure 1 shown. Among them, the target application may include an application with functions of displaying data information such as text, images, audio, and video. For example, the application may be an entertainment application (such as a game application, a video application, etc.), which can be used for users to upload data, view data, etc. (such as watching videos); the application may also be a navigation application, which can be used for users to perform navigation and positioning, and so on. Of course, the application may also be other applications with functions of displaying data information, which will not be exemplified one by one here. These applications can all be applications with voice interaction functions. When the user runs the application using the terminal device, the user can input voice data, and the terminal device can perform corresponding actions in the application according to the voice data. It should be understood that the service server 1000 in the present application can collect service data through the application. For example, the service data may be voice data input by the user. The service server 1000 can identify the voice data to obtain text data, and then return the text data to the terminal device, and the terminal device can generate corresponding action instructions based on the text data and execute the corresponding action instructions in the application.
[0089] An embodiment of the present application can select one terminal device as the target terminal device from multiple terminal devices. The terminal device may include: smart phones, tablets, laptop computers, desktop computers, desktop computers, smart watches, intelligent vehicle terminals, smart home appliances (such as smart TVs, smart speakers, etc.), intelligent voice interaction devices, etc., intelligent terminals with multimedia data processing functions (for example, video data playback function, music data playback function), but not limited thereto. For example, an embodiment of the present application can use Figure 1The terminal device 100a shown is used as the target terminal device. The target application can be integrated in the target terminal device. At this time, the target terminal device can perform data interaction with the service server 1000 through the target application.
[0090] For example, when a user is using the target application (such as a navigation application) in the target terminal device and a certain page is displayed on the target terminal device, and at this time the user issues a voice data, then the page displayed on the target terminal device can be called the current display interface, and the voice data issued by the user can be called the target audio data. Further, the service server 1000 can detect and collect the target audio data of the user through the bound account (which can be called the target object) of the user in the target application. The service server 1000 can perform audio recognition (i.e., speech recognition) on the target audio data to obtain text data, and use this text data as the audio recognition result of the target audio data. The service server 1000 can return the audio recognition result to the target terminal device. The target terminal device can generate a corresponding action instruction based on the audio recognition result and execute the action instruction in the target application. For example, taking the target audio data as "Go to Crescent Bay" as an example, after the service server 1000 obtains the target audio data, it recognizes the text data "Go to Crescent Bay", and this text data "Go to Crescent Bay" can be used as the audio recognition result of the target audio data; subsequently, the service server 1000 can return the audio recognition result "Go to Crescent Bay" to the target terminal device, and the target terminal device can generate a route planning instruction and plan a route from the current location to the destination "Crescent Bay" according to the route planning instruction. The target terminal device can display the route, so that the user can go from the current location to Crescent Bay under the guidance of the route.
[0091] Among them, the process of the service server 1000 recognizing the target audio data can be as follows: first, obtain the interface text included in the current display interface, and then generate an enhanced audio recognition model according to the interface text included in the current display interface; subsequently, the service server 1000 can jointly recognize the target audio data based on the enhanced audio recognition model and the general audio recognition model (the audio recognition model generated based on the general corpus) to obtain the audio recognition result. For the specific implementation of jointly recognizing the target audio data based on the enhanced audio recognition model and the general audio recognition model to obtain the audio recognition result, reference can be made to the description of steps S101 - S104 in the corresponding embodiments Figure 3 described later.
[0092] It can be understood that the method provided by the embodiments of the present application can be executed by a computer device, which includes but is not limited to a terminal device or a service server. Among them, the service server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0093] Among them, the terminal device and the service server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.
[0094] Optionally, it can be understood that the above computer device (such as the above service server 1000, terminal device 100a, terminal device 100b, etc.) can be a node in a distributed system. Among them, the distributed system can be a blockchain system, and the blockchain system can be a distributed system formed by connecting the multiple nodes through network communication. Among them, the nodes can form a peer-to-peer (P2P) network, and the P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In a distributed system, any form of computer device, such as a service server or a terminal device, can become a node in the blockchain system by joining the peer-to-peer network. For ease of understanding, the concept of blockchain will be described below: Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm, mainly used to sort data in chronological order, encrypt it into a ledger, make it impossible to be tampered with and forged, and at the same time can perform data verification, storage, and update. When the computer device is a blockchain node, due to the tamper-proof and anti-forgery characteristics of the blockchain, the data (such as audio recognition results) in the present application can be made authentic and secure, so that the results obtained after performing relevant data processing based on these data are more reliable.
[0095] The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc. For ease of understanding, please refer to Figure 2 , Figure 2 which is a scene diagram of audio recognition provided by the embodiments of the present application. Among them, as Figure 2 shown, the service server 1000 can be the above-mentioned Figure 1 shown service server 1000, and as Figure 2 shown, the terminal device 100a can be the one in the aboveFigure 1 Terminal device 100a in the terminal device cluster corresponding to the embodiment.
[0096] As Figure 2 shown, the terminal device 100a can be the terminal corresponding to object a (such as user a). When user a opens the target application (such as a navigation application) through the terminal device 100a, the terminal device 100a can display the home page of the navigation application, and this home page can be called the current display interface of the terminal device 100a. When the interface of the terminal device 100a is in the current display interface, user a issues voice data "I want to go to aa Park in City A", and the terminal device 100a can use this voice data "I want to go to aa Park in City A" as the target audio data. The terminal device 100a can send the target audio data and the current display interface to the service server 1000.
[0097] Further, the service server 1000 can obtain the interface text data included in the current display interface and generate an enhanced audio recognition model for the current display interface according to the interface text data. Subsequently, the service server 1000 can dynamically fuse the enhanced audio recognition model for the current display interface with a general audio recognition model (an audio recognition model generated from a large amount of general text data, and this general recognition model can be applicable to all display interfaces with a larger recognition range), that is, identify the target audio data through the enhanced audio recognition model and the general recognition model together to obtain the recognized text data "I want to go to aa Park in City A". The service server 1000 can use this recognized text data "I want to go to aa Park in City A" as the audio recognition result of the target audio data, and the service server 1000 can return this audio recognition result to the terminal device 100a.
[0098] Further, after the terminal device 100a obtains the audio recognition result, it can obtain the current location of user a (which can be called the current location) through the bound account of user a in the navigation application. The terminal device 100a can plan a route from the current location to aa Park in City A (such as planning a walking route from the current location to aa Park in City A). As Figure 2 shown, after planning the walking route, the terminal device 100a can jump from the current display interface to the route display interface, and the terminal device 100a can display the walking route in this route display interface. At the same time, as Figure 2As shown, the terminal device 100 can also display the total walking distance (i.e., 1 kilometer) and the walking duration (e.g., 10 minutes) for this walking route; the terminal device 100a can also display a navigation start control for this walking route (such as the start walking navigation control shown in Figure 2). When the user performs a trigger operation on the start walking navigation control, the terminal device 100a can enter the navigation interface. User a can walk from the current location to the aa Park in City A under the guidance of this walking route. Among them, the above trigger operation can include contact operations such as clicking or long-pressing, and can also include non-contact operations such as voice or gesture. It will not be limited here.
[0099] Among them, for the specific implementation method of the service server 1000 to dynamically fuse the enhanced audio recognition model and the general audio recognition model, and jointly perform audio recognition on the target audio data through the enhanced audio recognition model and the general audio recognition model, reference can be made to the description of steps S101 - S104 in the subsequent Figure 3 corresponding embodiments.
[0100] Further, please refer to Figure 3 , Figure 3 is a schematic flowchart of a data processing method provided by an embodiment of the present application. Among them, this method can be executed by a terminal device (such as any terminal device in the terminal device cluster in the above Figure 1 corresponding embodiment, such as the terminal device 100a); this method can also be executed by a service server (such as the service server 1000 in the above Figure 1 corresponding embodiment); this method can also be jointly executed by the terminal device and the service server. Taking the example that this method is executed by the service server, as Figure 3 shown, the method process can at least include the following steps S101 - S104:
[0101] Step S101, obtain the interface text included in the current display interface, and obtain the enhanced audio recognition model generated based on the interface text.
[0102] In the present application, the current display interface can refer to the currently displayed application interface of a certain application installed on the terminal device (such as a video application, a music application, a navigation application, a social application, etc.). The current display interface can be obtained through the application programming interface (API) of each application. After obtaining the current display interface, the interface text in the current display interface (that is, the text content included in the current display interface) can be obtained, and an enhanced audio recognition model can be constructed according to the interface text in the current display interface.
[0103] The specific method for constructing an enhanced audio recognition model based on interface text may be: valid text information included in the interface text may be extracted; then, the valid text information may be standardized to obtain standard text information; then, the standard text information may be segmented to obtain text segmentation, and a language processing model may be trained based on the text segmentation, and the trained language processing model may be determined as an enhanced audio recognition model.
[0104] Among them, the specific method for standardizing the valid text information and obtaining the standard text information may be: the auxiliary text identifier included in the valid text information may be obtained, the auxiliary text identifier may be deleted, and the filtered text information may be obtained; then, the character information included in the filtered text information and the text expression information corresponding to the character information may be obtained, the character information may be replaced with the text expression information, and the standard text information may be obtained.
[0105] It should be understood that the valid text content (which may be referred to as valid text information) in the current interface text can be extracted through the API, and then the valid text information can be standardized, wherein the specific processing method of the standardization processing may be: first, the silent identification information (i.e., unpronounced symbols, such as punctuation marks, which may be referred to as auxiliary text identification) in the valid text information may be obtained, and these silent identification information (i.e., auxiliary text identification) may be deleted, thereby obtaining the filtered text information; then, the character information included in the filtered text information may also be obtained (the character information here may refer to the pronunciation character information of non-Chinese characters, such as Arabic numeral information such as 1, 2, 3, 4, etc.), and the text expression information corresponding to these character information may be obtained, and the text expression information here may refer to the Chinese text information of one. For example, the text expression information corresponding to the Arabic numeral 1 may be the Chinese text "一"; for example, the text expression information corresponding to the English word "everyday" may be the Chinese text "天天" (i.e., the translated text). These character information may be replaced with the corresponding text expression information, thereby obtaining the standard text information.
[0106] Further, the standard text information can be segmented to obtain text segments. For example, for the text information "Xiaoming is seventeen years old", it may include the phrases "Xiaoming", "seventeen", and "years old". Then, the extraction can be carried out based on the phrases as units, and the extracted text segments include "Xiaoming", "seventeen", and "years old". Further, a language processing model can be trained according to the text segments corresponding to the interface text. That is to say, the text segments corresponding to the interface text can be used as training data to train the language processing model. Here, the language processing model can refer to a statistical language model (LanguageModel, LM). The modeling goal of the statistical language model is to learn the distribution of the training data, that is, how to more accurately predict the future words under the condition of a given correct historical word sequence (that is: using limited training data to establish a probability distribution model that better conforms to the real data distribution). The statistical language model can include grammar language models (N gram LM), language models based on neural networks (such as Long Short-Term Memory (LSTM) models, Convolutional Neural Networks (CNN), autoregressive models (Transformer)), etc. There will be no further examples here. That is to say, the language processing model here can refer to an N-gram grammar language model or a language model based on a neural network, and there will be no restrictions here. The language processing model can be trained through the text segments corresponding to the interface text, and the trained language processing model can be used as the enhanced audio recognition model corresponding to the current display interface.
[0107] In step S102, when the target audio data of the target object is received, audio recognition is performed on the target audio data according to the probability reference ratio corresponding to the enhanced audio recognition model, the enhanced audio recognition model, and the general audio recognition model generated based on the general corpus, to obtain the initial selection probabilities corresponding to one or more initial candidate texts, and the first target text is determined from the one or more initial candidate texts according to the one or more initial selection probabilities.
[0108] In this application, the target object may refer to the bound account of the target user in the application. When the application interface displayed on the terminal device is the current display interface, the target user can send voice data (such as a voice command), and this voice data can be referred to as target audio data. The terminal device can obtain this target audio data, and the terminal device can associate the target audio data with the bound account (target object) (that is to say, the target audio data is actually the voice data corresponding to the target user), and send it to the business server. When the business server receives the target audio data, it can obtain a general audio recognition model generated based on a general corpus. Here, the general audio recognition model can be of the same model type as the enhanced audio recognition model (for example, both are N-gram language models). The enhanced audio recognition model is trained using the interface text segmentation of the current display interface as training data. Then, the enhanced audio recognition model has pertinence in recognizing the current display interface; while the general audio recognition model can be trained using the general corpus as training data. The general corpus can include the interface text segmentation of the current display interface and can also include a large number of text segmentations other than the interface text segmentation of the current display interface (such as the interface text segmentation of other non-current display interfaces). For the recognition of the current display interface, the general audio recognition model does not have pertinence and has generality.
[0109] Furthermore, the business server can obtain the probability reference ratio corresponding to the enhanced audio recognition model, and during the recognition process of the target audio data, fuse the enhanced audio recognition model and the general audio recognition model according to this probability reference ratio (that is, fuse the model output results of the enhanced audio recognition model and the general audio recognition model). Thus, by combining the model output results of the enhanced audio recognition model and the general audio recognition model, a more accurate target text (which can be called the first target text) of the target audio data can be obtained. The specific method can be: recognize the target audio data according to the probability reference ratio, the enhanced audio recognition model, and the general audio recognition model, and thus initial selection probabilities corresponding to one or more initial candidate texts can be obtained. The first target text corresponding to the target audio data can be determined from one or more initial candidate texts according to one or more initial selection probabilities. Among them, the probability reference ratio can refer to the fusion ratio between the enhanced audio recognition model and the general audio recognition model. The upper and lower limits of this fusion ratio can be specified manually, and the probability reference ratio needs to be within the range composed of this lower limit and upper limit. For example, taking the lower limit value of 0.3 and the upper limit value of 0.8 as an example, the range composed of this lower limit and upper limit can be [0.3, 0.8], and the probability reference ratio can be within 0.3 to 0.8 (including 0.3 and 0.8).
[0110] That is to say, through the probability reference ratio, the enhanced audio recognition model and the general audio recognition model jointly recognize the target audio data, and thus one or more selection probabilities corresponding to the initial candidate texts can be obtained (which can be understood as the score values corresponding to each recognized text). The initial candidate text corresponding to the maximum selection probability among these selection probabilities can be used as the target text corresponding to the target audio data (that is: obtain the maximum initial selection probability among one or more initial selection probabilities; determine the initial candidate text corresponding to the maximum initial selection probability among one or more initial candidate texts as the first target text). Among them, for the specific implementation method of jointly recognizing the target audio data by the enhanced audio recognition model and the general audio recognition model through the probability reference ratio to obtain one or more initial selection probabilities corresponding to the initial candidate texts, reference can be made to the description in the subsequent Figure 5 corresponding embodiments.
[0111] Step S103, if the first target text does not match the interface text, then adjust the probability reference ratio. According to the adjusted probability reference ratio, the enhanced audio recognition model, and the general audio recognition model, perform audio recognition on the target audio data to obtain one or more target selection probabilities corresponding to the target candidate texts, and determine the second target text among one or more target candidate texts according to the target selection probabilities.
[0112] In this application, after obtaining the first target text corresponding to the target audio data through the enhanced audio recognition model and the general audio recognition model, the first target text can be matched with the interface text of the current display interface to determine whether the first target text matches the interface text of the current display interface. The specific implementation method can be: obtain the set of valid text statements corresponding to the interface text; subsequently, match the first target text with the set of valid text statements; if there is a valid text statement in the set of valid text statements that matches the first target text, it can be determined that the first target text matches the interface text; and if there is no valid text statement in the set of valid text statements that matches the first target text, it can be determined that the first target text does not match the interface text.
[0113] Among them, the method of matching the first target text with the set of valid text statements can be matched by similarity. For example, the first target text can be input into any encoder capable of semantic encoding. The encoder can be any deep learning model capable of semantic encoding, such as: Convolutional Neural Network (CNN) model, Long Short-Term Memory (LSTM) model, combined model of Long Short-Term Memory and attention mechanism (LSTM+Attention), Bidirectional Encoder Representations from Transformers (BERT), and so on. Here, no further examples will be given one by one. Through the encoder, the semantic vector expression features corresponding to the first target text can be output; similarly, each valid text statement in the set of valid text statements can be input into the encoder, and through the encoder, the statement vector expression features corresponding to each valid text statement can also be output.
[0114] Furthermore, the statement vector expression features corresponding to each valid text statement can be respectively matched with the semantic vector expression features corresponding to the first target text (calculate the similarity between the two vectors). If the similarity between a certain statement vector expression feature and the semantic vector expression feature is greater than (or equal to) the similarity threshold, it can be determined that the first target text is matched with the valid text statement corresponding to this statement vector expression feature, and then it can be determined that the first target text is matched with the interface text; if the similarity between each statement vector expression feature and the semantic vector expression feature is less than the similarity threshold, it can be determined that the first target text does not match any valid text statement, and then it can be determined that the first target text does not match the interface text. Among them, the method of calculating similarity can be to calculate the cosine similarity between two vectors, and use the cosine similarity as the similarity between the two vectors; in addition, the method of calculating similarity can also be other methods of calculating similarity, such as calculating the vector distance and using the vector distance as the similarity between the two vectors. This application does not limit this.
[0115] Further, after matching, if it is determined that the first target text matches the interface text, the first target text can be directly determined as the audio recognition result for the target audio data. If it is determined that the first target text does not match the interface text, the probability reference ratio (i.e., the fusion ratio, also referred to as the fusion weight) can be adjusted to make the fusion ratio more reasonable by adjusting the probability reference ratio. The specific implementation of adjusting the probability reference ratio can be as follows: if the first target text does not match the interface text, the current adjustment times corresponding to the probability reference ratio can be obtained; subsequently, the unit ratio increase step can be obtained, and the current ratio increase step corresponding to the probability reference ratio can be determined based on the current adjustment times and the unit ratio increase step; subsequently, the probability reference ratio and the ratio increase step can be processed arithmetically, and thus the target probability reference ratio can be obtained, and the target probability reference ratio can be used as the adjusted probability reference ratio.
[0116] Among them, it can be understood that here we can set the adjustment method to an increasing adjustment method (of course, it can also be other adjustment methods, and the increasing adjustment method is only an example), and the unit increase value (i.e., the unit ratio increase step) for each adjustment of the probability reference ratio can be set in advance, and the current ratio increase value (i.e., the current ratio increase step, for example, the product of the current adjustment times and the unit ratio increase step can be used as the current ratio increase step) can be determined based on the current adjustment times and the unit ratio increase step. For example, taking the current adjustment times as 2 times (i.e., the probability reference ratio has been adjusted 2 times) and the unit ratio increase step as 0.1 as an example, the current ratio increase step of the probability reference ratio can be: 2×0.1 = 0.2. Subsequently, the probability reference ratio and the current ratio increase step can be processed arithmetically (for example, a multiplication operation is performed), and the operation result (such as the product) can be used as the target probability reference ratio. Adjusting the probability reference ratio means replacing the probability reference ratio with the target probability reference ratio.
[0117] Optionally, in a feasible embodiment, the adjustment of the probability reference ratio may not be based on the current adjustment times. For example, taking the increasing adjustment as an example, the adjustment amount for each time can be set in advance. When adjusting each time, the current probability reference ratio can be added to the adjustment amount set each time to obtain the adjusted probability reference ratio. The method of adjusting the probability reference ratio is of course not limited to this, and it can also be other methods that can adjust and change the probability reference ratio (such as a numerical range can be given (such as 0.01 - 0.2), and the service server can randomly select a value from this numerical range as the adjustment amount each time it adjusts, and then add the probability reference ratio and the selected adjustment amount to obtain the adjusted probability reference ratio), and this application does not limit it.
[0118] It should be noted that a size range (which can be composed of the aforementioned upper limit value and lower limit value) can be set for the probability reference ratio in this application, and the probability reference ratio needs to be within this size range. Then, after adjusting the probability reference ratio to obtain the adjusted probability reference ratio, the adjusted probability reference ratio can be matched with the upper limit value (which can be called the ratio threshold). If the adjusted probability reference ratio is less than the ratio threshold, it can be indicated that the adjusted probability reference ratio meets this size range. At this time, based on the adjusted probability reference ratio, the enhanced audio recognition model, and the general audio recognition model, audio recognition can be performed on the target audio data to obtain the target selection probabilities corresponding to one or more target candidate texts, and the second target text can be determined from one or more target candidate texts according to the target selection probabilities. If the adjusted probability reference ratio is greater than (or equal to) the ratio threshold, it can be indicated that the adjusted probability reference ratio exceeds the specified size range. At this time, it can be stated that the probability reference ratio has been adjusted to the upper limit value, and the target texts recognized by the two models still cannot match the interface text in the current display interface. Then, it can be shown that the target audio data of the target object is not the audio data for the current display interface, but the audio data for a non-current display interface (for example, when the application is a music player application and the current display interface is a music recommendation interface, which includes 10 songs recommended to the target object. However, the target audio data of the target object is "I want to go to aa park. Help me plan the route." Obviously, this target audio data has nothing to do with the current display interface. Although the target user sees the music recommendation interface, the target audio data of the target user is related to a non-current display interface. Then, the recognition result of this target audio data cannot be matched with the current display interface. Even if the probability reference ratio (i.e., the fusion ratio) is increased to the upper limit value, the target text still cannot be matched with the current display interface).
[0119] It should be understood that when the adjusted probability reference ratio is greater than the ratio threshold, at this time, it can be determined that the target audio data has nothing to do with the current display interface and is related to a non-current display interface. Then, in order to ensure the general recognition effect of the target audio data, the enhanced audio recognition model for the current display interface can be cancelled. Subsequently, only the general audio recognition model is used to perform audio recognition on the target audio data, and the target text recognized by the general audio recognition model is used as the recognition result of the target audio data.
[0120] It can be understood that if the adjusted probability reference ratio is less than the ratio threshold, then when the second target text of the target audio data is recognized based on the adjusted probability reference ratio, the enhanced audio recognition model, and the general audio recognition model, the second target text can be matched with the interface text again.
[0121] In step S104, if the second target text matches the interface text, the second target text is determined as the audio recognition result for the target audio data.
[0122] In this application, as can be seen from the above, the second target text can be matched with the interface text. If the second target text matches the interface text, the second target text can be determined as the audio recognition result for the target audio data; if the second target text does not match the interface text, the probability reference ratio can be adjusted continuously, and then a new adjusted probability reference ratio is obtained. When the new adjusted reference ratio is within the specified range, the target audio data can be recognized based on the new adjusted probability reference ratio to obtain a new target text (which can be called the third target text), and then it can be matched with the interface text until it matches the interface text or the adjusted probability reference ratio exceeds the specified range.
[0123] It can be understood that after determining the audio recognition result corresponding to the target audio data, the service server can send the audio recognition result to the target terminal device (i.e., the terminal device corresponding to the current display interface), and the target terminal device can generate an action instruction based on the audio recognition result and execute the action instruction on the target display interface. For example, if the target audio data is related to the current display interface, the target terminal device can generate an action instruction for the current display interface and execute the action instruction; if the target audio data is not related to the current display interface, an action instruction for a non-current display interface can be generated and executed on the non-current display interface.
[0124] In the embodiment of the present application, an enhanced audio recognition model for the current display interface can be generated through the current display interface of the terminal device. The enhanced audio recognition model is generated according to the real-time updated content in the current display interface, which is highly targeted and can be combined with the general audio recognition model to more accurately recognize the target audio data related to the current display interface. At the same time, when the recognized text cannot be matched with the interface text, by dynamically adjusting the model fusion weight (i.e., the probability reference ratio), the problem of inaccurate recognized text caused by unreasonable fusion weight settings can be improved, and the accuracy of audio recognition can be further improved. In addition, if the recognized text still fails to match any valid statement in the current display interface after the weight adjustment reaches the upper limit value, the enhanced audio recognition model for the current display interface can be canceled to only use the general audio recognition model to recognize the target audio data, which can ensure the general recognition effect of the audio data. That is to say, the present application can strengthen the fusion effect between the enhanced audio recognition model and the general audio recognition model by dynamically adjusting the fusion weight, improve the recognition accuracy of the application interface content of the third-party application in the terminal device; at the same time, when the audio data is related to the current display interface, by dynamically adjusting the fusion weight, the problem that the recognized text fails to be corrected in time due to unreasonable weight settings can be improved, and the recognition accuracy can be further improved; at the same time, by setting an upper limit for the weight and canceling the enhanced audio recognition model when the upper limit is reached, the general recognition effect can be ensured when the audio data is not related to the content of the current display interface. In summary, the present application can improve the accuracy of audio recognition.
[0125] For ease of understanding the overall process, please refer to Figure 4 , Figure 4 which is a system flowchart provided by the embodiment of the present application. As Figure 4 shown, this process can at least include the following steps S501 - step S507:
[0126] Step S501, the service server fuses the general audio recognition model and the enhanced audio recognition model using the probability reference ratio to obtain the recognized text of the target audio data.
[0127] Step S502, return the recognized text to the terminal device so that the terminal device can match it with the interface text of the current display interface.
[0128] Step S503, determine whether the recognized text matches the current interface text.
[0129] Specifically, the recognized text can be matched with the current interface text to determine whether the recognized text matches the interface text. If it matches, the subsequent step S504 can be executed; if it does not match, the subsequent step S505 can be executed.
[0130] Step S504: Use the recognized text as the audio recognition result.
[0131] Specifically, after obtaining the audio recognition result of the target audio data, an action instruction can be generated based on the audio recognition result (such as an instruction to display the audio recognition result), and the terminal device can execute the corresponding instruction according to the action instruction (such as displaying the audio recognition result on the current display interface).
[0132] Step S505: Adjust the probability reference ratio.
[0133] Specifically, if they do not match, the probability reference ratio can be adjusted (such as increasing the adjustment) to obtain the adjusted probability reference ratio.
[0134] Step S506: Determine whether the adjusted probability reference ratio is greater than the upper limit value.
[0135] Specifically, the adjusted probability reference ratio can be matched with the upper limit value to determine whether the adjusted probability reference ratio exceeds the upper limit value. If the adjusted probability reference ratio is greater than the upper limit value, subsequent step S507 can be executed; if the adjusted probability reference ratio is less than the upper limit value, the process can return to step S501, that is, using the adjusted probability reference ratio to fuse the general audio recognition model and the enhanced audio recognition model, recognize the new recognized text of the target audio data, and repeat the subsequent steps.
[0136] Step S507: Deactivate the enhanced audio recognition model and only use the general audio recognition model for audio recognition to obtain the recognized text.
[0137] Specifically, after the adjusted probability reference ratio is greater than the upper limit value, the enhanced audio recognition model can be deactivated, and only the general audio recognition model is used to recognize the target audio data to obtain the recognized text, ensuring its general recognition effect.
[0138] Among them, for the specific implementation manners of steps S501 - S507, reference can be made to the descriptions in the corresponding embodiments above Figure 3 and will not be elaborated here.
[0139] In an embodiment of the present application, an enhanced audio recognition model for the current display interface can be generated through the current display interface of the terminal device. The enhanced audio recognition model is generated based on the real-time updated content in the current display interface, which is highly targeted and can be combined with the general audio recognition model to more accurately recognize the target audio data related to the current display interface. At the same time, when the recognized text cannot be matched with the interface text, by dynamically adjusting the model fusion weight (i.e., the probability reference ratio), the problem of inaccurate recognized text caused by unreasonable fusion weight settings can be improved, and the accuracy of audio recognition can be further improved. In addition, after the weight adjustment reaches the upper limit value and the recognized text still fails to match any valid statement in the current display interface, the enhanced audio recognition model for the current display interface can be canceled to recognize the target audio data only using the general audio recognition model, which can ensure the general recognition effect of the audio data. That is to say, the present application can strengthen the fusion effect between the enhanced audio recognition model and the general audio recognition model by dynamically adjusting the fusion weight, and improve the recognition accuracy of the application interface content of the third-party application in the terminal device. At the same time, when the audio data is related to the current display interface, by dynamically adjusting the fusion weight, the problem that the recognized text fails to be corrected in time due to unreasonable weight settings can be improved, and the recognition accuracy can be further improved. At the same time, by setting an upper limit for the weight and canceling the enhanced audio recognition model when the upper limit is reached, the general recognition effect can be ensured when the audio data is not related to the content of the current display interface. In summary, the present application can improve the accuracy of audio recognition.
[0140] Further, please refer to Figure 5 , Figure 5 which is a schematic diagram of an audio recognition process provided by an embodiment of the present application. Among them, this process can correspond to the process of Figure 3 in the corresponding embodiment, for performing audio recognition on the target audio data according to the probability reference ratio, the enhanced audio recognition model, and the general audio recognition model to obtain the initial selection probabilities corresponding to one or more initial candidate texts. As Figure 5 shown, this process can at least include the following steps S601 - step S605:
[0141] Step S601, obtain the phoneme features corresponding to the sub-audio data at time T i in the target audio data, and obtain the target candidate text words that match the phoneme features at time T i in the corpus. i is a positive integer.
[0142] Specifically, the target audio data can be feature extracted to obtain the phoneme features corresponding to the sub-audio data at each moment in the target audio data. Subsequently, the text words matching the phoneme features can be obtained in the general corpus in chronological order (such as the order of time) as candidate text words. For example, if the target audio data is the audio data "Have you eaten?", then the sub-audio data can include the sub-audio data "you", the sub-audio data "eat", the sub-audio data "rice", the sub-audio data "le", and the sub-audio data "may"; subsequently, the phoneme features corresponding to the sub-audio data "you" at the earliest moment can be obtained first, and then the text words matching it (such as the text word "you", the text word "ni", and the text word "ni") can be obtained in the general prediction library, and these text words can be used as candidate text words for the sub-audio data "you".
[0143] Step S602, obtaining the T i-1 The historical candidate text words corresponding to the sub-audio data at time T, and the first word selection probability corresponding to the target candidate text word is determined in the enhanced audio recognition model according to the historical candidate text words; i-1 Time is T i The moment before the moment.
[0144] Specifically, taking the target audio data as the audio data "Have you eaten?", the time order of its sub-audio data is as follows: the moment of the sub-audio data "you" is earlier than the moment of the sub-audio data "eat"; the moment of the sub-audio data "eat" is earlier than the moment of the sub-audio data "rice"; the moment of the sub-audio data "rice" is earlier than the moment of the sub-audio data "read"; the moment of the sub-audio data "read" is earlier than the moment of the sub-audio data "me?" Then the candidate text words corresponding to the sub-audio data "you" (such as you, you, Ni, Ni) can be obtained first, and the enhanced audio recognition model can determine the word selection probability corresponding to each candidate text word based on the number of occurrences of each candidate text word in the interface corpus corresponding to the interface text; similarly, the general audio recognition model can determine the word selection probability corresponding to each candidate text word based on the number of occurrences of each candidate text word in the interface text; the word selection probabilities of each candidate text word output by the two models can be fused through the probability reference ratio, so as to obtain a fused word selection probability corresponding to a candidate text word; and the candidate text words for which the fused word selection probability has been determined (that is, the candidate text word "you", the candidate text word "you", the candidate text word "Ni", the candidate text word "Ni") can be called historical candidate text words, and the sub-audio data that has been determined by probability (that is, the sub-audio data "you") can be called historical sub-audio data.
[0145] Further, the remaining sub-audio data can be recognized. After obtaining the target candidate text words corresponding to a certain sub-audio data, the word selection probability corresponding to the target candidate text words can be determined based on the historical candidate text words (which can be called prior words). For example, for the sub-audio data "chi", after obtaining that the target candidate text words corresponding to the sub-audio data "chi" include "chi" and "chi", the enhanced audio recognition model can determine the word selection probabilities corresponding to the target candidate text words "chi" and "chi" based on the historical candidate text words "ni", "nin", "ni", "ni", and the interface corpus; similarly, the general audio recognition model can also determine the word selection probabilities corresponding to the target candidate text words "chi" and "chi" based on the historical candidate text words and the general corpus; subsequently, the two word selection probabilities can be fused through a probability reference ratio to obtain the word fusion probability corresponding to each target candidate text word. That is to say, for each candidate word whose probability is to be determined, the probability is determined based on the prior words.
[0146] Taking the target candidate text words that match the phoneme features at time T i as an example, for the specific implementation method of outputting the first word selection probability corresponding to the target candidate text words through the enhanced audio recognition model and the historical candidate text words: the interface corpus corresponding to the interface text can be obtained; where, the interface corpus includes K text sample sentences associated with the interface text; the text sample sentences including the historical candidate text words can be obtained from the K text sample sentences as the set of text sample sentences to be counted; subsequently, in the set of text sample sentences to be counted, the text sample sentences whose next sample text word after the historical candidate text word is the target candidate text word can be used as the target text sample sentences; the total number of sentences corresponding to the set of text sample sentences to be counted and the target number corresponding to the target text sample sentences can be obtained; the first word selection probability is determined according to the ratio between the target number and the total number of sentences.
[0147] Among them, the K text sample sentences can refer to the aforementioned set of valid text sentences. For easy understanding, please refer to Table 1 together. Table 1 can refer to the K text sample sentences. The K text sample sentences can include 5 text sample sentences. For example, the text sample sentence "I want to go to the gym today", the text sample sentence "Today is sunny", the text sample sentence "Exercise every day", the text sample sentence "Drinking protein powder can build muscle", the text sample sentence "Want to drink water".
[0148] Table 1
[0149]
[0150]
[0151] The text words corresponding to the five text sample sentences shown in Table 1 can form an interface corpus. Taking the target audio data "I want to drink protein powder today" as an example, the sub-audio data "today" at time T0 can be obtained first. At this time, the phoneme features of the sub-audio data "today" at time T0 can be obtained, and the candidate text word "today" matching the phoneme features can be obtained in the interface corpus. Subsequently, the number of occurrences of the text word "today" in the interface corpus can be counted (i.e., 2 times). The total number of occurrences of the text words in the interface corpus is 29, so the probability of selecting the first word corresponding to the candidate text word "today" can be 2 / 29. Further, then, the sub-audio data at the next moment after T0 (such as T1) can be obtained as "天", and the phoneme features of the sub-audio data "天" can be obtained at this time, and the candidate text word "天" matching the phoneme features can be obtained in the interface corpus. Then, the candidate text word "今" whose word selection probability has been determined can be used as a historical candidate text word, and the text sample sentence containing the historical candidate text word "今" can be obtained from the above 5 text sample sentences as the text sample sentence to be counted. That is, "I want to go to the gym today" and "Today is sunny" can form a set of text sample sentences to be counted.
[0152] Furthermore, the next sample text word in the text sample sentence to be counted can be obtained. It can be seen that in the two text sample sentences to be counted, the sample text word located at the next position of the historical candidate text word is "天", so "I want to go to the gym today" and "Today is a sunny day" can both be used as target text sample sentences. Then the total number of sentences in the text sample sentence set to be counted is 2, and the target number of target text sample sentences is also 2. The ratio between the target number and the total number of sentences is 1, and the probability of selecting the first word of the target candidate text word is 1.
[0153] For ease of understanding, further, the sub-audio data at the next moment after T1 (such as T2) can be obtained as "想", and the phoneme features of the sub-audio data "想" can be obtained at this time, and the candidate text word "想" matching the phoneme features can be obtained in the interface corpus. Subsequently, the candidate text words "今" and "天" whose word selection probabilities have been determined can be used as historical candidate text words, and the historical candidate text word "天" with the latest moment can be obtained. In the above 5 text sample sentences, the text sample sentence containing the historical candidate text word "天" can be obtained as the text sample sentence to be counted. That is, "I want to go to the gym today", "Today is sunny", and "I have to exercise every day" can be composed of a set of text sample sentences to be counted.
[0154] Further, the next sample text word in the sample text sentence to be counted can be obtained. It can be seen that in the two sample text sentences to be counted, the sample text words at the next position after the historical candidate text word include "want", "is", "blank (i.e., the character 'tian' is the last text in the sentence)", "all". Then, the sentences "want to go to the gym today", "it is sunny today", and "exercise every day" can all be used as target text sample sentences. It can be determined that the total number of sentences corresponding to the set of sample text sentences to be counted is 4 (because the sample text sentence "it is sunny today" includes two 'tian' characters, and when counting the total number, the number of this sample text sentence to be counted needs to be determined as 2). The target number of the target text sample sentence (the sentence with the character 'xiang' immediately following the character 'tian') is 1. The ratio between this target number and the total number of sentences is 1 / 4. Then, the first-word selection probability of this target candidate text word is 1 / 4. Similarly, the target candidate text words corresponding to the remaining sub-audio data can also be determined, as well as the first-word selection probability corresponding to this target candidate text word.
[0155] Step S603: Determine the second-word selection probability corresponding to the target candidate text word in the general audio recognition model based on the historical candidate text word.
[0156] Specifically, the general audio recognition model can determine the second-word selection probability corresponding to the target candidate text word based on the general corpus and the historical candidate text word. For the determination method of the second-word selection probability, reference can be made to the determination method of the first-word selection probability, which will not be elaborated here.
[0157] Step S604: Combine the historical candidate text word and the target candidate text word in chronological order to obtain one or more initial candidate texts.
[0158] Specifically, the historical candidate text word and the target candidate text word can be combined in chronological order (such as the order of time from early to late), and one or more initial candidate texts can be obtained therefrom. For example, as described above, after determining the historical candidate text words "jin" and "tian", the historical candidate text word and the target candidate text word "xiang" can be combined in chronological order to obtain the initial candidate text "today want".
[0159] Step S605: Determine the initial selection probabilities corresponding to one or more initial candidate texts according to the first-word selection probability, the second-word selection probability, and the probability reference ratio.
[0160] Specifically, let the historical candidate text word include the historical candidate text word C j , and one or more of the initial candidate texts include the initial candidate text M constructed from the historical candidate text word C j and the target candidate text word jTake (j is a positive integer) as an example; for the specific implementation of determining the initial selection probabilities corresponding to one or more initial candidate texts according to the first word selection probability, the second word selection probability, and the probability reference ratio, it can be as follows: The first word selection probability and the second word selection probability can be fused according to the probability reference ratio to obtain the target fused word selection probability corresponding to the target candidate text word; subsequently, the historical fused word selection probability corresponding to the historical candidate text word C j can be obtained, and the historical fused word selection probability and the target fused word selection probability can be processed through an operation to obtain the initial selection probability corresponding to the initial candidate text M j .
[0161] Among them, the specific implementation of fusing the first word selection probability and the second word selection probability according to the probability reference ratio to obtain the target fused word selection probability corresponding to the target candidate text word can be: The first word selection probability and the probability reference ratio can be processed through an operation to obtain the first operation word selection probability (for example, the first word selection probability and the probability reference ratio can be multiplied to obtain the product as the first operation word selection probability); subsequently, the ratio difference between the model fusion coefficient (which can be a manually specified value, such as the value 1) and the probability reference ratio can be obtained, and the ratio difference and the second word selection probability can be processed through an operation to obtain the second operation word selection probability (for example, the second word selection probability and the ratio difference can be multiplied to obtain the product as the second operation word selection probability); subsequently, the maximum operation word selection probability can be obtained from the first operation word selection probability and the second operation word selection probability, and the maximum operation word selection probability can be determined as the target fused word selection probability corresponding to the target candidate text word.
[0162] The above is just one way of fusing word selection probabilities. For the probability fusion method, it can also be other fusion methods. For example, the first operation word selection probability and the second operation word selection probability can also be multiplied, added, etc., and the obtained operation result can be used as the target fused word selection probability.
[0163] In the embodiment of the present application, an enhanced audio recognition model for the current display interface can be generated through the current display interface of the terminal device. The enhanced audio recognition model is generated based on the real-time updated content in the current display interface, which is highly targeted and can be combined with the general audio recognition model to more accurately recognize the target audio data related to the current display interface. At the same time, when the recognized text cannot be matched with the interface text, by dynamically adjusting the model fusion weight (i.e., the probability reference ratio), the problem of inaccurate recognized text caused by unreasonable fusion weight settings can be improved, and the accuracy of audio recognition can be further improved. In addition, if the recognized text still fails to match any valid statement in the current display interface after the weight adjustment reaches the upper limit value, the enhanced audio recognition model for the current display interface can be cancelled to only use the general audio recognition model to recognize the target audio data, which can ensure the general recognition effect of the audio data. That is to say, the present application can strengthen the fusion effect between the enhanced audio recognition model and the general audio recognition model by dynamically adjusting the fusion weight, and improve the recognition accuracy of the application interface content of the third-party application in the terminal device. At the same time, when the audio data is related to the current display interface, by dynamically adjusting the fusion weight, the problem that the recognized text fails to be corrected in time due to unreasonable weight settings can be improved, and the recognition accuracy can be further improved. At the same time, by setting an upper limit for the weight and cancelling the enhanced audio recognition model when the upper limit is reached, the general recognition effect can be ensured when the audio data is not related to the content of the current display interface. In summary, the present application can improve the accuracy of audio recognition.
[0164] It can be understood that the above - mentioned method of fusing the enhanced audio recognition model and the general audio recognition model through the probability reference ratio is merely an exemplary fusion method. Of course, the method of fusing the enhanced audio recognition model and the general audio recognition model through the probability reference ratio is not limited to this. For example, when the enhanced audio recognition model and the general audio recognition model are N - gram language models, the recognition method of the N - gram language model can be adopted to determine the word selection probabilities corresponding to the target candidate text words respectively, and then the probability reference ratio is used to fuse the word selection probabilities corresponding to the target candidate text words respectively to obtain the fused word selection probabilities corresponding to each target candidate text word. Subsequently, each candidate text word is combined to obtain multiple candidate text paths (which can also be called candidate texts), and then the fused word selection probabilities corresponding to each word in the candidate text path are multiplied to obtain the selection probability corresponding to each candidate text path. The candidate text path corresponding to the maximum selection probability can be used as the target text jointly recognized by the two models. When the enhanced audio recognition model and the general audio recognition model are neural - network - based language models, the recognition method of the neural - network - based language model can be adopted to determine the word selection probabilities corresponding to the target candidate text words respectively (for example, the phoneme features of the sub - audio data can be vector - feature - matched with each text word in the corpus, so that the matching probability of each text word can be determined, and the text word corresponding to the maximum matching probability can be used as the recognition text of the sub - audio data. The recognition texts corresponding to each sub - audio data can be combined to obtain the audio recognition result of the model; subsequently, the probability reference ratio can be used to fuse the audio recognition results of the two models to obtain the final audio recognition result).
[0165] For the sake of easy understanding, taking the enhanced audio recognition model and the general audio recognition model as N - gram language models as an example, the method of fusing the enhanced audio recognition model and the general audio recognition model using the probability reference ratio can be shown as in formula (1):
[0166] P(w m+1 |w1,..., w m ) = Fusion{(1 - λ)P b (w m+1 |w1,..., w m ),
[0167] λP e (w m+1 |w1,..., w m )} Formula (1)
[0168] Among them, P(w m+1 |w1,..., w m)Can be used for a certain moment jointly identified by two models, the candidate text word w m+1 The corresponding probability of selecting a fused word; P b (w m+1 | w1,..., w m )Can be used to represent the probability of selecting a word (which can be called the second word selection probability) corresponding to the candidate text word w in the general audio recognition model; P m+1 (w e (w m+1 | w1,..., w m )Can be used to represent the probability of selecting a word (which can be called the first word selection probability) corresponding to the candidate text word w in the enhanced audio recognition model; λ can be used to represent the probability reference ratio (i.e., the fusion weight); 1 can be used to represent the model fusion coefficient; Fusion() can be used to represent the fusion function, and this fusion function can include the maximum interpolation function max(), the value addition operation function, and so on. Taking the fusion function as the maximum interpolation function max() as an example, the method of fusing the enhanced audio recognition model and the general audio recognition model using the probability reference ratio can be shown in formula (2): m+1 P(w
[0169] | w1,..., w m+1 | w1,..., w m ) = max{(1 - λ)P b (w m+1 | w1,..., w m ),
[0170] λP e (w m+1 | w1,..., w m )} Formula (2)
[0171] That is to say, the maximum value can be selected from (1 - λ)P b (w m+1 | w1,..., w m ) and λP e (w m+1 | w1,..., w m ) as the value of P(w m+1 | w1,..., w m ).
[0172] That is to say, whether in the general audio recognition model or in the enhanced audio recognition model, the word selection probability of each candidate text word is related to the fusion word selection probability of the historical candidate text words (prior words). It should be understood that, assuming that a complete candidate text path (that is, an initial candidate text) includes M candidate text words, then the selection probability of the candidate text path (which can be understood as the final path score) is as shown in formula (3):
[0173]
[0174] Among them, P(h) can be used to represent the selection probability of a candidate text path (initial candidate text); P(w i ) can be used to represent the probability of selecting a fusion word corresponding to a candidate text word in the candidate text path. In other words, the probability of selecting a fusion word corresponding to these candidate text words is multiplied to obtain the probability of selecting the candidate text path.
[0175] It should be understood that after determining the selection probability corresponding to each candidate text path (initial candidate text), the candidate text path corresponding to the maximum selection probability can be selected as the target text recognized by the two models. The specific implementation method is shown in formula (4):
[0176]
[0177] Among them, h z Can be used to characterize the candidate text path finally selected; P(h i ) can be used to characterize the selection probability of a candidate text path.
[0178] For easier understanding, please also refer to Figure 6 , Figure 6 Schematic diagram of an audio recognition scenario provided by an embodiment of the present application. Figure 6 As shown, taking the enhanced audio recognition model and the general audio recognition model as N original grammar language models as an example, taking the target audio data as "you have eaten" as an example, at time T0, the possible recognition results for the sub-audio data "you" include candidate text words "you", candidate text words "you", and candidate text words "ni". Then, through formula (1), the word fusion selection probability of the candidate text word "you", the word fusion selection probability of the candidate text word "you", and the word fusion selection probability of the candidate text word "ni" can be determined. It should be understood that the candidate text word "you", the candidate text word "you", and the candidate text word "ni" can be used as the starting position of the candidate text path.
[0179] Furthermore, at the next moment after T0 (such as T1), the possible recognition results for the sub-audio data "吃" include the candidate text words "吃" and "迟". At this time, the historical candidate text word "你" and the target candidate text word "吃" can be combined into a candidate text path (that is, the initial candidate text), and the historical candidate text word "你" and the target candidate text word "迟" can be combined into a candidate text path. Similarly, the same is true for the historical candidate text words "你" and "妮". At this time, 6 candidate text paths can be obtained. Taking the candidate text path "你吃" as an example, at this time, the word fusion selection probability corresponding to the target candidate text word "吃" in the candidate text path can be determined by formula (1); similarly, the word fusion selection probability corresponding to each target candidate text word in each candidate text path can be determined.
[0180] Similarly, the possible recognition results of the sub-audio data can be determined at each subsequent moment, and a new candidate text path can be formed with the previous historical candidate text words. The word fusion selection probability corresponding to each target candidate text word at each moment can be determined based on the historical candidate text words included in each candidate text path. Figure 6 As shown, 6 candidate text paths can be determined here in the end. Subsequently, for each candidate text path, the corresponding selection probability can be determined by the above formula (2). For example, taking the candidate text path "you have eaten" as an example, the word fusion selection probability of the candidate text word "you", the word fusion selection probability of the candidate text word "eat", the word fusion selection probability of the candidate text word "rice", and the word fusion selection probability of the candidate text word "le" in the candidate text path "you have eaten" can be multiplied, thereby obtaining the selection probability of the candidate text path "you have eaten" as 0.7. Similarly, the selection probabilities corresponding to the other candidate text paths can be obtained as 0.05, 0.1, 0.05, 0.05, and 0.05, respectively. Then, the candidate text path "you have eaten" corresponding to the maximum selection probability of 0.7 can be determined as the recognition text jointly recognized by the two models.
[0181] For further information, see Figure 7 , Figure 7 A system interaction diagram provided in an embodiment of the present application. Figure 7As shown, the system interaction flow can be the interaction between a smart terminal (terminal device) and a cloud server (which can be a business server). After receiving voice data, the smart terminal can send the voice data and the current display interface data (interface text) to the cloud server. The cloud server can recognize the voice data. After the recognition is completed and the voice recognition result is obtained, it can return the text result corresponding to the voice to the smart terminal. The smart terminal can present the text result to the user and execute the action corresponding to the voice data. Among them, for the specific implementation method of the cloud server to recognize voice data, a method of dynamically fusing and recognizing based on a general audio recognition model and an enhanced audio recognition model can be adopted. The specific implementation method can refer to the description in the corresponding embodiment above Figures 3 - 6 and will not be elaborated here. The beneficial effects brought by it will not be elaborated either.
[0182] Further, please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a data processing device provided by an embodiment of the present application. The data processing device can be a computer program (including program code) running in a computer device. For example, the data processing device is an application software. The data processing device can be used to execute Figure 3 the method shown. As Figure 8 shown, the data processing device 1 can include: a model acquisition module 11, a selection probability determination module 12, a target text determination module 13, a ratio adjustment module 14, and a result determination module 15.
[0183] The model acquisition module 11 is used to acquire the interface text included in the current display interface;
[0184] The model acquisition module 11 is further used to acquire an enhanced audio recognition model generated based on the interface text;
[0185] The selection probability determination module 12 is used to, when receiving the target audio data of the target object, perform audio recognition on the target audio data according to the probability reference ratio corresponding to the enhanced audio recognition model, the enhanced audio recognition model, and the general audio recognition model generated based on the general corpus, and obtain the initial selection probabilities corresponding to one or more initial candidate texts;
[0186] The target text determination module 13 is used to determine the first target text from one or more initial candidate texts according to one or more initial selection probabilities;
[0187] The ratio adjustment module 14 is used to adjust the probability reference ratio if the first target text does not match the interface text;
[0188] The target text determination module 13 is further configured to perform audio recognition on the target audio data according to the adjusted probability reference ratio, the enhanced audio recognition model, and the general audio recognition model, obtain the target selection probabilities corresponding to one or more target candidate texts respectively, and determine the second target text from the one or more target candidate texts according to the target selection probabilities;
[0189] The result determination module 15 is configured to, if the second target text matches the interface text, determine the second target text as the audio recognition result for the target audio data.
[0190] Among them, for the specific implementation manners of the model acquisition module 11, the selection probability determination module 12, the target text determination module 13, the ratio adjustment module 14, and the result determination module 15, reference can be made to the descriptions of steps S101 - S104 in the corresponding embodiments above. Figure 3 The descriptions of the corresponding embodiments.
[0191] In one embodiment, the model acquisition module 11 may include: a text extraction unit 111, a normalization processing unit 112, and a text tokenization unit 113.
[0192] The text extraction unit 111 is configured to extract the valid text information included in the interface text;
[0193] The normalization processing unit 112 is configured to perform normalization processing on the valid text information to obtain the standard text information;
[0194] The text tokenization unit 113 is configured to perform tokenization processing on the standard text information to obtain text tokens, train a language processing model according to the text tokens, and determine the trained language processing model as the enhanced audio recognition model.
[0195] Among them, for the specific implementation manners of the text extraction unit 111, the normalization processing unit 112, and the text tokenization unit 113, reference can be made to the description of step S101 in the corresponding embodiments above. Figure 3 The descriptions of the corresponding embodiments.
[0196] In one embodiment, the normalization processing unit 112 may include: a text filtering subunit 1121 and a character replacement subunit 1122.
[0197] The text filtering subunit 1121 is configured to obtain the auxiliary text identifiers included in the valid text information, delete the auxiliary text identifiers, and obtain the filtered text information;
[0198] The character replacement subunit 1122 is configured to obtain the character information included in the filtered text information and the text expression information corresponding to the character information, and replace the character information with the text expression information to obtain the standard text information.
[0199] Among them, for the specific implementation manners of the text filtering subunit 1121 and the character replacement subunit 1122, reference may be made to the description of step S101 in the corresponding embodiment above. Figure 3 described in the corresponding embodiment step S101.
[0200] In one embodiment, the selection probability determination module 12 may include: a candidate word acquisition unit 121, a first word probability determination unit 122, a second word probability determination unit 123, a word combination unit 124, and a selection probability determination unit 125.
[0201] The candidate word acquisition unit 121 is configured to obtain the phoneme features corresponding to the sub-audio data at time T in the target audio data, and obtain the target candidate text words that match the phoneme features at time T in the general corpus; i is a positive integer; i in the general corpus, obtain the target candidate text words that match the phoneme features at time T; i Time T is the previous moment of time T;
[0202] The first word probability determination unit 122 is configured to obtain the historical candidate text words corresponding to the sub-audio data at time T in the target audio data, and determine the first word selection probability corresponding to the target candidate text words according to the historical candidate text words in the enhanced audio recognition model; T i-1 Time T is the previous moment of time T; i-1 Time T is the previous moment of time T; i Time T is the previous moment of time T;
[0203] The second word probability determination unit 123 is configured to determine the second word selection probability corresponding to the target candidate text words according to the historical candidate text words in the general audio recognition model;
[0204] The word combination unit 124 is configured to combine the historical candidate text words and the target candidate text words in chronological order to obtain one or more initial candidate texts;
[0205] The selection probability determination unit 125 is configured to determine the initial selection probabilities corresponding to one or more initial candidate texts according to the first word selection probability, the second word selection probability, and the probability reference ratio.
[0206] Among them, for the specific implementation manners of the candidate word acquisition unit 121, the first word probability determination unit 122, the second word probability determination unit 123, the word combination unit 124, and the selection probability determination unit 125, reference may be made to the description of step S102 in the corresponding embodiment above. Figure 3 described in the corresponding embodiment step S102.
[0207] In one embodiment, the first word probability determination unit 122 may include: a corpus acquisition subunit 1221, a statement statistics subunit 1222, and a word probability determination subunit 1223.
[0208] A corpus acquisition subunit 1221 is configured to acquire an interface corpus corresponding to interface text; the interface corpus includes K text sample statements associated with the interface text;
[0209] A statement statistics subunit 1222 is configured to acquire, from the K text sample statements, the text sample statements including historical candidate text words as a set of text sample statements to be statistically analyzed;
[0210] The statement statistics subunit 1222 is further configured to, in the set of text sample statements to be statistically analyzed, use the text sample statements in which the next sample text word of the historical candidate text word is a target candidate text word as target text sample statements;
[0211] The statement statistics subunit 1222 is further configured to acquire the total number of statements corresponding to the set of text sample statements to be statistically analyzed and the target number corresponding to the target text sample statements;
[0212] A word probability determination subunit 1223 is configured to determine a first word selection probability according to the ratio between the target number and the total number of statements.
[0213] Among them, for the specific implementation manners of the corpus acquisition subunit 1221, the statement statistics subunit 1222, and the word probability determination subunit 1223, reference can be made to the description of step S102 in the corresponding embodiment above. Figure 3 the corresponding description of embodiment step S102.
[0214] In one embodiment, the historical candidate text word includes historical candidate text word C j , and one or more initial candidate texts include initial candidate text M constructed from historical candidate text word C j and the target candidate text word; j is a positive integer; j ; j is a positive integer;
[0215] The selection probability determination unit 125 may include: a probability fusion subunit 1251 and a probability operation subunit 1252.
[0216] The probability fusion subunit 1251 is configured to fuse the first word selection probability and the second word selection probability according to a probability reference ratio to obtain a target fusion word selection probability corresponding to the target candidate text word;
[0217] The probability operation subunit 1252 is configured to obtain the historical fusion word selection probability corresponding to historical candidate text word C j , perform an arithmetic process on the historical fusion word selection probability and the target fusion word selection probability to obtain an initial selection probability corresponding to initial candidate text M j ;
[0218] Among them, for the specific implementation manners of the probability fusion subunit 1251 and the probability operation subunit 1252, reference can be made to the aboveFigure 3 Description corresponding to step S102 of the embodiment
[0219] In one embodiment, the probability fusion subunit 1251 is further specifically configured to perform an arithmetic process on the first word selection probability and the probability reference ratio to obtain a first arithmetic word selection probability;
[0220] The probability fusion subunit 1251 is further specifically configured to obtain a ratio difference between the model fusion coefficient and the probability reference ratio, and perform an arithmetic process on the ratio difference and the second word selection probability to obtain a second arithmetic word selection probability;
[0221] The probability fusion subunit 1251 is further specifically configured to obtain the maximum arithmetic word selection probability from the first arithmetic word selection probability and the second arithmetic word selection probability, and determine the maximum arithmetic word selection probability as the target fusion word selection probability corresponding to the target candidate text word.
[0222] In one embodiment, the target text determination module 13 may include: a maximum probability acquisition unit 131 and a target text determination unit 132.
[0223] The maximum probability acquisition unit 131 is configured to obtain the maximum initial selection probability from one or more initial selection probabilities;
[0224] The target text determination unit 132 is configured to determine the initial candidate text corresponding to the maximum initial selection probability among one or more initial candidate texts as the first target text.
[0225] Wherein, for the specific implementation manners of the maximum probability acquisition unit 131 and the target text determination unit 132, reference may be made to the description of the corresponding embodiment step S102 above Figure 3 Description corresponding to step S102 of the embodiment
[0226] In one embodiment, the ratio adjustment module 14 may include: an adjustment times acquisition unit 141, an increment step determination unit 142, and a ratio adjustment unit 143.
[0227] The adjustment times acquisition unit 141 is configured to, if the first target text does not match the interface text, obtain the current adjustment times corresponding to the probability reference ratio;
[0228] The increment step determination unit 142 is configured to obtain a unit ratio increment step, and determine the current ratio increment step corresponding to the probability reference ratio according to the current adjustment times and the unit ratio increment step;
[0229] The ratio adjustment unit 143 is configured to perform an arithmetic process on the probability reference ratio and the ratio increment step to obtain a target probability reference ratio, and use the target probability reference ratio as the adjusted probability reference ratio.
[0230] Among them, for the specific implementation manners of the adjustment times obtaining unit 141, the growth step determining unit 142, and the ratio adjustment unit 143, reference can be made to the description of step S103 in the corresponding embodiment above. Figure 3 as described in the corresponding embodiment step S103.
[0231] In one embodiment, the data processing device 1 may further include: a statement set obtaining module 16 and a text matching module 17.
[0232] The statement set obtaining module 16 is configured to obtain a set of valid text statements corresponding to the interface text;
[0233] The text matching module 17 is configured to match the first target text with the set of valid text statements;
[0234] The text matching module 17 is further configured to, if there is a valid text statement in the set of valid text statements that matches the first target text, determine that the first target text matches the interface text;
[0235] The text matching module 17 is further configured to, if there is no valid text statement in the set of valid text statements that matches the first target text, determine that the first target text does not match the interface text.
[0236] Among them, for the specific implementation manners of the statement set obtaining module 16 and the text matching module 17, reference can be made to the description of step S102 in the corresponding embodiment above. Figure 3 as described in the corresponding embodiment step S102.
[0237] In one embodiment, the data processing device 1 may further include: an audio result determining module 18.
[0238] The audio result determining module 18 is configured to, if the first target text matches the interface text, determine the first target text as the audio recognition result for the target audio data.
[0239] Among them, for the specific implementation manner of the audio result determining module 18, reference can be made to the description of step S104 in the corresponding embodiment above. Figure 3 as described in the corresponding embodiment step S104.
[0240] In one embodiment, the data processing device 1 may further include: a ratio matching module 19, a step execution module 20, and a model cancellation module 21.
[0241] The ratio matching module 19 is configured to match the adjusted probability reference ratio with a ratio threshold;
[0242] A step execution module 20, configured to, if the adjusted probability reference ratio is less than the ratio threshold, perform the step of performing audio recognition on target audio data according to the adjusted probability reference ratio, the enhanced audio recognition model, and the general audio recognition model to obtain target selection probabilities corresponding to one or more target candidate texts, and determining a second target text from the one or more target candidate texts according to the target selection probabilities;
[0243] A model cancellation module 21, configured to, if the adjusted probability reference ratio is greater than the ratio threshold, cancel the enhanced audio recognition model, and perform audio recognition on the target audio data according to the general audio recognition model to obtain an audio recognition result for the target audio data.
[0244] Among them, for the specific implementation manners of the ratio matching module 19, the step execution module 20, and the model cancellation module 21, reference may be made to the description of step S103 in the corresponding embodiment above. Figure 3 as described in the corresponding embodiment step S103.
[0245] In one embodiment, the data processing device 1 may further include: a result sending module 22.
[0246] The result sending module 22 is configured to send the audio recognition result for the target audio data to the target terminal device, so that the target terminal device generates an action instruction based on the audio recognition result and executes the action instruction in the target display interface; the target terminal device is the terminal device corresponding to the current display interface; the target display interface includes the current display interface.
[0247] Among them, for the specific implementation manner of the result sending module 22, reference may be made to the description of step S104 in the corresponding embodiment above. Figure 3 as described in the corresponding embodiment step S104.
[0248] In the embodiments of the present application, an enhanced audio recognition model for the current display interface can be generated through the current display interface of the terminal device. The enhanced audio recognition model is generated based on the real-time updated content in the current display interface, which is highly targeted and can be combined with the general audio recognition model to more accurately recognize the target audio data related to the current display interface. At the same time, when the recognized text cannot match the interface text, by dynamically adjusting the model fusion weight (i.e., the probability reference ratio), the problem of inaccurate recognized text caused by unreasonable fusion weight settings can be improved, and the accuracy of audio recognition can be further improved. In addition, after the weight adjustment reaches the upper limit value and the recognized text still fails to match any valid statement in the current display interface, the enhanced audio recognition model for the current display interface can be cancelled to only use the general audio recognition model to recognize the target audio data, which can ensure the general recognition effect of the audio data. That is to say, the present application can strengthen the fusion effect between the enhanced audio recognition model and the general audio recognition model by dynamically adjusting the fusion weight, and improve the recognition accuracy of the application interface content of the third-party application in the terminal device. At the same time, when the audio data is related to the current display interface, by dynamically adjusting the fusion weight, the problem that the recognized text fails to be corrected in time due to unreasonable weight settings can be improved, and the recognition accuracy can be further improved. At the same time, by setting an upper limit for the weight and cancelling the enhanced audio recognition model when the upper limit is reached, the general recognition effect can be ensured when the audio data is not related to the content of the current display interface. In summary, the present application can improve the accuracy of audio recognition.
[0249] Further, please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 9 shown, the above Figure 8The data processing device 1 in the corresponding embodiment can be applied to the above computer device 8000. The computer device 8000 may include: a processor 8001, a network interface 8004, and a memory 8005. In addition, the computer device 8000 further includes: a user interface 8003 and at least one communication bus 8002. Among them, the communication bus 8002 is used to realize the connection and communication between these components. Among them, the user interface 8003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 8003 may further include a standard wired interface and a wireless interface. The network interface 8004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 8005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 8005 may further be at least one storage device located far from the aforementioned processor 8001. As Figure 9 shown, the memory 8005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0250] In Figure 9 the computer device 8000 shown, the network interface 8004 can provide network communication functions; while the user interface 8003 is mainly used to provide an input interface for users; and the processor 8001 can be used to call the device control application program stored in the memory 8005 to achieve:
[0251] Obtain the interface text included in the current display interface, and obtain an enhanced audio recognition model generated based on the interface text;
[0252] When receiving the target audio data of the target object, perform audio recognition on the target audio data according to the probability reference ratio corresponding to the enhanced audio recognition model, the enhanced audio recognition model, and the general audio recognition model generated based on the general corpus, to obtain the initial selection probabilities corresponding to one or more initial candidate texts, and determine the first target text from one or more initial candidate texts according to the one or more initial selection probabilities;
[0253] If the first target text does not match the interface text, adjust the probability reference ratio, perform audio recognition on the target audio data according to the adjusted probability reference ratio, the enhanced audio recognition model, and the general audio recognition model, to obtain the target selection probabilities corresponding to one or more target candidate texts, and determine the second target text from one or more target candidate texts according to the target selection probabilities;
[0254] If the second target text matches the interface text, the second target text is determined as the audio recognition result for the target audio data.
[0255] It should be understood that the computer device 8000 described in the embodiments of the present application can execute the description of the data processing method in the corresponding embodiments mentioned above, and can also execute the description of the data processing device 1 in the corresponding embodiments mentioned above, which will not be elaborated here. In addition, the beneficial effects of adopting the same method will not be elaborated either. Figures 3 to 6 The computer device 8000 described in the embodiments of the present application can execute the description of the data processing method in the corresponding embodiments mentioned above, and can also execute the description of the data processing device 1 in the corresponding embodiments mentioned above, which will not be elaborated here. In addition, the beneficial effects of adopting the same method will not be elaborated either. Figure 8 It should be understood that the computer device 8000 described in the embodiments of the present application can execute the description of the data processing method in the corresponding embodiments mentioned above, and can also execute the description of the data processing device 1 in the corresponding embodiments mentioned above, which will not be elaborated here. In addition, the beneficial effects of adopting the same method will not be elaborated either.
[0256] In addition, it should be noted here that: the embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the computer device 1000 for data processing mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, it can execute the description of the data processing method in the corresponding embodiments mentioned above. Therefore, it will not be elaborated here. In addition, the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application. Figures 3 to 6 In addition, it should be noted here that: the embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the computer device 1000 for data processing mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, it can execute the description of the data processing method in the corresponding embodiments mentioned above. Therefore, it will not be elaborated here. In addition, the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application.
[0257] The above computer-readable storage medium may be the data processing device provided in any of the foregoing embodiments or the internal storage unit of the above computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.
[0258] In one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in one aspect of the embodiments of the present application.
[0259] In the description, claims, and drawings of the embodiments of this application, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the term "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or units is not limited to the listed steps or modules, but may optionally further include steps or modules not listed, or may optionally further include other step units inherent to these processes, methods, devices, products, or equipment.
[0260] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0261] The methods and related devices provided in the embodiments of this application are described with reference to the method flowcharts and / or structural schematic diagrams provided in the embodiments of this application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks.
[0262] The above disclosure is only for the preferred embodiments of the present application. Of course, it cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A data processing method, characterized in that, Including: Obtain the interface text included in the current display interface, and obtain the enhanced audio recognition model generated based on the interface text; When receiving the target audio data of the target object, perform audio recognition on the target audio data according to the probability reference ratio corresponding to the enhanced audio recognition model, the enhanced audio recognition model, and the general audio recognition model generated based on the general corpus, to obtain the initial selection probabilities corresponding to one or more initial candidate texts, and determine the first target text from the one or more initial candidate texts according to the one or more initial selection probabilities; If the first target text does not match the interface text, adjust the probability reference ratio, perform audio recognition on the target audio data according to the adjusted probability reference ratio, the enhanced audio recognition model, and the general audio recognition model, to obtain the target selection probabilities corresponding to one or more target candidate texts, and determine the second target text from the one or more target candidate texts according to the target selection probabilities; If the second target text matches the interface text, determine the second target text as the audio recognition result for the target audio data; Among them, performing audio recognition on the target audio data according to the probability reference ratio corresponding to the enhanced audio recognition model, the enhanced audio recognition model, and the general audio recognition model generated based on the general corpus to obtain the initial selection probabilities corresponding to one or more initial candidate texts, including: obtaining the phoneme features corresponding to the sub-audio data at time T in the target audio data, and obtaining the target candidate text word that matches the phoneme features at time T in the general corpus; i is a positive integer; obtaining the historical candidate text word corresponding to the sub-audio data at time T in the target audio data, and determining the first word selection probability corresponding to the target candidate text word according to the historical candidate text word in the enhanced audio recognition model; time T is the previous time of time T; determining the second word selection probability corresponding to the target candidate text word according to the historical candidate text word in the general audio recognition model; combining the historical candidate text word and the target candidate text word in chronological order to obtain the one or more initial candidate texts; determining the initial selection probabilities corresponding to the one or more initial candidate texts respectively according to the first word selection probability, the second word selection probability, and the probability reference ratio. i Obtaining the phoneme features corresponding to the sub-audio data at time T in the target audio data, and obtaining the target candidate text word that matches the phoneme features at time T in the general corpus; i is a positive integer; i Obtaining the phoneme features corresponding to the sub-audio data at time T in the target audio data, and obtaining the target candidate text word that matches the phoneme features at time T in the general corpus; i is a positive integer; i-1 Obtaining the historical candidate text word corresponding to the sub-audio data at time T in the target audio data, and determining the first word selection probability corresponding to the target candidate text word according to the historical candidate text word in the enhanced audio recognition model; T i-1 Time is the previous time of time T; i Time is the previous time of time T; determining the second word selection probability corresponding to the target candidate text word according to the historical candidate text word in the general audio recognition model; combining the historical candidate text word and the target candidate text word in chronological order to obtain the one or more initial candidate texts; determining the initial selection probabilities corresponding to the one or more initial candidate texts respectively according to the first word selection probability, the second word selection probability, and the probability reference ratio.
2. The method according to claim 1, wherein The determining the first word selection probability corresponding to the target candidate text word according to the historical candidate text word in the enhanced audio recognition model includes: Obtain the interface corpus corresponding to the interface text; the interface corpus includes K text sample sentences associated with the interface text; Obtain the text sample sentences including the historical candidate text word from the K text sample sentences as the set of text sample sentences to be counted; In the set of text sample sentences to be counted, use the text sample sentence in which the next sample text word of the historical candidate text word is the target candidate text word as the target text sample sentence; Obtain the total number of sentences corresponding to the set of text sample sentences to be counted, and the target number corresponding to the target text sample sentence; Determine the first word selection probability according to the ratio between the target number and the total number of sentences.
3. The method according to claim 1, wherein The historical candidate text words include the historical candidate text word C j , and the one or more initial candidate texts include the initial candidate text M constructed from the historical candidate text word C j and the target candidate text word j ; j is a positive integer; The determining the initial selection probabilities corresponding to the one or more initial candidate texts according to the first word selection probability, the second word selection probability, and the probability reference ratio includes: Fuse the first word selection probability and the second word selection probability according to the probability reference ratio to obtain the target fused word selection probability corresponding to the target candidate text word; Obtain the historical candidate text word C j For the corresponding historical fusion word selection probability, perform an arithmetic operation on the historical fusion word selection probability and the target fusion word selection probability to obtain the initial candidate text M j The corresponding initial selection probability.
4. The method according to claim 3, wherein The fusing the first word selection probability and the second word selection probability according to the probability reference ratio to obtain the target fused word selection probability corresponding to the target candidate text word includes: Perform an arithmetic operation on the first word selection probability and the probability reference ratio to obtain the first arithmetic word selection probability; Obtain the ratio difference between the model fusion coefficient and the probability reference ratio, and perform an arithmetic operation on the ratio difference and the second word selection probability to obtain the second arithmetic word selection probability; Obtain the maximum operator selection probability from the first operator selection probability and the second operator selection probability, and determine the maximum operator selection probability as the target fusion word selection probability corresponding to the target candidate text word.
5. The method according to claim 1, wherein If the first target text does not match the interface text, the adjustment of the probability reference ratio includes: If the first target text does not match the interface text, obtain the current adjustment times corresponding to the probability reference ratio; Obtain the unit ratio growth step, and determine the current ratio growth step corresponding to the probability reference ratio according to the current adjustment times and the unit ratio growth step; Perform arithmetic processing on the probability reference ratio and the ratio growth step to obtain a target probability reference ratio, and use the target probability reference ratio as the adjusted probability reference ratio.
6. The method according to claim 1, wherein The method further includes: Match the adjusted probability reference ratio with a ratio threshold; If the adjusted probability reference ratio is less than the ratio threshold, execute the step of performing audio recognition on the target audio data according to the adjusted probability reference ratio, the enhanced audio recognition model, and the general audio recognition model to obtain target selection probabilities corresponding to one or more target candidate texts, and determining a second target text from the one or more target candidate texts according to the target selection probabilities; If the adjusted probability reference ratio is greater than the ratio threshold, deactivate the enhanced audio recognition model, and perform audio recognition on the target audio data according to the general audio recognition model to obtain an audio recognition result for the target audio data.
7. A computer device, characterized in that, Including: A processor, a memory, and a network interface; The processor is connected to the memory and the network interface. Among them, the network interface is used to provide network communication functions, the memory is used to store program codes, and the processor is used to call the program codes so that the computer device executes the method according to any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is adapted to be loaded and executed by the processor to execute the method according to any one of claims 1-6.
9. A computer program product or a computer program, characterized in that, The computer program product or computer program includes computer instructions, the computer instructions are stored in the computer-readable storage medium, and the computer instructions are adapted to be read and executed by the processor so that a computer device with the processor executes the method according to any one of claims 1-6.
Citation Information
Patent Citations
Method, device, apparatus for speech recognition, and storage medium
CN109243461A
Voice interaction method, device and system
CN111383631A
Audio data processing method and device, storage medium and equipment
CN111554300A