Text recognition method, device, computer-readable storage medium and electronic device
By generating and judging semantic complete interactive text in human-computer interaction, the problem of recognition errors caused by semantic incompleteness in human-computer interaction is solved, and the accuracy of semantic integrity judgment and the efficiency of user intention recognition are improved.
Patent Information
- Application Number
- CN202210765786.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-01
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-07-01
AI Technical Summary
In human-computer interaction, since the speech sent by the user may be paused, resulting in the device's intent recognition errors of the user, it is difficult for existing semantic judgment methods to obtain reliable semantic integrity judgment results.
By obtaining the first interactive text obtained by the user and the target device in human-computer interaction, a second interactive text is generated, and a semantic integrity judgment step is performed to determine whether the second interactive text meets the semantic integrity judgment conditions. If it is met, its semantic integrity information is determined. Based on this information and the second interactive text, a semantic complete third interactive text is obtained, and it is recognized to obtain the first text recognition result.
It improves the accuracy of semantic integrity judgment, can obtain semantic complete text for user intention recognition, effectively reducing the risk of recognition errors caused by user semantic incompleteness during human-computer interaction.
Smart Images

Figure CN115114935B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a text recognition method, apparatus, computer-readable storage medium, and electronic device. Background Art
[0002] With the development of artificial intelligence technologies, there are more and more application scenarios for human-computer interaction. Users can interact with devices in various ways such as voice, gestures, and eye contact. For example, when a user is in a vehicle, the user can control devices such as windows and air conditioners by voice, thus freeing the hands and avoiding traffic hazards; when the vehicle stops, the user can experience many novel functions of the intelligent devices in the cockpit through various ways such as gestures and eye contact.
[0003] When a user performs human-computer interaction with a device by voice, since there may be pauses in the voice emitted by the user, the device may misrecognize the user's intention. To avoid the device obtaining incomplete semantics of the user, currently, a semantic non-stop determination method can be adopted to reduce the risk of misrecognizing the user's intention by judging the integrity of the user's semantics.
[0004] The current semantic non-stop determination method usually performs multi-modal semantic non-stop determination by extracting audio vector features and text vector features from audio data and text data. However, since the semantic integrity is strongly correlated with the application scenario, and there may be ambiguity in judging whether the semantics is complete, it is difficult to obtain a reliable semantic integrity judgment result only by identifying through audio data or text data. Summary of the Invention
[0005] Embodiments of the present disclosure provide a text recognition method, apparatus, computer-readable storage medium, and electronic device.
[0006] Embodiments of the present disclosure provide a text recognition method, which includes: obtaining a first interaction text obtained by a user performing human-computer interaction with a target device; generating a second interaction text based on the first interaction text; performing a semantic integrity judgment step based on the second interaction text to determine whether the second interaction text meets the semantic integrity judgment condition; in response to the second interaction text meeting the semantic integrity judgment condition, determining semantic integrity information of the second interaction text; obtaining a third interaction text with complete semantics based on the semantic integrity information and the second interaction text; and recognizing the third interaction text to obtain a first text recognition result.
[0007] According to another aspect of the embodiments of the present disclosure, a text recognition device is provided. The device includes: an acquisition module configured to acquire a first interaction text obtained by a user's human-computer interaction with a target device; a first generation module configured to generate a second interaction text based on the first interaction text; a first determination module configured to perform a semantic integrity judgment step based on the second interaction text to determine whether the second interaction text meets the semantic integrity judgment condition; a second determination module configured to, in response to the second interaction text meeting the semantic integrity judgment condition, determine the semantic integrity information of the second interaction text; a second generation module configured to obtain a third interaction text with semantic integrity based on the semantic integrity information and the second interaction text; and a first recognition module configured to recognize the third interaction text to obtain a first text recognition result.
[0008] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program is used to execute the above-mentioned text recognition method.
[0009] According to another aspect of the embodiments of the present disclosure, an electronic device is provided. The electronic device includes: a processor; a memory for storing processor-executable instructions; and the processor is configured to read the executable instructions from the memory and execute the instructions to implement the above-mentioned text recognition method.
[0010] Based on the text recognition method, device, computer-readable storage medium, and electronic device provided in the above embodiments of the present disclosure, by generating a second interaction text based on the first interaction text obtained by a user's human-computer interaction with a target device, then determining whether the second interaction text meets the semantic integrity judgment condition, if it meets, determining the semantic integrity information of the second interaction text, and then obtaining a third interaction text with semantic integrity based on the semantic integrity information and the second interaction text, and finally recognizing the third interaction text to obtain a first text recognition result. It realizes not only recognizing the features of the interaction text, but also obtaining an interaction text with semantic integrity based on the current interaction text obtained by human-computer interaction through means such as multi-round interaction, thereby improving the accuracy of semantic integrity judgment, and being able to obtain an interaction text with semantic integrity for user intention recognition, effectively reducing the risk of recognition errors caused by incomplete semantics of the user during human-computer interaction.
[0011] The technical solution of the present disclosure will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0012] The above and other objects, features, and advantages of the present disclosure will become more apparent by describing the embodiments of the present disclosure in more detail with reference to the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation to the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.
[0013] Figure 1 is a system diagram applicable to the present disclosure.
[0014] Figure 2 is a schematic flowchart of a text recognition method provided by an exemplary embodiment of the present disclosure.
[0015] Figure 3 is a schematic flowchart of a text recognition method provided by another exemplary embodiment of the present disclosure.
[0016] Figure 4 is a schematic flowchart of a text recognition method provided by another exemplary embodiment of the present disclosure.
[0017] Figure 5 is a schematic flowchart of a text recognition method provided by another exemplary embodiment of the present disclosure.
[0018] Figure 6 is a schematic flowchart of a text recognition method provided by another exemplary embodiment of the present disclosure.
[0019] Figure 7 is a schematic flowchart of a text recognition method provided by another exemplary embodiment of the present disclosure.
[0020] Figure 8 is a schematic structural diagram of a text recognition device provided by an exemplary embodiment of the present disclosure.
[0021] Figure 9 is a schematic structural diagram of a text recognition device provided by another exemplary embodiment of the present disclosure.
[0022] Figure 10 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed Embodiments
[0023] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.
[0024] It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0025] Those skilled in the art can understand that terms such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.
[0026] It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0027] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, in the absence of a clear limitation or contrary indication in the context, it can generally be understood as one or more.
[0028] In addition, the term "and / or" in the present disclosure is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.
[0029] It should also be understood that the description of each embodiment in the present disclosure emphasizes the differences between the embodiments, and their similarities or similarities can be referred to each other. For the sake of brevity, they will not be elaborated one by one.
[0030] At the same time, it should be understood that for the convenience of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0031] The following description of at least one exemplary embodiment is actually merely illustrative and in no way limits the present disclosure and its application or use.
[0032] Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods and devices should be regarded as part of the specification.
[0033] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0034] Embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0035] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0036] Overview of the Application
[0037] Current human-computer interaction methods are usually affected by factors such as non-standard user speech and environmental noise, resulting in the device not being able to recognize the user's expression well, or there being a short pause during the user's expression, causing the device to cut the user's speech into two segments, so that the device cannot correctly parse the user's intention. For example, if the user wants to express "open the window", but pauses between "open" and "window", then the device will recognize two results: "open" and "window". For the device, the semantics of these two results cannot be parsed, which will have a greater impact on human-computer interaction.
[0038] Embodiments of the present disclosure aim to solve the problem of incorrect human-computer interaction caused by incomplete recognized user semantics during human-computer interaction, and overcome the problem of poor scenario adaptability caused by the current semantic judgment method only using the user's speech or text features for integrity judgment. By judging whether the semantics of the current interaction text is complete, an interaction text with complete semantics is obtained through multiple rounds of interaction for user intention recognition, thereby improving the accuracy of user intention recognition.
[0039] Exemplary System
[0040] Figure 1Fig. 0 shows an exemplary system architecture 100 of a text recognition method or a text recognition device to which embodiments of the present disclosure can be applied.
[0041] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0042] The user may use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications, such as a voice recognition application, an image recognition application, etc., may be installed on the terminal device 101.
[0043] The terminal device 101 may be various electronic devices, including but not limited to mobile terminals such as in-vehicle terminals, mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0044] The server 103 may be a server that provides various services, such as a background server that recognizes interactive text, audio, etc. uploaded by the terminal device 101. The background server may analyze the received interactive text to obtain a semantically complete text and recognize the semantically complete text.
[0045] It should be noted that the text recognition method provided by the embodiments of the present disclosure may be executed by the server 103 or the terminal device 101. Correspondingly, the text recognition device may be disposed in the server 103 or the terminal device 101.
[0046] It should be understood that Figure 1 the numbers of the terminal device, the network, and the server in
[0047] Exemplary Method
[0048] Figure 2 are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. In the case where information such as interactive text does not need to be obtained remotely, the above system architecture may not include a network and only include a server or a terminal device. Figure 1 shown, the terminal device 101 or the server 103), as Figure 2 shown, the method includes the following steps:
[0049] Step 201: Obtain a first interaction text resulting from the human-machine interaction between the user and the target device.
[0050] In this embodiment, the electronic device can obtain a first interaction text resulting from the human-machine interaction between the user and the target device. Herein, the target device can be various devices supporting human-machine interaction. For example, when this embodiment is applied to a vehicle, the target device can be a vehicle-mounted controller. When this embodiment is applied to household appliances, the target device can be smart TVs, smart air conditioners and other household appliances. It should be noted that the electronic device executing this embodiment can be the same device as the target device or a device separated from the target device. For example, the target device can be a controller in a vehicle or a remote server communicatively connected to the controller.
[0051] The first interaction text is the text (also known as a query) obtained by the target device through identifying the instruction issued by the user. The instructions issued by the user can be of various types, such as voice, gesture, line of sight direction, etc. As an example, when the user conducts human-machine interaction with the target device through voice, the ASR (Automatic Speech Recognition) technology can be adopted for speech recognition to obtain the first interaction text.
[0052] Step 202: Generate a second interaction text based on the first interaction text.
[0053] In this embodiment, the electronic device can generate a second interaction text based on the first interaction text. Specifically, as an example, when the user conducts human-machine interaction with the target device for the first time, the first interaction text can be directly determined as the second interaction text. When the user has conducted at least one interaction with the target device, that is, when this text recognition method has been executed at least once, if a semantically complete third interaction text cannot be obtained based on the second interaction text during the previous execution of this text recognition method, the previously obtained historically incomplete interaction text will be cached in a preset storage area. When this text recognition method is executed this time, if the historically incomplete interaction text has been cached in the preset storage area, the first interaction text and the historically incomplete interaction text can be merged into the second interaction text.
[0054] Step 203: Based on the second interaction text, execute a semantic integrity judgment step to determine whether the second interaction text meets the semantic integrity judgment conditions.
[0055] In this embodiment, the electronic device can determine whether the second interaction text meets the semantic integrity judgment condition based on the second interaction text. The semantic integrity judgment condition is a condition set in advance for screening the semantic integrity of the second interaction text. If the second interaction text meets the semantic integrity judgment condition, it is necessary to further analyze the semantic integrity of the second interaction text. If it does not meet the semantic integrity judgment condition, the second interaction text can be directly recognized to determine the user's intention.
[0056] The semantic integrity judgment condition can be set as needed. For example, it can be determined whether the number of characters included in the second interaction text exceeds a preset number. If it exceeds the preset number, it does not meet the semantic integrity judgment condition. At this time, it is determined that the semantics of the second interaction text is complete, and the second interaction text is directly recognized. If it does not exceed the preset number, the semantic integrity of the second interaction text is further analyzed.
[0057] Step 204: In response to the second interaction text meeting the semantic integrity judgment condition, determine the semantic integrity information of the second interaction text.
[0058] In this embodiment, the electronic device can determine the semantic integrity information of the second interaction text in response to the second interaction text meeting the semantic integrity judgment condition. The semantic integrity information is used to represent whether the semantics of the second interaction text is complete. The existing methods can be used to analyze the semantic integrity of the second interaction text. For example, a semantic integrity analysis model based on TCN (Temporal Convolutional Network).
[0059] Step 205: Based on the semantic integrity information and the second interaction text, obtain a third interaction text with complete semantics.
[0060] In this embodiment, the electronic device can obtain a third interaction text with complete semantics based on the semantic integrity information and the second interaction text. Specifically, as an example, when the semantic integrity information indicates that the semantics of the second interaction text is complete, the second interaction text can be determined as the third interaction text with complete semantics. When the semantic integrity information indicates that the semantics of the second interaction text is incomplete, the text obtained from the previous human-computer interaction with the user can be continuously obtained, merged with the second interaction text, and then it is determined again whether the merged text meets the semantic integrity judgment condition and whether the semantics is complete until a third interaction text with complete semantics is obtained.
[0061] As an example, the user conducts human-computer interaction with the target device through voice. If the first interaction text recognized is "The air conditioner temperature is high", and since this is the first time the first interaction text is obtained and there is no cached historical interaction text, then "The air conditioner temperature is high" is the second interaction text. The semantic integrity judgment condition includes: if the number of characters in the second interaction text is less than or equal to 15, it is determined that the second interaction text meets the semantic integrity judgment condition. Further semantic integrity analysis is performed on the second interaction text, and it is determined that the semantics of "The air conditioner temperature is high" is incomplete. At this time, the electronic device can cache "The air conditioner temperature is high" as the historical interaction text, and the electronic device prompts the user to continue sending voice signals. If the text "One degree higher" is obtained again and merged with the cached "The air conditioner temperature is high", the semantically complete third interaction text "The air conditioner temperature is one degree higher" is obtained.
[0062] Step 206, recognize the third interaction text to obtain the first text recognition result.
[0063] In this embodiment, the electronic device can recognize the third interaction text to obtain the first text recognition result. Among them, the first text recognition result is used to represent the user's intention. The method for recognizing the third interaction text can adopt existing methods, such as recognition based on NLU (Natural Language Understanding) technology.
[0064] Continuing with the above example, performing NLU recognition on the third interaction text "The air conditioner temperature is one degree higher", the first text recognition result obtained can include intention information indicating controlling the air conditioner to adjust the temperature and slot information indicating a one-degree increase in temperature. Subsequently, the target device controls the air conditioner temperature to increase by one degree.
[0065] The method provided in the above embodiments of the present disclosure generates the second interaction text based on the first interaction text obtained through human-computer interaction between the user and the target device, then determines whether the second interaction text meets the semantic integrity judgment condition. If it meets, the semantic integrity information of the second interaction text is determined. Then, based on the semantic integrity information and the second interaction text, a semantically complete third interaction text is obtained. Finally, the third interaction text is recognized to obtain the first text recognition result. It realizes not only recognizing the features of the interaction text, but based on the interaction text obtained from the current human-computer interaction, obtaining a semantically complete interaction text through means such as multi-round interaction, thereby improving the accuracy of semantic integrity judgment, and being able to obtain a semantically complete text for user intention recognition, effectively reducing the risk of recognition errors caused by incomplete semantics of the user during human-computer interaction.
[0066] In some alternative implementation manners, as Figure 3 shown, step 205 includes:
[0067] Step 2051: In response to the semantic integrity information indicating that the second interactive text is semantically complete, determine the second interactive text as the third interactive text.
[0068] Step 2052: In response to the semantic integrity information indicating that the second interactive text is semantically incomplete, perform semantic parsing on the second interactive text to obtain first semantic parsing information.
[0069] Among them, the purpose of semantic parsing is to determine the main meaning of the second interactive text, that is, to determine the specific intention of the user. Specifically, the method of intention recognition and slot recognition for the second interactive text can be used for semantic parsing to obtain intention information and slot information. The intention information indicates whether the user has a human-computer interaction intention, and the slot information refers to the meaningful keywords included in the text. Specifically, a set of text templates of the target type can be preset in advance. If the second interactive text matches the text template, it is determined that the intention of the second interactive text is clear. For example, when this embodiment is applied in the field of vehicle control, the set of text templates may include body control text templates, vehicle information query text templates, in-vehicle device control text templates, etc. As an example, if the second interactive text includes "The air conditioner temperature is a bit higher", the intention recognition result indicates that the second interactive text matches the in-vehicle device control text template, and it is determined that the user has a clear intention to adjust the in-vehicle device. The slot information may include various keywords in the field of vehicle control, such as entities like windows, sunroofs, air conditioners, verbs like "bigger", "brighter", "raise", and quantifiers like "third gear", "26 degrees".
[0070] Intention recognition and slot recognition can be implemented by using existing intention slot recognition methods. For example, the method based on LSTM (Long Short-Term Memory) + CRF (Conditional Random Field) can be used.
[0071] Step 2053: Based on the first semantic parsing information and the second interactive text, perform at least one round of question-and-answer recognition operations with the user to obtain user feedback texts respectively corresponding to the at least one round of question-and-answer recognition operations.
[0072] Specifically, questions can be generated and output through a preset question-and-answer template and / or a question generation model. The user provides feedback based on the output questions, thereby obtaining feedback text. For example, if the second interaction text is "The temperature is high", and the intention information included in the first semantic parsing information is empty and the slot information is incomplete, it is impossible to determine the precise meaning of "The temperature is high" in terms of human-computer interaction, that is, the second interaction text is ambiguous. Then questions such as "Please state which part of the temperature needs to be adjusted", "Is the temperature to be raised or lowered", "By how many degrees should it be raised or lowered" can be generated to prompt the user to supplement the intention information and slot information, thereby guiding the user to give further instructions and obtaining the user feedback text.
[0073] The above question-and-answer template can contain a set of multiple question texts. If the missing intention information and slot information in the first semantic parsing information match a certain question in the question-and-answer template, then that question is output. The above question generation model can be trained based on existing deep learning models, for example, trained based on existing NLG (Natural Language Generation) models.
[0074] Step 2054: Generate a third interaction text based on the user feedback text and the second interaction text respectively corresponding to at least one round of question-and-answer recognition operations.
[0075] Specifically, if the user feedback text and the second interaction text obtained after at least one round of question-and-answer recognition operations can supplement the missing intention information and slot information in the second interaction text, and if the supplemented text contains clear intention information and complete slot information, that is, the ambiguity is eliminated, then a semantically complete third interaction text can be generated based on the text after the ambiguity is eliminated. Usually, the method for judging semantic integrity described in step 204 can be used again to determine whether the supplemented text is semantically complete. If it is complete, the supplemented text is determined as the third interaction text; otherwise, the user can be prompted to continue the human-computer interaction to obtain a new interaction text, and the new interaction text is combined with the above supplemented text, and the semantic integrity of the combined text is judged again until a semantically complete third interaction text is obtained.
[0076] As an example, if the initial second interaction text is "The temperature is high", after at least one round of question-and-answer operations, the missing intention information and slot information in the second interaction text are supplemented to obtain the supplemented text "Raise the air conditioner temperature by one degree". Continuing to use the method in step 204 to determine that the supplemented text is semantically complete, then "Raise the air conditioner temperature by one degree" is determined as the semantically complete third interaction text.
[0077] In this embodiment, when the semantics of the second interaction text is incomplete, at least one round of question-and-answer recognition operation is performed, and a third interaction text with complete semantics is generated according to the user feedback text, so as to ensure that a text with complete semantics can be obtained during human-computer interaction, effectively avoiding the subsequent recognition of texts with incomplete semantics and improving the accuracy of human-computer interaction.
[0078] In some alternative implementation manners, as Figure 4 shown, step 2054 includes:
[0079] Step 20541, generate a fourth interaction text to be recognized for semantic integrity based on the user feedback text and the second interaction text respectively corresponding to at least one round of question-and-answer recognition operation.
[0080] Specifically, after at least one round of question-and-answer operations, the second interaction text can be supplemented according to the user feedback text (for example, supplementing intention information and slot information). If the supplemented text (for example, it can be the direct combination of the user feedback text and the second interaction text, or the text after supplementing the missing intention information and slot information in the second interaction text) can eliminate ambiguity (for example, the intention information indicates a clear human-computer interaction intention and the slot information is complete), then the supplemented text can be determined as the fourth interaction text. If the ambiguity cannot be eliminated after a preset number of rounds (for example, 3 rounds) of question-and-answer recognition operations, then the supplemented text obtained most recently is determined as the fourth interaction text.
[0081] Step 20542, determine whether the semantics of the fourth interaction text is complete.
[0082] Specifically, the method for determining semantic integrity described in step 204 above can be used to determine whether the fourth interaction text is semantically complete.
[0083] If it is complete, execute step 20543; if it is incomplete, execute step 20544.
[0084] Step 20543, determine the fourth interaction text as the third interaction text.
[0085] That is, if the semantics of the fourth interaction text is complete, the fourth interaction text is determined as the third interaction text.
[0086] Step 20544, store the fourth interaction text in a preset storage area.
[0087] That is, if the semantics of the fourth interaction text is incomplete, the fourth interaction text is stored in a preset storage area. The preset storage area can be a storage area pre-set for temporarily caching texts, which can be set locally in the above-mentioned electronic device or remotely.
[0088] Step 20545: Obtain the fifth interaction text obtained when the user interacts with the target device again.
[0089] Since the semantics of the fourth interaction text is incomplete at this time, the user can further interact with the target device, and determine the interaction text obtained this time as the fifth interaction text.
[0090] Generally, the target device can issue a prompt message, and the user interacts with the target device again according to the prompt message. For example, the target device can emit a prompt sound "Please say it again", or, based on the method of generating questions described in the above step 2053, emit a question prompt sound. For example, "Please tell which part's temperature to adjust", "Increase or decrease the temperature", "Increase or decrease by how many degrees", etc.
[0091] Step 20546: Merge the fifth interaction text with the text cached in the preset storage area, and update the second interaction text based on the merged text.
[0092] The preset storage area is a storage area dedicated to caching interaction texts, and the preset storage area will be emptied after obtaining the semantically complete third interaction text. Therefore, the text merged with the fifth interaction text can be all the texts cached in the preset storage area, that is, the texts cached after multiple rounds of interactions.
[0093] It should be noted that the method of merging the fifth interaction text with the text cached in the preset storage area can be direct merging or indirect merging. For example, extract the intent information and / or slot information from the fifth interaction text, use the extracted intent information and / or slot information to supplement the intent information and / or slot information of the text in the preset storage area, and determine the supplemented text as the merged text. The merged text is the updated second interaction text.
[0094] Step 20547: Based on the updated second interaction text, continue to execute the semantic integrity judgment step until a semantically complete third interaction text is obtained.
[0095] That is, the electronic device continues to execute the above step 203 to determine whether the updated second interaction text meets the semantic integrity judgment conditions. If it meets, execute the above steps 204 - 205, and so on in a loop until a semantically complete third interaction text is obtained.
[0096] In this embodiment, by generating a fourth interaction text after at least one round of question-and-answer recognition operations, when the semantics of the fourth interaction text is incomplete, caching the fourth interaction text, and cyclically executing steps 203 - 205 based on the cached text, it is possible to efficiently obtain a semantically complete third interaction text, thereby improving the accuracy and efficiency of human-computer interaction.
[0097] In some alternative implementations, such as Figure 5 shown, step 20541 includes:
[0098] Step 205411, determining whether the round of the question-and-answer recognition operation to be executed is less than or equal to a preset round.
[0099] If it is less than or equal to the preset round (for example, 3 rounds), execute step 205412, otherwise execute step 205417.
[0100] Step 205412, performing the question-and-answer recognition operation.
[0101] Step 205413, merging the obtained user feedback text and the second interaction text to generate a sixth interaction text to be semantically parsed.
[0102] Among them, the obtained user feedback text is a set of user feedback texts obtained from the first round of question-and-answer recognition operation to the current round of question-and-answer recognition operation. For example, if the current execution reaches the third round of question-and-answer recognition operation, all the user feedback texts obtained in the first three rounds are merged with the second interaction text to generate a sixth interaction text. It should be noted that the method of merging the obtained user feedback text and the second interaction text is the same as the merging method described in the above step 20546 and will not be elaborated here.
[0103] Step 205414, performing semantic parsing on the sixth interaction text to obtain second semantic parsing information.
[0104] Among them, the method of performing semantic parsing on the sixth interaction text is the same as the semantic parsing method described in the above step 2052 and will not be elaborated here.
[0105] Step 205415, determining whether the sixth interaction text meets the interaction intention condition based on the second semantic parsing information.
[0106] Among them, the second semantic parsing information is used to indicate whether the sixth interaction text has a clear human-computer interaction intention. If it has a clear human-computer interaction intention, it is determined that the sixth interaction text meets the interaction intention condition. For example, if the slot information included in the second semantic parsing information is complete, it can be determined that the interaction intention condition is met.
[0107] If the sixth interaction text meets the interaction intention condition, execute step 205416, otherwise continue to execute step 205411, that is, the maximum number of rounds of executing the question-and-answer recognition operation is the preset round.
[0108] Step 205416, determining the sixth interaction text as the fourth interaction text.
[0109] In this embodiment, during the process of performing at least one round of question-and-answer recognition operations, it is determined whether the sixth interaction text obtained from the current round of question-and-answer recognition operations meets the interaction intention condition. When it meets the interaction intention condition, the fourth interaction text to be subjected to semantic integrity recognition is obtained, thereby achieving the acquisition of the fourth interaction text that meets the interaction intention condition through performing at least one round of question-and-answer recognition operations. Based on the fourth interaction text that meets the interaction intention condition, the third interaction text with complete semantics can be further obtained quickly, improving the efficiency of human-computer interaction.
[0110] In some alternative implementation manners, step 205414 may be executed as follows:
[0111] Perform intention recognition and / or slot recognition on the sixth interaction text to obtain an intention recognition result and / or a slot recognition result.
[0112] Among them, for the methods of intention recognition and slot recognition, please refer to the content described in step 2052 above and will not be elaborated here.
[0113] Determining whether the sixth interaction text meets the interaction intention condition includes:
[0114] In response to the intention recognition result indicating that the user has a human-computer interaction intention and / or the slot recognition result indicating that the slots in the sixth interaction text are complete, it is determined that the sixth interaction text meets the interaction intention condition.
[0115] Generally, in order to obtain a sixth interaction text with a clearer human-computer interaction intention, the interaction intention condition is set to that the user has a human-computer interaction intention and the slots are complete. If in order to reduce the information processing amount of the electronic device and improve the information processing speed, the interaction intention condition can be set to that the user has a human-computer interaction intention or the slots are complete.
[0116] As an example, if the sixth interaction text includes "The air conditioner temperature is a bit higher", the intention recognition result is "Air conditioner temperature adjustment", then it can be determined that the user has a human-computer interaction intention, and the slot recognition result includes "Air conditioner", "Temperature", "Increase", then it is determined that the slots are complete.
[0117] In this embodiment, by performing intention recognition and / or slot recognition on the sixth interaction text, it can be accurately determined whether the sixth user interaction text has a clear human-computer interaction intention, thereby helping to more specifically prompt the user to make human-computer interaction actions, making the question-and-answer recognition operation more efficient and improving the accuracy of human-computer interaction.
[0118] In some alternative implementation manners, as Figure 5 shown, if the number of rounds of the question-and-answer recognition operations to be executed is greater than the preset number of rounds, execute step 205417:
[0119] Determine the most recently obtained sixth interaction text as the fourth interaction text.
[0120] For example, if the number of rounds of the executed question-and-answer recognition operation reaches three, and the sixth user interaction text obtained in the third round of the question-and-answer recognition operation still does not meet the interaction intention condition, step 205411 is executed again to determine that the round of the question-and-answer recognition operation to be executed is 4, which is less than the preset round of 3. Then, step 205417 is executed to determine the sixth interaction text obtained in the third round of the question-and-answer recognition operation as the fourth interaction text.
[0121] In this embodiment, by setting the maximum number of rounds of the question-and-answer recognition operation, it is possible to avoid the situation where the program cannot be fully executed for a long time when the ambiguity cannot be eliminated after multiple recognitions of the user's feedback text, which helps to cache the text as soon as possible, and then obtain the interaction text again, and combine the cached text to obtain the complete third interaction text, thereby helping to improve the efficiency of human-computer interaction.
[0122] In some alternative implementation manners, as Figure 6 shown, step 202 includes:
[0123] Step 2021, determine whether there is pre-cached text in the preset storage area.
[0124] The pre-cached text is the text with incomplete semantics obtained previously. That is, when the text recognition method is executed last time, if the complete third interaction text cannot be obtained based on the second interaction text, the text with incomplete semantics obtained last time (such as the second interaction text obtained last time, or the fourth interaction text described in step 20544 above) will be cached in the preset storage area and wait to be extracted from the preset storage area when the text recognition method is executed next time.
[0125] Step 2022, in response to determining that there is pre-cached text, merge the first interaction text and the pre-cached text into the second interaction text.
[0126] In this embodiment, by merging the first interaction text with the pre-cached text to obtain the second interaction text, the texts of multiple rounds of human-computer interaction are merged, which helps to obtain the complete third interaction text for text recognition. Compared with the existing semantic judgment method that can only judge whether the semantics is complete by analyzing voice and text features, this embodiment can greatly reduce the risk of obtaining text with incomplete semantics and improve the accuracy of human-computer interaction.
[0127] In some alternative implementation manners, as Figure 7 shown, step 203 includes:
[0128] Step 2031, determine whether the text length of the second interaction text is less than or equal to the preset length.
[0129] Among them, the text length refers to the number of characters included in the second interactive text. For example, the preset length is set to 15. Usually, the number of characters can be the total number of various types of characters, such as the total number of characters of types such as Chinese characters and English letters.
[0130] Step 2032, determine whether the second interactive text matches the complete text in the preset complete text set.
[0131] Among them, the complete text set can be a pre-set set containing a large number of semantically complete texts for human-computer interaction. If the second interactive text matches a certain complete text, it is determined that the semantics of the second interactive text is complete. The complete text set can also be called a whitelist.
[0132] The method for determining whether the second interactive text matches the complete text can be implemented by various current methods. For example, the method of determining the similarity between texts can be used. If the similarity between the second interactive text and a certain complete text is greater than or equal to the similarity threshold, it is determined that the second interactive text matches the complete text. Or, the keyword set of the second interactive text can be extracted, and the keyword set of the second interactive text is compared with the keyword sets included in each complete text respectively. If the keyword set of the second interactive text is exactly the same as the keyword set of a certain complete text or the proportion of the same keywords in the total number of keywords is greater than or equal to the preset proportion, it is determined that the second interactive text matches the complete text.
[0133] Step 2033, determine whether the type of the second interactive text is the target type.
[0134] Among them, the target type can be the type corresponding to the application scenario of this method. For example, if this method is applied in a vehicle control scenario, the target type can be the vehicle control type.
[0135] Step 2034, in response to determining that the text length is less than or equal to the preset length, and / or the second interactive text does not match the complete text in the complete text set, and / or the type of the second interactive text is the target type, determine that the second interactive text meets the integrity judgment condition.
[0136] Specifically, if the text length is less than or equal to a preset length, it is generally considered that the semantics of the text may be incomplete. If it is greater than the preset length, the semantics are generally considered complete. If the second interactive text does not match the complete texts in the complete text set, that is, the second interactive text is not in the whitelist, it is considered that the semantics of the second interactive text may be incomplete and the text integrity needs to be further determined. If the type of the second interactive text is the target type, it is determined that the second interactive text matches the scenario to which the method is applied, and the semantic integrity needs to be further judged. If it is not the target type, it is determined that the second interactive text has nothing to do with human-computer interaction, and the semantic integrity does not need to be judged. The recognition of the second interactive text will not involve human-computer interaction.
[0137] This embodiment provides multiple semantic integrity judgment conditions, which can quickly screen the second interactive text, avoiding semantic integrity analysis for each obtained second interactive text, thus helping to improve the efficiency of obtaining semantically complete texts.
[0138] In some alternative implementation manners, step 2033 may be executed as follows:
[0139] Based on the multimodal information collected for the user, determine whether the type of the second interactive text is the target type. Among them, the multimodal information includes but is not limited to at least one of the following: user image, user video, user voice.
[0140] Based on the multimodal information, type recognition can be combined with the second interactive text. For example, the user image can be recognized to determine the line-of-sight direction feature or face orientation feature of the user; the user video can be recognized to determine the lip movement feature of the user; the user voice can be recognized to determine the voice feature. The respective features are merged with the text features of the second interactive text, and based on a pre-trained classification model, the merged features are classified to determine whether the type of the second interactive text is the target type.
[0141] In this embodiment, by determining whether the type of the second interactive text is the target type based on the multimodal information, the accuracy of determining the type of the second interactive text can be greatly improved, which helps to improve the efficiency of human-computer interaction.
[0142] In some alternative implementation manners, step 203 includes:
[0143] In response to determining that the second interactive text does not meet the semantic integrity judgment condition, recognize the second interactive text to obtain a second text recognition result.
[0144] If the second interaction text does not meet the semantic integrity judgment condition, it means that the second interaction text is already a semantically complete text (for example, the second interaction text matches the complete text in the above complete text set), and there is no need to further perform semantic integrity analysis, and the text can be directly recognized; or, it means that the second interaction text is the text recognized from the chat voice between users and cannot be used for the target type of human-computer interaction (such as vehicle control). In this case, the second interaction text is recognized for other applications.
[0145] In this embodiment, when the second interaction text does not meet the semantic integrity judgment condition, the semantic integrity analysis is not performed on the second interaction text, realizing the screening of the second interaction text, avoiding performing semantic integrity analysis on each obtained second interaction text, and thus helping to improve the efficiency of obtaining semantically complete texts.
[0146] Exemplary Device
[0147] Figure 8 It is a schematic structural diagram of a text recognition device provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device, such as Figure 8 As shown, the text recognition device includes: an acquisition module 801, configured to acquire a first interaction text obtained by a user's human-computer interaction with a target device; a first generation module 802, configured to generate a second interaction text based on the first interaction text; a first determination module 803, configured to perform a semantic integrity judgment step based on the second interaction text to determine whether the second interaction text meets the semantic integrity judgment condition; a second determination module 804, configured to determine the semantic integrity information of the second interaction text in response to the second interaction text meeting the semantic integrity judgment condition; a second generation module 805, configured to obtain a semantically complete third interaction text based on the semantic integrity information and the second interaction text; a first recognition module 806, configured to recognize the third interaction text to obtain a first text recognition result.
[0148] In this embodiment, the acquisition module 801 can acquire a first interaction text obtained by a user's human-computer interaction with a target device. Among them, the target device can be various devices supporting human-computer interaction. For example, when this embodiment is applied to a vehicle, the target device can be an in-vehicle controller. When this embodiment is applied to household appliances, the target device can be a smart TV, a smart air conditioner and other household appliances. It should be noted that the electronic device executing this embodiment can be the same device as the target device, or can be a device separated from the target device. For example, the target device can be a controller on a vehicle, or a remote server communicatively connected to the controller.
[0149] The first interaction text is the text obtained by the target device recognizing the instructions issued by the user (also known as query). The instructions issued by the user can be of various types, such as voice, gesture, line of sight direction, etc. As an example, when the user conducts human-computer interaction with the target device through voice, ASR (Automatic Speech Recognition) technology can be used for speech recognition to obtain the first interaction text.
[0150] In this embodiment, the first generation module 802 can generate a second interaction text based on the first interaction text. Specifically, as an example, when the user first conducts human-computer interaction with the target device, the first interaction text can be directly determined as the second interaction text. When the user has conducted at least one interaction with the target device, that is, the text recognition method has been executed at least once, if the semantic complete third interaction text cannot be obtained based on the second interaction text when the text recognition method was previously executed, the previously obtained historical interaction text with incomplete semantics will be cached in the preset storage area. When the text recognition method is executed this time, if there is historical interaction text cached in the preset storage area, the first interaction text and the historical interaction text can be merged into the second interaction text.
[0151] In this embodiment, the first determination module 803 can determine whether the second interaction text meets the semantic integrity judgment condition based on the second interaction text. Among them, the semantic integrity judgment condition is a condition set in advance for screening the semantic integrity of the second interaction text. If the second interaction text meets the semantic integrity judgment condition, further analysis of the semantic integrity of the second interaction text is required. If it does not meet the semantic integrity judgment condition, the second interaction text can be directly recognized to determine the user's intention.
[0152] The semantic integrity judgment condition can be set as needed. For example, it can be determined whether the number of characters included in the second interaction text exceeds the preset number. If it exceeds the preset number, it does not meet the semantic integrity judgment condition. At this time, it is determined that the semantics of the second interaction text is complete, and the second interaction text is directly recognized; if it does not exceed the preset number, further semantic integrity analysis of the second interaction text is performed.
[0153] In this embodiment, the second determination module 804 can determine the semantic integrity information of the second interaction text in response to the second interaction text meeting the semantic integrity judgment condition. Among them, the semantic integrity information is used to characterize whether the semantics of the second interaction text is complete. Existing methods can be used for semantic integrity analysis of the second interaction text. For example, a semantic integrity analysis model based on TCN (Temporal Convolutional Network).
[0154] In this embodiment, the second generation module 805 may obtain a semantically complete third interaction text based on the semantic integrity information and the second interaction text. Specifically, as an example, when the semantic integrity information indicates that the semantics of the second interaction text is complete, the second interaction text may be determined as the semantically complete third interaction text. When the semantic integrity information indicates that the semantics of the second interaction text is incomplete, the text obtained from the human-machine interaction with the above user may be continuously acquired, merged with the second interaction text, and then it is determined again whether the merged text meets the semantic integrity judgment condition and whether the semantics is complete until a semantically complete third interaction text is obtained.
[0155] In this embodiment, the first recognition module 806 may recognize the third interaction text to obtain a first text recognition result. The first text recognition result is used to characterize the user's intention. The method for recognizing the third interaction text may adopt existing methods, such as recognition based on NLU (Natural Language Understanding) technology.
[0156] Continuing with the above example, performing NLU recognition on the third interaction text "The air conditioner temperature is one degree higher", the first text recognition result obtained may include intention information indicating controlling the air conditioner to adjust the temperature and slot information indicating that the temperature is increased by one degree. Subsequently, the target device controls the air conditioner temperature to increase by one degree.
[0157] Refer to Figure 9 , Figure 9 which is a schematic structural diagram of a text recognition device provided by another exemplary embodiment of the present disclosure.
[0158] In some optional implementation manners, the second generation module 805 includes: a first determination unit 8051, configured to determine the second interaction text as the third interaction text in response to the semantic integrity information indicating that the semantics of the second interaction text is complete; a parsing unit 8052, configured to perform semantic parsing on the second interaction text to obtain first semantic parsing information in response to the semantic integrity information indicating that the semantics of the second interaction text is incomplete; a question-and-answer unit 8053, configured to perform at least one round of question-and-answer recognition operations with the user based on the first semantic parsing information and the second interaction text to obtain user feedback texts respectively corresponding to the at least one round of question-and-answer recognition operations; a first generation unit 8054, configured to generate the third interaction text based on the user feedback texts respectively corresponding to the at least one round of question-and-answer recognition operations and the second interaction text.
[0159] In some alternative implementation manners, the first generation unit 8054 includes: a first generation subunit 80541, configured to generate a fourth interaction text to be subject to semantic integrity recognition based on the user feedback text and the second interaction text respectively corresponding to at least one round of question-and-answer recognition operations; a first determination subunit 80542, configured to determine the fourth interaction text as the third interaction text in response to determining that the semantics of the fourth interaction text is complete; a storage subunit 80543, configured to store the fourth interaction text in a preset storage area in response to determining that the semantics of the fourth interaction text is incomplete; an acquisition subunit 80544, configured to acquire a fifth interaction text obtained by the user interacting with the target device again; a merging subunit 80545, configured to merge the fifth interaction text with the text in the preset storage area, and update the second interaction text based on the merged text; a second generation subunit 80546, configured to continue to execute the semantic integrity judgment step based on the updated second interaction text until a semantically complete third interaction text is obtained.
[0160] In some alternative implementation manners, the first generation subunit 80541 is further configured to: determine whether the number of rounds of the question-and-answer recognition operations to be executed is less than or equal to a preset number of rounds; in response to being less than or equal to the preset number of rounds, execute the question-and-answer recognition operations; merge the obtained user feedback text and the second interaction text to generate a sixth interaction text to be subject to semantic parsing; perform semantic parsing on the sixth interaction text to obtain second semantic parsing information; determine whether the sixth interaction text meets the interaction intention condition based on the second semantic parsing information; in response to determining that the sixth interaction text meets the interaction intention condition, determine the sixth interaction text as the fourth interaction text; in response to determining that the sixth interaction text does not meet the interaction intention condition, continue to execute the next round of question-and-answer recognition operations.
[0161] In some alternative implementation manners, the first generation subunit 80541 is further configured to: perform intention recognition and / or slot recognition on the sixth interaction text to obtain an intention recognition result and / or a slot recognition result; in response to the intention recognition result indicating that the user has a human-computer interaction intention, and / or the slot recognition result indicating that the slots in the sixth interaction text are complete, determine that the sixth interaction text meets the interaction intention condition.
[0162] In some alternative implementation manners, the first generation subunit 80541 is further configured to: in response to the number of rounds of the question-and-answer recognition operations to be executed being greater than the preset number of rounds, determine the most recently obtained sixth interaction text as the fourth interaction text.
[0163] In some alternative implementation manners, the first generation module 802 includes: a second determination unit 8021, configured to determine whether there is pre-cached text in the preset storage area; a merging unit 8022, configured to merge the first interaction text and the pre-cached text into the second interaction text in response to determining that there is pre-cached text.
[0164] In some alternative implementations, the first determination module 803 includes: a third determination unit 8031, configured to determine whether the text length of the second interaction text is less than or equal to a preset length; a fourth determination unit 8032, configured to determine whether the second interaction text matches a complete text in a preset complete text set; a fifth determination unit 8033, configured to determine whether the type of the second interaction text is a target type; a sixth determination unit 8034, configured to determine that the second interaction text meets the integrity judgment condition in response to determining that the text length is less than or equal to the preset length, and / or the second interaction text does not match the complete text in the complete text set, and / or the type of the second interaction text is the target type.
[0165] In some alternative implementations, the fifth determination unit 8033 is further configured to: based on multi-modal information collected for the user, determine whether the type of the second interaction text is a target type, where the multi-modal information includes at least one of the following: user image, user video, user voice.
[0166] In some alternative implementations, the apparatus further includes: a second recognition module 807, configured to recognize the second interaction text to obtain a second text recognition result in response to determining that the second interaction text does not meet the semantic integrity judgment condition.
[0167] The text recognition apparatus provided in the above embodiments of the present disclosure generates a second interaction text based on the first interaction text obtained by a human-computer interaction between a user and a target device, then determines whether the second interaction text meets the semantic integrity judgment condition. If it meets, the semantic integrity information of the second interaction text is determined, and then based on the semantic integrity information and the second interaction text, a semantically complete third interaction text is obtained. Finally, the third interaction text is recognized to obtain a first text recognition result. It realizes not only recognizing the features of the interaction text, but also obtaining a semantically complete interaction text through means such as multi-round interaction based on the interaction text obtained by the current human-computer interaction, thereby improving the accuracy of semantic integrity judgment, and being able to obtain a semantically complete text for user intention recognition, effectively reducing the risk of recognition errors caused by incomplete semantics of the user during human-computer interaction.
[0168] Exemplary Electronic Device
[0169] Next, refer to Figure 10 to describe an electronic device according to an embodiment of the present disclosure. The electronic device may be any one or both of the terminal device 101 and the server 103 as shown in Figure 1 or a stand-alone device independent of them, and the stand-alone device can communicate with the terminal device 101 and the server 103 to receive the input signals collected from them.
[0170] Figure 10 The block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0171] As Figure 10 shown, the electronic device 1000 includes one or more processors 1001 and a memory 1002.
[0172] The processor 1001 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1000 to perform desired functions.
[0173] The memory 1002 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1001 may run the program instructions to implement the text recognition method of various embodiments of the present disclosure above and / or other desired functions. Various contents such as interactive text, first text recognition result, etc. may also be stored in the computer-readable storage medium.
[0174] In one example, the electronic device 1000 may further include: an input device 1003 and an output device 1004, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0175] For example, when the electronic device is a terminal device 101 or a server 103, the input device 1003 may be a device such as a microphone, camera, mouse, keyboard, etc., for inputting contents such as voice, image, text, etc. to be recognized during human-computer interaction. When the electronic device is a stand-alone device, the input device 1003 may be a communication network connector for receiving the input contents such as voice, image, text, etc. from the terminal device 101 and the server 103.
[0176] The output device 1004 may output various information to the outside, including the text recognition result. The output device 1004 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0177] Of course, for simplicity, Figure 10Only some of the components related to the present disclosure in the electronic device 1000 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 1000 may further include any other appropriate components.
[0178] Exemplary Computer Program Product and Computer Readable Storage Medium
[0179] In addition to the above methods and devices, embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when run by a processor, cause the processor to execute the steps in the text recognition methods according to various embodiments of the present disclosure described in the above "Exemplary Methods" section of this specification.
[0180] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0181] Furthermore, embodiments of the present disclosure may also be computer-readable storage media, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the text recognition methods according to various embodiments of the present disclosure described in the above "Exemplary Methods" section of this specification.
[0182] The computer-readable storage media may adopt any combination of one or more readable media. The readable media may be a readable signal media or a readable storage media. The readable storage media may, for example, include but are not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage media include: electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0183] The basic principles of the present disclosure have been described in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for illustrative and facilitating understanding purposes, rather than limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0184] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and reference can be made to the relevant part of the method embodiment for the relevant content.
[0185] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.
[0186] The methods and apparatuses of the present disclosure can be implemented in many ways. For example, the methods and apparatuses of the present disclosure can be implemented through software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the method is only for illustration, and the steps of the method of the present disclosure are not limited to the above specifically described order, unless otherwise specifically stated in other ways. In addition, in some embodiments, the present disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers the recording medium storing the programs for executing the methods according to the present disclosure.
[0187] It should also be noted that in the apparatuses, equipment, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0188] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0189] The above description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A text recognition method, include: Acquire a first interactive text obtained by human-computer interaction between a user and a target device; Based on the first interactive text, generate a second interactive text; Based on the second interactive text, executing a semantic integrity judgment step to determine whether the second interactive text meets a semantic integrity judgment condition, wherein the semantic integrity judgment condition is a condition for pre-screening and setting the semantic integrity of the second interactive text; In response to the second interactive text meeting the semantic integrity judgment condition, determining semantic integrity information of the second interactive text; Based on the semantic integrity information and the second interactive text, obtaining a semantically complete third interactive text; The third interactive text is recognized to obtain a first text recognition result.
2. The method according to claim 1, in, The obtaining, based on the semantic integrity information and the second interactive text, a semantically complete third interactive text includes: In response to the semantic integrity information indicating that the second interactive text is semantically complete, determining the second interactive text as a third interactive text; In response to the semantic integrity information indicating that the second interactive text is semantically incomplete, performing semantic parsing on the second interactive text to obtain first semantic parsing information; Based on the first semantic parsing information and the second interactive text, performing at least one round of question and answer recognition operation with the user, and obtaining user feedback texts corresponding to the at least one round of question and answer recognition operation; The third interaction text is generated based on the user feedback texts and the second interaction texts respectively corresponding to the at least one round of question and answer recognition operations.
3. The method according to claim 2, in, The generating the third interactive text based on the user feedback text and the second interactive text respectively corresponding to the at least one round of question and answer recognition operation comprises: Generate a fourth interactive text to be subjected to semantic integrity recognition based on the user feedback text and the second interactive text respectively corresponding to the at least one round of question and answer recognition operation; In response to determining that the fourth interactive text is semantically complete, determining the fourth interactive text as the third interactive text; In response to determining that the semantics of the fourth interactive text is incomplete, storing the fourth interactive text in a preset storage area; Acquire a fifth interaction text obtained by the user performing human-computer interaction with the target device again; Merging the fifth interactive text with the text in the preset storage area, and updating the second interactive text based on the merged text; Based on the updated second interactive text, the semantic integrity judgment step is continued until the semantically complete third interactive text is obtained.
4. The method according to claim 3, in, The step of generating a fourth interactive text to be subjected to semantic integrity recognition based on the user feedback text and the second interactive text respectively corresponding to the at least one round of question and answer recognition operation comprises: Determine whether the round number of the question-answer recognition operation to be performed is less than or equal to a preset round number; In response to being less than or equal to the preset round number, performing a question-answer recognition operation; Merge the obtained user feedback text and the second interaction text to generate a sixth interaction text to be semantically parsed; Semantically parse the sixth interaction text to obtain second semantic parsing information; Based on the second semantic parsing information, determine whether the sixth interaction text meets the interaction intention condition; In response to determining that it meets the interaction intention condition, determine the sixth interaction text as the fourth interaction text; In response to determining that it does not meet the interaction intention condition, continue to perform the next round of question-and-answer recognition operation.
5. The method according to claim 4, wherein, The semantically parsing the sixth interaction text to obtain second semantic parsing information includes: Perform intention recognition and / or slot recognition on the sixth interaction text to obtain an intention recognition result and / or a slot recognition result; The determining whether the sixth interaction text meets the interaction intention condition includes: In response to the intention recognition result indicating that the user has a human-computer interaction intention, and / or the slot recognition result indicating that the slots in the sixth interaction text are complete, determine that the sixth interaction text meets the interaction intention condition.
6. The method according to claim 4, wherein, The generating, based on the user feedback text and the second interaction text respectively corresponding to the at least one round of question-and-answer recognition operation, a fourth interaction text to be semantically integrity recognized further includes: In response to the number of rounds of the question-and-answer recognition operation to be executed being greater than a preset number of rounds, determine the sixth interaction text obtained by the most recent execution of the question-and-answer recognition operation as the fourth interaction text.
7. The method according to claim 1, wherein, The generating the second interaction text based on the first interaction text includes: Determine whether there is a pre-cached text in a preset storage area; In response to determining that there is a pre-cached text, merge the first interaction text and the pre-cached text into the second interaction text.
8. The method according to any one of claims 1-7, wherein, The determining whether the second interaction text meets the semantic integrity judgment condition includes: Determine whether the text length of the second interaction text is less than or equal to a preset length; Determine whether the second interaction text matches a complete text in a preset set of complete texts; Determine whether the type of the second interaction text is a target type; In response to determining that the text length is less than or equal to the preset length, and / or the second interaction text does not match the complete text in the set of complete texts, and / or the type of the second interaction text is the target type, determine that the second interaction text meets the integrity judgment condition.
9. The method according to any one of claims 1-7, wherein, The determining whether the second interaction text meets the semantic integrity judgment condition includes: In response to determining that the second interaction text does not meet the semantic integrity judgment condition, perform recognition on the second interaction text to obtain a second text recognition result.
10. A text recognition device, comprising: An acquisition module for acquiring a first interaction text obtained by a user's human-computer interaction with a target device; A first generation module, configured to generate a second interaction text based on the first interaction text; A first determination module, configured to perform a semantic integrity judgment step based on the second interaction text to determine whether the second interaction text meets the semantic integrity judgment condition, where the semantic integrity judgment condition is a condition pre-screened and set for the semantic integrity of the second interaction text; A second determination module, configured to determine the semantic integrity information of the second interaction text in response to the second interaction text meeting the semantic integrity judgment condition; A second generation module, configured to obtain a semantically complete third interaction text based on the semantic integrity information and the second interaction text; A first recognition module, configured to recognize the third interaction text to obtain a first text recognition result.
11. A computer-readable storage medium storing a computer program, where the computer program is executed by a processor to implement the method according to any one of claims 1-9 above.
12. An electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor, configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1-9 above.
Citation Information
Patent Citations
Intelligent dialogue method and device based on text intention recognition and storage medium
CN109840276A
Multi-modal semantic integrity recognition method and device and electronic equipment
CN112101045A