Intention recognition method and device based on generative pre-training GPT model

By combining the preset intent recognition model and the generative pre-trained GPT model in the intention recognition system, the problem of low intent recognition accuracy in the prior art is solved, and the device control efficiency and user experience are improved.

CN120126467APending Publication Date: 2025-06-10QINGDAO HAIER TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311683900.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The accuracy of intention recognition in the prior art leads to low equipment control efficiency.

Method used

The intent recognition method based on the generative pre-trained GPT model is adopted to initially identify user intentions through the pre-set intent recognition model, and when the identification is inaccurate, the generative pre-trained GPT model is used to correct user intention information carrying voice information and scene information.

Benefits of technology

Improves the accuracy of intention recognition and device control efficiency, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126467A_ABST
    Figure CN120126467A_ABST
Patent Text Reader

Abstract

The invention discloses an intention recognition method and device based on a generative pre-training GPT model, and the method comprises the steps: obtaining target voice information sent by a user in a target scene; control intention recognition is conducted on the target voice information through a preset intention recognition model, an initial user intention is obtained, and the preset intention recognition model is constructed according to equipment operation allowed to be executed by equipment in the target scene; and calling a generative pre-training GPT model to perform user intention correction on inquiry information carrying the target voice information and the scene information of the target scene under the condition that it is determined that the preset intention recognition model does not recognize the correct control intention of the user on the equipment in the target scene according to the initial user intention, thereby obtaining the target user intention. By adopting the technical scheme, the problem that the control efficiency of the equipment is relatively low due to low intention recognition accuracy in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of smart home, and in particular, to a method and device for intent recognition based on the generative pre-trained GPT model. Background Art

[0002] With the rapid development of Internet of Things technology, in daily life, replacing traditional home appliances with smart home appliances can optimize people's lifestyle and enhance the convenience of home life. Intent recognition of user instructions is an important prerequisite for realizing smart home. Currently, the basic methods of intent recognition include deep learning intent recognition and rule template intent recognition, etc. However, most of these methods can only recognize a part of the high-frequency direct intents according to the trained deep learning model or the preset rule template model, and the recognition accuracy of most of the user's indirect intents is relatively low, which in turn leads to low control efficiency of the device.

[0003] In view of the low accuracy of intent recognition in the related art, which in turn leads to low control efficiency of the device, no effective solution has been proposed yet. Summary of the Invention

[0004] The embodiments of this application provide a method and device for intent recognition based on the generative pre-trained GPT model to at least solve the problems such as low accuracy of intent recognition in the related art, which in turn leads to low control efficiency of the device.

[0005] According to an embodiment of the embodiments of this application, a method for intent recognition based on the generative pre-trained GPT model is provided, including: obtaining target voice information sent by a user in a target scenario, where one or more devices that are allowed to be controlled are deployed in the target scenario;

[0006] Performing control intent recognition on the target voice information through a preset intent recognition model to obtain an initial user intent, where the preset intent recognition model is constructed according to device operations allowed to be performed by the devices in the target scenario, and the initial user intent is used to indicate the initial control intent of the user for the devices in the target scenario recognized by the preset intent recognition model;

[0007] In the case where it is determined according to the initial user intent that the preset intent recognition model fails to recognize the correct control intent of the user for the devices in the target scenario, calling the generative pre-trained GPT model to correct the user intent for an inquiry message carrying the target voice information and the scenario information of the target scenario to obtain a target user intent, where the generative pre-trained GPT model is a generative language model trained using multi-domain information.

[0008] In an exemplary embodiment, calling the generative pre-trained GPT model to correct the user intention of the query information carrying the target voice information and the scene information of the target scene to obtain the target user intention includes: obtaining the environmental information of the target scene and the device information of the devices deployed in the target scene as the scene information, where the device information includes: device identifier, device location, and device status; converting the target voice information into target text information; constructing the query information according to the target text information, the environmental information, and the device information; and inputting the query information into the generative pre-trained GPT model to obtain the target user intention output by the generative pre-trained GPT model.

[0009] In an exemplary embodiment, the constructing the query information according to the target text information, the environmental information, and the device information includes: splicing the environmental information with the environmental information description to form an environmental field, splicing the device information with the device information description to form a device field, splicing the target text information with the command description to form a command field, and generating a query question corresponding to the initial user intention; and splicing the environmental field, the device field, the command field, and the query question to form the query information.

[0010] In an exemplary embodiment, before calling the generative pre-trained GPT model to correct the user intention of the query information carrying the target voice information and the scene information of the target scene to obtain the target user intention, the method further includes: detecting whether there is an initial control intention recognized by the preset intention recognition model in the initial user intention, where the initial control intention is an intention for expressing the user's control of the devices in the target scene; in the case where it is detected that there is no initial control intention in the initial user intention, determining that the preset intention recognition model fails to recognize the correct control intention of the user for the target scene; in the case where it is detected that there is an initial control intention in the initial user intention, detecting whether the initial control intention is correct; and in the case where it is detected that the initial control intention is incorrect, determining that the preset intention recognition model fails to recognize the correct control intention of the user for the target scene.

[0011] In an exemplary embodiment, the detecting whether the initial control intention is correct includes: controlling the devices in the target scene according to the initial control intention, and collecting the feedback information of the user in the target scene; performing emotion detection on the feedback information; and in the case where it is detected that there is a negative emotion in the feedback information, determining that the initial control intention is incorrect.

[0012] In an exemplary embodiment, the method of correcting a user intention for an inquiry message carrying the target voice information and the scene information of the target scene by invoking the generative pre-trained GPT model to obtain a target user intention includes: constructing the scene information according to the environmental information of the target scene and the device information of the devices deployed in the target scene; generating a first description field for describing the target voice information, a second description field for describing that the devices in the target scene have been controlled according to the initial control intention, and a third description field for describing that there is a negative emotion in the feedback information; constructing interaction information according to the first description field, the second description field and the third description field, where the interaction information is used to indicate an interaction process of interacting with the user using the initial user intention; constructing the inquiry message carrying the scene information and the interaction information; and inputting the inquiry message into the generative pre-trained GPT model to obtain the target user intention output by the generative pre-trained GPT model.

[0013] In an exemplary embodiment, after the method of correcting a user intention for an inquiry message carrying the target voice information and the scene information of the target scene by invoking the generative pre-trained GPT model to obtain a target user intention, the method further includes: when the target user intention is used to indicate that the generative pre-trained GPT model recognizes a target control intention corresponding to the target voice information, controlling the devices in the target scene according to the target user intention; and when the target user intention is used to indicate that the generative pre-trained GPT model does not recognize a control intention of the target voice information, playing a default reply message.

[0014] According to another embodiment of the embodiments of the present application, there is also provided an intention recognition device based on a generative pre-trained GPT model, including:

[0015] An acquisition module, configured to acquire target voice information sent by a user in a target scene, where one or more devices that can be controlled are deployed in the target scene;

[0016] An identification module, configured to perform control intention identification on the target voice information through a preset intention recognition model to obtain an initial user intention, where the preset intention recognition model is constructed according to device operations allowed to be performed by the devices in the target scene, and the initial user intention is used to indicate an initial control intention of the user for the devices in the target scene recognized by the preset intention recognition model;

[0017] A calling module, configured to, when it is determined according to the initial user intention that the preset intention recognition model fails to recognize the correct control intention of the user for the device in the target scenario, call the generative pre-trained GPT model to correct the user intention for the query information carrying the target voice information and the scenario information of the target scenario, so as to obtain the target user intention, where the generative pre-trained GPT model is a generative language model trained using multi-domain information.

[0018] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the above-mentioned intention recognition method based on the generative pre-trained GPT model when running.

[0019] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the above-mentioned processor executes the above-mentioned intention recognition method based on the generative pre-trained GPT model through the computer program.

[0020] Through the present application, the preset intention recognition model is used to recognize the target voice information to obtain the initial user intention. When the initial user intention indicates that the preset intention recognition model fails to recognize the correct control intention of the user, the generative pre-trained GPT model with a question-and-answer function trained using multi-domain information is used to reply to the query information to obtain the target user intention, that is, when the preset intention recognition model fails to recognize the correct control intention of the user, the target user intention that meets the target voice information and the target scenario is obtained by querying the generative pre-trained GPT model. By using the generative pre-trained GPT model after the preset intention recognition model, the coverage range of user intention recognition is increased, thereby improving the accuracy of intention recognition, and thus improving the control efficiency of the device. Therefore, the problem of low accuracy of intention recognition and low control efficiency of the device in the related art can be solved, the accuracy of intention recognition and the control efficiency of the device are improved, and the user experience is enhanced. Description of the Drawings

[0021] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0023] Figure 1 is a hardware structure block diagram of a mobile terminal for an intent recognition method based on the generative pre-trained GPT model according to an embodiment of the present application;

[0024] Figure 2 is a flowchart of an intent recognition method based on the generative pre-trained GPT model according to an embodiment of the present application;

[0025] Figure 3 is a flowchart of a process for recognizing an initial user intent according to an embodiment of the present application;

[0026] Figure 4 is a schematic diagram of a process for detecting a target user intent according to an embodiment of the present application;

[0027] Figure 5 is a schematic diagram of a method for detecting whether an initial user intent includes an initial control intent according to an embodiment of the present application;

[0028] Figure 6 is a schematic diagram of an intent recognition method based on the generative pre-trained GPT model according to an embodiment of the present application;

[0029] Figure 7 is a structure block diagram of an intent recognition device based on the generative pre-trained GPT model according to an embodiment of the present application;

[0030] Figure 8 is a schematic diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] The method embodiments provided by the embodiments of the present application can be executed on a computer terminal, a device terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 1It is a hardware block diagram of a mobile terminal for an intent recognition method based on the generative pre-trained GPT model according to an embodiment of the present application. As Figure 1 shown, the computer terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. In an exemplary embodiment, the above computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only illustrative and does not limit the structure of the above computer terminal. For example, the computer terminal may further include more or fewer components than Figure 1 shown in the figure, or have the same functions as Figure 1 shown in the figure or different configurations with more functions than Figure 1 shown in the figure.

[0033] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the intent recognition method based on the generative pre-trained GPT model in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories may be connected to the computer terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0034] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the computer terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0035] In this embodiment, an intent recognition method based on the generative pre-trained GPT model is provided, which is applied to the above device terminal, Figure 2It is a flowchart of an intent recognition method based on the generative pre-trained GPT model according to an embodiment of the present application. The process includes the following steps:

[0036] Step S202, obtain target voice information sent by the user in a target scenario, where one or more controllable devices are deployed in the target scenario;

[0037] Step S204, perform control intent recognition on the target voice information through a preset intent recognition model to obtain an initial user intent, where the preset intent recognition model is constructed based on device operations allowed to be performed by devices in the target scenario, and the initial user intent is used to indicate the initial control intent of the user on the devices in the target scenario recognized by the preset intent recognition model;

[0038] Step S206, in the case where it is determined according to the initial user intent that the preset intent recognition model fails to recognize the correct control intent of the user on the devices in the target scenario, call the generative pre-trained GPT model to correct the user intent for an inquiry message carrying the target voice information and the scenario information of the target scenario, and obtain a target user intent, where the generative pre-trained GPT model is a generative language model trained using multi-domain information.

[0039] Through the above steps, use the preset intent recognition model to recognize the target voice information to obtain the initial user intent. In the case where the initial user intent indicates that the preset intent recognition model fails to recognize the correct control intent of the user, then use the generative pre-trained GPT model with a question-and-answer function trained using multi-domain information to reply to the inquiry message and obtain the target user intent. That is, in the case where the preset intent recognition model fails to recognize the correct control intent of the user, obtain the target user intent that meets the target voice information and the target scenario by asking the generative pre-trained GPT model. By using the generative pre-trained GPT model after the preset intent recognition model, the coverage of user intent recognition is increased, thereby improving the accuracy of intent recognition and thus improving the control efficiency of the device. Therefore, it is possible to solve the problem in the related art that the accuracy of intent recognition is low, which in turn leads to low control efficiency of the device, improve the accuracy of intent recognition and the control efficiency of the device, and enhance the user experience.

[0040] Optionally, the execution subject of the above steps may be a background processor, or other devices with similar processing capabilities, or may also be a machine integrated with at least a voice acquisition device and a data processing device. The voice acquisition device may include a voice collection module such as a microphone, and the data processing device may include a terminal such as a computer, but is not limited thereto.

[0041] In the technical solution provided in step S202 above, the target voice information can be, but is not limited to, voice control instructions such as voice instructions issued by the user for controlling devices in the target scenario, voice instructions for controlling devices in the target scenario played through a playback device, etc., or can be input control instructions, such as text instructions input by the user in an APP (Application, application software), system, or website on a mobile phone or computer terminal that allows controlling devices in the target scenario, or can also be touch control instructions, such as touch instructions for controlling devices in the target scenario through touch, click, or corresponding options in an APP, system, or website on a mobile phone or computer terminal that allows controlling devices in the target scenario.

[0042] Optionally, in this embodiment, the above devices can be, but are not limited to, intelligent networked devices, such as audio-visual devices, lighting systems, curtains, air conditioners, security systems, digital cinema systems, audio-visual servers, media cabinet systems, network appliances, etc.

[0043] Optionally, in this embodiment, the target scenario can be, but is not limited to, a smart home device scenario that integrates smart devices related to home life using technologies such as integrated wiring technology, network communication technology, security prevention technology, automatic control technology, and audio-visual technology with a residence as the platform, or can be an integrated scenario that uses a smart home system as the platform, with household appliances and home appliance devices as the main control objects, which can enhance home intelligence, security, convenience, comfort, and achieve environmental protection control.

[0044] Optionally, in this embodiment, the user can be, but is not limited to, an authorized group or an authorized person who has the control authority of the devices deployed in the target scenario, etc. It can be, but is not limited to, obtaining information triggered by the user's voice in the target scenario for turning on or off the target device as the target voice information. For example, turning on or off the TV, speaker, air conditioner, etc., adjusting the temperature of the target device, for example, adjusting the water heater temperature to 42 degrees Celsius, adjusting the air conditioner temperature to 26 degrees Celsius, etc., navigation, querying information, etc.

[0045] Optionally, in this embodiment, the target voice information issued by the user in the target scenario may include, but is not limited to, the control intention for one or more devices deployed in the target scenario. For example: at night, when the user issues the target voice information of "It's so dark here", the user may need to turn on the light at this time; during the day, when the user issues the target voice information of "The light is so dazzling", the user may need to draw the curtains at this time, or may also need to increase the screen brightness; when the user issues the target voice information of "It's choking", the user may be cooking at this time and need to turn on the range hood, or there may be someone smoking in the room and need to turn on the air conditioner, purifier and other devices' purification modes; when the user issues the target voice information of "I'm going to take a bath later", there may be a series of control intentions at this time, such as turning on the water heater, raising the water temperature of the water heater, turning on the bathroom heater, and turning on the pre-bath mode of the bathroom heater, etc.

[0046] In the technical solution provided in the above step S204, the preset intention recognition model may be, but is not limited to, a model constructed according to the operation types, operation methods, etc. supported by one or more devices allowed to be controlled in the target scenario, and is used to identify the user's intention. For example: a rule classification model based on a dictionary template that constructs a template through word segmentation, part-of-speech and regular expressions, a statistical classification model that classifies according to the corresponding word frequency and subsequent behaviors through relevant behavior logs, a machine learning-based model that first manually annotates the data and then performs word segmentation, and then combines feature extraction methods such as chi-square test, and then performs feature vectorization, and finally uses classification methods such as SVM (Support Vector Machine), decision tree, etc. for training, or an intention recognition model based on deep learning: such as an LSTM (Long Short-Term Memory)-attention (attention mechanism) model or a BERT (Bidirectional Encoder Representations from Transformer) network classification intention recognition model, etc.

[0047] Optionally, in this embodiment, the control intention may, but is not limited to, be the user's control will, that is, what the user wants to do. The control intention may be a "Dialog Act", that is, the behavior of the information state or context shared by the user in the conversation and continuously updated. The intention may, but is not limited to, include direct intention and indirect intention. The direct intention may, but is not limited to, include the instruction directly issued by the user that contains the operation to be performed on a certain device or a certain behavior to be performed, which may be in the form of "verb + noun" composed of "control operation verb + specific device to be executed" or "control operation verb + target noun", such as home appliance control operations like turning on the air conditioner and turning off the TV, or querying the encyclopedia, querying the weather, etc. The indirect intention may, but is not limited to, be the instruction issued by the user that does not directly contain the operation to be performed on a certain device or a certain behavior to be performed. For example, when the user issues the instruction "It's so dark here", at this time, it may be necessary to turn on the light (at night), it may be necessary to draw the curtains (during the day and when the curtains are open), or it may be that the user is facing the screen and needs to increase the screen brightness; or "I'm choking", at this time, it may be cooking and it is necessary to turn on the range hood. Or when smoking in the room, it is necessary to turn on the purification mode of devices such as the air conditioner and purifier; or "I'm going to take a bath later", which may correspond to a series of intentions, such as turning on the water heater, raising the water temperature of the water heater, turning on the bathroom heater, and turning on the pre-bath mode of the bathroom heater.

[0048] Optionally, in this embodiment, intention recognition may, but is not limited to, classify information for the Query (query, inquiry) input or played by the user, and then perform reasonable operations for the input intention in the next step. Intention recognition is also called intention classification, that is, classifying the user's words into the previously defined intention categories according to the fields and intentions involved in the user's words. Intention recognition is a sub-module of speech understanding and is also the key to the composition of the human-computer dialogue system.

[0049] Optionally, in this embodiment, the control intention recognition of the target voice information may, but is not limited to, include successfully recognizing the control intention in the target voice information, being unable to recognize the control intention in the target voice information, and other error situations, etc. In different situations, the preset intention recognition model may, but is not limited to, be used to output different initial user intentions. For example, when the preset intention recognition model successfully recognizes the control intention, it outputs the initial user intention; when the preset intention recognition model fails to recognize the control intention, it outputs the corresponding reply, etc.

[0050] Optionally, in this embodiment, the preset intention recognition model outputs the initial user intention according to the target voice information issued by the user. For example, when the user issues the target voice information: Help me open the curtains in the living room, and the preset intention recognition model outputs the initial user intention to open the curtains located in the living room, it is considered that the preset intention recognition model has successfully recognized.

[0051] Optionally, in this embodiment, the pre-set intention recognition model may, but is not limited to, fail to recognize the user's correct control intention for the device in the target scenario in various situations, such as:

[0052] The target device for executing the control intention is misrecognized. For example, the user sends the target voice message "Little A is a bit cold". In fact, the control intention to be executed is to increase the temperature of the air conditioner named Little A. However, the pre-set intention recognition model may misrecognize Little A as a person's name and play the song "A Bit Cold" by the singer Little A, resulting in misrecognition of the target device for executing the control intention. Or, the target voice message sent by the user has a dialect, such as "Turn on the air conditioner". Or, there are problems of polysemy, sentence segmentation, and word segmentation. For example, the target voice message sent by the user is "Remind me to buy the boss's vegetables at one o'clock in the afternoon". The pre-set intention recognition model misrecognizes it as "buy the boss's vegetables at one o'clock", causing an error.

[0053] In an alternative implementation manner, an example of the recognition process of the initial user intention is provided. Figure 3 It is a flowchart of the recognition process of the initial user intention according to the embodiment of the present application, as Figure 3 shown. The query sent by the user is used as the target voice message, and the pre-set intention recognition model is used to recognize the user's control intention to obtain the initial user intention.

[0054] In the technical solution provided in step S206 above, the generative pre-trained GPT model is a generative language model trained with multi-domain information. The above multi-domain information may, but is not limited to, include Internet information such as network big data resources. The generative pre-trained GPT model may, but is not limited to, be an integrated model trained using artificial intelligence technology with the function of accepting questions and answering them. For example, Chat GPT (Chat Generative Pre-trained Transformer, generative pre-trained dialogue conversion model). Chat GPT is a deep learning model for text generation trained based on available data on the Internet. Chat GPT is a natural language processing tool driven by artificial intelligence technology. It can conduct conversations by understanding and learning human language and can also interact according to the context of the chat.

[0055] Optionally, in this embodiment, the generative pre-trained GPT model may, but is not limited to, be used for further recognition of the control intention of the target voice message. It may, but is not limited to, add certain prompt information on the basis of the original target voice message as the input of the generative pre-trained GPT model so that the generative pre-trained GPT model can correct the user's intention according to the prompt information. The above prompt information may, but is not limited to, include the scenario information of the target scenario, such as:

[0056] Taking the scenario information of the target scenario including 11 am, room temperature of 28 degrees, and humidity of 40%, and the devices deployed in the target scenario including an air conditioner (located in the master bedroom, alias is Xiao A, in the on state), a water heater (located in the kitchen, in the off state), a speaker (located in the living room, in the on state), and a washing machine (located in the bathroom, in the off state), and the target voice information being "Xiao A, blow cold air" as an example, one can, but is not limited to, use the scenario information and the target voice information to form an inquiry message and input it into the generative pre-trained GPT model, and let the generative pre-trained GPT model infer the user's control intention.

[0057] Or, taking the scenario information of the target scenario including 11 am, room temperature of 28 degrees, and humidity of 40%, and the devices deployed in the target scenario including an air conditioner (located in the master bedroom, alias is Xiao A, in the on state), a water heater (located in the kitchen, in the off state), a speaker (located in the living room, in the on state), and a washing machine (located in the bathroom, in the off state), and the target voice information being "What to do if it rains" as an example, one can, but is not limited to, use the scenario information and the target voice information to form an inquiry message and input it into the generative pre-trained GPT model, and let the generative pre-trained GPT model infer the user's control intention and give a reply.

[0058] Or, taking the scenario information of the target scenario including 11 am, room temperature of 28 degrees, and humidity of 40%, and the devices deployed in the target scenario including an air conditioner (located in the master bedroom, alias is Xiao A, in the on state), a water heater (located in the kitchen, in the off state), a speaker (located in the living room, in the on state), a washing machine (located in the bathroom, in the off state), a light (in the master bedroom, in the off state), and curtains (in the master bedroom, opening degree of 20%), and the target voice information being "It's too dark" as an example, one can, but is not limited to, use the scenario information and the target voice information to form an inquiry message and input it into the generative pre-trained GPT model, and let the generative pre-trained GPT model infer the user's control intention and give a reply.

[0059] Or, taking the scenario information of the target scenario including 11 pm, room temperature of 28 degrees, and humidity of 40%, and the devices deployed in the target scenario including an air conditioner (located in the master bedroom, alias is Xiao A, in the on state), a water heater (located in the kitchen, in the off state), a speaker (located in the living room, in the on state), a washing machine (located in the bathroom, in the off state), a light (in the master bedroom, in the off state), and curtains (in the master bedroom, opening degree of 20%), and the target voice information being "It's too dark" as an example, one can, but is not limited to, use the scenario information and the target voice information to form an inquiry message and input it into the generative pre-trained GPT model, and let the generative pre-trained GPT model infer the user's control intention and give a reply.

[0060] Alternatively, taking the scenario information of the target scenario including 11 PM, room temperature of 28 degrees, humidity of 40%, and the devices deployed in the target scenario including an air conditioner (located in the master bedroom, alias is Xiao A, in the on state), a water heater (located in the kitchen, in the off state), a speaker (located in the living room, in the on state), a washing machine (located in the bathroom, in the off state), a light (in the master bedroom, in the off state), and a curtain (in the master bedroom, with an opening degree of 20%), and the target voice information being "What's that smell? It's a bit choking" as an example, one can, but is not limited to, use the scenario information and the target voice information to form an inquiry message and input it into the generative pre-trained GPT model, and have the generative pre-trained GPT model infer the user's control intention and give a reply, etc.

[0061] In an exemplary embodiment, one can, but is not limited to, call the generative pre-trained GPT model to correct the user intention of the inquiry message carrying the target voice information and the scenario information of the target scenario to obtain the target user intention in the following manner: Obtain the environmental information of the target scenario and the device information of the devices deployed in the target scenario as the scenario information, where the device information includes: device identifier, device location, and device status; convert the target voice information into target text information; construct the inquiry message based on the target text information, the environmental information, and the device information; and input the inquiry message into the generative pre-trained GPT model to obtain the target user intention output by the generative pre-trained GPT model.

[0062] Optionally, in this embodiment, one can, but is not limited to, obtain the information of the target scenario where the user is located as the environmental information. The target scenario can be, but is not limited to, the smart home scenario where the user is located. The environmental information can be, but is not limited to, including information such as the time information, temperature information, humidity information, air quality parameters, etc. of the environment.

[0063] Optionally, in this embodiment, one can, but is not limited to, obtain the device information of each device deployed in the target scenario as the scenario information. The target scenario can be, but is not limited to, the smart home scenario where the user is located. The device information can be, but is not limited to, including the relevant information of all smart home devices in the smart home scenario, and can be, but is not limited to, including the device identifier, device location, device status, etc.

[0064] Optionally, in this embodiment, the devices deployed in the target scenario may but are not limited to having a unique device identifier, and the device identifier may but is not limited to indicating information such as device name, device number, etc. The device identifier may but is not limited to being an identification information composed of strings such as characters and graphics. The corresponding device may be determined from multiple devices deployed in the target scenario according to the device identifier of the device. For example: find the lamp with device number 1 according to device identifier 1, find the air conditioner with device number 2 according to device identifier 2, find the water heater with device number 3 according to device identifier 3, find the range hood with device number 4 according to device identifier 4, find the washing machine with device number 5 according to device identifier 5, and so on.

[0065] Optionally, in this embodiment, the device location may but is not limited to indicating the location of the device in the space where it is located. For example: the device is located in the master bedroom, the device is located in the living room, the device is located in the kitchen, the device is located in the bathroom, and so on. Further, a coordinate system may but is not limited to be established in the target scenario, and the coordinate is used to represent the device location of each device deployed in the target scenario.

[0066] Optionally, in this embodiment, the device status may but is not limited to being information on the current working status of the device determined according to the device characteristic status parameter, and may but is not limited to including that the device is powered on, the device is not powered on, and so on.

[0067] Optionally, in this embodiment, the target voice information sent by the user in the target scenario may but is not limited to be converted into target text information by using natural language processing algorithms. The natural language processing algorithms may but are not limited to including: text classification algorithms, such as deep learning models like Naive Bayes, Support Vector Machine, Convolutional Neural Network, and Recurrent Neural Network; named entity recognition algorithms, such as Conditional Random Field, Sequence Labeling Model, etc.; speech emotion recognition algorithms, such as recurrent neural network based on emotion recognition, etc.; speech recognition algorithms, such as Hidden Markov Model, Long Short-Term Memory Network, Transcription Attention Model, etc. based on acoustic model and language model.

[0068] Optionally, in this embodiment, the above query information may but is not limited to be constructed according to the language form recognizable by the Generative Pretrained Transformer (GPT) model. The Generative Pretrained Transformer (GPT) model may but is not limited to perform operations such as correcting and verifying the initial user intention output by the preset intention recognition model according to the target text information, environmental information, and device information carried by the query information.

[0069] Further, it is possible but not limited to carry the operations that need to be performed by the generative pre-trained GPT model in the query information. For example, in the case where the generative pre-trained GPT model needs to correct the initial user intention output by the preset intention recognition model, information indicating the correction operation is carried in the query information; in the case where the generative pre-trained GPT model needs to verify the initial user intention output by the preset intention recognition model, information indicating the verification operation is carried in the query information, etc.

[0070] Optionally, in this embodiment, in the case where the initial user intention output by the preset intention recognition model cannot correctly express the user's control intention for the device, it is possible but not limited to use the generative pre-trained GPT model to further correct the initial user intention output by the preset intention recognition model, and output the target user intention that correctly expresses the user's control intention for the device.

[0071] In an exemplary embodiment, it is possible but not limited to construct the query information according to the target text information, the environment information, and the device information in the following manner: concatenate the environment information with the environment information description to form an environment field, concatenate the device information with the device information description to form a device field, concatenate the target text information with the command description to form a command field, and generate the query question corresponding to the initial user intention; concatenate the environment field, the device field, the command field, and the query question to form the query information.

[0072] Optionally, in this embodiment, it is possible but not limited to set a description statement as the above-mentioned environment information description according to the language form supported by the generative pre-trained GPT model. The above description statement can be, but is not limited to, a sentence pattern composed of one or more words for describing environment information, such as: The current time is..., The current temperature is..., The current humidity is..., The current air quality parameter is..., etc.

[0073] Optionally, in this embodiment, it is possible but not limited to set a description statement as the above-mentioned device information description according to the language form supported by the generative pre-trained GPT model. The above description statement can be, but is not limited to, a sentence pattern composed of one or more words for describing device information, such as: The device identifier of... device is..., The device location of... device is..., The device status of... device is..., etc.

[0074] Optionally, in this embodiment, it is possible but not limited to set a description statement as the above-mentioned command description according to the language form supported by the generative pre-trained GPT model. The above description statement can be, but is not limited to, a fixed sentence pattern composed of one or more words for describing the target text information, such as: The user is saying... now, The user said..., A certain user said..., etc.

[0075] Optionally, in this embodiment, taking 11 am, room temperature of 28 degrees, and humidity of 40% as environmental information, and using "It is now..." as an example of environmental information description, the environmental information and the environmental information description are concatenated to obtain an environmental field: It is now 11 am, room temperature is 28 degrees, and humidity is 40%.

[0076] Optionally, in this embodiment, taking an air conditioner with the nickname of "Xia A", located in the master bedroom and not powered on as device information, and using "There is a... device" as an example of device information description, the device information and the device information description are concatenated to obtain a device field: There is an air conditioner device with the nickname of "Xia A", located in the master bedroom and not powered on.

[0077] Optionally, in this embodiment, taking "Xia A blows cold air" as the target text information, and using "The user now says..." as an example of command description, the target text information and the command description are concatenated to obtain a command field: The user now says Xia A blows cold air.

[0078] Optionally, in this embodiment, the generative pre-trained GPT model is a model with the function of accepting questions and giving answers. Therefore, it can, but is not limited to, generate inquiry questions that conform to the language form supported by the generative pre-trained GPT model according to the questions that need to be answered by the generative pre-trained GPT model. The inquiry questions can, but are not limited to, be sentence patterns that can be parsed by the generative pre-trained GPT model composed of one or more words. For example, carrying the inquiry question "Speculate the user's intention" in the inquiry information to instruct the generative pre-trained GPT model to execute the operation of speculating the user's intention; carrying the inquiry question "Speculate the user's intention and correct directly if there is an error" in the inquiry information to instruct the generative pre-trained GPT model to execute the operation of speculating the user's intention and correct directly if there is an error; carrying the inquiry question "We may have misunderstood the user's intention. Please help us correct it and output the correct intention" in the inquiry information to instruct the generative pre-trained GPT model to execute the operation of correcting and outputting the correct intention, etc.

[0079] Optionally, in this embodiment, taking the environmental field: It is now 11 am, room temperature is 28 degrees, and humidity is 40%; the device field: There is an air conditioner device with the nickname of "Xia A", located in the master bedroom and not powered on; the command field: The user now says Xia A blows cold air; and the inquiry question: Speculate the user's intention as an example, the environmental field, the device field, the command field, and the inquiry question are concatenated into an inquiry information: It is now 11 am, room temperature is 28 degrees, and humidity is 40%, there is an air conditioner device with the nickname of "Xia A", located in the master bedroom and not powered on, the user now says Xia A blows cold air, speculate the user's intention.

[0080] Further, in order to enable the generative pre-trained GPT model to answer the input query information more accurately, it is possible but not limited to annotate the query information to obtain the query information: It is 11 am now, the room temperature is 28 degrees, and the humidity is 40% [Environment field] There is an air conditioning device, with the nickname of Xiao A, located in the master bedroom and not powered on [Device field] The user now says that Xiao A blows cold air [Command field] Infer the user's intention [Query question]. Or, [Environment field] It is 11 am now, the room temperature is 28 degrees, and the humidity is 40% [Device field] There is an air conditioning device, with the nickname of Xiao A, located in the master bedroom and not powered on [Command field] The user now says that Xiao A blows cold air [Query question] Infer the user's intention.

[0081] In an exemplary embodiment, before calling the generative pre-trained GPT model to correct the user intention of the query information carrying the target voice information and the scene information of the target scene to obtain the target user intention, it is possible but not limited to detect by the following methods: Detect whether there is an initial control intention recognized by the preset intention recognition model in the initial user intention, where the initial control intention is an intention used to express the user's control of the device in the target scene; in the case where it is detected that there is no such initial control intention in the initial user intention, determine that the preset intention recognition model has not recognized the correct control intention of the user for the target scene; in the case where it is detected that there is such initial control intention in the initial user intention, detect whether the initial control intention is correct; in the case where it is detected that the initial control intention is incorrect, determine that the preset intention recognition model has not recognized the correct control intention of the user for the target scene.

[0082] Optionally, in this embodiment, it is possible but not limited to use the generative pre-trained GPT model to verify the initial user intention recognized by the preset intention recognition model. For example: input the initial user intention output by the preset intention recognition model and the query information into the generative pre-trained GPT model, and the above query information is used to instruct the generative pre-trained GPT model to detect whether there is an initial control intention recognized by the preset intention recognition model, and the detection is performed by the generative pre-trained GPT model. Or, input the initial user intention output by the preset intention recognition model and the query information into other devices with detection functions, and the detection is performed by the devices with detection functions, etc.

[0083] Optionally, in this embodiment, the initial control intention is an intention used to express the user's control of the device in the target scene. It is possible but not limited to determine whether the initial user intention output by the preset intention recognition model includes an operation capable of controlling the device deployed in the target scene, so as to determine whether the preset intention recognition model has recognized the initial control intention. For example:

[0084] When the initial user intention output by the preset intention recognition model includes an operation capable of controlling the devices deployed in the target scenario, it is determined that the preset intention recognition model has recognized the initial control intention; or, when the initial user intention output by the preset intention recognition model does not include an operation capable of controlling the devices deployed in the target scenario, it is determined that the preset intention recognition model has not recognized the initial control intention, etc.

[0085] In an exemplary embodiment, the initial control intention can be detected for correctness in the following ways, but not limited to: controlling the devices in the target scenario according to the initial control intention, and collecting the feedback information of the user in the target scenario; performing emotion detection on the feedback information; when negative emotion is detected in the feedback information, determining that the initial control intention is incorrect.

[0086] Optionally, in this embodiment, the feedback information can be, but not limited to, the reaction of the user to the execution of the initial control intention, and the feedback information can be, but not limited to, including: voice feedback, text feedback, etc. When an incorrect operation is performed, the user can, but not limited to, send feedback information with negative emotion, such as: How can you not understand this? How can you not know this? etc.

[0087] Optionally, in this embodiment, the feedback information can be detected for emotion in the following ways, but not limited to: using an algorithm based on rules for emotion detection, an emotion detection model relying on machine learning, etc. For example: the method based on a dictionary, analyzing the words in the feedback information using an emotion dictionary, and judging the emotional tendency of the feedback information according to the emotional polarity (such as positive, negative or neutral, etc.) of the words in the emotion dictionary; training an emotion classifier using a supervised learning algorithm to identify the emotion in the feedback information; using a model based on a recurrent neural network, a long short-term memory network model, a model based on an attention mechanism to identify the emotion in the feedback information, etc.

[0088] Optionally, in this embodiment, the devices in the target scenario can be controlled according to the initial control intention, and the feedback information of the user can be collected in real time, such as: obtaining the voice feedback information sent by the user through a device with a voice receiving function, receiving the text feedback information sent by the user through an APP, a system or a website, etc.

[0089] Optionally, in this embodiment, it is possible but not limited to determine whether there is an error in the initial control intention according to the user emotion existing in the feedback information. For example, when the feedback information indicates that the user is in a positive emotion or has no emotion, it is determined that the initial control intention is correct. At this time, it is determined that the initial user intention is successfully recognized, and there is no need to query the generative pre-trained GPT model. Instead, the initial user intention is directly used to execute the target voice information. When the feedback information indicates that the user has a negative emotion, it is determined that there is an error in the initial control intention, and it is determined that the initial user intention is misrecognized but the initial control intention has been executed. It is necessary to use the query information to query the generative pre-trained GPT model to obtain the reply information corresponding to the query information, so as to use the reply information to execute the target control intention again.

[0090] In an exemplary embodiment, it is possible but not limited to call the generative pre-trained GPT model to correct the user intention of the query information carrying the target voice information and the scene information of the target scene in the following manner to obtain the target user intention: construct the scene information according to the environmental information of the target scene and the device information of the devices deployed in the target scene; generate a first description field for describing the target voice information, a second description field for describing that the devices in the target scene have been controlled according to the initial control intention, and a third description field for describing the existence of negative emotion in the feedback information; construct interaction information according to the first description field, the second description field, and the third description field, where the interaction information is used to indicate the interaction process of interacting with the user using the initial user intention; construct the query information carrying the scene information and the interaction information; input the query information into the generative pre-trained GPT model to obtain the target user intention output by the generative pre-trained GPT model.

[0091] Optionally, in this embodiment, it is possible but not limited to reorganize the environmental information of the target scene and the device information of the devices deployed in the target scene to obtain the scene information.

[0092] Optionally, in this embodiment, it is possible but not limited to regenerate the first description field for describing the target voice information, or convert the target voice information into field information as the first description field.

[0093] Optionally, in this embodiment, it is possible but not limited to regenerate the second description field for describing that the devices in the target scene have been controlled according to the initial control intention, or obtain the information corresponding to the control of the devices in the target scene according to the initial control intention as the first description field.

[0094] Optionally, in this embodiment, it is possible but not limited to regenerate a third description field for describing the negative emotion in the feedback information, or obtain the information corresponding to the negative emotion as the third description field.

[0095] Optionally, in this embodiment, it is possible but not limited to combine the first description field, the second description field, and the third description field to obtain interaction information, or re-arrange the first description field, the second description field, and the third description field to obtain a statement that can describe the interaction process of interacting with the user using the initial user intention as the interaction information.

[0096] Optionally, in this embodiment, the query information carries scenario information and interaction information. It is possible but not limited to combine the scenario information and the interaction information as the query information, such as: [Scenario Information][Interaction Information]. Or, [Interaction Information][Scenario Information], etc.

[0097] In an alternative embodiment, an example of a process for detecting the target user intention is provided. Figure 4 It is a schematic diagram of a process for detecting the target user intention according to an embodiment of the present application, as Figure 4 shown. In the target scenario, an air conditioner (located in the master bedroom, alias is Xiao A, in standby state), a water heater (located in the kitchen, in shutdown state), a stereo (located in the living room, in on state), a washing machine (located in the bathroom, in standby state) are deployed, and the time is 11 am, the room temperature is 28 degrees, and the humidity is 40%. The user issues a target voice message in the target scenario: Xiao A, blow cold air. It is possible but not limited to detect the target user intention from the target voice message in the following ways:

[0098] First, use a pre-set intention recognition model to perform control intention recognition on the target voice message to obtain the initial user intention, and then detect whether there is an initial control intention recognized by the pre-set intention recognition model for expressing the user's control of the device in the target scenario in the initial user intention;

[0099] In the case where it is detected that there is no initial control intention in the initial user intention, or the detected initial control intention is incorrect, use the target text information, environmental information, and device information to construct query information, and input the query information into the generative pre-trained GPT model to obtain the target user intention output by the generative pre-trained GPT model.

[0100] In an exemplary embodiment, after the generative pre-trained GPT model is called to correct the user intention of the query information carrying the target voice information and the scene information of the target scene to obtain the target user intention, the target user intention can be responded to in the following ways, but not limited to: when the target user intention is used to indicate that the generative pre-trained GPT model recognizes the target control intention corresponding to the target voice information, control the device in the target scene according to the target user intention; when the target user intention is used to indicate that the generative pre-trained GPT model does not recognize the control intention of the target voice information, play the default reply information.

[0101] Optionally, in this embodiment, the semantics included in the target user intention can be parsed, but not limited to, to determine whether the generative pre-trained GPT model recognizes the target control intention corresponding to the target voice information. If the semantics included in the target user intention is an instruction for a certain operation on the device in the target scene, it can be determined that the generative pre-trained GPT model recognizes the target control intention, and the device in the target scene is controlled according to the target control intention included in the semantics.

[0102] Optionally, in this embodiment, if the semantics included in the target user intention is not an instruction for a certain operation on the device in the target scene, it can be determined that the generative pre-trained GPT model does not recognize the target control intention. At this time, according to the default settings, the system or platform deploying the above method based on the generative pre-trained GPT model can be controlled to play the default reply information, such as "Sorry, I didn't understand. Can you say it again?" and so on.

[0103] In an alternative embodiment, an example of a method for detecting whether the initial user intention includes an initial control intention is provided. Figure 5 It is a schematic diagram of a method for detecting whether the initial user intention includes an initial control intention according to an embodiment of the present application. As Figure 5 shown, it can be used, but not limited to, to detect whether there is an initial control intention recognized by the preset intention recognition model in the initial user intention using the entire semantic coverage range. The specific method is as follows:

[0104] The target voice information is input into a model capable of making a judgment (which can be, but not limited to, a preset intention recognition model), and the model judges whether the semantic intention included in the target voice information is within the preset semantic coverage range:

[0105] If the semantic intention included in the target voice information is within the semantic coverage range, it represents that the control intention in the target voice information has been covered, and it is determined that there is an initial control intention in the initial user intention;

[0106] If the semantic intention contained in the target voice information is outside the semantic coverage range, it represents that the control intention in the target voice information is not covered, and it is determined that the initial user intention cannot be recognized or replied. In this case, the generative pre-trained GPT model needs to be queried with an inquiry message.

[0107] Query the generative pre-trained GPT model with the inquiry message to obtain the corrected target user intention of the user intention.

[0108] In an alternative embodiment, Figure 6 is a schematic diagram of an intention recognition method based on the generative pre-trained GPT model according to an embodiment of the present application. As Figure 6 shown, the target user intention can be recognized in the following ways, but not limited to:

[0109] First, obtain the user's target voice information, input the user's target voice information into the preset intention recognition model to obtain the initial user intention output by the preset intention recognition model, and then determine whether the control intention contained in the user's target voice information is covered in the initial user intention. If it is covered, directly execute the control intention without passing through the generative pre-trained GPT model, and execute the control intention in the target voice information according to the initial user intention output by the preset intention recognition model.

[0110] During the process of executing the control intention in the target voice information according to the initial user intention output by the preset intention recognition model, judge the user's emotion based on the feedback information of the user on the operation result. When the user's emotion is positive or there is no emotion, end the operation; when the user's emotion is negative, return the situation where the initial user intention is recognized as incorrect but the target voice information has been executed, and an inquiry message can be sent to the generative pre-trained GPT model to find the reason, and then detect whether there is a correct control intention in the obtained target user intention again and obtain the result.

[0111] In the case where there is no correct control intention, it is considered that it cannot be recognized or replied, and the recognition is incorrect but the target voice information has not been executed. At this time, an inquiry message can be sent to the generative pre-trained GPT model to find the reason.

[0112] Further, taking the situation where the user issues a target voice message in the target scenario as: "Xia A, blow some wind", the preset intent recognition model is used to perform control intent recognition on the target voice message, and the initial user intent "Play, singer: Xia A, song 'Blow Some Wind'" is obtained. The environmental information of the target scenario includes the time being 11 o'clock, the room temperature being 28 degrees Celsius, and the humidity being 40%. The device information includes an air conditioner (located in the master bedroom, alias is Xia A, in the on state), a water heater (located in the kitchen, in the off state), a sound system (located in the living room, in the on state), a washing machine (located in the bathroom, in the off state), a light (in the master bedroom, not turned on), and a curtain (in the master bedroom, opening degree 20%). For example, after obtaining the initial user intent, detecting that there is an initial control intent in the initial user intent, and controlling the device in the target scenario according to the initial control intent, the user's feedback information is obtained: My air conditioner is called Xia A, why is the song playing?

[0113] Further, perform an emotion detection on the feedback information, and it is detected that there is a negative emotion in the feedback information, so it is determined that the initial control intent is incorrect, and the preset intent recognition model fails to recognize the correct control intent of the user for the target scenario.

[0114] Construct scenario information based on the environmental information of the target scenario and the device information of the devices deployed in the target scenario (for example: time 11 o'clock, room temperature 28 degrees Celsius, humidity 40%, device information includes an air conditioner (located in the master bedroom, alias is Xia A, in the on state), a water heater (located in the kitchen, in the off state), a sound system (located in the living room, in the on state), a washing machine (located in the bathroom, in the off state), a light (in the master bedroom, not turned on), and a curtain (in the master bedroom, opening degree 20%)), generate a first description field for describing the target voice message (for example: the target voice message issued by the user is "Xia A, blow some wind"), a second description field for describing that the device in the target scenario has been controlled according to the initial control intent (for example: according to the target voice message issued by the user, the initial control intent has been executed: Play, singer: Xia A, song 'Blow Some Wind'), and a third description field for describing that there is a negative emotion in the feedback information (for example: there is a negative emotion in the user's feedback information: Why is the song playing?).

[0115] Use the first description field, the second description field, and the third description field to construct interaction information, such as: The user issued a target voice message: "Xia A, blow some wind". The initial user intent "Play, singer: Xia A, song 'Blow Some Wind'" has been executed, and there is a negative emotion in the user's feedback information: Why is the song playing?

[0116] Construct an inquiry message carrying scenario information and interaction information, including: In the target scenario where there are an air conditioner (located in the master bedroom, alias is Little A, in the on state), a water heater (located in the kitchen, in the off state), a speaker (located in the living room, in the on state), a washing machine (located in the bathroom, in the off state), a light (in the master bedroom, not turned on), and a curtain (in the master bedroom, with an opening degree of 20%), the user sent a target voice message: Little A, blow some wind. The initial user intention "Play, singer: Little A, song 'Blow Some Wind'" was executed, and there is a negative emotion in the user's feedback message: Why play the song?

[0117] Input the inquiry message into the generative pre-trained GPT model to obtain the target user intention output by the generative pre-trained GPT model. It should be noted that it is expected that the commercial model will output "Correction result: Turn on the air conditioner with the alias Little A".

[0118] By using a preset intention recognition model, the generalization and semantic understanding capabilities of the model are enhanced. Using the generative pre-trained GPT model can enhance the intention recognition ability of the entire system through interactive learning, increase the coverage of user intention recognition, and thus improve the accuracy of intention recognition, thereby improving the control efficiency of the device. It can solve the problem of low accuracy of intention recognition in related technologies, which leads to low control efficiency of the device, improve the accuracy of intention recognition and the control efficiency of the device, and enhance the user experience.

[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), including several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0120] Figure 7 It is a structural block diagram of an intention recognition device based on the generative pre-trained GPT model according to an embodiment of the present application, as Figure 7 shown. The device includes:

[0121] An acquisition module 702, configured to acquire the target voice message sent by the user in the target scenario, where

[0122] One or more controllable devices are deployed in the target scenario;

[0123] An identification module 704, configured to perform control intention identification on the target voice information through a preset intention recognition model to obtain an initial user intention, where the preset intention recognition model is constructed based on device operations allowed to be performed by devices in the target scenario, and the initial user intention is used to indicate the initial control intention of the user for the devices in the target scenario recognized by the preset intention recognition model;

[0124] An invocation module 706, configured to, when it is determined according to the initial user intention that the preset intention recognition model fails to recognize the correct control intention of the user for the devices in the target scenario, invoke a generative pre-trained GPT model to correct the user intention of an inquiry message carrying the target voice information and the scenario information of the target scenario, so as to obtain a target user intention, where the generative pre-trained GPT model is a generative language model trained using multi-domain information.

[0125] With the above device, the preset intention recognition model is used to recognize the target voice information to obtain the initial user intention. When the initial user intention indicates that the preset intention recognition model fails to recognize the correct control intention of the user, the generative pre-trained GPT model with a question-and-answer function trained using multi-domain information is used to reply to the inquiry message to obtain the target user intention. That is, when the preset intention recognition model fails to recognize the correct control intention of the user, the target user intention that meets the target voice information and the target scenario is obtained by asking the generative pre-trained GPT model. By using the generative pre-trained GPT model after the preset intention recognition model, the coverage of user intention recognition is increased, thereby improving the accuracy of intention recognition, and thus improving the control efficiency of the device. Therefore, the problem of low accuracy of intention recognition in the related art, which leads to low control efficiency of the device, can be solved, the accuracy of intention recognition and the control efficiency of the device are improved, and the user experience is enhanced.

[0126] In an exemplary embodiment, the invocation module includes:

[0127] An acquisition unit, configured to acquire the environmental information of the target scenario and the device information of the devices deployed in the target scenario as the scenario information, where the device information includes: device identifier, device location, and device status;

[0128] A conversion unit, configured to convert the target voice information into target text information;

[0129] A construction unit, configured to construct the inquiry message according to the target text information, the environmental information, and the device information;

[0130] An input unit for inputting the query information into the generative pre-trained GPT model to obtain the target user intention output by the generative pre-trained GPT model.

[0131] In an exemplary embodiment, the construction unit is further configured to: splice the environment information and the environment information description into an environment field, splice the device information and the device information description into a device field, splice the target text information and the command description into a command field, and generate an inquiry question corresponding to the initial user intention; splice the environment field, the device field, the command field, and the inquiry question into the query information.

[0132] In an exemplary embodiment, the apparatus further includes:

[0133] A first detection module for detecting whether there is an initial control intention recognized by the preset intention recognition model in the initial user intention, where the initial control intention is an intention for expressing the user's control of a device in the target scenario;

[0134] A first determination module for determining that the preset intention recognition model fails to recognize the correct control intention of the user for the target scenario when it is detected that there is no such initial control intention in the initial user intention;

[0135] A second detection module for detecting whether the initial control intention is correct when it is detected that there is the initial control intention in the initial user intention;

[0136] A second determination module for determining that the preset intention recognition model fails to recognize the correct control intention of the user for the target scenario when it is detected that the initial control intention is incorrect.

[0137] In an exemplary embodiment, the second detection module includes:

[0138] A processing unit for controlling the device in the target scenario according to the initial control intention and collecting the feedback information of the user in the target scenario;

[0139] A detection unit for performing emotion detection on the feedback information;

[0140] A third determination unit for determining that the initial control intention is incorrect when it is detected that there is a negative emotion in the feedback information.

[0141] In an exemplary embodiment, the calling module is further configured to: construct the scenario information according to the environmental information of the target scenario and the device information of the devices deployed in the target scenario; generate a first description field for describing the target voice information, a second description field for describing that the devices in the target scenario have been controlled according to the initial control intention, and a third description field for describing that there is a negative emotion in the feedback information; construct interaction information according to the first description field, the second description field and the third description field, where the interaction information is used to indicate an interaction process of interacting with the user using the initial user intention; construct the query information carrying the scenario information and the interaction information; and input the query information into the generative pre-trained GPT model to obtain the target user intention output by the generative pre-trained GPT model.

[0142] In an exemplary embodiment, the apparatus further includes:

[0143] A control module, configured to control the devices in the target scenario according to the target user intention when the target user intention is used to indicate that the generative pre-trained GPT model recognizes the target control intention corresponding to the target voice information;

[0144] A playback module, configured to play a default reply message when the target user intention is used to indicate that the generative pre-trained GPT model does not recognize the control intention of the target voice information.

[0145] An embodiment of the present application further provides a storage medium, which includes a stored program, where the above program executes the method of any one of the above when running.

[0146] Optionally, in this embodiment, the above storage medium may be set to store program code for executing the following steps:

[0147] S1. Obtain target voice information sent by a user in a target scenario, where one or more devices that can be controlled are deployed in the target scenario;

[0148] S2. Perform control intention recognition on the target voice information through a preset intention recognition model to obtain an initial user intention, where the preset intention recognition model is constructed according to device operations allowed to be performed by the devices in the target scenario, and the initial user intention is used to indicate the initial control intention of the user for the devices in the target scenario recognized by the preset intention recognition model;

[0149] S3. In the case where it is determined according to the initial user intention that the preset intention recognition model fails to recognize the correct control intention of the user for the device in the target scenario, call the generative pre-trained GPT model to correct the user intention for the query information carrying the target voice information and the scenario information of the target scenario, so as to obtain the target user intention, where the generative pre-trained GPT model is a generative language model trained using multi-domain information.

[0150] An embodiment of the present application further provides an electronic device. Figure 8 It is a schematic diagram of an optional electronic device according to an embodiment of the present application, as Figure 8 shown. The electronic device includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0151] Optionally, the above electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0152] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:

[0153] S1. Obtain the target voice information issued by the user in the target scenario, where one or more devices that allow to be controlled are deployed in the target scenario;

[0154] S2. Perform control intention recognition on the target voice information through a preset intention recognition model to obtain an initial user intention, where the preset intention recognition model is constructed based on the device operations allowed to be performed by the devices in the target scenario, and the initial user intention is used to indicate the initial control intention of the user for the devices in the target scenario recognized by the preset intention recognition model;

[0155] S3. In the case where it is determined according to the initial user intention that the preset intention recognition model fails to recognize the correct control intention of the user for the device in the target scenario, call the generative pre-trained GPT model to correct the user intention for the query information carrying the target voice information and the scenario information of the target scenario, so as to obtain the target user intention, where the generative pre-trained GPT model is a generative language model trained using multi-domain information.

[0156] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes, such as USB flash drives, read-only memories (ROM), random access memories (RAM), external hard drives, magnetic disks, or optical discs.

[0157] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.

[0158] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to be implemented. In this way, the present application is not limited to any specific combination of hardware and software.

[0159] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. An intent recognition method based on the generative pre-trained GPT model, characterized in that, it includes: Obtain the target voice information sent by the user in the target scenario, where one or more controllable devices are deployed in the target scenario; Perform control intent recognition on the target voice information through a preset intent recognition model to obtain an initial user intent, where the preset intent recognition model is constructed based on the device operations allowed to be performed by the devices in the target scenario, and the initial user intent is used to indicate the initial control intent of the user for the devices in the target scenario recognized by the preset intent recognition model; In the case where it is determined according to the initial user intent that the preset intent recognition model fails to recognize the correct control intent of the user for the devices in the target scenario, call the generative pre-trained GPT model to correct the user intent for the query information carrying the target voice information and the scenario information of the target scenario, and obtain the target user intent, where the generative pre-trained GPT model is a generative language model trained using multi-domain information.

2. The method according to claim 1, characterized in that, The step of calling the generative pre-trained GPT model to correct the user intent for the query information carrying the target voice information and the scenario information of the target scenario to obtain the target user intent includes: Obtain the environmental information of the target scenario and the device information of the devices deployed in the target scenario as the scenario information, where the device information includes: device identifier, device location, and device status; Convert the target voice information into target text information; Construct the query information according to the target text information, the environmental information, and the device information; Input the query information into the generative pre-trained GPT model to obtain the target user intent output by the generative pre-trained GPT model.

3. The method according to claim 2, characterized in that, The step of constructing the query information according to the target text information, the environmental information, and the device information includes: Concatenate the environmental information with the environmental information description to form an environmental field, concatenate the device information with the device information description to form a device field, concatenate the target text information with the command description to form a command field, and generate an inquiry question corresponding to the initial user intent; Concatenate the environmental field, the device field, the command field, and the inquiry question to form the query information.

4. The method according to claim 1, characterized in that, Before the step of calling the generative pre-trained GPT model to correct the user intent for the query information carrying the target voice information and the scenario information of the target scenario to obtain the target user intent, the method further includes: Detect whether there is an initial control intent recognized by the preset intent recognition model in the initial user intent, where the initial control intent is an intent used to express the user's control of the devices in the target scenario. In the case that the initial control intention does not exist in the detected initial user intention, it is determined that the preset intention recognition model fails to recognize the correct control intention of the user for the target scenario; In the case that the initial control intention exists in the detected initial user intention, it is detected whether the initial control intention is correct; In the case that the detected initial control intention is incorrect, it is determined that the preset intention recognition model fails to recognize the correct control intention of the user for the target scenario.

5. The method according to claim 4, wherein, the detecting whether the initial control intention is correct includes: controlling the device in the target scenario according to the initial control intention, and collecting the feedback information of the user in the target scenario; performing emotion detection on the feedback information; in the case that negative emotion exists in the detected feedback information, determining that the initial control intention is incorrect.

6. The method according to claim 5, wherein, the calling the generative pre-trained GPT model to correct the user intention of the query information carrying the target voice information and the scenario information of the target scenario to obtain the target user intention includes: constructing the scenario information according to the environmental information of the target scenario and the device information of the devices deployed in the target scenario; generating a first description field for describing the target voice information, a second description field for describing that the device in the target scenario has been controlled according to the initial control intention, and a third description field for describing that negative emotion exists in the feedback information; constructing interaction information according to the first description field, the second description field and the third description field, wherein the interaction information is used to indicate the interaction process of interacting with the user using the initial user intention; constructing the query information carrying the scenario information and the interaction information; inputting the query information into the generative pre-trained GPT model to obtain the target user intention output by the generative pre-trained GPT model.

7. The method according to claim 1, wherein, after the calling the generative pre-trained GPT model to correct the user intention of the query information carrying the target voice information and the scenario information of the target scenario to obtain the target user intention, the method further includes: in the case that the target user intention is used to indicate that the generative pre-trained GPT model recognizes the target control intention corresponding to the target voice information, controlling the device in the target scenario according to the target user intention; in the case that the target user intention is used to indicate that the generative pre-trained GPT model fails to recognize the control intention of the target voice information, playing a default reply message.

8. An intention recognition device based on a generative pre-trained GPT model, wherein, it includes: an acquisition module, configured to acquire target voice information sent by a user in a target scenario, where one or more devices that allow to be controlled are deployed in the target scenario; An identification module, configured to perform control intention identification on the target voice information through a preset intention recognition model to obtain an initial user intention, where the preset intention recognition model is constructed based on device operations allowed to be performed by devices in the target scenario, and the initial user intention is used to indicate the initial control intention of the user for the devices in the target scenario recognized by the preset intention recognition model; A calling module, configured to, when it is determined according to the initial user intention that the preset intention recognition model fails to recognize the correct control intention of the user for the devices in the target scenario, call a generative pre-trained GPT model to correct the user intention of an inquiry message carrying the target voice information and the scenario information of the target scenario, so as to obtain a target user intention, where the generative pre-trained GPT model is a generative language model trained using multi-domain information.

9. A computer-readable storage medium, characterized in that, the computer-readable storage medium includes a stored program, where the program, when running, executes the method according to any one of claims 1 to 7.

10. An electronic device, including a memory and a processor, characterized in that, a computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.