Multi-modal analysis-based intention determination method and electronic equipment

By performing speech recognition and intention analysis on the voice signal to be recognized in the smart device, and using historical interactive data for intention verification, the problem of low accuracy of intention recognition in the smart device is solved, and the accuracy of device control is improved.

CN120071909APending Publication Date: 2025-05-30QINGDAO HAIER INTELLIGENT HOME APPLIANCE TECHNOLOGY CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311629316.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the voice interaction between users and smart devices in the prior art, the low accuracy of intention recognition leads to poor accuracy of device control.

Method used

The intention determination method based on multimodal analysis is adopted, and speech recognition and intention analysis are performed by treating the recognized speech signal, combined with object attribute information in the historical interaction data, matching historical interaction data is found to verify the current intention and improve the accuracy of intention recognition.

Benefits of technology

By recognizing the current voice interaction signal and comparing the historical interaction intention, the accuracy of intention recognition is improved, thereby improving the accuracy of device control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071909A_ABST
    Figure CN120071909A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intention determination method based on multi-modal analysis and electronic equipment, and the method comprises the steps: carrying out the voice recognition of a to-be-recognized voice signal of a target object, and obtaining a target voice text; performing object intention analysis on the target voice text to obtain an analyzed first object intention; based on the target voice text and object attribute information of the target object, searching for first historical interaction data matched with the to-be-recognized voice signal from a group of historical interaction data; under the condition that the first object intention and the second object intention are consistent, the first intelligent equipment is controlled to execute first equipment operation, the second object intention is the object intention recorded by the first historical interaction data, and the first intelligent equipment is the intelligent equipment matched with the first object intention. The first device operation is a device operation matched with the first object intention. Through the method and the device, the problem of poor equipment control accuracy caused by low object intention recognition accuracy in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology. Specifically, the embodiments of the present application relate to a method for determining an intention based on multimodal analysis and an electronic device. Background Art

[0002] Currently, a user can perform voice interaction with a voice device (such as a smart home device) in a smart device control system (such as a smart home system) by issuing a voice command, so that the smart device control system can identify the interaction intention of the user based on the voice signal of the user collected by the voice device, and then perform a corresponding device operation based on the identified interaction intention. In an interaction scenario such as a home scenario, the interaction intentions of the user are diverse. For example, playing music, announcing recipes, querying the weather, device control, life services, etc., so that multiple intentions or multiple ambiguous intentions will be identified in a single round of voice interaction between the user and the voice device. At this time, it is necessary to accurately infer the true intention of the user.

[0003] The multi-intention discrimination method in the related art usually performs intention judgment based on the NLP (Natural Language Processing) parsing result and according to the trained corpus. When multiple object intentions are identified from the voice signal of the user, scores are given to different object intentions to determine the most likely object intention. However, the intention discrimination result may have a large error, resulting in inaccurate recognition of the object intention. It can be seen that there is a problem in the related art that the accuracy of device control is poor due to the low accuracy of object intention recognition. Summary of the Invention

[0004] The embodiments of the present application provide a method for determining an intention based on multimodal analysis and an electronic device, so as to at least solve the problem in the related art that the accuracy of device control is poor due to the low accuracy of object intention recognition.

[0005] According to an embodiment of the present application, there is provided a method for determining an intention based on multimodal analysis, including: performing speech recognition on a speech signal to be recognized of a target object to obtain a target speech text, where the speech signal to be recognized is a speech signal of the intention of the object to be recognized; performing object intention parsing on the target speech text to obtain a parsed first object intention; based on the target speech text and the object attribute information of the target object, searching for first historical interaction data matching the speech signal to be recognized from a set of historical interaction data, where the object attribute information includes attribute information of a set of object attributes, and each historical interaction data in the set of historical interaction data is used to record the correspondence between the speech text of an object, the object intention of the object, and the object attribute information of the object obtained during a voice interaction process; when the first object intention and the second object intention Figure 1 are consistent, controlling a first intelligent device to perform a first device operation, where the second object intention is the object intention recorded in the first historical interaction data, the first intelligent device is an intelligent device matching the first object intention, and the first device operation is a device operation matching the first object intention.

[0006] According to another embodiment of the present application, there is provided a device for determining an intention based on multimodal analysis, including: a first recognition unit for performing speech recognition on a speech signal to be recognized of a target object to obtain a target speech text, where the speech signal to be recognized is a speech signal of the intention of the object to be recognized; an analysis unit for performing object intention parsing on the target speech text to obtain a parsed first object intention; a search unit for searching for first historical interaction data matching the speech signal to be recognized from a set of historical interaction data based on the target speech text and the object attribute information of the target object, where the object attribute information includes attribute information of a set of object attributes, and each historical interaction data in the set of historical interaction data is used to record the correspondence between the speech text of an object, the object intention of the object, and the object attribute information of the object obtained during a voice interaction process; a first control unit for when the first object intention and the second object intention Figure 1 are consistent, controlling a first intelligent device to perform a first device operation, where the second object intention is the object intention recorded in the first historical interaction data, the first intelligent device is an intelligent device matching the first object intention, and the first device operation is a device operation matching the first object intention.

[0007] According to another embodiment of the present application, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0008] According to another embodiment of the present application, there is also provided an electronic device including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0009] Through the embodiments of the present application, a method is adopted to perform voice recognition on the voice signal to be recognized of the target object based on historical interaction data to verify the interaction intention, so as to obtain the target voice text, where the voice signal to be recognized is the voice signal of the intention of the object to be recognized; perform object intention parsing on the target voice text to obtain the parsed first object intention; based on the target voice text and the object attribute information of the target object, search for the first historical interaction data matching the voice signal to be recognized from a set of historical interaction data, where the object attribute information includes the attribute information of a set of object attributes, and each historical interaction data in the set of historical interaction data is used to record the correspondence relationship among the voice text of an object, the object intention of an object, and the object attribute information of an object obtained during a voice interaction process; when the first object intention and the second object intention Figure 1 are consistent, control the first intelligent device to execute the first device operation, where the second object intention is the object intention recorded by the first historical interaction data, the first intelligent device is the intelligent device matching the first object intention, and the first device operation is the device operation matching the first object intention. Since the current interaction intention obtained by recognizing the current voice interaction signal is compared with the most matching historical interaction intention to verify the current interaction intention, the purpose of improving the accuracy of intention recognition is achieved, and the technical effect of improving the accuracy of device control is achieved, and the problem in the related art that the accuracy of device control is poor due to the low accuracy of object intention recognition is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a hardware structure block diagram of a computer terminal for intention determination based on multimodal analysis according to an embodiment of the present application;

[0011] Figure 2 is a flowchart of a method for intention determination based on multimodal analysis according to an embodiment of the present application;

[0012] Figure 3 is a schematic diagram of a method for intention determination based on multimodal analysis according to an embodiment of the present application;

[0013] Figure 4 is a schematic diagram of another method for determining intent based on multimodal analysis according to an embodiment of the present application;

[0014] Figure 5 is a schematic diagram of another method for determining intention based on multimodal analysis according to an embodiment of the present application;

[0015] Figure 6 is a schematic diagram of another method for determining intention based on multimodal analysis according to an embodiment of the present application;

[0016] Figure 7 is a schematic diagram of another method for determining intention based on multimodal analysis according to an embodiment of the present application;

[0017] Figure 8 is a schematic diagram of another method for determining intention based on multimodal analysis according to an embodiment of the present application;

[0018] Figure 9 is a schematic diagram of another method for determining intention based on multimodal analysis according to an embodiment of the present application;

[0019] Figure 10 It is a structural block diagram of an intention determination device based on multimodal analysis according to an embodiment of the present application. DETAILED DESCRIPTION

[0020] The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments. It should be noted that the terms "first", "second", etc. in the description and claims of the embodiments of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0021] According to one aspect of an embodiment of the present application, a method for determining intent based on multimodal analysis is provided. The method for determining intent based on multimodal analysis can be widely used in smart home (Smart Home), smart home, smart home device ecology, smart residential (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned method for determining intent based on multimodal analysis can be applied to Figure 1 In the hardware environment composed of the terminal device 102 and the server 104 shown in FIG. Figure 1 As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or a client installed on the terminal. A database can be set on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.

[0022] The above network may include, but is not limited to, at least one of the following: a wired network, a wireless network. The above wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The above wireless network may include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 is not limited to a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing device, a smart dishwasher, a smart projection device, a smart TV, a smart drying rack, a smart curtain, a smart audio and video, a smart socket, a smart speaker, a smart sound box, a smart fresh air device, a smart kitchen and bathroom device, a smart bathroom device, a smart floor sweeping robot, a smart window cleaning robot, a smart mopping robot, a smart air purification device, a smart steam box, a smart microwave oven, a smart kitchen water heater, a smart purifier, a smart water dispenser, a smart door lock, etc.

[0023] Optionally, the method for determining an intention based on multimodal analysis in this embodiment may be executed by the terminal device 102, or may be executed by the server 104, or may be jointly executed by the terminal device 102 and the server 104. Among them, the execution by the terminal device 102 may be executed by the client on the terminal device 102. This embodiment does not make any limitation in this regard.

[0024] According to one aspect of the embodiments of the present application, a method for determining an intention based on multimodal analysis is provided. The method for determining an intention based on multimodal analysis can be applied to the scenario of voice control of intelligent devices. The intelligent devices here can be intelligent home devices in a smart home system, or intelligent devices in other scenarios, as long as the device has a voice interaction function. Voice control of intelligent devices may refer to: parsing the intention of interactive voice data and performing device control based on the parsed intention.

[0025] Taking the voice control of intelligent home devices in a home scenario as an example, in the process of human-machine voice interaction, the user outputs a statement. After semantic understanding, the human-machine interaction system recognizes the user's intention, and then gives a corresponding reply to the user according to the intention. The user can ask follow-up questions through multiple rounds of interaction in a short time, and the human-machine voice interaction system will give a reply to the user according to the statement output by the user in the current round to obtain an accurate interaction result.

[0026] In a home scenario, the interaction intentions of users are diverse. For example, playing music, announcing recipes, querying the weather, device control, life services, etc. When users interact with smart home devices, multiple intentions or ambiguous intentions may be generated in a single round of interaction. At this time, it is necessary for smart home devices (or the background servers of smart home devices) to accurately infer the true intentions of users.

[0027] In related technologies, the method for multi-intention discrimination is usually as follows: Based on the parsing results of NLP, intention judgment is performed according to the trained corpus. When a user's query (inquiry, referring to the user's voice interaction instruction) contains multiple intentions, scores are given to different intentions to determine the most likely user intention. However, the above intention discrimination method only classifies intentions based on the ASR (Automatic Speech Recognition) parsing results and the user's voice text.

[0028] However, in the same home scenario, different users may have different intentions for the same query. For example, when a boy asks "Do I need to take an umbrella when I go out?", his intention is to inquire whether it will rain, while when a girl asks "Do I need to take an umbrella when I go out?", her intention may be: inquiring whether it will rain or inquiring about the outdoor ultraviolet intensity. If only text-based intention parsing is used, the above intentions cannot be distinguished.

[0029] To at least partially solve the above problems, in this embodiment, for a user of a smart device platform (a platform that provides services corresponding to multiple smart devices), the current voice text recognized from the voice signal of the current interaction object can be subjected to intention recognition to obtain the current object intention. Based on the object attributes of the current interaction object and the current voice text and historical interaction data, a matching object intention can be obtained, and the current object intention and the matching object intention are compared to verify the current object intention, which can achieve the purpose of improving the accuracy of intention recognition, thereby improving the accuracy of device control.

[0030] Here, providing feedback on text-based user intention inference in combination with the user's historical behavior can help improve the accuracy of user intention recognition. For example, when the user says "Do I need to bring an umbrella when I go out?", and the smart home appliance responds with "It doesn't rain today", but then the user asks "What is the ultraviolet intensity outside?", at this time, it can be determined that the user is concerned not only about rain but also about the ultraviolet intensity. Therefore, the user's intention here is to ask whether to bring an umbrella for sun protection. So, when this user asks again "Do I need to bring an umbrella when going out", it can be considered that there is a greater probability that they want to ask two intentions: "Is it raining" and "Is the ultraviolet intensity too high", that is, in the case of rain, an umbrella is needed, and when the ultraviolet intensity is high, an umbrella is also needed. Make a final response operation by fusing the two intentions.

[0031] As an alternative implementation, taking the computer terminal to execute the intention determination method based on multimodal analysis in this embodiment as an example, Figure 2 is a schematic flowchart of the intention determination method based on multimodal analysis according to the embodiment of the present application, as Figure 2 shown, this process includes the following steps:

[0032] Step S202, perform speech recognition on the speech signal to be recognized of the target object to obtain the target speech text, where the speech signal to be recognized is the speech signal of the intention of the object to be recognized.

[0033] For the current target object to be interacted with (corresponding to the current user), when it is necessary to use voice to control the intelligent device to perform the corresponding device operation, a control voice can be issued. The intelligent device (with a voice collection function, also called a voice device) can collect the speech signal of the target object to obtain the speech signal to be recognized, and the speech signal to be recognized is the speech signal of the intention of the object to be recognized. The speech signal to be recognized can be obtained by voice collection through the audio collection module of the intelligent device (for example, a microphone array). The speech signal to be recognized can be unprocessed speech data (that is, raw audio data), or speech data obtained after preprocessing the raw audio data, or can be obtained by other means.

[0034] The audio collection module can be installed in the intelligent device or can be an independent device. After the audio collection module in the intelligent device collects the speech signal, the collected speech signal to be recognized can be sent to the server for the server to perform subsequent processing operations. Optionally, without contradiction, all or part of the subsequent operations can also be performed locally by the intelligent device, which is not limited in this embodiment. Here, the intelligent device and the server can both belong to the smart home system or other human-computer interaction systems.

[0035] For example, as Figure 3As shown, taking a smart TV as an example of a smart device, the operating system of the smart TV (an example of a human-computer interaction system) may include, but is not limited to, at least one of the following: a TV set 302, an audio acquisition device 304, a speech recognition device 306, a voice interaction control device 308, a server 310, and so on. Among them, the audio acquisition device 304 may be a microphone or other components capable of acquiring audio signals. The speech recognition device 306 may be a device with functions such as processing of speech information and extraction of speech features. It may be an independently set device or a functional module integrated in the voice interaction control device 308 of the TV set. The voice interaction control device 308 may include, but is not limited to, at least one of the following: a processor 312, for example, a CPU (Central Processing Unit), and a memory 314, for example, a RAM (Random Access Memory), or it may also be a disk memory. The server 310 may be a cloud server or a local server built into the TV set. The description of the server is similar to that of the foregoing embodiments and will not be elaborated specifically here.

[0036] When the TV set is in a voice interaction state, the voice information of the user currently using the TV set can be obtained in real time through the audio acquisition device 304. The voice information may be a sound segment emitted by the user, voice command information, or others. The speech recognition device 306 then recognizes the user's voice information. Among them, the voice capture device may be built into the TV set or fixedly installed in the area where the TV set is located, or it may also be a mobile device independent of the TV set, such as a mobile phone, a smart watch, and other mobile devices with voice capture functions. Here, when the TV set is in a voice interaction state, the user can control the operation of the TV set by emitting voice information. The voice interaction state may be the default state of the TV set or can be turned on through the user's control command.

[0037] It should be noted that Figure 3 the structure of the TV set system does not limit the human-computer interaction system. The human-computer interaction system may include Figure 3 more or fewer components, or combine some components, or have different component arrangements.

[0038] For the speech signal to be recognized, in order to obtain the object intention, the speech signal to be recognized can be subjected to speech recognition, thereby obtaining the target speech text. The method of performing speech recognition on the speech signal to be recognized can be any speech recognition method in the related art. For example, it can be to perform ASR parsing on the speech signal to be recognized: input the speech signal to be recognized into an ASR model to obtain the speech text output by the ASR model, thereby obtaining the target speech text.

[0039] Step S204: Perform object intention parsing on the target voice text to obtain the parsed first object intention.

[0040] After obtaining the target voice text, object intention parsing can be performed on the target voice text to obtain the parsed first object intention. Performing object intention parsing on the target voice text can be: inputting the target voice text into a semantic parsing model, and the semantic parsing model outputs the parsed object intention. For example, the object intention is parsed through an NLP model. Any intention parsing method in the related technologies can be adopted for performing object intention parsing on the target voice text, and the method for object intention parsing in this embodiment is not specifically limited.

[0041] Here, the NLP model generally requires both NLU (Natural Language Understanding) and NLG (Natural Language Generation). For example, in common products such as voice assistants and smart speakers, in order to support users to call various skills of the machine using natural language voice, it is not only necessary to understand what the user is saying, but also necessary to perform specific actions to meet the user's needs, such as answering "The information you are looking for is in this list". When understanding the user's words and intentions, the machine needs to use NLU technology; when responding to the user in the form of text or language, the machine needs to use NLG technology. Therefore, generally, a method is not specifically divided into whether it is NLU or NLG.

[0042] Step S206: Based on the target voice text and the object attribute information of the target object, search for the first historical interaction data that matches the to-be-recognized voice signal from a set of historical interaction data.

[0043] To determine the matching object intention of the to-be-recognized voice signal, the first historical interaction data that matches the to-be-recognized voice signal can be found from a set of historical interaction data based on the target voice text and the object attribute information of the target object. The object attribute information includes the attribute information of a set of object attributes. A set of object attributes can be attributes for distinguishing different interaction scenarios of different objects, and can include at least one of the following: object attributes for identifying the object identity (such as age or age range, gender, etc.), object location, object emotion, etc. Each historical interaction data is used to record the corresponding relationship among the voice text of an object, the object intention of the same object, and the object attribute information of the same object obtained during a voice interaction process.

[0044] Here, for each piece of historical interaction data in the historical interaction database, the user's speech text or the processed speech text can be used as an index, and the indexed content can include, but is not limited to, object intent and object attribute information, etc. First, the target speech text is respectively matched with each piece of historical interaction data in the historical interaction database, and then the object attribute information of the target object is matched with the object attribute information in the matched historical interaction data, so as to determine whether there is historical interaction data that matches the target speech text and the object attribute information of the target object, that is, the first historical interaction data. The object intent recorded in the first historical interaction data is the second object intent.

[0045] Optionally, if there is no matched historical interaction data, the first object intent can be used as the final object intent to perform subsequent operations, and based on whether the target object updates this object intent, determine the historical interaction data to be saved, that is, if there is an update, the object intent in the saved historical interaction data is the updated object intent, and if there is no update, the object intent in the saved historical interaction data is the updated first object intent.

[0046] Step S208, when the first object intent and the second object intent Figure 1 are consistent, control the first intelligent device to perform the first device operation, where the first intelligent device is the intelligent device that matches the first object intent, and the first device operation is the device operation that matches the first object intent.

[0047] When obtaining the first object intent and the second object intent, the first object intent and the second object intent can be compared to determine whether they are consistent. The consistency of the object intent here Figure 1 can mean that the object intents are exactly the same or have the same semantics. If the first object intent and the second object intent Figure 1 are consistent, then the first intelligent device can be controlled to perform the device operation that matches any one of the two object intents, that is, the first device operation.

[0048] Among them, the first intelligent device can be the intelligent device that collected the speech signal to be recognized as described above. For example, the current execution is a voice question-and-answer interaction, or the intelligent device that the target object needs to control is the same as the intelligent device that collected the speech signal, or it can also be an intelligent device different from the intelligent device that collected the speech signal to be recognized as described above. For example, one intelligent device controls other intelligent devices to perform device operations (such as controlling the opening of curtains or a TV by interacting with a smart speaker). The first intelligent device can be one or more intelligent devices, and the corresponding first device operation can also be one or more device operations.

[0049] Optionally, for a scenario where the first object intention and the second object intention are inconsistent, the subsequent device operations can be directly performed according to the first object intention, or more information can be obtained through a new round of voice interaction to update the object intention, and the subsequent device operations can be performed based on the updated object intention. This embodiment does not make any limitations in this regard.

[0050] Through the above steps S202 to S208, speech recognition is performed on the speech signal to be recognized of the target object to obtain the target speech text. Among them, the speech signal to be recognized is the speech signal of the object intention to be recognized; object intention parsing is performed on the target speech text to obtain the parsed first object intention; based on the target speech text and the object attribute information of the target object, the first historical interaction data matching the speech signal to be recognized is searched from a set of historical interaction data. Among them, the object attribute information includes the attribute information of a set of object attributes, and each historical interaction data in the set of historical interaction data is used to record the correspondence relationship between the speech text of an object, the object intention of an object, and the object attribute information of an object obtained during a voice interaction process; when the first object intention and the second object intention Figure 1 are consistent, the first intelligent device is controlled to perform the first device operation. Among them, the second object intention is the object intention recorded in the first historical interaction data, the first intelligent device is the intelligent device matching the first object intention, and the first device operation is the device operation matching the first object intention, which solves the problem of poor accuracy of device control caused by low accuracy of object intention recognition in the related art and improves the accuracy of device control.

[0051] In an exemplary embodiment, one or more devices can be bound to the same object or object group at the same time. Correspondingly, the speech signal to be recognized includes the speech signals sent by a set of target devices corresponding to the same voice command (or the same interaction voice) issued by the target object, and the set of target devices belongs to multiple preset devices, that is, the set of target devices is the devices among the multiple preset devices that collect the same voice command issued by the target object. Here, the multiple preset devices can be the devices bound to the same object or object group, such as the smart home devices bound to the same family group.

[0052] For example, when the user issues a voice query at a certain location in the home, multiple nearby smart home devices receive the user's voice signal in turn according to the distance from the user.

[0053] Considering that the intentions of the same voice command issued by the same object at different positions may also vary, a set of configured object attributes may include the object's location. For example, when the user says, "It stinks," the intention parsed from the text may be to ventilate the room. However, when the user is in the bedroom, the user's intention may be to "open the window" or "turn on the fresh air mode of the air conditioner"; when the user is in the bathroom, the user's intention may be to "turn on the exhaust fan"; when the user is in the kitchen, the user's intention may be to "turn on the range hood." Therefore, the user's location helps to accurately identify the user's intention.

[0054] In this embodiment, for the convenience of determining the object's location, the collection area of the voice collection component (i.e., the aforementioned audio collection device) of each preset device can be divided into multiple non-overlapping areas, that is, multiple preset sound areas. For example, the smart devices in the user's home can divide the recognizable sound areas according to the built-in algorithm and number each sound area. As Figure 4 shown, smart device A can be divided into two sound areas, 1 and 2, smart device B can be divided into three sound areas, 1, 2, and 3, and smart device C can be divided into four sound areas, 1, 2, 3, and 4. In addition, the room where each smart device is located can also be marked, as Figure 5 shown.

[0055] For the target object, the object location of the target object can be determined separately based on the voice signals collected by each target device, and then the final location of the target object can be determined by fusing the determined object locations of the target object. Correspondingly, the above method further includes:

[0056] S11, separately extract the location information of the target object from the voice signals sent by each target device in a group of target devices, and obtain the location information corresponding to each target device, where the location information corresponding to each target device is used to indicate the target sound area where the target object is located among the multiple preset sound areas of each target device;

[0057] S12, determine the object location information of the target object according to the location information corresponding to each target device, where the object attribute information of the target object is the object location information of the target object.

[0058] For the voice signals sent by each target device in the signal to be recognized, the location information of the target object can be extracted from the voice signals sent by each target device, that is, the location information corresponding to each target device. Among them, the location information corresponding to each target device is used to indicate the target sound area where the target object is located among the multiple preset sound areas of each target device. Here, the location information corresponding to each target device may include the sound area identifier of the target sound area of each target device.

[0059] For example, the microphone array on each smart device scans the voice energy intensity in each sound area to determine the sound area number where the user is located. Each smart device outputs the sound area number in the format of {device id: sound area}, and the device id is the unique id (unique identifier) set before the device leaves the factory. For example, {Smart Device A: 2}, {Smart Device B: 1}.

[0060] After obtaining the location information corresponding to each target device, the object location information of the target object can be determined according to the location information corresponding to each target device. Here, the object location information of the target object can indicate that the target object is located in the target sound area of a certain target device, or it can be the object location determined by combining other information. This is not limited in this embodiment. Correspondingly, the object attribute information of the target object is the object location information of the target object. For example, the user's location can be determined according to the sound area numbers where the user is located given by multiple smart devices.

[0061] Optionally, determining the object location of the target object can be performed by a sound source localization module, which is a program module (i.e., software program) integrated with a sound source localization algorithm. It can be set on the server or on a certain smart device.

[0062] Through this embodiment, by configuring multiple sound areas of the smart device and determining the user location based on the device sound area where the recognized object is located, the accuracy of user intention recognition is improved.

[0063] In an exemplary embodiment, determining the object location information of the target object according to the location information corresponding to each target device includes:

[0064] S21, in the case where each preset device in multiple preset devices is located in a different room and a group of target devices only includes one target device, determine the room information of the room where one target device is located as the object location information of the target object.

[0065] In this embodiment, if each preset device is located in a different room and a group of target devices only includes one target device, the room information of the room where the one target device is located can be determined as the object location information of the target object. In addition, it can also be to determine the room information of the room where the one target device is located and the sound area information of the target sound area where the target object is located among the multiple preset sound areas of the one target device as the object location information of the target object.

[0066] For example, taking Figure 4 the placement position of the smart devices in as an example, if there is only one smart device in each room, such as Figure 5As shown in the figure, the intelligent device that receives the user's voice determines the room where the user is located, and the room where the intelligent device is located is the user's location. If intelligent device A receives the user's voice, it is determined that the user is in the bathroom; if intelligent device B receives the user's voice, it is determined that the user is in the bedroom; if intelligent device C receives the user's voice, it is determined that the user is in the living room.

[0067] Through this embodiment, for the scenario where there is only one intelligent device in each room, the room where the user is located is determined according to the intelligent device that receives the user's voice, so as to determine the user's location, which can more accurately and conveniently determine the user's location.

[0068] In an exemplary embodiment, if there are multiple intelligent devices in a room, in order to improve the accuracy of object location determination, the home environment within the sound area of each device can be marked, and at least one preset object is marked in each preset sound area. Here, the method of object identification can be manual marking or automatic mapping with a sweeping robot with visual recognition function. The constructed home environment can be furniture such as sofas, dining tables, dressing mirrors, etc. or commonly used locations set by the user. For example, as Figure 6 shown, within the 2nd sound area of intelligent device A and the 1st sound area of intelligent device B, there are an entrance door and a shoe cabinet; within the 1st sound area of intelligent device C, there is an entrance door, and within the 2nd sound area, there are a shoe cabinet and a dining table. These objects can be marked.

[0069] Correspondingly, according to the position information corresponding to each target device, the object position information of the target object is determined, including:

[0070] S31, in the case where at least two of the multiple preset devices are located in the same room and / or a group of target devices includes multiple target devices, determine a first preset object set, where the first preset object set includes the preset objects marked in the target sound area of each target device;

[0071] S32, determine a second preset object set, where the second preset object set includes the preset objects marked in the other sound areas except the target sound area of each target device among the multiple preset sound areas of each target device, and the preset objects marked in the multiple preset sound areas of the other preset devices except the group of target devices among the multiple preset devices;

[0072] S33, in the case where there is a group of target objects in the first preset object set, determine the object position information of the target object according to the group of target objects, where the group of target objects is the preset objects that belong to the first preset object set and do not belong to the second preset object set.

[0073] If at least two of the multiple preset devices are located in the same room and / or a set of target devices includes multiple target devices, it is possible to determine whether there are preset objects that belong to the target sound area of each target device and do not belong to other preset sound areas except the target sound area of each target device. If so, the object position of the target object can be determined based on such preset objects.

[0074] In this embodiment, the preset objects marked in the target sound area of each target device can be determined respectively to obtain a first set of preset objects. Here, the first set of preset objects may include one or more non-repeating preset objects; the preset objects marked in the other sound areas except the target sound area of each target device among the multiple preset sound areas of each target device, and the preset objects marked in the multiple preset sound areas of the other preset devices except the set of target devices among the multiple preset devices are determined as the second set of preset objects; if there is a set of target objects in the first set of preset objects that belong to the first set of preset objects and do not belong to the preset objects in the second set of preset objects, that is, a set of target objects, the object position information of the target object can be determined based on the set of target objects.

[0075] For example, when a user has a voice interaction with a smart device, first judge the sound area numbers of the user in each smart device according to the voice signals received by each smart device, and then take the intersection of the home furnishing types within the range of the sound area numbers of each smart device and subtract the home furnishings in the sound areas where no audio signal is received, so as to further infer the user's position.

[0076] Through this embodiment, by representing the objects in the preset sound area and determining the object position based on the preset objects that belong to the target sound area of the target device and do not belong to other preset sound areas except the target sound area of the target device, the accuracy of object position determination can be improved.

[0077] In an exemplary embodiment, when there is a set of target objects in the first set of preset objects, determining the object position information of the target object based on the set of target objects includes:

[0078] S41, when there is a set of target objects in the first set of preset objects and the set of target objects only includes one target object, output the object information of the one target object as the object position information of the target object;

[0079] S42, when there is a set of target objects in the first set of preset objects, the set of target objects includes at least two target objects, and the at least two target objects are located in the same target device, output the object information of the target object that is closest to the target sound area of the same target device among the at least two target objects as the object position information of the target object;

[0080] S43. When there is a set of target objects in the first preset object set, a set of target objects includes at least two target objects, and at least two target objects are located at at least two target devices, determine the target device with the maximum signal energy of the corresponding voice signal among the at least two target devices to obtain the maximum energy device; output the object information of the target object with the closest distance to the target sound area of the maximum energy device among the at least two target objects as the object position information of the target object.

[0081] In this embodiment, if there is a set of target objects in the first preset object set and a set of target objects includes only one target object, indicating that the target object is relatively close to the one target object, the position of the one target object can be directly output as the object position information of the target object. If there is a set of target objects in the first preset object set, a set of target objects includes at least two target objects, and at least two target objects are located at the same target device, it indicates that the target object is located within the same target device and is relatively close to the at least two target objects. To represent the object position more conveniently, the object information of the target object with the closest distance to the target sound area (or the same target device) of the same target device among the at least two target objects can be output as the object position information of the target object; if there is a set of target objects in the first preset object set, a set of target objects includes at least two target objects, and at least two target objects are located at at least two target devices, it can indicate that the target object is relatively close to at least two target devices at this time. At this time, the target device with the maximum signal energy of the corresponding voice signal among the at least two target devices can be determined, that is, the maximum energy device, and then the object information of the target object with the closest distance to the target sound area (or the maximum energy device) of the maximum energy device is selected from the at least two target objects and output as the object position information of the target object.

[0082] For example, as Figure 7As shown, when the user issues a query, the voice device collects the user's voice signal and generates the corresponding sound zone number pairs. For example, {Smart device A: 2}, {Smart device B: 1}. Mark the homes included in the sound zone number pairs of each voice device that receives the user's voice, and take the intersection as {P}. At the same time, mark the homes included in the sound zone number pairs of the voice devices that do not receive the user's voice to form the set {Q}. The homes that are in {P} and not in {Q} are formed into {K}. When there is only one home element in {K}, output the location of this home as the user's location; when {K} includes at least two home elements, find the list of devices that include the home elements in {K} and receive the user's voice signal. If there is only one device in this device list, output the home element closest to the sound zone of this device as the user's location; if there are at least two devices in this device list, they can be sorted according to the strength of the voice signal received by the devices, and the device that receives the strongest voice signal is selected, and the home element closest to the sound zone of this device is output as the user's location.

[0083] Optionally, combined with Figure 7 and Figure 8 , for the method of user sound source localization, its application examples can be as follows:

[0084] Example 1: When the user's voice is received in the 2nd sound zone of Smart device A, the 3rd sound zone of Smart device B, and the 2nd sound zone of Smart device C, it can be determined that the user is at Location 1;

[0085] Example 2: When the user's voice is received in the 2nd sound zone of Smart device A, the 2nd sound zone of Smart device B, and the 2nd sound zone of Smart device C, it can be determined that the user is at Location 2;

[0086] Example 3: When the user's voice is received in the 2nd sound zone of Smart device A, the 1st sound zone of Smart device B, and the 2nd sound zone of Smart device C, it can be determined that the user is at Location 3. At this time, there is a front door and a shoe cabinet in the 2nd sound zone of Smart device A and the 1st sound zone of Smart device B; there is a shoe cabinet and a dining table in the 2nd sound zone of Smart device C. Take the intersection as the shoe cabinet, and subtract the homes in other sound zones: the dining table. Therefore, it is determined that the user is near the shoe cabinet;

[0087] Example 4: When the user's voice is received in the 1st sound zone of Smart device A, the 1st sound zone of Smart device B, and the 1st sound zone of Smart device C, it can be determined that the user is at Location 4;

[0088] Example 5: When Smart device A and Smart device B do not receive sound and the user's voice is received in the 2nd sound zone of Smart device C, it can be determined that the user is at Location 5. The 2nd sound zone of Smart device C includes a shoe cabinet and a dining table. Subtract the homes included in the sound zones of Smart device A and Smart device B (shoe cabinet, front door). Therefore, it is determined that the user is at the dining table;

[0089] Example 6: When smart device A and smart device B do not receive sound and the first sound zone of smart device C receives user speech, it can be determined that the user is at position 6;

[0090] Example 7: When smart device A and smart device B do not receive sound and the fourth sound zone of smart device C receives user speech, it can be determined that the user is at position 7;

[0091] Example 8: When smart device A and smart device B do not receive sound and the third sound zone of smart device C receives user speech, it can be determined that the user is at position 8.

[0092] It should be noted that a preset coding method can be adopted to code the user's position. For example, the one-hot coding method or other methods. The coding result can be transmitted to the server or other devices.

[0093] Through this embodiment, by configuring the corresponding relationship between different scenarios and the user's position, the object position can be flexibly determined, and the accuracy of user position determination can be improved.

[0094] In an exemplary embodiment, the above method further includes:

[0095] S51, extracting speech features from the speech signal to be recognized to obtain the speech features of the target object;

[0096] S52, performing voiceprint recognition on the target object according to the speech features of the target object to obtain the identity recognition result of the target object, where the identity recognition result is used to indicate whether there is a matching object of the target object in a group of preset objects;

[0097] S53, when it is determined according to the identity recognition result that there is a matching object of the target object in a group of preset objects, determining the attribute information of a group of specified attributes of the matching object as the attribute information of a group of specified attributes of the target object;

[0098] S54, when it is determined according to the identity recognition result that there is no matching object of the target object in a group of preset objects, performing voiceprint attribute detection on the speech signal to be recognized to obtain the attribute information of a group of specified attributes of the target object;

[0099] S55, performing speech emotion recognition according to the speech features of the target object to obtain the emotion category of the object emotion of the target object.

[0100] In this embodiment, a set of object attributes includes a set of specified attributes and object emotions. The set of specified attributes includes at least one of the following: gender, age, and may also include other types of object attributes. Among them, object emotion recognition can be achieved by invoking a voice emotion recognition model, or can be achieved by using other voice emotion recognition methods. The set of specified attributes can be obtained by identifying the object identity of the target object through voiceprint attributes and matching based on the configured attribute information, or can be directly identified based on voiceprint attributes.

[0101] To obtain a set of specified attributes and object emotions of the target object, voice feature extraction can be first performed on the voice signal to be recognized to obtain the voice features of the target object. Here, for the voice signals of different target devices, voice feature extraction can be performed separately, or a part of them can be selected for voice feature extraction. The extracted voice features can be applied to subsequent processes such as attribute recognition and voice text recognition.

[0102] For example, as Figure 8 shown, after the user wakes up the intelligent device and issues a voice query, such as "Do I need to take an umbrella when going out?"; after the intelligent device receives the user voice signal, it can perform voice front-end processing steps such as frame addition and windowing, FFT (Fast Fourier Transformation), Mel filtering, and discrete cosine transform on the received user voice signal, and perform feature extraction.

[0103] Perform voiceprint recognition based on the extracted voice features to identify the user's identity: According to the voice features of the target object, voiceprint recognition can be performed on the target object to obtain the identity recognition result of the target object. The identity recognition result is used to indicate whether there is a matching object of the target object among a set of preset objects; if it is determined according to the identity recognition result that there is a matching object of the target object, the attribute information of a set of specified attributes of the matching object is determined as the attribute information of the set of specified attributes of the target object; if it is determined according to the identity recognition result that there is no matching object of the target object, voiceprint attribute detection is performed on the voice signal to be recognized to obtain the attribute information of a set of specified attributes of the target object. Voiceprint attribute detection can be performed using the voice features of the target object. Here, if some specified attributes cannot identify the corresponding attribute information, this specified attribute can be ignored, or it can be set to a specified value (for example, set to empty).

[0104] For example, perform voiceprint recognition based on the extracted voice features to identify the user's identity. If the owner of the voiceprint is not registered in the family, call the voiceprint attribute detection model to output the user's age and gender.

[0105] Meanwhile, voice emotion recognition can be performed according to the voice characteristics of the target object to obtain the emotion category of the object emotion of the target object. The voice emotion recognition can be performed by invoking the aforementioned voice emotion recognition model. The object emotion can include anger, anxiety, calmness, happiness, etc. For example, while performing ASR, the voice emotion recognition model can be invoked in parallel to determine the current emotion category of the user. Here, the ASR recognition, voiceprint attribute, and emotion recognition can be detected simultaneously.

[0106] It should be noted that before extracting the voice characteristics from the voice signal to be recognized, preprocessing operations such as pre-emphasis and framing are performed on it, which can effectively eliminate the influence of factors such as aliasing, high-order harmonic distortion, and high frequency caused by the human vocal organs themselves and the devices for collecting voice signals on the quality of the voice signal, ensure that the signals obtained from subsequent voice processing are more uniform and smooth, provide high-quality parameters for signal parameter extraction, and improve the quality of voice processing.

[0107] Optionally, the voiceprint attribute detection of the voice signal to be recognized can be performed by invoking a voiceprint attribute detection model. The voice signal to be recognized is input into the preset interface of the voiceprint attribute detection model to obtain the age and gender of the target object output by the voiceprint attribute detection model. The voiceprint attribute of the voice signal to be recognized of the target object can also be detected by other means. The voice emotion recognition of the voice characteristics of the target object can be performed by inputting the voice characteristics of the target object into the preset interface of the voice emotion recognition model to obtain the emotion characteristics of the target object output by the voice emotion recognition model. The voice emotion recognition of the voice characteristics of the target object can also be performed by other means.

[0108] For example, as Figure 9 shown, when the user issues an inquiry, after the device receives the user's voice signal, it will perform front-end processing of the voice signal to extract the voice characteristics of the user, and then input the extracted voice characteristics into the voiceprint attribute detection model and the voice emotion recognition model respectively, and output the age, gender of the user and the emotion of the user respectively.

[0109] Through this embodiment, by extracting object attributes such as the age, gender, and emotion of the object, and thus judging the user's intention according to the user's age, gender, and emotion, etc., the user's intention can be recognized more accurately.

[0110] In an exemplary embodiment, based on the target voice text and the object attribute information of the target object, searching for the first historical interaction data matching the voice signal to be recognized from a set of historical interaction data includes:

[0111] S61. Match the target voice text with the voice texts recorded in each piece of historical interaction data respectively to obtain a set of candidate historical interaction data. Among them, for each candidate historical interaction data in the set of candidate historical interaction data, the text similarity between the voice text it records and the target voice text is greater than or equal to a preset similarity threshold.

[0112] S62. Select from the set of candidate historical interaction data the candidate historical interaction data that matches the object attribute information of the target object to obtain the first historical interaction data.

[0113] When performing historical interaction data matching, historical interaction data retrieval can be first carried out using the voice text as an index: match the target voice text with the voice texts recorded in each piece of historical interaction data respectively to obtain a set of candidate historical interaction data. The text similarity between the voice text recorded in each candidate historical interaction data and the target voice text is greater than or equal to a preset similarity threshold. The preset similarity threshold can be flexibly set as needed, and its value can be 80%, 90% or other values.

[0114] For similarity calculation, the target voice text can be converted into an encoded form and the encodings of all texts in a set of historical data are used for similarity calculation. Among them, the form of converting into text can be achieved by inputting the voice text into a pre-trained model to obtain the encoding of the voice text output by the pre-trained model. The pre-trained model can adopt word2vec (a related model for generating word vectors), transformer (transformer), BERT (Bidirectional Encoder Representations from Transformers, bidirectional encoder representations from transformers), etc. The formula for similarity calculation can be as shown in formula (1):

[0115]

[0116] Among them, v 1 can represent the encoding of a voice text, and v 2 can represent the encoding of another voice text.

[0117] For example, when constructing a historical interaction database, a unique intent id can be assigned to each intent; input the ASR text of the user's voice into the pre-trained model to obtain the encoding of this text; store the compiled user text, the corresponding intent id, interaction time, user location encoding, user identity, user gender, user age, and user voice emotion as a data group in the database with the user text as the index.

[0118] When searching in the historical interaction database, the ASR text of the user's speech can be input into a pre-trained model to obtain the encoding of the text; calculate the similarity between the encoding of the text and all text encodings in the database, sort the calculated similarity values from largest to smallest, and find one or more user speech text encodings with the largest similarity; if only one similar text is found, output the intention corresponding to the text.

[0119] Through this embodiment, first match based on the speech text, and then match based on each object attribute, so as to determine the closest historical interaction data, which can improve the accuracy of historical interaction data matching.

[0120] In an exemplary embodiment, selecting candidate historical interaction data that matches the object attribute information of the target object from a group of candidate historical interaction data to obtain the first historical interaction data includes:

[0121] S71, in the case that a group of candidate historical interaction data contains multiple candidate historical interaction data, sequentially match the object attribute information with the object attribute information recorded in each candidate historical interaction data based on each object attribute in a group of object attributes to obtain the matching result of each candidate historical interaction data;

[0122] S72, determine the candidate historical interaction data with the highest matching degree indicated by the corresponding matching result of a group of candidate historical interaction data as the first historical interaction data.

[0123] If a group of candidate historical interaction data contains only one candidate historical interaction data, then this one candidate historical interaction data can be directly determined as the first historical interaction data. Optionally, the one candidate historical interaction data can also be verified based on the object attribute information of the target object. If the matching degree between the object attribute information recorded in this one candidate historical interaction data and the object attribute information of the target object reaches the set matching degree threshold, then execute the subsequent processing flow. If the matching degree between the object attribute information recorded in this one candidate historical interaction data and the object attribute information of the target object does not reach the set matching degree threshold, then the subsequent processing can be directly executed according to the first object intention, or the object intention can be determined through further interaction with the target object, and this embodiment does not make a limitation on this.

[0124] If a set of candidate historical interaction data contains multiple candidate historical interaction data, screening can be performed based on other information: the object attribute information can be sequentially matched with the object attribute information recorded in each candidate historical interaction data based on each object attribute to obtain the matching result of each candidate historical interaction data. The matching result of each candidate historical interaction data is used to indicate the matching degree between the object attribute information recorded in each candidate historical interaction data and the object attribute information of the target object. Here, the matching degree of the object attribute information can be the weighted sum of the matching degrees determined based on each object attribute, or the maximum value, minimum value, or median value of the matching degrees determined based on each object attribute, etc. It can also be other determination methods, which are not limited in this embodiment; the candidate historical interaction data with the highest matching degree indicated by the corresponding matching result in a set of candidate historical interaction data is determined as the first historical interaction data, or the candidate historical interaction data with the highest matching degree and a matching degree greater than the set matching degree threshold indicated by the corresponding matching result in a set of candidate historical interaction data is determined as the first historical interaction data.

[0125] For example, if multiple similar texts are found, the most appropriate set of data is sequentially matched according to the location, identity, gender, age, emotion, and interaction time of the current user, and the intention corresponding to the data is output.

[0126] Through this embodiment, the voice text is first matched, and then each object attribute is sequentially matched to determine the closest historical interaction data, which can improve the accuracy of historical interaction data matching.

[0127] In an exemplary embodiment, the above method further includes:

[0128] S81, when the first object intention and the second object intention are inconsistent, an intention prompt message is sent through the target voice device, where the intention prompt message is used to prompt the target object to confirm the object intention corresponding to the first object intention and the second object intention that corresponds to the voice signal to be recognized;

[0129] S82, when the response voice signal of the target object is received, the response voice signal is subjected to speech recognition to obtain a response voice text;

[0130] S83, when the response voice text indicates that the object intention corresponding to the first object intention and the second object intention that corresponds to the voice signal to be recognized is the target object intention, the second intelligent device is controlled to execute a second device operation, where the second intelligent device is an intelligent device that matches the target object intention, and the second device operation is a device operation that matches the target object intention.

[0131] To more accurately determine the user's intention, if the intention of the first object is inconsistent with that of the second object, an intention prompt message can be sent through the target voice device to prompt the target object to confirm the object intention corresponding to the first object intention and the second object intention in the to-be-recognized voice signal. The format of the above intention prompt message can be "Are you asking about Figure 1 or about Figure 2 ", "Do you need Figure 1 or need Figure 2 ", or others.

[0132] For example, when the user says a query, such as, "Do I need to take an umbrella when going out", the ways to recognize the user's intention can include:

[0133] Step 1: The sound source localization module (Module 1) determines the user's location based on the positions of the distributed terminal devices in the home and the sound source localization technology;

[0134] Step 2: Perform ASR parsing on the user's voice and output the recognized text;

[0135] Step 3: The voiceprint attribute recognition module (Module 2) recognizes the user's identity, age, gender, and emotion based on the user's acoustic signal information;

[0136] Step 4: According to all the user-related information obtained in Steps 1 and 2, search for query content with a relatively high similarity in the historical data of the home user interaction behavior to obtain a preset intention.

[0137] Step 5: Input the ASR recognition result of Step 2 into the NLU model and output the parsed user intention.

[0138] Step 6: If the user intentions output by Step 3 and Step 4 Figure 1 are consistent, execute according to the intention; if they are inconsistent, clarify the user intention.

[0139] The way to clarify the user intention can be: send an intention clarification query to the user. For example, if the user intention output by Step 3 is "Ask about the ultraviolet intensity" and the user intention output by Step 4 is "Ask whether it will rain", then send a clarification query "Are you asking about the ultraviolet intensity or whether it will rain" and wait for the user to confirm.

[0140] When a response voice signal of a target object is received, speech recognition can be performed on the received response voice signal to obtain a corresponding response voice text. When the response voice text indicates that the object intention corresponding to the voice signal to be recognized among the first object intention and the second object intention is the target object intention, the second intelligent device can be controlled to perform a second device operation. Wherein, the second intelligent device can be an intelligent device matching the target object intention, and the second device operation can be a device operation matching the target object intention. The target object intention can be any one of the first object intention and the second object intention. In addition, if the user indicates that the actual intention is not any one of the first object intention and the second object intention, the object intention can be re-recognized, and the way of recognizing the object intention is similar to that in the foregoing embodiments, which will not be elaborated herein.

[0141] Through this embodiment, when the recognized object intention and the matched object intention are inconsistent, an inquiry request for intention clarification is sent to the user, so that the user intention can be obtained more accurately.

[0142] In an exemplary embodiment, after controlling the second intelligent device to perform the second device operation, the above method further includes:

[0143] S91, adding second historical interaction data to a set of historical interaction data, where the second historical interaction record is used to record the correspondence between the target voice text, the target object intention, and the object attribute information of the target object.

[0144] After controlling the second intelligent device to perform the second device operation, a set of historical interaction data can be updated, and the update method can be adding second historical interaction data to a set of historical interaction data, where the second historical interaction record can be used to record the correspondence between the target voice text, the target object intention, and the object attribute information of the target object.

[0145] For example, after the user gives an answer, new data (ASR text, intention, location, identity, age, gender, emotion, interaction time) is written into the home user interaction behavior database for update.

[0146] Through this embodiment, updating the historical interaction database based on the clarified user intention can facilitate subsequent recognition of the user intention, thereby improving the accuracy of recognizing the user intention.

[0147] Exemplarily, such as Figure 9As shown, when the user asks "Do I need to bring an umbrella when I go out?", the voice acoustic signal can be input into Module 1 (sound source localization module), Module 2 (voiceprint attribute recognition module) simultaneously or sequentially for ASR parsing and recognition. In Module 1, the user's sound source can be located to obtain the sound area number of the user's location. In Module 2, the user's attribute information (such as identity, age, gender) and emotion can be recognized. Through ASR parsing and recognition, the user's voice text can be input into the NLU model, and the NLU model outputs the user's intention. Then, according to the user's sound area number, identity, age, gender, and emotion obtained from Module 1 and Module 2, the data with the highest similarity is searched in the historical interaction database based on the user's voice text recognized by ASR parsing. The intention corresponding to this data can be compared with the user intention output by the NLU model. If they are inconsistent, the user can clarify the intention, update the historical interaction database with the clarified user intention and the corresponding other data, and execute the intention according to the user's clarified intention.

[0148] Through this optional example, a multi-intention judgment method based on the home space and home user roles is proposed, which can achieve personalized intention parsing; utilize the advantages of multiple smart home devices in the home and use a distributed method for multi-intention clarification; and multi-intention clarification is not only related to the user information such as the current user's voice text and emotion, but also strongly related to the user's historical behavior, which can well avoid the situation where the intelligent home appliance service is unavailable and the recognition error cannot be corrected in time.

[0149] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0150] On the other hand, according to an embodiment of the present application, an intention determination device based on multi-modal analysis is further provided. This device is used to implement the intention determination method based on multi-modal analysis provided in the above embodiments, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0151] Figure 10 is a structural block diagram of an intention determination device based on multi-modal analysis according to an embodiment of the present application, as Figure 10As shown in the figure, the device includes: a first recognition unit 1002, configured to perform speech recognition on a speech signal to be recognized of a target object to obtain a target speech text, where the speech signal to be recognized is a speech signal of the intention of the object to be recognized; an analysis unit 1004, configured to perform object intention analysis on the target speech text to obtain a first object intention obtained by analysis; a search unit 1006, configured to search, based on the target speech text and the object attribute information of the target object, for first historical interaction data that matches the speech signal to be recognized from a set of historical interaction data, where the object attribute information includes the attribute information of a set of object attributes, and each historical interaction data in the set of historical interaction data is used to record the correspondence relationship among the speech text of an object, the object intention of an object, and the object attribute information of an object obtained during a speech interaction process; a first control unit 1008, configured to control a first intelligent device to perform a first device operation when the first object intention and the second object intention Figure 1 are consistent, where the second object intention is the object intention recorded in the first historical interaction data, the first intelligent device is an intelligent device that matches the first object intention, and the first device operation is a device operation that matches the first object intention.

[0152] Through the embodiments of the present application, speech recognition is performed on a speech signal to be recognized of a target object to obtain a target speech text, where the speech signal to be recognized is a speech signal of the intention of the object to be recognized; object intention analysis is performed on the target speech text to obtain a first object intention obtained by analysis; based on the target speech text and the object attribute information of the target object, first historical interaction data that matches the speech signal to be recognized is searched from a set of historical interaction data, where the object attribute information includes the attribute information of a set of object attributes, and each historical interaction data in the set of historical interaction data is used to record the correspondence relationship among the speech text of an object, the object intention of an object, and the object attribute information of an object obtained during a speech interaction process; when the first object intention and the second object intention Figure 1 are consistent, a first intelligent device is controlled to perform a first device operation, where the second object intention is the object intention recorded in the first historical interaction data, the first intelligent device is an intelligent device that matches the first object intention, and the first device operation is a device operation that matches the first object intention, which can solve the problem in the related art that the accuracy of device control is poor due to the low accuracy of object intention recognition, and improve the accuracy of device control.

[0153] Optionally, the speech signal to be recognized includes speech signals corresponding to the same speech instruction issued by a target object and sent by a set of target devices among a plurality of preset devices; the apparatus further includes: an extraction unit, configured to respectively extract the position information of the target object from the speech signals sent by each target device in the set of target devices, to obtain position information corresponding to each target device, where the position information corresponding to each target device is used to indicate the target sound area where the target object is located among a plurality of preset sound areas of each target device, and the plurality of preset sound areas of each target device are a plurality of non-overlapping areas into which the collection area of the speech collection component of each target device is divided; a first determination unit, configured to determine the object position information of the target object according to the position information corresponding to each target device, where a set of object attributes includes object position, and the object attribute information of the target object is the object position information of the target object.

[0154] Optionally, the first determination unit includes: a first determination module, configured to, when each of the plurality of preset devices is located in a different room and the set of target devices includes only one target device, determine the room information of the room where one target device is located as the object position information of the target object.

[0155] Optionally, at least one preset object is marked in each preset sound area of each of the plurality of preset devices; the first determination unit includes: a second determination module, configured to determine a first preset object set when at least two of the plurality of preset devices are located in the same room and / or the set of target devices includes a plurality of target devices, where the first preset object set includes the target objects marked in the target sound areas of each target device; a third determination module, configured to determine a second preset object set, where the second preset object set includes the preset objects marked in the other sound areas except the target sound areas of each target device among the plurality of preset sound areas of each target device, and the preset objects marked in the plurality of preset sound areas of the other preset devices except the set of target devices among the plurality of preset devices; a fourth determination module, configured to, when there is a set of target objects in the first preset object set, determine the object position information of the target object according to the set of target objects, where the set of target objects is the preset objects that belong to the first preset object set and do not belong to the second preset object set.

[0156] Optionally, the fourth determination module includes: a first output subunit, configured to output the object information of a target object as the object position information of the target object when there is a set of target objects in the first preset object set and the set of target objects includes only one target object; a second output subunit, configured to output the object information of the target object that is closest to the target sound area of the same target device among at least two target objects as the object position information of the target object when there is a set of target objects in the first preset object set, the set of target objects includes at least two target objects, and the at least two target objects are located in the same target device; an execution subunit, configured to determine, when there is a set of target objects in the first preset object set, the set of target objects includes at least two target objects, and the at least two target objects are located in at least two target devices, a target device with the largest signal energy of the corresponding voice signal among the at least two target devices to obtain a maximum energy device; and output the object information of the target object that is closest to the target sound area of the maximum energy device among the at least two target objects as the object position information of the target object.

[0157] Optionally, the above device further includes: an extraction unit, configured to perform voice feature extraction on the voice signal to be recognized to obtain the voice features of the target object; a second recognition unit, configured to perform voiceprint recognition on the target object according to the voice features of the target object to obtain an identity recognition result of the target object, where the identity recognition result is used to indicate whether there is a matching object of the target object in a set of preset objects; a second determination unit, configured to, when it is determined according to the identity recognition result that there is a matching object of the target object in the set of preset objects, determine the attribute information of a set of specified attributes of the matching object as the attribute information of a set of specified attributes of the target object; a detection unit, configured to, when it is determined according to the identity recognition result that there is no matching object of the target object in the set of preset objects, perform voiceprint attribute detection on the voice signal to be recognized to obtain the attribute information of a set of specified attributes of the target object; a third recognition unit, configured to perform voice emotion recognition according to the voice features of the target object to obtain the emotion category of the object emotion of the target object; where a set of object attributes includes a set of specified attributes and object emotion, and a set of specified attributes includes at least one of the following: gender, age.

[0158] Optionally, the search unit includes: a first matching module, configured to match the target voice text with the voice text recorded in each historical interaction data respectively to obtain a set of candidate historical interaction data, where the text similarity between the voice text recorded in each candidate historical interaction data in the set of candidate historical interaction data and the target voice text is greater than or equal to a preset similarity threshold; and a selection module, configured to select, from the set of candidate historical interaction data, the candidate historical interaction data that matches the object attribute information of the target object to obtain the first historical interaction data.

[0159] Optionally, the selection module includes: a matching sub-module, configured to, when a set of candidate historical interaction data includes multiple pieces of candidate historical interaction data, sequentially match the object attribute information with the object attribute information recorded in each piece of candidate historical interaction data based on each object attribute in a set of object attributes, to obtain a matching result for each piece of candidate historical interaction data, where the matching result for each piece of candidate historical interaction data is used to indicate the matching degree between the object attribute information recorded in each piece of candidate historical interaction data and the object attribute information of the target object; a determination sub-module, configured to determine, as the first historical interaction data, the piece of candidate historical interaction data with the highest matching degree indicated by the corresponding matching result in the set of candidate historical interaction data.

[0160] Optionally, the above-mentioned apparatus further includes: a sending unit, configured to, when the first object intention and the second object intention are inconsistent, send an intention prompt message through the target voice device, where the intention prompt message is used to prompt the target object to confirm the object intention corresponding to the to-be-recognized voice signal in the first object intention and the second object intention; a fourth recognition unit, configured to, when receiving the response voice signal of the target object, perform voice recognition on the response voice signal to obtain a response voice text; a second control unit, configured to, when the response voice text indicates that the object intention corresponding to the to-be-recognized voice signal in the first object intention and the second object intention is the target object intention, control the second intelligent device to perform a second device operation, where the second intelligent device is an intelligent device matching the target object intention, and the second device operation is a device operation matching the target object intention.

[0161] Optionally, the above-mentioned apparatus further includes: an adding unit, configured to add second historical interaction data to a set of historical interaction data, where the second historical interaction record is used to record the correspondence between the target voice text, the target object intention, and the object attribute information of the target object.

[0162] It should be noted that the above-mentioned each module can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited thereto: all the above-mentioned modules are located in the same processor; or, the above-mentioned each module is located in different processors in any combination form.

[0163] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above-mentioned method embodiments when running.

[0164] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, ROM (Read-Only Memory), RAM (Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0165] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments.

[0166] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0167] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.

[0168] Obviously, those skilled in the art should understand that the above modules or steps of the embodiments of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0169] The above is only the preferred embodiment of the present application and is not used to limit the embodiments of the present application. For those skilled in the art, the embodiments of the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.

Claims

1. A method for intent determination based on multimodal analysis, characterized in that, it includes: Performing speech recognition on the speech signal to be recognized of the target object to obtain the target speech text, wherein the speech signal to be recognized is the speech signal of the intent of the object to be recognized; Performing object intent parsing on the target speech text to obtain the parsed first object intent; Based on the target speech text and the object attribute information of the target object, searching for the first historical interaction data that matches the speech signal to be recognized from a set of historical interaction data, wherein the object attribute information includes the attribute information of a set of object attributes, and each historical interaction data in the set of historical interaction data is used to record the correspondence relationship among the speech text of an object, the object intent of the object, and the object attribute information of the object obtained during a speech interaction process; When the first object intent and the second object intent are consistent, controlling the first intelligent device to perform the first device operation, wherein the second object intent is the object intent recorded by the first historical interaction data, the first intelligent device is the intelligent device that matches the first object intent, and the first device operation is the device operation that matches the first object intent.

2. The method according to claim 1, characterized in that, the speech signal to be recognized includes the speech signals sent by a group of target devices among multiple preset devices and corresponding to the same speech instruction issued by the target object; the method further includes: Respectively extracting the position information of the target object from the speech signals sent by each target device in the group of target devices to obtain the position information corresponding to each target device, wherein the position information corresponding to each target device is used to indicate the target sound zone where the target object is located among the multiple preset sound zones of each target device, and the multiple preset sound zones of each target device are multiple non-overlapping regions divided from the collection area of the speech collection component of each target device; Determining the object position information of the target object according to the position information corresponding to each target device, wherein the set of object attributes includes the object position, and the object attribute information of the target object is the object position information of the target object.

3. The method according to claim 2, characterized in that, the determining the object position information of the target object according to the position information corresponding to each target device includes: When each of the multiple preset devices is located in a different room and the group of target devices includes only one target device, determining the room information of the room where the one target device is located as the object position information of the target object.

4. The method according to claim 2, characterized in that, at least one preset object is marked in each preset sound zone of each of the multiple preset devices; the determining the object position information of the target object according to the position information corresponding to each target device includes: In the case where at least two of the multiple preset devices are located in the same room and / or the set of target devices includes multiple target devices, determine a first preset object set, where the first preset object set includes the preset objects marked in the target sound area of each target device; Determine a second preset object set, where the second preset object set includes the preset objects marked in the other sound areas except the target sound area of each target device among the multiple preset sound areas of each target device, and the preset objects marked in the multiple preset sound areas of the other preset devices except the set of target devices among the multiple preset devices; In the case where a set of target objects exists in the first preset object set, determine the object position information of the target object according to the set of target objects, where the set of target objects is the preset objects that belong to the first preset object set and do not belong to the second preset object set.

5. The method according to claim 4, wherein, the step of determining the object position information of the target object according to the set of target objects in the case where a set of target objects exists in the first preset object set includes: in the case where the set of target objects exists in the first preset object set and the set of target objects only includes one target object, output the object information of the one target object as the object position information of the target object; in the case where the set of target objects exists in the first preset object set, the set of target objects includes at least two target objects, and the at least two target objects are located in the same target device, output the object information of the target object that is closest to the target sound area of the same target device among the at least two target objects as the object position information of the target object; in the case where the set of target objects exists in the first preset object set, the set of target objects includes at least two target objects, and the at least two target objects are located in at least two target devices, determine the target device with the largest signal energy of the corresponding voice signal among the at least two target devices to obtain the maximum energy device; output the object information of the target object that is closest to the target sound area of the maximum energy device among the at least two target objects as the object position information of the target object.

6. The method according to claim 1, wherein, the method further includes: extracting voice features from the voice signal to be recognized to obtain the voice features of the target object; performing voiceprint recognition on the target object according to the voice features of the target object to obtain an identity recognition result of the target object, where the identity recognition result is used to indicate whether there is a matching object of the target object among a set of preset objects; in the case where it is determined according to the identity recognition result that there is a matching object of the target object among the set of preset objects, determine the attribute information of a set of specified attributes of the matching object as the attribute information of the set of specified attributes of the target object. In the case that no matching object of the target object exists in the set of preset objects determined according to the identity recognition result, perform a voiceprint attribute detection on the voice signal to be recognized, and obtain the attribute information of the set of specified attributes of the target object; Perform a voice emotion recognition according to the voice characteristics of the target object to obtain the emotion category of the object emotion of the target object; Wherein, the set of object attributes includes the set of specified attributes and the object emotion, and the set of specified attributes includes at least one of the following: gender, age.

7. The method according to claim 1, wherein, The finding the first historical interaction data matching the voice signal to be recognized from a set of historical interaction data based on the target voice text and the object attribute information of the target object includes: Matching the target voice text with the voice text recorded in each of the historical interaction data respectively to obtain a set of candidate historical interaction data, wherein, for each candidate historical interaction data in the set of candidate historical interaction data, the text similarity between the voice text recorded therein and the target voice text is greater than or equal to a preset similarity threshold; Selecting, from the set of candidate historical interaction data, the candidate historical interaction data that matches the object attribute information of the target object to obtain the first historical interaction data.

8. The method according to claim 7, wherein, The selecting, from the set of candidate historical interaction data, the candidate historical interaction data that matches the object attribute information of the target object to obtain the first historical interaction data includes: In the case that the set of candidate historical interaction data contains multiple candidate historical interaction data, sequentially matching the object attribute information with the object attribute information recorded in each of the candidate historical interaction data based on each object attribute in the set of object attributes to obtain the matching result of each candidate historical interaction data, wherein, the matching result of each candidate historical interaction data is used to indicate the matching degree between the object attribute information recorded in each candidate historical interaction data and the object attribute information of the target object; Determining, as the first historical interaction data, the candidate historical interaction data with the highest matching degree indicated by the corresponding matching result in the set of candidate historical interaction data.

9. The method according to any one of claims 1 to 8, wherein, The method further includes: In the case that the first object intention and the second object intention are inconsistent, sending an intention prompt message through a target voice device, wherein the intention prompt message is used to prompt the target object to confirm the object intention corresponding to the first object intention and the second object intention and related to the voice signal to be recognized; In the case of receiving the response voice signal of the target object, performing a voice recognition on the response voice signal to obtain a response voice text; In the case where the response speech text indicates that the object intention corresponding to the to-be-recognized speech signal among the first object intention and the second object intention is the target object intention, control the second intelligent device to perform a second device operation, where the second intelligent device is an intelligent device matching the target object intention, and the second device operation is a device operation matching the target object intention.

10. The method according to claim 9, wherein, after controlling the second intelligent device to perform the second device operation, the method further includes: adding second historical interaction data to the set of historical interaction data, where the second historical interaction record is used to record the correspondence between the target speech text, the target object intention, and the object attribute information of the target object.

11. An electronic device, comprising a memory and a processor, wherein, a computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 10 through the computer program.