A method and device for driving a virtual image gaze direction based on a dialogue task
By acquiring user voice information to identify the location of business entities, generating gaze angle information, and controlling the virtual avatar system to adjust its gaze, the problem of virtual avatars being unable to interact in a human-like manner in existing technologies is solved, and the effect of adjusting the gaze according to the user's task is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA FAW CO LTD
- Filing Date
- 2023-01-04
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies cannot control virtual avatars to interact with users in a human-like manner based on specific business entities, nor can they dynamically adjust the virtual avatar's gaze direction according to the user's dialogue tasks.
By acquiring user voice information, recognizing semantic information, obtaining business entity information and their location information, and generating gaze angle information, the virtual avatar system is controlled to adjust the gaze direction.
This allows the virtual avatar to adjust its gaze based on the business entity interacting with the user, thus improving the humanized interaction effect.
Smart Images

Figure CN116107430B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of virtual reality interaction technology, specifically to a method and apparatus for driving the gaze direction of a virtual avatar based on a dialogue task. Background Technology
[0002] In existing voice systems, the orientation of a virtual avatar is determined based on wake-up location or sound source localization. For example, if the driver wakes the avatar with a wake-up word, the wake-up location is usually the location of the wake-up word. Sound source localization typically uses a microphone array to locate the speaker's position.
[0003] However, in some scenarios, virtual avatars should theoretically not always face the user, for example:
[0004] 1. When a user asks, "Where is the sunroof switch?", the virtual avatar replies, "The sunroof switch is directly above the vehicle." In this case, it would be more appropriate to look towards the sunroof.
[0005] 2. When a user asks, "How do I get to the intersection ahead?", the virtual avatar replies, "Turn left at the intersection ahead." In this situation, it would be more appropriate to look towards the left side of the vehicle.
[0006] 3. When a user asks "Play music," the virtual avatar replies "Okay, music will be playing soon." Directing the user's gaze towards the music player on the screen at this point would create a more dynamic effect.
[0007] 4. When a user asks, "Throw the passenger seat out," the virtual avatar replies, "The passenger seat is so cute, don't throw it out" (Entertainment Mode). Directing the view towards the passenger seat at this point will make the avatar appear more dynamic.
[0008] However, existing technology cannot control virtual avatars to interact with specific business entities in a more human-like manner.
[0009] Therefore, there is a need for a technical solution to address or at least mitigate the aforementioned shortcomings of existing technologies. Summary of the Invention
[0010] The purpose of this invention is to provide a method for driving the gaze direction of a virtual avatar based on a dialogue task to overcome or at least mitigate one of the aforementioned defects of the prior art.
[0011] One aspect of the present invention provides a method for driving the gaze direction of a virtual avatar based on a dialogue task, the method comprising:
[0012] Obtain user voice information;
[0013] Obtain business entity information based on user voice information;
[0014] Obtain the location information of the business entity based on the business entity information;
[0015] Generate line-of-sight angle information based on the location information;
[0016] The viewing angle information is sent to the virtual avatar system, so that the virtual avatar controlled by the virtual avatar system faces the position information.
[0017] Optionally, obtaining business entity information based on user voice information includes:
[0018] Identify the semantic information of the user's voice information;
[0019] Business entity information is obtained based on the semantic information.
[0020] Optionally, obtaining business entity information based on the semantic information includes:
[0021] Based on the semantic information, domain information, intent information, and slot information are obtained;
[0022] The domain information, intent information, and slot information are transmitted to the business application so that the business application can perform actions based on the domain information, intent information, and slot information.
[0023] The business entity information is obtained based on the actions performed by the business application.
[0024] Optionally, the step of obtaining the location information of the business entity information based on the business entity information includes:
[0025] Obtain a preset business entity location library, which includes at least one location information and preset business entity information, with one preset business entity information associated with one location information.
[0026] Obtain the location information associated with the preset business entity information that is identical to the business entity information.
[0027] Optionally, the method for driving the virtual avatar's gaze direction based on dialogue tasks further includes:
[0028] Generate the preset business entity location library.
[0029] Optionally, generating the preset business entity location database includes:
[0030] A three-dimensional coordinate system is established using the virtual image as the origin.
[0031] For each business entity information, a coordinate information is generated in the three-dimensional coordinate system as the position information of that business entity information.
[0032] Optionally, generating the line-of-sight angle information based on the location information includes:
[0033] Obtain the origin of the coordinate system;
[0034] Obtain location information;
[0035] The line-of-sight angle information is obtained using the location information and the origin of the coordinate system.
[0036] Optionally, the line-of-sight angle information obtained through the position information and the coordinate origin is obtained using the following formula:
[0037]
[0038]
[0039]
[0040] Optionally, before obtaining the location information of the business entity information based on the business entity information, the method for driving the virtual avatar's gaze direction based on the dialogue task further includes:
[0041] When the number of business entity information obtained exceeds one, select one of the business entity information to obtain the location information of that business entity information.
[0042] This application also provides a device for driving the gaze direction of a virtual avatar based on a dialogue task, the device comprising:
[0043] User voice information acquisition module, the user voice information acquisition module is used to acquire user voice information;
[0044] A business entity information acquisition module, which is used to acquire business entity information based on user voice information;
[0045] A location information acquisition module, wherein the location information acquisition module is used to acquire the location information of the business entity information based on the business entity information;
[0046] A line-of-sight angle information acquisition module, which is used to generate line-of-sight angle information based on the position information;
[0047] The sending module is used to send the line-of-sight angle information to the virtual avatar system, so that the virtual avatar controlled by the virtual avatar system faces the position information.
[0048] Beneficial effects
[0049] This application has the following advantages:
[0050] The method for driving the gaze direction of virtual avatars based on dialogue tasks in this application can provide location information for virtual avatars according to the business entity that the user wants to interact with, thereby changing the gaze or orientation of the location information based on the business entity that the user wants to interact with, thus making the virtual entity more human-like. Attached Figure Description
[0051] Figure 1 This is a flowchart illustrating the method for driving the gaze direction of a virtual avatar based on a dialogue task, as described in the first embodiment of this application.
[0052] Figure 2 It is an electronic device used to achieve Figure 1 The method shown is based on dialogue task-driven virtual avatar gaze direction.
[0053] Figure 3 yes Figure 1 The diagram shows the entity coordinates of a method for driving the gaze direction of a virtual avatar based on a dialogue task. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below with reference to the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, embodiments of this application. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0055] Figure 1 This is a flowchart illustrating the method for driving the gaze direction of a virtual avatar based on a dialogue task, as described in the first embodiment of this application.
[0056] like Figure 1 The methods shown for driving the gaze direction of virtual avatars based on dialogue tasks include:
[0057] Step 1: Obtain user voice information;
[0058] Step 2: Obtain business entity information based on user voice information;
[0059] Step 3: Obtain the location information of the business entity based on the business entity information;
[0060] Step 4: Generate line-of-sight angle information based on the location information;
[0061] Step 5: Send the viewing angle information to the virtual avatar system, so that the virtual avatar controlled by the virtual avatar system faces the position information.
[0062] The method for driving the gaze direction of virtual avatars based on dialogue tasks in this application can provide location information for virtual avatars according to the business entity that the user wants to interact with, thereby changing the gaze or orientation of the location information based on the business entity that the user wants to interact with, thus making the virtual entity more human-like.
[0063] In this embodiment, obtaining business entity information based on user voice information includes:
[0064] Identify the semantic information of the user's voice information;
[0065] Business entity information is obtained based on the semantic information.
[0066] In this embodiment, obtaining business entity information based on the semantic information includes:
[0067] Based on the semantic information, domain information, intent information, and slot information are obtained;
[0068] The domain information, intent information, and slot information are transmitted to the business application so that the business application can perform actions based on the domain information, intent information, and slot information.
[0069] The business entity information is obtained based on the actions performed by the business application.
[0070] In this embodiment, obtaining the location information of the business entity information based on the business entity information includes:
[0071] Obtain a preset business entity location library, which includes at least one location information and preset business entity information, with one preset business entity information associated with one location information.
[0072] Obtain the location information associated with the preset business entity information that is identical to the business entity information.
[0073] In this embodiment, the method for driving the virtual avatar's gaze direction based on dialogue tasks further includes:
[0074] Generate the preset business entity location library.
[0075] In this embodiment, generating the preset business entity location library includes:
[0076] A three-dimensional coordinate system is established using the virtual image as the origin.
[0077] For each business entity information, a coordinate information is generated in the three-dimensional coordinate system as the position information of that business entity information.
[0078] In this embodiment, generating the line-of-sight angle information based on the location information includes:
[0079] Obtain the origin of the coordinate system;
[0080] Obtain location information;
[0081] The line-of-sight angle information is obtained using the location information and the origin of the coordinate system.
[0082] In this embodiment, the line-of-sight angle information obtained through the position information and the coordinate origin is obtained using the following formula:
[0083]
[0084]
[0085]
[0086] In this embodiment, before obtaining the location information of the business entity information based on the business entity information, the method for driving the virtual avatar's gaze direction based on the dialogue task further includes:
[0087] When the number of business entity information obtained exceeds one, select one of the business entity information to obtain the location information of that business entity information.
[0088] The following examples further illustrate this application in detail. It is understood that these examples do not constitute any limitation on this application.
[0089] Obtain user voice information;
[0090] The business entity information is obtained based on the user's voice information, specifically as follows: The voice semantic recognition engine recognizes the content of the user input, identifies the Domain Intention Slot of the user input information, and sends the result to the dialogue management module. The dialogue management module sends semantic instructions to the business application based on the Domain information. The business application executes the semantic content, confirms the action entity, and feeds back the entity to the dialogue management module.
[0091] For example: Domain = CarInformation Intention = Searching Slot = "TopWindowKey" (Where is the sunroof switch?). After receiving this message, the vehicle control module confirms that the entity is "TopWindowKey" and then returns this entity to the dialogue management.
[0092] Sometimes the semantic slot and the entity slot are inconsistent. For example, if a user says "open the car window", the semantic slot information is Slot = "Window", but the vehicle is opening the driver's side window, then the returned entity information is "FRWindow".
[0093] Sometimes, entity information does not have corresponding slot information. For example, when a map performs a query for intersection turning information, with Domain=map, Intention=RouteInquire, and Slot=“null”, the map application still needs to return the entity “LeftSide”.
[0094] Sometimes the entity information is a business app, such as a media player (MedioPlayer).
[0095] Sometimes entity information is slot information contained in NLU, such as "FrontRightPosition".
[0096] The dialogue management module queries entity coordinate information in the entity database.
[0097] The entity location information database records the coordinate information (x, y, z) of different entities.
[0098] See Figure 3 The line-of-sight angle information is obtained using the following method:
[0099] Get the coordinates (x0, y0, z0) of the virtual avatar;
[0100] Obtain the location information (x1, y1, z1) of the business entity;
[0101] Calculate angle
[0102] Understandably, in some embodiments, if entity coordinates are not obtained, or if there are too many entities to identify a unique entity coordinate, the angle is not calculated and a default angle, i.e., the angle facing the user, is used.
[0103] This application also provides a device for driving the gaze direction of a virtual avatar based on a dialogue task. The device includes a user voice information acquisition module, a business entity information acquisition module, a location information acquisition module, a gaze angle information acquisition module, and a sending module. The user voice information acquisition module acquires user voice information; the business entity information acquisition module acquires business entity information based on the user voice information; the location information acquisition module acquires the location information of the business entity information based on the business entity information; the gaze angle information acquisition module generates gaze angle information based on the location information; and the sending module sends the gaze angle information to the virtual avatar system, thereby causing the virtual avatar controlled by the virtual avatar system to face the location information.
[0104] It should be noted that the foregoing explanation of the method embodiments also applies to the system of this embodiment, and will not be repeated here.
[0105] This application also provides an electronic device. In this embodiment, the electronic device is an edge server, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, it implements the above-described method for driving the gaze direction of a virtual avatar based on a dialogue task.
[0106] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the implementation of the above-described method for driving the gaze direction of a virtual avatar based on a dialogue task.
[0107] Figure 2 This is an exemplary structural diagram of an electronic device capable of implementing the method for driving the gaze direction of a virtual avatar based on a dialogue task according to an embodiment of this application.
[0108] like Figure 2As shown, the electronic device includes an input device 501, an input interface 502, a central processing unit 503, a memory 504, an output interface 505, and an output device 506. The input interface 502, central processing unit 503, memory 504, and output interface 505 are interconnected via a bus 507. The input device 501 and output device 506 are connected to the bus 507 via the input interface 502 and output interface 505, respectively, and thus connected to other components of the electronic device. Specifically, the input device 504 receives input information from the outside and transmits it to the central processing unit 503 via the input interface 502. The central processing unit 503 processes the input information based on computer-executable instructions stored in the memory 504 to generate output information, temporarily or permanently storing the output information in the memory 504, and then transmitting the output information to the output device 506 via the output interface 505. The output device 506 outputs the output information to the outside of the electronic device for user use.
[0109] In other words, Figure 2 The illustrated electronic device may also be implemented as including: a memory storing computer-executable instructions; and one or more processors, which can be coupled when executing the computer-executable instructions. Figure 1 The method described is based on dialogue task-driven virtual avatar gaze direction.
[0110] In one embodiment, Figure 2 The electronic device shown can be implemented as including: a memory 504 configured to store executable program code; and one or more processors 503 configured to run the executable program code stored in the memory 504 to perform the method for driving the gaze direction of a virtual avatar based on a dialogue task in the above embodiments.
[0111] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0112] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0113] Computer-readable media include both permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, DVD or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0114] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0115] Furthermore, it is clear that the word "comprising" does not exclude other units or steps. Multiple units, modules, or devices recited in a device claim may also be implemented by a single unit or overall device through software or hardware. The terms "first," "second," etc., are used to identify names, not to indicate any specific order.
[0116] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutively marked blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or the overall flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0117] In this embodiment, the processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0118] Memory can be used to store computer programs and / or modules. The processor implements various functions of the device / terminal equipment by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area can store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). In addition, memory can include high-speed random access memory, and can also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0119] In this embodiment, if the modules / units integrated into the device / terminal equipment are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by hardware related to computer program instructions. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0120] It should be noted that the content contained in a computer-readable medium may be appropriately added to or reduced according to the requirements of legislation and patent practice in the jurisdiction. Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.
[0121] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A method for driving the gaze direction of a virtual avatar based on a dialogue task, characterized in that, The method for driving the virtual avatar's gaze direction based on dialogue tasks includes: Obtain user voice information; Obtain business entity information based on user voice information; Obtain the location information of the business entity based on the business entity information; Generate line-of-sight angle information based on the location information; The viewing angle information is sent to the virtual avatar system, so that the virtual avatar controlled by the virtual avatar system faces the position information.
2. The method for driving the gaze direction of a virtual avatar based on a dialogue task as described in claim 1, characterized in that, The step of obtaining business entity information based on user voice information includes: Identify the semantic information of the user's voice information; Business entity information is obtained based on the semantic information.
3. The method for driving the gaze direction of a virtual avatar based on dialogue tasks as described in claim 2, characterized in that, The step of obtaining business entity information based on the semantic information includes: Based on the semantic information, domain information, intent information, and slot information are obtained; The domain information, intent information, and slot information are transmitted to the business application so that the business application can perform actions based on the domain information, intent information, and slot information. The business entity information is obtained based on the actions performed by the business application.
4. The method for driving the gaze direction of a virtual avatar based on a dialogue task as described in claim 3, characterized in that, The location information for obtaining business entity information based on business entity information includes: Obtain a preset business entity location library, which includes at least one location information and preset business entity information, with one preset business entity information associated with one location information. Obtain the location information associated with the preset business entity information that is identical to the business entity information.
5. The method for driving the gaze direction of a virtual avatar based on a dialogue task as described in claim 4, characterized in that, The method for driving the gaze direction of a virtual avatar based on dialogue tasks further includes: Generate the preset business entity location library.
6. The method for driving the gaze direction of a virtual avatar based on a dialogue task as described in claim 5, characterized in that, The generation of the preset business entity location library includes: A three-dimensional coordinate system is established using the virtual image as the origin. For each business entity information, a coordinate information is generated in the three-dimensional coordinate system as the position information of that business entity information.
7. The method for driving the gaze direction of a virtual avatar based on a dialogue task as described in claim 6, characterized in that, The step of generating the line-of-sight angle information based on the location information includes: Obtain the origin of the coordinate system; Obtain location information; The line-of-sight angle information is obtained using the location information and the origin of the coordinate system.
8. The method for driving the gaze direction of a virtual avatar based on a dialogue task as described in claim 7, characterized in that, The line-of-sight angle information obtained through the location information and the origin of the coordinate system is obtained using the following formula:
9. The method for driving the gaze direction of a virtual avatar based on a dialogue task as described in claim 8, characterized in that, Before obtaining the location information of the business entity information based on the business entity information, the method for driving the virtual avatar's gaze direction based on the dialogue task further includes: When the number of business entity information obtained exceeds one, select one of the business entity information to obtain the location information of that business entity information.
10. A device for driving the gaze direction of a virtual avatar based on a dialogue task, characterized in that, The device for driving the virtual avatar's gaze direction based on dialogue tasks includes: User voice information acquisition module, the user voice information acquisition module is used to acquire user voice information; A business entity information acquisition module, which is used to acquire business entity information based on user voice information; A location information acquisition module, wherein the location information acquisition module is used to acquire the location information of the business entity information based on the business entity information; A line-of-sight angle information acquisition module, which is used to generate line-of-sight angle information based on the position information; The sending module is used to send the line-of-sight angle information to the virtual avatar system, so that the virtual avatar controlled by the virtual avatar system faces the position information.