Virtual character adjustment method and apparatus, vehicle, electronic device, storage medium and product

By acquiring user location information and using sound positioning algorithms to adjust the orientation of the virtual avatar, the problem of the virtual avatar not being able to follow changes in user location has been solved, improving the naturalness and experience of human-computer interaction.

WO2025247240A1PCT designated stage Publication Date: 2025-12-04BEIJING CO WHEELS TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/097567
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-27
Filing Date
2025-05-27
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

In existing technologies, virtual avatars can only look in a few fixed directions and cannot follow the user's position changes in real time, resulting in an unnatural experience when people talk to voice assistants.

Method used

By acquiring the location information of the user's voice interaction with the virtual avatar, and using preset sound positioning and angle algorithms, the virtual avatar's facial orientation is adjusted so that it can follow the user's position changes in real time.

Benefits of technology

It improves the experience of human-voice assistant dialogue, ensuring that the virtual avatar can turn towards the speaking user in real time based on the user's location, thus enhancing the naturalness of the interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025097567_04122025_PF_FP_ABST
    Figure CN2025097567_04122025_PF_FP_ABST
Patent Text Reader

Abstract

A virtual character adjustment method and apparatus, a vehicle, an electronic device, a storage medium, and a program product. The method comprises: acquiring position information of a user performing voice interaction with a virtual character; and, on the basis of the position information, adjusting the face orientation of the virtual character. On the basis of the position information of the user performing voice interaction the virtual character, the face orientation of the virtual character is adjusted in real time, so that the virtual character turns to the speaking user in real time on the basis of a position change of the user, thereby improving the experience of conversations between people and voice assistants.
Need to check novelty before this filing date? Find Prior Art

Description

Virtual avatar adjustment methods and devices, vehicles, electronic devices, storage media and products

[0001] Cross-reference of related applications

[0002] This disclosure claims priority to Chinese Patent Application No. 202410669038.1, filed on May 27, 2024, entitled “Virtual Image Adjustment Method and Apparatus, Vehicle, Electronic Device and Storage Medium”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to the field of sound processing technology, and includes, but is not limited to, a method and apparatus for adjusting virtual avatars, vehicles, electronic devices, storage media, and program products. Background Technology

[0004] With the rapid development of artificial intelligence technology, voice assistants have become an indispensable part of people's daily lives. People can interact with voice assistants to perform various convenient operations and access services. Currently, most voice assistants interact through text, but with advancements in speech recognition technology, more and more voice assistants are beginning to support voice interaction.

[0005] The process of a person conversing with a voice assistant can be understood as a conversation between two people. When one person is attentively listening to the other, their eyes are likely to be looking at the speaker. Therefore, in order to improve the conversation experience between a person and a voice assistant, a virtual image of the voice assistant will be summoned on the in-vehicle screen when a person converses with the voice assistant, to simulate the conversation between two people.

[0006] However, in related technologies, the virtual avatars correspond to one voice zone per direction, meaning the virtual avatars can only look in a few fixed directions. When a user activates the voice assistant, if the user moves while speaking, the virtual avatar's face cannot be facing the user in real time (for example, if the user moves from the center of the steering wheel to the armrest during a conversation, the virtual avatar is still looking at the steering wheel), resulting in an unnatural experience of dialogue between the user and the voice assistant. Summary of the Invention

[0007] This disclosure provides a method and apparatus for adjusting a virtual avatar, as well as a vehicle, electronic device, storage medium, and program product. Its main purpose is to solve the problem in related technologies where virtual avatars are designed with one voice register corresponding to one virtual avatar's direction, limiting the avatar to a few fixed directions. Furthermore, if the user moves while speaking after activating the voice assistant, the virtual avatar's face cannot be in real-time facing the user, resulting in an unnatural user experience during conversations.

[0008] This disclosure provides a method for adjusting a virtual avatar, including:

[0009] Obtain location information of the user's voice interaction with the virtual avatar;

[0010] The facial orientation of the virtual avatar is adjusted based on the location information.

[0011] This disclosure provides a virtual avatar adjustment device, including:

[0012] The acquisition unit is configured to acquire location information of the user's voice interaction with the virtual avatar.

[0013] The adjustment unit is configured to adjust the facial orientation of the virtual avatar based on the location information.

[0014] This disclosure provides a vehicle, wherein the vehicle includes the virtual avatar adjustment device as described above.

[0015] This disclosure provides an electronic device, including:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described above.

[0019] This disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to perform the methods described above.

[0020] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the method described above.

[0021] The virtual avatar adjustment method, apparatus, vehicle, electronic device, storage medium, and program product disclosed herein acquire location information of a user interacting with a virtual avatar via voice; and adjust the facial orientation of the virtual avatar based on the location information. Compared with related technologies, the embodiments of this disclosure adjust the facial orientation of the virtual avatar in real time based on the location information of the user interacting with the virtual avatar via voice, enabling the virtual avatar to turn towards the speaking user in real time according to changes in the user's location, thereby improving the experience of dialogue between humans and voice assistants.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0023] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0024] Figure 1 is a flowchart illustrating a virtual avatar adjustment method provided in an embodiment of this disclosure;

[0025] Figure 2 is a schematic diagram illustrating the principle of virtual image adjustment provided in an embodiment of this disclosure;

[0026] Figure 3 is a schematic diagram of a virtual image provided in an embodiment of this disclosure;

[0027] Figure 4 is a schematic flowchart of a face-turning angle calculation provided in an embodiment of this disclosure;

[0028] Figure 5 is a schematic diagram of a process for determining sound coordinates provided in an embodiment of this disclosure;

[0029] Figure 6A is a schematic diagram of a virtual avatar adjustment device provided in an embodiment of this disclosure;

[0030] Figure 6B is a schematic diagram of the structure of a virtual image adjustment device provided in an embodiment of this disclosure;

[0031] Figure 7 is a schematic diagram of the structure of a virtual image adjustment device provided in an embodiment of this disclosure;

[0032] Figure 8 is a schematic block diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0033] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0034] The following description, with reference to the accompanying drawings, outlines a method and apparatus for adjusting virtual avatars, as well as a vehicle, electronic device, storage medium, and program product according to embodiments of the present disclosure.

[0035] Figure 1 is a flowchart illustrating a virtual avatar adjustment method provided in an embodiment of this disclosure.

[0036] As shown in Figure 1, the method includes the following steps:

[0037] Step 101: Obtain the location information of the user's voice interaction with the virtual avatar.

[0038] Here, location information may include the user's sound source location or the location of key points on the head; the sound source location is the area where the user emits sound information, and the sound information refers to the sound information of the user's voice interaction with the virtual avatar.

[0039] It should be noted that the location of the sound source can be determined using a preset sound localization algorithm, while the location of key head points can be determined using an image recognition algorithm. These key head points can include the user's eye and mouth positions, among others, and are not specifically limited here. The following explanation will use location information as an example to illustrate the sound source location.

[0040] In this embodiment of the disclosure, before obtaining the location of the sound source of the user's voice interaction with the virtual avatar, the sound information of the user's voice interaction with the virtual avatar can be obtained; the location of the sound source can be determined by a preset sound localization algorithm based on the sound information.

[0041] In this embodiment of the disclosure, the user may include a first user, and the above method may further include: receiving a wake-up command from the first user, the wake-up command being used to instruct the virtual avatar to be woken up; displaying the virtual avatar on the screen; and adjusting the facial orientation of the virtual avatar according to the location information, which includes: adjusting the facial orientation of the virtual avatar to a first position according to the location information of the first user.

[0042] Here, there are no restrictions on the input method for the wake-up command; for example, it can be voice input or gesture input.

[0043] For example, the first user can input a wake-up command via voice to wake up the virtual avatar; at this time, the virtual avatar will be displayed on the screen; here, the screen may include the driver's screen, passenger's screen or rear screen of the vehicle.

[0044] Here, there are no specific restrictions on the display position of the virtual image on the screen. For example, the virtual image can be displayed at the top, middle, or left of the screen.

[0045] In this embodiment of the disclosure, the face of the virtual avatar can be adjusted to face a first position based on the location information of the first user; here, the first position may include: the location information of the first user when the first user inputs a wake-up command, which includes the location of the sound source or the location of key points on the head.

[0046] It should be noted that if the location information of the first user is the location of the sound source, then the first location is the location of the sound source when the first user inputs the wake-up command; if the location information of the first user is the location of the head key points, then the first location is the location of the head key points when the first user inputs the wake-up command.

[0047] In this embodiment of the disclosure, the method may further include: receiving a control command from the first user; adjusting the face of the virtual avatar from facing a first position to facing a second position according to the position information of the first user, wherein the second position includes: the position information of the first user when the first user inputs the control command, the position information including the sound source position or the head key point position.

[0048] Here, there are no restrictions on the input method of control commands; for example, it can be voice input or gesture input.

[0049] For example, the control command is used to instruct the virtual avatar to perform a response operation, which may be a voice reply, that is, a voice reply to the user's question, such as checking the weather; or a control action, such as opening a car window; no specific limitation is made here.

[0050] In this embodiment of the disclosure, if the first user continues to input control commands via voice after waking up the virtual avatar, and the first user moves during the input of control commands, then the face of the virtual avatar will be adjusted from facing the first position to facing the second position. It is understood that since the user's position before and after the movement may be the same or different, the first position may be the same as the second position or different from the second position.

[0051] It should be noted that if the location information of the first user is the location of the sound source, then the second location is the location of the sound source when the first user inputs the control command; if the location information of the first user is the location of the head key point, then the second location is the location of the head key point when the first user inputs the control command.

[0052] For example, when the first user moves from the center of the steering wheel to the armrest during the input of control commands, the virtual avatar's face is adjusted from facing the center of the steering wheel to facing the armrest. That is, the virtual avatar's face orientation changes according to the first user's position. The purpose of this adjustment is to ensure that the virtual avatar can look at the user in real time during the interaction, thereby improving the user's interactive experience.

[0053] In this embodiment of the disclosure, the user may further include a second user. After adjusting the face of the virtual avatar from facing a first position to facing a second position, the method further includes: receiving a control command from the second user; adjusting the face of the virtual avatar to face a third position according to the position information of the second user, wherein the third position includes: the position information of the second user when the second user inputs the control command; and after the virtual avatar executes the control command of the second user, adjusting the face of the virtual avatar from facing the third position to facing the position information of the first user.

[0054] Here, location information includes the location of the sound source or the location of key points on the head; the second user is a different user from the first user mentioned above; for example, if the first user is the primary driver, then the second user can be the passenger or a rear passenger, without specific limitations here.

[0055] For example, while the first user is interacting with the virtual avatar, the second user can also interact with the virtual avatar by inputting control commands through voice or gestures.

[0056] For example, after receiving the control command from the second user, if the virtual avatar is currently broadcasting interactive information to the first user, the control command from the second user will not be executed. However, a notification sound can be generated at this time, for example, playing a "ding" notification sound to remind the second user. If no interactive information is being broadcast to the first user at this time, the control command from the second user will be executed. After executing the control command from the second user, the virtual avatar's face will be adjusted to face a third position according to the position information of the second user. Here, the third position includes: the position of the second user's voice source or the position of the key point of the head when the second user inputs the control command.

[0057] It should be noted that if the location information of the second user is the location of the sound source, then the third location is the location of the sound source when the second user inputs the wake-up command; if the location information of the second user is the location of a key point on the head, then the third location is the location of a key point on the head when the second user inputs the wake-up command.

[0058] Furthermore, after the virtual avatar executes the control command, the virtual avatar's face is adjusted from facing the third position to facing the first user. In other words, after the virtual avatar faces the second user, it will turn back to face the first user after executing the control command.

[0059] For example, the size of the virtual avatar can change during the interaction between the user and the virtual avatar; for instance, when the user is far away from the virtual avatar, the virtual avatar can be enlarged appropriately during the interaction to improve the user's visual experience.

[0060] In this embodiment of the disclosure, the virtual avatar includes eyes and a facial contour. Adjusting the facial orientation of the virtual avatar according to the position information may include: adjusting the position of the virtual avatar's eyes relative to the facial contour and / or adjusting the shape of the facial contour, so that the virtual avatar's face is oriented towards the position information.

[0061] For example, the facial orientation of the virtual avatar can be adaptively adjusted by adjusting the position of the virtual avatar's eyes relative to the facial contour and / or adjusting the shape of the facial contour.

[0062] In this embodiment of the disclosure, the angle adjustment range corresponding to the orientation is 0-360 degrees. That is to say, the face of the virtual image can face various directions, such as left, right, up, down, etc., and can rotate at different angles in each direction.

[0063] Step 102: Adjust the facial orientation of the virtual avatar according to the location information.

[0064] In this embodiment of the disclosure, the image location coordinates of the virtual image can also be obtained, and the facial orientation of the virtual image can be adjusted according to the location information. This can include: adjusting the facial orientation of the virtual image according to the location information and the image location coordinates.

[0065] For example, as can be seen from the above, the location information can be the location of the sound source. Therefore, the facial orientation of the virtual image can be adjusted according to the location of the sound source and the coordinates of the image's location.

[0066] In this embodiment of the disclosure, adjusting the facial orientation of the virtual image based on the sound source location and the image location coordinates may include: determining sound coordinates based on the sound source location; calculating the turning angle of the virtual image using a preset angle algorithm based on the sound coordinates and the image location coordinates; and adjusting the facial orientation of the virtual image based on the turning angle.

[0067] In this embodiment of the disclosure, the sound source position can be transformed by a preset coordinate system to obtain sound coordinates, wherein the sound coordinates are the position coordinates corresponding to the area where the sound information is emitted, that is, the position coordinates corresponding to the sound source position.

[0068] In this embodiment, the image position coordinates are the position coordinates of the virtual image in a preset coordinate system. The preset coordinate system is a custom-defined coordinate system that can be pre-configured in the vehicle. Once the preset coordinate system is configured, it will not change. The subsequent image position coordinates and sound coordinates are all position coordinates in the preset coordinate system, that is, the image position coordinates and sound coordinates are in the same coordinate system. There are no restrictions on the form of the image position coordinates and sound coordinates. For example, image position coordinates (x1, y1, z1), sound coordinates (x2, y2, z2), etc. Regarding the setting of the preset coordinate system, for example, a spatial coordinate system can be set with the vehicle's center console position as the origin, or a spatial coordinate system can be set with the middle position of the vehicle's front as the origin, etc. Specifically, this embodiment does not limit the setting of the preset coordinate system.

[0069] It should be noted that different vehicles may have different coordinate systems, but once the vehicle's coordinate system is set, it will not be changed. Also, since the display position of the virtual avatar can be determined intuitively, after determining the preset coordinate system, the avatar's position coordinates can be directly determined based on the preset coordinate system. In summary, the avatar's position coordinates are also pre-configured in the vehicle. That is, after determining the preset coordinate system, the preset coordinate system and the avatar's position coordinates can be directly configured in the vehicle for subsequent use.

[0070] The sound information is the voice emitted by the user when interacting or communicating with the virtual avatar. The method of acquiring the sound information may include, but is not limited to, acquiring it through a microphone array (MIC array). The MIC array is a mic array formed by deploying multiple microphones (mic) at various points in the vehicle. The mic array can locate the approximate location of the sound source based on the distance of the sound and output the (X,Y,Z) spatial coordinates of the sound source.

[0071] The preset sound localization algorithm is a custom-set algorithm, such as beamforming, multiple signal classification (MUSIC) algorithm, cross-spectral estimation, etc. The specific implementation of this disclosure is not limited.

[0072] In this embodiment of the disclosure, the preset angle algorithm is a custom-set algorithm, such as the arctangent function (ArcTan), etc. Specifically, this embodiment of the disclosure does not limit the preset angle algorithm.

[0073] To facilitate understanding of the implementation process of the embodiments of this disclosure, a schematic diagram of the principle of virtual image adjustment is provided, as shown in Figure 2. In this diagram, the voice image represents the virtual image, the sound coordinate space point represents the area where the sound information is emitted, and the preset coordinate system is set with the center position of the front of the vehicle as the origin. The turning angle includes angle A and angle B.

[0074] As shown in Figure 2, it should be noted that the display position of the voice avatar, i.e., the virtual avatar, is not fixed. The possible display positions include, but are not limited to, displaying it through the rear screen, the driver's side screen, and the passenger side screen. When displaying the virtual avatar, only one virtual avatar will be displayed in one position at a time; multiple virtual avatars will not appear simultaneously. Furthermore, the display position of the virtual avatar can be determined by the location of the wake-up command. For example, if the wake-up command is uttered from the driver's side, the virtual avatar will be displayed on the driver's side screen; if the wake-up command is uttered from the passenger side, the virtual avatar will be displayed on the passenger side screen; if the wake-up command is not uttered from either the driver's or passenger side screen, the virtual avatar will be displayed on the rear screen.

[0075] It should be noted that the coordinates of the virtual avatar's position at each display location have been pre-configured in the vehicle. Once the display location of the virtual avatar is determined, the corresponding coordinates can be directly accessed.

[0076] In this embodiment of the disclosure, the virtual image's turning angle is calculated using a preset angle algorithm based on the sound coordinates and the image's position coordinates; the virtual image's facial orientation is then adjusted according to the turning angle.

[0077] In this embodiment of the disclosure, after rotating the face of the virtual image according to the face-turning angle, the face of the virtual image can be adjusted to the target direction.

[0078] For ease of understanding, this disclosure provides a schematic diagram of a virtual avatar display, as shown in Figure 3. In this diagram, looking left indicates a leftward turn of the face, with the target direction to the left; looking right indicates a rightward turn of the face, with the target direction to the right; looking down indicates a downward turn of the face, with the target direction to the bottom; and looking down to the right indicates a downward turn of the face, with the target direction to the bottom right. By adjusting the facial orientation of the virtual avatar, during a conversation between the user and the voice assistant, even if the user changes position while speaking or reclines their seat, the avatar will always be facing the user.

[0079] The virtual avatar adjustment method disclosed herein obtains the avatar's position coordinates and the sound information of the voice interaction with the virtual avatar, and determines the sound coordinates based on the sound information using a preset sound positioning algorithm; wherein, the sound coordinates are the position coordinates corresponding to the area where the sound information is emitted; the virtual avatar's face-turning angle is calculated based on the sound coordinates and the avatar's position coordinates using a preset angle algorithm; and the virtual avatar's facial orientation is adjusted based on the face-turning angle. Compared with related technologies, the embodiments of this disclosure determine the relative position between the voice-emitting position and the virtual avatar based on the sound coordinates of the sound information and the virtual avatar's position coordinates, and calculate the avatar's rotation angle in real time based on the sound position and the avatar's position, enabling the virtual avatar to turn towards the speaking user in real time according to the rotation angle, thereby improving the experience of human-voice assistant dialogue.

[0080] In one possible implementation of this embodiment, as a refinement of step 102 above, when calculating the face-turning angle, it is necessary to calculate it using the coordinate values ​​of the sound coordinates and the coordinate values ​​of the image position coordinates. Therefore, in order to accurately calculate the face-turning angle, it can be implemented in the following way, but is not limited to: extracting the first spatial coordinate value of the sound coordinates and extracting the second spatial coordinate value of the image position coordinates; and calculating the face-turning angle of the virtual image using the preset angle algorithm based on the first spatial coordinate value and the second spatial coordinate value.

[0081] Related to the above embodiments, since the face-turning angle is not a single-direction angle but includes both lateral and longitudinal face-turning angles, the calculation of the face-turning angle requires calculating angles in both directions, namely the lateral and longitudinal face-turning angles. Simultaneously, the coordinate values ​​of the sound coordinates and the image position coordinates are determined by the coordinate system in which they reside. Therefore, to understand the face-turning angle calculation process, this disclosure provides a flowchart illustrating the face-turning angle calculation, as shown in Figure 4, including:

[0082] Step 401: Obtain the first horizontal coordinate value, the first vertical coordinate value, and the first vertical coordinate value from the first spatial coordinate value, and obtain the second horizontal coordinate value, the second vertical coordinate value, and the second vertical coordinate value from the second spatial coordinate value.

[0083] In this embodiment, when calculating the face-turning angle, it is necessary to calculate based on the coordinate values ​​of the voice coordinates and the image position coordinates. The coordinate values ​​can be directly extracted from the voice coordinates and image position coordinates. The representation of the coordinate values ​​is determined by the coordinate system in which the voice coordinates and image position coordinates are located, which includes, but is not limited to, a Cartesian coordinate system. Specifically, this embodiment does not impose any limitations on the coordinate system in which the voice coordinates and image position coordinates are located.

[0084] To facilitate understanding of the implementation process of the embodiments of this disclosure, a Cartesian coordinate system will be used as an example for the following description.

[0085] In a Cartesian coordinate system, regarding the first spatial coordinate value and the second spatial coordinate value, for example: if the image position coordinates are determined to be (x1, y1, z1) and the sound coordinates are (x2, y2, z2), then the first horizontal coordinate value in the first spatial coordinate value is x2, the first vertical coordinate value is y2, and the first vertical coordinate value is z2; the second horizontal coordinate value in the second spatial coordinate value is x1, the second vertical coordinate value is y1, and the second vertical coordinate value is z1.

[0086] Step 402: Calculate the lateral face-turning angle of the virtual image based on the first horizontal coordinate value, the second horizontal coordinate value, the first vertical coordinate value, and the second vertical coordinate value, using a preset correspondence between coordinates and face-turning angles; wherein, the lateral face-turning angle is used to adjust the facial orientation of the virtual image laterally.

[0087] In this embodiment of the disclosure, the preset correspondence between coordinates and turning angle is a custom-selected correspondence formula, such as the arctangent function formula. For ease of understanding, the arctangent function (ArcTan) will be used as an example for the following explanation.

[0088] ArcTan is a trigonometric function, an inverse function, whose domain is all real numbers and whose range is all angles whose non-terminal sides coincide with the coordinate axes. In computer graphics and some mathematical applications, ArcTan is often used to calculate angles from given coordinates (such as the coordinates of a point on the unit circle). ArcTan can also be extended to multiple quadrants, and angles across the entire coordinate plane can be obtained by adding or subtracting kπ (where k is an integer).

[0089] The lateral turning angle can be calculated using, but is not limited to, formula (1): Angle A = (ArcTan((y2-y1) / (x2-x1)))*180 / π Formula (1)

[0090] Wherein, angle A is the horizontal turning angle, x2 is the first horizontal coordinate value, x1 is the second horizontal coordinate value, y2 is the first vertical coordinate value, and y1 is the second vertical coordinate value.

[0091] Step 403: Calculate the vertical face-turning angle of the virtual image based on the first horizontal coordinate value, the second horizontal coordinate value, the first vertical coordinate value, and the second vertical coordinate value, using the preset correspondence between coordinates and face-turning angles; wherein, the vertical face-turning angle is used to vertically adjust the facial orientation of the virtual image; wherein, the face-turning angle includes the horizontal face-turning angle and the vertical face-turning angle.

[0092] In this embodiment of the disclosure, the longitudinal turning angle can be calculated using, but is not limited to, formula (2): Angle B = (ArcTan((z2-z1) / (x2-x1)))*180 / π Formula (2)

[0093] Wherein, angle B is the longitudinal turning angle, x2 is the first horizontal coordinate value, x1 is the second horizontal coordinate value, z2 is the first vertical coordinate value, and z1 is the second vertical coordinate value.

[0094] It should be noted that the lateral turning angle and the longitudinal turning angle together constitute the turning angle.

[0095] In one possible implementation of this disclosure, since it is necessary to calculate the turning angle using sound coordinates and image position coordinates, and the image position coordinates are fixed and can be directly retrieved from the vehicle database, different turning angles are mainly determined using different sound coordinates to adjust the virtual image. To accurately determine the turning angle, the sound source location needs to be accurately determined. Regarding the determination of the sound source location, this disclosure provides a flowchart illustrating the determination of sound coordinates, as shown in Figure 5, including:

[0096] Step 501: Obtain the time difference of arrival of the sound information and the device layout information of the preset radio equipment array; wherein, the time difference of arrival is the time difference between the arrival of the sound information at each radio equipment in the preset radio equipment array, and the preset radio equipment array is used to obtain the sound information.

[0097] In this embodiment of the disclosure, the preset audio recording device array is a custom-configured device array, which consists of multiple audio recording devices arranged in a predetermined layout, such as a microphone array (MIC array). The MIC array is formed by deploying multiple microphones (mic) at various points inside the vehicle. Specifically, this embodiment of the disclosure does not limit the audio recording devices.

[0098] The time difference of arrival (TDOA) is a parameter used to determine the time difference between the arrival of a wave (such as a sound wave or radio wave) at two or more receivers. The TDOA includes multiple time differences, i.e., the time difference between the sound information from the sound source and each receiving device. When the wave source (the source of the sound information) emits a sound wave, the sound wave is received by multiple receivers (preset receiving devices). Since the distance between the receivers and the speed of sound propagation in the medium are known, the time difference between the sound wave arriving at each receiver can be calculated. By analyzing these time differences, the location of the wave source can be determined.

[0099] The device layout information refers to the layout information of each preset radio device in the preset radio device array, including but not limited to: the placement position of each preset radio device, the distance between each preset radio device, etc. The device layout information can be obtained directly from the vehicle database.

[0100] Step 502: Calculate the sound source location of the sound information using the preset geometric relationship algorithm based on the device layout information and the time difference of arrival.

[0101] In this embodiment of the disclosure, the preset geometric relationship algorithm is a custom-selected algorithm that includes one or more equations. All equations are determined by the difference in sound wave arrival times (which can be determined by the wave arrival time difference) and the distance between the receiving devices (which can be determined by the device layout information). Algorithms, such as maximum likelihood estimation or least squares method, can then be used to optimize the estimation of the sound source location.

[0102] The time difference of arrival (TDOA) can be used to determine the distance between the area from which the sound information is emitted and each preset receiving device. The device layout information can be used to determine the location of the area from which the sound information is emitted, i.e., the location of the sound source, based on the layout relationship of each preset receiving device.

[0103] Furthermore, the sound source location can be transformed using preset coordinate system information to obtain the sound coordinates.

[0104] In this embodiment of the disclosure, the preset coordinate system information is a preset coordinate system pre-configured in the vehicle. After determining the location of the sound source, the coordinates of the sound source location, i.e., the sound coordinates, can be directly determined in the preset coordinate system.

[0105] In one possible implementation of this disclosure, since the display position of the virtual avatar is not fixed and can be displayed on multiple screens, it is necessary to determine the display position of the virtual avatar by the position of the wake-up command in order to improve the experience of human-voice assistant dialogue. Regarding the determination of the display position of the virtual avatar, the following methods can also be used, but are not limited to: responding to the wake-up command of the virtual avatar, determining the command coordinates of the wake-up command by the preset sound positioning algorithm; wherein, the command coordinates are the position coordinates corresponding to the area where the wake-up command is issued; displaying the virtual avatar on the screen in the area corresponding to the command coordinates, and displaying the face of the virtual avatar facing the direction of the command coordinates.

[0106] In this embodiment of the disclosure, the preset sound localization algorithm is a custom-set algorithm. For details on the preset sound localization algorithm and the determination process of the command coordinates, please refer to the description of steps 501-502, and therefore will not be repeated here.

[0107] The area corresponding to the command coordinates includes, but is not limited to, the driver's seat area, the passenger seat area, and the rear seat area. If the area corresponding to the command coordinates is the driver's seat area, a virtual image will be displayed on the driver's seat screen; if the area corresponding to the command coordinates is the passenger seat area, a virtual image will be displayed on the passenger seat screen; and if the area corresponding to the command coordinates is the rear seat area, a virtual image will be displayed on the rear seat screen.

[0108] In this embodiment of the disclosure, the screen includes at least a driver's side screen, a passenger side screen, or a rear-seat screen, and the range of change of the virtual image's turning angle on the driver's side screen and the passenger side screen is greater than the range of change of the virtual image's turning angle on the rear-seat screen.

[0109] Understandably, compared to the distance between rear-seat users and the rear screen, the distance between the driver's side user and the driver's side screen, and the distance between the passenger side user and the passenger side screen, are closer. The closer the distance, the wider the field of view; the farther the distance, the narrower the field of view of the virtual avatar. Therefore, the range of changes in the virtual avatar's turning angle on the driver's side screen and the passenger side screen is greater than the range of changes in the virtual avatar's turning angle on the rear screen.

[0110] In one possible implementation of this disclosure, in order to accurately acquire sound information and locate the sound source, the following method may also be used, but is not limited to: acquiring the sound information of the user's voice interaction with the virtual image through a preset radio equipment array.

[0111] In summary, the embodiments disclosed herein can achieve the following effects:

[0112] This embodiment of the present disclosure determines the relative position between the speaking position and the virtual image based on the sound coordinates of the sound information and the image position coordinates of the virtual image, and calculates the rotation angle of the image in real time based on the sound position and the image position, so that the virtual image can turn towards the speaking user in real time according to the rotation angle, thereby improving the experience of human-voice assistant dialogue.

[0113] Corresponding to the virtual avatar adjustment method described above, this disclosure also proposes a virtual avatar adjustment device. Since the device embodiments of this disclosure correspond to the method embodiments described above, details not disclosed in the device embodiments can be referred to the method embodiments described above, and will not be repeated here.

[0114] Figure 6A is a schematic diagram of a virtual avatar adjustment device provided in an embodiment of this disclosure. As shown in Figure 6A, it includes:

[0115] Acquisition unit 61 is configured to acquire location information of the user's voice interaction with the virtual avatar;

[0116] The adjustment unit 64 is configured to adjust the facial orientation of the virtual image based on the location information.

[0117] The virtual avatar adjustment device provided in this disclosure acquires location information of the user interacting with the virtual avatar via voice; and adjusts the facial orientation of the virtual avatar according to the location information. Compared with related technologies, the embodiments of this disclosure adjust the facial orientation of the virtual avatar in real time based on the location information of the user interacting with the virtual avatar via voice, so that the virtual avatar can turn towards the speaking user in real time according to changes in the user's position, thereby improving the experience of dialogue between humans and voice assistants.

[0118] Furthermore, in one possible implementation of this embodiment, the acquisition unit 61 is further configured to: acquire the image position coordinates of the virtual image;

[0119] The adjustment unit 64 is further configured to: adjust the facial orientation of the virtual image according to the location information and the image location coordinates.

[0120] Furthermore, in one possible implementation of this disclosure embodiment, the user includes a first user, as shown in FIG7. The device further includes a display unit 65, which is configured as follows:

[0121] Receive a wake-up command from a first user, the wake-up command being used to instruct the virtual avatar to be woken up;

[0122] The virtual image is displayed on the screen;

[0123] The adjustment unit 64 is further configured to: adjust the facial orientation of the virtual image to a first position according to the location information of the first user, wherein the first position includes: the location information of the first user when the first user inputs a wake-up command.

[0124] Furthermore, in one possible implementation of this embodiment, the adjustment unit 64 is further configured to: receive a control command from the first user, the control command being used to instruct the virtual avatar to perform a control action;

[0125] Based on the location information of the first user, the face of the virtual avatar is adjusted from facing a first position to facing a second position, where the second position includes the location information of the first user when the first user inputs a control command.

[0126] Furthermore, in one possible implementation of this disclosure embodiment, the user further includes a second user, and the adjustment unit 64 is further configured to:

[0127] Receive control commands from the second user;

[0128] Based on the location information of the second user, the facial orientation of the virtual avatar is adjusted to a third position, wherein the third position includes: the location information of the second user when the second user inputs control commands;

[0129] After the virtual avatar executes the control command of the second user, the virtual avatar's face is adjusted from facing the third position to facing the first user.

[0130] Furthermore, in one possible implementation of this disclosure embodiment, the virtual avatar includes eyes and a facial contour, and the adjustment unit 64 is further configured to: adjust the position of the virtual avatar's eyes relative to the facial contour and / or adjust the shape of the facial contour according to the position information, so that the virtual avatar's face faces the position information.

[0131] Furthermore, in one possible implementation of this disclosure embodiment, the angle adjustment range is 0-360 degrees.

[0132] Furthermore, in one possible implementation of this disclosure embodiment, the location information includes the user's sound source location or head key point location.

[0133] Furthermore, in one possible implementation of this embodiment, the adjustment unit 64 is further configured as follows:

[0134] The facial orientation of the virtual avatar is adjusted based on the location of the sound source and the coordinates of the avatar's location.

[0135] Furthermore, in one possible implementation of this embodiment, as shown in FIG6B, the device further includes a determining unit 62, which is configured as follows:

[0136] Acquire audio information of the user's voice interaction with the virtual avatar;

[0137] The location of the sound source is determined by a preset sound localization algorithm based on the sound information; wherein, the location of the sound source is the area where the sound information is emitted.

[0138] Furthermore, in one possible implementation of this embodiment, the adjustment unit 64 is further configured as follows:

[0139] The sound coordinates are determined based on the location of the sound source; wherein, the sound coordinates are the location coordinates corresponding to the region emitting the sound information;

[0140] Based on the sound coordinates and the image position coordinates, the turning angle of the virtual image is calculated using a preset angle algorithm;

[0141] The virtual avatar's facial orientation is adjusted according to the stated turning angle.

[0142] Furthermore, in one possible implementation of this embodiment, as shown in FIG6B, the device further includes a computing unit 63, which is configured as follows:

[0143] Extract the first spatial coordinate value of the sound coordinates and the second spatial coordinate value of the image position coordinates; calculate the turning angle of the virtual image based on the first spatial coordinate value and the second spatial coordinate value using the preset angle algorithm.

[0144] Furthermore, in one possible implementation of this disclosure embodiment, as shown in FIG7, the computing unit 63 includes:

[0145] Extraction module 631 is configured to obtain the first horizontal coordinate value, the first vertical coordinate value, and the first vertical coordinate value in the first spatial coordinate value, and to obtain the second horizontal coordinate value, the second vertical coordinate value, and the second vertical coordinate value in the second spatial coordinate value;

[0146] The calculation module 632 is configured to calculate the lateral face-turning angle of the virtual image based on the first horizontal coordinate value, the second horizontal coordinate value, the first vertical coordinate value, and the second vertical coordinate value, through a preset correspondence between coordinates and face-turning angles; wherein, the lateral face-turning angle is used to adjust the facial orientation of the virtual image laterally;

[0147] The calculation module 632 is further configured to calculate the vertical face-turning angle of the virtual image based on the first horizontal coordinate value, the second horizontal coordinate value, the first vertical coordinate value, and the second vertical coordinate value, through the preset correspondence between the coordinates and the face-turning angle; wherein, the vertical face-turning angle is used to vertically adjust the facial orientation of the virtual image;

[0148] Furthermore, in one possible implementation of this disclosure embodiment, as shown in FIG7, the determining unit 62 includes:

[0149] The acquisition module 621 is configured to acquire the time difference of arrival of the sound information and the device layout information of a preset radio equipment array; wherein, the time difference of arrival is the time difference between the arrival of the sound information at each radio equipment in the preset radio equipment array, and the preset radio equipment array is used to acquire the sound information;

[0150] The calculation module 622 is configured to calculate the sound source location of the sound information based on the device layout information and the time difference of arrival using a preset geometric relationship algorithm; wherein, the sound source location is the region emitting the sound information.

[0151] Furthermore, in one possible implementation of this disclosure embodiment, before acquiring the sound information of the user's voice interaction with the virtual avatar, the determining unit 62 is further configured to, in response to the wake-up command of the virtual avatar, calculate the command coordinates of the wake-up command using the preset sound positioning algorithm; wherein, the command coordinates are the location coordinates corresponding to the area where the wake-up command was issued;

[0152] The display unit 65 is configured to display the virtual image in the screen area corresponding to the command coordinates, and to display the face of the virtual image as facing the command coordinates.

[0153] Furthermore, in one possible implementation of this disclosure embodiment, the screen includes a driver's side screen, a passenger side screen, or a rear-seat screen, and the range of change in the turning angle of the virtual image on the driver's side screen and the passenger side screen is greater than the range of change in the turning angle of the virtual image on the rear-seat screen.

[0154] Furthermore, in one possible implementation of this embodiment, the acquisition unit 61 is further configured to acquire the sound information of the user interacting with the virtual image via a preset radio equipment array.

[0155] In this disclosure embodiment, a vehicle is also provided, wherein the vehicle is equipped with a virtual avatar adjustment device. It should be noted that the foregoing explanation of the method embodiments also applies to the devices in this disclosure embodiment, as the principles are the same, and this disclosure embodiment is not further limited thereto. According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0156] Figure 8 illustrates a schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0157] As shown in Figure 8, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 802 or a computer program loaded from storage unit 808 into RAM (Random Access Memory) 803. RAM 803 can also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. I / O (Input / Output) interface 805 is also connected to bus 804.

[0158] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0159] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the virtual avatar adjustment method. For example, in some embodiments, the virtual avatar adjustment method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the aforementioned virtual avatar adjustment method by any other suitable means (e.g., by means of firmware).

[0160] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0161] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0163] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0164] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0165] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0166] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0167] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0168] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A virtual figure adjusting method, comprising: obtaining position information of a user interacting with a virtual figure by voice; adjusting a face orientation of the virtual figure according to the position information.

2. The method of claim 1, wherein, The method further comprises: obtaining a figure position coordinate of the virtual figure; adjusting the face orientation of the virtual figure according to the position information, comprising: adjusting the face orientation of the virtual figure according to the position information and the figure position coordinate.

3. The method according to claim 1 or 2, characterized in that, The user comprises a first user, and the method further comprises: receiving a wake-up instruction of the first user, the wake-up instruction being used to instruct to wake up the virtual figure; displaying the virtual figure on a screen; adjusting the face orientation of the virtual figure according to the position information, comprising: adjusting the face orientation of the virtual figure to a first position according to the position information of the first user, the first position comprising position information of the first user when the first user inputs the wake-up instruction.

4. The method of claim 3, wherein, The method further comprises: receiving a control instruction of the first user; adjusting the face orientation of the virtual figure from the first position to a second position according to the position information of the first user, the second position comprising position information of the first user when the first user inputs the control instruction.

5. The method according to claim 3 or 4, characterized in that, The user further comprises a second user, and the method further comprises: receiving a control instruction of the second user; adjusting the face orientation of the virtual figure to a third position according to the position information of the second user, the third position comprising position information of the second user when the second user inputs the control instruction; adjusting the face orientation of the virtual figure from the third position to the position information of the first user after the virtual figure executes the control instruction of the second user.

6. The method according to any one of claims 1 to 5, wherein, The virtual figure comprises eyes and a face contour, and adjusting the face orientation of the virtual figure according to the position information comprises: adjusting a position of the eyes relative to the face contour and / or adjusting a face contour shape according to the position information, so that the face of the virtual figure is oriented to the position information.

7. The method according to any one of claims 1 to 6, wherein, The corresponding angle adjustment range of the orientation is 0-360 degrees.

8. The method according to any one of claims 1 to 7, wherein, The position information comprises a sound source position or a head key point position of the user.

9. The method of claim 8, wherein, Adjusting the face orientation of the virtual figure according to the position information and the figure position coordinate comprises: adjusting the face orientation of the virtual figure according to the sound source position and the figure position coordinate.

10. The method of claim 9, wherein, The method further comprises: obtaining sound information of the user interacting with the virtual figure by voice; determining a sound source position by a preset sound positioning algorithm according to the sound information; wherein the sound source position is a region emitting the sound information.

11. The method of claim 9 or 10, wherein, Adjusting the face orientation of the virtual figure according to the sound source position and the figure position coordinate comprises: determining a sound coordinate according to the sound source position; wherein the sound coordinate is a position coordinate corresponding to the region emitting the sound information; calculating a face turning angle of the virtual figure by a preset angle algorithm according to the sound coordinate and the figure position coordinate; adjusting the face orientation of the virtual figure according to the face turning angle.

12. The method of claim 11, wherein, According to the sound coordinates and the image position coordinates, the turning head angle of the virtual image is calculated by a preset angle algorithm, including: extracting a first spatial coordinate value of the sound coordinates and a second spatial coordinate value of the image position coordinates; calculating the turning head angle of the virtual image by the preset angle algorithm according to the first spatial coordinate value and the second spatial coordinate value.

13. The method of claim 12, wherein, The turning head angle includes a horizontal turning head angle and a vertical turning head angle, and the turning head angle of the virtual image is calculated by the preset angle algorithm according to the first spatial coordinate value and the second spatial coordinate value, including: obtaining a first horizontal coordinate value, a first vertical coordinate value and a first longitudinal coordinate value in the first spatial coordinate value, and obtaining a second horizontal coordinate value, a second vertical coordinate value and a second longitudinal coordinate value in the second spatial coordinate value; calculating the horizontal turning head angle of the virtual image by a preset coordinate and turning head angle correspondence according to the first horizontal coordinate value, the second horizontal coordinate value, the first vertical coordinate value and the second vertical coordinate value; wherein the horizontal turning head angle is used to adjust the face direction of the virtual image horizontally; calculating the vertical turning head angle of the virtual image by the preset coordinate and turning head angle correspondence according to the first horizontal coordinate value, the second horizontal coordinate value, the first longitudinal coordinate value and the second longitudinal coordinate value; wherein the vertical turning head angle is used to adjust the face direction of the virtual image vertically.

14. The method of claim 10, wherein, The sound source position is determined by a preset sound positioning algorithm according to the sound information, including: obtaining a time difference of arrival of the sound information and device layout information of a preset sound device array; wherein the time difference of arrival is the time difference of arrival of the sound information at each sound device in the preset sound device array, and the preset sound device array is used to obtain the sound information; calculating the sound source position of the sound information by a preset geometric relationship algorithm according to the device layout information and the time difference of arrival.

15. The method of claim 10, wherein, Before obtaining the sound information of the user interacting with the virtual image by voice, the method further includes: in response to the wake-up instruction of the virtual image, determining the instruction coordinates of the wake-up instruction by the preset sound positioning algorithm; wherein the instruction coordinates are the position coordinates corresponding to the region where the wake-up instruction is issued; displaying the virtual image in the screen of the region corresponding to the instruction coordinates, and displaying the face direction of the virtual image as facing the direction of the instruction coordinates.

16. The method of claim 15, wherein, The screen includes a main driver screen, a co-driver screen or a rear screen, and the change range of the turning head angle of the virtual image on the main driver screen and the co-driver screen is greater than the change range of the turning head angle of the virtual image on the rear screen.

17. The method of any one of claims 10, 14-16, wherein, Obtaining the sound information of the user interacting with the virtual image by voice includes: obtaining the sound information of the user interacting with the virtual image by voice by a preset sound device array.

18. A virtual image adjusting device, comprising: an acquisition unit configured to obtain position information of a user interacting with a virtual image by voice; an adjusting unit configured to adjust a face orientation of the virtual image according to the position information. 19.A vehicle comprising the virtual image adjusting apparatus according to claim 18.

20. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-17. 21.A non-transitory computer readable storage medium storing computer instructions for causing a computer to perform the method of any one of claims 1-17. 22.A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-17.

Citation Information

Patent Citations

  • Voice interaction method and device based on virtual robot image and vehicle-mounted equipment intelligent control system

    CN111124123A

  • Virtual digital human sight line following interaction method

    CN114265543A

  • Virtual digital human sight line following system and method based on sound and vehicle

    CN116643713A

  • Virtual character orientation adjusting method and device, electronic equipment and storage medium

    CN116870475A