Method, device, vehicle-mounted device and medium for generating expressions of in-vehicle virtual avatars

By collecting and processing the key points of the sample face, binding it to the face of the vehicle virtual image, and driving expression generation based on the target face key points, the technical problem of virtual image expression generation in the vehicle scene is solved, and a customized, simple and efficient expression generation effect is achieved.

CN115116106BActive Publication Date: 2025-06-13GREAT WALL MOTOR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210043607.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-14
Publication Date
2025-06-13
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

The prior art cannot be directly applied to the expressions of virtual avatars in vehicle-mounted scenarios, because there are usually no devices on the car that can obtain in-depth information.

Method used

By collecting sample faces of multiple sample users, extracting key points of the sample face and obtaining their coordinates, generating faces to be bound, and binding them to the faces of the pre-created car virtual image. When the target face key points of the target user are collected, the face key points of the vehicle virtual image are driven according to the target face key points to generate the expression of the vehicle virtual image.

Benefits of technology

It realizes the customized generation of vehicle virtual image expressions in vehicle-mounted scenarios, improving the interactive experience of drivers and passengers and the fun of the interactive process, and the process is simple and does not require large-scale calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116106B_ABST
    Figure CN115116106B_ABST
Patent Text Reader

Abstract

The embodiments of this application are applicable to the field of intelligent vehicle technologies, and provide a method, a device, an in-vehicle device, and a medium for generating expressions of an in-vehicle virtual image. The method includes: collecting sample faces of a plurality of sample users; extracting a plurality of sample face key points from the sample faces, and obtaining the coordinates of each of the sample face key points; generating a face to be bound according to the coordinates of the plurality of sample face key points; binding the face to be bound to the face of a pre-generated in-vehicle virtual image; when target face key points of a target user are collected, driving the face key points of the in-vehicle virtual image according to the target face key points to generate an expression of the in-vehicle virtual image. By adopting the above method, the expression of the in-vehicle virtual image can be customized, the interaction experience of the driver and passengers can be improved, and the fun of the interaction process can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application belong to the technical field of intelligent vehicles, and particularly relate to a method, device, vehicle-mounted device and medium for generating expressions of in-vehicle virtual avatars. Background Technique

[0002] The vehicle-mounted device can interact with the driver and passengers through the virtual avatar, improving the interaction experience of the driver and passengers. For example, the virtual avatar can announce navigation information to the driver, or greet the passengers when they get on and off the vehicle.

[0003] In the prior art, 3D modeling of the human face can be performed through an infrared camera and 3D structured light to generate the expression of the virtual avatar. However, obtaining 3D depth information is required to generate the expression of the virtual avatar using the above method. In a vehicle, devices capable of obtaining depth information are often not equipped, so the above method cannot be directly applied to the vehicle-mounted scenario. Therefore, in the prior art, the expressions of virtual avatars in the vehicle-mounted scenario are usually preset and manually adjusted by the driver, and the expressions are relatively single and fixed. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide a method, device, vehicle-mounted device and medium for generating expressions of in-vehicle virtual avatars, which can customize the expressions of in-vehicle virtual avatars to improve the interaction experience of the driver and passengers and enhance the interest of the interaction process.

[0005] The first aspect of the embodiments of the present application provides a method for generating an expression of an in-vehicle virtual avatar, including:

[0006] Collecting sample human faces of multiple sample users;

[0007] Extracting multiple sample human face key points from the sample human faces, and obtaining the coordinates of each sample human face key point;

[0008] Generating a to-be-bound human face according to the coordinates of multiple sample human face key points;

[0009] Binding the to-be-bound human face to the face of a pre-generated in-vehicle virtual avatar;

[0010] When the target human face key points of the target user are collected, driving the face key points of the in-vehicle virtual avatar according to the target human face key points to generate the expression of the in-vehicle virtual avatar.

[0011] The second aspect of the embodiments of the present application provides an apparatus for generating an expression of an in-vehicle virtual avatar, including:

[0012] A sample human face collection module, configured to collect sample human faces of multiple sample users;

[0013] The sample face key point extraction module is used to extract multiple sample face key points from the sample face;

[0014] The sample face key point coordinate acquisition module is used to acquire the coordinates of each sample face key point;

[0015] The to-be-bound face generation module is used to generate a to-be-bound face according to the coordinates of multiple sample face key points;

[0016] The vehicle-mounted virtual image binding module is used to bind the to-be-bound face to the face of a pre-generated vehicle-mounted virtual image;

[0017] The expression generation module is used to drive the face key points of the vehicle-mounted virtual image according to the target face key points of the target user when the target face key points of the target user are collected, so as to generate the expression of the vehicle-mounted virtual image.

[0018] In a third aspect of the embodiments of the present application, a vehicle-mounted device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the expression generation method of the vehicle-mounted virtual image as described in the first aspect above is implemented.

[0019] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the expression generation method of the vehicle-mounted virtual image as described in the first aspect above is implemented.

[0020] In a fifth aspect of the embodiments of the present application, a computer program product is provided. When the computer program product runs on a computer, the computer is enabled to execute the expression generation method of the vehicle-mounted virtual image as described in the first aspect above.

[0021] Compared with the prior art, the embodiments of the present application have the following advantages:

[0022] In the embodiments of the present application, by collecting the sample faces of multiple sample users, multiple sample face key points can be extracted from the sample faces, and the coordinates of each sample face key point can be obtained. After generating a to-be-bound face according to the coordinates of multiple sample face key points, the to-be-bound face can be bound to the face of a pre-generated vehicle-mounted virtual image. In this way, when the target face key points of the target user are collected, the vehicle-mounted device can drive the face key points of the vehicle-mounted virtual image according to the target face key points to generate the expression of the vehicle-mounted virtual image that is the same as the current expression of the target user. The process of generating the expression of the vehicle-mounted virtual image in the embodiments of the present application is simple and does not require large-scale calculation, and can be customized and applied to the vehicle-mounted scenario. Description of the Drawings

[0023] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0024] Figure 1 It is a schematic flowchart of generating an expression of a virtual image in the prior art;

[0025] Figure 2 It is a schematic diagram of a method for generating an expression of an in-vehicle virtual image provided by an embodiment of the present application;

[0026] Figure 3 It is a schematic diagram of an implementation manner of S203 in a method for generating an expression of an in-vehicle virtual image provided by an embodiment of the present application;

[0027] Figure 4 It is a schematic diagram of an implementation manner of S2031 in a method for generating an expression of an in-vehicle virtual image provided by an embodiment of the present application;

[0028] Figure 5 It is a schematic diagram of another implementation manner of S2031 in a method for generating an expression of an in-vehicle virtual image provided by an embodiment of the present application;

[0029] Figure 6 It is a schematic diagram of an implementation manner of S204 in a method for generating an expression of an in-vehicle virtual image provided by an embodiment of the present application;

[0030] Figure 7 It is a schematic diagram of facial key points of an in-vehicle virtual image provided by an embodiment of the present application;

[0031] Figure 8 It is a schematic diagram of an implementation manner of S205 in a method for generating an expression of an in-vehicle virtual image provided by an embodiment of the present application;

[0032] Figure 9 It is a schematic diagram of a device for generating an expression of an in-vehicle virtual image provided by an embodiment of the present application;

[0033] Figure 10 It is a schematic diagram of an in-vehicle device provided by an embodiment of the present application. Detailed implementation manners

[0034] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, in order to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present application.

[0035] See Figure 1 , which is a schematic flow diagram of a process for generating expressions of virtual avatars in the prior art. When generating the expressions of virtual avatars according to the process shown in Figure 1 , a device capable of acquiring depth information can be used to adopt a face detection algorithm to track the change of the user's face position in real time, calculate the position relationship of the face in space, and synchronously capture the expression; then, the captured expression is matched with the current face expression database. Finally, the position mapping relationship between the virtual avatar and the user's face is calculated in real time, and the expression of the virtual avatar is rendered based on the matched expression and the position mapping relationship. To solve the problem that in-vehicle scenarios do not have devices that can acquire depth information, and the above method cannot be directly used to generate the expressions of virtual avatars, the embodiments of the present application provide a method for generating expressions of in-vehicle virtual avatars, which can simply and efficiently generate the expressions of in-vehicle virtual avatars based on the face images collected by ordinary in-vehicle cameras.

[0036] The technical solution of the present application will be described below through specific embodiments.

[0037] Refer to Figure 2 , which shows a schematic diagram of a method for generating expressions of in-vehicle virtual avatars provided by the embodiments of the present application. The method may specifically include the following steps:

[0038] S201. Collect the face images of multiple sample users.

[0039] This method can be applied to in-vehicle devices, which can be electronic devices installed in various types or models of vehicles. Exemplarily, the in-vehicle device can be an in-vehicle computer device, an in-vehicle voice assistant, etc. The embodiments of the present application do not limit the specific type of the in-vehicle device.

[0040] In the embodiments of the present application, in order to customize the generation of the expressions of in-vehicle virtual avatars, the face images of multiple sample users can be collected first for modeling. The above sample users can be any users, and their face images are the face images of the sample users. The face images of the sample users can be collected by an in-vehicle camera, and the in-vehicle camera can transmit the face images of multiple sample users to the in-vehicle device, and the in-vehicle device completes the subsequent modeling process.

[0041] In a possible implementation of the embodiment of the present application, the sample facial images of multiple sample users can also be collected by other camera devices. The sample facial images collected by these camera devices can be transmitted to other types of terminal devices, and the terminal devices complete the modeling. The model obtained by the terminal device can be deployed in the vehicle-mounted device for generating expressions for the vehicle-mounted virtual avatar subsequently.

[0042] S202. Extract multiple sample facial key points from the sample facial images, and obtain the coordinates of each of the sample facial key points.

[0043] In the embodiment of the present application, the sample facial key points can be the facial key points in the sample facial images, such as the eyes, nose, lips of the sample user, and the key points on the outer contour of the face, and can also include multiple key points in the middle area of the face. The embodiment of the present application does not limit this.

[0044] It should be noted that for each sample facial image, the above multiple sample facial key points can be extracted. Exemplarily, for sample user A, the key points such as eyes, nose, lips, and the key points on the outer contour of the face can be extracted from his sample facial image a; for sample user B, the above key points such as eyes, nose, lips, and the key points on the outer contour of the face can also be extracted from his sample facial image b.

[0045] Each sample facial key point can have corresponding coordinates, and the coordinates can include the x-axis coordinate and the y-axis coordinate. That is, the coordinates of each sample facial key point can be expressed as (x, y).

[0046] In the embodiment of the present application, MediaPipe or a similar algorithm can be used to obtain the three-dimensional depth map of the sample facial image, so as to extract the coordinates of the sample facial key points.

[0047] MediaPipe is a data stream processing machine learning application development framework developed and open-sourced by Google. It is a graph-based data processing pipeline for building data sources in various forms, such as video, audio, sensor data, and any time series data. MediaPipe is cross-platform and can run on embedded platforms (such as Raspberry Pi), mobile devices (iOS and Android), workstations, and servers, and supports mobile GPU acceleration. Using MediaPipe, machine learning tasks can be built into a data stream pipeline represented by a graph module, which can include an inference model and a streaming media processing function.

[0048] Using MediaPipe, the vehicle-mounted device can control an ordinary vehicle-mounted camera to collect facial images of the user or sample users, and then the corresponding three-dimensional depth map can be obtained without the need to rely on other depth devices.

[0049] S203. Generate a face to be bound based on the coordinates of multiple sample face key points.

[0050] In the embodiments of the present application, the face to be bound may be a face obtained after normalizing multiple sample faces. By normalizing multiple sample faces, the influence of different face differences on the final effect can be avoided.

[0051] In a possible implementation manner of the embodiments of the present application, as Figure 3 shown, specifically generating the face to be bound based on the coordinates of multiple sample face key points in S203 may include the following sub-steps S2031 - S2032:

[0052] S2031. Normalize the coordinates of multiple sample face key points to obtain the normalized coordinates of the sample face key points.

[0053] In the embodiments of the present application, normalizing the sample face can be achieved by normalizing the coordinates of the sample face key points.

[0054] In specific implementation, since the coordinates of the sample face key points include the x-axis coordinate and the y-axis coordinate, when normalizing, the x-axis coordinate and the y-axis coordinate can be normalized respectively.

[0055] In a possible implementation manner of the embodiments of the present application, as Figure 4 shown, specifically normalizing the coordinates of multiple sample face key points in S2031 to obtain the normalized coordinates of the sample face key points may include the following sub-steps S311 - S314:

[0056] S311. Determine the length and width of the face frame of each sample face, and determine the length and width of the face frame of the face to be bound.

[0057] In the embodiments of the present application, the length and width of the face frame of the sample face and the length and width of the face frame of the face to be bound may be determined first. The face frame of the face to be bound is the face frame of the normalized sample face.

[0058] S312. Calculate the first ratio between the length of the face frame of the face to be bound and the length of the face frame of each sample face respectively, and calculate the second ratio between the width of the face frame of the face to be bound and the width of the face frame of each sample face respectively.

[0059] For the lengths of the face bounding boxes of the face to be bound and the sample faces, the first ratio between the length of the face bounding box of the face to be bound and the length of the face bounding box of each sample face, and the second ratio between the width of the face bounding box of the face to be bound and the width of the face bounding box of each sample face can be calculated.

[0060] In the embodiments of the present application, the above first ratio and second ratio can be calculated using the following formulas:

[0061] First ratio = h avg / h i ……(1)

[0062] Second ratio = w avg / w i ……(2)

[0063] Wherein, h avg is the length of the face bounding box of the face to be bound, w avg is the width of the face bounding box of the face to be bound, h i is the length of the face bounding box of the i-th sample face, and w i is the width of the face bounding box of the i-th sample face.

[0064] S313. Calculate the x-axis coordinates of the normalized sample face key points according to the first ratio and the x-axis coordinates of the multiple sample face key points.

[0065] S314. Calculate the y-axis coordinates of the normalized sample face key points according to the second ratio and the y-axis coordinates of the multiple sample face key points; wherein, the x-axis coordinates and y-axis coordinates of the normalized sample face key points together constitute the coordinates of the normalized sample face key points.

[0066] Then, the x-axis coordinates of the normalized sample face key points can be calculated according to the first ratio and the x-axis coordinates of the multiple sample face key points, and the y-axis coordinates of the normalized sample face key points can be calculated according to the second ratio and the y-axis coordinates of the multiple sample face key points.

[0067] In a specific implementation, the x-axis coordinates and y-axis coordinates of the normalized sample face key points can be calculated using the following formulas:

[0068] x = x i *(h avg / h i )……(3)

[0069] y = y i *(w avg / w i )……(4)

[0070] Wherein, x is the x-axis coordinate of the normalized sample face key point, y is the y-axis coordinate of the normalized sample face key point, x i is the x-axis coordinate of the i-th sample face key point, y i is the y-axis coordinate of the i-th sample face key point.

[0071] In this way, the x-axis coordinate and y-axis coordinate of the normalized sample face key point together constitute the coordinate of the normalized sample face key point. By normalizing multiple sample face key points in the embodiment of the present application, the influence on the modeling process caused by different face differences can be avoided, so that subsequent face binding and expression generation can be carried out according to a relatively unified standard, reducing the difficulty for the vehicle-mounted device to recognize and process face key points and improving the efficiency of expression generation.

[0072] In a possible implementation manner of the embodiment of the present application, the coordinate of the sample face key point may further include a depth value, that is, the z-axis coordinate. As Figure 5 shown, the normalization process of the coordinates of multiple sample face key points in S2031 to obtain the coordinates of the normalized sample face key point may further include the following sub-steps S315-S317:

[0073] S315. Determine the depth of the face frame of each sample face, and determine the depth of the face frame of the face to be bound.

[0074] S316. Calculate the third ratio between the depth of the face frame of the face to be bound and the depth of the face frame of each sample face respectively.

[0075] S317. Calculate the z-axis coordinate of the normalized sample face key point according to the third ratio and the z-axis coordinates of multiple sample face key points; wherein, the x-axis coordinate, y-axis coordinate and z-axis coordinate of the normalized sample face key point together constitute the coordinate of the normalized sample face key point.

[0076] Similar to calculating the x-axis coordinate and y-axis coordinate of the normalized sample face key point, when calculating the z-axis coordinate of the normalized sample face key point, the depth of the face frame of each sample face and the depth of the face frame of the face to be bound can be determined first, and the third ratio between the depth of the face frame of the face to be bound and the depth of the face frame of each sample face can be calculated, so that the z-axis coordinate of the normalized sample face key point can be calculated according to the third ratio and the z-axis coordinates of multiple sample face key points. In this way, the x-axis coordinate, y-axis coordinate and z-axis coordinate of the normalized sample face key point together constitute the coordinate of the normalized sample face key point, that is, (x, y, z).

[0077] S2032. Generate the to-be-bound human face according to the coordinates of the sample human face key points after normalization.

[0078] In the embodiment of the present application, the to-be-bound human face can be generated according to the coordinates (x, y, z) of the sample human face key points after normalization. The above to-be-bound human face can be regarded as the average face obtained by normalizing the sample human faces of multiple sample users. Since the z-axis coordinate can represent the movement of the human face key points in the depth direction of the human face, in the embodiment of the present application, by adding the z-axis coordinate to calibrate the coordinates of the sample human face key points, the accuracy and vividness of the expressions of the subsequent generated in-vehicle virtual avatars can be further improved.

[0079] S204. Bind the to-be-bound human face to the face of the pre-generated in-vehicle virtual avatar.

[0080] In the embodiment of the present application, the in-vehicle virtual avatar can be pre-produced and configured in the in-vehicle device. The in-vehicle virtual avatar can be any type of virtual avatar, such as humanoid, animal or object shape, etc. The embodiment of the present application does not limit the specific type of the in-vehicle virtual avatar.

[0081] In a possible implementation manner of the embodiment of the present application, binding the to-be-bound human face to the face of the pre-generated in-vehicle virtual avatar can be achieved by binding the human face key points of the to-be-bound human face to the face key points of the in-vehicle virtual avatar.

[0082] Specifically, as Figure 6 shown, the binding of the to-be-bound human face to the face of the pre-generated in-vehicle virtual avatar in S204 can specifically include the following sub-steps S2041 - S2043:

[0083] S2041. Calibrate the face key points of the in-vehicle virtual avatar and obtain the human face key points of the to-be-bound human face.

[0084] In the embodiment of the present application, the human face key points of the to-be-bound human face can include multiple key points of the eyes, nose, lips, outer contour of the human face, and the middle area of the human face; correspondingly, the face key points of the in-vehicle virtual avatar can include multiple key points of the eyes, nose, lips, outer contour of the face, and the middle area of the face.

[0085] As Figure 7As shown, it is a schematic diagram of facial key points of an in-vehicle virtual image provided by an embodiment of the present application. The facial key points shown in this schematic diagram include not only the key points on the eyes, nose, lips, the external contour of the human face or the face of the in-vehicle virtual image, but also the key points in the middle area of the face. In this way, compared with the prior art where extracting human face key points only involves the key points on the eyes, nose, lips, the external contour of the human face or the face, the embodiment of the present application can more precisely control the expressions of the subsequent generated in-vehicle virtual image by extracting the key points in the middle area of the human face or the face.

[0086] It should be noted that the facial key points of the in-vehicle virtual image can be calibrated when the in-vehicle virtual image is produced and generated. And the human face key points of the to-be-bound human face can be obtained by normalizing the sample human face key points of the sample user through algorithms such as MediaPipe.

[0087] S2042. Determine the corresponding relationship between the facial key points of the in-vehicle virtual image and the human face key points of the to-be-bound human face.

[0088] In the embodiment of the present application, determining the corresponding relationship between the facial key points of the in-vehicle virtual image and the human face key points of the to-be-bound human face may refer to determining which position relationship of the human face key points of the to-be-bound human face is consistent with each facial key point of the in-vehicle virtual image. Specifically, the key points on the eyes of the in-vehicle virtual image should have a corresponding relationship with the key points on the eyes of the to-be-bound human face, the key points on the nose of the in-vehicle virtual image should have a corresponding relationship with the key points on the nose of the to-be-bound human face, the key points on the lips of the in-vehicle virtual image should have a corresponding relationship with the key points on the lips of the to-be-bound human face, and so on.

[0089] S2043. Bind the human face key points of the to-be-bound human face having the corresponding relationship with the facial key points of the in-vehicle virtual image.

[0090] After determining the corresponding relationship between the facial key points of the in-vehicle virtual image and the human face key points of the to-be-bound human face, the human face key points having the above relationship can be bound to the facial key points, so as to achieve the binding between the in-vehicle virtual image and the to-be-bound human face.

[0091] After completing the above binding, the eyes of the in-vehicle virtual image can correspond to the eyes of the to-be-bound human face. Correspondingly, the nose, lips, the external contour of the human face, and the middle area of the human face of the in-vehicle virtual image can also correspond to the nose, lips, the external contour of the face, and the middle area of the face of the to-be-bound human face.

[0092] S205. When the target face key points of the target user are collected, drive the face key points of the in-vehicle virtual image according to the target face key points to generate the expression of the in-vehicle virtual image.

[0093] In the embodiment of the present application, the target user may be a driver or passenger in the vehicle whose face image is collected in real time by the in-vehicle device controlling the in-vehicle camera. For example, a car driver, a passenger, etc. The in-vehicle device can collect the target face key points of the target user in the manner introduced above. That is, the in-vehicle device can control the in-vehicle camera to capture the face image of the driver or passenger to obtain the target face. Then, the in-vehicle device can use algorithms such as MediaPipe to extract key points such as the eyes, nose, lips, the outer contour of the face, and the middle area of the face of the target face, and obtain the key point coordinates of these key points. Generally, since the middle area of the face occupies a relatively large area in the whole face, by extracting the key points of the middle area of the face, key points at different positions on the face can be obtained more evenly, which helps to control the expression of the in-vehicle virtual image more precisely.

[0094] When the in-vehicle device collects the target face key points, it can drive the face key points of the in-vehicle virtual image according to the target face key points, thereby generating the expression of the in-vehicle virtual image.

[0095] In a possible implementation manner of the embodiment of the present application, as Figure 8 shown, driving the face key points of the in-vehicle virtual image according to the target face key points in S205 to generate the expression of the in-vehicle virtual image may specifically include the following sub-steps S2051 - S2053:

[0096] S2051. Determine the positional relationship between the target face key points and the face key points of the in-vehicle virtual image.

[0097] In the embodiment of the present application, determining the positional relationship between the target face key points and the face key points of the in-vehicle virtual image may be to determine the relationship between the target face key points and the corresponding face key points. For example, determine the correspondence between the eye key points in the target face key points and the eye key points in the face key points of the in-vehicle virtual image.

[0098] S2052. Obtain the key point coordinates of the target face key points.

[0099] In the embodiment of the present application, algorithms such as MediaPipe can be used to obtain the key point coordinates of these target face key points.

[0100] S2053. For any facial key point of the in-vehicle virtual image, adjust the coordinate of the facial key point to be the same as the key point coordinate of the target human face key point with the same positional relationship, so as to generate an expression of the in-vehicle virtual image; wherein, the expression is the same as the current expression of the target user.

[0101] In practical applications, the coordinates of the facial key points of the in-vehicle virtual image can be adjusted to be the same as the key point coordinates of the target human face key points with the same positional relationship, thereby generating an expression of the in-vehicle virtual image. In this way, the generated expression of the in-vehicle virtual image can be the same as the current expression of the target user.

[0102] In the embodiment of the present application, by collecting the sample human faces of multiple sample users, multiple sample human face key points can be extracted from the sample human faces, and the coordinates of each sample human face key point can be obtained. After generating the to-be-bound human face according to the coordinates of multiple sample human face key points, the to-be-bound human face can be bound to the face of the pre-generated in-vehicle virtual image. In this way, when the target human face key points of the target user are collected, the vehicle-mounted device can drive the facial key points of the in-vehicle virtual image according to the target human face key points to generate an expression of the in-vehicle virtual image that is the same as the current expression of the target user. The process of generating the expression of the in-vehicle virtual image in the embodiment of the present application is simple and does not require large-scale calculations, and can be customized for application in the vehicle-mounted scenario. For example, in the welcome scenario when the vehicle-mounted device is powered on, the driver can customize it with the expressions of his relatives, making the welcome interface more familiar and cordial. Another example is that when a passenger gets on the vehicle, the expression of the in-vehicle virtual image can be generated in real time according to the current expression of the passenger, increasing the interaction fun.

[0103] It should be noted that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or posterior, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0104] Refer to Figure 9 , which shows a schematic diagram of an expression generation device for an in-vehicle virtual image provided by an embodiment of the present application. The device may specifically include a sample human face collection module 901, a sample human face key point extraction module 902, a sample human face key point coordinate acquisition module 903, a to-be-bound human face generation module 904, an in-vehicle virtual image binding module 905, and an expression generation module 906, where:

[0105] The sample human face collection module 901 is configured to collect sample human faces of multiple sample users;

[0106] The sample human face key point extraction module 902 is configured to extract multiple sample human face key points from the sample human faces;

[0107] The sample face key point coordinate acquisition module 903 is used to acquire the coordinates of each of the sample face key points;

[0108] The to-be-bound face generation module 904 is used to generate a to-be-bound face according to the coordinates of multiple sample face key points;

[0109] The vehicle-mounted virtual image binding module 905 is used to bind the to-be-bound face to the face of a pre-generated vehicle-mounted virtual image;

[0110] The expression generation module 906 is used to drive the face key points of the vehicle-mounted virtual image according to the target face key points of the target user when the target face key points of the target user are collected, so as to generate an expression of the vehicle-mounted virtual image.

[0111] In an embodiment of the present application, the to-be-bound face generation module 904 may specifically be used to: perform normalization processing on the coordinates of multiple sample face key points to obtain the coordinates of the normalized sample face key points; and generate the to-be-bound face according to the coordinates of the normalized sample face key points.

[0112] In an embodiment of the present application, the coordinates of the sample face key points include an x-axis coordinate and a y-axis coordinate, and the to-be-bound face generation module 904 may further be used to: determine the length and width of the face frame of each sample face, and determine the length and width of the face frame of the to-be-bound face; calculate a first ratio between the length of the face frame of the to-be-bound face and the length of the face frame of each sample face respectively, and calculate a second ratio between the width of the face frame of the to-be-bound face and the width of the face frame of each sample face respectively; calculate the x-axis coordinate of the normalized sample face key point according to the first ratio and the x-axis coordinates of multiple sample face key points; calculate the y-axis coordinate of the normalized sample face key point according to the second ratio and the y-axis coordinates of multiple sample face key points; wherein, the x-axis coordinate and the y-axis coordinate of the normalized sample face key point together constitute the coordinates of the normalized sample face key point.

[0113] In the embodiments of the present application, the coordinates of the sample human face key points further include the z-axis coordinates, and the to-be-bound human face generation module 904 can further be used to: determine the depth of the face frame of each sample human face, and determine the depth of the face frame of the to-be-bound human face; calculate the third ratio between the depth of the face frame of the to-be-bound human face and the depth of the face frame of each sample human face respectively; calculate the normalized z-axis coordinates of the sample human face key points according to the third ratio and the z-axis coordinates of multiple sample human face key points; wherein, the normalized x-axis coordinates, y-axis coordinates and z-axis coordinates of the sample human face key points together constitute the coordinates of the normalized sample human face key points.

[0114] In the embodiments of the present application, the vehicle-mounted virtual image binding module 905 can specifically be used to: calibrate the facial key points of the vehicle-mounted virtual image, and obtain the human face key points of the to-be-bound human face; determine the corresponding relationship between the facial key points of the vehicle-mounted virtual image and the human face key points of the to-be-bound human face; bind the human face key points of the to-be-bound human face with the corresponding relationship to the facial key points of the vehicle-mounted virtual image.

[0115] In the embodiments of the present application, the human face key points include multiple key points of the eyes, nose, lips, outer contour of the human face and the middle area of the human face of the to-be-bound human face; correspondingly, the facial key points include multiple key points of the eyes, nose, lips, outer contour of the face and the middle area of the face of the vehicle-mounted virtual image.

[0116] In the embodiments of the present application, the expression generation module 906 can specifically be used to: determine the positional relationship between the target human face key points and the facial key points of the vehicle-mounted virtual image; obtain the key point coordinates of the target human face key points; for any facial key point of the vehicle-mounted virtual image, adjust the coordinates of the facial key point to be the same as the key point coordinates of the target human face key points with the same positional relationship, so as to generate the expression of the vehicle-mounted virtual image; wherein, the expression is the same as the current expression of the target user.

[0117] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the related parts, reference can be made to the description in the method embodiment section.

[0118] Refer to Figure 10 , which shows a schematic diagram of a vehicle-mounted device provided by the embodiments of the present application. As Figure 10As shown in the figure, the in-vehicle device 1000 in the embodiments of the present application includes: a processor 1010, a memory 1020, and a computer program 1021 stored in the memory 1020 and executable on the processor 1010. When the processor 1010 executes the computer program 1021, the steps in each of the embodiments of the above-described method for generating expressions of in-vehicle virtual avatars are implemented, such as Figure 2 the steps S201 to S205 shown in the figure. Alternatively, when the processor 1010 executes the computer program 1021, the functions of each module / unit in each of the above-described device embodiments are implemented, such as Figure 9 the functions of the modules 901 to 906 shown in the figure.

[0119] Exemplarily, the computer program 1021 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 1020 and executed by the processor 1010 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments can be used to describe the execution process of the computer program 1021 in the in-vehicle device 1000. For example, the computer program 1021 can be divided into a sample face collection module, a sample face key point extraction module, a sample face key point coordinate acquisition module, a face to be bound generation module, an in-vehicle virtual avatar binding module, and an expression generation module. The specific functions of each module are as follows:

[0120] The sample face collection module is used to collect sample faces of multiple sample users;

[0121] The sample face key point extraction module is used to extract multiple sample face key points from the sample faces;

[0122] The sample face key point coordinate acquisition module is used to acquire the coordinates of each sample face key point;

[0123] The face to be bound generation module is used to generate a face to be bound according to the coordinates of multiple sample face key points;

[0124] The in-vehicle virtual avatar binding module is used to bind the face to be bound to the face of a pre-generated in-vehicle virtual avatar;

[0125] The expression generation module is used to drive the face key points of the in-vehicle virtual avatar according to the target face key points of the target user when the target face key points of the target user are collected, so as to generate the expression of the in-vehicle virtual avatar.

[0126] The vehicle-mounted device 1000 may be a vehicle-mounted computer device, a vehicle-mounted voice assistant, or other devices in the foregoing various embodiments. The vehicle-mounted device 1000 may include, but is not limited to, a processor 1010 and a memory 1020. Those skilled in the art can understand that Figure 10 This is merely an example of the vehicle-mounted device 1000 and does not constitute a limitation on the vehicle-mounted device 1000. It may include more or fewer components than those shown in the figure, or combine certain components, or have different components. For example, the vehicle-mounted device 1000 may further include an input / output device, a network access device, a bus, etc.

[0127] The processor 1010 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0128] The memory 1020 may be an internal storage unit of the vehicle-mounted device 1000, such as a hard disk or memory of the vehicle-mounted device 1000. The memory 1020 may also be an external storage device of the vehicle-mounted device 1000, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the vehicle-mounted device 1000. Further, the memory 1020 may also include both the internal storage unit and the external storage device of the vehicle-mounted device 1000. The memory 1020 is used to store the computer program 1021 and other programs and data required by the vehicle-mounted device 1000. The memory 1020 may also be used to temporarily store data that has been output or is to be output.

[0129] An embodiment of the present application also discloses a vehicle-mounted device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for generating expressions of the vehicle-mounted virtual avatar as described in the foregoing various embodiments.

[0130] The embodiments of the present application also disclose a computer-readable storage medium storing a computer program, which when executed by a processor, implements the method for generating expressions of in-vehicle virtual avatars as described in the foregoing respective embodiments.

[0131] The embodiments of the present application also disclose a computer program product, which when running on a computer, causes the computer to execute the method for generating expressions of in-vehicle virtual avatars as described in the foregoing respective embodiments.

[0132] In the embodiments of the present application, if the various functions implemented by the in-vehicle device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, the steps of the above respective method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to the in-vehicle virtual avatar expression generation device / in-vehicle device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0133] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0134] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0135] In the embodiments provided in the present application, it should be understood that the disclosed device / vehicle-mounted device and method can be implemented in other ways. For example, the device / vehicle-mounted device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0136] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0137] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for generating expressions of in-vehicle avatars, characterized in that, it includes: Collecting sample human faces of multiple sample users; Extracting multiple sample human face key points from the sample human faces, and obtaining the coordinates of each of the sample human face key points; Generating a to-be-bound human face according to the coordinates of the multiple sample human face key points; Binding the to-be-bound human face to the face of a pre-generated in-vehicle avatar; wherein, the in-vehicle avatar is pre-produced and configured in an in-vehicle device; When target human face key points of a target user are collected, driving the facial key points of the in-vehicle avatar according to the target human face key points to generate an expression of the in-vehicle avatar.

2. The method according to claim 1, characterized in that, the generating a to-be-bound human face according to the coordinates of the multiple sample human face key points includes: Performing normalization processing on the coordinates of the multiple sample human face key points to obtain the normalized coordinates of the sample human face key points; Generating the to-be-bound human face according to the normalized coordinates of the sample human face key points.

3. The method according to claim 2, characterized in that, the coordinates of the sample human face key points include x-axis coordinates and y-axis coordinates, and the performing normalization processing on the coordinates of the multiple sample human face key points to obtain the normalized coordinates of the sample human face key points includes: Determining the length and width of the human face frame of each of the sample human faces, and determining the length and width of the human face frame of the to-be-bound human face; Calculating a first ratio between the length of the human face frame of the to-be-bound human face and the length of the human face frame of each of the sample human faces respectively, and calculating a second ratio between the width of the human face frame of the to-be-bound human face and the width of the human face frame of each of the sample human faces respectively; Calculating the x-axis coordinates of the normalized sample human face key points according to the first ratio and the x-axis coordinates of the multiple sample human face key points; Calculating the y-axis coordinates of the normalized sample human face key points according to the second ratio and the y-axis coordinates of the multiple sample human face key points; wherein, the x-axis coordinates and y-axis coordinates of the normalized sample human face key points jointly constitute the coordinates of the normalized sample human face key points.

4. The method according to claim 3, characterized in that, the coordinates of the sample human face key points further include z-axis coordinates, and the performing normalization processing on the coordinates of the multiple sample human face key points to obtain the normalized coordinates of the sample human face key points further includes: Determining the depth of the human face frame of each of the sample human faces, and determining the depth of the human face frame of the to-be-bound human face; Calculating a third ratio between the depth of the human face frame of the to-be-bound human face and the depth of the human face frame of each of the sample human faces respectively; Calculating the z-axis coordinates of the normalized sample human face key points according to the third ratio and the z-axis coordinates of the multiple sample human face key points; wherein, the x-axis coordinates, y-axis coordinates and z-axis coordinates of the normalized sample human face key points jointly constitute the coordinates of the normalized sample human face key points.

5. The method according to any one of claims 1 - 4, wherein, the binding of the to - be - bound human face to the face of a pre - generated in - vehicle virtual image includes: calibrating the facial key points of the in - vehicle virtual image and obtaining the human face key points of the to - be - bound human face; determining the correspondence between the facial key points of the in - vehicle virtual image and the human face key points of the to - be - bound human face; binding the human face key points of the to - be - bound human face having the correspondence to the facial key points of the in - vehicle virtual image.

6. The method according to claim 5, wherein, the human face key points include multiple key points of the eyes, nose, lips, outer contour of the human face, and the middle area of the human face of the to - be - bound human face; correspondingly, the facial key points include multiple key points of the eyes, nose, lips, outer contour of the face, and the middle area of the face of the in - vehicle virtual image.

7. The method according to any one of claims 1 - 4 or 6, wherein, the driving of the facial key points of the in - vehicle virtual image according to the target human face key points to generate the expression of the in - vehicle virtual image includes: determining the positional relationship between the target human face key points and the facial key points of the in - vehicle virtual image; obtaining the key point coordinates of the target human face key points; for any facial key point of the in - vehicle virtual image, adjusting the coordinates of the facial key point to be the same as the key point coordinates of the target human face key point having the same positional relationship to generate the expression of the in - vehicle virtual image; wherein, the expression is the same as the current expression of the target user.

8. An in - vehicle virtual image expression generation device, wherein, it includes: a sample human face acquisition module for acquiring sample human faces of multiple sample users; a sample human face key point extraction module for extracting multiple sample human face key points from the sample human faces; a sample human face key point coordinate acquisition module for acquiring the coordinates of each sample human face key point; a to - be - bound human face generation module for generating a to - be - bound human face according to the coordinates of multiple sample human face key points; an in - vehicle virtual image binding module for binding the to - be - bound human face to the face of a pre - generated in - vehicle virtual image; wherein, the in - vehicle virtual image is pre - fabricated and configured in an in - vehicle device; an expression generation module for driving the facial key points of the in - vehicle virtual image according to the target human face key points when the target human face key points of the target user are acquired to generate the expression of the in - vehicle virtual image.

9. An in - vehicle device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, it implements the in - vehicle virtual image expression generation method according to any one of claims 1 - 7.

10. A computer - readable storage medium storing a computer program, wherein, when the computer program is executed by a processor, it implements the in - vehicle virtual image expression generation method according to any one of claims 1 - 7.

Citation Information

Patent Citations

  • Dynamic image generation method, device, apparatus, and storage medium

    CN109147017A