An image configuration method and device, a vehicle terminal and a storage medium

By collecting and analyzing images of target objects within the vehicle terminal, a realistic virtual image of the voice assistant is generated, solving the problem of a monotonous voice assistant image in the vehicle terminal and improving user experience and system performance.

CN116342764BActive Publication Date: 2026-03-20CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-21
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

The virtual avatars of the voice assistants in vehicles are poorly designed, lack a realistic human feel, and fail to provide users with a visually dynamic and lifelike experience.

Method used

By collecting images of target objects within the vehicle terminal, extracting their features, and generating virtual avatar configuration files for the voice assistant based on these features, the virtual avatar features are matched and expanded using an avatar database to generate a realistic virtual avatar, and action data is set according to the usage scenario.

Benefits of technology

The generated voice assistant avatar is more dynamic and realistic, meeting users' diverse and personalized needs, improving user experience, and optimizing CPU and memory usage to avoid high usage affecting other operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116342764B_ABST
    Figure CN116342764B_ABST
Patent Text Reader

Abstract

The application relates to an image configuration method and device, a vehicle terminal and a storage medium, and relates to the technical field of artificial intelligence. The method is applied to a vehicle terminal and comprises the following steps: collecting an image of a target object in the vehicle terminal; extracting an image feature of the target object from the image of the target object; and generating a virtual image configuration file of a voice assistant based on the image feature of the target object. Thus, the virtual image of the voice assistant configured in this way is lifelike, lively, stereoscopic, and has a strong stereoscopic effect, thereby solving the problems of single configuration, few types and lack of real virtual human feeling in related technologies when a virtual image of a voice assistant of a vehicle terminal is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, in particular to the field of vehicle terminal intelligent voice recognition technology and three-dimensional engine image processing technology, and specifically relates to an image configuration method and device, a vehicle terminal and a storage medium. BACKGROUND

[0002] With the continuous development of intelligent voice technology, the scenarios in which users use voice assistants are becoming more and more rich. With the development of the automotive industry, users have higher and higher requirements for the technology of vehicle terminals. In order to improve the user experience, manufacturers will add a voice assistant of a vehicle terminal to the vehicle terminal for users to use. The voice assistant of the vehicle terminal provides many convenient operations for users on the vehicle.

[0003] In related technologies, there are many types of construction of virtual images of voice assistants of vehicle terminals, such as two-dimensional icon types, three-dimensional icon types, or two-dimensional virtual human types, but the constructed images are relatively single, do not have a real virtual human feeling, and cannot bring a visually dynamic and realistic feeling to users. SUMMARY

[0004] The present application provides an image configuration method and device, a vehicle terminal and a storage medium to at least solve the technical problem of single virtual image of a voice assistant of a vehicle terminal in related technologies. The technical solution of the present application is as follows:

[0005] According to a first aspect of the present application, an image configuration method is provided, applied to a vehicle terminal, and the method comprises: collecting an image of a target object in the vehicle terminal; extracting an image feature of the target object from the image of the target object; and generating a virtual image configuration file of a voice assistant based on the image feature of the target object.

[0006] According to the above technical means, when constructing a virtual image of a voice assistant of a vehicle terminal, the image of a target object is collected, the image feature of the target object is extracted from the image of the target object, and then a virtual image configuration file of a voice assistant is generated based on the image feature of the target object. As can be seen, the image configuration method provided by the present application can generate a configuration file according to the image feature of the target object, and then generate a virtual image of a voice assistant according to the configuration file. In this way, the virtual image of the voice assistant obtained by configuration has a virtual human feeling and is more dynamic and realistic, which enriches the virtual image of the voice assistant and can solve the problem of single image construction, few types and lack of real virtual human feeling when constructing a virtual image of a voice assistant of a vehicle terminal in related technologies.

[0007] In a possible implementation, the method further includes: in response to modification of the image feature of the virtual image by the user, obtaining the modified image feature of the virtual image; and updating the virtual image configuration file according to the modified image feature of the virtual image.

[0008] According to the technical means described above, the number of the image feature of the virtual image of the voice assistant that matches the image feature of the target object can be found from the image database; and the virtual image configuration file of the voice assistant can be generated according to the number of the image feature of the virtual image. Different from the technical problem that the image of the virtual image of the voice assistant is single and does not have a real virtual human feeling in the related art, the image feature of the target object is matched with the image feature of the virtual image of the voice assistant stored in the image database, so that the virtual image of the voice assistant generated has a certain similarity with the target object, the real feeling of the virtual image of the voice assistant is improved, and the virtual image of the voice assistant is more lifelike. Meanwhile, the image feature of the virtual image of the voice assistant in the image database is expandable, and the type of the image feature in the image database can be expanded and the number of the image feature in the image database can be increased from time to time, so as to improve the similarity between the image feature of the virtual image of the voice assistant and the image feature of the target object, make the virtual image of the voice assistant generated more flexible and lifelike, and improve the accuracy of the virtual image configuration of the voice assistant.

[0009] In a possible implementation, the method further includes: obtaining action data of the virtual image in multiple scenes; and generating the virtual image configuration file based on the image feature of the target object and the action data of the virtual image in the multiple scenes.

[0010] According to the technical means described above, different action data can be set for the virtual image of the voice assistant according to different use scenes, so that the image of the voice assistant is more lively and flexible, and the use experience of the user is improved.

[0011] In a possible implementation, the method further includes: in response to modification of the image feature of the virtual image by the user, obtaining the modified image feature of the virtual image; and updating the virtual image configuration file according to the modified image feature of the virtual image.

[0012] According to the technical means, unlike the virtual image of the voice assistant in the related art, which is the same or only a few appearance types are available for the user to select, the technical means can meet the diversified and personalized needs of the user, and can provide the user with a selection interface for configuring the virtual image of the voice assistant. In addition to the virtual image of the voice assistant matched by the vehicle terminal, the user can also freely customize the virtual image of the voice assistant of the vehicle terminal according to his / her preferences. Meanwhile, the technical means is simple to operate and convenient to use. The user only needs to perform a few simple operations to realize a series of complex operations such as scanning of the image of the user, feature extraction, and image generation, thereby reducing the learning cost of the user, improving the participation of the user, and bringing a good user experience to the user.

[0013] In a possible implementation, the method further includes: in a first scenario, in response to the operation of the user waking up the voice assistant, generating three-dimensional image data of the virtual image of the voice assistant based on the virtual image configuration file; displaying the three-dimensional virtual image of the voice assistant on the display interface of the vehicle terminal based on the three-dimensional image data; the first scenario is a scenario in which the CPU and / or memory of the vehicle terminal has an occupation rate less than or equal to a preset threshold; or in a second scenario, in response to the operation of the user waking up the voice assistant, generating two-dimensional image data of the virtual image of the voice assistant based on the virtual image configuration file; displaying the two-dimensional virtual image of the voice assistant on the display interface based on the two-dimensional image data; the second scenario is a scenario in which the CPU and / or memory of the vehicle terminal has an occupation rate greater than the preset threshold.

[0014] According to the technical means, the virtual image of the voice assistant can be displayed in different dimensions on the display interface according to different use scenarios of the virtual image of the voice assistant, thereby optimizing the occupation rate of the CPU and / or memory of the vehicle terminal, and avoiding the CPU and / or memory from being occupied too much to affect other operations of the user.

[0015] In a possible implementation, generating the two-dimensional image data of the virtual image of the voice assistant based on the virtual image configuration file includes: constructing a three-dimensional model of the virtual image based on the virtual image configuration file; and performing two-dimensional rendering on the three-dimensional model to obtain the two-dimensional image data.

[0016] According to the technical means, the two-dimensional image data is obtained by rendering the three-dimensional model, and the two-dimensional virtual image of the voice assistant can be displayed on the display interface when the CPU and / or memory of the vehicle terminal has an occupation rate greater than the preset threshold, thereby avoiding the CPU and / or memory from being occupied too much to affect other operations of the user, while ensuring that the virtual image of the voice assistant has a vivid and realistic visual effect.

[0017] In one possible implementation, the two-dimensional image data includes multiple consecutive frames of two-dimensional images; the above-mentioned displaying a two-dimensional virtual image of the voice assistant on the display interface based on the two-dimensional image data includes: displaying a two-dimensional virtual image of the voice assistant on the display interface based on multiple consecutive frames of two-dimensional images and a preset frame rate.

[0018] Based on the aforementioned technical means, this application can display the two-dimensional virtual image of the voice assistant in a second scenario, at a preset frame rate, in the form of continuous video frames. At this frame rate, the actions displayed by the two-dimensional virtual image are continuous. This solves the problem of excessive CPU and / or memory usage while ensuring that the virtual image of the voice assistant has a vivid visual effect.

[0019] According to a second aspect provided in this application, an image configuration device is provided, comprising: a acquisition module for acquiring an image of a target object within a vehicle terminal; an extraction module for extracting the image features of the target object from the image of the target object; and a generation module for generating a virtual image configuration file for a voice assistant based on the image features of the target object.

[0020] In one possible implementation, the image configuration device further includes a search module; the search module is used to search the image feature number of the virtual image of the voice assistant that matches the image features of the target object from the image database; the image database is used to store multiple image features of the virtual image of the voice assistant, and the number of each image feature of the virtual image; the generation module is specifically used to generate a virtual image configuration file of the voice assistant based on the number of the image feature of the virtual image of the voice assistant that matches the image features of the target object.

[0021] In one possible implementation, the image configuration device further includes an acquisition module; the acquisition module is used to acquire motion data of the virtual image in multiple scenarios; and the generation module is specifically used to generate a virtual image configuration file based on the image features of the target object and the motion data of the virtual image in multiple scenarios.

[0022] In one possible implementation, the avatar configuration device further includes an acquisition module; the acquisition module is further configured to acquire the modified avatar features in response to a user's modification of the avatar features; and the generation module is further configured to update the avatar configuration file based on the modified avatar features.

[0023] In a possible implementation, the generating module is further configured to, in a first scenario, generate, in response to the operation of the user waking up the voice assistant, three-dimensional image data of the virtual image of the voice assistant based on the virtual image profile; display, based on the three-dimensional image data, the three-dimensional virtual image of the voice assistant on a display interface of the vehicle terminal; the first scenario is a scenario in which an occupation rate of a central processing unit (CPU) and / or a memory of the vehicle terminal is less than or equal to a preset threshold; or, in a second scenario, generate, in response to the operation of the user waking up the voice assistant, two-dimensional image data of the virtual image of the voice assistant based on the virtual image profile; display, based on the two-dimensional image data, the two-dimensional virtual image of the voice assistant on the display interface; the second scenario is a scenario in which the occupation rate of the CPU and / or the memory of the vehicle terminal is greater than the preset threshold.

[0024] In a possible implementation, the generating module is specifically configured to construct a three-dimensional model of the virtual image based on the virtual image profile; and perform two-dimensional rendering on the three-dimensional model to obtain the two-dimensional image data.

[0025] In a possible implementation, the two-dimensional image data includes a plurality of frames of continuous two-dimensional images; and the generating module is specifically configured to display, based on the plurality of frames of continuous two-dimensional images and a preset frame rate, the two-dimensional virtual image of the voice assistant on the display interface.

[0026] According to a third aspect provided in the present application, a vehicle terminal is provided, which includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor executes the program to implement the method in the first aspect and any possible implementation thereof.

[0027] According to a fourth aspect provided in the present application, a computer-readable storage medium is provided, and when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method in the first aspect and any possible implementation thereof.

[0028] Therefore, the above technical features of the present application have the following beneficial effects:

[0029] (1) In constructing the virtual image of the voice assistant of the vehicle terminal, the image of the target object is collected, the image features of the target object are extracted from the image of the target object, and then the virtual image configuration file of the voice assistant is generated based on the image features of the target object. It can be seen that the image configuration method provided by the application can generate a configuration file according to the image features of the target object, and then generate a virtual image of the voice assistant according to the configuration file. In this way, the virtual image of the voice assistant configured has a virtual human feeling, is more lively and realistic, enriches the virtual image of the voice assistant, and solves the problem of single image construction, few types and lack of real virtual human feeling in related technologies when constructing the virtual image of the voice assistant of the vehicle terminal.

[0030] (2) The application can find the image features of the virtual image of the voice assistant matched with the image features of the target object from the image database; and generate a virtual image configuration file of the voice assistant according to the image features of the virtual image. Unlike the single image of the virtual image of the voice assistant in related technologies, which does not have the real virtual human feeling, the application matches the image features of the target object with the image features of the virtual image of the voice assistant stored in the image database, so that the generated virtual image of the voice assistant has a certain similarity with the target object, improves the realism of the virtual image of the voice assistant, and makes the virtual image of the voice assistant more realistic. At the same time, the image features of the virtual image of the voice assistant in the image database are expandable, and the types of the image features in the image database can be expanded and the number can be increased from time to time, so as to improve the similarity between the image features of the virtual image of the voice assistant and the image features of the target object, make the generated virtual image of the voice assistant more lively and realistic, and improve the accuracy of the virtual image configuration of the voice assistant.

[0031] (3) The application can set different action data for the virtual image of the voice assistant according to different use scenarios, so that the image of the voice assistant of the vehicle terminal is more lively and flexible, and the user's use experience is improved.

[0032] (4) Unlike the virtual image of the voice assistant in related technologies, which is the same or only has a few appearance types for users to choose from, the application can provide a configuration selection interface for the virtual image of the voice assistant for users. In addition to using the virtual image of the voice assistant matched by the vehicle terminal, users can also freely customize the virtual image of the voice assistant of the vehicle terminal according to their own preferences from the selection interface. At the same time, the operation of the application is simple and convenient, and users only need to perform a few simple operations to realize a series of complex operations such as image scanning, feature extraction and image generation, which reduces the learning cost of users, improves the participation of users, and brings good use experience to users.

[0033] (5) The application can configure different actions for the virtual character according to different use scenarios, make the image of the voice assistant of the vehicle terminal more vivid and flexible, and improve the user experience.

[0034] (6) The application can make the virtual image of the voice assistant display different dimensions of images on the display interface according to different use scenarios of the virtual image of the voice assistant, thereby optimizing the CPU and / or memory occupancy rate of the vehicle terminal, avoiding too high CPU and / or memory occupancy rate, and affecting other operations of the user.

[0035] (7) The application obtains two-dimensional image data by rendering a three-dimensional model, and can display a two-dimensional virtual image of the voice assistant on the display interface when the CPU and / or memory occupancy rate of the vehicle terminal is greater than a preset threshold, thereby avoiding too high CPU and / or memory occupancy rate, affecting other operations of the user, and ensuring that the virtual image of the voice assistant has vivid and realistic visual effects.

[0036] (8) The application can display a two-dimensional virtual image of the voice assistant in the form of continuous video frames at a preset frame rate in the second scenario. The action displayed by the two-dimensional virtual image is continuous at this frame rate. In this way, the problem of too high CPU and / or memory occupancy rate can be solved, and the virtual image of the voice assistant can have vivid visual effects.

[0037] It should be noted that the technical effects brought by any one of the implementation manners of the second aspect to the fourth aspect can be referred to the technical effects brought by the corresponding implementation manners in the first aspect, which will not be repeated here.

[0038] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the application. BRIEF DESCRIPTION OF DRAWINGS

[0039] The accompanying drawings incorporated in the specification and constituting a part of it illustrate embodiments consistent with the application and, together with the specification, serve to explain the principles of the application, and do not constitute an undue limitation on the application.

[0040] Figure 1 is a system structure diagram of an image configuration system according to an exemplary embodiment;

[0041] Figure 2 is a flowchart of an image configuration method according to an exemplary embodiment;

[0042] Figure 3 is a flowchart of another image configuration method according to an exemplary embodiment;

[0043] Figure 4 is a flow chart of another image configuration method according to an exemplary embodiment

[0044] Figure 5 is a structural diagram of an image configuration device according to an exemplary embodiment;

[0045] Figure 6 is a structural diagram of a vehicle terminal according to an exemplary embodiment.

[0046] 700-image configuration device, 701-acquisition module, 702-extraction module, 703-generation module, 704-search module, 705-obtaining module, 800-vehicle terminal, 801-processor, 802-memory. DETAILED DESCRIPTION

[0047] In order to make the ordinary person skilled in the art better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings.

[0048] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Rather, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0049] An embodiment of the present application is described below with reference to the accompanying drawings. There are many types of construction of the virtual image of the voice assistant of the vehicle terminal in the related art mentioned in the background art, such as two-dimensional icon type, three-dimensional icon type, or two-dimensional virtual person type, but the constructed image is relatively single, does not have a real virtual person feeling, and does not bring a visually dynamic and realistic feeling to the user. The embodiment of the present application provides an image configuration method, which can construct the virtual image of the voice assistant of the vehicle terminal by collecting the image of the target object, extracting the image features of the target object from the image of the target object, and then generating a virtual image configuration file of the voice assistant based on the image features of the target object. It can be seen that the image configuration method provided by the present application can generate a configuration file according to the image features of the target object, and then generate a virtual image of the voice assistant according to the configuration file. In this way, the virtual image of the voice assistant configured has a virtual person feeling and is more dynamic and realistic, which enriches the virtual image of the voice assistant and solves the problems of single image construction, few types, and no real virtual person feeling in the related art when constructing the virtual image of the voice assistant of the vehicle terminal.

[0050] For ease of understanding, the image configuration method provided by the present application is specifically introduced below in combination with the accompanying drawings.

[0051] Figure 1 An image configuration system according to an exemplary embodiment is shown, as shown in Figure 1 The image configuration system includes a screen 100, an in-vehicle camera 200, an image acquisition subsystem 300, an image replication subsystem 400, a three-dimensional engine 500, and an image manual selection subsystem 600.

[0052] The screen 100 is used to display the virtual image of the voice assistant.

[0053] The in-vehicle camera 200 is used to collect the image of the target object in the vehicle terminal.

[0054] The image acquisition subsystem 300 is used to extract the image features of the target object from the image of the target object.

[0055] The image replication subsystem 400 is used to generate a virtual image configuration file of the voice assistant based on the image features of the target object.

[0056] The image replication subsystem 400 is specifically configured to search for a number of image features of a virtual image of the voice assistant that match the image features of the target object from an image database; the image database is configured to store a plurality of image features of the virtual image of the voice assistant, and the number of each image feature of the virtual image; and generate a virtual image profile of the voice assistant based on the number of image features of the virtual image of the voice assistant that match the image features of the target object.

[0057] The image replication subsystem 400 is specifically configured to obtain action data of the virtual image in a plurality of scenes; and generate a virtual image profile based on the image features of the target object and the action data of the virtual image in the plurality of scenes.

[0058] The image replication subsystem 400 is further configured to generate a configuration file according to the collected image features in combination with the numbers of the image features in the image database, and mark the new name given by the user to the virtual image; and then import the marked configuration file into the voice assistant. When the user wakes up the voice assistant with the new name, the image presented on the screen 100 is the virtual image generated by the configuration.

[0059] The three-dimensional engine 500 is configured to display the virtual image of the voice assistant on the screen 100 based on the virtual image profile of the voice assistant.

[0060] As a possible implementation, the three-dimensional engine 500 is specifically configured to, in a first scenario, generate three-dimensional image data of the virtual image of the voice assistant based on the virtual image profile in response to an operation of the user waking up the voice assistant; and display the three-dimensional virtual image of the voice assistant on a display interface of the vehicle terminal based on the three-dimensional image data.

[0061] The first scenario is a scenario in which the CPU and / or memory of the vehicle terminal has an occupancy rate less than or equal to a preset threshold.

[0062] As another possible implementation, the three-dimensional engine 500 is specifically configured to, in a second scenario, generate two-dimensional image data of the virtual image of the voice assistant based on the virtual image profile in response to an operation of the user waking up the voice assistant; and display the two-dimensional virtual image of the voice assistant on the display interface based on the two-dimensional image data.

[0063] The second scenario is a scenario in which the CPU and / or memory of the vehicle terminal has an occupancy rate greater than a preset threshold.

[0064] Illustratively, the three-dimensional engine 500 is specifically configured to construct a three-dimensional model of the virtual image based on the virtual image profile; and perform two-dimensional rendering on the three-dimensional model to obtain two-dimensional image data.

[0065] The image manual selection subsystem 600 is configured to obtain modified image features of the virtual image in response to the user modifying the image features of the virtual image, and update the virtual image profile according to the modified image features of the virtual image.

[0066] The image manual selection subsystem 600 is further configured to mark the new name given by the user to the virtual image, and then import the marked profile into the voice assistant. When the user wakes up the voice assistant with the new name given, the three-dimensional engine 500 invokes the image modeling data based on the profile, and then loads and runs the image to be presented, which is the virtual image generated by the profile.

[0067] Figure 2 is a flowchart of an image configuration method according to an example embodiment, as shown in Figure 2 The image configuration method includes the following steps:

[0068] S101, collect an image of a target object in a vehicle terminal.

[0069] As a possible implementation, the above step S101 can be implemented as follows: after the user opens the image collection interface, the image of the target object is collected by the in-vehicle camera.

[0070] As another possible implementation, the above step S101 can be implemented as follows: the image of the target object is obtained from a video stream captured by the in-vehicle camera; wherein the image of the target object is any picture frame in the above video stream.

[0071] In some embodiments, when the image or picture frame collected by the in-vehicle camera includes multiple objects, the user can manually select any object in the image or picture frame as the target object to obtain the image of the target object.

[0072] In addition, in order to improve the accuracy of image collection of the target object, the shooting performance of the in-vehicle camera has certain requirements. For example, the camera should have high definition so as to accurately extract the face shape, hairstyle, whether wearing glasses and other information of the face; the shooting delay of the camera should be short, so that the target object does not need to deliberately maintain the modeling to ensure the shooting effect, and the collected image will not produce ghosting due to changes in action; the camera should have good adaptability to light, and should not be too sensitive to light changes to ensure that the target object can normally collect images in most scenarios.

[0073] S102, extract image features of the target object from the image of the target object.

[0074] The image features of the target object include multiple categories of image features, such as face shape, hair style, eye shape, and accessories.

[0075] In some embodiments, the image features of the target object are extracted from the image of the target object by inputting the image of the target object into an image recognition model to obtain the image features of the target object. The image recognition model is configured to extract the image features of the target object from the image of the target object. For example, the image features of the target object extracted by the image recognition model include round face, long hair, phoenix eyes, and ear studs.

[0076] As a possible implementation, the training process of the image recognition model can be implemented as follows: a plurality of face images are obtained; each face image is configured with a label according to the image features of the face image; for example, the hair style label, the face shape label, the label, the eye shape label, and the accessory label are configured for the face image; a training sample is constructed according to the plurality of face images and the labels corresponding to each face image, and the image recognition model is trained according to the training sample to obtain the trained image recognition model.

[0077] For example, the image recognition model can recognize the hair style of the target object, such as long hair, short hair, or ponytail. The image recognition model can extend the hair style according to the design needs, such as adding a sheep horn braid and a ponytail.

[0078] In some embodiments, the image recognition model can recognize the hair style of the target object from different image backgrounds. It can be understood that the recognition of the hair style is relatively difficult, because the image background of the target object is an important factor. If the image background and the hair color are quite different, the hair style of the target object can be easily recognized. However, if the image background and the hair color are similar, the hair style of the target object cannot be accurately recognized. Therefore, the image recognition model is trained by using images of target objects with different image backgrounds, which can improve the accuracy of the image recognition model in recognizing the hair style.

[0079] For example, the image recognition model can recognize the face shape of the target object, such as bun face, melon seed face, and long face.

[0080] For example, if the target object is wearing glasses in the image, the image recognition model will recognize the image feature of wearing glasses when recognizing. The image recognition model can also extend the details of the glasses according to the design needs, such as whether the glasses are square or round.

[0081] For example, the image recognition model can identify that the target object has a high nose bridge, a low nose bridge, a large nose tip, a small nose tip, large eyes, small eyes, close-set eyes, or wide-set eyes, and so on.

[0082] It can be understood that in the process of collecting the image of the target object, it will be affected by many environmental factors, such as light and shade, background, and the like. Therefore, when training the image recognition model, images of the target object in different environments need to be collected for calibration to train the image recognition model, so as to improve the recognition accuracy of the image recognition model.

[0083] It can be understood that according to the design needs and actual needs of users, the types and details that can be recognized by the image recognition model can be continuously expanded. Each added type can enable the image simulation model to simulate as many use scenarios as possible, and the image recognition model can be continuously adapted to improve the accuracy of the image recognition model in recognizing the image features of the target object.

[0084] It should be noted that the above image features and the image recognition model are only examples and do not constitute a specific limitation on the image configuration method provided by the embodiments of the present application. In other embodiments, the image features can also include other types, and the image recognition model can also recognize other types of image features. It can be understood that the types of image features that can be recognized by the image recognition model can be adapted according to different use scenarios and are not limited to the above-mentioned several types. For example, the types of image features that can be recognized by the image recognition model can be flexibly selected according to system settings and user needs, and the embodiments of the present application do not limit this.

[0085] S103, generating a virtual image configuration file of the voice assistant based on the image features of the target object.

[0086] In some embodiments, the above step S103 can be implemented as: searching, from an image database, a number of image features of a virtual image of the voice assistant that match the image features of the target object; the image database is configured to store a plurality of image features of the virtual image of the voice assistant, and a number of each image feature of the virtual image; and generating the virtual image configuration file of the voice assistant based on the number of the image features of the virtual image of the voice assistant that match the image features of the target object.

[0087] As a possible implementation, the image features of the target object can include face shape, hairstyle, and accessories.

[0088] In some embodiments, the image database includes a plurality of image feature sets of the virtual image of the voice assistant. Each image feature set includes a plurality of different types of image features of the virtual image.

[0089] For example, it is assumed that the image database includes a first image feature set. The first image feature set is a set of different types of image features corresponding to a first image feature of a virtual image. For example, the first image feature set can be a set of face shapes of a virtual image, and the first image feature set includes different types of face shapes of a virtual image, such as a melon seed face image feature, a long face image feature, and a round face image feature, etc.

[0090] In some embodiments, when the image database stores the image features of the virtual image of the voice assistant, each image feature corresponds to a different number for identifying the image features of the virtual image of the voice assistant. For example, the image database includes an image feature set for storing the image features and numbers of the virtual image of the voice assistant. For example, the image feature set A includes a first image feature set A1, a second image feature set A2, and a third image feature set A3. The first image feature set A1 is a set of different types of image features corresponding to a first image feature (face shape); for example, the first image feature set A1 includes a melon seed face A1-1, a round face A1-2, and other image features. The second image feature set A2 is a set of different types of image features corresponding to a second image feature (hair style); for example, the second image feature set A2 includes long hair A2-1, short hair A2-2, and other image features. The third image feature set A3 is a set of different types of image features corresponding to a third image feature (accessory); for example, the third image feature set includes glasses A3-1, earrings A3-2, and other image features.

[0091] For example, if the image recognition model identifies that the image features of the target object are a melon seed face and short hair, the numbers of the image features of the virtual image of the voice assistant that match the image features of the target object can be determined as A1-1 and A2-2 according to the numbers of the image features of the virtual image of the voice assistant corresponding to the image features.

[0092] As a possible implementation, after the image replication subsystem finds the numbers of the image features of the virtual image of the voice assistant that match the image features of the target object, a configuration file of the virtual image of the voice assistant can be generated. The configuration file includes the image features and numbers of the virtual image of the voice assistant that match the image features of the target object. For example, the configuration file includes [round face A1-2, long hair A2-1, glasses A3-1].

[0093] In some embodiments, the image database is also used to store action data and numbers of the virtual image of the voice assistant; and the step S103 can be implemented by obtaining action data of the virtual image in multiple scenes; and generating a virtual image configuration file based on the image features of the target object and the action data of the virtual image in multiple scenes.

[0094] As a possible implementation, in order to make the virtual image of the voice assistant visually look lively and flexible, the manufacturer designs multiple sets of action poses for the virtual image of the voice assistant in advance according to different use scenarios when using the voice assistant, each set of action poses corresponding action data is stored in the storage, when the user uses the voice assistant in different scenarios, the virtual image of the voice assistant can show different action poses. Among them, the action pose includes the mouth action when speaking, the hand action, the leg action, the body part action, the head action, the turning action, etc.

[0095] In some embodiments, the action data stored in the image database includes a first action set; the first action set is a set of different types of action data corresponding to a first part of the virtual image. For example, the first part of the virtual image is any one of the multiple parts of the virtual image. For example, assuming that the multiple parts of the virtual image include: face, hand and mouth, etc.; the first part of the virtual image can be the mouth, and the first action set can be a set of different types of action data corresponding to the mouth, for example, the first action set can include: speaking action data and smiling action data, etc.

[0096] For example, assuming that the action pose of the virtual image of the mouth is designed in advance, the action data of the virtual image of the voice assistant is determined from the first action set. Assuming that the action pose designed in advance is smiling, the action data configured by the virtual image of the voice assistant is the action data of smiling.

[0097] In some embodiments, the action data stored in the image database all correspond to different numbers. For example, assuming that the image database includes an action pose set for storing action data. For example, the action pose set B includes: the first action set B1 and the second action set B2. Among them, the first action set B1 is a set of different types of action data corresponding to a first part (mouth) of the virtual image; for example, the first action set B1 includes: speaking B1-1, smiling B1-2, etc. The second action set B2 is a set of different types of action data corresponding to a second part (hand) of the virtual image; for example, the second action set B2 includes: waving B2-1, waving B2-2, etc.

[0098] As a possible implementation, the image replication subsystem generates a virtual image configuration file based on the image features of the target object and the action data of the virtual image in multiple scenarios. Among them, the virtual image configuration file includes: action data and corresponding number. For example, the virtual image configuration file includes: [speaking B1-1, waving B2-1].

[0099] In addition, these action models are not adjusted according to the image configuration, but are matched with the use scene, so that the voice assistant image looks more lifelike. For example, when the voice assistant is broadcasting, the processor calls the modeling data of the mouth action; when the voice assistant is navigating, the processor calls the modeling data of the hand action in addition to the modeling data of the mouth action and displays it on the virtual image of the voice assistant.

[0100] It can be understood that each image feature and action data corresponds to a different number and is stored in the image database; each image feature or action data is numbered layer by layer for classification, which facilitates subsequent retrieval of image features or action data by the image database; at the same time, when the type and content of the image feature or action data are expanded, the above-mentioned numbering rules can also be continued for numbering, which facilitates the expansion of the image features or action data stored in the image database.

[0101] In some embodiments, the virtual image of the voice assistant can also be manually configured by the user. For example, the following steps can be implemented:

[0102] Step a1, in response to the user's modification of the image features of the virtual image, the modified image features of the virtual image are obtained.

[0103] As a possible implementation, the user can manually select the image features such as accessories of the voice assistant that he or she likes in the image collection interface, and the virtual image of the voice assistant will be updated to the image designed by the user.

[0104] Step a2, updating the virtual image configuration file according to the modified image features of the virtual image.

[0105] It can be understood that the virtual image configuration file is updated according to the image features of the virtual image manually selected by the user, and then the three-dimensional engine displays the virtual image of the voice assistant modified by the user through the car screen. For example, the user customizes and selects a melon seed face, long hair, and glasses, and the virtual image of the voice assistant displayed on the car screen will display the corresponding image features.

[0106] In some embodiments, as Figure 3 described above, after step S103, the above method further includes the following steps:

[0107] S104, in response to the user's operation of waking up the voice assistant, displaying the virtual image of the voice assistant on the display interface of the vehicle terminal.

[0108] For example, the operation of the user waking up the voice assistant includes a voice operation or a touch screen operation, etc. For example, the user can call the name of the virtual image of the voice assistant to wake up the voice assistant. For another example, the user can click a button for waking up the voice assistant on a car screen to wake up the voice assistant.

[0109] As a possible implementation, in the first scenario, in response to the operation of the user waking up the voice assistant, three-dimensional image data of the virtual image of the voice assistant is generated based on the virtual image profile; and the three-dimensional virtual image of the voice assistant is displayed on the display interface of the vehicle terminal based on the three-dimensional image data.

[0110] For example, the first scenario can be a navigation scenario. The preset threshold can be 80%.

[0111] As another possible implementation, in response to the operation of the user waking up the voice assistant, two-dimensional image data of the virtual image of the voice assistant is generated based on the virtual image profile; and the two-dimensional virtual image of the voice assistant is displayed on the display interface based on the two-dimensional image data.

[0112] For example, the second scenario can be a voice broadcast scenario.

[0113] For example, when the use scenario of the voice assistant is the voice broadcast scenario, the three-dimensional image can be displayed to make the virtual image of the voice assistant more lifelike and have a real virtual human feeling. When the use scenario of the voice assistant is the navigation scenario, the two-dimensional image can be displayed to avoid high CPU and / or memory occupancy.

[0114] In some embodiments, the two-dimensional image data includes a plurality of frames of continuous two-dimensional images. Then, the displaying the two-dimensional virtual image of the voice assistant on the display interface based on the two-dimensional image data includes: displaying the two-dimensional virtual image of the voice assistant on the display interface based on the plurality of frames of continuous two-dimensional images and a preset frame rate. For example, the preset frame rate can be 20 frames per second. It can be understood that, in the second scenario, the two-dimensional virtual image of the voice assistant can be displayed in the form of continuous video frames at the preset frame rate. The motion displayed by the two-dimensional virtual image at this frame rate is continuous. In this way, the problem of high CPU and / or memory occupancy can be solved, and the virtual image of the voice assistant can have a lively visual effect.

[0115] In some embodiments, based on the virtual image profile, generating the two-dimensional image data of the virtual image of the voice assistant includes: based on the virtual image profile, constructing a three-dimensional model of the virtual image; and performing two-dimensional rendering on the three-dimensional model to obtain the two-dimensional image data.

[0116] It can be understood that, unlike mobile phone 3D games, after opening the mobile phone 3D game software, the user will not open other video software, and there is almost no situation of simultaneous use of multiple software; and the voice assistant of the vehicle terminal and other functions of the vehicle terminal often have a situation of long-time use in the same time period, such as awakening the voice assistant in the video playing interface, and how to solve the long-term high load occupation of CPU and / or memory on the vehicle terminal is a difficult problem. The present application can solve this problem by using a three-to-two method (for example, constructing a three-dimensional model and then performing two-dimensional rendering on the three-dimensional model, which is provided in the embodiments of the present application), and in a scene with high CPU and / or memory occupation rate, the virtual image of the voice assistant is switched from a three-dimensional virtual image to a two-dimensional virtual image; at the same time, in order to ensure that the virtual image of the voice assistant is lifelike, the two-dimensional virtual image is provided with consecutive actions in a continuous frame manner, and the frame rate is 20 frames per second, and the action seen by the naked eye under this frame rate is almost continuous. This scheme can not only solve the problem of hardware resources, but also ensure the lifelike effect of the virtual image of the voice assistant.

[0117] It can be understood that the more the dimensions of the virtual image are, the more lifelike the virtual image will be; and the higher the resolution of each dimension is, the higher the definition of the virtual image will be. At the same time, in the case that the use scene of the voice assistant is high energy consumption, configuring the virtual image of the voice assistant as a two-dimensional image can optimize the CPU and memory occupation rate of the processor, and avoid that the high CPU and / or memory occupation rate brings bad user experience.

[0118] In order to facilitate understanding, the following will exemplarily illustrate an image configuration method provided by the embodiments of the present application.

[0119] Exemplarily, the user needs to configure the virtual image of the voice assistant in the vehicle terminal, and the vehicle terminal has already installed the voice assistant image configuration software used to configure the virtual image of the voice assistant, such as Figure 4 As shown in the figure, the image configuration process can be implemented as the following steps:

[0120] Step a1, the user clicks the voice assistant image configuration software on the screen.

[0121] Step a2, the voice assistant image configuration software judges whether the voice assistant has a configured virtual image.

[0122] Step a3, if there is a configured virtual image, the car machine screen directly jumps to the interface, displays the configured virtual image, and pops up a virtual image modification button to ask and guide the user to select virtual image modification; if there is no configured virtual image, step a5 is directly executed.

[0123] Step a4, the user clicks the virtual image modification button to modify the virtual image to be configured.

[0124] Step a5, the car machine screen pops up a gender selection interface to guide the user to select male or female.

[0125] Step a6, after completing the gender selection, the interface jumps to the image displayed by the in-vehicle camera.

[0126] Step a7, the image collection subsystem recognizes and searches the face, and when the face is recognized, it is framed out with a picture frame to prompt the user to select the face to be copied.

[0127] Step a8, after the user selects, the image collection subsystem extracts the features of the selected face, including hairstyle, face shape, glasses, decorative accessories, etc.; then forms the image features and sends them to the image copying subsystem.

[0128] Step a9, the image copying subsystem will create a configuration file according to the received image features, and import the configuration file into the voice assistant.

[0129] Step a10, the three-dimensional engine loads and displays the configured virtual image of the voice assistant according to the configuration file.

[0130] It can be understood that compared with the mobile phone three-dimensional image customization game software, the three-dimensional image configuration of the mobile phone three-dimensional image customization game software is mainly a standalone entertainment software, which mainly displays modeling effects and action effects; and the image configuration method provided in the embodiment has all the functions of the voice assistant after the virtual image of the voice assistant is configured, and can recognize the user's instructions and perform corresponding operations.

[0131] The above describes the solutions provided by the embodiments of the present application from the method aspect. In order to implement the above functions, the image configuration apparatus or the electronic device comprises hardware structures and / or software modules corresponding to the respective functions. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0132] The embodiments of the present application can divide the functional modules of the image configuration apparatus or the electronic device according to the above method. For example, the image configuration apparatus or the electronic device can comprise functional modules corresponding to the respective functions, or two or more functions can be integrated into one processing module. The integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical functional division. Actual implementation can have another division manner.

[0133] Figure 5 is a block diagram of an image configuration apparatus according to an example embodiment. Referring to Figure 5 The image configuration apparatus 700 comprises a collection module 701, an extraction module 702, and a generation module 703.

[0134] The collection module 701 is configured to collect an image of a target object in a vehicle terminal.

[0135] The extraction module 702 is configured to extract an image feature of the target object from the image of the target object.

[0136] The generation module 703 is configured to generate a virtual image configuration file of a voice assistant based on the image feature of the target object. In some embodiments,

[0137] In some embodiments, the image configuration apparatus 700 further comprises a search module 704. The search module 704 is configured to search, from an image database, a number of an image feature of a virtual image of the voice assistant that matches the image feature of the target object. The image database is configured to store a plurality of image features of the virtual image of the voice assistant and the number of each image feature of the virtual image. The generation module 703 is specifically configured to generate the virtual image configuration file of the voice assistant based on the number of the image feature of the virtual image of the voice assistant that matches the image feature of the target object.

[0138] In some embodiments, the image configuration device 700 further includes an acquisition module 705; the acquisition module 705 is used to acquire action data of the virtual image in multiple scenarios; the generation module 703 is specifically used to generate a virtual image configuration file based on the image features of the target object and the action data of the virtual image in multiple scenarios.

[0139] In some embodiments, the avatar configuration device 700 further includes an acquisition module 705; the acquisition module 705 is further configured to acquire the modified avatar features of the virtual avatar in response to the user's modification of the avatar features; the generation module 703 is further configured to update the virtual avatar configuration file according to the modified avatar features of the virtual avatar.

[0140] In some embodiments, the generation module 703 is further configured to, in a first scenario, in response to a user's operation to wake up the voice assistant, generate three-dimensional image data of the virtual image of the voice assistant based on the virtual image configuration file; and display the three-dimensional virtual image of the voice assistant on the display interface of the vehicle terminal based on the three-dimensional image data; the first scenario is a scenario where the CPU and / or memory occupancy rate of the vehicle terminal is less than or equal to a preset threshold; or, in a second scenario, in response to a user's operation to wake up the voice assistant, generate two-dimensional image data of the virtual image of the voice assistant based on the virtual image configuration file; and display the two-dimensional virtual image of the voice assistant on the display interface based on the two-dimensional image data; the second scenario is a scenario where the CPU and / or memory occupancy rate of the vehicle terminal is greater than a preset threshold.

[0141] In some embodiments, the generation module 703 is specifically used to construct a three-dimensional model of the virtual avatar based on the virtual avatar configuration file; and to perform two-dimensional rendering on the three-dimensional model to obtain two-dimensional image data.

[0142] In some embodiments, the two-dimensional image data includes multiple consecutive two-dimensional images; the generation module 703 is specifically used to display a two-dimensional virtual image of the voice assistant on the display interface based on the multiple consecutive two-dimensional images and a preset frame rate.

[0143] According to the technical means, the image configuration method provided in the application can generate a configuration file according to the image features of the target object, and then generate a virtual image of the voice assistant according to the configuration file. In this way, the virtual image of the voice assistant configured can have a virtual human feeling, be more lively and realistic, and enrich the virtual image of the voice assistant, thereby solving the problems of single configuration image, few types and lack of real virtual human feeling when a virtual image of a voice assistant of a vehicle terminal is constructed in the related art.

[0144] As to the device in the above-mentioned embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be described in detail here.

[0145] Figure 6 is a block diagram of a vehicle terminal according to an exemplary embodiment. As shown in Figure 6 , the vehicle terminal 800 includes but is not limited to a processor 801 and a memory 802.

[0146] The memory 802 described above is used to store executable instructions of the processor 801. It can be understood that the processor 801 is configured to execute the instructions to implement the image configuration method in the above-mentioned embodiments.

[0147] It should be noted that those skilled in the art can understand Figure 6 that the structure of the vehicle terminal 800 shown in the above-mentioned embodiments does not constitute a limitation on the vehicle terminal 800, and the vehicle terminal 800 can include more or fewer components than those shown in the above-mentioned embodiments, or combine certain components, or different component arrangements. Figure 6

[0148] The processor 801 is the control center of the vehicle terminal 800, and connects various parts of the vehicle terminal 800 through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 802, and calling data stored in the memory 802, the processor 801 performs various functions and processes data of the vehicle terminal 800, thereby overall monitoring the vehicle terminal 800. The processor 801 can include one or more processing units. Optionally, the processor 801 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and application programs, and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 801.

[0149] ​The memory 802 can be used to store software programs and various data. The memory 802 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs (such as a determination unit, a processing unit, etc.) required by at least one function module, and the like. In addition, the memory 802 can include a high-speed random access memory, and can also include a non-volatile memory, for example, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device.

[0150] In the example embodiment, a computer readable storage medium including instructions, for example, the memory 802 including instructions, is also provided, and the instructions can be executed by the processor 801 of the vehicle terminal 800 to implement the image configuration method in the above embodiment.

[0151] In actual implementation, Figure 5 The functions of the collection module 701, the extraction module 702, the generation module 703, the lookup module 704, and the acquisition module 705 in the above embodiment can be implemented by the processor 801 calling the computer program stored in the memory 802. Figure 6 The specific execution process can refer to the description of the image configuration method in the above embodiment, and will not be described here.

[0152] Alternatively, the computer readable storage medium can be a non-transitory computer readable storage medium, for example, a Read-Only Memory (ROM), a Random Access Memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0153] In the example embodiment, the embodiment of the present application also provides a computer program product including one or more instructions, which can be executed by the processor 801 of the vehicle terminal 800 to complete the image configuration method in the above embodiment.

[0154] It should be noted that the instructions in the above computer readable storage medium or the one or more instructions in the computer program product are executed by the processor of the vehicle terminal 800 to implement each process of the above image configuration method embodiment, and can achieve the same technical effect as the above image configuration method. To avoid repetition, it will not be described here.

[0155] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete the above described full classification part or part of the function.

[0156] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the division of the apparatus embodiments is merely an example, and for example, the division of the modules or units can be different, and for example, multiple modules or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0157] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, i.e., may be located in one place, or may be distributed in multiple different places. Some or all of the classification units can be selected according to actual needs to achieve the purpose of the embodiments.

[0158] In addition, each functional unit in the various embodiments of the present application can be integrated into a processing unit, or each unit can be physically present separately, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0159] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application essentially or the part that contributes to the prior art or the whole classification or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions to make a device (which can be a single chip, chip, etc.) or processor (processor) execute all or part of the steps of the various embodiments of the method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, and various media that can store program codes.

[0160] The above is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any changes or replacements within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for configuring an image, characterized in that, Applied to a vehicle terminal, the method includes: Acquire images of the target object within the vehicle terminal; Extract the image features of the target object from its image; Acquire action data of the virtual avatar in multiple scenarios; the multiple scenarios include: navigation scenario and voice broadcast scenario; The image database is used to search for the image feature number of the virtual avatar of the voice assistant that matches the image features of the target object, and the action data number that matches the virtual avatar of the voice assistant; the image database is used to store multiple image features of the virtual avatar of the voice assistant, the number of each image feature of the virtual avatar, multiple action data of the virtual avatar of the voice assistant, and the number of each action data. Based on the image characteristics of the target object, the number of the image characteristics that match the image characteristics of the target object, the action data of the virtual image of the target object in multiple scenarios, and the number that matches the action data, the virtual image configuration file is generated; In response to a user waking up the voice assistant, image data of the virtual avatar of the voice assistant is generated based on the virtual avatar configuration file; the image data includes: three-dimensional image data and two-dimensional image data; When the CPU and / or memory utilization is less than or equal to a preset threshold, a three-dimensional virtual image of the voice assistant is displayed on the display interface of the vehicle terminal based on the three-dimensional image data; or, If the CPU and / or memory usage exceeds the preset threshold, a two-dimensional virtual image of the voice assistant is displayed on the display interface based on the two-dimensional image data.

2. The method according to claim 1, characterized in that, The method further includes: In response to a user's modification of the virtual avatar's visual characteristics, the modified visual characteristics of the virtual avatar are obtained; Update the virtual avatar configuration file based on the modified avatar characteristics.

3. The method according to claim 1, characterized in that, Based on the virtual avatar configuration file, two-dimensional image data of the voice assistant's virtual avatar is generated, including: Based on the virtual avatar configuration file, a three-dimensional model of the virtual avatar is constructed; The three-dimensional model is rendered in two dimensions to obtain the two-dimensional image data.

4. The method according to claim 3, characterized in that, The two-dimensional image data includes multiple consecutive frames of two-dimensional images; the step of displaying the two-dimensional virtual image of the voice assistant on the display interface based on the two-dimensional image data includes: Based on the multiple consecutive two-dimensional images and the preset frame rate, a two-dimensional virtual image of the voice assistant is displayed on the display interface.

5. A visual configuration device, characterized in that, include: The acquisition module is used to acquire images of target objects within the vehicle terminal. The extraction module is used to extract the image features of the target object from the image of the target object; The acquisition module is used to acquire the action data of the virtual avatar in multiple scenarios; The multiple scenarios include: navigation scenarios and voice broadcasting scenarios; The search module is used to search the image database for the image feature number of the virtual image of the voice assistant that matches the image features of the target object, and the action data number that matches the virtual image of the voice assistant; the image database is used to store multiple image features of the virtual image of the voice assistant, the number of each image feature of the virtual image, multiple action data of the virtual image of the voice assistant, and the number of each action data. The generation module is used to generate the virtual image configuration file based on the image features of the target object, the number of the image features that match the image features of the target object, the action data of the virtual image of the target object in multiple scenarios, and the number that matches the action data. The generation module is further configured to, in response to a user's operation of waking up the voice assistant, generate image data of the virtual avatar of the voice assistant based on the virtual avatar configuration file; the image data includes: three-dimensional image data and two-dimensional image data; when the CPU and / or memory occupancy rate is less than or equal to a preset threshold, the three-dimensional virtual avatar of the voice assistant is displayed on the display interface of the vehicle terminal based on the three-dimensional image data; or, when the CPU and / or memory occupancy rate is greater than the preset threshold, the two-dimensional virtual avatar of the voice assistant is displayed on the display interface based on the two-dimensional image data.

6. The apparatus according to claim 5, characterized in that, The image configuration device also includes an acquisition module; The acquisition module is also used to acquire the modified image features of the virtual image in response to the user's modification of the image features of the virtual image. The generation module is also used to update the virtual avatar configuration file based on the modified avatar features.

7. A vehicle terminal, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the image configuration method as described in any one of claims 1 to 4.

8. A computer-readable storage medium, characterized in that, When the computer-executable instructions stored in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is capable of performing the image configuration method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method and device for controlling multimedia equipment to work

    CN108881706A

  • Vehicle-mounted assistant virtual image self-defining method

    CN114327705A

  • Human-computer interaction method and electronic equipment

    CN115205917A