Method and device for generating vehicle-mounted virtual image, vehicle, storage medium and computer program product

By generating personalized in-vehicle virtual images through frame splitting and restoration technology, the problems of fixed virtual images and poor interactivity in existing technologies are solved, and the user experience and dynamic continuity of the image are improved.

CN120769129APending Publication Date: 2025-10-10FAW VOLKSWAGEN AUTOMOTIVE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510890485.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In the existing technology, the in-vehicle virtual images are fixed or limited, the user interaction experience is poor, and there is a lack of personalization.

Method used

By splitting the original comic video into frames and replacing it with the target image frame by frame, character contour control, face repair and key point detection are performed, and a dynamic repair image is synthesized to finally generate a personalized in-vehicle virtual image video.

Benefits of technology

It realizes personalized customized in-vehicle virtual image, improves the interactive experience between users and vehicle computers, improves the continuity and stability of movements, and enhances user intimacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120769129A_ABST
    Figure CN120769129A_ABST
Patent Text Reader

Abstract

The invention provides a method and equipment for generating a vehicle-mounted virtual image, a vehicle, a storage medium and a computer program product. The method comprises the following steps: carrying out frame splitting processing on an original cartoon video to generate a frame-by-frame original image sequence; according to the image of the target image, replacing the character image in each image in the frame-by-frame original image sequence with the target image, and generating a frame-by-frame face replacement image; executing figure contour control on each image in the frame-by-frame face changing images to generate frame-by-frame static restoration images; carrying out character key point detection and posture recognition on the frame-by-frame static restoration images, calibrating key point positions and postures of characters, executing multi-frame action coherence restoration according to the key point positions and postures, and generating frame-by-frame dynamic restoration images; and performing video synthesis on the frame-by-frame dynamic restoration images to generate a vehicle-mounted virtual image video. The personalized customization of the vehicle-mounted virtual image is realized, the generated virtual image is friendly, natural and smooth, and the user experience is greatly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention generally relate to the field of smart cockpit technology, and more specifically, to a method, device, vehicle, storage medium, and computer program product for generating an in-vehicle virtual image. Background Art

[0002] With the development of smart cockpit technology, the cockpit has become the third mobile space. Drivers and passengers rely on its convenient functions, making in-vehicle entertainment and interactive features increasingly popular with users. The smart cockpit has become a key factor in improving the user experience. In human-computer interaction, the virtual image of the vehicle computer allows users to truly feel the presence of the interactive object, which is a core factor in the smart cockpit.

[0003] However, in previous human-computer interaction scenarios, the in-vehicle virtual images of smart cockpits are often relatively fixed or have only a few preset images, and the virtual images have little connection with users, resulting in a poor interactive experience. Summary of the Invention

[0004] In order to solve the above-mentioned problems in the prior art, in a first aspect, an embodiment of the present invention provides a method for generating a vehicle-mounted virtual image, the method comprising: de-framing an original comic video to generate a frame-by-frame original image sequence; replacing the character image in each image in the frame-by-frame original image sequence with the target image according to the image of the target image to generate a frame-by-frame face-changing image; performing character contour control on each image in the frame-by-frame face-changing image to generate a frame-by-frame static repair image; performing character key point detection and posture recognition on the frame-by-frame static repair image to calibrate the character's key point position and posture, performing multi-frame action continuity repair according to the key point position and the posture to generate a frame-by-frame dynamic repair image; performing video synthesis on the frame-by-frame dynamic repair image to generate a vehicle-mounted virtual image video.

[0005] In some embodiments, according to the image of the target image, the character image in each image in the frame-by-frame original image sequence is replaced with the target image, and the frame-by-frame face-changing image is generated, which includes: performing face replacement, face detection and face repair on each image in the frame-by-frame original image sequence according to the target image, thereby generating the frame-by-frame face-changing image.

[0006] In some embodiments, the method further includes: performing background removal on the frame-by-frame dynamically restored images to generate frame-by-frame background-free images; and performing video synthesis on the frame-by-frame background-free images to generate a background-free vehicle-mounted virtual image video.

[0007] In some embodiments, performing character contour control on each image in the frame-by-frame face-changing images to generate frame-by-frame static repair images further includes: performing face redrawing and repair on the image after performing character contour control to generate the frame-by-frame static repair image.

[0008] In some embodiments, based on the image of the target image, the character image in each image in the frame-by-frame original image sequence is replaced with the target image to generate a frame-by-frame face-changing image, which includes: obtaining a target source image; identifying and numbering the character's face in the target source image to generate an alternative image with a serial number; receiving a serial number input by a user; and determining the target image from the alternative images with a serial number based on the serial number input by the user.

[0009] In some embodiments, the key points of the character include facial contour points and / or facial structure points.

[0010] In a second aspect, an embodiment of the present invention proposes a device for generating a vehicle-mounted virtual image, the device comprising a memory and a processor, the memory storing a computer program, and implementing the method for generating a vehicle-mounted virtual image described in any of the above embodiments when the computer program is executed by the processor.

[0011] In a third aspect, an embodiment of the present invention provides a vehicle, comprising the device for generating an in-vehicle virtual image as described in the above embodiment.

[0012] In a fourth aspect, an embodiment of the present invention provides a storage medium storing computer-readable instructions, which, when executed by a processor, executes the method for generating a vehicle-mounted virtual image described in any of the above embodiments.

[0013] In a fifth aspect, an embodiment of the present invention provides a computer program product comprising computer-readable instructions, which, when executed by a processor, executes the steps of the method for generating a vehicle-mounted virtual image described in any of the above embodiments.

[0014] The main problems currently faced by image cloning in the smart cockpit domain include: ① being rigid or having only a few preset images; ② having little user interaction and a poor user experience. To address these issues, the present invention proposes a method, device, vehicle, storage medium, and computer program product for generating an in-vehicle virtual image. This method implements an image cloning solution that can cartoonize characters while preserving their movements by decomposing the original animation video, performing frame-by-frame character face replacement, controlling the character's outline and redrawing the face, detecting key points, and restoring the continuity of multiple frames of motion. Finally, multiple-frame image synthesis creates a dynamic cloned image. This allows for personalized customization of in-vehicle virtual images, significantly improving the user experience of image cloning and interaction.

[0015] Based on the implementation manner of the present invention, the user's own elements can be integrated into the image setting, and supplemented by coherent movements, the image cloning is natural and smooth, which will make the user feel more intimate, strengthen the connection between the user and the car computer, and experience the feeling of unity between man and car.

[0016] The image cloning algorithm introduces multiple stability-enhancing links, which not only ensures the stability of facial features after face-changing, but also makes the movement connection of the deframed image more consistent and coordinated, ensuring the quality of the dynamic cloned image. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The above and other objects, features and advantages of the embodiments of the present invention will become readily understood by reading the following detailed description with reference to the accompanying drawings, in which several embodiments of the present invention are shown by way of example and not limitation, in which:

[0018] Figure 1 A flowchart of a method for generating a vehicle-mounted virtual image according to an embodiment of the present invention is shown;

[0019] Figure 2 A schematic diagram illustrating character outline control according to an embodiment of the present invention is shown;

[0020] Figure 3 A schematic diagram of detecting key points of a person according to an embodiment of the present invention is shown;

[0021] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION

[0022] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way.

[0023] In one aspect, an embodiment of the present invention provides a method for generating a vehicle-mounted virtual image. Figure 1 , which shows a flow chart of a method for generating a vehicle-mounted virtual image according to an embodiment of the present invention. The method includes steps S101-S105.

[0024] In step S101, the original comic video is deframed to generate a frame-by-frame original image sequence. The video deframe tool is used to deframe the video into a frame-by-frame image sequence to prepare for the subsequent algorithm processing.

[0025] In step S102, based on the target image, the character image in each image of the original image sequence is replaced with the target image to generate a frame-by-frame face-changing image. In this step, the character image face-changing is performed frame by frame.

[0026] As one embodiment of the present invention, step S102 may include performing face replacement, face detection, and face restoration on each frame of the original image sequence based on the target image, thereby generating frame-by-frame face-swapped images. In this embodiment, the final face-swapped image is achieved through the collaboration of three modules: a face replacement module, a face detection module, and a face restoration module.

[0027] As an example only, this step can be implemented using the Reactor face-changing algorithm module.

[0028] For example, the face replacement model might use ins wapper_128, the face detection model might use retinaface_resnet50, and the face restoration model might use GFPGAN v1.4. By combining these three models, the final image face swap can be completed. After setting the model parameters and the corresponding input and output, the image face swap can be completed. This face swapping involves swapping each frame of the video one by one, ultimately completing the entire video frame sequence.

[0029] As an embodiment of the present invention, the target image can be determined in step S102 by the following steps: obtaining a target source image; identifying and numbering the faces of the characters in the target source image to generate alternative images with serial numbers; receiving the serial number input by the user; and determining the target image from the alternative images with serial numbers based on the serial number input by the user.

[0030] As an example, take a source image and select the face number in the source image to be swapped, which is 0, 1, 2, 3, etc. from left to right. Then, for the target image to be swapped, set the target face number in the selected image. If both the source and target images have only one task, use the default number of 0.

[0031] In step S103, character contour control is performed on each image in the face-changing image frame by frame to generate a static restoration image frame by frame.

[0032] refer to Figure 2 , which shows a schematic diagram of character contour control according to an embodiment of the present invention.

[0033] As an embodiment of the present invention, after performing the character contour control, face redrawing and restoration may be performed on the image after performing the character contour control to generate a frame-by-frame static restoration image.

[0034] As an example, the ControlNet algorithm architecture can be used to draw the outline of a person in a sequence of images after video frames are deframed. By controlling the outline and combining it with the impact face refinement and redrawing algorithm component, facial details can be optimized to make the person image in each frame more vivid and clear.

[0035] ControlNet is a deep learning architecture that extends existing pre-trained models (such as StableDiffusion) to allow users to exert more precise control over the generation process. By adding additional conditional input layers, ControlNet can adjust the output results according to the user's specific needs, such as converting line drawings into photorealistic images, pose-guided character image generation, and key point detection of characters.

[0036] Impact is a facial refinement algorithm component. Its main inputs include the image to be processed, the algorithm model, the text constraint description of the text image, the VAE algorithm model, the bbox detection model, the SAM model, and some corresponding algorithm parameters. The main parameters include: denoising, iteration step number, scheduler, sampler, etc.

[0037] In step S104, character key point detection and posture recognition are performed on the static restoration images frame by frame, the key point positions and postures of the characters are calibrated, and multi-frame motion continuity restoration is performed based on the key point positions and postures to generate dynamic restoration images frame by frame.

[0038] refer to Figure 3 , which shows a schematic diagram of character key point detection according to an embodiment of the present invention.

[0039] As an example, the ControlNet algorithm architecture can be used to perform key point detection and posture recognition on the image sequence after video frame decomposition, calibrate the key point positions and upper body postures of the characters, and then use the AnimateDiff algorithm framework to create high-quality inter-frame transitions, thereby making the action connection between sequence frames smoother and achieving a smooth and natural animation effect.

[0040] AnimateDiff is a technology that combines diffusion models with animation generation. Its key idea is to leverage the powerful generation capabilities of diffusion models to create high-quality frame transitions, resulting in smooth and natural animation effects. AnimateDiff can be considered a method for converting static images into dynamic sequences, making it particularly suitable for creating artistically inspired animated content.

[0041] As an embodiment of the present invention, the key points of the character may include facial contour points and / or facial structure points.

[0042] In step S105, the frame-by-frame dynamic restoration images are synthesized into a video to generate a vehicle-mounted virtual image video. The sequence frames of the dynamic restoration images are synthesized into a video and finally merged to obtain a motion-coherent clone image.

[0043] As an embodiment of the present invention, the method may further include: performing background removal on each frame of the dynamically restored image to generate a frame-by-frame background-free image; performing video synthesis on the frame-by-frame background-free images to generate a background-free in-vehicle virtual image video; and performing video synthesis on the sequence of image frames with the background removed to ultimately merge and produce a cloned image with coherent movements and no background.

[0044] As an example, you can use the BiRefNet algorithm module, selecting the BiRefNet-DIS_ep580 model, to perform background removal on a video frame sequence. BiRefNet is a comfyui algorithm module for image background removal. Its input parameters are the image to be processed and the background removal algorithm model. For example, the BiRefNet-DIS_ep580 model can be selected to convert an image into an image with a transparent background.

[0045] By performing background removal, a virtual image without background can be obtained, which is convenient for embedding into the display screen of the vehicle computer.

[0046] On the other hand, an embodiment of the present invention proposes a device for generating a vehicle-mounted virtual image, the device including a memory and a processor, the memory storing a computer program, and implementing the method for generating a vehicle-mounted virtual image described in any of the above embodiments when the computer program is executed by the processor.

[0047] In yet another aspect, an embodiment of the present invention provides a vehicle, which includes the device for generating an in-vehicle virtual image according to the above embodiment.

[0048] In yet another aspect, an embodiment of the present invention provides a storage medium storing computer-readable instructions. When the instructions are executed by a processor, the method for generating a vehicle-mounted virtual image described in any of the above embodiments is executed.

[0049] In yet another aspect, an embodiment of the present invention provides a computer program product comprising computer-readable instructions, which, when executed by a processor, executes the steps of the method for generating a vehicle-mounted virtual image described in any of the above embodiments.

[0050] The main problems currently faced by image cloning in the smart cockpit domain include: ① being rigid or having only a few preset images; ② having little user interaction and a poor user experience. To address these issues, the present invention proposes a technical solution for generating in-vehicle virtual images, namely, an image cloning solution. This solution involves decomposing the original animation video, performing frame-by-frame character face swapping, character contour control and facial redrawing, character key point detection and multi-frame motion continuity restoration, and finally synthesizing multiple frames to create a dynamic cloned image. This ultimately creates an image cloning solution that can cartoonize characters while preserving their movements, enabling personalized customization of in-vehicle virtual images and significantly improving the user experience of image cloning and interaction.

[0051] Based on the implementation manner of the present invention, the user's own elements can be integrated into the image setting, and supplemented by coherent movements, the image cloning is natural and smooth, which will make the user feel more intimate, strengthen the connection between the user and the car computer, and experience the feeling of unity between man and car.

[0052] The image cloning algorithm introduces multiple stability-enhancing links, which not only ensures the stability of facial features after face-changing, but also makes the movement connection of the deframed image more consistent and coordinated, ensuring the quality of the dynamic cloned image.

[0053] In addition to being applicable to vehicle computers, the vehicle computer virtual image generated by the embodiments of the present invention can also be applied to intelligent customer service, FAQ parts of intelligent dialogue robots, etc.

[0054] For illustrative purposes, the foregoing description of the embodiments of the present invention has been given, which is not exhaustive nor intended to limit the present invention to disclosed exact forms. It will be appreciated by those skilled in the art that various changes may be made without departing from the scope of the present invention, and that elements therein may be replaced with equivalents. In addition, without departing from the basic scope of the present invention, many modifications may be made so that specific situations or materials are adapted to the teachings of the present invention. Therefore, the present invention is not intended to be limited to the specific embodiments disclosed as the best mode for realizing the present invention, and the present invention will include all embodiments within the scope of the appended claims.

Claims

1. A method for generating a vehicle-mounted virtual image, characterized in that: The method comprises: Deframe the original comic video to generate a frame-by-frame original image sequence; According to the image of the target image, the character image in each image in the frame-by-frame original image sequence is replaced with the target image to generate a frame-by-frame face-changing image; Performing character contour control on each of the frame-by-frame face-changing images to generate a frame-by-frame static restoration image; Performing character key point detection and posture recognition on the frame-by-frame static restoration images, calibrating the key point positions and postures of the character, performing multi-frame motion continuity restoration based on the key point positions and postures, and generating frame-by-frame dynamic restoration images; The frame-by-frame dynamic restoration images are subjected to video synthesis to generate a vehicle-mounted virtual image video.

2. The method according to claim 1, characterized in that According to the target image, the character image in each image in the frame-by-frame original image sequence is replaced with the target image to generate the frame-by-frame face-swapped image, which includes: Face replacement, face detection and face restoration are performed on each image in the frame-by-frame original image sequence according to the target image, thereby generating the frame-by-frame face-changing image.

3. The method according to claim 1, characterized in that The method further comprises: Performing background removal on the frame-by-frame dynamically restored images to generate frame-by-frame background-free images; The background-free images are synthesized frame by frame to generate a vehicle-mounted virtual image video without background.

4. The method according to any one of claims 1 to 3, characterized in that Performing character contour control on each image in the frame-by-frame face-changing image to generate a frame-by-frame static restoration image further includes: Performing face redrawing and restoration on the image after performing character contour control to generate the frame-by-frame static restoration image.

5. The method according to any one of claims 1 to 3, characterized in that According to the target image, the character image in each image in the frame-by-frame original image sequence is replaced with the target image to generate the frame-by-frame face-swapped image, which includes: Get the target source image; Marking and numbering the faces of the people in the target source image to generate candidate images with serial numbers; Receive the serial number entered by the user; The target image is determined from the candidate images with serial numbers according to the serial number input by the user.

6. The method according to any one of claims 1 to 3, characterized in that The key points of a character include facial feature contour points and / or facial structure points.

7. A device for generating a vehicle-mounted virtual image, characterized in that: The device includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method for generating a vehicle-mounted virtual image according to any one of claims 1 to 6 is implemented.

8. A vehicle, characterized in that: The vehicle includes the device for generating an in-vehicle virtual image according to claim 7.

9. A storage medium storing computer-readable instructions, wherein when the instructions are executed by a processor, the method for generating a vehicle-mounted virtual image according to any one of claims 1 to 6 is executed.

10. A computer program product comprising computer-readable instructions, characterized in that When the instructions are executed by a processor, the steps of the method for generating a vehicle-mounted virtual image according to any one of claims 1 to 6 are executed.