Vehicle type wallpaper generation method and device, electronic equipment and vehicle

By extracting and analyzing the feature information of the target style diagram, line drawing and foreground model diagram, the generation model can more accurately understand and process the generation intention of model wallpaper, solve the problem of low quality of model wallpaper generation in the existing technology, and achieve high-quality model wallpaper generation.

CN120235974APending Publication Date: 2025-07-01GREAT WALL MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510260707.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

When generating vehicle model wallpapers, the prior art has problems such as loss of details, inaccurate or confusing structural outlines, and unnatural visual effects when facing background images of complex geometric structures, which affects the generation quality and limits the application of AIGC technology in high-demand fields.

Method used

By obtaining the target style diagram, target line drawing and target prospect model diagram, the first feature extraction model is used to extract style feature information, the second feature extraction model is used to extract structural outline feature information, and combining the third feature extraction model to analyze key elements and correlation, generate target prompt words, and enter the generated model for image generation to ensure the accuracy of style, structural outline and model features.

Benefits of technology

The generation quality of model wallpaper is improved, the coordination of style characteristics, structural outlines and model characteristics is ensured, and high-quality target model wallpaper is generated, avoiding common problems of blurred details and structural distortion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235974A_ABST
    Figure CN120235974A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle type wallpaper generation method and device, electronic equipment and a vehicle, and belongs to the technical field of image processing, and the method comprises the steps: obtaining a target style graph, a target line draft graph and a target foreground vehicle type graph; extracting style feature information of the target style graph through a first feature extraction model; extracting structure contour feature information of the target line draft graph through a second feature extraction model; inputting the target style graph, the target line draft graph and the target foreground vehicle model graph into a third feature extraction model for feature extraction to obtain a target cue word; and inputting the style feature information, the structure contour feature information, the target cue word and the target foreground vehicle model graph into a generative model for image generation to obtain target vehicle model wallpaper. According to the method, the multi-dimensional feature information is extracted through the feature extraction models, rich generation basis is provided for the target vehicle type wallpaper, the generation model can understand the features of the target vehicle type wallpaper to be generated from multiple aspects, and therefore the quality of the generated target vehicle type wallpaper is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to a method, device, electronic device, and vehicle for generating vehicle wallpapers. Background Art

[0002] Generating vehicle wallpapers based on Artificial Intelligence Generated Content (AIGC) technology has shown great potential in improving efficiency and creative design. Current related technologies can generate images with relatively good quality and effects in some cases. However, when dealing with background images containing complex geometric structures, such as variously shaped rocks, trees, flowers, etc. in urban buildings and natural landscapes, these objects often contain complex components, and there are still problems such as loss of detailed content, inaccurate or chaotic structural outlines, and unnatural visual effects. These problems not only affect the overall quality of the generated vehicle wallpapers but also limit the application of AIGC technology in higher - requirement fields to a certain extent.

[0003] Therefore, how to improve the quality of vehicle wallpapers generated by AIGC technology has become an urgent technical problem for those skilled in the art. Summary of the Invention

[0004] Embodiments of this application provide a method, device, electronic device, and vehicle for generating vehicle wallpapers, which can solve the problems of loss of detailed content, inaccurate or chaotic structural outlines, and unnatural visual effects in generating vehicle wallpapers in related technologies, and improve the quality of vehicle wallpapers.

[0005] In a first aspect, embodiments of this application provide a method for generating vehicle wallpapers. The method includes: obtaining a target style map, a target line drawing, and a target foreground vehicle map; extracting style feature information corresponding to the target style map through a first feature extraction model; extracting structural outline feature information corresponding to the target line drawing through a second feature extraction model; inputting the target style map, the target line drawing, and the target foreground vehicle map into a third feature extraction model for feature extraction to obtain a target prompt; inputting the style feature information, the structural outline feature information, the target prompt, and the target foreground vehicle map into a generation model for image generation to obtain a target vehicle wallpaper.

[0006] Optionally, the step of inputting the target style map, the target line drawing map, and the target foreground vehicle model map into a third feature extraction model for feature extraction to obtain a target prompt includes: inputting the target style map, the target line drawing map, and the target foreground vehicle model map into a vision-language large model for feature extraction to obtain the target prompt. In this embodiment, by analyzing the input target style map, target line drawing map, and target foreground vehicle model map through the vision-language large model, the key elements therein and the potential correlations between various types of input maps are identified, and these elements are transformed into target prompts suitable for the generation model to enrich the constraints for generating the target vehicle model wallpaper. The understanding and feature fusion ability of the vision-language large model can extract the key elements in the target style map, target line drawing map, and target foreground vehicle model map, and perform correlation analysis based on the currently input maps, combining and expanding the prompts to generate rich target prompts.

[0007] Optionally, the step of inputting the style feature information, the structural contour feature information, the target prompt, and the target foreground vehicle model map into a generation model for image generation to obtain the target vehicle model wallpaper includes: inputting the style feature information, the structural contour feature information, the target prompt, and the target foreground vehicle model map into an Inpaint diffusion model generation model for image generation to obtain the target vehicle model wallpaper. In this embodiment, the style feature information, the structural contour feature information, and the target prompt can clearly instruct the Inpaint diffusion model. The Inpaint diffusion model generates an image corresponding to and associated with the style, structural contour, and target prompt based on these feature information and the target prompt, achieving precise control of the generation result.

[0008] Optionally, the step of inputting the style feature information, the structural contour feature information, the target prompt, and the target foreground vehicle model map into a generation model for image generation to obtain the target vehicle model wallpaper includes: extracting the foreground mask area of the target vehicle model wallpaper from the target foreground vehicle model map; extracting the background mask area of the target vehicle model wallpaper from the target line drawing map according to the structural contour feature information; filling the foreground mask area and the background mask area according to the style feature information and the target prompt to obtain the target vehicle model wallpaper. In this embodiment, the foreground mask area defines the foreground area to be filled, the background mask area defines the background area to be filled, and the style feature information and the target prompt define what content to fill into the foreground mask area and the background mask area. Thus, the generation model is guided to clarify the areas and content to be filled, thereby generating a high-quality target vehicle model wallpaper that meets the style requirements, structural contour requirements, and coordination requirements.

[0009] Optionally, obtaining the target line drawing includes: inputting the target foreground vehicle model image and the initial line drawing into the vision-language large model for fusion to obtain the target line drawing; wherein, the initial line drawing is used to describe the background contour of the target vehicle model wallpaper; the target line drawing is used to describe the background contour and the foreground contour of the target vehicle model wallpaper. In this embodiment, by optimizing the target line drawing input to the model, the target foreground vehicle model image representing the foreground object is fused with the initial line drawing, and the obtained target line drawing is not simply a sketch of the foreground and background contours, but a composite map integrating the spatial topological relationship between the foreground and the background, describing the specific layout and correlation between the background and the foreground. Thereby enabling the subsequent generation model to quickly determine the foreground mask area and the background mask area according to the target line drawing, without the need to analyze in detail the logical relationship between the target foreground vehicle model image and the initial line drawing to judge their specific layout and correlation in the target vehicle model wallpaper, improving the generation efficiency of the target vehicle model wallpaper.

[0010] Optionally, the step of inputting the target foreground vehicle model image and the initial line drawing into the vision-language large model for fusion to obtain the target line drawing includes: extracting the vehicle contour of the target vehicle model according to the target foreground vehicle model image; fusing the vehicle contour into the target area of the initial line drawing to obtain the target line drawing. In this embodiment, the vision-language large model analyzes the logical relationship and the mutual correlation between the vehicle in the target foreground vehicle model image and the initial line drawing, determines the layout and position during fusion based on their logical relationship and mutual correlation, fuses the target foreground vehicle model image and the initial line drawing, and obtains the target line drawing integrating the spatial topological relationship between the foreground and the background, enabling the subsequent generation model to quickly determine the foreground mask area and the background mask area according to the target line drawing, and improving the generation efficiency of the target vehicle model wallpaper.

[0011] Optionally, after obtaining the target vehicle model wallpaper, the method further includes: obtaining the filled foreground mask area or the filled background mask area; comparing the filled foreground mask area with the foreground mask area, or comparing the filled background mask area with the background mask area; and outputting the target vehicle model wallpaper when the difference between the filled foreground mask area and the foreground mask area is less than a first preset threshold, or when the difference between the filled background mask area and the background mask area is less than a second preset threshold. In this embodiment, by comparing the filled foreground mask area with the foreground mask area before filling, or comparing the filled background mask area with the background mask area before filling, the effect of the generated target vehicle model wallpaper is evaluated, so as to save the target vehicle model wallpaper for use when the expected effect is met, or dynamically adjust each step of generating the target vehicle model wallpaper when the expected effect is not met, thereby further optimizing the target vehicle model wallpaper.

[0012] In a second aspect, an embodiment of the present application provides a vehicle model wallpaper generation device, which includes: an acquisition module for acquiring a target style map, a target line drawing, and a target foreground vehicle model map; a feature extraction module for extracting style feature information corresponding to the target style map through a first feature extraction model; extracting structural contour feature information corresponding to the target line drawing through a second feature extraction model; and inputting the target style map, the target line drawing, and the target foreground vehicle model map into a third feature extraction model for feature extraction to obtain a target prompt; and a generation module for inputting the style feature information, the structural contour feature information, the target prompt, and the target foreground vehicle model map into a generation model for image generation to obtain a target vehicle model wallpaper.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0015] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the steps of the method described in the first aspect.

[0016] In a sixth aspect, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes programs or instructions that, when executed, implement the steps of the method described in the first aspect.

[0017] In a seventh aspect, an embodiment of the present application provides a vehicle. The vehicle includes a memory and a processor. An executable program code is stored in the memory. The processor is configured to call and execute the executable program code, and when the executable program code is executed by the processor, it implements the steps of the method described in the first aspect.

[0018] In the embodiment of the present application, by obtaining a target style map, a target line drawing, and a target foreground vehicle model map, the style feature information corresponding to the target style map is extracted by a first feature extraction model, so that the style of the target vehicle model wallpaper is determined through this style feature information. The structural contour feature information corresponding to the target line drawing is extracted by a second feature extraction model, so that the structure and contour of the target vehicle model wallpaper are determined through this structural contour feature information. Then, the target style map, the target line drawing, and the target foreground vehicle model map are input into a third feature extraction model for feature extraction to obtain a target prompt word. The extracted target prompt word synthesizes the style features corresponding to the target style map, the structural contour features corresponding to the target line drawing, and the vehicle model features corresponding to the target foreground vehicle model map, and describes the features of the target vehicle model wallpaper to be generated from multiple dimensions including style, structural contour, and vehicle model. The style feature information describing the style of the target vehicle model wallpaper to be generated, the structural contour feature information describing the structural contour of the target vehicle model wallpaper to be generated, and the target prompt word describing the comprehensive multi-faceted features of the target vehicle model wallpaper to be generated are used as inputs to the generation model. The generation model is richly and accurately guided by these feature information and the target prompt word, and the generation of the target vehicle model wallpaper is constrained, so that the generation model can more accurately understand and process the generation intention of the target vehicle model wallpaper, providing rich and accurate basis for generating the target vehicle model wallpaper. Furthermore, the style feature information, the structural contour feature information, the target prompt word, and the target foreground vehicle model map are input into the generation model for image generation. It is possible to separately constrain and detail the style, structural contour, and vehicle model of the target vehicle model wallpaper, and also to increase the coordination of generating the target vehicle model wallpaper by synthesizing the features of the three, thereby obtaining a high-quality target vehicle model wallpaper and avoiding the generation model from generating wallpapers with common elements and blurred details based on the input fuzzy or simple feature information.

[0019] The above description is only an overview of the technical solution of this application. In order to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically illustrates the specific implementation manners of this disclosure. Brief Description of the Drawings

[0020] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of this disclosure. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 A flowchart showing a method for generating a vehicle wallpaper provided by an embodiment of this application; Figure 2 A structural diagram showing a ControlNet model provided by an embodiment of this application; Figure 3 A structural diagram showing an IP-Adapter model provided by an embodiment of this application; Figure 4 A structural diagram showing an Inpaint diffusion model provided by an embodiment of this application; Figure 5 A flowchart showing another method for generating a vehicle wallpaper provided by an embodiment of this application; Figure 6 A schematic diagram showing a method for generating a vehicle wallpaper provided by an embodiment of this application; Figure 7 A structural diagram showing a device for generating a vehicle wallpaper provided by an embodiment of this application; Figure 8 A structural diagram showing an electronic device provided by an embodiment of this application; Figure 9 A structural diagram showing a vehicle provided by an embodiment of this application. Detailed Description of the Preferred Embodiments

[0021] The exemplary embodiments of this disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that this disclosure can be more thoroughly understood and the scope of this disclosure can be fully conveyed to those skilled in the art.

[0022] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0023] In the related art, the artificial intelligence-generated content technology has shown significant efficiency advantages and creative breakthroughs in the intelligent generation of vehicle wallpapers. Relying on the image generation algorithm of deep learning, this technology can quickly complete the full-process design from concept conception to visual presentation, and has achieved high-definition image quality output and diverse style adaptation in conventional application scenarios, greatly shortening the creation cycle in the traditional design process. Currently, through the collaborative optimization of Generative Adversarial Networks (GAN) and Diffusion Model, in the scenario of a monochromatic background or simple composition, it has been able to stably output the rendering effect of the vehicle body that meets industrial standards.

[0024] However, when faced with complex scenarios with multiple geometric structures, the existing technology still faces certain technical bottlenecks. For example, in the urban landscape scenario, logical paradoxes often occur in the processing of the structural contour relationship of the building complex, which may be manifested as structural overlap or unclear contours. For natural terrain scenarios, when depicting rocks or trees, their textures are ignored and the hierarchical relationship is misinterpreted. The above-mentioned misalignment phenomena not only cause a sense of disharmony in visual perception, but also form multiple restrictions at the professional application level: for example, in the field of automotive digital marketing, it may affect consumers' perception of product positioning; in the pre-production of film and television special effects, it will reduce the credibility of the scene; and in the industrial design review link, it may lead to misjudgment of the spatial perception of the design scheme.

[0025] Therefore, in order to solve the above problems and improve the quality of the generated vehicle wallpapers, the embodiments of this application provide a vehicle wallpaper generation method, device, electronic device, and vehicle, which can at least solve the above problems.

[0026] The vehicle wallpaper generation method, device, electronic device, and vehicle provided by the embodiments of this application will be described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0027] Figure 1The flowchart of a method for generating a vehicle wallpaper provided by an embodiment of the present application is shown. This method can be executed by an electronic device. Refer to Figure 1 , and the method may include the following steps.

[0028] Step 110, obtain a target style map, a target line drawing, and a target foreground vehicle map.

[0029] The target style map among them can be selected from a style map database, and at least one style map is stored in the style map database for selection. The target style map is used to represent the style features to be filled. That is to say, the target style map provides style features or a style tone for the generation of the target vehicle wallpaper. For example, features in aspects such as the color, light and shadow, and texture of the target vehicle wallpaper.

[0030] The target line drawing among them can be selected from a line drawing database, and at least one line drawing is stored in the line drawing database for selection. The target line drawing is used to represent the basic geometric structure and contour of the background to be filled. That is to say, the target line drawing provides the basic geometric structure and contour for the generation of the target vehicle wallpaper, and can clarify information such as the shape and boundary of the target vehicle wallpaper.

[0031] The target foreground vehicle map is used to represent the foreground object in the target vehicle wallpaper. In the embodiment of the present application, the target foreground vehicle map is a car.

[0032] Step 120, extract the style feature information corresponding to the target style map through a first feature extraction model; extract the structural contour feature information corresponding to the target line drawing through a second feature extraction model; input the target style map, the target line drawing, and the target foreground vehicle map into a third feature extraction model for feature extraction to obtain a target prompt.

[0033] Among them, for the target style map representing the style features to be filled, the style feature information corresponding to the target style map is extracted through a first feature extraction model. This first feature extraction model is used to extract visual features such as color distribution, texture pattern, and composition style from the target style map, and can be fused with the generation model to transfer these visual features to the target vehicle wallpaper to be generated, as constraints in terms of style for the generated vehicle wallpaper. For example, feature information such as color, light and shadow, and texture, that is, encoding the style features in the target style map into feature information that can be processed by the subsequent generation model, as an additional condition for generating the target vehicle wallpaper. Through this implementation method, both the style diversity of the target vehicle wallpaper and the accuracy and naturalness of style transfer of the target vehicle wallpaper are improved.

[0034] Among them, for the target line drawing representing the basic geometric structure and contour of the background to be filled, the second feature extraction model extracts the structural contour feature information corresponding to the target line drawing. The second feature extraction model is used to extract the structural contour feature information in the target line drawing and can be fused with the generation model as a constraint condition for the generated vehicle model wallpaper in terms of structural contour, ensuring that the generated vehicle model wallpaper is geometrically consistent with the target line drawing. Figure 1 For example, feature information such as edges, shapes, and ranges. This structural contour feature information serves as an additional condition for generating the target vehicle model wallpaper to guide the integrity of the target vehicle model wallpaper in terms of structural contour features, improve the structural contour accuracy of the generated vehicle model wallpaper, and avoid common problems such as structural distortion or detail loss in traditional generation models.

[0035] Among them, through the third feature extraction model, comprehensive analysis is performed on the target style map, the target line drawing, and the target foreground vehicle model map to extract appropriate target prompt words. The third feature extraction model is used to analyze the input target style map, target line drawing, and target foreground vehicle model map, identify key elements therein (such as color features, texture features, artistic features, etc. in the target style map, and structural features, contour features, layout features, etc. of objects in the target line drawing), as well as potential correlations between various types of input maps, and transform these elements into target prompt words suitable for the generation model. These target prompt words are also used as constraint conditions for generating the vehicle model wallpaper. For example, through comprehensive analysis of the target style map, target line drawing, and target foreground vehicle model map, detailed description information about the target style map, target line drawing, and target foreground vehicle model map is obtained, and then this detailed description information is analyzed to select prompt words with higher frequencies of occurrence, prompt words with specific indicative meanings, etc. These prompt words are used as target prompt words. Prompt words with higher adaptability can also be selected as target prompt words. For example, the prompt words include the first type of prompt words A1 and A2, and the second type of prompt word B. The frequency of occurrence of A1 is higher than that of A2, but after analysis, the adaptability between A2 and B is higher than that of A1, so A2 is preferentially selected as the target prompt word. Thus, by inputting these target prompt words into the subsequent generation model, rich and accurate basis is provided for the generation model to generate the target vehicle model wallpaper.

[0036] Step 130, input the style feature information, the structural contour feature information, the target prompt words, and the target foreground vehicle model map into the generation model for image generation to obtain the target vehicle model wallpaper.

[0037] Among them, the style feature information, the structural contour feature information, the target prompt word, and the target foreground vehicle model diagram are all used as inputs to the generation model, enabling the generation model to more accurately understand the generation intention of the target vehicle model wallpaper based on rich features. The generation model integrates the above-mentioned first feature extraction model, second feature extraction model, and third feature extraction model, and uses the style feature information extracted by the first feature extraction model, the structural contour feature information extracted by the second feature extraction model, the target prompt word generated by the third feature extraction model, and the target foreground vehicle model diagram as inputs to the generation model. Through the constraints of the input target foreground vehicle model diagram, various types of feature information, and the target prompt word, a high-quality target vehicle model wallpaper that integrates various types of features and is constrained by the target prompt word is generated. The feature information in multiple dimensions enables the generation model to coordinate and unify the style features, structural contour features, and foreground vehicle model diagram, and meticulously depict the target vehicle model wallpaper.

[0038] In the embodiment of the present application, by obtaining the target style diagram, the target line drawing diagram, and the target foreground vehicle model diagram, the style feature information corresponding to the target style diagram is extracted by the first feature extraction model, so that the style of the target vehicle model wallpaper is determined through this style feature information. The structural contour feature information corresponding to the target line drawing diagram is extracted by the second feature extraction model, so that the structure and contour of the target vehicle model wallpaper are determined through this structural contour feature information. Then, the target style diagram, the target line drawing diagram, and the target foreground vehicle model diagram are input into the third feature extraction model for feature extraction to obtain a target prompt word. This target prompt word synthesizes multi-dimensional features such as the style features corresponding to the target style diagram, the structural contour features corresponding to the target line drawing diagram, and the vehicle model features corresponding to the target foreground vehicle model diagram, and describes the features of the target vehicle model wallpaper to be generated currently from multiple aspects of style, structural contour, and vehicle model. The style feature information describing the style of the target vehicle model wallpaper to be generated, the structural contour feature information describing the structural contour of the target vehicle model wallpaper to be generated, and the target prompt word describing the comprehensive multi-aspect features of the target vehicle model wallpaper to be generated are used as inputs to the generation model. Through these feature information and the target prompt word, the generation model is richly and accurately guided, and the generation of the target vehicle model wallpaper is constrained, enabling the generation model to more accurately understand and process the generation intention of the target vehicle model wallpaper, providing rich and accurate basis for generating the target vehicle model wallpaper. Furthermore, the style feature information, the structural contour feature information, the target prompt word, and the target foreground vehicle model diagram are input into the generation model for image generation, which can not only separately constrain and detail the style, structural contour, and vehicle model of the target vehicle model wallpaper, but also comprehensively integrate the features of the three to increase the coordination of generating the target vehicle model wallpaper, thereby obtaining a high-quality target vehicle model wallpaper.

[0039] In an exemplary embodiment, based on the characteristic that the second feature extraction model in the above embodiment is used to extract the structural contour feature information in the target line drawing as a constraint condition for generating the vehicle model wallpaper in terms of the structural contour, an exemplary second feature extraction model, namely the ControlNet model, is provided in the embodiment of the present application. Extracting the structural contour feature information corresponding to the target line drawing through the second feature extraction model in the above embodiment may include: extracting the structural contour feature information corresponding to the target line drawing through the ControlNet model.

[0040] Among them, Figure 2 shows a structural schematic diagram of a ControlNet model provided in the embodiment of the present application. The ControlNet model is an "auxiliary" model. It supports more input conditions to control the image generation process by adding an auxiliary module to the diffusion model. These input conditions can be information in the form of the edges, sketches, depth maps, semantic segmentation maps, key points, scribbles, etc. of the image. Users can control certain attributes of the image, such as the painting style, actions, colors, etc., by manually editing the source image or providing specific input conditions, so as to generate new images. In the embodiment of the present application, by introducing the ControlNet model into the generation model as a control condition in the generation model, refined guidance for the generation process is realized. Specifically, referring to Figure 2 , the input signals of the model include: a target prompt (Prompt c), which is converted into a text embedding vector through the TextEncoder; a time step (Time t), which is encoded into a time condition vector by the Time Encoder; a control condition (Condition c), such as a line drawing, which is input as a structural constraint. The SD encoder block (Encoder Block) captures the multi-dimensional features of the input signals step by step from the high-resolution layer (64×64) to the low-resolution layer (8×8). The ControlNet branch processes the control condition in parallel to generate the structural contour feature information, which is fused with the generation model through zero convolution. The SD decoder block (Decoder Block) then goes from the low-resolution layer (8×8) to the high-resolution layer (64×64) step by step to reconstruct the final image. The generation model fused with the ControlNet model can gradually transform the structural contour feature information of the target line drawing into image details that are close to reality and conform to physical laws. The ControlNet model includes an encoder block from the high-resolution layer to the low-resolution layer and a decoder block from the low-resolution layer to the high-resolution layer, analyzing features layer by layer and then integrating features layer by layer to guide the generation of the target vehicle model wallpaper, ensuring the geometric accuracy of the target vehicle model wallpaper and assisting the generation model to generate high-quality target vehicle model wallpapers.

[0041] In an exemplary embodiment, based on the above embodiments, the third feature extraction model is used to analyze the input target style map, target line drawing, and target foreground vehicle model map, identify the key elements therein and the potential correlations between various types of input maps, and convert these elements into the characteristics of the target prompt words suitable for the generation model. In the embodiments of the present application, an exemplary third feature extraction model is provided, namely Vision-Language Large Models (VLLMs). The above-mentioned input of the target style map, the target line drawing, and the target foreground vehicle model map into the third feature extraction model for feature extraction to obtain the target prompt words may include: inputting the target style map, the target line drawing, and the target foreground vehicle model map into the vision-language large model for feature extraction to obtain the target prompt words.

[0042] Among them, VLLMs is a large-scale model that combines visual and language processing capabilities. VLLMs is jointly trained with a large amount of image data and text data, can understand and generate content containing visual information and language information, can more accurately understand and process images, and extract the target prompt words for generating the target vehicle model wallpaper.

[0043] In an exemplary embodiment, based on the above embodiments, the first feature extraction model is used to extract visual features such as color distribution, texture pattern, and composition style from the target style map, and transfer these visual features to the characteristics of the target vehicle model wallpaper to be generated. In the embodiments of the present application, an exemplary first feature extraction model is provided, namely the IP-Adapter model. The above embodiments extract the style feature information corresponding to the target style map through the first feature extraction model, including: extracting the style feature information corresponding to the target style map through the IP-Adapter model.

[0044] Among them, Figure 3The figure shows a schematic structural diagram of an IP-Adapter model provided by an embodiment of the present application. The IP-Adapter model is a model for enhancing image processing capabilities, especially the ability to integrate specific styles or attributes in a generative model. By adding one or more layers of adapters to an existing generative model, the generative model can learn new functions while maintaining its original functions. These adapters can enhance the flexibility and adaptability of the model, so that the style feature information corresponding to the target style map can be quickly adapted to different style types through the IP-Adapter model, reducing the demand for computing resources. Specifically, the main role of IP-Adapter is to encode the style information of the input image into a form that can be understood by the existing generative model, that is, style feature information or vectors. In the embodiment of the present application, by adding the IP-Adapter model to the generative model, the generative model integrated with IP-Adapter encodes the input target style map through an image encoder to obtain style feature information (Image Features), encodes the input text (target prompt words, etc.) through a text encoder to obtain text feature information (Text Features), and deeply fuses the style feature information and the text feature information through a decoupled cross-attention layer. The fused feature information is transmitted to the generative model, and the noise in the noisy target vehicle wallpaper to be generated (x t ) is gradually removed through a denoising layer (Denoising U-Net) to generate the target vehicle wallpaper (x t-1 ). For the generation requirements of vehicle wallpapers with different styles, it not only avoids the high cost of retraining the entire generative model, but also can quickly transfer a variety of different styles to the target vehicle wallpaper to be generated, providing accurate style features for the generative model to generate the target vehicle wallpaper and assisting the generative model to generate high-quality target vehicle wallpapers.

[0045] In an exemplary embodiment, based on the generation model in the above embodiment, the first feature extraction model, the second feature extraction model, and the third feature extraction model are fused. Through the input target foreground vehicle model image, various types of feature information, and the constraint of the target prompt, the characteristics of the target vehicle model wallpaper that fuses various types of features and is constrained by the target prompt are generated. In the embodiment of the present application, an exemplary generation model is provided, namely the Inpaint diffusion model. In the above embodiment, the style feature information, the structural contour feature information, the target prompt, and the target foreground vehicle model image are input into the generation model for image generation to obtain the target vehicle model wallpaper, which may include: inputting the style feature information, the structural contour feature information, the target prompt, and the target foreground vehicle model image into the Inpaint diffusion model generation model for image generation to obtain the target vehicle model wallpaper.

[0046] Among them, referring to Figure 4 , Figure 4 shows a schematic structural diagram of an Inpaint diffusion model provided in the embodiment of the present application. The Inpaint diffusion model is an image generation model based on deep learning. By gradually adding noise, the image is transformed from the original distribution to a simple distribution (such as a Gaussian distribution), and then the original image is recovered from the noise through an inverse process. Specifically, it can be divided into two stages: the forward process (diffusion process): gradually add noise to the image until the image completely becomes random noise (from x0 to x T ). The inverse process (denoising process): gradually recover the original image from the noise (from x T to x0). After multiple steps of denoising, the Inpaint diffusion model generates a complete repaired image, where the missing area is filled with content consistent with the surrounding image. By further fusing the multi-dimensional constraint conditions of the style feature information, the structural contour feature information, the target prompt, and the target foreground vehicle model image through the Inpaint diffusion model to generate the vehicle model wallpaper for repair, it can not only ensure that the generated target vehicle model wallpaper meets the style requirements and the requirements of the target prompt, but also be close to reality and conform to physical laws, making the target vehicle model wallpaper have better coordination, thereby improving the quality of the target vehicle model wallpaper.

[0047] In the above embodiment, by using the style feature information, the structural contour feature information, the target prompt, and the target foreground vehicle model image as the input of the generation model, the generation of the target vehicle model wallpaper is constrained from the aspects of style, structural contour, and comprehensively. Specifically, in an exemplary embodiment, the above step 130 inputs the style feature information, the structural contour feature information, the target prompt, and the target foreground vehicle model image into the generation model for image generation to obtain the target vehicle model wallpaper, which may include the following steps.

[0048] Step 131: Extract the foreground mask area of the target vehicle model wallpaper from the target foreground vehicle model image.

[0049] The target foreground vehicle model image is used to represent the foreground object in the target vehicle model wallpaper, and the foreground object is extracted therefrom as the foreground mask area to be filled in the target vehicle model wallpaper. The mask area includes a foreground mask area and a background mask area. The foreground and background are distinguished by binary values, with the foreground being white (1 or 255) and the background being black (0). In the generated target vehicle model wallpaper, the target foreground vehicle model image needs to be located in the foreground mask area, and the target line drawing needs to be located in the background mask area. The foreground mask area to be filled in the target vehicle model wallpaper can be determined through the target foreground vehicle model image. When generating the target vehicle model wallpaper, the relevant content of the target foreground vehicle model image is filled into this foreground mask area.

[0050] Step 132: Extract the background mask area of the target vehicle model wallpaper from the target line drawing according to the structural contour feature information.

[0051] The target line drawing is used to represent the basic geometric structure and contour of the background to be filled. The extracted structural contour feature information describes the contour and structure of the background of the target vehicle model wallpaper. The structural contour of the background is extracted from the target line drawing as the background mask area to be filled in the target vehicle model wallpaper. When generating the target vehicle model wallpaper, the relevant content of the target line drawing is filled into this background mask area.

[0052] Step 133: Fill the foreground mask area and the background mask area according to the style feature information and the target prompt words to obtain the target vehicle model wallpaper.

[0053] Among them, the foreground mask area defines the foreground area to be filled, the background mask area defines the background area to be filled, and the style feature information and the target prompt words define what content to fill into the foreground mask area and the background mask area, thus constituting the complete target vehicle model wallpaper.

[0054] In the embodiment of the present application, the generation model generates a target vehicle model wallpaper based on the input style feature information, the structural contour feature information, the target prompt, and the target foreground vehicle model image. Specifically, the foreground mask area to be filled can be determined through the target foreground vehicle model image, the background mask area to be filled can be determined through the structural contour feature information, and then through the understanding and processing of the style feature information and the target prompt, the foreground mask area and the background mask area are filled based on the features described by the style feature information and the target prompt, so as to realize the filling of the foreground mask area and the background mask area and obtain the target vehicle model wallpaper. The generation of the target vehicle model wallpaper inputs multi-dimensional features into the generation model, enriching the input of the generation model, enabling the generation model to more clearly define the features of each dimension of the target vehicle model wallpaper, and thus accurately generating a high-quality target vehicle model wallpaper.

[0055] To further improve the generation efficiency and accuracy of the target vehicle model wallpaper and quickly obtain the layout and correlation relationship between the foreground and background in the target vehicle model wallpaper, in an exemplary embodiment, the initial line drawing is adjusted by VLLMs capable of understanding and processing images, and the target foreground vehicle model image is fused with the initial line drawing. The step of obtaining the target line drawing in step 110 above may include step 111: inputting the target foreground vehicle model image and the initial line drawing into the VLLMs for fusion to obtain the target line drawing.

[0056] Among them, the initial line drawing is used to describe the background contour of the target vehicle model wallpaper; the target line drawing is used to describe the background contour and the foreground contour of the target vehicle model wallpaper. The target line drawing provides a structured guidance for the subsequent generation model through precise geometric constraints. The target line drawing is not simply a sketch of the foreground and background contours, but a composite map that integrates the spatial topological relationship between the foreground and the background, describing the specific layout and correlation relationship between the background and the foreground.

[0057] In the embodiment of the present application, the target foreground vehicle model image is used to represent the foreground object in the target vehicle model wallpaper, and the initial line drawing is used to describe the background contour of the target vehicle model wallpaper. That is to say, the initial line drawing does not include the foreground object. Inputting the target foreground vehicle model image and the initial line drawing into the VLLMs for fusion to obtain a target line drawing including the foreground contour and the background contour, the target line drawing describes the specific layout of the foreground and background in the line drawing and the correlation relationship between the background and the foreground, so that the subsequent generation model can quickly determine the foreground mask area and the background mask area according to the target line drawing, without having to analyze the logical relationship between the target foreground vehicle model image and the initial line drawing in detail to judge their specific layout and correlation relationship in the target vehicle model wallpaper, improving the generation efficiency and accuracy of the target vehicle model wallpaper.

[0058] In an exemplary embodiment, the understanding and image processing capabilities of a vision-language large model can also be utilized to understand the target foreground vehicle model diagram and the initial line drawing diagram, analyze the logic and relevance when fusing the two, and fuse the target foreground vehicle model diagram and the initial line drawing diagram to obtain the target line drawing diagram. The following steps may be included.

[0059] Step 1111: Extract the vehicle contour of the target vehicle according to the target foreground vehicle model diagram.

[0060] The vehicle contour describes the external boundary shape of the vehicle in the target foreground vehicle model diagram, outlining the overall shape of the vehicle.

[0061] Step 1112: Fuse the vehicle contour into the target area of the initial line drawing diagram to obtain the target line drawing diagram.

[0062] The target area is determined by the vision-language large model through analyzing the target foreground vehicle model diagram and the initial line drawing diagram. For example, analyze the logical relationship and mutual association between the vehicle in the target foreground vehicle model diagram and the initial line drawing diagram, and determine the layout and position during fusion through their logical relationship and mutual association to ensure the coordination and physical law compliance after their fusion.

[0063] In the embodiment of the present application, the target foreground vehicle model diagram and the initial line drawing diagram are fused through the vision-language large model, and the vehicle contour extracted from the target foreground vehicle model diagram is fused into the target area of the initial line drawing diagram to obtain the target line drawing diagram. For example, if the target foreground vehicle model diagram is a car and the initial line drawing diagram is a street view diagram, by analyzing the target foreground vehicle model diagram and the initial line drawing diagram, it can be known that the fusion area of the car in the street view diagram is the street. Therefore, the car contour is fused into the area where the street is located in the street view diagram, making the target line drawing diagram logical and reasonable. It should be noted that in other implementation manners, the fusion results of the two are somewhat different due to different inputs, and the above example is only an exemplary embodiment.

[0064] After generating the target vehicle model wallpaper in the above embodiment, it is also necessary to evaluate the effect of the generated target vehicle model wallpaper to confirm whether the generated target vehicle model wallpaper can meet the expected effect. In an exemplary embodiment, after obtaining the target vehicle model wallpaper in step 133 above, the method further includes the following steps.

[0065] Step 1331: Obtain the filled foreground mask area or the filled background mask area.

[0066] Among them, the filled foreground mask region and the filled background mask region are obtained by filling based on the style feature information and the target prompt word through the generation model. Due to the situation of errors or inaccuracies in the generation model, there may be differences between the filled foreground mask region and the foreground mask region before filling, and there may also be differences between the filled background mask region and the background mask region before filling.

[0067] Step 1332: Compare the filled foreground mask region with the foreground mask region, or compare the filled background mask region with the background mask region.

[0068] Among them, the difference between the foreground mask region before and after filling or the background mask region before and after filling is compared to determine the filling effect.

[0069] Step 1333: Output the target vehicle model wallpaper when the difference between the filled foreground mask region and the foreground mask region is less than a first preset threshold, or when the difference between the filled background mask region and the background mask region is less than a second preset threshold.

[0070] In an embodiment of the present application, after the target vehicle model wallpaper is generated, by obtaining the filled foreground mask region or the filled background mask region, and comparing the filled foreground mask region with the foreground mask region before filling, or comparing the filled background mask region with the background mask region before filling, by comparing the difference between the two, and setting the first preset threshold or the second preset threshold, when the difference between the filled foreground mask region and the foreground mask region is less than the first preset threshold, or when the difference between the filled background mask region and the background mask region is less than the second preset threshold, it can be determined that the difference between the foreground mask region before and after filling or the background mask region before and after filling is small, so as to determine that the effect of the currently generated target vehicle model wallpaper is acceptable, and the target vehicle model wallpaper is output for use.

[0071] In an exemplary embodiment, generating the target vehicle model wallpaper and evaluating the effect of the target vehicle model wallpaper as an implementation manner of evaluating the effect of the target vehicle model wallpaper, the embodiment of the present application also provides the following another implementation manner of evaluating the effect of the target vehicle model wallpaper. The method further includes: saving the target vehicle model wallpaper in response to a confirmation instruction for the target vehicle model wallpaper.

[0072] In an embodiment of the present application, after the target vehicle model wallpaper is output, the user can also check the target vehicle model wallpaper. If the user confirms that the effect of the target vehicle model wallpaper meets the expectation and issues a confirmation instruction, then in response to the confirmation instruction for the target vehicle model wallpaper, the target vehicle model wallpaper is saved for subsequent acquisition and use. It should be noted that the above two implementation manners for evaluating the effect of the target vehicle model wallpaper can be used in combination or one of them can be used, and no specific limitation is made here.

[0073] The implementation manner of the vehicle model wallpaper generation method of the present application will be described below through the following specific embodiments.

[0074] Figure 5 Shown is a schematic flowchart of another vehicle model wallpaper generation method provided in an embodiment of the present application in a specific embodiment. Refer to Figure 5 , and this method may include the following steps.

[0075] Step 501, obtain a target style map, a target line drawing, and a target foreground vehicle model map.

[0076] Step 502, extract style feature information from the target style map through the IP-Adapter model; extract structural contour feature information from the target line drawing through the ControlNet model; extract prompt words from the target style map, the target line drawing, and the target foreground vehicle model map through the VLLMs.

[0077] Step 503, perform image generation on the style feature information, the structural contour feature information, the prompt words, and the target foreground vehicle model map through the Inpaint diffusion model to obtain the target vehicle model wallpaper.

[0078] In an embodiment of the present application, the style feature information in the target style map is extracted through the IP-Adapter model, the structural contour feature information in the target line drawing is extracted through the ControlNet model, the prompt words in the target style map, the target line drawing, and the target foreground vehicle model map are comprehensively extracted through the VLLMs, and the style feature information, the structural contour feature information, the prompt words, and the target foreground vehicle model map are used as the input of the Inpaint diffusion model, which greatly enriches the basis for the Inpaint diffusion model to generate the target vehicle model wallpaper. It can not only separately constrain and detail the style, structural contour, and vehicle model of the target vehicle model wallpaper, but also comprehensively combine the features of the three to increase the coordination of generating the target vehicle model wallpaper, so that the Inpaint diffusion model can more clearly define the features of each dimension of the target vehicle model wallpaper, constrain the multi-dimensional features of the target vehicle model wallpaper, and thus accurately generate a high-quality target vehicle model wallpaper.

[0079] Figure 6Shown in another specific embodiment is a schematic diagram of a method for generating a vehicle wallpaper provided by an embodiment of the present application. Refer to Figure 6 , the method may include the following steps.

[0080] Step 601, input a reference style map. Select a style map from the style map database, and this style map represents the style features to be filled. The reference style map provides the style tone for the subsequent generation of the target vehicle wallpaper, such as features in aspects like color, light and shadow, and texture.

[0081] Step 602, input a line drawing. Select a line drawing from the line drawing database. The line drawing provides the basic geometric structure of the background to be filled for the subsequent generation of the target vehicle wallpaper, and clarifies the shape and boundary information.

[0082] Step 603, input a foreground vehicle image and a mask. Input a foreground vehicle image (for example, a car) and a mask. The mask is used to mark the areas that need to be processed, which can be parts on the foreground object that need to be redrawn or repaired, or parts in the background that need to be redrawn or repaired.

[0083] Step 604, VLLMs extract information. VLLMs output appropriate prompt words according to the reference style map, the line drawing, and the foreground vehicle image.

[0084] Step 605, optimize the line drawing. Adjust the line drawing, extract the vehicle contour in the foreground vehicle image, and fuse the vehicle contour into the target area of the line drawing to output a new line drawing. Among them, the target area is determined by VLLMs analyzing the line drawing and the foreground vehicle image.

[0085] Step 606, the IP-Adapter model processes the reference style map. Process the reference style map through the IP-Adapter model to extract the style features therein, such as color, texture, and light and shadow effects, and use the style features as the style guidance for the subsequent generation of the target vehicle wallpaper.

[0086] Step 607, the ControlNet model processes the line drawing. Process the line drawing through the ControlNet model to extract the structural contour features therein, and use the structural contour features as the structural contour guidance for the subsequent generation of the vehicle wallpaper.

[0087] Step 608, the Inpaint diffusion model generates the target vehicle wallpaper. Input the style features, the structural contour features, the prompt words, and the foreground vehicle image into the Inpaint diffusion model. The Inpaint diffusion model intelligently fills the masked area to obtain the target vehicle wallpaper.

[0088] Step 609: Check the generated target vehicle model wallpaper and confirm whether the output effect meets the expectations. If satisfied, save the final result and end the image processing process.

[0089] In an alternative embodiment, in the case where the output effect of the target vehicle model wallpaper is not satisfactory, the input reference style map, line drawing, or foreground vehicle model map can be adjusted or reselected according to the unsatisfactory part of the target vehicle model wallpaper, and a new target vehicle model wallpaper can be generated again through the above steps.

[0090] In the embodiments of the present application, the reference style map, line drawing, and foreground vehicle model map are respectively subjected to feature extraction through different feature extraction models to obtain corresponding style features, structural contour features, and prompt words. The prompt words integrate multi-source information such as the reference style map, line drawing, and foreground vehicle model map, and deep extraction and integration of information are performed through VLLMs, providing rich and accurate basis for the subsequent generation model to generate the target vehicle model wallpaper. The features extracted by the ControlNet model introduced into the Inpaint diffusion model make the geometric structural contour of the generated target vehicle model wallpaper more coordinated, improving the overall visual effect of the target vehicle model wallpaper. The Inpaint diffusion model generates the target vehicle model wallpaper based on the constraints of style features, structural contour features, prompt words, and the foreground vehicle model map, ensuring good connection between the background and the foreground, enhancing the integrity of the target vehicle model wallpaper, and improving the overall quality of the target vehicle model wallpaper.

[0091] It should be noted that for the vehicle model wallpaper generation method provided in the embodiments of the present application, the execution subject can be a vehicle model wallpaper generation device, or a control module in the vehicle model wallpaper generation device for executing the vehicle model wallpaper generation method. In the embodiments of the present application, the vehicle model wallpaper generation method is executed by the vehicle model wallpaper generation device as an example to illustrate the vehicle model wallpaper generation device provided in the embodiments of the present application.

[0092] Figure 7 The structural schematic diagram of a vehicle model wallpaper generation device provided in the embodiments of the present application is shown. As Figure 7 shown, the device 700 includes: an acquisition module 71, a feature extraction module 72, and a generation module 73.

[0093] Among them, the acquisition module 71 is used to acquire a target style map, a target line drawing, and a target foreground vehicle model map; the feature extraction module 72 is used to extract the style feature information corresponding to the target style map through a first feature extraction model; extract the structural contour feature information corresponding to the target line drawing through a second feature extraction model; input the target style map, the target line drawing, and the target foreground vehicle model map into a third feature extraction model for feature extraction to obtain a target prompt; the generation module 73 is used to input the style feature information, the structural contour feature information, the target prompt, and the target foreground vehicle model map into a generation model for image generation to obtain a target vehicle model wallpaper.

[0094] In a specific embodiment, the feature extraction module 72 may be used to extract the structural contour feature information corresponding to the target line drawing through a ControlNet model.

[0095] In a specific embodiment, the feature extraction module 72 may be used to input the target style map, the target line drawing, and the target foreground vehicle model map into a vision-language large model for feature extraction to obtain the target prompt.

[0096] In a specific embodiment, the feature extraction module 72 may be used to extract the style feature information corresponding to the target style map through an IP-Adapter model.

[0097] In a specific embodiment, the feature extraction module 72 may be used to input the style feature information, the structural contour feature information, the target prompt, and the target foreground vehicle model map into an Inpaint diffusion model generation model for image generation to obtain a target vehicle model wallpaper.

[0098] In a specific embodiment, the above-mentioned generation module 73 may be used to determine the foreground mask area of the target vehicle model wallpaper according to the target foreground vehicle model map; determine the background mask area of the target vehicle model wallpaper according to the structural contour feature information; fill the foreground mask area and the background mask area according to the style feature information and the target prompt to obtain the target vehicle model wallpaper.

[0099] In a specific embodiment, the above-mentioned acquisition module 71 may be used to input the target foreground vehicle model map and an initial line drawing into the vision-language large model for fusion to obtain the target line drawing; wherein, the initial line drawing is used to describe the background contour of the target vehicle model wallpaper; the target line drawing is used to describe the background contour and the foreground contour of the target vehicle model wallpaper.

[0100] In a specific embodiment, the above-mentioned acquisition module 71 can be used to extract the vehicle contour of the target vehicle according to the target foreground vehicle map; and fuse the vehicle contour into the target area of the initial line drawing to obtain a target line drawing.

[0101] In a specific embodiment, the above-mentioned device 700 may further include an output module, which is configured to obtain the filled foreground mask area or the filled background mask area; compare the filled foreground mask area with the foreground mask area, or compare the filled background mask area with the background mask area; and output the target vehicle wallpaper when the difference between the filled foreground mask area and the foreground mask area is less than a first preset threshold, or when the difference between the filled background mask area and the background mask area is less than a second preset threshold.

[0102] In a specific embodiment, the above-mentioned device 700 may further include a saving module, which is configured to save the target vehicle wallpaper in response to a confirmation instruction for the target vehicle wallpaper.

[0103] The vehicle wallpaper generation device in the embodiments of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0104] The vehicle wallpaper generation device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0105] The vehicle wallpaper generation device provided in the embodiments of the present application can implement each process implemented in the above Figures 1 to 6 vehicle wallpaper generation method embodiment. To avoid repetition, it will not be elaborated here.

[0106] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, which is used to execute the above-mentioned vehicle wallpaper generation method. Figure 8 FIG. is a schematic structural diagram of an electronic device for implementing various embodiments of the present application. The electronic device may vary greatly due to different configurations or performances, and may include a processor 801, a communication interface 802, a memory 803, and a communication bus 804. Among them, the processor 801, the communication interface 802, and the memory 803 complete mutual communication through the communication bus 804. The processor 801 can call a computer program stored in the memory 803 and executable on the processor 801 to execute the various steps of the above-mentioned vehicle wallpaper generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0107] The above electronic device structure does not limit the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. For example, the input unit may include a Graphics Processing Unit (GPU) and a microphone, and the display unit may be configured with a display panel in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit includes at least one of a touch panel and other input devices. The touch panel is also called a touch screen. Other input devices may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0108] The memory can be used to store software programs and various types of data. The memory may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory can include volatile memory or non-volatile memory, or the memory can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synchlink DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM).

[0109] The processor may include one or more processing units; optionally, the processor integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor either.

[0110] Figure 9 It is a schematic structural diagram of a vehicle provided by an embodiment of the present application.

[0111] Exemplarily, as Figure 9 shown, the vehicle includes: a memory 901 and a processor 902. Among them, an executable program code 9011 is stored in the memory 901, and the processor 902 is used to call and execute the executable program code 9011 to execute each process of implementing the above-mentioned embodiment of the vehicle wallpaper generation method.

[0112] In this embodiment, the vehicle can be divided into functional modules according to the above method examples. For example, it can correspond to each functional module, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0113] In the case of dividing each functional module according to each function, the vehicle may include: an acquisition module, a feature extraction module, etc. It should be noted that all relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be repeated here.

[0114] The vehicle provided in this embodiment is used to execute the above method for generating a vehicle wallpaper, so the same effect as the above implementation method can be achieved.

[0115] In the case of adopting an integrated unit, the vehicle may include a processing module and a storage module. Among them, the processing module can be used to control and manage the actions of the vehicle. The storage module can be used to support the vehicle to execute mutual program codes and data, etc.

[0116] Among them, the processing module can be a processor or a controller, which can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor can also be a combination that realizes computing functions, such as including a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc. The storage module can be a memory.

[0117] This application embodiment also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it realizes each process of the above method embodiment for generating a vehicle wallpaper and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0118] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc.

[0119] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run programs or instructions to implement each process of the above-described embodiment of the vehicle wallpaper generation method and can achieve the same technical effects. To avoid repetition, details are not described here again.

[0120] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-a-chip, etc.

[0121] The embodiments of the present application further provide a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes programs or instructions. When the programs or instructions are executed, they implement each process of the above-described embodiment of the vehicle wallpaper generation method and can achieve the same technical effects. To avoid repetition, details are not described here again.

[0122] Among them, the beneficial effects of the above embodiments can be referred to the beneficial effects in the corresponding methods provided above, and details are not described here again.

[0123] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and simplicity of description, only the above-mentioned division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0124] In the embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0125] In the description of the present disclosure, it should be understood that if terms such as "upper", "lower", "front", "rear", "left", and "right" are used to indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the indicated position or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of the present disclosure.

[0126] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.

[0127] The above are only embodiments of the present disclosure and are not used to limit the present disclosure. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the scope of the claims of the present disclosure.

Claims

1. A method for generating a vehicle model wallpaper, characterized in that: include: Obtain the target style image, target line drawing image, and target foreground vehicle model image; Extracting style feature information corresponding to the target style map through a first feature extraction model; Extracting structural contour feature information corresponding to the target line drawing by a second feature extraction model; Inputting the target style image, the target line drawing image, and the target foreground vehicle model image into a third feature extraction model for feature extraction to obtain a target prompt word; The style feature information, the structure outline feature information, the target prompt word and the target foreground vehicle model image are input into a generation model to generate an image, thereby obtaining a target vehicle model wallpaper.

2. The method according to claim 1, characterized in that The step of inputting the target style map, the target line drawing, and the target foreground vehicle model map into a third feature extraction model for feature extraction to obtain a target prompt word includes: The target style map, the target line drawing, and the target foreground vehicle model map are input into the visual language large model for feature extraction to obtain the target prompt word.

3. The method according to claim 1, characterized in that The step of inputting the style feature information, the structure outline feature information, the target prompt word, and the target foreground vehicle model image into a generation model to generate an image to obtain a target vehicle model wallpaper includes: The style feature information, the structure outline feature information, the target prompt word and the target foreground vehicle model image are input into the Inpaint diffusion model generation model for image generation to obtain the target vehicle model wallpaper.

4. The method according to claim 1, characterized in that The step of inputting the style feature information, the structure outline feature information, the target prompt word, and the target foreground vehicle model image into a generation model to generate an image to obtain a target vehicle model wallpaper includes: Extracting a foreground mask area of ​​the target vehicle model wallpaper from the target foreground vehicle model image; Extracting a background mask area of ​​the target vehicle model wallpaper from the target line drawing according to the structural contour feature information; The foreground mask area and the background mask area are filled according to the style feature information and the target prompt word to obtain the target vehicle model wallpaper.

5. The method according to claim 2, characterized in that: Obtaining the target line drawing includes: Inputting the target foreground vehicle model image and the initial line drawing image into the visual language large model for fusion to obtain the target line drawing image; The initial line drawing is used to describe the background outline of the target vehicle model wallpaper; the target line drawing is used to describe the background outline and foreground outline of the target vehicle model wallpaper.

6. The method according to claim 5, characterized in that The step of inputting the target foreground vehicle model image and the initial line drawing image into the visual language large model for fusion to obtain the target line drawing image includes: Extracting a vehicle outline of the target vehicle model according to the target foreground vehicle model image; The vehicle outline is fused into a target area of ​​the initial line drawing to obtain a target line drawing.

7. The method according to claim 4, characterized in that The method further comprises: Get the filled foreground mask area or the filled background mask area; Comparing the filled foreground mask area with the foreground mask area, or comparing the filled background mask area with the background mask area; When the difference between the filled foreground mask area and the foreground mask area is less than a first preset threshold, or when the difference between the filled background mask area and the background mask area is less than a second preset threshold, the target vehicle model wallpaper is output.

8. A vehicle model wallpaper generating device, characterized in that: include: An acquisition module is used to acquire a target style image, a target line drawing image, and a target foreground vehicle model image; A feature extraction module is used to extract style feature information corresponding to the target style map through a first feature extraction model; extract structural contour feature information corresponding to the target line drawing through a second feature extraction model; input the target style map, the target line drawing, and the target foreground vehicle model map into a third feature extraction model for feature extraction to obtain a target prompt word; The generation module is used to input the style feature information, the structure outline feature information, the target prompt word and the target foreground vehicle model image into the generation model to generate an image and obtain the target vehicle model wallpaper.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the vehicle model wallpaper generation method as described in any one of claims 1 to 7.

10. A vehicle, characterized in that: The vehicle comprises: a memory and a processor, wherein the memory stores executable program code, and the processor is used to call and execute the executable program code, and when the executable program code is executed by the processor, the steps of the vehicle model wallpaper generation method as described in any one of claims 1 to 7 are implemented.