Vehicle machine wallpaper generation method and device, vehicle and storage medium

By processing multi-round voice commands and vehicle system signal combinations using a multi-mode large model, personalized vehicle system wallpapers are generated, solving the problems of lack of personalization and cumbersome operation of vehicle system wallpapers, thus improving user experience and satisfaction.

CN121170076APending Publication Date: 2025-12-19BEIJING AUTOMOBILE RES GENERAL INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511116940.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

In existing technologies, in-vehicle wallpapers lack personalization, are cumbersome to operate, have limited wallpaper selection from the cloud, and rely on user descriptions for voice-generated wallpapers, which cannot be modified, leading to decreased user satisfaction.

Method used

By using multi-modal large-scale model fusion to process users' multi-turn voice commands, text and prompt word combination commands and/or vehicle system signal combination commands, semantic parsing, prompt word extraction and enhancement are performed to generate specific wallpaper operation commands, and finally display the vehicle system wallpaper.

Benefits of technology

It achieves highly intelligent and personalized in-vehicle wallpaper generation, accurately understands user intent, dynamically adjusts wallpaper content, improves user experience and satisfaction, reduces operational burden, and stimulates purchase desire.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170076A_ABST
    Figure CN121170076A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle machine intelligent display, in particular to a vehicle machine wallpaper generation method and device, a vehicle and a storage medium, and the method comprises the steps: receiving a multi-round voice generation instruction of a user, a text and prompt word combination instruction of a vehicle machine and / or a vehicle machine signal combination instruction of the vehicle; inputting the instruction into a pre-constructed multi-mode large model to obtain a semantic analysis result, extracting at least one cue word according to the semantic analysis result, and performing enhancement processing on the cue word to obtain at least one wallpaper action of content generation, content addition, content deletion and content modification of the vehicle machine wallpaper; and based on the multi-mode large model, according to the wallpaper action, generating final display wallpaper of the vehicle-mounted terminal, and displaying the final display wallpaper on the vehicle-mounted terminal. Therefore, the problems that in the related technology, wallpaper preset by a system lacks personalization, the operation of importing pictures from external equipment is tedious, the selection of obtaining the wallpaper from the cloud is limited, and the wallpaper generated by voice depends on user description and cannot be modified are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle-mounted intelligent display technology, and in particular to a method, apparatus, vehicle, and storage medium for generating vehicle-mounted wallpapers. Background Technology

[0002] With the rapid development of automotive intelligence, the role of in-vehicle infotainment systems is becoming increasingly important. In addition to providing drivers with navigation and entertainment functions, the display interface has become a crucial factor influencing user experience. In-vehicle wallpapers, as an important component of the infotainment display interface, can create different visual atmospheres for users and enhance the driving experience.

[0003] In related technologies, the main ways to set wallpapers for in-vehicle systems include using preset fixed wallpapers, importing images from external devices, and retrieving wallpapers from the cloud. With the development of large-scale artificial intelligence models, users can also generate and set wallpapers via voice commands.

[0004] However, among the related technologies, the number of fixed wallpapers is limited and lacks personalization. Importing images from external devices requires users to spend time and effort to find and adapt them to the car screen, which is cumbersome and increases the user's workload. Retrieving wallpapers from the cloud can only provide a limited selection of categories and quantities, and relies on operators to manually change the wallpapers, resulting in a slow update speed. Voice-generated wallpapers rely on user language descriptions and cannot be modified, leading to a decline in user satisfaction with car screen wallpapers, which urgently needs improvement. Summary of the Invention

[0005] This application provides a method, apparatus, vehicle, and storage medium for generating in-vehicle wallpapers to solve problems in related technologies, such as the lack of personalization of system-preset wallpapers, cumbersome operation of importing images from external devices, limited selection of wallpapers obtained from the cloud, and voice-generated wallpapers relying on user descriptions and being unmodifiable.

[0006] The first aspect of this application provides a method for generating in-vehicle infotainment wallpaper, comprising the following steps: receiving a user's multi-turn voice generation command, a text and prompt word combination command from the in-vehicle infotainment system, and / or a vehicle's in-vehicle infotainment system signal combination command; inputting the multi-turn voice generation command, the text and prompt word combination command, and / or the vehicle infotainment system signal combination command into a pre-constructed multi-modal model to obtain semantic parsing results, extracting at least one prompt word based on the semantic parsing results, and enhancing the at least one prompt word to obtain at least one wallpaper action among content generation, content addition, content deletion, and content modification; generating the final display wallpaper of the in-vehicle infotainment system based on the multi-modal model and the at least one wallpaper action, and displaying the final display wallpaper on the in-vehicle infotainment system.

[0007] This application embodiment can use a multi-modal large model to fuse and process user's multi-turn voice commands, text and prompt word combination commands, and / or vehicle system signal combination commands, perform semantic parsing, prompt word extraction and enhancement, generate specific wallpaper operation commands, and finally generate and display vehicle system wallpapers. This achieves highly intelligent and personalized vehicle system wallpaper generation, accurately understands user intentions and dynamically adjusts wallpaper content, significantly improves the user experience and personalization level of the vehicle system, brings convenience and surprise to users, and the real-time generated exquisite wallpapers enhance user satisfaction with the vehicle system and stimulate users' desire to purchase.

[0008] Optionally, in one embodiment of this application, before generating the final display wallpaper of the vehicle system, the method further includes: obtaining the multimodal dataset of the multimodal large model, wherein the multimodal dataset includes wallpapers that meet user needs generated from real vehicle images, landscape images, images of various styles, user-inputted text and prompts, multimodal data in vehicle system signal data, and various information sources; and training a model using the multimodal dataset to construct the multimodal large model.

[0009] This application embodiment can utilize a diverse, multimodal dataset to train and construct a core multimodal large model, thereby ensuring the model's generalization ability and understanding depth. It can effectively integrate multi-source information such as vision, text, and vehicle status to generate high-quality wallpapers that meet user needs and are highly compatible with the in-vehicle environment and user preferences, providing strong basic model support for the entire wallpaper generation process.

[0010] Optionally, in one embodiment of this application, after obtaining the multimodal dataset, the method further includes: cleaning the multimodal dataset to obtain a multimodal dataset that meets preset training conditions; and labeling the images in the multimodal dataset that meet the preset conditions to adjust the model parameters based on the labeled images.

[0011] This application embodiment can perform data cleaning on multimodal datasets to meet training requirements, and annotate key images to guide model parameter adjustment, thereby improving the quality and effectiveness of training data. By cleaning, noise and unqualified data are removed, and by annotation, more accurate learning signals are provided, thus optimizing the model training process and ultimately improving the accuracy, relevance, and visual effect of the wallpapers generated by the model.

[0012] Optionally, in one embodiment of this application, displaying the final display wallpaper on the vehicle infotainment system includes: obtaining the generation type of the final display wallpaper; if the generation type is the generation type corresponding to the multi-turn voice generation instruction or the text and prompt word combination instruction, receiving a voice instruction or a manual setting instruction from the user, and displaying the final display wallpaper according to the voice instruction or the manual setting instruction; if the generation type is the generation type corresponding to the vehicle infotainment system signal combination instruction, controlling the vehicle infotainment system to display the final display wallpaper.

[0013] This application embodiment can adopt different display triggering mechanisms according to the source type of the wallpaper generation instruction. User-initiated wallpapers require user confirmation before display, while those generated by vehicle signal instructions are displayed automatically. This provides a flexible and scenario-appropriate wallpaper display strategy. For user-initiated modifications, the user is given the final confirmation right, improving controllability and satisfaction. Automatic updates triggered by vehicle status enable seamless and real-time wallpaper switching, enhancing the immersive environment and intelligent experience.

[0014] Optionally, in one embodiment of this application, generating the final display wallpaper of the vehicle system based on the at least one wallpaper action includes: determining at least one of an image address link, an instruction to generate an image, a prompt word and an enhanced prompt word, and an image name based on the at least one wallpaper action; and generating the final display wallpaper based on the at least one of the prompt words.

[0015] This application embodiment can determine specific generation elements based on wallpaper actions, including image address links, generation instructions, prompts and enhanced prompts, image names, etc., and generate the final wallpaper. It provides a clear, flexible and operable wallpaper generation implementation path, enhances the system's adaptability and functional richness, and meets users' diverse wallpaper customization needs.

[0016] Optionally, in one embodiment of this application, the method further includes: when the vehicle is powered on, acquiring at least one relevant information of the vehicle system for each preset duration; determining whether the at least one relevant information satisfies a preset change condition with the previous relevant information; and if the preset change condition is satisfied, generating the vehicle system signal combination command.

[0017] This application embodiment can periodically monitor relevant information of the vehicle's infotainment system after the vehicle is powered on. When a change in the relevant information is detected, it automatically triggers the generation of a combination of vehicle infotainment system signals, thereby realizing intelligent linkage between the vehicle's wallpaper and the vehicle's status, environment, and journey. The wallpaper can be automatically and timely updated according to the real-time changing vehicle context without the need for manual operation by the user, which greatly improves the system's proactive service capabilities and contextualized experience.

[0018] A second aspect of this application provides an apparatus for generating in-vehicle infotainment wallpapers, comprising: a receiving module for receiving a user's multi-turn voice generation command, a text and prompt word combination command from the in-vehicle infotainment system, and / or a vehicle infotainment system signal combination command; an extraction module for inputting the multi-turn voice generation command, the text and prompt word combination command, and / or the vehicle infotainment system signal combination command into a pre-constructed multi-modal model to obtain a semantic parsing result, extracting at least one prompt word based on the semantic parsing result, and enhancing the at least one prompt word to obtain at least one wallpaper action among content generation, content addition, content deletion, and content modification; and a generation module for generating the final display wallpaper of the in-vehicle infotainment system based on the multi-modal model and the at least one wallpaper action, and displaying the final display wallpaper on the in-vehicle infotainment system.

[0019] This application embodiment can use a multi-modal large model to fuse and process user's multi-turn voice commands, text and prompt word combination commands, and / or vehicle system signal combination commands, perform semantic parsing, prompt word extraction and enhancement, generate specific wallpaper operation commands, and finally generate and display vehicle system wallpapers. This achieves highly intelligent and personalized vehicle system wallpaper generation, accurately understands user intentions and dynamically adjusts wallpaper content, significantly improves the user experience and personalization level of the vehicle system, brings convenience and surprise to users, and the real-time generated exquisite wallpapers enhance user satisfaction with the vehicle system and stimulate users' desire to purchase.

[0020] Optionally, in one embodiment of this application, it further includes: an acquisition module, configured to acquire the multimodal dataset of the multimodal large model before generating the final display wallpaper of the vehicle system, wherein the multimodal dataset includes wallpapers that meet user needs generated from real vehicle images, landscape images, images of various styles, text and prompts input by the user, multimodal data in vehicle system signal data, and various information sources; and a construction module, configured to train a model using the multimodal dataset to construct the multimodal large model before generating the final display wallpaper of the vehicle system.

[0021] This application embodiment can utilize a diverse, multimodal dataset to train and construct a core multimodal large model, thereby ensuring the model's generalization ability and understanding depth. It can effectively integrate multi-source information such as vision, text, and vehicle status to generate high-quality wallpapers that meet user needs and are highly compatible with the in-vehicle environment and user preferences, providing strong basic model support for the entire wallpaper generation process.

[0022] Optionally, in one embodiment of this application, it further includes: a cleaning module, used to clean the multimodal dataset after acquiring it to obtain a multimodal dataset that meets preset training conditions; and an annotation module, used to annotate the images in the multimodal dataset that meet preset conditions after acquiring it, so as to adjust the model parameters based on the annotated images.

[0023] This application embodiment can perform data cleaning on multimodal datasets to meet training requirements, and annotate key images to guide model parameter adjustment, thereby improving the quality and effectiveness of training data. By cleaning, noise and unqualified data are removed, and by annotation, more accurate learning signals are provided, thus optimizing the model training process and ultimately improving the accuracy, relevance, and visual effect of the wallpapers generated by the model.

[0024] Optionally, in one embodiment of this application, the generation module includes: a type acquisition unit, configured to acquire the generation type of the final displayed wallpaper; a display unit, configured to receive a voice command or a manual setting command from the user when the generation type is the generation type corresponding to the multi-turn voice generation command or the text and prompt word combination command, and display the final displayed wallpaper according to the voice command or manual setting command; and a control unit, configured to control the vehicle system to display the final displayed wallpaper when the generation type is the generation type corresponding to the vehicle system signal combination command.

[0025] This application embodiment can adopt different display triggering mechanisms according to the source type of the wallpaper generation instruction. User-initiated wallpapers require user confirmation before display, while those generated by vehicle signal instructions are displayed automatically. This provides a flexible and scenario-appropriate wallpaper display strategy. For user-initiated modifications, the user is given the final confirmation right, improving controllability and satisfaction. Automatic updates triggered by vehicle status enable seamless and real-time wallpaper switching, enhancing the immersive environment and intelligent experience.

[0026] Optionally, in one embodiment of this application, the generation module includes: a determining unit, configured to determine at least one of an image address link, an instruction to generate an image, a prompt word and an enhanced prompt word, and an image name based on the at least one wallpaper action; and a wallpaper generation unit, configured to generate the final display wallpaper based on the at least one of the above-mentioned actions.

[0027] This application embodiment can determine specific generation elements based on wallpaper actions, including image address links, generation instructions, prompts and enhanced prompts, image names, etc., and generate the final wallpaper. It provides a clear, flexible and operable wallpaper generation implementation path, enhances the system's adaptability and functional richness, and meets users' diverse wallpaper customization needs.

[0028] Optionally, in one embodiment of this application, it further includes: an information acquisition module, configured to acquire at least one relevant information of the vehicle system for each preset duration when the vehicle is powered on; a judgment module, configured to judge whether the at least one relevant information satisfies a preset change condition with the previous relevant information; and an instruction generation module, configured to generate a vehicle system signal combination instruction if the preset change condition is satisfied.

[0029] This application embodiment can periodically monitor relevant information of the vehicle's infotainment system after the vehicle is powered on. When a change in the relevant information is detected, it automatically triggers the generation of a combination of vehicle infotainment system signals, thereby realizing intelligent linkage between the vehicle's wallpaper and the vehicle's status, environment, and journey. The wallpaper can be automatically and timely updated according to the real-time changing vehicle context without the need for manual operation by the user, which greatly improves the system's proactive service capabilities and contextualized experience.

[0030] A third aspect of this application provides a vehicle, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for generating vehicle wallpaper as described in the above embodiments.

[0031] A fourth aspect of this application provides a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for generating in-vehicle wallpapers.

[0032] A fifth aspect of this application provides a computer program product that stores a computer program that, when executed by a processor, implements the above-described method for generating in-vehicle wallpapers.

[0033] This application embodiment utilizes a multi-modal large model to fuse and process multi-turn voice commands, text and prompt word combinations, and / or vehicle system signal combinations. It performs semantic parsing, prompt word extraction and enhancement to generate specific wallpaper operation commands, ultimately generating and displaying the vehicle system wallpaper. This achieves highly intelligent and personalized vehicle system wallpaper generation, accurately understanding user intent and dynamically adjusting wallpaper content. This significantly improves the user experience and personalization level of the vehicle system, bringing convenience and a sense of surprise to users. The real-time generated, beautiful wallpapers enhance user satisfaction with the vehicle system and stimulate their desire to purchase. Therefore, it solves problems in related technologies such as the lack of personalization in pre-set wallpapers, cumbersome operations for importing images from external devices, limited wallpaper selection from the cloud, and the reliance on user descriptions for voice-generated wallpapers that cannot be modified.

[0034] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0035] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0036] Figure 1 This is a flowchart illustrating a method for generating in-vehicle wallpaper according to an embodiment of this application;

[0037] Figure 2 This is an architecture diagram of a method for generating in-vehicle wallpapers according to an embodiment of this application;

[0038] Figure 3 This is a flowchart illustrating a method for generating in-vehicle wallpapers according to an embodiment of this application.

[0039] Figure 4 A flowchart illustrating the process of generating wallpaper from vehicle-mounted system signals according to an embodiment of this application;

[0040] Figure 5 This is a flowchart of a multi-turn voice wallpaper generation process according to one embodiment of this application;

[0041] Figure 6 A flowchart for generating wallpapers by combining text and prompts according to an embodiment of this application;

[0042] Figure 7 This is a flowchart illustrating the training process for custom vehicle image data according to one embodiment of this application;

[0043] Figure 8 This is a schematic diagram of a device for generating in-vehicle wallpapers according to an embodiment of this application;

[0044] Figure 9 This is a structural schematic diagram of a vehicle provided according to an embodiment of this application. Detailed Implementation

[0045] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0046] The following description, with reference to the accompanying drawings, outlines a method, apparatus, vehicle, and storage medium for generating in-vehicle wallpapers according to embodiments of this application. Addressing the issues mentioned in the background art, such as the lack of personalization in pre-set wallpapers, cumbersome processes for importing images from external devices, limited wallpaper selection from the cloud, and the dependence of voice-generated wallpapers on user descriptions that cannot be modified, this application provides a method for generating in-vehicle wallpapers. This method utilizes a multi-modal large model to fuse and process multi-turn voice commands, text and prompt word combinations, and / or in-vehicle signal combinations. Semantic parsing, prompt word extraction and enhancement are performed to generate specific wallpaper operation instructions, ultimately generating and displaying the in-vehicle wallpaper. This achieves highly intelligent and personalized in-vehicle wallpaper generation, accurately understanding user intent and dynamically adjusting wallpaper content. This significantly improves the user experience and personalization level of the in-vehicle system, providing convenience and a sense of surprise. The real-time generated, beautiful wallpapers enhance user satisfaction with the in-vehicle system and stimulate purchasing desire. Therefore, this method solves the problems in related technologies, such as the lack of personalization in pre-set wallpapers, cumbersome processes for importing images from external devices, limited wallpaper selection from the cloud, and the dependence of voice-generated wallpapers on user descriptions that cannot be modified.

[0047] Specifically, Figure 1 This is a schematic flowchart illustrating a method for generating in-vehicle wallpaper according to an embodiment of this application.

[0048] like Figure 1 As shown, the method for generating this car infotainment wallpaper includes the following steps:

[0049] In step S101, the system receives multi-turn voice generation commands from the user, text and prompt word combination commands from the vehicle's infotainment system, and / or vehicle's infotainment system signal combination commands.

[0050] It is understood that the multi-turn voice generation command in this application embodiment can be the wallpaper generation request conveyed by the user through multiple voice interactions, such as first requesting "change to a landscape wallpaper" and then adding "with snow-capped mountains"; the text and prompt word combination command can be the combination of text description and style, element and other limiting words entered by the user on the vehicle interface, such as "sunset + oil painting style"; the vehicle signal combination command can be the automatic generation command formed by the vehicle system obtaining the city information (obtained after user authorization), weather information and festival and solar term information, such as "sunny day + Valentine's Day" triggering the request for a soft and romantic wallpaper.

[0051] In actual implementation, this embodiment can collect user voice in real time through the vehicle's built-in microphone, convert it into a text sequence using speech-to-text technology, and if the user is not satisfied with the wallpaper content generated by the first voice, they can add, delete, or modify elements, colors, styles, etc., of the obtained wallpaper through multiple rounds of voice commands, thereby receiving multiple rounds of voice generation commands. The system receives user text input information and prompt word information through the vehicle's touchscreen or input panel, obtaining combined text and prompt word commands from the vehicle. Text input information refers to information manually entered by the user, who can input the desired content into the corresponding text box; prompt word information provides users with image inspiration, while style information provides users with multiple style choices, such as Van Gogh style, cyberpunk style, paper-cut style, etc. Both prompt word information and style information are stored in the cloud and can be updated online.

[0052] Furthermore, the vehicle's infotainment system (V2S) signals acquire data from vehicle sensors, such as GPS (Global Positioning System) positioning, light sensor readings, and weather information, via the CAN (Controller Area Network) bus. It obtains relevant V2S signals in real time and determines whether this information has changed. When any information changes, the V2S receives the information and integrates it into a combined V2S signal command. If a radio station or music is playing when the information changes, the station name or music track information is simultaneously integrated into the combined V2S signal command. This embodiment can receive multi-turn voice commands from the user, combined text and prompts from the V2S, or a combination of the user's multi-turn voice commands, the V2S's text and prompts, and the vehicle's V2S signal command, or simply receive the vehicle's V2S signal command. For example, if a user repeatedly says "I want a live wallpaper" and "It should have wave elements," the system will integrate these two voice commands into a multi-turn voice command. When the V2S detects "rain," it will automatically combine them into a combined V2S signal command.

[0053] The embodiments of this application can support multiple types of command input, covering user-initiated customization and automatic adaptation to vehicle status scenarios, breaking the limitations of a single interaction method and improving the flexibility and scene adaptability of vehicle wallpaper generation.

[0054] Optionally, in one embodiment of this application, the method further includes: when the vehicle is powered on, acquiring at least one relevant information of the vehicle system for each preset duration; determining whether the at least one relevant information satisfies a preset change condition with the previous relevant information; and if the preset change condition is satisfied, generating a vehicle system signal combination command.

[0055] It is understood that the relevant information of the vehicle system in this embodiment may be city information, city weather information, festival and solar term information, etc. The preset duration may be 30 minutes, and the preset change condition may be that the city weather information changes and rainfall begins. The preset duration and preset change condition may be set by those skilled in the art according to the actual situation, and no specific restrictions are made here.

[0056] For example, in this embodiment of the application, when the vehicle's infotainment system is powered on, it retrieves city information, city weather information, and holiday / solar term information every 30 minutes and compares this information with the information retrieved by the vehicle's infotainment system the last time it was powered on to determine whether preset change conditions are met. If the preset change conditions are met, a vehicle's infotainment system signal combination command is generated. If the vehicle's infotainment system is powered off for more than 30 minutes, it compares this information with the last information retrieved before the vehicle's infotainment system was powered off the next time it is powered on to determine whether any changes have occurred.

[0057] The embodiments of this application can monitor changes in the vehicle's status in real time and automatically generate signal commands, enabling the wallpaper to dynamically adapt to the vehicle scene, reducing manual operation by the user, and improving the intelligence and automation level of the vehicle system.

[0058] In step S102, the multi-turn voice generation instruction, the text and prompt word combination instruction and / or the vehicle system signal combination instruction are input into the pre-constructed multi-mode large model to obtain the semantic parsing result. At least one prompt word is extracted based on the semantic parsing result, and at least one prompt word is enhanced to obtain at least one wallpaper action among the vehicle system wallpaper content generation, content addition, content deletion and content modification.

[0059] It is understood that the prompt words in this application embodiment can be keywords or phrases that provide users with inspiration for creating images during the generation of in-vehicle wallpapers, and are used to guide the multi-model large model to generate wallpapers with specific themes and styles.

[0060] In actual implementation, this application embodiment can input received multi-turn voice generation instructions, text and prompt word combination instructions, and / or vehicle signal combination instructions into a multi-modal large model for semantic parsing to obtain semantic parsing results: for multi-turn voice generation instructions, multi-turn information can be integrated through context association processing; for text and prompt word combination instructions, keywords and limiting words can be directly extracted; for vehicle signal combination instructions, they can be mapped to corresponding scene descriptions. Subsequently, core prompt words such as "snow mountain" and "oil painting" are extracted from the semantic parsing results and enhanced by combining the model's built-in style library and element library. The prompt words are extracted and expanded to make them richer and more targeted, so that the multi-modal large model can better understand user needs and generate more suitable vehicle wallpapers, such as adding "high saturation" and "obvious brushstrokes". Finally, wallpaper actions such as content generation, content addition, content deletion, and content modification are output. For example, in response to the instruction "change the snow mountain wallpaper to a traditional Chinese ink painting style", the model analyzes and extracts "snow mountain wallpaper" and "traditional Chinese ink painting style", enhances it into "snow mountain subject + ink painting brushstrokes + blank space artistic conception", and outputs the wallpaper action with modified content.

[0061] This application embodiment can transform vague and fragmented instructions into precise and executable operation directions through semantic parsing and prompt word enhancement of multi-modal large models, solve the problem of insufficient understanding of complex instructions in traditional systems, and improve the matching degree between wallpaper generation and user needs.

[0062] In step S103, based on the multi-mode large model, the final display wallpaper of the vehicle system is generated according to at least one wallpaper action, and the final display wallpaper is displayed on the vehicle system.

[0063] In actual implementation, this application embodiment can generate the final display wallpaper for the vehicle's infotainment system based on a multi-model large model, according to wallpaper actions such as content generation, content addition, content deletion, and content modification. For example, if the action is content generation, the model directly generates a completely new image based on enhanced prompts; if the action is content modification, elements are added or removed, or the style is changed, on the existing wallpaper. After the generated wallpaper is adapted to the resolution and display ratio by the vehicle's infotainment system, it is transmitted to the vehicle's infotainment system for final display.

[0064] The embodiments of this application can realize the rapid conversion from user needs to finished wallpapers, reduce intermediate operation steps, generate wallpapers in real time to meet in-depth customization needs, thereby improving the efficiency and intuitiveness of wallpaper generation, optimizing the display layer to ensure visual smoothness in driving scenarios, and enhancing the user experience.

[0065] Optionally, in one embodiment of this application, before generating the final display wallpaper of the vehicle system, the method further includes: obtaining a multimodal dataset of a multimodal large model, wherein the multimodal dataset includes wallpapers that meet user needs generated from real vehicle images, landscape images, images of various styles, user-inputted text and prompts, multimodal data in vehicle system signal data, and various information sources; and training a model using the multimodal dataset to construct a multimodal large model.

[0066] In practical implementation, this application embodiment can construct a large-scale multimodal dataset in the refined generation of in-vehicle wallpapers using a multimodal large model. This dataset organically integrates data from different modalities, including real vehicle images, landscape images, images of various styles, user-inputted text and prompts, and in-vehicle signal data. The data is then trained and fine-tuned using this large-scale multimodal dataset, enabling the multimodal large model to master various features and patterns. When specific tasks or new data emerge, the model parameters are adjusted and optimized to improve the model's performance and accuracy in specific scenarios such as in-vehicle wallpaper generation. Furthermore, by integrating multiple information sources, highly realistic wallpapers that meet user needs are generated.

[0067] For example, this application embodiment can organize a real vehicle image dataset. First, it needs to collect real vehicle images covering different car models, angles, lighting conditions, and various environmental backgrounds. Then, the real vehicle images are preprocessed, including processing the car model outline, processing the car logo, and deleting license plate text. Image edge detection algorithms are used to accurately extract the car model outline, ensuring clear lines that reflect the vehicle's shape characteristics. Image recognition and segmentation techniques are used to extract and optimize the car logo area separately, highlighting details and brand features. License plate text is deleted to protect user privacy and improve image aesthetics. Finally, the processed car model images are organized into a dataset, categorized and labeled according to factors such as car model, environment, and lighting, to build a structured dataset. For example, the real vehicle image dataset can be divided into sub-datasets for different brand car models, and each sub-dataset contains image samples under different lighting conditions (e.g., daytime, nighttime, cloudy, sunny) and environmental backgrounds (e.g., city streets, rural roads, mountainous areas).

[0068] Furthermore, this application embodiment can generate customized vehicle model wallpapers. Based on a dataset of real vehicle images, model training, fine-tuning, and optimization are performed to ensure that after the vehicle's infotainment system receives the instruction to generate a vehicle model wallpaper, it can reproduce the vehicle's shape and brand logo to the greatest extent possible, ensuring the adaptability of the vehicle to its surrounding environment, lighting, and other elements, and allowing users to customize the vehicle's color. During model training, key parts of the vehicle, such as the logo and body lines, are emphasized to improve the accuracy of the vehicle's shape reproduction. Simultaneously, the model learns the adaptability of the vehicle to its surrounding environment, lighting, and other elements. For example, by learning the light and shadow changes of the vehicle under different lighting conditions, as well as the color and texture adjustment methods of the vehicle in different environmental backgrounds, the generated in-vehicle wallpaper can blend naturally in various scenarios.

[0069] In some embodiments, in addition to collecting real vehicle images, customized vehicle wallpapers can also be generated using 3D model reconstruction technology, or by matching vehicle models to templates. Using 3D model reconstruction technology involves acquiring 3D vehicle data through multi-angle photography or laser scanning to construct a 3D model, or directly using a multi-million-level industrial 3D model of the vehicle. Custom operations are then performed on the model in 3D modeling software (such as changing the body color or adding decorations), and the data from the 3D model is rendered into a 2D wallpaper image. Alternatively, matching vehicle models to templates involves pre-establishing a template library for different vehicle models. When a user selects a customized vehicle wallpaper, the system uses image recognition technology to match the user's vehicle image with the template library. After determining the vehicle model, the wallpaper is generated based on the basic framework of the template and the user's custom requirements (such as color and background).

[0070] The embodiments of this application can construct a multimodal dataset and train a model, enabling the model to understand the relationship between text, signals and images, providing core technical support for accurately parsing instructions and generating wallpapers that meet requirements, and improving the model's generalization ability.

[0071] Optionally, in one embodiment of this application, after obtaining the multimodal dataset, the method further includes: cleaning the multimodal dataset to obtain a multimodal dataset that meets preset training conditions; and labeling the images in the multimodal dataset that meet the preset conditions to adjust the model parameters based on the labeled images.

[0072] It is understood that data cleaning in this application embodiment can be the removal of invalid information from the dataset; image annotation can be the marking of key elements and style features in the image; preset training conditions can be the basic quality standards that the multimodal dataset needs to meet, and preset conditions can be that the image has clear main elements. The preset training conditions and preset conditions can be set by those skilled in the art according to the actual situation, and no specific restrictions are made here.

[0073] For example, in this embodiment, after acquiring the multimodal dataset, model fine-tuning and optimization can be performed. The dataset is continuously cleaned to remove low-quality or duplicate photos, resulting in a multimodal dataset that meets preset training conditions. Simultaneously, images in the multimodal dataset that meet the preset conditions are labeled, and model parameters are adjusted based on test feedback to improve image generation quality. Model parameters are then adjusted based on the labeled images. For instance, if a user repeatedly selects a "starry sky + dark tone" wallpaper, the model will strengthen the association weight between the "starry sky" element and the "low saturation" and "high contrast" parameters, while reducing the generation probability of the rarely chosen "bright cartoon" style. Furthermore, by setting a user preference tag library, user-inputted text commands such as "minimalist style" and "avoid dense patterns" can be converted into parameters that the model can recognize, and these parameters are prioritized during the generation process.

[0074] The embodiments of this application can perform targeted data cleaning, fine-grained annotation, and model fine-tuning based on user feedback, enabling the multi-model large model to continuously learn and strengthen the user's personalized preferences, reduce the generation of wallpapers that do not meet the needs, thereby improving the quality of model output and ensuring that the output car wallpapers are highly consistent with the user's needs in terms of elements, style, and scene adaptability, significantly enhancing the personalized experience.

[0075] Optionally, in one embodiment of this application, displaying the final display wallpaper on the vehicle infotainment system includes: obtaining the generation type of the final display wallpaper; if the generation type is a multi-turn voice generation instruction or a text and prompt word combination instruction, receiving a voice instruction or a manual setting instruction from the user, and displaying the final display wallpaper according to the voice instruction or the manual setting instruction; if the generation type is a vehicle infotainment system signal combination instruction, controlling the vehicle infotainment system to display the final display wallpaper.

[0076] For example, this application embodiment can obtain the generation type of the final displayed wallpaper. The wallpaper returned to the vehicle's infotainment system has two setting methods: If the generation type corresponds to a multi-turn voice generation command or a combination of text and prompt words, it can be applied as the vehicle's infotainment system wallpaper after user voice command or manual setting. For example, after generation, a preview interface pops up, requiring user voice confirmation or manual click to display the final wallpaper. If the generation type corresponds to a combination of vehicle signal commands, it is directly applied as the vehicle's infotainment system wallpaper. For example, as the vehicle travels from sunset to night, the wallpaper gradually changes to dark mode, surprising the user. After being applied as the vehicle's infotainment system wallpaper, the wallpaper is automatically saved in the wallpaper store for easy searching or regeneration by the user.

[0077] This application embodiment can adopt a differentiated display strategy by distinguishing the generation type, which can not only ensure the user's control over the actively customized results, but also realize the seamless switching of the system to automatically adapt to the scenario, thus balancing personalized needs and driving safety.

[0078] Optionally, in one embodiment of this application, generating the final display wallpaper of the vehicle system based on at least one wallpaper action includes: determining at least one of the following based on at least one wallpaper action: an image address link, an instruction to generate an image, a prompt word and an enhanced prompt word, and an image name; and generating the final display wallpaper based on at least one of the following.

[0079] It is understood that the image address link in this application embodiment can be the access path of the wallpaper already stored locally in the vehicle or in the cloud.

[0080] In actual implementation, this embodiment can return the generated wallpaper to the vehicle's infotainment system. Based on at least one wallpaper action, the returned content is determined. The returned content includes an image address link, an instruction to generate the image, prompts and enhanced prompts, and an image name. The final displayed wallpaper is generated based on the returned content. The image address link is used by the vehicle to download the image generated by the multi-model cloud. The prompts and enhanced prompts are used to record previously generated wallpapers when the user modifies the wallpaper. The image name is the wallpaper name returned to the vehicle's infotainment system. When the user's instruction input does not exceed 8 Chinese characters or 16 characters, the wallpaper name is the content of the user's instruction input. When the user's instruction input exceeds 8 Chinese characters or 16 characters, the wallpaper name is generated by the large model.

[0081] The embodiments of this application can clearly define the core elements of wallpaper generation, transform abstract actions into specific executable steps, ensure that the generation process is controllable and traceable, and improve the accuracy and efficiency of wallpaper generation.

[0082] In some embodiments, in addition to multi-turn voice generation commands, text and prompt word generation commands, and vehicle system signal combination generation commands, gesture recognition-based wallpaper generation commands and user driving habit-based wallpaper generation commands can also be introduced. Gesture recognition-based wallpaper generation commands introduce gesture recognition and facial expression recognition as trigger methods for wallpaper generation. Users express their wallpaper needs by making specific gestures or corresponding facial expressions in front of the vehicle system screen. After recognizing the gesture or expression, the vehicle system generates the corresponding wallpaper based on the preset correspondence between the gesture or expression and wallpaper features. For example, a user drawing a circle gesture indicates the generation of a wallpaper with many circular elements; a user laughing happily generates a cheerful wallpaper. User driving habit-based wallpaper generation automatically generates commands based on user driving habit data (such as driving speed, driving time, driving route, etc.). For example, generating a simple and clear wallpaper when driving at high speeds to avoid distraction; generating low-brightness, soft-colored prompt words when driving at night to reduce visual fatigue. Through analysis of users' long-term driving habits, the model can automatically adjust the wallpaper style and features to provide a more suitable visual experience for the driving context.

[0083] In some embodiments, dynamic wallpaper generation can also be implemented, including dynamic processing of elements, composition and output of dynamic wallpapers, etc. Element dynamic processing can dynamically construct user-input elements; for example, for a request for "a waterfall cascading in a forest," the model uses algorithms to simulate the dynamic process of water flowing down from a height, including the speed of the water flow, changes in water volume, and splashing effects. Dynamic wallpaper composition and output involves the model adjusting the proportions, positions, and hierarchical relationships between various elements to make the entire image look natural and harmonious. Finally, the generated dynamic wallpaper is output to the vehicle's infotainment system in a format suitable for in-vehicle playback (such as MP4 video format or animated image sequence format), and the system displays it according to the user's settings.

[0084] Specifically, it can be combined with Figures 2 to 7 As shown, the working principle of the in-vehicle wallpaper generation method in this application embodiment is explained in detail with a specific example.

[0085] Figure 2 This is an architecture diagram of a method for generating in-vehicle wallpapers according to an embodiment of this application.

[0086] like Figure 2As shown, the embodiments of this application may include: acquiring data from two sources: vehicle system signal input and user information input. Vehicle system signals include weather information, city information, and holiday / solar term information; user information includes voice input, text input, and style input. The data is aggregated to the cockpit domain controller and simultaneously processed in the cloud. Cloud processing includes cloud output, a multi-modal large model, and cloud storage. Cloud transmission ensures real-time data interaction between the vehicle system and the cloud, addressing the issue of insufficient local computing power in the vehicle system. The multi-modal large model is used for semantic parsing, extracting and enhancing prompt words to obtain vehicle system wallpaper actions, enabling personalized customization. Cloud storage stores multi-modal datasets, including real vehicle images, style materials, and user history commands, providing data support for model training and rapid retrieval. The wallpaper is output to the vehicle system desktop wallpaper and wallpaper store. After being applied as a vehicle system wallpaper, the wallpaper is automatically saved in the wallpaper store for easy user searching or regeneration.

[0087] Figure 3 This is a flowchart illustrating a method for generating in-vehicle wallpapers according to an embodiment of this application.

[0088] like Figure 3 As shown, embodiments of this application may include the following steps:

[0089] Step S301: Input the command to generate wallpaper.

[0090] Step S302: Semantic analysis of large artificial intelligence models.

[0091] Step S303: Extraction, expansion, and enhancement of prompt words for the AI ​​multi-modal large model.

[0092] Step S304: The AI ​​multi-modal large model generates images based on the enhanced prompt words.

[0093] Step S305: Return to the vehicle's infotainment system.

[0094] Step S306: Apply as car infotainment wallpaper.

[0095] Figure 4 A flowchart illustrating the process of generating wallpaper from vehicle-mounted signals according to an embodiment of this application.

[0096] like Figure 4 As shown, embodiments of this application may include the following steps:

[0097] Step S401: Input vehicle-mounted signal combination.

[0098] Step S402: Determine whether the vehicle-mounted signal combination has changed. If yes, proceed to step S403; otherwise, end.

[0099] Step S403: Generate images using an AI multi-modal large model.

[0100] Step S404: Return to the vehicle's infotainment system.

[0101] Step S405: Apply as car infotainment wallpaper.

[0102] Figure 5 This is a flowchart of a multi-turn voice wallpaper generation process according to one embodiment of this application.

[0103] like Figure 5 As shown, embodiments of this application may include the following steps:

[0104] Step S501: User voice input.

[0105] Step S502: Generate images using an AI multi-modal large model.

[0106] Step S503: Return to the vehicle's infotainment system.

[0107] Step S504: The user judges whether the image is satisfactory. If yes, proceed to step S505; otherwise, proceed to step S506.

[0108] Step S505: Apply as car infotainment wallpaper.

[0109] Step S506: The user modifies the wallpaper voice input.

[0110] Figure 6 A flowchart for generating wallpapers by combining text and prompts according to an embodiment of this application is provided.

[0111] like Figure 6 As shown, embodiments of this application may include the following steps:

[0112] Step S601: User text output.

[0113] Step S602: Select and input the cloud-based prompt word.

[0114] Step S603: Input the cloud style selection.

[0115] Step S604: Generate images using an AI multi-modal large model.

[0116] Step S605: Return to the vehicle's infotainment system.

[0117] Step S606: Apply as car infotainment wallpaper.

[0118] Figure 7 This is a flowchart illustrating the training process for custom vehicle image data according to one embodiment of this application.

[0119] like Figure 7 As shown, embodiments of this application may include the following steps:

[0120] Step S701: Collect real vehicle images.

[0121] Step S702: Processing the line drawing of the vehicle model framework.

[0122] Step S703: Car logo processing.

[0123] Step S704: License plate processing.

[0124] Step S705: Organize the processed vehicle model image dataset.

[0125] Step S706: Classify and label data to create a structured dataset.

[0126] Step S707: Adapt different styles of vehicle images.

[0127] The in-vehicle wallpaper generation method proposed in this application can utilize a multi-modal large model to fuse and process multi-turn voice commands, text and prompt word combination commands, and / or in-vehicle signal combination commands. Semantic parsing, prompt word extraction and enhancement are performed to generate specific wallpaper operation commands, ultimately generating and displaying the in-vehicle wallpaper. This achieves highly intelligent and personalized in-vehicle wallpaper generation, accurately understanding user intent and dynamically adjusting wallpaper content, significantly improving the user experience and personalization level of the in-vehicle system. It brings convenience and a sense of surprise to users, and the real-time generated, beautiful wallpapers enhance user satisfaction with the in-vehicle system, stimulating their desire to purchase. Therefore, it solves the problems in related technologies, such as the lack of personalization in system-preset wallpapers, the cumbersome operation of importing images from external devices, the limited selection of wallpapers obtained from the cloud, and the reliance on user descriptions for voice-generated wallpapers that cannot be modified.

[0128] Next, the apparatus for generating in-vehicle wallpapers according to an embodiment of this application is described with reference to the accompanying drawings.

[0129] Figure 8 This is a schematic diagram of the structure of the in-vehicle wallpaper generation device according to an embodiment of this application.

[0130] like Figure 8 As shown, the in-vehicle wallpaper generation device 10 includes: a receiving module 100, an extraction module 200, and a generation module 300.

[0131] The receiving module 100 is used to receive multi-turn voice generation commands from users, text and prompt word combination commands from the vehicle's infotainment system, and / or vehicle infotainment system signal combination commands from the vehicle.

[0132] The extraction module 200 is used to input multi-turn voice generation instructions, text and prompt word combination instructions and / or vehicle signal combination instructions into a pre-constructed multi-mode large model to obtain semantic parsing results, extract at least one prompt word based on the semantic parsing results, and perform enhancement processing on at least one prompt word to obtain at least one wallpaper action among vehicle wallpaper content generation, content addition, content deletion and content modification.

[0133] The generation module 300 is used to generate the final display wallpaper of the vehicle system based on the multi-model large model and according to at least one wallpaper action, and to display the final display wallpaper on the vehicle system.

[0134] Optionally, in one embodiment of this application, the in-vehicle wallpaper generation device 10 further includes an acquisition module and a construction module.

[0135] The acquisition module is used to acquire a multimodal dataset of a large multimodal model before generating the final display wallpaper of the vehicle system. The multimodal dataset includes multimodal data from real vehicle images, landscape images, images of various styles, text and prompts input by the user, multimodal data from vehicle system signal data, and wallpapers generated from various information sources that meet the user's needs.

[0136] A building block is used to train a model using a multimodal dataset to build a large multimodal model before generating the final display wallpaper for the vehicle infotainment system.

[0137] Optionally, in one embodiment of this application, the in-vehicle wallpaper generation device 10 further includes a cleaning module and a labeling module.

[0138] The cleaning module is used to clean the multimodal dataset after it has been acquired, so as to obtain a multimodal dataset that meets the preset training conditions.

[0139] The annotation module is used to annotate images in the multimodal dataset that meet preset conditions after the multimodal dataset is acquired, so as to adjust the model parameters based on the annotated images.

[0140] Optionally, in one embodiment of this application, the generation module 300 includes: a type acquisition unit, a display unit, and a control unit.

[0141] The type acquisition unit is used to obtain the generation type of the final displayed wallpaper.

[0142] The display unit is used to receive user feedback voice commands or manual setting commands when the generation type is a multi-turn voice generation command or a text and prompt word combination command, and to display the final wallpaper according to the voice command or manual setting command.

[0143] The control unit is used to control the vehicle's display to show the final wallpaper when the generation type is the same as the generation type corresponding to the vehicle's signal combination instruction.

[0144] Optionally, in one embodiment of this application, the generation module 300 includes a determination unit and a wallpaper generation unit.

[0145] The determining unit is used to determine at least one of the following based on at least one wallpaper action: image address link, instruction to generate image, prompt word and enhanced prompt word, and image name.

[0146] A wallpaper generation unit is used to generate a final display wallpaper based on at least one of the following:

[0147] Optionally, in one embodiment of this application, the in-vehicle wallpaper generation device 10 further includes: an information acquisition module, a judgment module, and an instruction generation module.

[0148] The information acquisition module is used to acquire at least one relevant information from the vehicle's infotainment system at preset intervals when the vehicle is powered on.

[0149] The judgment module is used to determine whether at least one relevant information meets the preset change conditions compared with the previous relevant information.

[0150] The instruction generation module is used to generate vehicle-mounted signal combination instructions if preset change conditions are met.

[0151] It should be noted that the foregoing explanation of the method for generating in-vehicle wallpapers also applies to the device for generating in-vehicle wallpapers in this embodiment, and will not be repeated here.

[0152] The in-vehicle wallpaper generation device proposed in this application can use a multi-mode large model to fuse and process multi-turn voice commands, text and prompt word combination commands, and / or in-vehicle signal combination commands. It performs semantic analysis, prompt word extraction and enhancement to generate specific wallpaper operation commands, and finally generates and displays the in-vehicle wallpaper. This achieves highly intelligent and personalized in-vehicle wallpaper generation, accurately understanding user intent and dynamically adjusting wallpaper content, significantly improving the user experience and personalization level of the in-vehicle system. It brings convenience and a sense of surprise to users, and the real-time generated beautiful wallpapers enhance user satisfaction with the in-vehicle system and stimulate users' desire to purchase. Therefore, it solves the problems in related technologies, such as the lack of personalization in system-preset wallpapers, the cumbersome operation of importing images from external devices, the limited selection of wallpapers obtained from the cloud, and the reliance on user descriptions for voice-generated wallpapers that cannot be modified.

[0153] Figure 9 A schematic diagram of the structure of a vehicle provided in an embodiment of this application. The vehicle may include:

[0154] The memory 901, the processor 902, and the computer program stored on the memory 901 and capable of running on the processor 902.

[0155] When the processor 902 executes the program, it implements the method for generating in-vehicle wallpapers provided in the above embodiments.

[0156] Furthermore, the vehicle also includes:

[0157] Communication interface 903 is used for communication between memory 901 and processor 902.

[0158] The memory 901 is used to store computer programs that can run on the processor 902.

[0159] The memory 901 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0160] If the memory 901, processor 902, and communication interface 903 are implemented independently, then the communication interface 903, memory 901, and processor 902 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0161] Optionally, in a specific implementation, if the memory 901, processor 902, and communication interface 903 are integrated on a single chip, then the memory 901, processor 902, and communication interface 903 can communicate with each other through an internal interface.

[0162] The processor 902 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0163] This embodiment also provides a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for generating in-vehicle wallpapers.

[0164] This application also provides a computer program product storing a computer program that, when executed by a processor, implements the above-described method for generating in-vehicle wallpapers.

[0165] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0166] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0167] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0168] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0169] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0170] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0171] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0172] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for generating in-vehicle infotainment wallpaper, characterized in that, Includes the following steps: Receive multi-turn voice generation commands from users, text and prompt word combination commands from the vehicle's infotainment system, and / or vehicle infotainment system signal combination commands; The multi-turn speech generation instruction, the text and prompt word combination instruction, and / or the vehicle signal combination instruction are input into a pre-constructed multi-mode large model to obtain semantic parsing results. At least one prompt word is extracted based on the semantic parsing results, and the at least one prompt word is enhanced to obtain at least one wallpaper action among the following: content generation, content addition, content deletion, and content modification of the vehicle wallpaper. Based on the multi-mode large model, the final display wallpaper of the vehicle system is generated according to the at least one wallpaper action, and the final display wallpaper is displayed on the vehicle system.

2. The method according to claim 1, characterized in that, Before generating the final display wallpaper for the vehicle's infotainment system, the following steps are also included: Obtain the multimodal dataset of the multimodal large model, wherein the multimodal dataset includes multimodal data generated from real vehicle images, landscape images, images of various styles, user-inputted text and prompt words, vehicle signal data, and wallpapers that meet user needs from various information sources; The model is trained using the multimodal dataset to construct the large multimodal model.

3. The method according to claim 2, characterized in that, After acquiring the multimodal dataset, the following is also included: The multimodal dataset is cleaned to obtain a multimodal dataset that meets preset training conditions; Images in the multimodal dataset that meet preset conditions are labeled, and model parameters are adjusted based on the labeled images.

4. The method according to claim 1, characterized in that, The process of displaying the final wallpaper on the vehicle's infotainment system includes: Obtain the generation type of the final displayed wallpaper; When the generation type is the generation type corresponding to the multi-turn voice generation instruction or the text and prompt word combination instruction, the system receives voice instructions or manual setting instructions from the user and displays the final wallpaper according to the voice instructions or manual setting instructions. When the generation type is the same as the generation type corresponding to the vehicle system signal combination instruction, the vehicle system is controlled to display the final displayed wallpaper.

5. The method according to claim 1, characterized in that, The step of generating the final display wallpaper for the vehicle system based on the at least one wallpaper action includes: Based on the at least one wallpaper action, determine at least one of the following: image address link, instruction to generate image, prompt word and enhanced prompt word, and image name; The final display wallpaper is generated based on at least one of the above.

6. The method according to claim 1, characterized in that, Also includes: When the vehicle is powered on, at least one relevant information of the vehicle system is acquired at each preset time interval; Determine whether the at least one relevant information satisfies a preset change condition compared to the previous relevant information; If the preset change conditions are met, the vehicle-mounted signal combination command is generated.

7. A device for generating in-vehicle wallpaper, characterized in that, include: The receiving module is used to receive multi-turn voice generation commands from users, text and prompt word combination commands from the vehicle's infotainment system, and / or vehicle infotainment system signal combination commands from the vehicle. The extraction module is used to input the multi-turn voice generation instruction, the text and prompt word combination instruction and / or the vehicle signal combination instruction into a pre-constructed multi-mode large model to obtain semantic parsing results, extract at least one prompt word based on the semantic parsing results, and perform enhancement processing on the at least one prompt word to obtain at least one wallpaper action among the vehicle wallpaper content generation, content addition, content deletion and content modification. A generation module is used to generate the final display wallpaper of the vehicle system based on the multi-model large model and according to the at least one wallpaper action, and to display the final display wallpaper on the vehicle system.

8. A vehicle, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the method for generating in-vehicle wallpapers as described in any one of claims 1-6.

9. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method for generating in-vehicle wallpapers as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, The computer program is executed to implement the method for generating in-vehicle wallpapers as described in any one of claims 1-6.