Model training and intelligent cockpit visible-to-speak method

By training a multimodal smart cockpit interaction model, the problems of passiveness in information acquisition, high resource consumption and poor interface compatibility in the existing technology are solved, and the accurate recognition of the interface element functions of the smart cockpit and the accurate understanding of voice intentions are achieved, which improves the accuracy of interaction and user experience.

CN119993147AActive Publication Date: 2025-05-13YUN ZHI SHENG (HANG ZHOU) ZHI NENG KE JI YOU XIAN GONG SI

Patent Information

Application Number
CN202510179347.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-13
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing smart cockpit has the problem of passiveness in information acquisition, high resource consumption and poor interface compatibility, and the cost of third-party application development is high.

Method used

By training a multimodal intelligent cockpit interaction model, the first sample set containing the interface element attribute information samples and the second sample set of voice command samples are obtained, and the model is trained to identify the functional meaning of the interface element and map the speech intention to the interface element operation.

Benefits of technology

It realizes accurate identification of the functional meaning of interface elements of the target smart cockpit system and accurate understanding of voice intentions, reducing resource consumption and development costs, and improving interaction accuracy and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993147A_ABST
    Figure CN119993147A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode intelligent cockpit interaction large model training method and an intelligent cockpit visible-to-speak method. And training the multi-modal basic large model having the multi-modal data processing capability and the natural language understanding capability by using the first sample set, and converting the multi-modal basic large model into a multi-modal intelligent cockpit interaction large model. The process enables the model to accurately recognize the functional meanings of the interface elements contained in each interface in the target intelligent cockpit system, and lays a solid foundation for subsequent interaction with the user. And based on the second sample set, training a multi-modal intelligent cockpit interaction large model, wherein the multi-modal intelligent cockpit interaction large model can learn a mapping relationship between the voice intention and the interface elements of the interface images in the target intelligent cockpit system. When a user sends a voice instruction subsequently, the multi-mode intelligent cockpit interaction large model can accurately understand the intention of the user and find the corresponding interface element to execute the operation, so that the accuracy and effectiveness of voice interaction are greatly enhanced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent driving and deep learning technology, and in particular to a training method for a multimodal smart cockpit interaction large model and a smart cockpit visible and audible method. Background Art

[0002] With the rapid development of smart cockpit technology, how to provide drivers with a safer, more efficient and natural interactive experience has become a core issue that needs to be solved in the industry. Although traditional interaction methods, such as manually touching the screen or operating physical buttons, meet basic operating needs to a certain extent, they will inevitably distract the driver's attention, thereby increasing the potential risks during driving. Therefore, it is particularly important to develop a new interaction method that can reduce driver distraction and improve driving safety.

[0003] The See-and-Say feature is a highly promising interactive solution that has emerged in this context. This feature allows drivers to directly operate elements on the screen through voice commands, thus achieving a more intuitive and natural interactive experience. However, the existing methods of implementing the See-and-Say feature face many difficulties.

[0004] On the one hand, although the implementation based on Android auxiliary services can use the Android system's own mechanism to monitor screen drawing changes and traverse all texts in the current interface for registration, this method has problems such as passive information acquisition, high resource consumption, and limited interface compatibility. Because it relies on system-level auxiliary services, it cannot actively obtain or control interface element information, which may lead to compatibility issues in certain specific interfaces or applications. At the same time, this method will consume a lot of system resources when processing a large number of interface elements, affecting overall performance.

[0005] On the other hand, although the implementation method based on the third-party application registration solution allows developers to actively write text for registration and achieve more flexible interactive control, the development cost and difficulty of this method are relatively high. Developers need to deeply customize each application and register the corresponding text on different interfaces by themselves, which not only increases the development workload, but may also lead to poor application adaptability. Since the interface layout and elements of different applications may be different, developers need to adapt each application separately, which undoubtedly increases the development cost and cycle. Summary of the invention

[0006] The present application provides a multimodal smart cockpit interaction large model training and smart cockpit visible and speakable method, which is used to solve the problems of passive and uncontrollable information acquisition, poor interface compatibility, high third-party application development cost and high resource consumption in the existing smart cockpit visible and speakable function implementation methods.

[0007] In a first aspect, the present application provides a method for training a multimodal smart cockpit interaction large model, the method comprising:

[0008] Obtaining a first sample set including attribute information samples corresponding to a plurality of interface elements respectively; wherein the attribute information samples of any interface element include one or more of the following: operation-related information of the interface element, the interface element's own characteristics, text descriptions related to the interface element, and context information of the interface element, and each of the attribute information samples corresponds to a real functional meaning label;

[0009] Based on each of the attribute information samples in the first sample set and the real functional meaning labels corresponding to each of the attribute information samples, the pre-trained multimodal basic big model is trained to obtain a trained multimodal smart cockpit interaction big model; wherein the multimodal basic big model is a big model that has multimodal data processing capabilities and natural language understanding capabilities, and the multimodal smart cockpit interaction big model has the ability to identify the functional meanings of the interface elements contained in each interface in the target smart cockpit system;

[0010] Acquire a second sample set including a plurality of voice command samples for interacting with the target smart cockpit system; wherein any voice command sample corresponds to a voice intent label and an interface element label;

[0011] Based on each of the voice command samples in the second sample set, and the voice intent label and interface element label of each of the voice command samples, the multimodal smart cockpit interaction big model is trained so that the multimodal smart cockpit interaction big model learns the mapping relationship between voice intent and the interface elements of each interface image in the target smart cockpit system.

[0012] In a second aspect, the present application also provides a method for enabling a smart cockpit to be visible and audible based on the model trained by the above method, the method comprising:

[0013] Get the voice commands to be processed;

[0014] Determine, by using a pre-trained multimodal smart cockpit interaction model, a target interface element corresponding to the voice intent of the voice command to be processed based on the voice command to be processed;

[0015] Determine the target coordinate information corresponding to the target interface element according to the pre-saved coordinate information corresponding to each interface element;

[0016] According to the target coordinate information, the target interface element is operated in the current interface of the target smart cockpit system.

[0017] In a third aspect, the present application provides a training device for a multimodal intelligent cockpit interaction large model, the device comprising:

[0018] A first acquisition unit is used to acquire a first sample set including attribute information samples corresponding to a plurality of interface elements respectively; wherein the attribute information samples of any interface element include one or more of the following: operation-related information of the interface element, the interface element's own characteristics, text descriptions related to the interface element, and context information of the interface element, and each of the attribute information samples corresponds to a real functional meaning label;

[0019] A first training unit is used to train the pre-trained multimodal basic big model based on each of the attribute information samples in the first sample set and the real functional meaning labels corresponding to each of the attribute information samples, so as to obtain a trained multimodal smart cockpit interaction big model; wherein the multimodal basic big model is a big model that has multimodal data processing capabilities and natural language understanding capabilities, and the multimodal smart cockpit interaction big model has the ability to identify the functional meanings of interface elements contained in each interface in the target smart cockpit system;

[0020] A second acquisition unit is used to acquire a second sample set including a plurality of voice command samples for interacting with the target smart cockpit system; wherein any voice command sample corresponds to a voice intent label and an interface element label;

[0021] The second training unit is used to train the multimodal smart cockpit interaction big model based on each of the voice command samples in the second sample set, and the voice intent label and interface element label of each of the voice command samples, so that the multimodal smart cockpit interaction big model learns the mapping relationship between voice intent and the interface elements of each interface image in the target smart cockpit system.

[0022] In a fourth aspect, the present application provides a smart cockpit visible and audible device, the device comprising:

[0023] An acquisition module, used to acquire voice instructions to be processed;

[0024] A training module, for determining a target interface element corresponding to the voice intent of the voice command to be processed based on the voice command to be processed by using a pre-trained multimodal smart cockpit interaction large model;

[0025] A determination module, used to determine the target coordinate information corresponding to the target interface element according to the pre-saved coordinate information corresponding to each interface element;

[0026] An operating module is used to operate the target interface element in the current interface of the target smart cockpit system according to the target coordinate information.

[0027] In a fifth aspect, the present application provides a computer device, comprising a processor, wherein the processor is used to implement the steps of the training method of the multimodal smart cockpit interaction large model as described above when executing a computer program stored in a memory, or to implement the steps of the smart cockpit visible-and-speakable method as described above.

[0028] In a sixth aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the training method of the multimodal smart cockpit interaction large model as described above, or implements the steps of the smart cockpit visible-and-speakable method as described above.

[0029] The beneficial effects of this application are as follows:

[0030] 1. By obtaining the first sample set based on the target smart cockpit system interface image, each attribute information sample corresponds to a real functional meaning label, providing rich and accurate original materials for the multimodal smart cockpit interaction model. The multimodal smart cockpit interaction model can learn the functional meanings of various interface elements in different scenarios, build a comprehensive smart cockpit interface knowledge system, and accurately identify the functional meanings of interface elements contained in each interface in the target smart cockpit system, laying a solid foundation for subsequent interaction with users.

[0031] 2. Use the first sample set to train the multimodal basic large model that already has multimodal data processing capabilities and natural language understanding capabilities, and transform it into a multimodal smart cockpit interaction large model. This process allows the model to deeply understand the characteristics of the smart cockpit system and user interaction needs, significantly improve its intelligent interaction level, and be able to respond quickly and accurately according to user operations or instructions, achieving more efficient and intelligent human-computer interaction.

[0032] 3. Obtain a second sample set containing voice command samples and their corresponding voice intent labels and interface element labels, further enriching the learning data of the multimodal smart cockpit interaction model. The multimodal smart cockpit interaction model is trained based on the second sample set. The multimodal smart cockpit interaction model can learn the mapping relationship between voice intent and the interface elements of each interface image in the target smart cockpit system. Subsequently, when the user issues a voice command, the multimodal smart cockpit interaction model can accurately understand the user's intention and find the corresponding interface elements to perform the operation, which greatly enhances the accuracy and effectiveness of voice interaction and improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0034] Figure 1 A schematic diagram of the training process of a multimodal smart cockpit interaction large model provided in an embodiment of the present application;

[0035] Figure 2 A schematic diagram of a specific multi-modal intelligent cockpit interaction large model training process provided in an embodiment of the present application;

[0036] Figure 3 A schematic diagram of a process of a smart cockpit that can be seen and said provided in an embodiment of the present application;

[0037] Figure 4 A schematic diagram of a process flow of a smart cockpit that can be seen and said provided in an embodiment of the present application;

[0038] Figure 5 A schematic diagram of the structure of a training device for a multimodal intelligent cockpit interaction large model provided in an embodiment of the present application;

[0039] Figure 6 A schematic diagram of the structure of a smart cockpit visible and speaking device provided in an embodiment of the present application;

[0040] Figure 7 It is a structural schematic diagram of a computer device provided in an optional embodiment of the present application. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0042] In order to improve the intelligence level of the interaction between the smart cockpit system and the user and achieve a more accurate and efficient human-computer interaction experience, this application provides a multimodal smart cockpit interaction large model training and a smart cockpit visible and audible method.

[0043] Embodiment 1:

[0044] This application provides a training method for a multimodal intelligent cockpit interaction large model. Figure 1A schematic diagram of a training process of a multimodal intelligent cockpit interaction large model provided in an embodiment of the present application, the process includes:

[0045] S101: Obtain a first sample set including attribute information samples corresponding to multiple interface elements respectively; wherein the attribute information samples of any interface element include one or more of the following: operation-related information of the interface element, inherent characteristics of the interface element, text description related to the interface element and context information of the interface element, and each of the attribute information samples corresponds to a real functional meaning label.

[0046] The training method of the multimodal smart cockpit interaction large model provided in this application is applied to a computer device, which can be a smart terminal, such as an in-vehicle central control, an in-vehicle processor, etc., or a server, such as an in-vehicle business server, etc.

[0047] In order to improve the level of intelligence of the interaction between the smart cockpit system and the user and achieve a more accurate and efficient human-computer interaction experience, in this application, it is necessary to pre-train a multimodal smart cockpit interaction large model for the target smart cockpit system. In order to train the multimodal smart cockpit interaction large model, it is necessary to pre-acquire a sample set (referred to as the first sample set) of relevant information of each interface element in the target smart cockpit system for model learning. Among them, the first sample set contains attribute information samples corresponding to multiple interface elements. An interface element is a visual component with independent functions or interactions on the interface of the target smart cockpit system. The types of the interface elements include one or more of the following: buttons, text boxes, such as text boxes in the car navigation interface, play / pause buttons in the music playback interface, etc. The attribute information sample is a series of feature information extracted for each interface element. For example, the attribute information sample of the volume adjustment slider covers the coordinate information, size, and current value representing the volume of the slider. This information can accurately describe the state and characteristics of the interface elements in the interface. At the same time, the attribute information samples also correspond to real functional meaning labels, which are used to accurately describe the functions of interface elements. For example, the real functional meaning label of the route planning button in the navigation interface is "plan navigation route".

[0048] In a possible implementation, the attribute information sample of any interface element includes one or more of the following: operation-related information of the interface element, the interface element's own characteristics, the environment information where the interface element is located, and the context information of the interface element. Among them, the operation-related information of the interface element is used to guide how to operate the interface element. For example, button-type interface elements usually correspond to click operations, and clicking the play button can realize the music playback function; slider-type interface elements correspond to sliding operations, and sliding the volume adjustment slider up and down can change the volume; long pressing a point on the map interface can call up a special function for setting the destination; text box-type interface elements usually correspond to click and text input operations, and enter text content in the text box after clicking the text box. The interface element's own characteristics cover the shape, color, texture, text content, coordinate information and other characteristics of the interface element. The context information of the interface element involves rich information such as the previous and subsequent operation processes associated with the interface element, closely related related functions, and the specific page scene in which it is located.

[0049] Among them, these attribute information samples are all obtained based on the interface images in the target smart cockpit system.

[0050] There are several ways to obtain interface images:

[0051] 1. Screenshot acquisition

[0052] It can be obtained by taking screenshots of the interface of the target smart cockpit system. For example, the screenshots are taken according to a preset period or time point. Exemplarily, each interface image of the target smart cockpit system is obtained by periodic screenshots. For example, a suitable time period is set, such as every 100 milliseconds, to automatically trigger the screenshot command, so as to capture the screen once. In this way, the interface status of the smart cockpit system at different use stages can be captured, including the system startup, user operation process, and the interface during stable operation.

[0053] It should be noted that when setting the screenshot cycle, different values ​​can be set according to different scenarios. If you want to fully capture the interface status of the smart cockpit system at different stages of use, including system startup, user operation process, and interface during stable operation, the screenshot cycle can be set shorter. However, the screenshot operation will take up certain system resources, including CPU, memory, and storage reading and writing. If the screenshot cycle is too short, such as taking a screenshot every 1 second or even shorter, it may cause the system to run slowly and affect the normal functions of the smart cockpit system, such as real-time navigation route calculation and smooth multimedia playback. Therefore, if you want the screenshot operation to occupy as little system resources as possible and affect system operation, the screenshot cycle can be set longer.

[0054] 2. Screen recording acquisition

[0055] Screen recording software can be used to continuously record the interface display of the target smart cockpit system in different operating scenarios, and key frames can be extracted from the recorded video as interface images according to time nodes.

[0056] As a possible implementation method, in order to reduce the resources consumed by the subsequent processing of the interface image, in the present application, after the interface image in the target smart cockpit system is obtained based on the above embodiment, the interface image can be preprocessed. Among them, the image preprocessing includes one or more of the following: grayscale processing, image noise reduction. Among them, grayscale processing can convert a color image into a grayscale image, retaining only the brightness information of the image and removing the color information. After the grayscale processing, the image has a reduced amount of data, and at the same time simplifies the calculation process of the subsequent element recognition model, improves the processing efficiency, and can highlight the structure and contour of the image, which is conducive to element recognition. The image after noise reduction processing can provide clearer and more accurate image data for subsequent element recognition and reduce the interference of noise on the recognition results. The use of data enhancement technology (such as random cropping, rotation, and flipping of button images) can increase the diversity of attribute data samples and improve the multimodal smart cockpit interaction model. The ability to recognize and understand interface elements in different scenarios.

[0057] In one example, attribute information samples may be obtained from interface images in a target smart cockpit system by manual collection.

[0058] In another example, the attribute information sample can also be automatically obtained based on the interface image in the target smart cockpit system in the following manner:

[0059] For information related to the operation of interface elements, an interaction log recording module can be implanted in the target smart cockpit system. When the user interacts with the interface elements, the system records the interaction behavior in real time, including the timestamp of the click, the start and end positions of the slide, etc. For example, for the volume adjustment slider, the log recording module will record the time each time the user slides the slider, the range of change of the slider position, and other information. By analyzing these log data and combining them with the interface image, the operation-related information of the volume adjustment slider (such as the sliding operation and its direction and amplitude) can be obtained.

[0060] The characteristics of the interface elements themselves can be obtained in different ways according to the different types of interface elements.

[0061] Exemplarily, for any of the interface images, various types of interface elements contained in the interface image are determined by an image recognition algorithm (such as the YOLO algorithm or the Faster R-CNN algorithm).

[0062] For interface elements of button type in the interface image, the visual features (such as button shape, color, texture, etc.) and bounding box coordinates of each button in the interface image are determined by a pre-trained button recognition algorithm. Among them, the pre-trained button recognition algorithm is trained by constructing a data set with a large number of image samples containing various types of buttons. The data set constructed with image samples containing various types of buttons collects interface images of smart cockpit systems of different brands and models to ensure the diversity of buttons, covering various shapes (circular, square, polygonal, etc.), colors (solid color, gradient color, etc.) and different textures (smooth, frosted, patterned, etc.). These image samples are annotated to accurately mark the category, bounding box coordinates and corresponding visual feature attributes of each button.

[0063] For the interface elements of the text box type in the interface image, the text box in the interface image is converted into an editable text string through optical character recognition technology, and the boundary box coordinates of the area where the text box is located are recorded.

[0064] For text descriptions related to interface elements, for any interface element, the bounding box coordinates of the interface element can be determined from the interface image. Then, the image area containing the bounding box coordinates is recognized by optical character recognition technology to determine the text contained in the image, and the text is determined as the text description related to the interface element.

[0065] For the context information of interface elements, the operation log and user behavior data of the target smart cockpit system can be combined. The operation log records the order of user operations on various interface elements at different time points. By analyzing these data, the previous and subsequent operation processes associated with the interface element can be restored. At the same time, based on the system's functional architecture and page jump logic, the specific page scene in which the interface element is located and the closely related associated functions are determined. For example, in the navigation page, the context information of the map interface element includes its association with elements such as the destination input box and route planning button, as well as information such as which stage of the navigation function is currently in.

[0066] To improve the generalization ability of the model, the first sample set should cover as many different types of interfaces and interface elements in the target smart cockpit system as possible, including not only the common navigation interface and multimedia playback interface, but also the vehicle setting interface, driving information display interface, etc. And for each interface type, it is necessary to collect interface images in multiple different states, such as images of the navigation interface at different zoom levels and different route planning, to ensure that the model can learn comprehensive and accurate information about interface elements.

[0067] S102: Based on each of the attribute information samples in the first sample set and the real functional meaning labels corresponding to each of the attribute information samples, the pre-trained multimodal basic big model is trained to obtain a trained multimodal smart cockpit interaction big model; wherein the multimodal basic big model is a big model that already has multimodal data processing capabilities and natural language understanding capabilities, and the multimodal smart cockpit interaction big model has the ability to identify the functional meanings of the interface elements contained in each interface in the target smart cockpit system.

[0068] After obtaining the first sample set, it can be used to train the pre-trained multimodal basic large model to obtain a trained multimodal smart cockpit interaction large model. Among them, the multimodal basic large model is built based on a deep learning framework, and has the ability to process multimodal data and understand natural language. It can process various types of data such as images and texts, and can also understand the semantics of natural language. For example, some neural network models based on the Transformer architecture that have been trained with a large amount of image and text data have the ability to extract image features and understand text semantics.

[0069] Exemplarily, the attribute information samples in the first sample set and their corresponding true functional meaning labels are input into the multimodal basic large model. Through the multimodal basic large model, the functional meaning of the interface element corresponding to the attribute information sample can be predicted based on the input attribute information sample. The stochastic gradient descent algorithm and its variants are used to minimize the error between the model prediction result and the true functional meaning label. At the same time, regularization techniques (such as L1 and L2 regularization) are used to prevent overfitting, and hyperparameters such as the number of training rounds and batch size are reasonably set.

[0070] After multiple rounds of training, the multimodal basic large model gradually learns the functional meaning of each interface element in the target smart cockpit system, and is thus transformed into a trained multimodal smart cockpit interaction large model, which has the ability to identify the functional meaning of the interface elements contained in each interface in the target smart cockpit system.

[0071] The multimodal smart cockpit interaction model obtained after training the multimodal basic model on the first sample set is a model specifically used for the interaction of the smart cockpit system. It not only has the natural language understanding ability and multimodal data processing ability of the multimodal basic model, but also can accurately identify various interface elements in the target smart cockpit system interface and understand the functional meaning of these elements. For example, when the multimodal smart cockpit interaction model receives the navigation interface image, it can accurately determine the functions of each button and area, knowing that the text box is used to enter the destination and the route planning button is used to plan the itinerary route.

[0072] S103: Obtain a second sample set including a plurality of voice command samples for interacting with the target smart cockpit system; wherein any voice command sample corresponds to a voice intent label and an interface element label.

[0073] In order to further enhance the interactive capability of the multimodal smart cockpit interaction model, so that it can better understand the user's voice intention to interact with the target smart cockpit system through voice, it is necessary to obtain a second sample set, so that the multimodal smart cockpit interaction model can learn the mapping relationship between the voice intention and the interface elements of each interface image in the target smart cockpit system based on the second sample set. Among them, the second sample set contains multiple voice command samples about interacting with the target smart cockpit system. The voice command sample is used to characterize the voice content issued by the user when interacting with the target smart cockpit system by voice, such as "help me find a nearby hot pot restaurant", "play xxx's songs", "open the car window", etc. Each voice command sample corresponds to a voice intention label and an interface element label. The voice intention label is used to clarify the purpose of the user's voice command. For example, the voice intention label of "help me find a nearby hot pot restaurant" is "search for nearby hot pot restaurants"; the voice intention label of "play xxx's songs" is "play songs of a specified singer". The interface element label identifies the specific interface element in the target smart cockpit system interface involved in the voice command. For example, in the voice command "Help me find a nearby hot pot restaurant", the interface element label may be a text box in the navigation interface, because the relevant information of the hot pot restaurant needs to be entered in this interface element to perform the search operation; for "Play xxx's songs", the interface element label may be the song search input box or song list area in the music playback interface.

[0074] In this application, the method of obtaining voice command samples can be to collect voice commands in actual vehicle usage scenarios, or to invite users to issue voice commands and simultaneously record relevant information through simulation experiments, or the users can upload them during the use of the target smart cockpit.

[0075] In one example, considering that the collected voice signal is usually affected by the complex environment in the car, there is a lot of environmental noise and echo interference, which will seriously affect the quality and accuracy of the voice command. Therefore, it is necessary to perform voice preprocessing on the collected voice signal. Among them, the voice preprocessing includes but is not limited to one or more of the following: voice enhancement processing, framing of voice signals, and windowing operations. The voice enhancement processing can remove environmental noise and echo interference. The framing and windowing operations of voice signals can convert continuous voice signals into a series of short-time and high-quality voice frames.

[0076] In a possible implementation, the voice command sample can be a collected voice signal or a preprocessed voice signal. Of course, it can also be a voice feature vector, such as Mel-frequency cepstral coefficients (MFCC), linear prediction cepstral coefficients (LPCC), etc., by converting the voice signal into a set of feature vectors that can characterize the voice content. These feature vectors contain information such as the frequency and amplitude of the voice, which can reflect the acoustic characteristics of the voice, provide basic data for subsequent voice intent understanding, and reduce the amount of calculation for subsequent model training. In the specific implementation process, it can be flexibly set according to needs, and no specific limitation is made here.

[0077] As a possible implementation method, when obtaining voice command samples, the diversity and comprehensiveness of the voice command samples can be ensured as much as possible, covering not only common voice interaction scenarios such as navigation, music playback, and vehicle control, but also some more special or personalized voice command scenarios. For example, users use dialects or customized voice commands to interact with the smart cockpit system. By reflecting these situations in the sample as much as possible, the adaptability of the multimodal smart cockpit interaction model to various voice interaction situations can be improved.

[0078] S104: Based on each of the voice command samples in the second sample set, and the voice intent label and interface element label of each of the voice command samples, the multimodal smart cockpit interaction big model is trained so that the multimodal smart cockpit interaction big model learns the mapping relationship between voice intent and the interface elements of each interface image in the target smart cockpit system.

[0079] After obtaining the second sample set, it is used to further train the multimodal smart cockpit interaction model that has been trained with the first sample set. Exemplarily, during training, the voice command samples in the second sample set are input into the multimodal smart cockpit interaction model in an orderly manner. The multimodal smart cockpit interaction model will calculate the prediction results through forward propagation based on the input data, and then use the cross-entropy loss function to calculate the differences between the prediction results and the actual voice intent labels and interface element labels, and then update the parameters of the multimodal smart cockpit interaction model through the back-propagation algorithm, so that the multimodal smart cockpit interaction model can continuously learn the mapping relationship between voice intent and interface elements in the interface image. For example, for the voice intent of "zoom in on the map", the multimodal smart cockpit interaction model will learn the interface elements (such as the zoom in button) corresponding to the map zoom operation in the navigation interface.

[0080] In a possible implementation, the multimodal smart cockpit interaction big model is trained based on each of the voice command samples in the second sample set, and the voice intent label and the interface element label of each of the voice command samples, so that the multimodal smart cockpit interaction big model learns the mapping relationship between the voice intent and the interface elements of each interface image in the target smart cockpit system, including:

[0081] Obtain any of the voice command samples and its corresponding voice intent label and interface element label;

[0082] Outputting, through the multimodal smart cockpit interaction macro model, a first probability distribution of the predicted voice intent of the voice command sample and a second probability distribution of the interface elements corresponding to the predicted voice intent based on the voice command sample and the learned functional meanings of the interface elements contained in each interface of the target smart cockpit system;

[0083] Based on the first probability distribution and the voice intent label, and the second probability distribution and the interface element label, the multimodal smart cockpit interaction big model is trained so that the multimodal smart cockpit interaction big model learns the mapping relationship between voice intent and interface elements.

[0084] In any round of training of the multimodal smart cockpit interaction big model, first, any voice command sample and its corresponding voice intent label and interface element label are obtained from the second sample set. For example, the voice command sample "turn on the music and play my favorite playlist" is obtained, and its corresponding voice intent label is "play the specified playlist music", and the interface element label may be the playlist selection button and the play button in the music playback interface. Then, the obtained voice command sample is input into the multimodal smart cockpit interaction big model. The multimodal smart cockpit interaction big model conducts in-depth analysis and reasoning based on the voice command sample and the functional meaning of the interface elements contained in each interface of the target smart cockpit system previously learned through the first sample set. The multimodal smart cockpit interaction big model will output the first probability distribution of the predicted voice intent of the voice command sample. For example, for "turn on the music and play my favorite playlist", the multimodal smart cockpit interaction big model may output the probability of the predicted voice intent of "playing the specified playlist music" as 0.85, the probability of "turn on the music" as 0.1, and the probability of other irrelevant intents as 0.05, etc. At the same time, the multimodal smart cockpit interaction model will also output the second probability distribution of the interface elements corresponding to the predicted voice intent, that is, for the predicted voice intent of "playing music from a specified playlist", the probability of outputting the playlist selection button is 0.9, the probability of the play button is 0.08, and the probability of other interface elements is 0.02, etc. Then, the loss value is determined by calculating the difference between the first probability distribution and the voice intent label, and the difference between the second probability distribution and the interface element label. According to the loss value, the parameters of the multimodal smart cockpit interaction model are adjusted so that the multimodal smart cockpit interaction model can more accurately output a probability distribution that matches the true label in subsequent predictions, thereby allowing the multimodal smart cockpit interaction model to gradually learn the mapping relationship between voice intent and interface elements.

[0085] Since the acquired second sample set contains several voice command samples, the above operation is performed for each voice command sample to train the multimodal smart cockpit interaction model. When the preset convergence condition is met, the training of the multimodal smart cockpit interaction model is completed.

[0086] Among them, the preset convergence condition can be that the sum of the loss values ​​determined based on each voice command sample is less than the pre-configured loss threshold, or the sum of the determined loss values ​​has been on a downward trend and tends to be flat, or the number of iterations for training the multimodal smart cockpit interaction large model reaches the set maximum number of iterations, etc. It can be set flexibly in the specific implementation and is not specifically limited here.

[0087] As a possible implementation method, when training the multimodal smart cockpit interaction big model, the voice command samples can be divided into training samples and test samples. The multimodal smart cockpit interaction big model is first trained based on the training samples, and then the reliability of the trained multimodal smart cockpit interaction big model is verified based on the test samples.

[0088] After multiple rounds of training and optimization, the multimodal smart cockpit interaction model can understand and respond to voice commands more accurately. When a user issues a more complex voice command such as "help me find a Chinese restaurant", the multimodal smart cockpit interaction model can quickly locate the text box elements in the navigation interface or the life service recommendation interface.

[0089] The beneficial effects of this application are as follows:

[0090] 1. By obtaining the first sample set based on the target smart cockpit system interface image, each attribute information sample corresponds to a real functional meaning label, providing rich and accurate original materials for the multimodal smart cockpit interaction model. The multimodal smart cockpit interaction model can learn the functional meanings of various interface elements in different scenarios, build a comprehensive smart cockpit interface knowledge system, and accurately identify the functional meanings of interface elements contained in each interface in the target smart cockpit system, laying a solid foundation for subsequent interaction with users.

[0091] 2. Use the first sample set to train the multimodal basic large model that already has multimodal data processing capabilities and natural language understanding capabilities, and transform it into a multimodal smart cockpit interaction large model. This process allows the model to deeply understand the characteristics of the smart cockpit system and user interaction needs, significantly improve its intelligent interaction level, and be able to respond quickly and accurately according to user operations or instructions, achieving more efficient and intelligent human-computer interaction.

[0092] 3. Obtain a second sample set containing voice command samples and their corresponding voice intent labels and interface element labels, further enriching the learning data of the multimodal smart cockpit interaction model. The multimodal smart cockpit interaction model is trained based on the second sample set. The multimodal smart cockpit interaction model can learn the mapping relationship between voice intent and the interface elements of each interface image in the target smart cockpit system. Subsequently, when the user issues a voice command, the multimodal smart cockpit interaction model can accurately understand the user's intention and find the corresponding interface elements to perform the operation, which greatly enhances the accuracy and effectiveness of voice interaction and improves the user experience.

[0093] Embodiment 2:

[0094] The following is an explanation of the training method of the multimodal intelligent cockpit interaction large model provided by the present application through a specific embodiment. Figure 2A schematic diagram of the specific multimodal smart cockpit interaction large model training process provided in the embodiment of the present application, the process mainly includes a sample collection stage, a function meaning recognition stage, and a voice command and interface element debugging stage, and each stage is introduced separately:

[0095] Phase 1: Sample collection phase (including the first sample set and the second sample set).

[0096] For the first sample set:

[0097] S201: Obtain various interface images of the target smart cockpit system by periodically taking screenshots.

[0098] S202: performing image preprocessing on each interface image; wherein the image preprocessing includes one or more of the following: grayscale processing, image noise reduction, and data enhancement.

[0099] S203: For each preprocessed interface image, obtain a sample of attribute information of each interface element in the preprocessed interface image.

[0100] Among them, the attribute information sample of any interface element includes one or more of the following: operation-related information of the interface element, the interface element's own characteristics, text description related to the interface element, and context information of the interface element.

[0101] S204: Determine the real functional meaning labels corresponding to the attribute information samples of each interface element, and save the attribute information samples of each interface element and the real functional meaning labels corresponding to each attribute information sample.

[0102] For the second sample set:

[0103] S205: Collecting voice signals of the vehicle interacting with the target smart cockpit system.

[0104] S206: Perform speech preprocessing on the collected speech signal.

[0105] The speech preprocessing includes speech enhancement processing, framing and windowing operations.

[0106] S207: Based on the preprocessed voice signals, obtain the voice feature vector of each voice signal, and determine the voice feature vector of each voice signal as a voice command sample.

[0107] S208: Determine the voice intent label and the interface element label corresponding to each voice command sample, and save each voice command sample, as well as the voice intent label and the interface element label corresponding to each voice command sample.

[0108] It should be noted that the order of the steps of collecting the first sample set and the second sample set is not fixed, and S201 to S204 may be before S205 to S208, or after S201 to S204, or S201 to S204 and S205 to S208 may be performed simultaneously.

[0109] Stage 2: Functional meaning identification stage.

[0110] S209: Obtain a pre-trained multimodal basic large model.

[0111] Among them, the multimodal basic large model is a large model that already has multimodal data processing capabilities and natural language understanding capabilities.

[0112] S210: Obtain any attribute information sample and its corresponding true functional meaning label from the first sample set.

[0113] S211: Obtain the prediction function meaning based on the attribute information sample through the pre-trained multimodal basic large model.

[0114] S212: Based on the predicted functional meaning and the real functional meaning label, the multimodal basic large model is trained to obtain a trained multimodal smart cockpit interaction large model.

[0115] Among them, the multimodal smart cockpit interaction model has the ability to identify the functional meanings of the interface elements contained in each interface of the target smart cockpit system. In the third stage, based on the multimodal smart cockpit interaction model trained in the second stage, the mapping relationship between the multimodal smart cockpit interaction model learning voice commands and interface elements is further optimized.

[0116] Phase three: voice command and interface element debugging phase.

[0117] S213: Obtain any voice command sample and its corresponding voice intent label and interface element label.

[0118] S214: Through the multimodal smart cockpit interaction model, based on the voice command samples and the learned functional meanings of the interface elements contained in each interface of the target smart cockpit system, the first probability distribution of the predicted voice intent of the voice command samples and the second probability distribution of the interface elements corresponding to the predicted voice intent are output.

[0119] S215: Based on the first probability distribution and the voice intent label, and the second probability distribution and the interface element label, the multimodal smart cockpit interaction big model is trained so that the multimodal smart cockpit interaction big model learns the mapping relationship between voice intent and interface elements.

[0120] Embodiment 3:

[0121] The present application also provides a method for making a smart cockpit visible and audible based on the model trained by the above embodiments. Figure 3 A schematic diagram of a process of a smart cockpit that can be seen and said is provided in an embodiment of the present application, and the process includes:

[0122] S301: Obtaining a voice instruction to be processed.

[0123] S302: Determine the target interface element corresponding to the voice intent of the voice instruction to be processed based on the voice instruction to be processed through a pre-trained multimodal smart cockpit interaction model.

[0124] S303: Determine target coordinate information corresponding to the target interface element according to pre-saved coordinate information corresponding to each interface element.

[0125] S304: According to the target coordinate information, operate the target interface element in the current interface of the target smart cockpit system.

[0126] The smart cockpit visible and speakable method provided in this application is applied to a computer device, which can be an intelligent terminal, such as an in-vehicle central control, an in-vehicle processor, etc., or a server, such as an in-vehicle business server, etc. Among them, the computer device for performing the smart cockpit visible and speakable in this application can be the same as the computer device for performing the multimodal smart cockpit interaction large model training mentioned above, or it can be different.

[0127] In a possible implementation, the training of the multimodal smart cockpit interaction model is generally performed offline. After the trained multimodal smart cockpit interaction model is obtained, the multimodal smart cockpit interaction model can be deployed to the computer device used for the smart cockpit visual communication.

[0128] In this application, sensors (such as microphones, speakers, etc.) for collecting voice signals have been arranged in an array form and installed in the vehicle equipped with the target smart cockpit system according to the size and shape of the vehicle equipped with the target smart cockpit system, so as to ensure that sound signals from all directions inside the vehicle can be accurately captured. For example, microphones are installed in sound areas such as around the driver's seat, in the passenger area, and even in the trunk to achieve comprehensive monitoring of all important sound sources in the vehicle. Later, during the operation of the vehicle equipped with the target smart cockpit system, the microphones in different areas of the vehicle equipped with the target smart cockpit system can continuously collect voice signals from various areas in the vehicle.

[0129] In a possible implementation, the collected speech signal may be subjected to speech preprocessing. The speech preprocessing includes, but is not limited to, one or more of the following: speech enhancement processing, framing of speech signals, and windowing operations. The speech enhancement processing may remove environmental noise and echo interference. The framing and windowing operations of speech signals may convert a continuous speech signal into a series of short-duration and high-quality speech frames.

[0130] As a possible implementation, the voice signal can be directly determined as the voice command to be processed, or the preprocessed voice signal can be determined as the voice command to be processed. The preprocessing includes but is not limited to one or more of the following: voice enhancement processing, voice signal framing, and windowing operations.

[0131] As another possible implementation, a speech feature vector (for example, Mel-frequency cepstral coefficients (MFCC), linear prediction cepstral coefficients (LPCC), etc.) can also be determined as a voice instruction to be processed to reduce the amount of calculation of the subsequent multimodal smart cockpit interaction large model. The speech feature vector can be a speech feature vector of a speech signal, or it can be a speech feature vector of a preprocessed speech signal. Exemplarily, the collected speech signal is preprocessed. Among them, the preprocessing includes but is not limited to one or more of the following: speech enhancement processing, framing of speech signals, and windowing operations. Based on each preprocessed speech frame, the speech feature vector of each speech signal is obtained, and the speech feature vector of each speech signal is determined as a voice instruction to be processed.

[0132] After obtaining the voice instructions to be processed based on the above-mentioned embodiments, the voice instructions to be processed can be input into the pre-trained multimodal smart cockpit interaction model. Through the multimodal smart cockpit interaction model, the target interface element corresponding to the voice intent of the voice instructions to be processed is determined based on the voice instructions to be processed. In this process, the multimodal smart cockpit interaction model will perform semantic analysis on the voice instructions to be processed, and combine the previously learned mapping relationship between the voice intent and the interface elements of each interface image in the target smart cockpit system to comprehensively determine the target interface element that best matches the voice intent of the voice instruction.

[0133] After determining the target interface element and target key information, the target coordinate information corresponding to the target interface element is determined according to the pre-saved coordinate information corresponding to each interface element. For example, for the play and pause button in the music playback interface, the target coordinate information corresponding to the target interface element can be obtained by querying the pre-saved coordinate information corresponding to each interface element and determining the target coordinate information corresponding to the target interface element.

[0134] With the target coordinate information, the target interface elements can be operated in the current interface of the target smart cockpit system. Exemplarily, the target interface elements are operated at the target coordinate position through the automated operation simulation tool. For example, in the target smart cockpit system of the Android system, the Android Instrumentation framework or Accessibility Service can be used to simulate click events and trigger corresponding interface jumps or function operations.

[0135] In a possible implementation, after determining the target coordinate information corresponding to the target interface element according to the pre-saved coordinate information corresponding to each interface element, before operating the target interface element in the current interface of the target smart cockpit system according to the target coordinate information, the method further includes:

[0136] Determine whether the target coordinate information is valid in the current interface.

[0137] In order to ensure the accuracy of the operation, in the present application, after the target coordinate information is determined, it can be determined whether the target coordinate information is valid in the current interface of the target smart cockpit system. If it is determined that the target coordinate information is valid in the current interface, the subsequent steps of operating the target interface elements in the current interface of the target smart cockpit system according to the target coordinate information are performed to avoid coordinate errors of the interface elements due to interface changes or other abnormal situations. Among them, the valid here means that the target coordinate information is within the valid range of the current interface. Exemplarily, after obtaining the target coordinate information, it can be determined whether the target coordinate information is within the valid range of the current interface. If it is determined that the target coordinate information is within the valid range of the current interface, the target coordinate information is determined to be valid; if it is determined that the target coordinate information is not within the valid range of the current interface, the target coordinate information is determined to be invalid.

[0138] As a possible implementation manner, if it is determined that the target coordinate information is invalid in the current interface, the method further includes:

[0139] Obtain the latest interface image of the current interface;

[0140] Re-determining the latest coordinate information corresponding to each interface element in the latest interface image;

[0141] According to the newly determined latest coordinate information corresponding to each interface element in the latest interface image, the coordinate information corresponding to each interface element is updated;

[0142] Determine the target coordinate information corresponding to the target interface element based on the updated coordinate information corresponding to each interface element;

[0143] The target interface element is operated in the current interface according to the target coordinate information.

[0144] Based on the above embodiment, if it is determined that the target coordinate information is invalid, the latest interface image of the current interface is obtained, for example, by taking a screenshot or recording a screen.

[0145] Then, the latest coordinate information corresponding to each interface element in the latest interface image is re-determined. When determining the latest coordinate information corresponding to each interface element in the latest interface image, each interface element must first be re-identified and then its coordinate information is determined.

[0146] For example, the latest interface image acquired is first preprocessed. The image preprocessing includes one or more of the following: grayscale processing, image noise reduction, and data enhancement. After the image preprocessing is completed, the target detection and recognition algorithm, such as the convolutional neural network (CNN) algorithm based on deep learning, is used to identify each interface element in the image. These algorithms can learn and extract the features of interface elements, so as to accurately determine the type of each element and determine their latest coordinate information in the latest interface image.

[0147] When updating the coordinate information corresponding to each interface element according to the latest coordinate information corresponding to each interface element in the newly determined latest interface image, the newly acquired coordinate information can be compared and replaced with the original coordinate information. If the original coordinate information table is built based on a relational database, the data can be updated through SQL statements to ensure that the coordinate information of each interface element is the latest and accurate.

[0148] After the coordinate information is updated, the target coordinate information corresponding to the target interface element is determined again based on the updated coordinate information corresponding to each interface element. At this time, the system will query the accurate coordinates of the target interface element from the updated coordinate information library based on the multimodal smart cockpit interaction model's understanding of the voice command and the positioning results of the target interface element.

[0149] Finally, according to the determined target coordinate information, the target interface element is operated in the current interface.

[0150] As a possible implementation manner, after operating the target interface element in the interface of the target smart cockpit system according to the target coordinate information, the method further includes:

[0151] Determining whether the operation on the target interface element is successful according to the feedback information of the current interface;

[0152] If the operation is successful, the successful operation log is recorded;

[0153] If the operation fails, the cause of the failure is determined; the pre-configured repair measures corresponding to the failure cause are determined and executed; and the failed operation log is recorded.

[0154] In order to ensure the accuracy and stability of the operation, during the operation, the feedback information of the current interface will be monitored in real time to determine whether the operation of the target interface element is successful. Among them, the feedback information of the current interface may include the state change of the interface element, system prompt information, etc. For example, in the example of adjusting the air conditioner temperature, if the air conditioner temperature display value changes to 26 degrees, or the system pops up a prompt box "The temperature has been adjusted to 26 degrees", it can be determined that the operation is successful.

[0155] If the operation is successful, a successful operation log is recorded. The log content of the successful operation log may include the operation time, the voice command of the operation, the corresponding target interface elements, and the specific operation content. For example, it records "[specific time], the user issued the command 'adjust the air conditioning temperature to 26 degrees', operated the temperature adjustment slider on the air conditioning control interface, and set the temperature to 26 degrees". The successful operation log can be used to further optimize the multimodal smart cockpit interaction model in the future.

[0156] If the operation fails, determine the cause of the failure. There may be many reasons for the failure, such as the target interface element being blocked and unable to operate, a temporary system failure, a network connection problem, etc. The specific cause is determined by analyzing the feedback information of the current interface, the system log, and the status data of the relevant hardware devices. For example, if the system prompts "the network connection is abnormal and the operation cannot be completed", the failure cause is determined to be a network problem. Determine and execute the repair measures corresponding to the pre-configured failure cause. If the failure cause is a network problem, you can try to reconnect to the network, such as automatically switching the network connection mode from 4G to Wi-Fi, or restarting the network module; if the target interface element is blocked, you can guide the user to manually adjust the interface display through the interface prompt, or try to automatically adjust the interface layout to display the blocked element. Record the failed operation log. The log content of the failed operation log can record information such as the failure time, the voice command of the operation, the corresponding target interface element, the specific operation content, the failure cause, and the repair measures performed. For example, record "[specific time], the user issued the 'adjust the air conditioner temperature to 26 degrees' command, operated the temperature adjustment slider on the air conditioner control interface, and the operation failed due to abnormal network connection. I have tried to reconnect to the network." The failed operation log helps in the subsequent troubleshooting and repair of the target smart cockpit system faults, and continuously improves the stability and reliability of the target smart cockpit system.

[0157] In practical applications, in order to further improve the performance and user experience of the smart cockpit see-and-say method, by collecting actual feedback data from users during use, analyzing the model's processing effects in different scenarios, such as voice command recognition errors and inaccurate positioning of interface elements, regularly optimize and update the multimodal smart cockpit interaction model, and adjust the model's parameters and structure in a targeted manner to improve the model's accuracy and robustness. At the same time, with the upgrading of vehicle hardware equipment and the emergence of new application scenarios, the coordinate information database and related knowledge of interface elements are constantly updated to adapt to the dynamic changes of the smart cockpit system.

[0158] Embodiment 4:

[0159] The following is an explanation of the smart cockpit visible and audible method provided by the present application through a specific embodiment. Figure 4 A schematic diagram of a process of a smart cockpit that can be seen and said is provided in an embodiment of the present application, and the process includes:

[0160] S401: Collecting a voice signal to be processed.

[0161] S402: Perform speech preprocessing on the speech signal to be processed.

[0162] The speech preprocessing includes but is not limited to one or more of the following: speech enhancement processing, framing of speech signals, and windowing operations.

[0163] S403: Determine the voice instruction to be processed based on the preprocessed voice feature vector of the voice signal to be processed.

[0164] S404: Determine the target interface element corresponding to the voice intent of the voice command to be processed based on the voice command to be processed through the pre-trained multimodal smart cockpit interaction model.

[0165] S405: Determine target coordinate information corresponding to the target interface element according to the pre-saved coordinate information corresponding to each interface element.

[0166] S406: Determine whether the target coordinate information is valid in the current interface of the target smart cockpit system. If so, execute S412; otherwise, execute S407.

[0167] S407: Obtain the latest interface image of the current interface by taking screenshot.

[0168] S408: Perform image preprocessing on the latest interface image.

[0169] The image preprocessing includes one or more of the following: grayscale processing, image noise reduction, and data enhancement.

[0170] S409: Re-identify each interface element and then determine its coordinate information.

[0171] S410: updating the coordinate information corresponding to each interface element according to the latest coordinate information corresponding to each interface element in the newly determined latest interface image.

[0172] S411: Based on the updated coordinate information corresponding to each interface element, the target coordinate information corresponding to the target interface element is determined again, and S406 is executed.

[0173] S412: operating the target interface element in the current interface according to the target coordinate information.

[0174] S413: Determine whether the operation on the target interface element is successful based on the feedback information of the current interface, if so, execute S414, otherwise, execute S415.

[0175] S414: Record a successful operation log.

[0176] S415: Determine the failure cause, determine and execute the repair measures corresponding to the pre-configured failure cause, and record the failure operation log.

[0177] Embodiment 5:

[0178] Based on the same inventive concept, the present application also provides a training device for a multi-modal intelligent cockpit interaction large model. Figure 5 A schematic diagram of a structure of a training device for a multimodal intelligent cockpit interaction large model provided in an embodiment of the present application, the device comprising:

[0179] The first acquisition unit 51 is used to acquire a first sample set including attribute information samples corresponding to a plurality of interface elements respectively; wherein the attribute information samples of any interface element include one or more of the following: operation-related information of the interface element, the interface element's own characteristics, text descriptions related to the interface element, and context information of the interface element, and each of the attribute information samples corresponds to a real functional meaning label;

[0180] The first training unit 52 is used to train the pre-trained multimodal basic big model based on each of the attribute information samples in the first sample set and the real functional meaning labels corresponding to each of the attribute information samples, so as to obtain a trained multimodal smart cockpit interaction big model; wherein the multimodal basic big model is a big model that has multimodal data processing capabilities and natural language understanding capabilities, and the multimodal smart cockpit interaction big model has the ability to identify the functional meanings of the interface elements contained in each interface in the target smart cockpit system;

[0181] A second acquisition unit 53 is used to acquire a second sample set including a plurality of voice command samples for interacting with the target smart cockpit system; wherein any voice command sample corresponds to a voice intent label and an interface element label;

[0182] The second training unit 54 is used to train the multimodal smart cockpit interaction big model based on each of the voice command samples in the second sample set, and the voice intent label and interface element label of each of the voice command samples, so that the multimodal smart cockpit interaction big model learns the mapping relationship between voice intent and the interface elements of each interface image in the target smart cockpit system.

[0183] The training device of the multimodal smart cockpit interaction large model in this embodiment is presented in the form of functional modules, where the modules refer to application specific integrated circuits (ASICs), processors and memories that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0184] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0185] Embodiment 6:

[0186] Based on the same inventive concept, the present application also provides a smart cockpit visible and audible device. Figure 6 A schematic diagram of a smart cockpit visible and speaking device provided in an embodiment of the present application, the device comprising:

[0187] An acquisition module 61 is used to acquire a voice instruction to be processed;

[0188] An interface element determination module 62 is used to determine a target interface element corresponding to the voice intent of the voice instruction to be processed based on the voice instruction to be processed by using a pre-trained multi-modal smart cockpit interaction model;

[0189] A coordinate determination module 63, used to determine target coordinate information corresponding to the target interface element according to pre-stored coordinate information corresponding to each interface element;

[0190] The operating module 64 is used to operate the target interface element in the current interface of the target smart cockpit system according to the target coordinate information.

[0191] The smart cockpit in this embodiment can be seen to be presented in the form of a functional module, where the module refers to an application specific integrated circuit (ASIC), a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0192] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0193] Embodiment 7:

[0194] See also Figure 7 , Figure 7 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present application, such as Figure 7 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 7 A processor 10 is taken as an example.

[0195] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0196] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0197] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of a computer device based on the presentation of a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0198] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0199] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 7 The example of connecting through bus is taken in the following.

[0200] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0201] Embodiment 8:

[0202] On the basis of the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor implements the following steps when executing:

[0203] Obtaining a first sample set including attribute information samples corresponding to a plurality of interface elements respectively; wherein the attribute information samples of any interface element include one or more of the following: operation-related information of the interface element, the interface element's own characteristics, text descriptions related to the interface element, and context information of the interface element, and each of the attribute information samples corresponds to a real functional meaning label;

[0204] Based on each of the attribute information samples in the first sample set and the real functional meaning labels corresponding to each of the attribute information samples, the pre-trained multimodal basic big model is trained to obtain a trained multimodal smart cockpit interaction big model; wherein the multimodal basic big model is a big model that has multimodal data processing capabilities and natural language understanding capabilities, and the multimodal smart cockpit interaction big model has the ability to identify the functional meanings of the interface elements contained in each interface in the target smart cockpit system;

[0205] Acquire a second sample set including a plurality of voice command samples for interacting with the target smart cockpit system; wherein any voice command sample corresponds to a voice intent label and an interface element label;

[0206] Based on each of the voice command samples in the second sample set, and the voice intent label and interface element label of each of the voice command samples, the multimodal smart cockpit interaction big model is trained so that the multimodal smart cockpit interaction big model learns the mapping relationship between voice intent and the interface elements of each interface image in the target smart cockpit system.

[0207] Since the principle of solving the problem by the above-mentioned computer-readable storage medium is similar to the training method of the multimodal smart cockpit interaction large model, the implementation of the above-mentioned computer-readable storage medium can refer to Examples 1-2 of the method, and the repeated parts will not be repeated.

[0208] Embodiment 9:

[0209] On the basis of the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor implements the following steps when executing:

[0210] Get the voice commands to be processed;

[0211] Determine, by using a pre-trained multimodal smart cockpit interaction model, a target interface element corresponding to the voice intent of the voice command to be processed based on the voice command to be processed;

[0212] Determine the target coordinate information corresponding to the target interface element according to the pre-saved coordinate information corresponding to each interface element;

[0213] According to the target coordinate information, the target interface element is operated in the current interface of the target smart cockpit system.

[0214] Since the principle of solving the problem by the above-mentioned computer-readable storage medium is similar to the smart cockpit visible-as-it-can method, the implementation of the above-mentioned computer-readable storage medium can refer to Examples 3-4 of the method, and the repeated parts will not be repeated.

[0215] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A training method for a multimodal intelligent cockpit interaction large model, characterized in that: The method comprises: Obtaining a first sample set including attribute information samples corresponding to a plurality of interface elements respectively; wherein the attribute information samples of any interface element include one or more of the following: operation-related information of the interface element, the interface element's own characteristics, text descriptions related to the interface element, and context information of the interface element, and each of the attribute information samples corresponds to a real functional meaning label; Based on each of the attribute information samples in the first sample set and the real functional meaning labels corresponding to each of the attribute information samples, the pre-trained multimodal basic big model is trained to obtain a trained multimodal smart cockpit interaction big model; wherein the multimodal basic big model is a big model that has multimodal data processing capabilities and natural language understanding capabilities, and the multimodal smart cockpit interaction big model has the ability to identify the functional meanings of the interface elements contained in each interface in the target smart cockpit system; Acquire a second sample set including a plurality of voice command samples for interacting with the target smart cockpit system; wherein any voice command sample corresponds to a voice intent label and an interface element label; Based on each of the voice command samples in the second sample set, and the voice intent label and interface element label of each of the voice command samples, the multimodal smart cockpit interaction big model is trained so that the multimodal smart cockpit interaction big model learns the mapping relationship between voice intent and the interface elements of each interface image in the target smart cockpit system.

2. The method according to claim 1, characterized in that The inherent characteristics of each interface element are obtained from the interface image in the following manner: For any of the interface images, various types of interface elements contained in the interface image are determined through an image recognition algorithm; for interface elements in the interface image that are buttons, the visual features and bounding box coordinates of each button in the interface image are determined through a pre-trained button recognition algorithm; for interface elements in the interface image that are text boxes, the text boxes in the interface image are converted into editable text strings through optical character recognition technology, and the bounding box coordinates of the area where the text boxes are located are recorded.

3. The method according to claim 1 or 2, characterized in that Acquiring an interface image in the target smart cockpit system, including: The interface image of the target smart cockpit system is obtained by periodically taking screenshots or recording the screen.

4. The method according to claim 3, characterized in that The method further comprises: Perform image preprocessing on the interface image; wherein the image preprocessing includes one or more of the following: grayscale processing, image noise reduction, and data enhancement.

5. The method according to claim 1, characterized in that The multimodal smart cockpit interaction big model is trained based on each of the voice command samples in the second sample set, and the voice intent label and the interface element label of each of the voice command samples, so that the multimodal smart cockpit interaction big model learns the mapping relationship between the voice intent and the interface elements of each interface image in the target smart cockpit system, including: Obtain any of the voice command samples and its corresponding voice intent label and interface element label; Outputting, through the multimodal smart cockpit interaction macro model, a first probability distribution of the predicted voice intent of the voice command sample and a second probability distribution of the interface elements corresponding to the predicted voice intent based on the voice command sample and the learned functional meanings of the interface elements contained in each interface of the target smart cockpit system; Based on the first probability distribution and the voice intent label, and the second probability distribution and the interface element label, the multimodal smart cockpit interaction big model is trained so that the multimodal smart cockpit interaction big model learns the mapping relationship between voice intent and interface elements.

6. The method according to claim 1, characterized in that Any of the voice command samples is obtained in the following manner: Performing speech preprocessing on the acquired speech signal; wherein the speech preprocessing includes one or more of the following: speech enhancement processing, framing of speech signals, and windowing operations; A speech feature vector of the preprocessed speech signal is obtained, and the speech feature vector is determined as the speech instruction sample.

7. A method for making a smart cockpit visible and audible based on a model trained by the method described in any one of claims 1 to 6, characterized in that: The method comprises: Get the voice commands to be processed; Determine, by using a pre-trained multimodal smart cockpit interaction model, a target interface element corresponding to the voice intent of the voice command to be processed based on the voice command to be processed; Determine the target coordinate information corresponding to the target interface element according to the pre-saved coordinate information corresponding to each interface element; According to the target coordinate information, the target interface element is operated in the current interface of the target smart cockpit system.

8. The method according to claim 7, characterized in that After determining the target coordinate information corresponding to the target interface element according to the pre-saved coordinate information corresponding to each interface element, and before operating the target interface element in the current interface of the target smart cockpit system according to the target coordinate information, the method further includes: Determine whether the target coordinate information is valid in the current interface.

9. The method according to claim 8, characterized in that If it is determined that the target coordinate information is invalid in the current interface, the method further includes: Obtain the latest interface image of the current interface; Re-determining the latest coordinate information corresponding to each interface element in the latest interface image; According to the newly determined latest coordinate information corresponding to each interface element in the latest interface image, the coordinate information corresponding to each interface element is updated; Determine the target coordinate information corresponding to the target interface element based on the updated coordinate information corresponding to each interface element; The target interface element is operated in the current interface according to the target coordinate information.

10. The smart cockpit visible and audible method according to claim 7 or 9, characterized in that: After operating the target interface element in the interface of the target smart cockpit system according to the target coordinate information, the method further includes: Determining whether the operation on the target interface element is successful according to the feedback information of the current interface; If the operation is successful, the successful operation log is recorded; If the operation fails, the cause of the failure is determined; the pre-configured repair measures corresponding to the failure cause are determined and executed; and the failed operation log is recorded.

Citation Information

Patent Citations

  • Dialogue type intelligent interaction method and system based on natural language processing

    CN112487142A

  • Voice control method of user interface, control device and computer readable medium

    CN115775557A

  • UI interface design and man-machine interaction method based on voice instruction

    CN116243826A

  • Voice interaction method, server and computer readable storage medium

    CN117373456A

  • Control positioning model training method, control positioning method and device and control triggering method and device

    CN118734244A

Cited By

  • Interface anomaly detection method based on multi-modal large model comparison and dual verification

    CN120472282A

  • Interface operation instruction generation method, electronic equipment, storage medium and program product

    CN120704792A

  • Voice interaction method and system based on multi-stage large model

    CN121011180A