Cabin control method and device, vehicle-mounted equipment and computer program product
By locally deploying neural network models in the smart cockpit and using user historical data for training, the privacy security and response delay problems caused by network performance limitations in the prior art are solved, and faster, secure and intelligent cockpit services are achieved.
Patent Information
- Application Number
- CN202510583883.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
AI Technical Summary
The existing smart cockpit technology is limited by network performance when using large models to provide cockpit services, resulting in user privacy security and delayed response.
Deploy neural network models locally in the cockpit, and use the user's historical data for local training to generate cockpit services, thereby avoiding dependence on the cloud.
Through local training models, the response speed and privacy and security of cockpit services are improved, and more intelligent and personalized cockpit services are provided.
Smart Images

Figure CN120096606A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a vehicle intelligent cockpit technology, and in particular to a cockpit control method, device, vehicle-mounted equipment and computer program product. Background Art
[0002] At present, compared with traditional fuel vehicles, the outstanding changes of new energy vehicles are reflected in the application of intelligent technology in the cockpit. The progress of intelligent cockpit helps users control the car more conveniently. Among them, the diverse interaction methods can be said to be a major feature of intelligence. In addition to traditional button operation, the interaction of voice signals and multi-modal data has come on stage.
[0003] As a new way of human-computer interaction, the superiority of the interaction of voice signals and multimodal data lies in its ease of use and parallel task processing capabilities. In some cases where both hands are occupied, it can still help users perform operations, while reducing the user's learning cost for hard switches in the cockpit. However, precisely because of the convenience of these interactive methods, users have higher and higher requirements for services in the cockpit, and the current interactive methods are limited by network performance requirements. When using large models for cockpit services, large models can only be obtained by calling cloud services, which has brought people's concerns about privacy security and response delays. Summary of the invention
[0004] The embodiments of the present invention provide a cockpit control method, device, vehicle-mounted equipment and computer program product, which can improve the intelligence of cockpit services without relying on the cloud.
[0005] The technical solution of the present invention is achieved in this way: An embodiment of the present invention provides a cockpit control method, including: Acquire multimodal data of users in the vehicle cabin; Inputting the multimodal data of the user into the current neural network model to generate a cockpit service; Execute the cockpit function corresponding to the cockpit service; The current neural network model is obtained by locally training a preset neural network model using the historical data of the user in the cockpit; the historical data of the user is: the multimodal data of the user generated locally and the cockpit service corresponding to the multimodal data of the user generated locally; Among them, the user's multimodal data includes at least: the user's personal information and the user's environmental information; the user's personal information includes at least one or more of the following dimensions: information on the user's personal statistical dimension, information on the user's behavioral dimension, and information on the user's emotional dimension.
[0006] In this way, by deploying the neural network model locally in the cockpit and using the user's historical data to train the model, the model used in the cockpit control is not limited by the network performance. Not only can the cockpit control ensure the privacy and security of the user, but it can also respond to the user's needs in a timely manner, thereby improving the response speed of the cockpit service. In the cockpit, the model is trained locally using the generated user's personal information and user's environmental information, and the personal information includes one or more of the user's personal statistical dimension information, the user's behavioral dimension information, and the user's emotional dimension information. Therefore, the obtained model can take into account each data in the multimodal data to provide users with more intelligent cockpit services.
[0007] Furthermore, the method further comprises: In the case where there are multiple users, determining a user portrait of each user according to multimodal data in historical data of each of the users; The historical data of each user is updated according to the similarity of the user portraits between the users.
[0008] In this way, the user's historical data can be further expanded through user profiling to optimize local training data, thereby further improving the accuracy of the current neural network model obtained by pre-fusion and post-fusion, and providing users with more intelligent cockpit services.
[0009] Further, updating the historical data of each user according to the similarity of the user portraits among the users includes: In a case where the users include a first user and a second user and historical data of the first user is generated, determining a similarity between a user profile of the first user and a user profile of the second user; When the similarity is greater than a first preset threshold, the historical data of the first user is added to the historical data of the second user.
[0010] In this way, by comparing the similarity with the first preset threshold, it is determined whether to add the historical data of the first user to the historical data of the second user. In this way, the historical data between users with similar user portraits can be shared, so that the historical data of each user can be accumulated as soon as possible, which accelerates the local training process of the neural network model, optimizes the training data of the local model training, and improves the accuracy of the current neural network model.
[0011] Furthermore, the method further comprises: In the case where the user is a plurality of users and historical data is generated, determining the adoption degree of the generated historical data for each of the users; When the adoption degree is greater than a second preset threshold, the historical data is added to the historical data of the user corresponding to the adoption degree.
[0012] In this way, by comparing the degree of adoption with the second preset threshold, it is determined whether to add historical data to the user's historical data. In this way, the historical data that the user can adopt can be added to his historical data, thereby optimizing the user's historical data, and then optimizing the training data for local model training, thereby improving the accuracy of the current neural network model.
[0013] Furthermore, the method further comprises: Performing statistics on the historical data of the user to obtain statistical results; In the case where the statistical result satisfies the first preset condition, the preset neural network model is trained using the historical data of the user corresponding to the statistical result to obtain the current neural network model.
[0014] In this way, by statistics on the user's historical data and setting the first preset condition, local training of the preset neural network model is achieved, so that the model can be trained based on the user's historical data after the vehicle is put into use, so that the model can accurately provide users with intelligent cockpit services.
[0015] Furthermore, the method further comprises: When the statistical result indicates that the number of historical data generated for the user reaches a third preset threshold, determining that the statistical result satisfies the first preset condition; When the statistical result indicates that the duration of generating the historical data of the user reaches a preset duration, it is determined that the statistical result meets the first preset condition.
[0016] In this way, whether to train the preset neural network model is determined by measuring whether the number of historical data of the user reaches the third preset threshold, so that the on-board equipment can only perform training when the number of historical data of the user reaches a certain number, so that the training can be started after the training data used in model training reaches a certain scale, thereby improving the accuracy of model training; whether to train the preset neural network model is determined by measuring whether the generation time of the historical data of the user reaches the preset time, so that the on-board equipment can only perform training when the generation time of the historical data of the user reaches the preset time, so that the training can be started after the training data used in model training reaches a certain scale, thereby improving the accuracy of model training.
[0017] Furthermore, the method further comprises: In the case of obtaining the current neural network model, clearing the statistical result, returning to the step of performing statistics on the historical data of the user to obtain the statistical result, until the statistical result meets the second preset condition; When the statistical result satisfies the second preset condition, the local neural network model is trained using the historical data of the user to obtain the current neural network model; When the statistical result satisfies the first preset condition for a number of times reaching a preset number, it is determined that the statistical result satisfies the second preset condition.
[0018] In this way, by cyclically counting the user's historical data, multiple training of the preset neural network model is achieved to update the current neural network model, so that the training data of the current neural network model is the user's historical data generated in the recent period, so that the accuracy of the current neural network model can be improved in a short period of time; by using the user's historical data to train the preset neural network model when the above statistical results meet the second preset condition, it is possible to train the preset neural network model using training data of a sufficient scale, and the accuracy of the current neural network model can be further improved by using the front fusion and post fusion methods; by counting the statistical results, whether the number of times the first preset condition is met meets the preset number determines whether the statistical results meet the second preset condition, and then determines whether the user's historical data reaches a certain scale, and when a certain scale is reached, the preset neural network model is trained using the user's historical data accumulated by the front fusion, thereby further improving the accuracy of the current neural network model obtained.
[0019] Furthermore, the method further comprises: After acquiring multimodal data of a user in a cabin of a vehicle, determining the current neural network model from a set of neural network models; Among them, the neural network model of each user in the neural network model set is obtained by locally training a preset neural network model using the historical data of each user in the cabin; the historical data of each user is: the multimodal data of each user generated locally and the cabin service corresponding to the multimodal data of each user generated locally.
[0020] In this way, by using the historical data of each user to train the preset neural network model, the neural network model of each user can be obtained separately. In this way, when processing the user's multimodal data, a model can be determined from the neural network model set to process the user's multimodal data to generate cockpit services. In this way, the required cockpit services can be provided to the user in a targeted manner, thereby improving the intelligence of the cockpit services.
[0021] Furthermore, the method further comprises: In the case where there are multiple users in the cabin, determining a target user from the multiple users; Determine the neural network model of the target user from the set of neural network models; The neural network model of the target user is determined as the current neural network model.
[0022] In this way, the current neural network model is determined by determining the target user, so that the current network model is related to the determined target user, which is conducive to realizing cockpit services with a certain user as the core in multiple user scenarios, so that suitable neural network models can be selected under multiple users to provide cockpit services for multiple users.
[0023] Further, the determining the target user from the multiple users includes at least one of the following: Determining the target user from the multiple users according to the multimodal data of the user; Determining the target user according to the priority of the user; In a case where the multimodal data of the user includes a voice signal, the user who sends the voice signal is determined as the target user.
[0024] In this way, the target user is determined from the users through the multimodal data of each user mentioned above, so that the determined target user is related to the user's multimodal data. After the target user is determined through the priority of each user, the determined neural network model is related to the priority of each user, so that the needs of specific users can be given priority, so that the smart cockpit can select the appropriate neural network model according to the priority to provide cockpit services to users; users who send voice signals can also be used as target users in passive services. In this way, a suitable current neural network model can be determined for this service, which is conducive to providing users with more accurate cockpit services and further improving the user experience in the cockpit.
[0025] An embodiment of the present invention provides a cockpit control device, including: An acquisition module, used to acquire multimodal data of a user in a vehicle cabin; A generation module, used for inputting the multimodal data of the user into the current neural network model to generate a cockpit service; An execution module, used for executing a cockpit function corresponding to the cockpit service; The current neural network model is obtained by locally training a preset neural network model using the historical data of the user in the cockpit; the historical data of the user is: the multimodal data of the user generated locally and the cockpit service corresponding to the multimodal data of the user generated locally; Among them, the user's multimodal data includes at least: the user's personal information and the user's environmental information; the user's personal information includes at least one or more of the following dimensions: information on the user's personal statistical dimension, information on the user's behavioral dimension, and information on the user's emotional dimension.
[0026] An embodiment of the present invention provides a vehicle-mounted device, comprising: a processor and a storage medium storing instructions executable by the processor, wherein the storage medium relies on the processor to perform operations via a communication bus, and when the instructions are executed by the processor, the cockpit control method described in one or more of the above embodiments is executed.
[0027] An embodiment of the present invention further provides a computer program product, including a computer program or instructions, wherein when the computer program or instructions are executed by a processor, the steps of the cockpit control method described in one or more of the above embodiments are implemented.
[0028] Beneficial effects of the present invention: (1) By deploying the neural network model locally in the cockpit and using the user's historical data to train the model, the model used in the cockpit control is not limited by network performance. This not only ensures the privacy and security of the user, but also responds to the user's needs in a timely manner, thus improving the response speed of the cockpit service. (2) The model is trained locally in the cockpit using the generated personal information and environmental information of the user, and the personal information includes one or more of the information of the user's personal statistical dimension, the user's behavioral dimension, and the user's emotional dimension, so that the obtained model can take into account each data in the multimodal data to provide the user with a more intelligent cockpit service. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A schematic flow chart of an optional cockpit control method provided in an embodiment of the present invention; Figure 2 A schematic diagram of a user portrait provided in an embodiment of the present application; Figure 3 A flowchart of Example 1 of an optional cockpit control method provided by an embodiment of the present invention; Figure 4 A flowchart of a second example of an optional cockpit control method provided by an embodiment of the present invention; Figure 5 A flowchart of Example 3 of an optional cockpit control method provided by an embodiment of the present invention; Figure 6 A flowchart of a fourth example of an optional cockpit control method provided by an embodiment of the present invention; Figure 7 A flowchart of Example 5 of an optional cockpit control method provided by an embodiment of the present invention; Figure 8 A schematic diagram of the structure of an optional cockpit control device provided in an embodiment of the present invention; Fig. 9 A schematic structural diagram of an optional vehicle-mounted device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0031] In view of the problem that cockpit control in the related art is limited by network performance, an embodiment of the present invention provides a cockpit control method. Figure 1 A flowchart of an optional cockpit control method provided by an embodiment of the present invention is shown as follows: Figure 1 As shown, the cockpit control method may include: S101: Acquire multimodal data of a user in a vehicle cabin; For cockpit control, usually a large model trained in the cloud is called to control the cockpit and provide cockpit services to users, thereby meeting the needs of users in the cockpit. However, this method is limited by network performance. When the network performance is poor, the response speed of cockpit control will be affected, and the privacy of users cannot be effectively protected.
[0032] In order to improve the response speed of cockpit services while improving the intelligence of cockpit services, in an embodiment of the present invention, a neural network model is deployed locally. Here, after acquiring the neural network model, electronic devices other than the vehicle-mounted equipment train it to obtain the neural network model, and deploy it in the vehicle-mounted equipment of the vehicle as a local neural network model to enable users to control the cockpit in the initial stage of the vehicle being put into use.
[0033] In S101, the vehicle-mounted device obtains multimodal data of users in the vehicle cabin. Here, the users in the cabin may be one user or multiple users. Here, multiple users refer to two or more users, which is not specifically limited in the embodiment of the present invention.
[0034] Among them, the user's multimodal data may include: text data, image data, audio data, video data, behavior data and physiological data, etc. collected by the vehicle-mounted equipment through components on the vehicle. Here, the embodiment of the present invention does not make specific limitations on this.
[0035] For multiple users, the in-vehicle equipment needs to obtain the multimodal data of all users. For example, there are user A and user B in the cabin, then the multimodal data of the users obtained include: user A's multimodal data and user B's multimodal data. At this time, user A's multimodal data should include user B's multimodal data, and user B's multimodal data should include user A's multimodal data.
[0036] In addition, the user's multimodal data may include the user's personal information, and may also include environmental information of the user's environment, and of course, may also include vehicle status information, which is not specifically limited in the embodiments of the present invention.
[0037] S102: Inputting the user's multimodal data into the current neural network model to generate a cockpit service; After the vehicle-mounted device acquires the multimodal data of the user through the above-mentioned S101, in S102, the vehicle-mounted device can input the multimodal data of the user into the current neural network model to generate a cockpit service.
[0038] Among them, the current neural network model is obtained by locally training the preset neural network model using the historical data of the users in the cockpit. That is to say, after obtaining the multimodal data of the user, the multimodal data of the user can be input into the model obtained by locally training the preset neural network model using the historical data of the users in the cockpit. The historical data of the user is: the multimodal data of the user generated locally and the cockpit service corresponding to the multimodal data of the user generated locally. That is, the preset neural network model is trained using the multimodal data of the user that has been generated and the cockpit service corresponding to the multimodal data of the user that has been generated locally to obtain the current neural network model, and the cockpit service is generated for the acquired multimodal data of the user using the current neural network model.
[0039] It should be noted that here, in addition to directly inputting the user's multimodal data into the current neural network model to generate a cockpit service, the user's multimodal data may also be preprocessed, for example, data cleaning, and then the preprocessed multimodal data is input into the current neural network model to generate a cockpit service, which is not specifically limited in the embodiments of the present invention.
[0040] Among them, the above-mentioned current neural network model can be the only model deployed locally, or it can be a model selected from a set of locally deployed models. Here, the embodiment of the present invention does not make specific limitations on this.
[0041] Furthermore, the above-mentioned training of the preset neural network model may be obtained by a single training method, multiple training methods, or continuous training methods, and the embodiments of the present invention do not specifically limit this.
[0042] In addition, for a single user, a neural network model of the user can be obtained, and for multiple users, a neural network model of each user can be obtained.
[0043] In addition, with respect to the cockpit service, if the user's multimodal data contains imperative data, for example, a imperative voice signal, the cockpit service is a passive service; if the user's multimodal data does not contain imperative data, the cockpit service is an active service. It can be seen that the above-mentioned cockpit service can be a passive service or an active service. Here, the embodiments of the present invention do not make specific limitations on this.
[0044] S103: Execute the cockpit function corresponding to the cockpit service.
[0045] After the on-board device generates the cabin service through the above S102, the on-board device controls the components corresponding to the cabin service according to the cabin function corresponding to the cabin service to realize the function corresponding to the cabin service. For example, if the cabin service is to turn on the air conditioning for cooling, then the on-board device controls the air conditioning device to turn on and cool.
[0046] Regarding the above-mentioned multimodal data of the user, in an optional embodiment, the multimodal data of the user includes at least: personal information of the user and environmental information of the user; The user's personal information includes at least one or more of the following dimensions: information on the user's personal statistics dimension, information on the user's behavior dimension, and information on the user's emotion dimension.
[0047] It is understandable that the multimodal data of the above-mentioned user can be collected from sensors, microphones, cameras and other components deployed on the vehicle, so that the user's personal information and the user's environmental information can be collected. The user's personal information can include information in multiple dimensions, such as information on the user's personal statistics dimension, information on the user's behavior dimension, and information on the user's emotion dimension.
[0048] The information of the user's personal statistics dimension may include: user age group, gender, permanent residence area, religious belief, family structure and health status, etc. The information of the user's behavior dimension may include: in-car habitual behavior, car usage preference, operation preference, function preference, driving preference, driving habitual behavior, travel habits and travel purpose, etc. The information of the emotional dimension may include: emotional reactions under various conditions, such as traffic conditions, driving pressure, time pressure, vehicle failure, communication, navigation error, long-distance driving, in-car environment and bad weather, etc.
[0049] In addition, the user's environmental information may include temperature, humidity and noise inside and outside the cabin, and may also include window information and seat information in the cabin, etc. Here, the embodiment of the present invention does not make specific limitations on this.
[0050] Of course, the vehicle status information can also be obtained through the above multimodal data. For example, the vehicle speed, the state of the brake pedal, the state of the brake pedal, the state of the steering wheel, etc. can all be obtained from the user's multimodal data.
[0051] In this way, by acquiring the user's multimodal data and using the current neural network model to generate cockpit services for the user, the model can take into account multiple aspects of the user's data in the cockpit, thereby further improving the intelligence of the cockpit service generated by the model.
[0052] In order to implement local training of a preset neural network model, in an optional embodiment, the method may further include: Perform statistics on the user's historical data to obtain statistical results; When the statistical result satisfies the first preset condition, the preset neural network model is trained using the historical data of the user corresponding to the statistical result to obtain the current neural network model.
[0053] It can be understood that after the vehicle is put into use, historical data of users will be generated one by one in the cockpit. For a single user, the historical data of the single user can be counted to obtain the statistical results of the single user, and judge whether the statistical results of the single user meet the first preset condition. If so, the historical data of the single user corresponding to the statistical results of the single user is used to train the preset neural network model, so as to obtain the current neural network model of the single user.
[0054] For multiple users, the historical data of each user among the multiple users is counted to obtain the statistical results of each user, and it is determined whether the statistical results of each user meet the first preset condition. If so, the preset neural network model is trained using the historical data of each user corresponding to the statistical results of each user, so as to obtain the current neural network model of each user.
[0055] The historical data of the user corresponding to the above statistical result is: the historical data of the user used to obtain the statistical result.
[0056] The first preset condition may be a condition of a parameter that limits the user's historical data. When the first preset condition is met, the preset neural network model is trained using the user's historical data corresponding to the statistical result, thereby obtaining the current neural network model. When the first preset condition is not met, the user's historical data continues to be counted to obtain statistical results until the obtained statistical results meet the first preset condition.
[0057] In this way, by statistics on the user's historical data and setting the first preset condition, local training of the preset neural network model is achieved, so that the model can be trained based on the user's historical data after the vehicle is put into use, so that the model can accurately provide users with intelligent cockpit services.
[0058] Regarding the above statistical results satisfying the first preset condition, in an optional embodiment, the above method may further include at least one of the following: When the statistical result indicates that the number of historical data of the generated user reaches a third preset threshold, determining that the statistical result satisfies the first preset condition; When the statistical result indicates that the duration of generating the historical data of the user reaches a preset duration, it is determined that the statistical result satisfies the first preset condition.
[0059] It can be understood that the above-mentioned statistics on the user's historical data can be a statistics on the number of the user's historical data. When the statistical result indicates that the number of the user's historical data generated reaches a third preset threshold, it means that the user's historical data has met the training conditions. Therefore, it is determined that the statistical result meets the first preset condition, so that the preset neural network model is trained using the user's historical data to obtain the current neural network model.
[0060] It should be noted that the third preset threshold value can be flexibly set according to the user's usage frequency, usage habits and / or user needs.
[0061] In this way, by counting whether the number of historical data of the user reaches the first preset threshold, it is determined whether to train the preset neural network model, so that the vehicle-mounted equipment can only perform training when the number of historical data of the user reaches a certain number, so that the training is started after the training data used in the model training reaches a certain scale, thereby improving the accuracy of the model training.
[0062] In addition, the above-mentioned statistics on the user's historical data may be statistics on the generation time of the user's historical data. When the statistical result indicates that the generation time of the user's historical data reaches a preset time, it means that the user's historical data has met the training conditions. Therefore, it is determined that the statistical result meets the first preset condition, so that the preset neural network model is trained using the user's historical data to obtain the current neural network model.
[0063] It should be noted that the above preset duration can be flexibly set according to the user's usage frequency, usage habits and / or user needs.
[0064] In this way, by calculating whether the generation time of the user's historical data reaches the preset time, it is determined whether to train the preset neural network model, so that the vehicle-mounted equipment can only perform training when the generation time of the user's historical data reaches the preset time, so that the training is started after the training data used in the model training reaches a certain scale, thereby improving the accuracy of the model training.
[0065] In order to further improve the accuracy of the model, in an optional embodiment, the above method may further include: When the current neural network model is obtained, the statistical result is cleared, and the step of performing statistics on the user's historical data to obtain the statistical result is returned until the statistical result meets the second preset condition.
[0066] It can be understood that after the preset neural network model is trained to obtain the current neural network model after the first preset condition is met, the statistical result can be cleared and the statistical process of the user's historical data can be returned to be performed until the statistical result meets the second preset condition. In other words, the statistical process of the user's historical data is performed cyclically, and after each training of the preset neural network model, the statistical result is cleared and the statistical process is performed again, so that the preset neural network model can be trained each time the statistical result meets the first preset condition to update the current neural network model.
[0067] In addition, the termination condition of the above loop condition is that the statistical result satisfies the second preset condition, that is, after the loop executes the training of the preset neural network model, the statistical result reaches the second preset condition to end the loop.
[0068] In this way, through the cyclic statistics of the user's historical data, the preset neural network model can be trained multiple times to update the current neural network model, so that the training data of the current neural network model is the user's historical data generated in the recent period, thereby improving the accuracy of the current neural network model in a short period of time.
[0069] In the case where the statistical result meets the second preset condition, in an optional embodiment, the method may further include: When the statistical result meets the second preset condition, the preset neural network model is trained using the user's historical data to obtain the current neural network model; When the number of times that the statistical result satisfies the first preset condition reaches a preset number, it is determined that the statistical result satisfies the second preset condition.
[0070] It can be understood that by judging the statistical results and determining that the statistical results meet the second preset condition, it means that the user's historical data has reached a certain scale, and the accumulated user's historical data can be used to train the preset neural network model, thereby updating the current neural network model.
[0071] It should be noted that when the above statistical results meet the first preset condition for training the preset neural network model, it can be called pre-fusion, and when the above statistical results meet the second preset condition for training the preset neural network model, it can be called post-fusion. It can be seen that in the embodiment of the present invention, after the preset neural network model is pre-fused multiple times, the preset neural network model is post-fused, and the current neural network model is updated by using this pre-fusion and post-fusion method.
[0072] In addition, for single-user or multi-user scenarios, multiple pre-fusions are performed on the historical data of each user, and then post-fusion is performed to obtain the current neural network model of each user.
[0073] In this way, by using the user's historical data to train the preset neural network model when the above statistical results meet the second preset condition, it is possible to train the preset neural network model using training data of a sufficient scale, and the accuracy of the current neural network model can be further improved by using the pre-fusion and post-fusion methods.
[0074] In order to determine whether the statistical result meets the second preset condition, here, the number of times the statistical result meets the first preset condition is counted, and it is determined whether the number of times the statistical result meets the second preset condition reaches the preset number. If so, it means that the number of front fusions is sufficient and the post fusion can be triggered. Therefore, here, it is determined that the statistical result meets the second preset condition. If not, it means that the number of front fusions is not enough and the front fusion needs to continue. Therefore, it is determined that the statistical result does not meet the second preset condition.
[0075] In this way, by counting the statistical results, whether the number of times the first preset condition is met meets the preset number of times will be used to determine whether the statistical results meet the second preset condition, and then determine whether the user's historical data has reached a certain scale. When a certain scale is reached, the historical data of the user accumulated by the previous fusion is used to train the preset neural network model, thereby further improving the accuracy of the current neural network model.
[0076] Since the user in the embodiment of the present invention can be a single user or multiple users, for multiple users, in order to expand the historical data of the users so as to realize the preset neural network model as soon as possible, in an optional embodiment, the above method may also include: In the case where there are multiple users, determining a user profile of each user based on multimodal data in historical data of each of the users; Update each user's historical data based on the similarity of user portraits between users in each user group.
[0077] It can be understood that in the case where there are multiple users, in order to update the historical data of each user, here, the user portrait of each user can be first determined based on the multimodal data in the historical data of each user, and then the historical data of each user can be updated based on the similarity of the user portraits between users.
[0078] The reason for adopting the above method is to take into account the correlation between users in the cockpit and the similarity of habits between users. Therefore, the historical data generated by a certain user can be used as the historical data of other users. Here, the historical data of each user can be determined by the similarity of user portraits between users.
[0079] Among them, the multimodal data of each user's historical data can be used to determine the user portrait. Figure 2 A schematic diagram of an optional user portrait provided by an embodiment of the present invention, such as Figure 2 As shown, user portraits can be divided into at least three dimensions of information, which may include: demographic dimension, behavioral dimension, and emotional response dimension under various conditions (equivalent to the above-mentioned emotional dimension).
[0080] exist Figure 2 In the data, the personal statistics dimension information may include: user age group (e.g., 20-30 years old, 30-40 years old, etc.), gender, permanent residence area, religious belief, family structure (e.g., single-child, elderly family), health status (e.g., cold, health). This dimension information is mainly obtained by analyzing the user's facial information, voiceprint, travel data, liveness detection and other information.
[0081] Among them, the information of the behavior dimension may include: in-car habitual behavior (e.g., sleeping, smoking, eating fast food, etc.), car use preferences (e.g., often sitting in the front passenger seat, etc.), operation preferences (e.g., preference to use voice), function preferences (e.g., preference to use AAA to listen to music instead of BBB), driving preferences (e.g., preference for power recovery), driving habitual behavior (e.g., sight habit deviation, one-handed driving), travel habits (e.g., short-distance, long-distance, etc.), travel purpose (e.g., commuting, traveling). This dimension information is mainly obtained by analyzing cockpit logs, in-cabin image recognition, sight recognition, travel data and other information.
[0082] Here, emotional responses under various conditions may include: traffic conditions (e.g., irritability in traffic jams, nervousness when merging lanes), driving stress (e.g., reversing, unfamiliar routes), time pressure, vehicle failure, communication (e.g., irritability when talking on the phone, annoyance when talking to children in the car), navigation errors, long-distance driving, in-car environment (e.g., high humidity in the car), and bad weather. This dimension of information is mainly obtained by analyzing voiceprints, facial information, weather, and other information.
[0083] After obtaining the user portrait of each user, the similarity between the user portraits of two users can be calculated to evaluate whether the cockpit service automatically issued by the vehicle computer under the same conditions is suitable for different users. After obtaining the similarity, the user portrait labels that need to complement each other and the applicable training data can be determined based on the similarity.
[0084] Among them, in order to obtain the similarity of the user portraits of two users, after obtaining the user portraits, each user portrait can be vectorized to obtain the vector corresponding to each user portrait, and then the similarity between the vectors corresponding to the two user portraits can be calculated. The purpose of vectorization is to transform the background and behavioral tendencies of different users from non-quantifiable information dimensions into quantifiable coordinate dimensions. For each variable in the user portrait, its attributes are first transformed into discrete variables or continuous variables, and then the vector is normalized or standardized based on the transformed vectorized user portrait samples to reduce the noise impact of some variables due to fluctuations in their own value ranges. After completing vectorization, the similarity of the user portraits between the two users can be obtained by calculating the Euclidean distance between the two vectors, cosine similarity, Pearson correlation coefficient, and other methods to evaluate the vector distance / direction consistency in the vector space.
[0085] Of course, the two user portraits can also be classified by using a classification model. If a classification model is used, the output conclusion is whether the users have common characteristics: if the two user portraits are classified into one category, then the users have the same or similar tendencies for the cockpit services issued by the car computer, then the similarity of the user portraits can be S; if the two user portraits are not classified into one category, then the two users may have obvious differences in their tendencies for cockpit services, and the similarity of the two user portraits is N. For example, S is a value greater than the first preset threshold, and N is a value less than the first preset threshold. Here, the embodiments of the present invention do not make specific limitations on this.
[0086] In this way, the user's historical data can be further expanded through user profiling to optimize local training data, thereby further improving the accuracy of the current neural network model obtained by pre-fusion and post-fusion, and providing users with more intelligent cockpit services.
[0087] In order to update the historical data of each user, in an optional embodiment, updating the historical data of each user according to the similarity of user portraits between users may include: When the users include a first user and a second user and historical data of the first user is generated, determining a similarity between a user profile of the first user and a user profile of the second user; When the similarity is greater than a first preset threshold, the historical data of the first user is added to the historical data of the second user.
[0088] It can be understood that when the historical data of the first user is generated for the vehicle cabin, the similarity between the user portraits of the first user and the second user can be calculated, and then the similarity can be compared with the first preset threshold. If it is greater than, it means that the similarity between the first user and the second user is high, so the generated historical data of the first user is added to the historical data of the second user. If not, it means that the similarity between the first user and the second user is low, so the generated historical data of the first user is not added to the historical data of the second user.
[0089] In this way, by comparing the similarity with the first preset threshold, it is determined whether to add the historical data of the first user to the historical data of the second user. In this way, the historical data between users with similar user portraits can be shared, so that the historical data of each user can be accumulated as soon as possible, which accelerates the local training process of the neural network model, optimizes the training data of the local model training, and improves the accuracy of the current neural network model.
[0090] In addition, in a scenario with multiple users, in order to obtain historical data of a sufficient number of users as quickly as possible, in an optional embodiment, the above method may further include: In the case where the user is a plurality of users and historical data is generated, determining the adoption degree of the generated historical data for each of the users; When the adoption degree is greater than a second preset threshold, the historical data is added to the historical data of the user corresponding to the adoption degree.
[0091] It can be understood that for the scenario of multiple users, each time a piece of historical data is generated, it is necessary to determine the adoption level of the generated historical data for each user. It should be noted that for the historical data, it may be a piece of historical data under the cockpit service actively provided by the vehicle-mounted equipment, or it may be a piece of historical data of the cockpit service passively provided by the vehicle-mounted equipment. Therefore, here, after the historical data is generated, it is necessary to determine the adoption level of the historical data for each user.
[0092] For example, the historical data is a service provided for the voice signal of the first user and can meet the needs of the first user, so the adoption degree for the first user is relatively high. However, for the second user, the service provided for the voice signal of the first user is not the service that the second user wants. Therefore, the adoption degree for the second user is relatively low.
[0093] Here, in order to determine the adoption level of each user in the generated historical data, the user's adoption level can be determined based on the user's multimodal data after the cockpit service is provided in the generated historical data, the user's follow-up correction voice instructions and the user's direct vehicle operation (including buttons, soft switch lights).
[0094] The user response is also vectorized through multimodal data. For example, facial reactions can be analyzed to obtain three types of reactions: negative, no reaction, and positive. There are also different degrees of emotional fluctuations under the three types of reactions. The user's follow-up correction voice commands and direct vehicle computer operations are used to directly reflect the user's attitude towards the cabin service.
[0095] For example, if after a cockpit service (for example, the air conditioning temperature is adjusted to 25 degrees), the user follows up with other response instructions of the same function or directly adjusts the cockpit service (for example, the air conditioning temperature is adjusted to 23 degrees), it can be considered that the user does not fully agree with the cockpit service; if after the cockpit service, the user follows up with an elimination instruction of the same function (for example, turning off the air conditioning), it can be considered that the user does not agree with the cockpit service at all. Among them, there is a mapping relationship between the user's follow-up correction voice instructions and the user's direct vehicle operation and the degree of recognition. The final mapping result can be a discrete or continuous vectorized dimension, and the specific dimension classification shall be based on the richness of the mapping result. The above three types of user feedback data (user's multimodal data, user's follow-up correction voice instructions, and user's direct vehicle operation) jointly determine the user's adoption degree. The three are used to obtain the final user's adoption degree using linear weighting, nonlinear effect or correlation analysis as appropriate.
[0096] Here, the user's adoption level is taken into account, and a supervised learning + reward mechanism can be used to measure the impact of different dimensions in historical data on user adoption: taking the cockpit service as input and the adoption level of each user as output, supervised learning is performed to predict the user's adoption level for the generated historical data. At the same time, dimensions with a higher impact on the adoption level are rewarded. In subsequent service recommendations, the rewards are converted into actual vehicle function responses to continuously verify users' different attitudes towards specific recommended cockpit services.
[0097] In addition, the score of each user may be received for each piece of historical data, and the score may be determined as the adoption degree of the historical data for each user. Here, the embodiment of the present invention does not make any specific limitation to this.
[0098] After determining the adoption degree of the generated historical data for each user, the historical data may be added to the historical data of users whose adoption degree is greater than a second preset threshold.
[0099] Of course, the historical data and the degree of adoption of the historical data for each user may also be added to the historical data of each user for use in training the preset neural network model.
[0100] In this way, by comparing the degree of adoption with the second preset threshold, it is determined whether to add historical data to the user's historical data. In this way, the historical data that the user can adopt can be added to his historical data, thereby optimizing the user's historical data, and then optimizing the training data for local model training, thereby improving the accuracy of the current neural network model.
[0101] For the training of the preset neural network model, for the scenario of multiple users, in an optional embodiment, the above method may further include: After acquiring the multimodal data of the user in the cabin of the vehicle, the current neural network model is determined from the neural network model set.
[0102] It can be understood that when the historical data generated in the vehicle cabin includes historical data of multiple users, it is necessary to determine the neural network model of each user to form a neural network model set.
[0103] The neural network model of each user in the neural network model set is obtained by locally training the preset neural network model using the historical data of each user in the cabin; the historical data of each user is: the multimodal data of each user generated locally and the cabin service corresponding to the multimodal data of each user generated locally. That is to say, for each user, after obtaining the historical data, the cabin control method described in one or more of the above embodiments can be used to obtain the neural network model of each user.
[0104] The historical data of each user mentioned above can be obtained by using the cockpit control method described in one or more of the above embodiments.
[0105] In this way, by using the historical data of each user to train the preset neural network model, the neural network model of each user can be obtained separately. In this way, when processing the user's multimodal data, a model can be determined from the neural network model set to process the user's multimodal data to generate cockpit services. In this way, the required cockpit services can be provided to the user in a targeted manner, thereby improving the intelligence of the cockpit services.
[0106] In order to determine the current neural network model from the neural network model set, in an optional embodiment, determining the current neural network model from the neural network model set may include: In the case where the user in the cabin is a single user, determining the neural network model of the user from the neural network model set; The user's neural network model is determined as the current neural network model.
[0107] It can be understood that, for the case where there is only one user in the cabin, after obtaining the multimodal data of a certain user, the neural network model of the user can be determined from the neural network model, and the neural network model of the user can be determined as the current neural network model, and then the user's multimodal data can be input into the current neural network model to generate a cabin service and execute the cabin function corresponding to the cabin service.
[0108] In this way, the neural network model corresponding to the user is used to generate cockpit services for the user's multimodal data. The neural network model of the user can be used for the user, thereby providing the user with targeted cockpit services, thereby improving the user's experience in the cockpit.
[0109] In the case where there are multiple users in the cockpit, in order to determine the current neural network model, in an optional embodiment, the method may further include: When there are multiple users in the cockpit, determining a target user from among the users; Determine a neural network model of a target user from a set of neural network models; The neural network model of the target user is determined as the current neural network model.
[0110] It can be understood that when there are multiple users in the cabin, the multimodal data of the users obtained at this time may include multimodal data of multiple users. Here, a target user can be determined from these users, and then the neural network model of the target user can be determined from the neural network model set, and the neural network model of the target user can be used as the current neural network model.
[0111] Here, the target user may be determined from the users by random method, or according to preset rules, or according to an artificial intelligence (AI) model, and the embodiments of the present invention do not specifically limit this.
[0112] In this way, the current neural network model is determined by determining the target user, so that the current network model is related to the determined target user, which is conducive to realizing cockpit services with a certain user as the core in multiple user scenarios, so that suitable neural network models can be selected under multiple users to provide cockpit services for multiple users.
[0113] In order to determine the target user from the users, in an optional embodiment, determining the target user from the users may include: Determine target users from users based on their multimodal data.
[0114] It is understandable that the target user can be determined from the users according to the multimodal data of each user, wherein the age of each user can be known according to the multimodal data of each user, and the target user can be determined from the users according to the age of each user, the emotion of each user can be known according to the multimodal data of each user, and the target user can be determined from the users according to the emotion of each user, and the location of each user can be known according to the multimodal data of each user, and the target user can be determined from the users according to the location of each user. Here, the embodiments of the present invention do not make specific limitations on this.
[0115] In this way, the target user is determined from the users through the multimodal data of each user mentioned above, so that the determined target user is related to the user's multimodal data, and a suitable neural network model can be selected for this cockpit service, which helps to provide suitable cockpit services to multiple users and improve the user experience in the cockpit.
[0116] In addition, in order to determine the target user from the users, in an optional embodiment, determining the target user from the users may include: Determine target users based on their priorities.
[0117] It is understandable that in a multiple user scenario, a priority can be set for each user, wherein the priority of each user can be flexibly set by the user according to his or her own needs, or can be set by the vehicle-mounted device according to the user's age, or can be set by the vehicle-mounted device according to the position in the cabin, or can be set according to the frequency of use of the vehicle. Here, the embodiment of the present invention does not specifically limit this.
[0118] After the priority of each user is preset, the target user can be determined according to the priority of each user. Here, the user with the highest priority can be determined as the target user, and the user with the lowest priority can also be determined as the target user. Here, the embodiment of the present invention does not make any specific limitation on this.
[0119] In this way, after the target user is determined by the priority of each user, the determined neural network model is related to the priority of each user, so that the needs of specific users can be given priority, so that the smart cockpit can select the appropriate neural network model according to the priority and provide cockpit services to the user.
[0120] Furthermore, in order to determine the target user from the users, in an optional embodiment, determining the target user from the users may include: In the case where the multimodal data of the user includes a voice signal, the user who emits the voice signal is determined as the target user.
[0121] It can be understood that in the embodiment of the present invention, the cockpit service may include active services and passive services. For active services, the cockpit service is usually obtained by the vehicle-mounted equipment through collecting the user's multimodal data and inputting it into the current neural network model. For passive services, the cockpit service is usually obtained by the user sending a voice signal as the user's multimodal data and inputting it into the current neural network model. For this passive service, in the embodiment of the present invention, if the user's multimodal data includes a voice signal, it means that the user who sent the voice signal has put forward a demand for passive service.
[0122] Therefore, at this time, when there is a voice signal in the user's multimodal data, the user who sends the voice signal is taken as the target user. For example, the driver in the cockpit sends a voice signal: It's hot, turn on the air conditioner. Then, when the vehicle-mounted device obtains the user's multimodal data, the driver can be taken as the target user.
[0123] In this way, the user who sends the voice signal can be used as the target user in the passive service. In this way, the appropriate current neural network model can be determined for this service, which is conducive to providing users with more accurate cockpit services and further improving the user experience in the cockpit.
[0124] The cockpit control method described in one or more of the above embodiments is described below with examples.
[0125] Based on the convenience of using the interactive methods of cockpit services in related technologies, users have higher and higher requirements for services in the cockpit, but the interactive methods in related technologies may not be able to cope with the following practical problems: 1. It is impossible to accurately describe the functions to be used through voice, especially some non-high-frequency functions. For example, the low-speed warning sound outside the car, the user may describe it as the warning sound outside the car; 2. Incorrect execution after the command is given. This may be caused by a variety of reasons, such as the user speaking with an accent, complex voice commands, gestures used being mistakenly recognized as other gestures, etc. These reasons result in incorrect commands actually transmitted to the vehicle computer; 3. After the command is given, it is not executed or there is no response. This may be caused by many reasons, such as the acceptable generalization of the command is not enough, the gesture used is different from the trained gesture, etc., which results in the failure to actually transmit effective commands to the vehicle computer; 4. It can only respond to user needs passively, requiring users to speak actively before operations can be performed, and it is unable to infer and process users' potential needs.
[0126] At the same time, users also face many obstacles when describing interaction problems. Common problem descriptions may be "voice is not easy to use" or "no response when used", but the specific problem cannot be located, resulting in the efficiency of problem solving to be optimized. In addition, there are many foreseeable problems in using voice and multi-modal interaction. The current solution is to patch known problems, but the efficiency is questionable and the effect of solving unforeseen problems cannot be determined. The most reasonable way is to make the car computer better understand human language and absorb a wider range of commands. Interaction through large models is a good way.
[0127] The rise of big models in recent years has brought new possibilities for innovation. Big models have emerged in an endless stream at home and abroad. The exponential increase in parameters and the competition among big models have promoted the development of artificial intelligence. However, when it comes to actual application, the following multiple obstacles have emerged: 1. Limited by performance requirements, large models can only interact by calling cloud services, which raises concerns about privacy security and response delays; 2. High-frequency calls to cloud services may bring higher costs to users. Considering the input-output ratio of using large models to complete tasks, C-end users may be unwilling to bear such costs.
[0128] The application of a large model on the end can alleviate the resistance caused by the above problems, because the large model on the end reduces the call of cloud services, thereby greatly improving the response speed; in addition, it does not rely on the network, and the large model on the end can continue to respond in weak or disconnected network conditions; most importantly, user data does not need to be uploaded to the cloud, but stays on the end to ensure data security. In addition, the large model on the end can better solve the interaction problems described above, because most of the needs in the cockpit have high timeliness requirements, and users want to interact with the car in a flexible way. Compared with the large model on the cloud, the large model on the end can free up some computing power and continue to learn user habits, allowing users to interact more naturally, freely and smoothly when using the cockpit.
[0129] At present, the application of edge-side models is mainly concentrated on mobile devices such as mobile phones, while the application of edge-side large models has not yet been realized in the cockpit. Combined with the user needs in the cockpit and the services that can be provided, the cockpit edge-side large model can realize information penetration in the form of integrating various applications that need to process information sources, thereby improving the efficiency of various applications. At the same time, various information sources or applications can also communicate with each other and work together, from extracting the call content into a to-do schedule after a call to continuously learning user behavior and actively opening entertainment methods suitable for all users in the cockpit while the user is driving. All kinds of automated, self-learning, and adaptive behaviors are within the capability of the edge-side large model.
[0130] This example proposes an end-side multimodal big model that continuously learns cockpit user behavior. It can access multimodal data in the cockpit (for example, including: user biometric information collected by the camera, high-frequency user operation behaviors such as voice interaction and music playing, environmental information during driving, etc.), form memories after learning, generate behavioral profiles for each user based on different users, and cluster each user category; at the same time, the end-side multimodal big model is deeply integrated with the cockpit system, integrating the historical tasks, execution results and capability boundaries of different business information, and distributing applicable multi-source information to different businesses, realizing the business penetration of multimodal information input and system functions, and providing services for different types of users passively or actively. Using the end-side multimodal big model proposed in this example, you can provide a thousand-faceted service in the cockpit, making the cockpit a car assistant that truly understands you.
[0131] Specifically, the following methods can be used to implement the application of the large end-side model in the cockpit: Figure 3 A flowchart of an example 1 of an optional cockpit control method provided by an embodiment of the present invention is shown as follows: Figure 3 As shown, the method may include: S301: Obtain user information, environment information and task response results; S302: Inputting user information, environment information, and task response results into the terminal-side multimodal large model; S303: Perform pre-fusion and post-fusion to train a large multi-modal model on the end side.
[0132] Among them, a hybrid fusion end-side multimodal large model is used: Since the tasks processed by the large model include two types of active services and passive services, and the model needs to process multimodal information input such as vision and hearing, hybrid fusion is chosen as the information fusion solution.
[0133] When the model is learning for individual users with insufficient data, the pre-fusion solution is used, that is, before the model is trained, the data from different sources and types are pre-processed and integrated to form a unified data set. This method is conducive to the model learning the characteristics of different data sources at the same time during the training process, which helps to improve the generalization ability and performance of the model; when learning for individual users with sufficient data, the post-fusion solution is used, that is, after the model training is completed, the output results of different models or different data sources are integrated to obtain the final prediction results, so that the model can be allowed to independently learn the characteristics of each data source.
[0134] exist Figure 3In the process of pre-fusion, training data 1 is collected. When the first preset condition is met, the large model is trained using training data 1. Training data 2 is collected. When the first preset condition is met, the large model is trained using training data 2. Training data 3 is collected. When the first preset condition is met, the large model is trained using training data 3. In post-fusion, when training data 1, training data 2, and training data 3 meet the second preset condition, the large model is trained using training data 1, training data 2, and training data 3, thereby realizing the training of the hybrid fusion mode of pre-fusion and post-fusion.
[0135] Self-learning solution for the large multimodal model on the device side (equivalent to the training of the large multimodal model on the device side): Figure 4 A flowchart of a second example of an optional cockpit control method provided by an embodiment of the present invention is shown in FIG. Figure 4 As shown, the method may include: S401: Power on the vehicle computer; execute S402; S402: sensor background starts; execute S403; S403: Obtain user instructions; execute S404; S404: Obtain user information and environment information, analyze the number of users and user locations; execute S405; S405: input the data obtained in S404 into the end-side multimodal large model; execute S406; S406: Obtain response status from the multimodal large model on the end side; execute S407; S407: The vehicle computer responds and executes.
[0136] Figure 5 A flowchart of an example 3 of an optional cockpit control method provided by an embodiment of the present invention is shown as follows: Figure 5 As shown, the method may include: S501: User data input; Execute S501; S502: data preprocessing; executing S502; S503: User portrait construction; Execute S503; When the car computer is powered on, microphones, cameras and other sensor devices inside and outside the car are started in the background to record the location of each user when they get on the car.
[0137] For a single / multiple user who performs a function operation in the cockpit and stops the operation, the sensor equipment is used to collect the user's biometric information (for example, user voiceprint, line of sight, head deflection, emotional state, stress level, etc.), time information (for example, the year, month, and day when the operation is performed), environmental information (for example, weather and road conditions obtained by vehicle sensors, etc.) and response status (for example, success or failure of function execution, user feedback on the execution result, etc.), and associate the user's position in the cabin. When the knowledge of the early large model is insufficient, the user's needs are processed through the large model, and the processing logic of the large model itself shall prevail.
[0138] In addition, the user's portrait is repeatedly supplemented based on the information of the same user. The supplemented information dimensions can be divided into: original information dimension and inferred information dimension. At the same time, the information dimension fluctuates with the user's situation, such as the above Figure 2 The description is not repeated here.
[0139] S504: service provision and adoption rate backtracking; executing S505 or executing S506; S505: Adoption rate ≥ 90%, user portrait training data; execute S507; S506: The adoption rate is less than 90%, and the user profile is self-learned to optimize the profile; execute S503; Here, since the purpose of the big model is to better adapt to users, it is necessary to provide different services and emotional values for different users. In other words, the big model needs to understand which different types of people use the cockpit. The previous model adaptation method only learns user behavior through single-mode information. The end-side big model in this example combines the above multi-modal information for learning and integrates the perception of user emotions. Therefore, it is extremely sensitive to user emotions. This behavior is mapped to the classification processing of the big model, which significantly increases the proportion of the emotional dimension in the overall analysis. When the emotional dimension is used as the main variable, other variables belong to environmental variables. The combination of different environmental variables and main variables is the core point to distinguish user classification. The big model classifies user data containing the above dimensions, and then performs reinforcement learning based on the classification results. Based on rewards, the key dimension factors of different user classifications are confirmed, and at the same time, a typical image label is generated for each user based on the clustering results.
[0140] The premise of user classification is to ensure the accuracy of classification. The accuracy is strongly related to the credibility of the data used for training classification. The usage of each user portrait data includes self-learning of the current user and self-learning of other users using part of the user data. When the proportion of services provided to a certain part of users adopted by the user is ≥90%, the user's portrait will be put into the training of the overall user prediction service.
[0141] S507: overall user prediction service training; executing S508; S508: Model training and tuning; executing S509; S509: knowledge extraction and model evaluation; executing S510; S510: model optimization; executing S511; S511: Generate prediction / analysis results; execute S512; S512: user feedback; executing S513; S513: Feedback analysis and processing; return to execute S508.
[0142] Figure 6 A flowchart of a fourth example of an optional cockpit control method provided by an embodiment of the present invention is shown in FIG. Figure 6 As shown, the overall model training process may include: S601: Data collection phase; In S601 , family population data, driving behavior data of family members, and emotional data of people are collected.
[0143] S602: data preprocessing stage; At this stage, the collected data is cleaned, normalized / standardized, and completed, and missing values can be handled during data completion.
[0144] S603: feature extraction stage; In S603, relevant features are selected from the preprocessed data to achieve feature extraction, and a feature vector is constructed. Dimensionality reduction may be required to reduce computational complexity.
[0145] S604: using a clustering algorithm (K-means), determining the number of clusters, and performing cluster analysis; Here, you can select a clustering algorithm (e.g., K-means), determine the number of clusters, and perform cluster analysis; you can also perform pattern recognition to identify patterns and abnormal behaviors in the data.
[0146] S605: association rule learning; Among them, association rules between data can be discovered, for example, using the Apriori algorithm to learn association rules.
[0147] S606: Generate model; Here, data representation can be generated using an autoencoder or the like.
[0148] S607: Model evaluation; In S607 , clustering effect and pattern recognition accuracy are evaluated, and model parameters are adjusted as needed.
[0149] S608; Model optimization; Among them, the model can be optimized based on feedback and the algorithm parameters can be adjusted.
[0150] S609: Model loading; Here, in S609, the model can be automatically loaded and operated.
[0151] It should be pointed out that the user personalized model and the general model can coordinate and assist each other, so that the whole mechanism can be better integrated with the actual user situation.
[0152] In addition, in order for the model to better understand the user, it can provide targeted services to the user. Figure 7 A flowchart of Example 5 of an optional cockpit control method provided by an embodiment of the present invention is shown in FIG. Figure 7 As shown, the method may also include: S701: Acquire multimodal information during human-machine-user interaction; S702: Input multi-mode information into the large model; execute S703 and S704: S703: Active service reasoning; Execute S705: S704: Passive service request to provide passive service; execute S711: S705: Analyze using historical data; execute S706; S706: learning and understanding process; executing S707; Among them, in S706, the current variable learning and understanding process can also be combined to obtain the inference result.
[0153] S707: Providing active service; executing S708; S708: Strengthen learning adjustment; execute S709; S709: Optimize service strategy; execute S710; S710: User feedback is collected into the large model.
[0154] Here, in S710, user feedback is collected into a large model, which can be used for emotion / behavior analysis.
[0155] S711: real-time service processing; executing S712; Here, immediate service processing can be performed based on current variables.
[0156] S712: Respond to user needs.
[0157] It should be noted that the services provided in this example can be classified into passive services and active services; among them, passive services are the needs actively proposed by users. This part of the needs is highly immediate, and the task direction can be summarized based on the current main variables and environmental variables. Active services are services that users may need based on the analysis and reasoning of the large model. This part of the needs has a strong delay and requires a certain learning and understanding process. The services provided are obtained by learning the historical passive service response results, the main variables and environmental variables when the historical passive services were generated, and combining the current main variables and environmental variables. The purpose of the service emphasizes emotional positive feedback, and the effect of key reward factors can still be emphasized through reinforcement learning.
[0158] In this example, compared with the language big model that triggers business penetration through voice commands, the terminal-side multimodal big model can trigger business penetration through user emotions. This will affect the execution of future established tasks after the big model is deeply integrated with the system.
[0159] The following are the applications of the above examples in different scenarios: For a family of four (for example, a couple, a child, and an elderly person) on a short trip: 1. The car computer is powered on, and the microphones, cameras and other sensor devices inside and outside the car are started in the background to record the location of each family member when they get in the car; Among them, the camera can identify the couple in the front and rear seats when there is no obstruction. The rear seat can identify the elderly person in the right rear seat fastening the seat belt for the user in the left rear seat. At the same time, the microphones inside and outside the car detect voiceprints including children's voiceprints, so it is determined that the user in the left rear seat is a child. After the recognition is completed, the user composition and associated location information are passed to the large model. The overall process is as follows: Figure 4 .
[0160] 2. As soon as the vehicle leaves the basement, each sensor inputs the user signal just collected into the large model; Among them, the driving speed increased to 60km / h within 5 seconds, the wipers were not turned on, the temperature and humidity sensors identified the specific temperature and humidity as 28 degrees and 35% respectively, and the 26-degree air conditioner with 2nd wind speed was automatically turned on in the car; the co-pilot user tilted his head towards the back row, and the middle-aged female voiceprint was detected at the same time. No images were recognized on the left and right rear, but the voiceprints of elderly women and children were detected. The co-pilot turned his head and tilted his head towards the driver. The driver moved his lips and produced a middle-aged male voiceprint, his eyes still looking straight ahead, and then the co-pilot moved his lips, and the car computer received a voice command to set the air conditioner to internal circulation and lower the front windows a little.
[0161] Based on the received signals, the big model initially infers that the current climate is relatively pleasant and the road conditions are good (environmental information). The users do not have any radical behaviors or reactions. Among them, the co-driver takes the most actions, including: asking the back seat users about their requirements for the air conditioning in the car, listening to the driver's requirements and issuing voice commands; the two back seat users may not be satisfied with the current air conditioning settings in the car; the driver is not clear about his attitude towards the current environment. The overall process includes: Figure 3 and Figure 4 The anterior fusion part of 3. The vehicle traveled for a total of 3 days, including multiple trips within City A. The above information was input into the big model for each trip. The big model output the portrait of each user based on the data, with the portrait of a middle-aged woman as a representative as follows: a) Demographic dimension: aged between 35 and 40, resident in City A, no coughing or other behaviors, and considered healthy; b) Behavioral dimension: He would sleep for 30 minutes during the trip around 1pm every day. He frequently used his mobile phone in the car. The destination of each trip was a scenic spot around City A. He frequently used voice interaction. The most commonly used voice command was "play XXXX". He always sat in the passenger seat. c) Emotional reactions under various conditions: encountered traffic jams of more than 30 minutes twice, with no behavioral differences during the traffic jams; had frequent conversations with the back row during the 3-day trip, and the overall emotional reaction of the voiceprint was neutral; the voiceprint in the evening of the second day of the trip was quite different from the previous one, with relatively excited and happy emotional reactions, and issued a shooting task of "help me take a photo of the sunset" through voice commands, when the location was about to reach the destination of the navigation trip, scenic spot X, and the sunset landscape of scenic spot X was a publicity focus. Figure 2 .
[0162] 4. After analysis: the overall trip is a short-distance family trip. The portraits of the four users in the car are quite different, but the emotional changes of the two female users to special situations are similar. Among them, the middle-aged female user reacts more to the natural beauty. In addition, although the middle-aged female user issued the "play XXXX" command many times, it was actually played for the child passenger; due to driving behavior, the driver has emotional fluctuations in some situations that occur during driving, such as traffic jams, but after playing YYYY's music, the tension in his facial expression is relieved; the emotional fluctuations of the children during the trip are mainly caused by the video content played on the back screen. The overall flow chart is as follows Figure 3 The post-fusion part.
[0163] For daily travel of a family of three (for example, a couple and a child): 1. After a short trip, the most common use of the car is for a family of three (couple and child). Analysis shows that the purpose of the trip is to send children to school and to commute. Combined with the previous analysis conclusions, the large model began to actively ask the co-pilot if he needs to play XXXX for the back seat after multiple trips with children in the car, and capture sunset photos in good weather conditions, display them to the co-pilot seat and ask if he needs to transfer the photos to the co-pilot user's mobile phone; when the female co-pilot user is not in the car, even if there is an environment that meets the conditions for taking photos, the shooting behavior will not be triggered. The overall process is as follows: Figure 5 .
[0164] 2. The big model on the end side will adjust the execution of the given task through multi-mode information input: if the front row passengers quarrel, even if it is the time when XXXX is usually played for the back row users, the big model will not play the video and try to adjust the in-car environment to a state where the users feel relatively comfortable. The overall process is as follows: Figure 7 .
[0165] The end-side multimodal large model provided in this example can achieve insights into user emotions through multimodal information input, thereby providing smarter and more humane services.
[0166] An embodiment of the present invention provides a cockpit control method. By deploying a neural network model locally in the cockpit, cockpit control is not limited by network performance. Not only can the cockpit control ensure the privacy and security of users, but it can also respond to users' needs in a timely manner, thereby improving the response speed of cockpit services. On the basis of locally deploying the neural network model in the cockpit, the neural network model is locally trained using locally generated historical data, thereby using the historical data of local users to train the model, which can further improve the accuracy of the model and thus improve the intelligence of the cockpit service.
[0167] Based on the same inventive concept as the above-mentioned embodiment, an embodiment of the present invention provides a cockpit control device, Figure 8 A schematic diagram of the structure of an optional cockpit control device provided by an embodiment of the present invention, such as Figure 8 As shown, the cockpit control device 800 may include: An acquisition module 81 is used to acquire multimodal data of a user in a vehicle cabin; A generating module 82, for inputting the multimodal data of the user into the current neural network model to generate a cockpit service; An execution module 83, used for executing a cockpit function corresponding to the cockpit service; The current neural network model is obtained by locally training a preset neural network model using historical data of users in the cabin; the historical data of the user is: the multimodal data of the user generated locally and the cabin service corresponding to the multimodal data of the user generated locally; The multimodal data of the user includes at least: the user's personal information and the user's environmental information; the user's personal information includes at least one or more of the following dimensions: the user's personal statistics dimension information, the user's behavior dimension information, and the user's emotional dimension information; The device also includes: an updating module, which is used to: when there are multiple users, determine the user portrait of each user based on the multimodal data in the historical data of each user; and update the historical data of each user based on the similarity of the user portraits between the users.
[0168] In an optional embodiment, the device updates the historical data of each user according to the similarity of user portraits between users among the users, including: when the users include a first user and a second user and the historical data of the first user is generated, determining the similarity between the user portrait of the first user and the user portrait of the second user; when the similarity is greater than a first preset threshold, adding the historical data of the first user to the historical data of the second user.
[0169] In an optional embodiment, the device is also used to: when there are multiple users and historical data is generated, determine the adoption level of the generated historical data for each user among the users; when the adoption level is greater than a second preset threshold, add the historical data to the historical data of the user corresponding to the adoption level.
[0170] In an optional embodiment, the device is also used to: perform statistics on the user's historical data to obtain statistical results; when the statistical results meet the first preset condition, use the user's historical data corresponding to the statistical results to train a preset neural network model to obtain the current neural network model.
[0171] In an optional embodiment, the device is also used to: determine that the statistical result satisfies the first preset condition when the statistical result indicates that the number of historical data generated for the user reaches a third preset threshold; determine that the statistical result satisfies the first preset condition when the statistical result indicates that the duration of the historical data generated for the user reaches a preset duration.
[0172] In an optional embodiment, the device is also used to: when the current neural network model is obtained, clear the statistical results, return to the step of performing statistics on the user's historical data to obtain the statistical results, until the statistical results meet the second preset condition; when the statistical results meet the second preset condition, use the user's historical data to train the preset neural network model to obtain the current neural network model; wherein, when the statistical results meet the first preset condition for a preset number of times, it is determined that the statistical results meet the second preset condition.
[0173] In an optional embodiment, the device is also used to: after acquiring multimodal data of users in the vehicle's cabin, determine the current neural network model from a set of neural network models; wherein the neural network model of each user in the neural network model set is obtained by locally training a preset neural network model using historical data of each user in the cabin; the historical data of each user is: the multimodal data of each user generated locally and the cabin service corresponding to the multimodal data of each user generated locally.
[0174] In an optional embodiment, the device determines the current neural network model from a set of neural network models, including: when there is only one user in the cabin, determining the neural network model corresponding to the user from the set of neural network models; and determining the user's neural network model as the current neural network model.
[0175] In an optional embodiment, the device is also used to: when there are multiple users in the cabin, determine a target user from the users; determine a neural network model corresponding to the target user from a set of neural network models; and determine the neural network model corresponding to the target user as the current neural network model.
[0176] In an optional embodiment, the device determines the target user from the users, including at least one of the following: determining the target user from the users based on the users' multimodal data; determining the target user based on the users' priorities; and when the users' multimodal data includes a voice signal, determining the user who sends the voice signal as the target user.
[0177] In practical applications, the acquisition module 81, generation module 82 and execution module 83 mentioned above may be implemented by a processor located on the cockpit control device 800, specifically a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA).
[0178] Fig. 9 A schematic diagram of the structure of an optional vehicle-mounted device provided in an embodiment of the present invention, such as Fig. 9 As shown, an embodiment of the present invention provides a vehicle-mounted device 900, including: A processor 91 and a storage medium 92 storing executable instructions of the processor 91, wherein the storage medium 92 relies on the processor 91 to perform operations through a communication bus 93, and when the instructions are executed by the processor 91, the cockpit control method described in one or more of the above embodiments is executed.
[0179] It should be noted that, in actual application, the various components in the vehicle-mounted device 900 are coupled together via the communication bus 93. It is understandable that the communication bus 93 is used to realize the connection and communication between these components. In addition to the data bus, the communication bus 93 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Fig. 9 Various buses are labeled as communication buses 93.
[0180] An embodiment of the present invention provides a computer storage medium storing executable instructions. When the executable instructions are executed by one or more processors, the processors execute the cockpit control method as described in one or more of the above embodiments.
[0181] An embodiment of the present invention provides a computer program product, including a computer program or instructions. When the computer program or instructions are executed by a processor, the steps of the cockpit control method described in one or more embodiments are implemented.
[0182] Among them, the computer readable storage medium can be a ferromagnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface storage, an optical disk, or a compact disc read-only memory (CD-ROM) and other memories.
[0183] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program codes.
[0184] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0185] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0186] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0187] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.
Claims
1. A cockpit control method, characterized in that: include: Acquire multimodal data of users in the vehicle cabin; Inputting the multimodal data of the user into the current neural network model to generate a cockpit service; Execute the cockpit function corresponding to the cockpit service; The current neural network model is obtained by locally training a preset neural network model using the historical data of the user in the cockpit; the historical data of the user is: the multimodal data of the user generated locally and the cockpit service corresponding to the multimodal data of the user generated locally; The multimodal data of the user includes at least: the personal information of the user and the environmental information of the user; the personal information of the user includes at least one or more dimensions of information: information of the personal statistics dimension of the user, information of the behavior dimension of the user, and information of the emotional dimension of the user; Wherein, the method further comprises: In the case where there are multiple users, determining a user portrait of each user according to multimodal data in historical data of each of the users; The historical data of each user is updated according to the similarity of the user portraits between the users.
2. The method according to claim 1, characterized in that The updating of the historical data of each user according to the similarity of the user portraits among the users includes: In a case where the users include a first user and a second user and historical data of the first user is generated, determining a similarity between a user profile of the first user and a user profile of the second user; When the similarity is greater than a first preset threshold, the historical data of the first user is added to the historical data of the second user.
3. The method according to claim 1, characterized in that The method further comprises: In the case where the user is a plurality of users and historical data is generated, determining the adoption degree of the generated historical data for each of the users; When the adoption degree is greater than a second preset threshold, the historical data is added to the historical data of the user corresponding to the adoption degree.
4. The method according to claim 1, characterized in that: The method further comprises: Performing statistics on the historical data of the user to obtain statistical results; In the case where the statistical result satisfies the first preset condition, the preset neural network model is trained using the historical data of the user corresponding to the statistical result to obtain the current neural network model.
5. The method according to claim 4, characterized in that The method further comprises at least one of the following: When the statistical result indicates that the number of historical data generated for the user reaches a third preset threshold, determining that the statistical result satisfies the first preset condition; When the statistical result indicates that the duration of generating the historical data of the user reaches a preset duration, it is determined that the statistical result meets the first preset condition.
6. The method according to claim 4, characterized in that The method further comprises: In the case of obtaining the current neural network model, clearing the statistical result, returning to the step of performing statistics on the historical data of the user to obtain the statistical result, until the statistical result meets the second preset condition; When the statistical result satisfies the second preset condition, the local neural network model is trained using the historical data of the user to obtain the current neural network model; When the statistical result satisfies the first preset condition for a number of times reaching a preset number, it is determined that the statistical result satisfies the second preset condition.
7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: After acquiring multimodal data of a user in a cabin of a vehicle, determining the current neural network model from a set of neural network models; Among them, the neural network model of each user in the neural network model set is obtained by locally training a preset neural network model using the historical data of each user in the cabin; the historical data of each user is: the multimodal data of each user generated locally and the cabin service corresponding to the multimodal data of each user generated locally.
8. The method according to claim 7, characterized in that The method further comprises: In the case where there are multiple users in the cabin, determining a target user from the multiple users; Determine the neural network model of the target user from the set of neural network models; The neural network model of the target user is determined as the current neural network model.
9. The method according to claim 8, characterized in that The determining of the target user from the multiple users includes at least one of the following: Determining the target user from the multiple users according to the multimodal data of the user; Determining the target user according to the priority of the user; In a case where the multimodal data of the user includes a voice signal, the user who sends the voice signal is determined as the target user.
10. A cockpit control device, characterized in that: include: An acquisition module, used to acquire multimodal data of a user in a vehicle cabin; A generation module, used for inputting the multimodal data of the user into the current neural network model to generate a cockpit service; An execution module, used for executing a cockpit function corresponding to the cockpit service; The current neural network model is obtained by locally training a preset neural network model using the historical data of the user in the cockpit; the historical data of the user is: the multimodal data of the user generated locally and the cockpit service corresponding to the multimodal data of the user generated locally; The multimodal data of the user includes at least: the personal information of the user and the environmental information of the user; the personal information of the user includes at least one or more dimensions of information: information of the personal statistics dimension of the user, information of the behavior dimension of the user, and information of the emotional dimension of the user; The device further includes an updating module, which is used to: In the case where there are multiple users, determining a user portrait of each user according to multimodal data in historical data of each of the users; The historical data of each user is updated according to the similarity of the user portraits between the users.
11. A vehicle-mounted device, characterized in that: include: A processor and a storage medium storing instructions executable by the processor, wherein the storage medium relies on the processor to perform operations through a communication bus, and when the instructions are executed by the processor, the cockpit control method described in any one of claims 1 to 9 is executed.
12. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the cockpit control method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Adjustment method and device for emotion of user, equipment and readable storage medium
CN111724880A
Content request method and device, electronic equipment and storage medium
CN114021694A
System and method for data security in autonomous vehicle
CN115812315A
Music recommendation method, recommendation system, intelligent cabin and vehicle
CN116450946A
Vehicle thermal management method and device and storage medium
CN118438858A
Cited By
Vehicle-mounted application control method and device, vehicle, medium and product
CN120697783A
A method, device, vehicle, medium, and product for vehicle-mounted applications.
CN120697783B
Automobile cabin software control method and system based on multi-modal interaction
CN121764330A