Vehicle control service recommendation method and device, vehicle and storage medium
Through multimodal information fusion and intelligent model recommendation, the problem of inaccurate vehicle control service recommendation in the intelligent car cockpit system is solved, and more efficient and accurate service recommendation is achieved, improving user experience.
Patent Information
- Application Number
- CN202510419606.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-11
AI Technical Summary
The existing smart car cockpit system lacks a comprehensive understanding of user intentions and environment, resulting in the inadequate recommendation of vehicle control services.
By obtaining voice input data, image information and parameter information, a multimodal recommendation model is used to recommend services, including multimodal semantic fusion decision model, intelligent model and collaborative recommendation model, the vehicle-controlled atomic service set is determined, and screened and executed.
It improves the recommendation accuracy of vehicle control services, shortens the recommendation link, protects user data privacy, and improves user experience.
Smart Images

Figure CN120301937A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of intelligent vehicles, and particularly relates to a method, device, vehicle, and storage medium for recommending vehicle control services. Background Art
[0002] With the rapid development of intelligent vehicle technology, users' demand for the intelligence of automotive cockpits is increasing day by day. However, most existing intelligent vehicle cockpit systems only provide single voice control or visual assistance functions, lacking comprehensive understanding of user intentions and the environment, resulting in limited service experience.
[0003] Although some systems of existing in-vehicle vehicle control services can collect multimodal information, they lack advanced intelligent algorithms and models for comprehensive decision-making, resulting in less intelligent and accurate service recommendations. Summary of the Invention
[0004] Embodiments of this application provide a method, device, vehicle, and storage medium for recommending vehicle control services, which can effectively improve the accuracy of vehicle control service recommendations.
[0005] The technical solution of this application is implemented as follows:
[0006] Embodiments of this application provide a method for recommending vehicle control services. The method for recommending vehicle control services includes:
[0007] In response to a preset operation on the vehicle, based on the obtained voice input data, determine multimodal information; wherein, the multimodal information includes voice information, image information, and parameter information; the preset operation is used to trigger multimodal recognition;
[0008] Based on the voice information, the image information, and the parameter information, through a preset multimodal recommendation model, perform service recommendation to determine a set of vehicle control atomic services; wherein, the preset multimodal recommendation model is used to recommend vehicle control services through multimodal information;
[0009] In response to a screening operation on the set of vehicle control atomic services, determine the final vehicle control atomic service; and execute the final vehicle control atomic service.
[0010] It can be understood that, on the one hand, in response to a preset operation for the vehicle, based on the acquired voice input data, multimodal information is determined. Since the preset operation is used to trigger multimodal recognition, after triggering multimodal recognition, voice information, image information, and parameter information can be obtained at one time, avoiding the interception of voice information and image information. At the same time, it is wake-up free, voice distribution free, NLU interception free, and the recommendation link is shortened, improving the recommendation efficiency. On the other hand, based on the voice information, image information, and parameter information, service recommendation is performed through a preset multimodal recommendation model to determine the vehicle control atomic service set, and in response to a screening operation for the vehicle control atomic service set, the final vehicle control atomic service is determined, which can improve the accuracy of vehicle control service recommendation.
[0011] In the above solution, the determining of multimodal information based on the acquired voice input data in response to a preset operation for the vehicle includes:
[0012] In response to the preset operation for the vehicle, the voice input data is acquired;
[0013] Based on the preset operation and the voice input data, the voice information, image data, and the parameter information are acquired;
[0014] The image data is recognized to determine the image information corresponding to the image data.
[0015] It can be understood that in response to a preset operation for the vehicle, voice input data is acquired; based on the preset operation and the voice input data, voice information, image data, and parameter information are acquired; the image data is recognized to determine the image information corresponding to the image data, thereby determining multimodal information and enriching the multimodal information.
[0016] In the above solution, the preset multimodal recommendation model includes: a multimodal semantic fusion decision model, an intelligent model, and a collaborative recommendation model;
[0017] The determining of the vehicle control atomic service set by performing service recommendation through a preset multimodal recommendation model based on the voice information, the image information, and the parameter information includes:
[0018] Through the multimodal semantic fusion decision model, the voice information, the image information, and the parameter information are sorted out for intent to determine the comprehensive intent;
[0019] Based on the comprehensive intent and the parameter information, through the intelligent model and the collaborative recommendation model, service recommendation is performed to determine the vehicle control atomic service set.
[0020] It can be understood that through the multi-modal semantic fusion decision-making model, the voice information, image information, and parameter information are sorted out for intent, and the comprehensive intent is determined; based on the comprehensive intent and parameter information, through the intelligent model and collaborative recommendation model, service recommendations are made to determine the vehicle control atomic service set. Since the service recommendations are made through the trained model, the accuracy of the vehicle control service recommendations can be improved.
[0021] In the above solution, the process of sorting out the intent of the voice information, image information, and parameter information through the multi-modal semantic fusion decision-making model to determine the comprehensive intent includes:
[0022] Through the multi-modal semantic fusion decision-making model, feature extraction is respectively performed on the voice information, image information, and parameter information to determine the respective feature points of the voice information, image information, and parameter information;
[0023] Based on the respective feature points of the voice information, image information, and parameter information, intent sorting is performed to determine the comprehensive intent corresponding to the voice information.
[0024] It can be understood that through the multi-modal semantic fusion decision-making model, feature extraction is respectively performed on the voice information, image information, and parameter information to determine the respective feature points of the voice information, image information, and parameter information; based on the respective feature points of the voice information, image information, and parameter information, intent sorting is performed to determine the comprehensive intent corresponding to the voice information, which helps to clearly understand the user's intent requirements, thereby further improving the accuracy of the vehicle control service recommendations.
[0025] In the above solution, the process of making service recommendations through the intelligent model and collaborative recommendation model based on the comprehensive intent and parameter information to determine the vehicle control atomic service set includes:
[0026] Based on the comprehensive intent, parameter information, and the obtained vehicle whole atomic service list, service recommendations are made through the intelligent model to determine the initial vehicle control atomic service set;
[0027] The initial vehicle control atomic service set is screened through the collaborative recommendation model and the obtained vehicle-end real-time data to determine the vehicle control atomic service set.
[0028] It can be understood that service recommendations are made through the intelligent model to determine the initial vehicle control atomic service set; the initial vehicle control atomic services are screened through the collaborative recommendation model, making the determined vehicle control atomic service set more accurate.
[0029] In the above solution, the process of making service recommendations through the preset multi-modal recommendation model based on the voice information, image information, and parameter information to determine the vehicle control atomic service set includes:
[0030] Encrypt the voice information, the image information, and the parameter information to obtain encrypted voice information, encrypted image information, and encrypted parameter information;
[0031] Upload the encrypted voice information, the encrypted image information, and the encrypted parameter information to a server, so that the server performs service recommendation through the preset multi-modal recommendation model to determine the vehicle control atomic service set;
[0032] Receive the vehicle control atomic service set sent by the server.
[0033] It can be understood that by encrypting the voice information, the image information, and the parameter information to obtain encrypted voice information, encrypted image information, and encrypted parameter information; uploading the encrypted voice information, the encrypted image information, and the encrypted parameter information to a server, so that the server performs service recommendation through the preset multi-modal recommendation model to determine the vehicle control atomic service set; receiving the vehicle control atomic service set sent by the server; through data encryption, the transmission and storage security of user data are ensured, and at the same time, the user is allowed to customize the data sharing scope to protect the user's personal privacy.
[0034] In the above solution, before determining the vehicle control atomic service set by performing service recommendation through the preset multi-modal recommendation model based on the voice information, the image information, and the parameter information, the method further includes:
[0035] Obtain historical multi-modal information of different vehicles;
[0036] Based on the historical multi-modal information, train the initial multi-modal recommendation model until the output of the initial multi-modal recommendation model reaches a preset condition to determine the preset multi-modal recommendation model.
[0037] It can be understood that by training the initial multi-modal recommendation model with the historical multi-modal information of different vehicles until the output of the initial multi-modal recommendation model reaches a preset condition to determine the preset multi-modal recommendation model, the accuracy of the preset multi-modal recommendation model can be improved.
[0038] In the above solution, the method further includes:
[0039] Store the voice information, the image information, the parameter information, and the vehicle control atomic service set;
[0040] Train the preset multi-modal recommendation model with the voice information, the image information, the parameter information, and the vehicle control atomic service set to obtain an updated preset multi-modal recommendation model.
[0041] It can be understood that by using voice information, image information, parameter information, and vehicle control atomic service sets to train a preset multi-modal recommendation model, an updated preset multi-modal recommendation model is obtained, further improving the accuracy of the preset multi-modal recommendation model.
[0042] An embodiment of the present application provides a recommendation device for vehicle control services, including a determination unit, a recommendation unit, and an execution unit; wherein,
[0043] The determination unit is configured to, in response to a preset operation for the vehicle, determine multi-modal information based on the acquired voice input data; wherein, the multi-modal information includes voice information, image information, and parameter information; the preset operation is used to trigger multi-modal recognition.
[0044] The recommendation unit is configured to, based on the voice information, the image information, and the parameter information, perform service recommendation through a preset multi-modal recommendation model to determine a vehicle control atomic service set; wherein, the preset multi-modal recommendation model is used to recommend vehicle control services through multi-modal information; and is configured to, in response to a screening operation for the vehicle control atomic service set, determine a final vehicle control atomic service.
[0045] The execution unit is configured to execute the final vehicle control atomic service.
[0046] An embodiment of the present application provides a vehicle, including:
[0047] A memory for storing executable data instructions;
[0048] A processor, when executing the executable instructions stored in the memory, implements the vehicle control service recommendation method described above.
[0049] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a processor to implement the vehicle control service recommendation method when executed.
[0050] An embodiment of the present application provides a method, device, vehicle and computer-readable storage medium for recommending vehicle control services, wherein the method for recommending vehicle control services includes: in response to a preset operation for a vehicle, determining multimodal information based on acquired voice input data; wherein the multimodal information includes voice information, image information and parameter information; based on the voice information, the image information and the parameter information, recommending services through a preset multimodal recommendation model, and determining a vehicle control atomic service set; wherein the preset multimodal recommendation model is used to recommend vehicle control services through multimodal information; in response to a screening operation for the vehicle control atomic service set, determining a final vehicle control atomic service; and executing the final vehicle control atomic service. By adopting the above scheme, on the one hand, in response to a preset operation for the vehicle, multimodal information is determined based on the acquired voice input data. Since the preset operation is used to trigger multimodal recognition, voice information, image information and parameter information can be obtained at one time after the multimodal recognition is triggered, thereby avoiding the interception of voice information and image information. At the same time, there is no need for wake-up, voice distribution, or NLU interception, and the recommendation link is shortened, thereby improving the recommendation efficiency. On the other hand, based on voice information, image information and parameter information, service recommendations are made through a preset multimodal recommendation model, the vehicle control atomic service set is determined, and in response to the screening operation on the vehicle control atomic service set, the final vehicle control atomic service is determined, thereby improving the accuracy of vehicle control service recommendations. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 An optional process diagram of a method for recommending a vehicle control service is provided for an embodiment of the present application Figure 1 ;
[0052] Figure 2 An optional process diagram of a method for recommending a vehicle control service is provided for an embodiment of the present application Figure 2 ;
[0053] Figure 3 An optional process diagram of a method for recommending a vehicle control service is provided for an embodiment of the present application Figure 3 ;
[0054] Figure 4 An optional process diagram of a method for recommending a vehicle control service is provided for an embodiment of the present application Figure 4 ;
[0055] Figure 5 An optional process diagram of a method for recommending a vehicle control service is provided for an embodiment of the present application Figure 5 ;
[0056] Figure 6 An optional process diagram of a method for recommending a vehicle control service is provided for an embodiment of the present application Figure 6 ;
[0057] Figure 7 This is a schematic structural diagram of a recommendation device for vehicle control services provided by an embodiment of the present application;
[0058] Figure 8 This is a schematic structural diagram of a vehicle provided by an embodiment of the present application. Detailed implementation manners
[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the present application in detail with reference to the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not intended to limit the scope of the present application.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0061] In the following descriptions, terms such as "some embodiments", "this embodiment", "embodiments of the present application", and examples are involved, which describe subsets of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0062] If similar descriptions such as "first / second" appear in the application documents, the following explanations are added. In the following descriptions, the terms "first\second\third" only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0063] The embodiments of the present application provide a method for recommending vehicle control services. Figure 1 This is an optional process schematic of a method for recommending vehicle control services provided by an embodiment of the present application. Figure 1 will be described in combination with Figure 1 the steps shown.
[0064] S101. In response to a preset operation on the vehicle, determine multimodal information based on the acquired voice input data; wherein, the multimodal information includes voice information, image information, and parameter information; the preset operation is used to trigger multimodal recognition.
[0065] In the embodiments of the present application, the multimodal information includes voice information, image information, and parameter information; wherein, the parameter information is the real-time state parameter of the vehicle.
[0066] Exemplarily, the parameter information may be vehicle speed, temperature, etc.
[0067] In some embodiments of the present application, the execution subject of the recommendation method for vehicle control services is a vehicle.
[0068] In some embodiments of the present application, the recommendation method for vehicle control services is applicable to the recommendation scenario of vehicle control services in an intelligent vehicle cockpit.
[0069] In some embodiments of the present application, the preset operation is used to trigger multimodal recognition, and the preset operation is set in advance. The preset operation may be a long press on the voice assistant button on the steering wheel. The specific manner of the preset operation is not specifically limited in the embodiments of the present application.
[0070] In some embodiments of the present application, the preset operation can trigger multimodal recognition of voice data and image data.
[0071] In some embodiments of the present application, in response to a preset operation on the vehicle, the vehicle obtains voice input data; based on the preset operation and the voice input data, the vehicle obtains voice information, image data, and parameter information; and the vehicle performs recognition on the image data to determine the image information corresponding to the image data.
[0072] In some embodiments of the present application, in response to a preset operation on the vehicle, the vehicle obtains voice input data through a sound collection device, and based on the voice input data, the vehicle obtains voice information, image data, and real-time state parameters of the vehicle. By performing recognition on the image data, the vehicle determines the image information corresponding to the image data.
[0073] In some embodiments of the present application, the voice input data is natural language with voice requirements input by the user. For example, the user says: "There is a sprinkler truck ahead."
[0074] It can be understood that in response to a preset operation on the vehicle, voice input data is obtained; based on the preset operation and the voice input data, voice information, image data, and parameter information are obtained; and the image data is recognized to determine the image information corresponding to the image data, thereby determining multimodal information and enriching the multimodal information.
[0075] S102. Based on the voice information, image information, and parameter information, perform service recommendation through a preset multimodal recommendation model to determine a set of vehicle control atomic services; wherein, the preset multimodal recommendation model is used to recommend vehicle control services through multimodal information.
[0076] In some embodiments of the present application, the preset multimodal recommendation model includes: a multimodal semantic fusion decision model, an intelligent model, and a collaborative recommendation model.
[0077] In some embodiments of the present application, the multi-modal semantic fusion decision-making model is mainly used to sort out the comprehensive intention according to multi-modal information.
[0078] In some embodiments of the present application, the intelligent model is mainly used to recommend a vehicle control atomic service set based on the vehicle's atomic service list, multi-modal intention requirements, and vehicle real-time parameters.
[0079] In some embodiments of the present application, the preset multi-modal recommendation model can be deployed on the local vehicle or on the server. If the preset multi-modal recommendation model is deployed on the server, after the vehicle obtains multi-modal information, it will encrypt the multi-modal information and upload it to the server, and the server will determine the vehicle control atomic service set and then send it down to the vehicle.
[0080] In some embodiments of the present application, the collaborative recommendation model is mainly used to match the functional service parameters output by the intelligent model with the vehicle-end real-time data to screen the vehicle control atomic services; and in the future, after forming a user profile through intelligent learning of user behavior, it can perform multi-vehicle and multi-account collaborative recommendations, which are ever-changing. For example, during the morning rush hour congestion, when multiple identical vehicles use the AI-recommended navigation function, the multi-vehicle can collaboratively allocate the navigation routes for each vehicle to reach the fastest overall path, so as to achieve the purpose of alleviating traffic congestion.
[0081] It should be noted that the multi-modal intention requirement is the comprehensive intention.
[0082] In some embodiments of the present application, the vehicle performs service recommendation through a preset multi-modal recommendation model based on voice information, image information, and parameter information, and determines a vehicle control atomic service set from the vehicle's atomic service list.
[0083] In some embodiments of the present application, the vehicle sorts out the intention of voice information, image information, and parameter information through a multi-modal semantic fusion decision-making model to determine the comprehensive intention; based on the comprehensive intention and parameter information, through the intelligent model and the collaborative recommendation model, it performs service recommendation to determine the vehicle control atomic service set.
[0084] In some embodiments of the present application, the vehicle's atomic service list is obtained by service-ifying and decoupling various vehicle function parameters into individual services, which can be freely pieced together like individual atoms. The atomic service list is to summarize the already deconstructed atomic services into a list table. Currently, thousands of vehicle-end atomic services have been deconstructed, such as the range of 0-100% for opening and closing the window. This is the basic ability of vehicle-end service intelligence.
[0085] In some embodiments of the present application, the vehicle control atomic service set is a service set combined by the intelligent model by calling multiple atomic services in the vehicle's atomic service list, that is, a service set selected from the above-mentioned vehicle's atomic service list.
[0086] Exemplarily, for instance, if there are 1,000 vehicle-level atomic service lists, the 5 services selected by the intelligent model based on the comprehensive intention are collectively referred to as the vehicle control atomic service set output this time.
[0087] S103. In response to a screening operation for the vehicle control atomic service set, determine the final vehicle control atomic service; and execute the final vehicle control atomic service.
[0088] In some embodiments of the present application, the final vehicle control atomic service is the vehicle control atomic service that finally needs to be executed after user feedback.
[0089] In some embodiments of the present application, the vehicle responds to a screening operation for the vehicle control atomic service set, determines the final vehicle control atomic service; and executes the final vehicle control atomic service.
[0090] In some embodiments of the present application, the screening operation for the vehicle control atomic service set is an operation in which the user can give secondary feedback to select to add or delete services therein.
[0091] In some embodiments of the present application, after the vehicle control atomic service set is transmitted to the vehicle end and displayed to the user through the central control, the user can give secondary feedback to select to add or delete services therein. All the secondary feedback data of the users can be used to feed back the model. The final vehicle control atomic service refers to the atomic service that the vehicle finally needs to execute after the user's secondary feedback.
[0092] It should be noted that the vehicle control atomic service set can be sent by the server to the vehicle, or can be generated by the vehicle itself. Specifically, it is determined according to the deployment location of the preset multi-modal recommendation model.
[0093] It can be understood that, on the one hand, in response to a preset operation for the vehicle, based on the acquired voice input data, multi-modal information is determined. Since the preset operation is used to trigger multi-modal recognition, after triggering multi-modal recognition, voice information, image information, and parameter information can be acquired at one time, avoiding the interception of voice information and image information. At the same time, it is wake-up-free, voice-distribution-free, NLU-interception-free, and shortens the recommendation link, improving the recommendation efficiency; on the other hand, based on the voice information, image information, and parameter information, service recommendation is performed through a preset multi-modal recommendation model to determine the vehicle control atomic service set, and in response to a screening operation for the vehicle control atomic service set, the final vehicle control atomic service is determined, which can improve the accuracy of vehicle control service recommendation.
[0094] In some embodiments of the present application, as Figure 2 shown, S102 can be implemented through S201 and S202, as follows:
[0095] S201. Through a multi-modal semantic fusion decision model, organize the intentions of the voice information, image information, and parameter information to determine the comprehensive intention.
[0096] In some embodiments of the present application, through a multi-modal semantic fusion decision model, feature extraction is respectively performed on voice information, image information, and parameter information to determine the feature points of the voice information, image information, and parameter information respectively; based on the feature points of the voice information, image information, and parameter information respectively, intention sorting is performed to determine the comprehensive intention corresponding to the voice information.
[0097] In some embodiments of the present application, the comprehensive intention is the intention of the problem that the user may need to solve determined based on multi-modal information. Generally, if the user only inputs unimodal information, such as the user only says what they need by voice, the intention is clear. However, if multi-modal and multi-information are input simultaneously, the user's intention will become blurred, and perhaps the user doesn't even know what they most need. The multi-modal semantic fusion decision model can integrate the feature points of the three input sources of voice information, image information, and vehicle real-time parameters (i.e., parameter information) to comprehensively judge what the user's general intention is in this scenario.
[0098] Exemplarily, the voice information is that it is raining outside, without a clear instruction to turn on the windshield wipers, nor specifying whether the rain is light rain or heavy rain. The vehicle can combine the fact that it is raining and visually recognize that it is heavy rain to judge that the user's comprehensive intention is to solve the problem of not being able to see the front field of vision clearly due to the rain.
[0099] S202: Based on the comprehensive intention and parameter information, through an intelligent model and a collaborative recommendation model, perform service recommendation to determine the vehicle control atomic service set.
[0100] In some embodiments of the present application, the vehicle can perform service recommendation through an intelligent model based on the comprehensive intention and parameter information to determine the initial vehicle control atomic service set; through the collaborative recommendation model, screen the initial vehicle control atomic service set to determine the vehicle control atomic service set.
[0101] In some embodiments of the present application, the vehicle performs service recommendation through an intelligent model based on the comprehensive intention, parameter information, and the obtained vehicle atomic service list to determine the initial vehicle control atomic service set; through the collaborative recommendation model and the obtained vehicle-end real-time data, screen the initial vehicle control atomic service set to determine the vehicle control atomic service set.
[0102] In some embodiments of the present application, the vehicle performs service recommendation through an intelligent model based on the comprehensive intention, parameter information, and the obtained vehicle atomic service list to determine the initial vehicle control atomic service set corresponding to the comprehensive intention; based on the obtained vehicle-end real-time data, filter the initial vehicle control atomic service set through the collaborative recommendation model to remove the vehicle control atomic services that are repeated with the vehicle-end real-time data, and determine the vehicle control atomic service set.
[0103] Exemplarily, some existing states on the vehicle side are filtered out through the collaborative recommendation model. For example, if the output service is to adjust the air conditioner temperature to 24°C, but it is detected that the vehicle-side air conditioner temperature is already 24°C after being transmitted back to the vehicle side from the intelligent model, then this service is automatically filtered out.
[0104] It can be understood that through the multi-modal semantic fusion decision model, the voice information, image information, and parameter information are sorted out for intentions to determine the comprehensive intention; based on the comprehensive intention and parameter information, through the intelligent model and the collaborative recommendation model, service recommendations are made to determine the vehicle control atomic service set. Since the service recommendations are made through the trained models, the accuracy of the vehicle control service recommendations can be improved.
[0105] In some embodiments of the present application, as Figure 3 shown, S102 can also be implemented through S301, S302, and S303, as follows:
[0106] S301. Encrypt the voice information, image information, and parameter information to obtain the encrypted voice information, encrypted image information, and encrypted parameter information.
[0107] In some embodiments of the present application, the vehicle encrypts the voice information, image information, and parameter information to obtain the encrypted voice information, encrypted image information, and encrypted parameter information.
[0108] S302. Upload the encrypted voice information, encrypted image information, and encrypted parameter information to the server so that the server makes service recommendations through a preset multi-modal recommendation model to determine the vehicle control atomic service set.
[0109] In some embodiments of the present application, the vehicle uploads the encrypted voice information, encrypted image information, and encrypted parameter information to the server so that the server makes service recommendations through a preset multi-modal recommendation model to determine the vehicle control atomic service set.
[0110] In some embodiments of the present application, after receiving the encrypted voice information, encrypted image information, and encrypted parameter information, the server decrypts them to obtain the voice information, image information, and parameter information. The server makes service recommendations through a preset multi-modal recommendation model based on the voice information, image information, and parameter information to determine the vehicle control atomic service set.
[0111] In some embodiments of the present application, the preset multi-modal recommendation model includes a multi-modal semantic fusion decision model, an intelligent model, and a collaborative recommendation model.
[0112] In some embodiments of the present application, the server organizes the intent of voice information, image information, and parameter information through a multi-modal semantic fusion decision model to determine the comprehensive intent; based on the comprehensive intent and parameter information, through an intelligent model and a collaborative recommendation model, service recommendations are made to determine the vehicle control atomic service set.
[0113] S303. Receive the vehicle control atomic service set sent by the server.
[0114] In some embodiments of the present application, the vehicle receives the vehicle control atomic service set sent by the server.
[0115] It can be understood that by encrypting the voice information, image information, and parameter information, the encrypted voice information, encrypted image information, and encrypted parameter information are obtained; the encrypted voice information, encrypted image information, and encrypted parameter information are uploaded to the server so that the server makes service recommendations through a preset multi-modal recommendation model to determine the vehicle control atomic service set; receive the vehicle control atomic service set sent by the server; through data encryption, ensure the security of user data transmission and storage, and at the same time allow users to customize the data sharing scope to protect user personal privacy.
[0116] In some embodiments of the present application, the method for recommending vehicle control services further includes:
[0117] Store the voice information, image information, parameter information, and vehicle control atomic service set;
[0118] Train a preset multi-modal recommendation model with the voice information, image information, parameter information, and vehicle control atomic service set to obtain an updated preset multi-modal recommendation model.
[0119] In some embodiments of the present application, the vehicle stores the voice information, image information, parameter information, and vehicle control atomic service set; trains a preset multi-modal recommendation model with the voice information, image information, parameter information, and vehicle control atomic service set to obtain an updated preset multi-modal recommendation model. The updated preset multi-modal recommendation model is used for subsequent vehicle control service recommendations.
[0120] It can be understood that by encrypting the voice information, image information, and parameter information, the encrypted voice information, encrypted image information, and encrypted parameter information are obtained; the encrypted voice information, encrypted image information, and encrypted parameter information are uploaded to the server so that the server makes service recommendations through a preset multi-modal recommendation model to determine the vehicle control atomic service set; receive the vehicle control atomic service set sent by the server; through data encryption, ensure the security of user data transmission and storage, and at the same time allow users to customize the data sharing scope to protect user personal privacy.
[0121] In some embodiments of the present application, before executing S102, S104 and S105 are also executed, as follows:
[0122] S104. Obtain the historical multimodal information of different vehicles.
[0123] In some embodiments of the present application, a vehicle can obtain the historical multimodal information of different vehicles.
[0124] In some embodiments of the present application, the historical multimodal information includes historical voice information, historical image information, and historical parameter information; there is a corresponding relationship among the historical voice information, historical image information, and historical parameter information, that is, the historical voice information, historical image information, and historical parameter information of the same vehicle are used as a sample data for model training.
[0125] It should be noted that the historical voice information includes at least two historical voice sub-informations; the historical image information includes at least two historical image sub-informations; the historical parameter information includes at least two historical image sub-informations. Different historical voice sub-informations can be of the same vehicle or different vehicles.
[0126] S105. Based on the historical multimodal information, train the initial multimodal recommendation model until the output of the initial multimodal recommendation model reaches a preset condition, and determine the preset multimodal recommendation model.
[0127] In some embodiments of the present application, the preset condition is that the accuracy rate of the output of the initial multimodal recommendation model reaches a preset threshold.
[0128] In some embodiments of the present application, a vehicle trains the initial multimodal recommendation model based on the historical multimodal information until the accuracy rate of the output of the initial multimodal recommendation model reaches a preset threshold, outputs the initial multimodal recommendation model at this time, and determines it as the preset multimodal recommendation model.
[0129] It can be understood that by training the initial multimodal recommendation model with the historical multimodal information of different vehicles until the output of the initial multimodal recommendation model reaches a preset condition and determining the preset multimodal recommendation model, the accuracy of the preset multimodal recommendation model can be improved.
[0130] In some embodiments of the present application, such as Figure 4As shown in the figure, the recommended method for vehicle control services includes: S1. The user inputs semantics. Specifically, the user inputs voice, or long-presses the steering wheel voice key to trigger the vehicle body camera to obtain image data, forming semantic input. The voice and image model identifies whether the user input semantics contains strong feature 1. If it contains strong feature 1, it continues to filter and screen the secondary feature 3 involved. At the same time, the real-time image 6 is identified and the environmental information 4 is determined according to the non-strong feature 2. After sorting out all the feature information (i.e., strong feature 1, non-strong feature 2, and secondary feature 3) and environmental information 4 (i.e., personnel, weather, road conditions, status, light), it is input to the intelligent model 5 through the prompting project 8 (prompt). The intelligent model 5 outputs the functions and parameters 7 to be provided by the vehicle (i.e., the initial vehicle control atomic service set). After these service functions and parameters are transmitted back to the vehicle end, S2. Real-time parameter service filtering is executed. After executing S2, a vehicle control atomic service set 9 is generated, which is interacted with the interface and fed back to the user. Then S3. The user feedback (i.e., secondary feedback) is executed. Finally, the atomic service after feedback is executed.
[0131] It should be noted that strong features are artificially defined. Once a strong feature appears, it will definitely affect the service set recommended by the vehicle, such as a sprinkler truck, heavy rain, etc.; non-strong features are other features not included in the strong features; secondary features are artificially defined secondary detailed features under the strong features, such as the sprinkler truck is of the high gun type, sweeping type, etc.
[0132] It can be understood that by using voice information, image information, parameter information, and the vehicle control atomic service set, the preset multi-modal recommendation model is trained to obtain an updated preset multi-modal recommendation model, further improving the accuracy of the preset multi-modal recommendation model.
[0133] In some embodiments of the present application, the recommended process of the vehicle control service in the related technology is as Figure 5 shown. When the voice input is used as the trigger condition currently, there are several problems in the technical link: ① It needs to be awakened and triggered; ② The current NLU distribution will intercept some user requirements (i.e., NLU will intercept); ③ The determination model 12 can only take the voice input rejected by the voice module 11 for secondary determination to see if the user has hidden requirements, and then pass these requirements through the recommendation model 13. It is very difficult to determine and there are no features; ④ The entire link is too long, and the output delay of the large model cannot be guaranteed when walking again, and the user experience is poor.
[0134] In some embodiments of the present application, the recommended process of the vehicle control service in the present application is as Figure 6As shown, through the interaction method of long - pressing the voice button on the steering wheel, the triggering method is decoupled from the voice wake - up trigger. Through the new triggering method, the link of the voice module 11 can be skipped, and all data can be directly handed over to the recommendation model 13 for work. It has the advantages of being able to avoid wake - up, avoid voice module interception, and avoid secondary determination by the determination model 12. At the same time, it is an independent channel, which can clearly give a trigger signal for the multi - modal model to work, so as not to occupy visual resources by enabling multi - modal recognition work for a long time.
[0135] The three - dimensional implementation of the recommended method for vehicle control services in this application includes seven modules, specifically as follows:
[0136] S1: Voice interaction module;
[0137] S2: Image recognition module;
[0138] S3: Vehicle - wide real - time parameter acquisition module;
[0139] S4: Intelligent decision - making module;
[0140] S5: Recommended interaction module;
[0141] S6: Service execution module;
[0142] S7: Security and privacy protection module.
[0143] The specific working methods of the seven modules in this application are as follows:
[0144] S1: Voice interaction module: This module includes the user long - pressing the voice assistant button on the steering wheel to wake up the voice assistant without voice, and inputting the voice requirements of natural language. For example, when the user says: "There is a sprinkler in front", the voice interaction module accepts and recognizes the user's intention and description, converts the user's natural language description into the atomic requirements of the vehicle control service that may be required, and sends the information to the intelligent decision - making module after sorting.
[0145] It should be noted that the information here is voice information.
[0146] S2: Image recognition module: This module includes the user sending a signal to call the front - view, surround - view, and in - vehicle IMS cameras outside the vehicle to collect the current in - vehicle and out - of - vehicle image data after pressing the voice assistant button on the steering wheel, and analyzing and describing the information in the current image through the AI image recognition model, and sending the information to the intelligent decision - making module after sorting.
[0147] It should be noted that the information here is image information.
[0148] S3: Vehicle Real-time Parameter Acquisition Module: When the user presses the steering wheel voice assistant button, a signal is sent to this module, which is responsible for collecting real-time parameter information related to the vehicle, such as vehicle speed, temperature, torque and other vehicle parameters, and sending the sorted data information to the intelligent decision-making module.
[0149] It should be noted that the information here is parameter information.
[0150] S4: Intelligent Decision-making Module: This module consists of a vehicle atomic service list, a multi-modal semantic fusion decision-making model, an intelligent model, and a collaborative recommendation model. When the intelligent decision-making module receives the three types of information: voice, vision, and vehicle real-time parameters, it sorts out the comprehensive intention through the multi-modal semantic fusion decision-making model, and then combines through the prompt engineering, inputs the vehicle atomic service list, multi-modal intention requirements, and vehicle real-time parameters into the intelligent model at the same time. The intelligent model outputs the set of vehicle control atomic services that can be provided to the user at this time and transmits it back to the local EDC recommendation interaction module.
[0151] S5: Recommendation Interaction Module: This module includes an intelligent recommendation HMI window interaction and feedback system. After the intelligent model transmits back the set of vehicle control atomic services, it recommends the set of vehicle control atomic services to the user in the form of HMI of TTS and AI functions, seeking the user's feedback; then the user can use natural language voice or touch click to feedback to the interaction module, and can choose to cancel or add atomic services in the set of vehicle control atomic services, and then send a signal to the service execution module.
[0152] S6: Service Execution Module: Make data points in advance, store the data of the user's feedback, record the current vehicle interior and exterior environment information and the values of the user's feedback, and store and retain these values for the backfeeding training of the multi-modal recommendation model; at the same time, after receiving the signal of the user's feedback, execute the corresponding atomic vehicle control service.
[0153] S7: Security and Privacy Protection Module: Ensure the security of the transmission and storage of user data through data encryption and desensitization, and at the same time allow users to customize the data sharing scope to protect the user's personal privacy; at the same time, the model can be deployed locally to keep the data in the vehicle and not go to the cloud.
[0154] It can be understood that by comprehensively analyzing the user's voice commands, the images of the in-vehicle and out-of-vehicle cameras, and the vehicle status information, personalized vehicle control service recommendations are provided for the user, improving driving safety and the user experience. At the same time, the trigger interaction method is to long-press the steering wheel voice button, opening a new channel, which can achieve wake-up-free, voice distribution-free, NLU interception-free, shortening the recommendation link, and improving the recommendation efficiency.
[0155] This application embodiment also provides a recommendation device for vehicle control services, as Figure 7 shown Figure 7Schematic structural diagram of a recommended device for vehicle control services provided by an embodiment of the present application. The recommended device 7 for vehicle control services includes: a determination unit 701, a recommendation unit 702, and an execution unit 703; wherein,
[0156] The determination unit 701 is configured to, in response to a preset operation on the vehicle, determine multimodal information based on the acquired voice input data; wherein, the multimodal information includes voice information, image information, and parameter information; the preset operation is used to trigger multimodal recognition.
[0157] The recommendation unit 702 is configured to perform service recommendation through a preset multimodal recommendation model based on the voice information, the image information, and the parameter information, and determine a set of vehicle control atomic services; wherein, the preset multimodal recommendation model is used to recommend vehicle control services through multimodal information; and is configured to determine the final vehicle control atomic service in response to a screening operation on the set of vehicle control atomic services.
[0158] The execution unit 703 is configured to execute the final vehicle control atomic service.
[0159] In some embodiments of the present application, the recommended device 7 for vehicle control services further includes: an acquisition unit 704; wherein,
[0160] The acquisition unit 704 is configured to, in response to the preset operation on the vehicle, acquire the voice input data; based on the preset operation and the voice input data, acquire the voice information, image data, and the parameter information.
[0161] The determination unit 701 is further configured to identify the image data to determine the image information corresponding to the image data.
[0162] In some embodiments of the present application, the preset multimodal recommendation model includes: a multimodal semantic fusion decision model, an intelligent model, and a collaborative recommendation model.
[0163] The determination unit 701 is further configured to, through the multimodal semantic fusion decision model, organize the intentions of the voice information, the image information, and the parameter information to determine a comprehensive intention; based on the comprehensive intention and the parameter information, perform service recommendation through the intelligent model and the collaborative recommendation model to determine the set of vehicle control atomic services.
[0164] In some embodiments of the present application, the determining unit 701 is further configured to extract features from the voice information, the image information, and the parameter information respectively through the multi-modal semantic fusion decision model, and determine the feature points of the voice information, the image information, and the parameter information respectively; based on the feature points of the voice information, the image information, and the parameter information respectively, perform intention collation to determine the comprehensive intention corresponding to the voice information.
[0165] In some embodiments of the present application, the determining unit 701 is further configured to, based on the comprehensive intention, the parameter information, and the obtained vehicle whole vehicle atomic service list, perform service recommendation through the intelligent model to determine an initial vehicle control atomic service set; screen the initial vehicle control atomic service set through the collaborative recommendation model and the obtained vehicle-end real-time data to determine the vehicle control atomic service set.
[0166] In some embodiments of the present application, the obtaining unit 704 is further configured to encrypt the voice information, the image information, and the parameter information to obtain encrypted voice information, encrypted image information, and encrypted parameter information;
[0167] The determining unit 701 is further configured to upload the encrypted voice information, the encrypted image information, and the encrypted parameter information to the server, so that the server performs service recommendation through the preset multi-modal recommendation model to determine the vehicle control atomic service set; receive the vehicle control atomic service set sent by the server.
[0168] In some embodiments of the present application, the obtaining unit 704 is further configured to obtain historical multi-modal information of different vehicles before determining the vehicle control atomic service set by performing service recommendation through the preset multi-modal recommendation model based on the voice information, the image information, and the parameter information;
[0169] The determining unit 701 is further configured to train the initial multi-modal recommendation model based on the historical multi-modal information until the output of the initial multi-modal recommendation model reaches a preset condition, and determine the preset multi-modal recommendation model.
[0170] In some embodiments of the present application, the obtaining unit 704 is further configured to store the voice information, the image information, the parameter information, and the vehicle control atomic service set; train the preset multi-modal recommendation model through the voice information, the image information, the parameter information, and the vehicle control atomic service set to obtain an updated preset multi-modal recommendation model.
[0171] Based on the vehicle control service recommendation method of the above embodiments, an embodiment of the present application further provides a vehicle, as Figure 8 shownFigure 8 A structural schematic diagram of a vehicle provided by an embodiment of the present application. The vehicle 10 includes: a processor 801 and a memory 802. The memory 802 is used to store a computer program. The processor 801 is used to call and run the computer program from the memory 802 to execute the recommendation method for vehicle control services as described in the above embodiments.
[0172] In the embodiments of the present application, the above-mentioned processor 801 may be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic devices for implementing the above-mentioned processor functions may also be others, and the embodiments of the present application do not make specific limitations.
[0173] The embodiments of the present application provide a computer-readable storage medium storing a computer program, which is used to implement the recommendation method for vehicle control services as described in any of the above embodiments when executed by a first processor or a second processor.
[0174] Exemplarily, the program instructions corresponding to a recommendation method for vehicle control services in this embodiment may be stored on a storage medium such as an optical disc, a hard disk, or a USB flash drive. When the program instructions corresponding to a recommendation method for vehicle control services in the storage medium are read or executed by an electronic device, the recommendation method for vehicle control services as described in any of the above embodiments can be implemented.
[0175] In addition, in the embodiments of the present application, each functional module may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional module.
[0176] In addition, in the embodiments of the present application, each functional module may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional module.
[0177] When the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable the vehicle to execute all or part of the steps of the method of this embodiment.
[0178] It should be understood that the "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" or "in some embodiments" that appear throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitude of the serial numbers of the above processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments. The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated herein.
[0179] The modules described above as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules; they may be located in one place or distributed to multiple network units; some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0180] In addition, in each embodiment of the present application, each functional module can be fully integrated in a processing unit, or each module can be separately used as a unit, or two or more modules can be integrated in a unit; the above integrated modules can be implemented in the form of hardware, or in the form of hardware plus software functional units.
[0181] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium, and when the program is executed, it executes the steps including the above method embodiments.
[0182] The methods disclosed in several method embodiments provided by the embodiments of the present application can be arbitrarily combined without conflict to obtain new method embodiments.
[0183] The features disclosed in several product embodiments provided by the embodiments of the present application can be arbitrarily combined without conflict to obtain new product embodiments.
[0184] The features disclosed in several method or device embodiments provided by the embodiments of the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0185] As mentioned above, it is only the implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claimed rights.
Claims
1. A method for recommending vehicle control services, characterized in that, The recommended method for vehicle control services includes: In response to a preset operation on the vehicle, based on the obtained voice input data, determine multimodal information; wherein, the multimodal information includes voice information, image information, and parameter information; the preset operation is used to trigger multimodal recognition; Based on the voice information, the image information, and the parameter information, perform service recommendation through a preset multimodal recommendation model to determine a set of vehicle control atomic services; wherein, the preset multimodal recommendation model is used to recommend vehicle control services through multimodal information; In response to a screening operation on the set of vehicle control atomic services, determine the final vehicle control atomic service; and execute the final vehicle control atomic service.
2. The recommended method for vehicle control services according to claim 1, characterized in that The step of, in response to a preset operation on the vehicle, based on the obtained voice input data, determining multimodal information includes: In response to the preset operation on the vehicle, obtain the voice input data; Based on the preset operation and the voice input data, obtain the voice information, image data, and the parameter information; Identify the image data to determine the image information corresponding to the image data.
3. The recommended method for vehicle control services according to claim 1, wherein, The preset multimodal recommendation model includes: a multimodal semantic fusion decision model, an intelligent model, and a collaborative recommendation model; The step of, based on the voice information, the image information, and the parameter information, performing service recommendation through a preset multimodal recommendation model to determine a set of vehicle control atomic services includes: Through the multimodal semantic fusion decision model, organize the intentions of the voice information, the image information, and the parameter information to determine a comprehensive intention; Based on the comprehensive intention and the parameter information, perform service recommendation through the intelligent model and the collaborative recommendation model to determine the set of vehicle control atomic services.
4. The recommended method for vehicle control services according to claim 3, characterized in that, The step of, through the multimodal semantic fusion decision model, organizing the intentions of the voice information, the image information, and the parameter information to determine a comprehensive intention includes: Through the multimodal semantic fusion decision model, perform feature extraction on the voice information, the image information, and the parameter information respectively to determine the respective feature points of the voice information, the image information, and the parameter information; Based on the respective feature points of the voice information, the image information, and the parameter information, organize the intentions to determine the comprehensive intention corresponding to the voice information.
5. The recommended method for vehicle control services according to claim 3, wherein, The step of, based on the comprehensive intention and the parameter information, performing service recommendation through the intelligent model and the collaborative recommendation model to determine the set of vehicle control atomic services includes: Based on the comprehensive intention, the parameter information, and the obtained list of vehicle atomic services, perform service recommendation through the intelligent model to determine an initial set of vehicle control atomic services; Through the collaborative recommendation model and the obtained vehicle-end real-time data, screen the initial set of vehicle control atomic services to determine the set of vehicle control atomic services.
6. The recommended method for vehicle control services according to claim 1, wherein, The step of, based on the voice information, the image information, and the parameter information, performing service recommendation through a preset multimodal recommendation model to determine a set of vehicle control atomic services includes: Encrypt the voice information, the image information, and the parameter information to obtain encrypted voice information, encrypted image information, and encrypted parameter information; Upload the encrypted voice information, the encrypted image information, and the encrypted parameter information to a server, so that the server performs service recommendation through the preset multi-modal recommendation model to determine the vehicle control atomic service set; Receive the vehicle control atomic service set sent by the server.
7. The recommended method for vehicle control services according to any one of claims 1-6, characterized in that, Before performing service recommendation through the preset multi-modal recommendation model based on the voice information, the image information, and the parameter information to determine the vehicle control atomic service set, the method further includes: Obtain historical multi-modal information of different vehicles; Train an initial multi-modal recommendation model based on the historical multi-modal information until the output of the initial multi-modal recommendation model reaches a preset condition, and determine the preset multi-modal recommendation model.
8. The recommended method for vehicle control services according to any one of claims 1-6, characterized in that The method further includes: Store the voice information, the image information, the parameter information, and the vehicle control atomic service set; Train the preset multi-modal recommendation model through the voice information, the image information, the parameter information, and the vehicle control atomic service set to obtain an updated preset multi-modal recommendation model.
9. A recommendation device for vehicle control services, characterized in that, Includes a determination unit, a recommendation unit, and an execution unit; wherein, The determination unit is configured to, in response to a preset operation for a vehicle, determine multi-modal information based on acquired voice input data; wherein the multi-modal information includes voice information, image information, and parameter information; the preset operation is used to trigger multi-modal recognition; The recommendation unit is configured to perform service recommendation through a preset multi-modal recommendation model based on the voice information, the image information, and the parameter information to determine a vehicle control atomic service set; wherein the preset multi-modal recommendation model is used to recommend vehicle control services through multi-modal information; and is configured to determine a final vehicle control atomic service in response to a screening operation for the vehicle control atomic service set; The execution unit is configured to execute the final vehicle control atomic service.
10. A vehicle, characterized in that, Includes: A memory for storing executable data instructions; A processor, when executing the executable instructions stored in the memory, implements the vehicle control service recommendation method according to claims 1-8.
11. A computer-readable storage medium, characterized in that, Stores executable instructions, which when executed by a processor, implement the vehicle control service recommendation method according to any one of claims 1 to 8.