Vehicle scenario service orchestration method and apparatus, device, medium, and program product
By determining the user's voice intent and behavioral habits, the system intelligently matches atomic vehicle services, solving the problems of unfriendly vehicle service orchestration and high usage threshold, and realizing personalized vehicle service customization.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CHONGQING CHANGAN TECH CO LTD
- Filing Date
- 2025-09-26
- Publication Date
- 2026-07-30
AI Technical Summary
The existing vehicle service orchestration function requires manual operation by users, which is unfriendly, has a high barrier to entry, and makes it difficult for users to accurately describe their service needs.
By determining the explicitness of the user's voice information intent, matching vehicle atomized services, configuring scenario services based on user behavior habits, and using intent classification models and service orchestration models for intelligent and personalized service orchestration.
It enables intelligent conversion and personalized customization of user voice commands into vehicle services, improving the convenience and applicability of vehicle services and reducing operational difficulty.
Smart Images

Figure CN2025124341_30072026_PF_FP_ABST
Abstract
Description
Vehicle scenario service orchestration methods, devices, equipment, media, and program products
[0001] Cross-references
[0002] This application claims priority to Chinese Patent Application No. 202510099196.2, filed on January 22, 2025, entitled "Vehicle Scene Service Orchestration Method, Apparatus, Equipment, Medium, and Program Product", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to, but is not limited to, the field of automotive intelligent control technology, specifically to a vehicle scene service orchestration method, device, equipment, medium, and program product. Background Technology
[0004] Vehicle scenario services refer to specific actions a user takes when using a vehicle in a particular situation. For example, when driving in summer, users might activate voice navigation, play music, or turn on the air conditioning. Currently, the vehicle service orchestration function requires users to manually select the desired services by dragging, pulling, or dropping them. Summary of the Invention
[0005] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.
[0006] This disclosure provides a method, apparatus, device, medium, and program product for orchestrating vehicle scene services.
[0007] According to an embodiment of this disclosure, a vehicle scene service orchestration method is provided. The vehicle scene service orchestration method includes: determining the explicitness of the intent of the voice information input by the user; determining the vehicle atomic service that matches the voice information based on the explicitness of the intent; and configuring the vehicle atomic service based on the user's behavioral habits to obtain the currently required vehicle scene service.
[0008] Optionally, determining the clarity of the intent of the user-inputted voice information includes: obtaining the user-inputted voice information; converting the voice information based on the user's voice characteristics to obtain text information corresponding to the voice information; the voice characteristics include: accent, speech rate, and pronunciation habits; and determining the clarity of the intent of the voice information based on the text information.
[0009] Optionally, the above-mentioned determination of the explicitness of the intent of the speech information based on text information includes: sequentially inputting the text information into multiple feature extraction modules in the intent classification model, and outputting the text features corresponding to each feature extraction module; performing feature fusion on the text features corresponding to multiple feature extraction modules to obtain fused text features; inputting the fused text features into the classification module in the intent classification model, and outputting the classification result; the classification result characterizes whether the intent of the text information is explicit.
[0010] Optionally, the above-mentioned determination of the vehicle atomic service matching the voice information based on the clarity of intent includes: when the intent is clear, inputting the text information corresponding to the voice information into the first service orchestration model and outputting the vehicle atomic service matching the text information; when the intent is unclear, inputting the text information corresponding to the voice information into the second service orchestration model and outputting the vehicle atomic service matching the text information.
[0011] Optionally, the first service orchestration model is trained through the following steps: selecting the first vehicle scene service with a clear intent from the historical vehicle scene services orchestrated before the current moment; preprocessing the first vehicle scene service to obtain the processed first vehicle scene service; training the first fine-tuning model of the attention layer of the base model based on the processed first vehicle scene service to obtain the trained first fine-tuning model; and determining the first service orchestration model based on the trained first fine-tuning model and the base model.
[0012] Optionally, the preprocessing of the first vehicle scenario service mentioned above includes at least one of the following: data cleaning of the first vehicle scenario service; anomaly removal of the first vehicle scenario service; and format conversion of the first vehicle scenario service.
[0013] Optionally, the above-mentioned first fine-tuning model of the attention layer of the training base model based on the processed first vehicle scene service, to obtain the trained first fine-tuning model, includes: inputting the processed first vehicle scene service into the first fine-tuning model and outputting the processing result of the first fine-tuning model; determining the overall weight matrix of the first fine-tuning model and the base model, and the increment matrix of the first fine-tuning model based on the output result of the first fine-tuning model; adjusting the parameter matrix of the first fine-tuning model based on the overall weight matrix and the increment matrix until the iteration stopping condition is reached, to obtain the trained first fine-tuning model.
[0014] Optionally, the second service orchestration model described above is trained through the following steps: filtering out second vehicle scene services with unclear intentions from the historical vehicle scene services orchestrated before the current moment; preprocessing the second vehicle scene services to obtain processed second vehicle scene services; training a second fine-tuning model of the attention layer of the base model based on the processed second vehicle scene services to obtain a trained second fine-tuning model; and determining the second service orchestration model based on the trained second fine-tuning model and the base model.
[0015] Optionally, the above configuration of vehicle atomization services based on user behavior habits to obtain the currently required vehicle scenario services includes: obtaining commonly used parameters of the user regarding vehicle atomization services from the service parameter management system based on the user's identifier and the functional category of the vehicle atomization services; the service parameter management system updating in real time based on changes in user behavior and changes in the vehicle's environment; commonly used parameters representing user behavior habits; and configuring vehicle atomization services based on commonly used parameters to obtain the currently required vehicle scenario services.
[0016] Optionally, the vehicle scene service orchestration method further includes: performing security checks on the vehicle scene service; performing conflict checks on the vehicle scene service; using security checks and conflict checks to determine the executability of the vehicle scene service; and, if both security checks and conflict checks indicate that the vehicle scene service is free of anomalies, calling the application programming interface of the vehicle scene service to execute the vehicle scene service.
[0017] According to an embodiment of this disclosure, a vehicle scene service orchestration apparatus is provided. The vehicle scene service orchestration apparatus includes: a determining module, used to determine the explicitness of the intent of the voice information input by the user; a processing module, used to determine the vehicle atomic service matching the voice information based on the explicitness of the intent; and the processing module is further used to configure the vehicle atomic service based on the user's behavioral habits to obtain the currently required vehicle scene service.
[0018] According to another aspect of the present disclosure, the present application provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor, and the processor executes the program to implement some or all of the steps in the above-described method.
[0019] According to another aspect of the present disclosure, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above-described method.
[0020] According to another aspect of the present disclosure, the present application provides a computer program including computer-readable code, which, when executed in a computer device, causes a processor in the computer device to perform some or all of the steps in the above-described method.
[0021] According to another aspect of the present disclosure, the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above method.
[0022] The beneficial effects of the embodiments disclosed herein are as follows:
[0023] The process involves determining the explicitness of the intent behind the user's voice input; based on this explicitness, identifying the corresponding vehicle-specific service; and configuring this service based on the user's behavioral habits to obtain the desired vehicle scenario service. This intelligently converts user voice into vehicle-specific services and enables personalized customization of vehicle scenario services. Therefore, it not only supports voice-based vehicle scenario service creation but also allows for personalized customization of these services, providing users with practical and comfortable vehicle scenario services. Attached Figure Description
[0024] The accompanying drawings, which are included to provide an understanding of embodiments of this disclosure and form part of this application, are used to explain this application and do not constitute an undue limitation thereof. In the drawings:
[0025] Figure 1 is a schematic diagram of the implementation process of a vehicle scene service orchestration method provided in an embodiment of this application;
[0026] Figure 2 is a schematic diagram of the implementation process of a vehicle scene service orchestration method provided in an embodiment of this application;
[0027] Figure 3 is a schematic diagram of the implementation process of a vehicle scene service orchestration method provided in an embodiment of this application;
[0028] Figure 4 is a schematic diagram of the implementation process of a vehicle scene service orchestration method provided in an embodiment of this application;
[0029] Figure 5 is a schematic diagram of the implementation process of the orchestration big model in a vehicle scene service orchestration method provided in an embodiment of this application;
[0030] Figure 6 is a schematic diagram of the composition structure of a vehicle scene service orchestration device provided in an embodiment of this application;
[0031] Figure 7 is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this application. Embodiments of the present invention
[0032] The technical solution of this application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] To enable those skilled in the art to better understand the embodiments of this application, the technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of this application. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.
[0034] The terms "first," "second," and "third," etc., in the specification, embodiments, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0035] The embodiments of this application will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be understood that the preferred embodiments are only for illustrating this application and are not intended to limit the scope of protection of this application.
[0036] Currently, the vehicle service orchestration function suffers from at least the following problems: First, the manual orchestration method is unfriendly; second, user descriptions of services vary greatly, posing a significant barrier to entry, making it difficult for users to fluently articulate the required scenario services and dependencies; third, users are largely unaware of the parameters of services they frequently use, making it difficult for them to accurately utilize services. Because the visual interface cannot display all the atomic services a vehicle possesses at once, users are unclear about which atomic services the vehicle offers, resulting in a high barrier to entry and time-consuming operation.
[0037] Therefore, this application provides a vehicle scene service orchestration method that can be applied to computer devices, including but not limited to fixed devices and / or mobile devices. For example, the fixed devices include, but are not limited to, personal computers (PCs) or servers, where the server can be a cloud server or a regular server. The computer device can be a device located in a vehicle or a device outside of a vehicle.
[0038] Figure 1 is a schematic diagram of the implementation process of a vehicle scene service orchestration method provided in an embodiment of this application. As shown in Figure 1, the method may include the following steps 101 to 103:
[0039] Step 101: Determine the clarity of the intent behind the user's voice input.
[0040] Voice information refers to the voice commands conveyed by the user to the vehicle, or the voice signals collected by the vehicle. The intent conveyed by the voice information refers to the user's current service request. The clarity of intent refers to whether the intent is explicit.
[0041] In some implementations, step 101 can be specifically implemented by: recognizing the user's input voice information based on the user's voice characteristics to determine the clarity of the intent conveyed by the voice information. Voice characteristics include at least: accent, speech rate, and vocalization habits.
[0042] In some implementations, step 101 can also be implemented as follows: converting the user's voice input based on the user's voice characteristics to obtain corresponding text information; recognizing the text information to determine the clarity of the intent of the voice information.
[0043] Step 102: Based on the explicitness of the intent, determine the vehicle atomization service that matches the voice information.
[0044] Vehicle atomization services refer to breaking down complex vehicle functions into small, independent basic units. These units are implemented as independent and reusable functional modules through Application Programming Interfaces (APIs). Vehicle atomization services that match voice information refer to one or more vehicle atomization services that match user needs.
[0045] In some implementations, step 102 can be specifically implemented by: determining the vehicle atomic service that matches the voice information using different matching methods based on the explicitness of the intent. The matching methods include at least service orchestration model-based matching methods, manual matching methods, and database-based matching methods.
[0046] In one feasible implementation, if the matching method is a service orchestration model-based matching method, then when the intent is clear, the first service orchestration model is used to determine the vehicle atomic service matching the voice information; when the intent is unclear, the second service orchestration model is used to determine the vehicle atomic service matching the voice information. Both the first and second service orchestration models can be large language models (LLMs).
[0047] In one feasible implementation, if the matching method is a database-based matching method, then when the intent is clear, a first database matching method is used to determine the vehicle atomic service matching the voice information; when the intent is unclear, a second database matching method is used to determine the vehicle atomic service matching the voice information. The first database matching method is used to explicitly filter out vehicle atomic services matching the voice information from the database based on the intent. The second database matching method is used to filter out vehicle atomic services matching the voice information from the database using a fuzzy method; specifically, it first determines text information similar to the voice information, and then filters out vehicle atomic services matching that similar text information from the database.
[0048] Step 103: Configure the vehicle atomic service based on the user's behavior habits to obtain the vehicle scenario service required at the moment.
[0049] In some implementations, step 103 can be specifically implemented as follows: based on the user's behavioral habits, determine the user's commonly used parameters for vehicle atomization services; based on the commonly used parameters, configure the vehicle atomization services and Xining to obtain the currently required vehicle scenario services.
[0050] In this embodiment, the clarity of the intent conveyed by the user's voice input is determined; based on the clarity of the intent, a vehicle-specific service matching the voice input is determined; and the vehicle-specific service is configured based on the user's behavioral habits to obtain the currently required vehicle scenario service. Thus, determining the vehicle-specific service matching the voice input based on its clarity enables intelligent conversion from user voice to vehicle-specific services; configuring the vehicle-specific service based on the user's behavioral habits enables personalized customization of vehicle scenario services. Therefore, this not only supports users specifying vehicle scenario services using voice but also enables personalized customization of vehicle scenario services, creating practical and comfortable vehicle scenario services for users.
[0051] Figure 2 is a schematic diagram of the implementation process of a vehicle scene service orchestration method provided in an embodiment of this application. As shown in Figure 2, the method may include the following steps 201 to 2022:
[0052] Step 201: Obtain the voice information input by the user.
[0053] In some implementations, step 201 can be specifically implemented by: receiving voice information conveyed by the user to the vehicle through an application on a mobile phone. Alternatively, it can be implemented by receiving voice information conveyed by the user to the vehicle through the vehicle's visual interface.
[0054] When implemented, when a user issues a voice command, the vehicle system quickly activates the voice acquisition module to collect voice information.
[0055] Step 202: Convert the voice information based on the user's voice characteristics to obtain the text information corresponding to the voice information; the voice characteristics include: accent, speech rate, and pronunciation habits.
[0056] In some implementations, step 202 can be specifically implemented as follows: noise reduction algorithms are used to process the speech information; speech recognition technology is used to analyze the processed speech signal based on the user's speech characteristics to obtain the text information corresponding to the speech information. The speech recognition technology may include, but is not limited to, using speech recognition models, speech recognition engines, etc.
[0057] Speech recognition technology can be categorized as Automatic Speech Recognition (ASR). Noise reduction processing of speech information is performed to eliminate interference from environmental noise and ensure the accuracy of the collected speech information.
[0058] In some implementations, step 202 can be specifically implemented as follows: slicing the speech information based on the user's speech characteristics to obtain multiple speech slices; converting the multiple speech slices based on the user's speech characteristics to obtain text information of the multiple speech slices; and integrating the multiple speech slices to obtain the text information corresponding to the speech information.
[0059] For example, when a user says, "Turn the screen toward me when you start navigation," the vehicle system can accurately recognize the text information in the voice command without misinterpreting extraneous noise text.
[0060] It should be noted that the voice processing section adopts industry-leading voice recognition technology to provide users with efficient, accurate, and convenient voice recognition, enabling the vehicle system to hear and understand clearly, which greatly improves the practicality of the vehicle scene service orchestration method provided in this application embodiment.
[0061] Step 203: Based on the text information, determine the clarity of the intent of the voice information.
[0062] Here, steps 201 to 203 correspond to step 101 mentioned above, and can be implemented with reference to the specific implementation of step 101 mentioned above.
[0063] In some implementations, step 203 can be specifically implemented by using a binary classification model to process the text information and determine whether the intent of the voice information is clear.
[0064] In some implementations, step 203 can be achieved through the following steps 2031 to 2033:
[0065] Step 2031: Input the text information sequentially into multiple feature extraction modules in the intent classification model, and output the text features corresponding to each feature extraction module.
[0066] The intent classification model is a pre-trained binary classification model. It is used to determine whether the intent of the speech information is explicit. In one feasible implementation, the intent classification model may include: multiple feature extraction modules, a feature fusion layer, and a classification layer. The feature extraction modules can be Transformer layers used to extract text features from the text information; each Transformer layer may include a multi-head attention mechanism and a feedforward neural network; the multi-head attention mechanism is used to capture the relationships between different features in the text information, and the feedforward neural network is used to further extract and transform features from the output of the multi-head attention mechanism. The feature fusion layer is used to fuse the text features corresponding to multiple feature extraction modules. The classification layer is used to determine whether the intent of the text information is explicit based on the fused text features.
[0067] Step 2032: Perform feature fusion on the feature representations corresponding to the multiple feature extraction modules to obtain fused text features.
[0068] In some implementations, a feature fusion layer can be used to fuse the outputs of multiple Transformer layers to obtain fused text features.
[0069] Step 2033: Input the fused text features into the classification module of the intent classification model and output the classification result; the classification result indicates whether the intent of the text information is clear.
[0070] The classification module refers to the classification layer. For example, the classification layer can be a softmax layer.
[0071] Step 204: Based on the explicitness of the intent, determine the vehicle atomization service that matches the voice information.
[0072] In some implementations, step 204 can be achieved through the following steps 2041 to 2042:
[0073] Step 2041: If the intent is clear, input the text information corresponding to the voice information into the first service orchestration model and output the vehicle atomization service that matches the text information.
[0074] The first service orchestration model is used to process explicit intents. The first service orchestration model can be trained through steps a through d below.
[0075] Step 2042: If the intent is unclear, input the text information corresponding to the voice information into the second service orchestration model and output the vehicle atomization service that matches the text information.
[0076] Ambiguous intent, also known as vague intent, is addressed by the first service orchestration model. The second service orchestration model can be trained using steps e through h.
[0077] It should be noted that the first service orchestration model and the second service orchestration model have two capabilities: first, to understand the relationship between the user's plain language (text information that recognizes the user's speech) and the vehicle-side atomic services; second, to understand and infer the service combination that the user may want (vehicle atomic services that match the text information) based on the user's plain language.
[0078] Step 205: Based on the user's identifier and the functional category of the vehicle atomic service, obtain the user's commonly used parameters related to the vehicle atomic service from the service parameter management system; the service parameter management system updates in real time based on changes in the user's behavior and changes in the vehicle's environment; the commonly used parameters represent the user's behavioral habits.
[0079] A user's identifier can be an identification code (ID) used to identify the user. The service parameter management system is used to collect users' driving habits and adaptively correct the parameters of the generated services when automatically generating and creating vehicle scenario services, so as to achieve personalized vehicle scenario services and truly make vehicle scenario orchestration services different for each individual.
[0080] In some implementations, the specific method for constructing a service parameter management system can be as follows: Information about user usage of vehicle-side services is collected through event tracking and uploaded to the cloud; data mining and machine learning techniques are used to analyze and extract the collected information, generating user tags with semantic information; and the user tags are classified, stored, and indexed to obtain the service parameter management system. It should be noted that this application only collects legally recognized information and does not involve user privacy. Specifically, user tags are classified, stored, and indexed according to user identifiers and the functional categories of vehicle-based atomic services.
[0081] Update user tags in real time or periodically. Adjust and add new tags promptly as user behavior and preferences change, and delete tags that are no longer applicable. For example, the user's frequently used air conditioning fan speed and temperature settings over the past seven days may change with the seasons.
[0082] In some implementations, step 205 can be specifically implemented as follows: obtaining user tags based on user identifiers, and obtaining commonly used service parameters of the user regarding the vehicle atomization service in the recent period based on user tags and the functional category of the vehicle atomization service.
[0083] Step 206: Configure the vehicle atomic service based on the commonly used parameters to obtain the vehicle scenario service required at the moment.
[0084] Here, steps 205 to 206 correspond to step 103 mentioned above, and can be implemented with reference to the specific implementation of step 103 mentioned above.
[0085] In some embodiments, after step 206, the following steps may also be performed: performing security checks on the vehicle scene service; performing conflict checks on the vehicle scene service; the security checks and the conflict checks are used to determine the executability of the vehicle scene service; if both the security checks and the conflict checks indicate that the vehicle scene service is without abnormalities, the application programming interface of the vehicle scene service is invoked to execute the vehicle scene service.
[0086] Performing security and conflict detection on vehicle-related services ensures their safe execution in the current environment. Both security and conflict detection are built upon expert systems to prevent security issues during service execution. For example, opening the sunroof is incompatible with heavy rain outside; similarly, service execution conflicts are avoided, such as opening a car window conflicting with activating the external air circulation.
[0087] In some embodiments, the first service orchestration model in step 2041 can be obtained by training through the following steps a to d:
[0088] Step a: Select the first vehicle scenario service with a clear intent from the historical vehicle scenario services arranged before the current moment.
[0089] Historical vehicle scenario services refer to vehicle scenario services scheduled before the current moment. First vehicle scenario services refer to vehicle scenario services with a clear intent.
[0090] In some implementations, step a can be specifically implemented by: using a manual screening method to select the first vehicle scenario service with a clear intent from the historical vehicle scenario services.
[0091] In some implementations, step a can be further implemented as follows: using the above intent classification model to determine the explicitness of intent for multiple historical vehicle scenario services; and taking the historical vehicle scenario service with explicit intent as the first vehicle scenario service.
[0092] Step b: Preprocess the first vehicle scene service to obtain the processed first vehicle scene service.
[0093] In some implementations, step b, "preprocessing the first vehicle scene service", includes at least one of the following: cleaning the first vehicle scene service; removing anomalies from the first vehicle scene service; and converting the format of the first vehicle scene service.
[0094] Data cleaning of the first vehicle scenario service aims to remove duplicate, meaningless, and incorrectly formatted records. Anomaly removal of the first vehicle scenario service removes records that do not conform to the actual context. Format conversion of the first vehicle scenario service transforms it into a question-and-answer format for easier model processing. For example: Question (user voice) - Answer (matched vehicle atomic service).
[0095] Step c: Based on the first fine-tuned model of the attention layer of the first vehicle scene service training base model after processing, the trained first fine-tuned model is obtained.
[0096] The first fine-tuning model can refer to the LoRA fine-tuning model. The LoRA fine-tuning model consists of two trainable low-rank matrices (parameter matrices), and the LoRA matrix can be inserted into the attention layer.
[0097] In some implementations, step c can be specifically implemented as follows: adjust the parameter matrix of the first fine-tuned model of the attention layer of the base model based on the processed first vehicle scene service until the iteration stopping condition is met, and obtain the trained first fine-tuned model.
[0098] In some implementations, step c can be achieved through steps c1 to c3 as follows:
[0099] Step c1: Input the processed first vehicle scene service into the first fine-tuning model, and output the processing result of the first fine-tuning model.
[0100] Step c2: Based on the output of the first fine-tuning model, determine the overall weight matrix of the first fine-tuning model and the base model, as well as the increment matrix of the first fine-tuning model.
[0101] Step c3: Based on the overall weight matrix and the incremental matrix, adjust the parameter matrix of the first fine-tuned model until the iteration stopping condition is met, and obtain the trained first fine-tuned model.
[0102] In a standard neural network model, the computation of a linear layer can be represented as: y = Wx; where W is the weight matrix, x is the input, and y is the output. Traditional fine-tuning requires updating this large matrix W, while LoRA keeps W constant and adds a trainable increment matrix ∆W, resulting in the fine-tuned output as: y = (W + ∆W)x.
[0103] To reduce the number of parameters in ∆W, LoRA decomposes ∆W into two low-rank matrices A and B, i.e., ∆W = AB; where the dimensions of A and B are much smaller than W, thus reducing the number of parameters that need to be updated during fine-tuning.
[0104] During training, the gradient descent algorithm is used to update the parameters of the A and B matrices introduced by LoRA. The W of the original model is not updated. The forward propagation of the model will still calculate the complete Wx, and at the same time, it will also calculate the increment ∆Wx=(AB)x, and then add them together. In this way, LoRA fine-tuning can adapt to new task data.
[0105] Step d: Based on the trained first fine-tuning model and the base model, determine the first service orchestration model.
[0106] In some embodiments, the second service orchestration model in step 2042 can be trained through steps e to h as follows:
[0107] Step e: Filter out the second vehicle scenario service with unclear intent from the historical vehicle scenario services arranged before the current moment.
[0108] It should be noted that the training method of the second service orchestration model is similar to that of the first service orchestration model, and the specific implementation method of the first service orchestration model can be referred to in the implementation.
[0109] The second type of vehicle scenario service refers to vehicle scenario services with unclear intent.
[0110] In some implementations, step e can be specifically implemented by: using a manual screening method to select the first vehicle scenario service with unclear intent from the historical vehicle scenario services.
[0111] In some implementations, step e can be further implemented as follows: using the above intent classification model to determine the clarity of intent for multiple historical vehicle scenario services; and taking historical vehicle scenario services with unclear intent as the first vehicle scenario service.
[0112] Step f: Preprocess the second vehicle scene service to obtain the processed second vehicle scene service.
[0113] In some implementations, the "preprocessing the second vehicle scene service" in step f includes at least one of the following: data cleaning of the second vehicle scene service; anomaly removal of the second vehicle scene service; and format conversion of the second vehicle scene service.
[0114] Step g: Based on the processed second vehicle scene service training base model, the second fine-tuning model of the attention layer is obtained to obtain the trained second fine-tuning model.
[0115] The second fine-tuning model can also be the LoRA fine-tuning model.
[0116] In some implementations, step g can be specifically implemented as follows: adjust the parameter matrix of the second fine-tuned model of the attention layer of the base model based on the processed second vehicle scene service until the iteration stopping condition is met, and obtain the trained second fine-tuned model.
[0117] In some implementations, step g can be further implemented as follows: inputting the processed second vehicle scene service into the second fine-tuning model and outputting the processing result of the second fine-tuning model; based on the output result of the second fine-tuning model, determining the overall weight matrix of the second fine-tuning model and the base model, as well as the increment matrix of the second fine-tuning model; and adjusting the parameter matrix of the second fine-tuning model based on the overall weight matrix and the increment matrix until the iteration stopping condition is reached, thereby obtaining the trained second fine-tuning model.
[0118] Step h: Based on the trained second fine-tuning model and the base model, determine the second service orchestration model.
[0119] It should be noted that this application is based on the same base model and performs dynamic scheduling on different LoRA fine-tuning models to achieve vehicle scenario service generation under different needs.
[0120] In this embodiment, the impact of different accents, speaking speeds, and pronunciation habits on speech recognition is fully considered during the conversion of user speech into text information, achieving high-accuracy speech recognition. Based on the clarity of intent inherent in the user's input speech information, different service orchestration models are used to determine the vehicle atomic services that match the speech information. This intelligently realizes the understanding and creation relationship between user speech and vehicle atomic services. Based on user behavior habits, the parameters of vehicle atomic services are personalized, enabling fast, accurate, and personalized vehicle scene services. Therefore, it not only supports users in defining vehicle scene services using voice but also enables personalized customization of vehicle scene services, creating practical and comfortable vehicle scene services for users.
[0121] The following describes the application of the vehicle scene service orchestration method provided in this application embodiment in a real-world scenario.
[0122] Background research reveals that current technologies have the following three main shortcomings: 1. Existing solutions only define and cloud-based orchestration functions related to pre-processing, triggering, and execution of orchestration business functions, and do not support the ability to create scenario services based on voice; 2. Existing solutions cannot support service creation capabilities; 3. Existing solutions cannot support service personalization capabilities.
[0123] To address the above issues, this application proposes a creative and personalized voice scene orchestration scheme (corresponding to the aforementioned vehicle scene service orchestration method), which is a technical solution based on a generative large model. This application utilizes user voice input, employs a large model for scene modeling, and integrates the service creation capabilities of the large model's Chain of Thought (CoT) to automatically generate multiple service combinations of vehicle scene services for the user's required driving needs. Furthermore, by combining this with the user's driving habits, personalized scene services are provided to the user. In this way, the understanding and creation relationship between user voice and vehicle scene services is realized. Furthermore, by using the user's historical driving habits, service parameters are configured personally, enabling users to quickly, accurately, creatively, and personally generate vehicle scene services.
[0124] This application addresses the following three major pain points: First, it improves convenience by allowing users to directly speak through voice, thus overcoming the unfriendly nature of manual service arrangement. Second, it supports scenario arrangement in plain language, solving the problem of users having a significant learning curve with standard service descriptions. Third, it supports personalized service generation, enabling generated services to match users' driving habits, representing a breakthrough technology for intelligent interaction in smart cars.
[0125] This application includes four modules: speech understanding, service generation and creation, personalized service adaptation, and function invocation.
[0126] I. Speech comprehension;
[0127] In this application, speech processing plays a crucial role. Existing high-accuracy speech recognition technology accurately converts user-input speech into text information. When a user issues a voice command, the system quickly activates the speech acquisition module to collect sound signals and uses efficient algorithms for noise reduction, eliminating environmental noise interference and ensuring clear and accurate acquired speech information. Next, the speech recognition engine analyzes the processed speech information and converts it into understandable text information. In this process, the system fully considers different accents, speaking speeds, and pronunciation habits, achieving high-accuracy speech recognition through extensive training data and continuously optimized models.
[0128] For example, when a user says "Turn the screen toward me when you start navigation", the system can accurately recognize the instruction and provide the text to the next module. It will not misrecognize extraneous noise text. The system will accurately recognize and convey the instruction to the service generation and creation module.
[0129] The voice processing function, by adopting industry-leading voice recognition technology, provides users with efficient, accurate, and convenient voice recognition and understanding, enabling the system to hear and understand clearly, greatly improving the product's usability.
[0130] II. Service Generation and Creation;
[0131] The main tasks of service generation and creation are to enable the big oracle model to have two capabilities: first, to understand the relationship between the user's plain language and the vehicle's atomized services; and second, to understand and infer the service combinations that the user may want based on the user's plain language.
[0132] Based on the scenario requirements, a multi-layered Transformer architecture binary classification model is designed to distinguish whether a user's statement has a clear intent. For example: a clear intent—"When I start navigation, the screen turns towards me"; a vague intent—"I want to sleep." The specific implementation is as follows:
[0133] 1. Data Input: Receives data to be classified, with the data type being text.
[0134] 2. Multi-layer Transformer Processing: The input data is processed sequentially through multiple Transformer layers. Each Transformer layer includes a multi-head attention mechanism and a feedforward neural network. The multi-head attention mechanism is used to capture different feature relationships in the data, and the feedforward neural network further extracts and transforms features from the output of the attention mechanism.
[0135] 3. Feature Fusion: The text features output from multiple Transformer layers are fused to obtain fused text features.
[0136] 4. Classification output: Input the fused text features into the classification layer to obtain the binary classification result.
[0137] Depending on whether the user's voice intent is clear, different model paths are used, as follows:
[0138] Route 1 clarifies intent: Based on the LoRA_1 secondary fine-tuning orchestration model (corresponding to the first service orchestration model mentioned above), the recognized user's voice input is understood, and the corresponding atomic service on the vehicle side is output after deciding on the output. For example: input = When I start navigation, the screen turns towards me; output = {"trigger":"Navigation service, navigation parameters", "action":"Screen service, screen parameters"}.
[0139] Route Two: Fuzzy Intent. This route takes the user's voice input and uses a LoRA_2-based secondary fine-tuning orchestration model (corresponding to the second service orchestration model mentioned above). After understanding the user's vague, colloquial description, it creatively outputs corresponding atomic services for the vehicle based on CoT capabilities. For example: input = "I want to sleep"; output = {"action": "Ambient lighting service, ambient lighting parameters", "Music service, music parameters", "Seat service, seat parameters", "Window service, window parameters"}.
[0140] Technical Implementation Solution:
[0141] 1. Sample Data Collection and Preprocessing. Text data related to scenario arrangement was collected from multiple channels, including historical records of vehicle arrangement functions, product documentation, and manual annotations, resulting in explicit intent samples and fuzzy intent samples. The data volume for both explicit and fuzzy intent samples was no less than 20,000 records. After collecting the sample data, it was cleaned to remove duplicate, meaningless, and incorrectly formatted records. The samples were formatted according to the question-and-answer format used for fine-tuning the large model to facilitate model processing.
[0142] 2. LoRA Fine-tuning Model. A large open-source base model of a selected size is chosen, and a LoRA fine-tuning architecture is designed, including two trainable low-rank matrices that can be inserted into the attention layer. The detailed process is shown below:
[0143] In a standard neural network layer, the computation of a linear layer can be represented as: y = Wx; where W is the weight matrix, x is the input, and y is the output. Traditional fine-tuning requires updating this large matrix W, while LoRA keeps W constant and adds a trainable increment ∆W, resulting in the fine-tuned output: y = (W + ∆W)x.
[0144] To reduce the number of parameters in ∆W, LoRA decomposes ∆W into two low-rank matrices A and B, i.e., ∆W = AB; where the dimensions of A and B are much smaller than W, thus reducing the number of parameters that need to be updated during fine-tuning.
[0145] During training, the gradient descent algorithm is used to update the parameters of the A and B matrices introduced by LoRA. The W of the original model is not updated. The forward propagation of the model will still calculate the complete Wx and the incremental ∆Wx=(AB)x, and then add them together. In this way, LoRA fine-tuning can adapt to new task data.
[0146] 3. Once the two fine-tuned LoRA models are trained, they can be used in conjunction with the base model. When the preorder classification model categorizes the user input as a clear intent, the system will automatically call LoRA_1 and combine it with the base model to infer the input and output a vehicle scene service. When the preorder classification model categorizes the user input as a vague intent, the system will automatically call LoRA_2 and combine it with the base model to infer the input and create services, and output a vehicle scene service. Thus, this application can dynamically schedule different fine-tuned LoRA models based on the same base model to generate vehicle scene services for different needs.
[0147] III. Personalized service adapts to individual needs;
[0148] By building a tag system (corresponding to the service parameter management system mentioned above) to collect users' driving habits (corresponding to the user behavior habits mentioned above), when automatically generating and creating vehicle scene services, the tag system is called and the parameters of the generated services are adaptively corrected to achieve personalized vehicle scene services. This ensures that the vehicle scene orchestration service is tailored to each individual and is unique to each person.
[0149] The detailed technical solution is as follows:
[0150] 1. User Tag Collection. Signals from in-vehicle user service usage are collected and uploaded to the cloud via event tracking (excluding confidential user information; only legally permissible data is tracked). Data mining and machine learning techniques are used to analyze and extract semantic information from the collected data, generating user tags. Examples include: frequently used air conditioning fan speed and temperature settings over the past seven days.
[0151] 2. Tag Management and Updates. Establish a tag system to classify, store, and index user tags. Update user tags in real-time or periodically. As user behavior and preferences change, promptly adjust and add new tags, and delete tags that are no longer applicable. For example, as the seasons change, users' settings for air conditioning temperature and fan speed will also change.
[0152] 3. Personalized Service Generation Module. As shown in Figure 4, after the model generates the corresponding service, the system will query the user's frequently used service parameters for the corresponding vehicle scenario service in the recent period based on the user's identifier, and assign these parameters to the service parameters of the vehicle scenario service, thereby realizing personalized service generation.
[0153] The purpose of personalized service adaptation is to accurately grasp user needs by deeply mining and analyzing user tags, and to provide users with highly personalized services, thereby improving user satisfaction and service quality.
[0154] IV. Function Call;
[0155] After generating corresponding vehicle scenario services based on user descriptions, security and conflict detection must be performed on the service composition to ensure that the services can execute securely in the current environment. Both the security and conflict detection modules are built based on expert systems to avoid security issues during service execution; for example, opening the sunroof while it's raining heavily outside. They also avoid service execution conflicts; for example, opening a car window and activating the external air circulation.
[0156] In summary, this application provides users with a creative and personalized voice scene arrangement function. It not only supports users to define scenes via voice, but also allows for automatic service creation and combination through user-interacted speech. Furthermore, it automatically personalizes service settings based on user habits. This empowers users to become the product defining the scene, creating practical and comfortable driving scenarios for themselves.
[0157] As shown in Figure 3, the personalized voice scene service orchestration scheme of this application embodiment includes the following steps 301 to 304:
[0158] Step 301: Obtain user voice input.
[0159] Step 302: The large-scale service model automatically generates corresponding services based on user voice, and uses the general knowledge capabilities of the large-scale model to produce creative services based on prompting engineering.
[0160] Step 303: Obtain tag data of user's driving habits, switch service parameters in the generated service, and realize personalized service.
[0161] Step 304: Make service calls and execute based on the generated services.
[0162] As shown in Figure 4, the personalized voice scene service orchestration scheme of this application embodiment specifically includes four parts: voice input, orchestration big model, tagging system, and service execution. Voice input includes: voice input and speech recognition (ASR recognition). The orchestration big model includes: text input distribution, low-rank matrix, orchestration big model, and service composition; specifically, text input is distributed to the orchestration big model to obtain service composition; during this process, the low-rank matrix can be updated based on the text input. The tagging system includes: user identifier, user service tag, and new service composition; specifically, user service tag is obtained based on the user identifier to obtain a new service composition. Service execution includes: service API library, function calls, security and conflict detection, and vehicle scene service generation; specifically, function calls are performed through the service API library to perform security and conflict detection on the service composition output by the big model, and after detection, vehicle scene services are generated.
[0163] As shown in Figure 5, the processing flow of the orchestration big model includes: inputting the text input into the binary classification model; if the classification result is that the intent is clear, then the text input is processed based on LoRA_1 and the language big model (corresponding to the above base model) to obtain the service composition; if the classification result is that the intent is unclear, then the text input is processed based on LoRA_2 and the language big model to obtain the service composition.
[0164] The points to be protected in this application include at least: 1. A vehicle atomization service that automatically generates plain language speech based on a large model; 2. A vehicle scene service that combines a hundred plain language speech automatic creation services based on a large model; 3. A model architecture that dynamically solves two types of problems using a single large model; 4. A personalized scene service solution.
[0165] Based on the foregoing embodiments, this application provides a vehicle scene service orchestration device. The device includes various units and modules included in each unit, which can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0166] Figure 6 is a schematic diagram of the composition structure of a vehicle scene service orchestration device provided in an embodiment of this application. As shown in Figure 6, the vehicle scene service orchestration device 600 includes: a determining module 610 and a processing module 620, wherein:
[0167] The determining module 610 is used to determine the explicitness of the intent of the voice information input by the user;
[0168] Processing module 620 is used to determine a vehicle atomization service that matches the voice information based on the explicitness of the intent;
[0169] The processing module 620 is further configured to configure the vehicle atomic service based on the user's behavioral habits to obtain the currently required vehicle scenario service.
[0170] In some embodiments, the determining module 610 is further configured to obtain the voice information input by the user; convert the voice information based on the user's voice characteristics to obtain text information corresponding to the voice information; the voice characteristics include: accent, speech rate, and pronunciation habits; and determine the clarity of the intent of the voice information based on the text information.
[0171] In some embodiments, the determining module 610 is further configured to sequentially input the text information into multiple feature extraction modules in the intent classification model, and output the text features corresponding to each feature extraction module; perform feature fusion on the text features corresponding to multiple feature extraction modules to obtain fused text features; input the fused text features into the classification module in the intent classification model, and output the classification result; the classification result characterizes whether the intent of the text information is clear.
[0172] In some embodiments, the processing module 620 is further configured to, when the intent is clear, input the text information corresponding to the voice information into a first service orchestration model and output a vehicle atomization service matching the text information; when the intent is unclear, input the text information corresponding to the voice information into a second service orchestration model and output a vehicle atomization service matching the text information.
[0173] In some embodiments, the processing module 620 is further configured to: filter out a first vehicle scene service with a clear intent from the historical vehicle scene services orchestrated before the current time; preprocess the first vehicle scene service to obtain a processed first vehicle scene service; train a first fine-tuning model of the attention layer of the base model based on the processed first vehicle scene service to obtain a trained first fine-tuning model; and determine the first service orchestration model based on the trained first fine-tuning model and the base model.
[0174] In some embodiments, the processing module 620 is further configured to perform data cleaning on the first vehicle scene service; remove anomalies from the first vehicle scene service; and perform format conversion on the first vehicle scene service.
[0175] In some embodiments, the processing module 620 is further configured to input the processed first vehicle scene service into the first fine-tuning model and output the processing result of the first fine-tuning model; based on the output result of the first fine-tuning model, determine the overall weight matrix of the first fine-tuning model and the base model, as well as the increment matrix of the first fine-tuning model; based on the overall weight matrix and the increment matrix, adjust the parameter matrix of the first fine-tuning model until the iteration stopping condition is reached, and obtain the trained first fine-tuning model.
[0176] In some embodiments, the processing module 620 is further configured to: filter out a second vehicle scene service with unclear intent from the historical vehicle scene services orchestrated before the current time; preprocess the second vehicle scene service to obtain a processed second vehicle scene service; train a second fine-tuning model of the attention layer of the base model based on the processed second vehicle scene service to obtain a trained second fine-tuning model; and determine the second service orchestration model based on the trained second fine-tuning model and the base model.
[0177] In some embodiments, the processing module 620 is further configured to obtain common parameters of the user regarding the vehicle atomic service from the service parameter management system based on the user's identifier and the functional category of the vehicle atomic service; the service parameter management system updates in real time based on changes in the user's behavior and changes in the vehicle's environment; the common parameters characterize the user's behavioral habits; and the vehicle atomic service is configured based on the common parameters to obtain the currently required vehicle scenario service.
[0178] In some embodiments, the processing module 620 is further configured to perform security checks on the vehicle scene service; perform conflict checks on the vehicle scene service; the security checks and the conflict checks are used to determine the executability of the vehicle scene service; and, if both the security checks and the conflict checks indicate that the vehicle scene service is without abnormalities, call the application programming interface of the vehicle scene service to execute the vehicle scene service.
[0179] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this application can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0180] It should be noted that, in the embodiments of this application, if the above-described vehicle scenario service orchestration method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0181] This application provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.
[0182] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium can be transient or non-transient.
[0183] This application provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.
[0184] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0185] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referred to interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0186] It should be noted that Figure 7 is a schematic diagram of a hardware entity of a computer device in an embodiment of this application. As shown in Figure 7, the hardware entity of the computer device 700 includes: a processor 701, a communication interface 702, and a memory 703, wherein:
[0187] Processor 701 typically controls the overall operation of computer device 700.
[0188] Communication interface 702 enables computer devices to communicate with other terminals or servers over a network.
[0189] The memory 703 is configured to store instructions and applications executable by the processor 701, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 701 and various modules in the computer device 700. It can be implemented using flash memory or random access memory (RAM). Data transfer between the processor 701, the communication interface 702, and the memory 703 can be performed via bus 704.
[0190] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0191] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0192] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0193] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0194] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0195] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0196] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.
[0197] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for orchestrating vehicle scene services, comprising: Determine the explicitness of the intent behind the user's voice input; Based on the explicitness of the intent, determine the vehicle atomization service that matches the voice information; The vehicle atomic service is configured based on the user's behavioral habits to obtain the vehicle scenario service required at the moment.
2. The vehicle scene service orchestration method according to claim 1, wherein, The determination of the explicitness of the intent of the user's input voice information includes: Obtain the voice information input by the user; The voice information is converted based on the user's voice characteristics to obtain the corresponding text information; the voice characteristics include: accent, speech rate, and pronunciation habits; Based on the text information, the clarity of the intent conveyed by the voice information is determined.
3. The vehicle scene service orchestration method according to claim 2, wherein, The determination of the explicitness of the intent of the voice information based on the text information includes: The text information is sequentially input into multiple feature extraction modules in the intent classification model, and the text features corresponding to each feature extraction module are output. The text features corresponding to multiple feature extraction modules are fused to obtain fused text features; The fused text features are input into the classification module of the intent classification model, and the classification result is output; the classification result indicates whether the intent of the text information is clear.
4. The vehicle scene service orchestration method according to any one of claims 1 to 3, wherein, The determination of the vehicle atomization service matching the voice information based on the explicitness of the intent includes: When the intent is clear, the text information corresponding to the voice information is input into the first service orchestration model, and the vehicle atomization service matching the text information is output. In cases where the intent is unclear, the text information corresponding to the voice information is input into the second service orchestration model, and the vehicle atomization service matching the text information is output.
5. The vehicle scene service orchestration method according to claim 4, wherein, The first service orchestration model is trained through the following steps: From the historical vehicle scenario services arranged before the current moment, select the first vehicle scenario service with a clear intent; The first vehicle scene service is preprocessed to obtain the processed first vehicle scene service. Based on the first fine-tuned model of the attention layer of the processed first vehicle scene service training base model, the trained first fine-tuned model is obtained. Based on the trained first fine-tuning model and the base model, the first service orchestration model is determined.
6. The vehicle scene service orchestration method according to claim 5, wherein, The preprocessing of the first vehicle scenario service includes at least one of the following: Perform data cleaning on the first vehicle scenario service; Perform anomaly removal on the first vehicle scenario service; The format of the first vehicle scenario service is converted.
7. The vehicle scene service orchestration method according to claim 5, wherein, The first fine-tuned model based on the attention layer of the processed first vehicle scene service training base model, to obtain the trained first fine-tuned model, includes: The processed first vehicle scenario service is input into the first fine-tuning model, and the processing result of the first fine-tuning model is output. Based on the output of the first fine-tuning model, the overall weight matrix of the first fine-tuning model and the base model, as well as the increment matrix of the first fine-tuning model, are determined. Based on the overall weight matrix and the incremental matrix, the parameter matrix of the first fine-tuned model is adjusted until the iteration stopping condition is met, thus obtaining the trained first fine-tuned model.
8. The vehicle scene service orchestration method according to claim 4, wherein, The second service orchestration model is trained through the following steps: From the historical vehicle scenario services arranged before the current moment, filter out the second vehicle scenario service with unclear intent; The second vehicle scenario service is preprocessed to obtain the processed second vehicle scenario service; The second fine-tuned model is obtained by using the attention layer of the processed second vehicle scene service training base model as the second fine-tuned model. Based on the trained second fine-tuning model and the base model, the second service orchestration model is determined.
9. The vehicle scene service orchestration method according to claim 1, wherein, The configuration of the vehicle atomic service based on the user's behavioral habits to obtain the currently required vehicle scenario service includes: Based on the user's identifier and the functional category of the vehicle atomization service, the service parameter management system retrieves the user's commonly used parameters related to the vehicle atomization service from the service parameter management system; the service parameter management system updates in real time based on changes in the user's behavior and changes in the vehicle's environment; the commonly used parameters characterize the user's behavioral habits. Configure the vehicle atomic service based on the commonly used parameters to obtain the vehicle scenario service required at the moment.
10. The vehicle scene service orchestration method according to claim 1, wherein, The vehicle scenario service orchestration method also includes: Perform security checks on the vehicle scene services; Conflict detection is performed on the vehicle scene service; the security detection and the conflict detection are used to determine the executability of the vehicle scene service. If both the security detection and the conflict detection indicate that the vehicle scene service is normal, the application programming interface of the vehicle scene service is invoked to execute the vehicle scene service.
11. A vehicle scene service orchestration device, wherein, The vehicle scene service orchestration device includes: The determination module is used to determine the explicitness of the intent behind the user's voice input. A processing module is used to determine a vehicle atomization service that matches the voice information based on the explicitness of the intent; The processing module is also used to configure the vehicle atomic service based on the user's behavioral habits to obtain the vehicle scenario service required at the moment.
12. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, wherein, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, wherein, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 10.
14. A computer program product comprising a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is read and executed by a computer, it implements the steps of the method according to any one of claims 1 to 10.