Instruction processing method and apparatus, vehicle, and storage medium

CN122451130BActive Publication Date: 2026-09-29CHONGQING LANDIAN AUTOMOBILE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610945291.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-29
Estimated Expiration
2046-06-29

AI Technical Summary

Technical Problem

[0003]相关的路由分发方案通常采用固定规则或简单的二分类方式进行任务分发,分发速度快,但无法适应复杂多变的人机交互意图

Benefits of technology

[0032]响应于所述动作识别结果与目标动作相匹配,以及所述关键词识别结果与目标关键词相匹配中至少一者,基于所述车辆本地的第一图像处理模型对所述待处理图像进行图像识别,得到所述待处理图像对应的第一图像识别结果;所述第一图像识别结果,用于与预设的图像识别结果进行匹配,以判断是否将所述第一图像识别结果发送至所述指令文本对应的处理模型进行处理。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122451130B_ABST
    Figure CN122451130B_ABST
Patent Text Reader

Abstract

The application relates to an instruction processing method and device, a vehicle and a storage medium. The method comprises the following steps: acquiring an instruction text corresponding to an interaction instruction of a vehicle; in response to the instruction text matching a preset intention text, taking a text processing model of the vehicle as a processing model corresponding to the instruction text; in response to the instruction text not matching the preset intention text, performing uncertainty evaluation on the instruction text to obtain a first evaluation result, performing complexity evaluation on the instruction text to obtain a second evaluation result, and determining the processing model corresponding to the instruction text based on the first evaluation result and the second evaluation result; and processing the instruction text based on the processing model corresponding to the instruction text to obtain a processing result of the instruction text. The method can balance the processing efficiency and the processing accuracy of the vehicle interaction instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent vehicle technology, and in particular to an instruction processing method, apparatus, vehicle, and storage medium. Background Technology

[0002] With the rapid development of vehicle-to-everything (V2X) technology, vehicles have evolved from traditional transportation tools into mobile intelligent terminals with powerful sensing, communication, and computing capabilities. To address the limitations of computing and sensing capabilities during vehicle operation, vehicles can utilize networks and sensors to establish various communication modes, including vehicle-to-infrastructure, vehicle-to-vehicle, and vehicle-to-server communication. In the V2X environment, the routing and distribution technology for human-machine interaction commands between drivers and vehicles is one of the core technologies for achieving efficient communication.

[0003] Current routing and distribution schemes typically employ fixed rules or simple binary classification methods for task distribution. While these methods offer fast distribution speeds, they cannot adapt to complex and ever-changing human-machine interaction intentions. Therefore, how to process vehicle human-machine interaction commands more quickly and accurately has become one of the important research directions. Summary of the Invention

[0004] Based on this, this application addresses the aforementioned technical problems by providing an instruction processing method, apparatus, vehicle, and storage medium that can balance the processing efficiency and accuracy of vehicle interaction instructions.

[0005] In a first aspect, this application provides an instruction processing method, including:

[0006] Obtain the command text corresponding to the vehicle's interaction commands;

[0007] In response to the instruction text matching a preset intent text, the vehicle's local text processing model is used as the processing model corresponding to the instruction text.

[0008] In response to a mismatch between the instruction text and the preset intent text, an uncertainty assessment is performed on the instruction text to obtain a first assessment result, and a complexity assessment is performed on the instruction text to obtain a second assessment result. Based on the first assessment result and the second assessment result, the processing model corresponding to the instruction text is determined.

[0009] The instruction text is processed based on the processing model corresponding to the instruction text to obtain the processing result of the instruction text; the processing includes one or more of intent recognition, response text generation and instruction splitting; wherein, the response text generation includes performing intent recognition on the instruction text and generating response text based on the intent recognition result; the instruction splitting includes performing intent recognition on the instruction text, and if multiple intent recognition results are obtained, the instruction text is split into multiple instructions, with one instruction corresponding to one intent recognition result.

[0010] The above-mentioned instruction processing method has at least the following beneficial effects: For the instruction text of vehicle interaction instructions, it differentiates and matches the corresponding processing model. When the instruction text matches the preset intent text, it calls the vehicle's local text processing model to process the instruction text, improving the processing speed of conventional interaction instructions. When the instruction text does not match the preset intent text, it selects one or more suitable processing models by comprehensively evaluating the uncertainty and complexity of the instruction text. This achieves accurate differentiation of fuzzy, complex, and other unconventional interaction instructions and assigns them appropriate processing models, effectively balancing the response speed of conventional instruction text with the processing accuracy of unconventional instruction text. This improves both the processing efficiency and accuracy of vehicle interaction instructions.

[0011] In an optional embodiment of the first aspect, determining the processing model corresponding to the instruction text based on the first evaluation result and the second evaluation result includes:

[0012] The first evaluation result and the second evaluation result are merged to obtain a fused evaluation result;

[0013] In response to the fusion evaluation result satisfying the first evaluation condition, the local text processing model of the vehicle is used as the processing model corresponding to the instruction text;

[0014] In response to the fusion evaluation result satisfying the second evaluation condition, both the vehicle-local text processing model and the server-deployed text processing model are used as the processing model corresponding to the instruction text; the number of parameters of the server-deployed text processing model is greater than the number of parameters of the vehicle-local text processing model.

[0015] In response to the fusion evaluation result satisfying the third evaluation condition, the text processing model deployed on the server is used as the processing model corresponding to the instruction text.

[0016] In this embodiment, at least the following beneficial effects are achieved: the destination processing end of the instruction text is determined based on the fusion evaluation results, realizing the routing and distribution decision of the instruction text; by utilizing the complexity and uncertainty of the instruction text, it fully considers whether the text processing model deployed locally on the vehicle is sufficient to process the instruction text, and whether it is necessary to use the text processing model deployed on the server to process the instruction text together; by combining the processing results of the two text processing models, it helps to improve the accuracy and reliability of the instruction text processing results, thereby improving the processing accuracy of the vehicle's interactive instructions.

[0017] In an optional embodiment of the first aspect, the uncertainty evaluation of the instruction text to obtain a first evaluation result includes:

[0018] The instruction text is segmented based on the vehicle's local word segmentation processing model to obtain word units in the instruction text, and the uncertainty of the word units is evaluated to obtain the first evaluation result.

[0019] The uncertainty assessment of the lexical unit to obtain the first assessment result includes:

[0020] Obtain the probability of the word appearing in the instruction text as predicted by the local text processing model of the vehicle;

[0021] The Shannon entropy is obtained based on the probability of multiple consecutive terms appearing in the instruction text, and the Shannon entropy is determined as the first evaluation result.

[0022] In this embodiment, at least the following beneficial effects are achieved: the instruction text is segmented by the vehicle's local word segmentation processing model to obtain the word units in the instruction text, and then the Shannon entropy is calculated by combining the probability of the occurrence of multiple consecutive word units in the instruction text predicted by the text processing model. The Shannon entropy is used as the first evaluation result to characterize the degree of semantic ambiguity, providing a reliable processing basis for subsequent steps to determine the processing model corresponding to the instruction text.

[0023] In an optional embodiment of the first aspect, the complexity of the instruction text is evaluated to obtain a second evaluation result, including:

[0024] Based on the vehicle's local feature extraction model, feature extraction is performed on the instruction text to obtain feature data of the instruction text, and the complexity of the feature data is evaluated to obtain the second evaluation result; the feature data includes length features, time features, fuzzy word features, modal word features, and question features;

[0025] The step of evaluating the complexity of the feature data to obtain the second evaluation result includes:

[0026] Obtain the length evaluation result of the length feature, the time evaluation result of the time feature, the fuzzy word evaluation result of the fuzzy word feature, the modal word evaluation result of the modal word feature, and the question evaluation result of the question feature;

[0027] The second evaluation result is obtained based on at least one of the length evaluation result, the time evaluation result, the fuzzy word evaluation result, the modal word evaluation result, and the question evaluation result, as well as the corresponding feature weights.

[0028] In this embodiment, at least the following beneficial effects are achieved: multi-dimensional feature data including length, time, ambiguous words, modal words, and questions are extracted from the instruction text by the vehicle's local feature extraction model, and then a second evaluation result representing the complexity of the text is calculated based on at least one feature data and the feature weight corresponding to the feature data, providing a reliable processing basis for subsequent steps to determine the processing model corresponding to the instruction text.

[0029] In an optional embodiment of the first aspect, after obtaining the instruction text corresponding to the vehicle's interaction instruction, the method further includes:

[0030] The image to be processed acquired by the vehicle is subjected to action recognition to obtain the action recognition result of the image to be processed;

[0031] The instruction text is subjected to keyword recognition to obtain the keyword recognition result of the instruction text;

[0032] In response to at least one of the action recognition result matching the target action and the keyword recognition result matching the target keyword, the image to be processed is image-recognized based on the vehicle's local first image processing model to obtain a first image recognition result corresponding to the image to be processed; the first image recognition result is used to match with a preset image recognition result to determine whether to send the first image recognition result to the processing model corresponding to the instruction text for processing.

[0033] In this embodiment, at least the following beneficial effects are achieved: by performing action recognition processing on the image to be processed, keyword recognition processing on the instruction text, or context analysis processing on the instruction text, the action recognition results, keyword recognition results, and / or context information can be used to determine whether the image to be processed contains information that can reflect the user's intent (such as target keywords, target actions, target information). Then, if the image to be processed contains information that can reflect the user's intent, the image to be processed is image recognized based on the vehicle's local first image processing model to obtain the first image recognition result corresponding to the image to be processed. This allows the processing model to perform multimodal joint processing on the first image recognition result and the instruction text, thereby improving the accuracy of the processed result.

[0034] In an optional embodiment of the first aspect, after obtaining the first image recognition result corresponding to the image to be processed, the method further includes:

[0035] In response to the first image recognition result matching a preset image recognition result, the first image recognition result and the instruction text are processed based on the processing model corresponding to the instruction text to obtain the processing result corresponding to the interaction instruction;

[0036] In response to the mismatch between the first image recognition result and the preset image recognition result, the image to be processed is performed on the image based on the second image processing model deployed on the server to obtain the second image recognition result corresponding to the image to be processed.

[0037] Furthermore, the second image recognition result and the instruction text are processed based on the processing model corresponding to the instruction text to obtain the processing result corresponding to the interaction instruction.

[0038] In this embodiment, at least the following beneficial effects are achieved: When the first image recognition result matches the preset image recognition result, the local image recognition result and the instruction text are jointly parsed using the multimodal processing model corresponding to the instruction text. This eliminates the need to upload the image to be processed to the server, significantly reducing the amount of image data transmission, lowering interaction latency, and reducing network resource consumption. When the first image recognition result does not match the preset image recognition result, the second image processing model deployed on the server is called to re-complete the high-precision image recognition and output the second image recognition result. Then, the processing model matching the instruction text is used to jointly process and analyze the second image recognition result and the instruction text, outputting the processed result of the instruction text. This fully leverages the advantage of the fast recognition response speed of the first image processing model to cope with conventional image interaction scenarios, while also using the high-performance second image processing model on the server to compensate for the insufficient recognition capability of the first image processing model. This achieves collaborative parsing of images and text, effectively improving the recognition accuracy of interactive instructions in complex scenarios.

[0039] In an optional embodiment of the first aspect, after processing the instruction text based on the processing model corresponding to the instruction text to obtain the processing result of the instruction text, the method further includes:

[0040] Obtain feedback information returned regarding the processing result;

[0041] Based on the feedback information, update the preset intent text and / or the feature weights associated with the complexity assessment.

[0042] In this embodiment, at least the following beneficial effects are achieved: after obtaining the processing result corresponding to the interactive command, the preset feature weights associated with the intent text and complexity assessment can be continuously optimized based on the actual feedback information of the processing result, or even the weights associated with the fusion assessment result can be optimized, thereby continuously improving the accuracy of determining the processing model corresponding to the command text, that is, improving the accuracy of routing and distribution, and thus improving the processing accuracy of the command text.

[0043] Secondly, this application also provides an instruction processing apparatus, comprising:

[0044] The text acquisition module is used to acquire the instruction text corresponding to the vehicle's interaction commands;

[0045] The text matching module is used to respond to the instruction text matching with the intent text containing a preset value, and to use the local text processing model of the vehicle as the processing model corresponding to the instruction text.

[0046] The routing and scheduling module is used to respond to the mismatch between the instruction text and the preset intent text by performing uncertainty evaluation on the instruction text to obtain a first evaluation result, and performing complexity evaluation on the instruction text to obtain a second evaluation result, and determining the processing model corresponding to the instruction text based on the first evaluation result and the second evaluation result.

[0047] The instruction processing module is used to process the instruction text based on the processing model corresponding to the instruction text to obtain the processing result of the instruction text; the processing includes one or more of intent recognition, response text generation, and instruction splitting; wherein, the response text generation includes performing intent recognition on the instruction text and generating response text based on the intent recognition result; the instruction splitting includes performing intent recognition on the instruction text, and if multiple intent recognition results are obtained, then splitting the instruction text into multiple instructions, with one instruction corresponding to one intent recognition result.

[0048] Thirdly, this application also provides a new energy vehicle, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods described above.

[0049] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0050] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in any of the above aspects.

[0051] Regarding the beneficial effects of any of the technical solutions in the second to fifth aspects mentioned above, refer to the beneficial effects of the corresponding technical solutions in the first aspect; repeated examples will not be listed here. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 This is a schematic diagram of an optional application environment for the instruction processing method in one embodiment;

[0054] Figure 2 This is a schematic diagram of an optional instruction processing method in one embodiment;

[0055] Figure 3 This is an optional flowchart illustrating the steps of determining at least one destination processing end of instruction text in one embodiment;

[0056] Figure 4 This is a schematic diagram of an optional instruction processing method in another embodiment;

[0057] Figure 5 This is a schematic diagram of an optional instruction processing method in yet another embodiment;

[0058] Figure 6 This is a schematic diagram of an optional structure of the instruction processing device in one embodiment;

[0059] Figure 7 This is a schematic diagram of an optional internal structure of a new energy vehicle in one embodiment. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application.

[0061] The terms "first," "second," etc., used in this application may be used to describe various elements, but these elements are not limited by these terms. These terms are used only to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0062] The instruction processing method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, vehicle terminal 101 communicates with server terminal 102 via a network. A data storage system stores the data that server terminal 102 needs to process. This data storage system can be integrated into vehicle terminal 101 or located in the cloud or on other network servers. Vehicle terminal 101 obtains the instruction text corresponding to the vehicle's interaction commands; in response to a match between the instruction text and a preset intent text, it uses the vehicle's local text processing model as the processing model corresponding to the instruction text; in response to a mismatch between the instruction text and the preset intent text, it performs an uncertainty assessment on the instruction text to obtain a first assessment result, and a complexity assessment on the instruction text to obtain a second assessment result, and determines the processing model corresponding to the instruction text based on the first and second assessment results; and processes the instruction text based on the processing model corresponding to the instruction text. Vehicle terminal 101 can be a system installed on a vehicle (such as a new energy vehicle), used to provide various functions and services to the vehicle. Server terminal 102 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0063] In one exemplary embodiment, such as Figure 2 As shown, an instruction processing method is provided, which is applied to... Figure 1 The following explanation will be based on vehicle-side unit 101. Specifically:

[0064] Step S201: Obtain the instruction text corresponding to the vehicle's interaction instructions.

[0065] In this context, the instruction text refers to the text data describing vehicle interaction commands. For example, the instruction text could be "Open the driver's side window." Interaction commands can be instructions issued by the driver or passengers to the vehicle.

[0066] For example, the driver and passengers in the vehicle interact with the vehicle, the user issues commands to the vehicle, and the vehicle parses the commands to obtain the corresponding command text.

[0067] For example, users can issue interactive commands to the vehicle via voice. The vehicle's onboard system uses a locally deployed speech recognition model to process the commands and obtain the corresponding text. The speech recognition model is an artificial intelligence model that converts speech signals (such as interactive commands) into text data; it can be built based on Automatic Speech Recognition (ASR) technology. Users can also input interactive commands directly into the vehicle, which will then provide the corresponding text.

[0068] Step S202: In response to the matching of the instruction text with the preset intent text, the local text processing model of the vehicle is used as the processing model corresponding to the instruction text.

[0069] Intent text refers to keywords that represent a user's intent. Simply put, intent text describes what the user wants to do or what their purpose is. For example, intent text could be text like "open the car window" or "turn off the air conditioning."

[0070] For example, the vehicle-side deployment includes an intent word library, which stores multiple high-frequency, simple intent texts. This intent text library can also be seen as a whitelist of command texts. The vehicle-side performs hierarchical matching processing on all intent texts in the intent word library against the command texts. The first priority is exact matching, which detects whether the command text contains words that are exactly the same as the intent text (e.g., "open the window," "turn off the air conditioner," etc.). The second priority is fuzzy prefix matching, which detects whether the prefix action words of the command text are the same as the prefix action words of the intent text (e.g., "open," "close," and "adjust," etc.). If the command text contains a target word that is exactly the same as the intent text, and / or if the prefix action word of the command text is the same as the prefix action word of the intent text, then the command text is determined to contain the preset intent text. This command text can be considered a high-frequency, simple command text (or, in other words, a command text with a clear intent). The vehicle then uses its local text processing model as the processing model for the command text, routing it to that model for processing. The vehicle then processes the command text to obtain the processing result. Based on the processing result, the vehicle can also execute control actions and / or display the processing result.

[0071] Among them, the text processing model refers to the language processing model (LLM) used to understand, generate and process instruction text.

[0072] Step S203: In response to the mismatch between the instruction text and the preset intent text, perform uncertainty evaluation on the instruction text to obtain a first evaluation result, and perform complexity evaluation on the instruction text to obtain a second evaluation result, and determine the processing model corresponding to the instruction text based on the first evaluation result and the second evaluation result.

[0073] The first evaluation result can be an index describing the uncertainty of the instruction text. Uncertainty is used to characterize the probability that the text processing model deployed locally on the vehicle can correctly understand the user's intent. The higher the uncertainty, the lower the probability that the text processing model deployed locally on the vehicle can correctly understand the user's intent; the lower the uncertainty, the higher the probability that the text processing model deployed locally on the vehicle can correctly understand the user's intent.

[0074] The second evaluation result can be an indicator describing the complexity of the instruction text. Complexity is used to characterize the complexity of the user intent corresponding to the interactive instruction.

[0075] For example, if it is detected that the instruction text does not contain the same target word as the intent text, and it is detected that the prefix action word of the instruction text is different from the prefix action word of the intent text, that is, it is determined that the instruction text does not contain the preset intent text, and it can be considered that the instruction text does not belong to the high-frequency and simple instruction. Then, the complexity and uncertainty of the instruction text can be analyzed to obtain the first evaluation result and the second evaluation result respectively. Based on the first evaluation result and the second evaluation result, the processing model corresponding to the instruction text can be determined, that is, the routing and distribution strategy of the instruction text can be determined. Finally, at least one destination processing end (such as vehicle end and / or server end) of the instruction text can be determined, and the text processing model deployed locally on the destination processing end can be used as the processing model corresponding to the instruction text.

[0076] Step S204: Process the instruction text based on the processing model corresponding to the instruction text to obtain the processing result of the instruction text; the processing includes one or more of intent recognition, response text generation and instruction splitting; wherein, response text generation includes performing intent recognition on the instruction text and generating response text based on the intent recognition result; instruction splitting includes performing intent recognition on the instruction text, and if multiple intent recognition results are obtained, the instruction text is split into multiple instructions, with one instruction corresponding to one intent recognition result.

[0077] For example, the vehicle-mounted terminal routes the command text to one or more locally deployed processing models for processing. The command text is processed by one or more processing models to obtain the processing result. The vehicle-mounted terminal can then perform control actions on the vehicle based on the processing result and / or display the processing result.

[0078] For example, suppose the processing model corresponding to the instruction text infers that the intention of the instruction text is to open the car window. Then, it generates and sends a window opening instruction to the relevant vehicle components, so as to instruct the relevant vehicle components to open the vehicle window.

[0079] It should be noted that the processing of instruction text in this embodiment and the following text includes, but is not limited to:

[0080] (1) Perform intent recognition on the instruction text to obtain the intent recognition results of the interactive instructions. For example, the intent recognition results of in-vehicle comfort control (such as controlling the in-vehicle temperature, in-vehicle humidity, etc.), vehicle status query (such as vehicle condition query), driving navigation, entertainment audio and video, consultation, etc. are obtained.

[0081] (2) Perform intent recognition and response text generation on the instruction text in sequence: First, perform intent recognition on the instruction text to obtain the intent recognition result of the interactive instruction, and then generate the response text based on the intent recognition result. For example, when a user queries the vehicle status, after recognizing the content to be queried, the corresponding response text is generated.

[0082] (3) Perform intent recognition and instruction splitting on the instruction text in sequence: Perform intent recognition on the instruction text to obtain the intent recognition result of the interactive instruction. If multiple intent recognition results are obtained through intent recognition, the instruction text is split into multiple instructions. Each instruction corresponds to an intent recognition result. Generate an independent execution instruction for each intent recognition result to instruct the vehicle to perform the corresponding control action based on the execution instruction.

[0083] In the above-mentioned instruction processing method, the corresponding processing model is matched differentially to the instruction text of the vehicle's interaction instructions. When the instruction text matches the preset intent text, the vehicle's local text processing model is called to process the instruction text, which improves the processing speed of the instruction text of regular interaction instructions. When the instruction text does not match the preset intent text, one or more suitable processing models are selected by comprehensively evaluating the uncertainty and complexity of the instruction text. This achieves accurate differentiation of unconventional interaction instructions such as fuzzy and complex ones and assigns them appropriate processing models. It effectively balances the response speed of regular instruction text with the processing accuracy of unconventional instruction text, thereby improving both the processing efficiency and accuracy of the vehicle's interaction instructions.

[0084] In one exemplary embodiment, such as Figure 3 As shown, step S203 above, which determines the processing model corresponding to the instruction text based on the first evaluation result and the second evaluation result, includes the following:

[0085] Step S301: The first evaluation result and the second evaluation result are merged to obtain the merged evaluation result.

[0086] The fusion evaluation results are used to determine the routing and distribution object of the instruction text (i.e., the target processing end), thereby obtaining the processing model corresponding to the instruction text.

[0087] For example, the vehicle-side terminal can combine the first evaluation result and the second evaluation result of the instruction text to determine the fusion evaluation result of the instruction text. For instance, the vehicle-side terminal can perform a weighted summation of the first evaluation result and the second evaluation result based on the first weight (e.g., 0.5) corresponding to the first evaluation result and the second weight (e.g., 0.5) corresponding to the second evaluation result of the instruction text to obtain the fusion evaluation result of the instruction text. For example, the fusion evaluation result = 0.5 * first evaluation result + 0.5 * second evaluation result.

[0088] Step S302: In response to the fusion evaluation result satisfying the first evaluation condition, the vehicle's local text processing model is used as the processing model corresponding to the instruction text.

[0089] The first evaluation condition refers to the judgment condition used to determine the fusion evaluation result. For example, the first evaluation condition could be whether the fusion evaluation result meets the first evaluation threshold range.

[0090] For example, if the fusion evaluation result is detected to meet the first evaluation threshold range (e.g., the fusion evaluation result is within the range of [0, 0.3)), it means that the model deployed locally on the vehicle is sufficient to process the instruction text. Then the vehicle can be identified as the destination processing end for the instruction text, and the vehicle will route the instruction text to the locally deployed text processing model for processing.

[0091] In step S303, in response to the fusion evaluation result satisfying the second evaluation condition, both the local text processing model of the vehicle and the text processing model deployed on the server are used as the processing models corresponding to the instruction text; the number of parameters of the text processing model deployed on the server is greater than the number of parameters of the local text processing model of the vehicle.

[0092] Among them, the vehicle-local text processing model refers to a language processing model with a small number of parameters. For example, the number of parameters in a vehicle-local text processing model can be 2B (2 Billion), and the response latency is less than 100ms, making it suitable for handling simple vehicle interaction commands (such as "open the window" or "turn down the air conditioning") and other deterministic tasks.

[0093] In this context, a server-deployed text processing model refers to a language processing model with a large number of parameters. Server-deployed text processing models possess stronger semantic understanding, reasoning and planning, knowledge-based question answering, and text generation capabilities than local text processing models in vehicles. For example, a server-deployed text processing model can have 400 bytes of parameters, making it suitable for handling complex tasks such as trip planning and multi-turn dialogues.

[0094] The second evaluation condition refers to the criteria used to determine the fusion evaluation result. For example, the second evaluation condition could be whether the fusion evaluation result meets the second evaluation threshold range.

[0095] For example, if the fusion evaluation result is detected to meet the second evaluation threshold range (e.g., the fusion evaluation result is within the range of [0.3, 0.7)), it indicates that the text processing model deployed locally on the vehicle may not be sufficient to process the instruction text. It is best to also refer to the processing result of the text processing model deployed on the server. In this case, both the text processing model on the vehicle and the text processing model deployed on the server can be determined as the destination processing end for the instruction text. The vehicle then routes the instruction text to the locally deployed text processing model for processing to obtain the vehicle-side processing result of the instruction text, and routes the instruction text to the text processing model deployed on the server for processing to obtain the server-side processing result of the instruction text. The vehicle receives the server-side processing result returned by the server. Since the text processing model deployed on the server has a larger number of parameters and stronger semantic understanding, reasoning planning, and knowledge question answering capabilities than the text processing model on the vehicle's local side, the text processing model that the server is not familiar with is more suitable for handling complex tasks such as trip planning and multi-turn dialogue. Therefore, the vehicle can prioritize executing control actions on the vehicle based on the processing result returned by the server and / or displaying the server-side processing result.

[0096] In practical applications, the vehicle-side processing results can also be used as backup results. For example, if a user reports that the server-side processing result is inaccurate, the vehicle-side processing result can be used to replace the server-side processing result, and the user can be prompted to upgrade their experience. Furthermore, the vehicle-side processing results can be used as enhancement results. For instance, based on the vehicle-side processing results, the server-side processing results can be optimized to make them more accurate and comprehensive.

[0097] By deploying text processing models of varying complexity (e.g., parameter scale) on both the vehicle and server sides to process the command text, we can obtain processing results from both the vehicle and server sides. This achieves distributed collaborative parsing of command text between the vehicle and cloud, improving both the accuracy and efficiency of the processing results. The vehicle and server sides process simultaneously, eliminating the need for separate model inference times as in sequential processing. The local processing model on the vehicle side can quickly complete lightweight local command processing, reducing response latency and ensuring real-time basic interactions. The processing model deployed on the server side leverages stronger computing power to achieve deep semantic refinement, compensating for the limitations of vehicle-side computing power and model capabilities. Simultaneously, the separate output of processing results from both ends improves the accuracy of vehicle human-machine interaction command parsing, balancing the dual requirements of local interaction immediacy and high-precision server-side processing.

[0098] Step S304: In response to the fusion evaluation result meeting the third evaluation condition, the text processing model deployed on the server is used as the processing model corresponding to the instruction text.

[0099] The third evaluation condition refers to the criteria used to determine the fusion evaluation result. For example, the third scoring condition could be to detect whether the fusion evaluation result meets the third evaluation threshold range.

[0100] For example, if the fusion evaluation result is detected to meet the third evaluation threshold range (e.g., the fusion evaluation result is within the range of [0.7, 1.0] points), it indicates that the text processing model deployed locally on the vehicle is insufficient to process the instruction text, and it is recommended to use the text processing model deployed on the server to process the instruction text. In this case, the server can be identified as the destination processing end for the instruction text, and the vehicle will then route the instruction text to the text processing model deployed on the server for processing.

[0101] In this embodiment, the destination processing end of the instruction text is determined based on the fusion evaluation results, and the routing and distribution decision of the instruction text is realized. By taking into full account the complexity and uncertainty of the instruction text, it is considered whether the text processing model deployed locally on the vehicle is sufficient to process the instruction text, and whether the text processing model deployed on the server should be used to process the instruction text together. Combining the processing results of the two text processing models helps to improve the accuracy and reliability of the instruction text processing results, thereby improving the processing accuracy of the vehicle's interactive instructions.

[0102] In an exemplary embodiment, step S203 above, which involves evaluating the uncertainty of the instruction text to obtain a first evaluation result, includes: segmenting the instruction text into words based on a vehicle-local word segmentation processing model to obtain word units in the instruction text, and evaluating the uncertainty of the word units to obtain the first evaluation result. Specifically, evaluating the uncertainty of the word units to obtain the first evaluation result includes: obtaining the probability of word units appearing in the instruction text predicted by the vehicle-local text processing model; obtaining Shannon entropy based on the probability of multiple consecutive word units appearing in the instruction text, and determining the Shannon entropy as the first evaluation result.

[0103] For example, the vehicle inputs the instruction text into a locally deployed word segmentation model. The model segments the instruction text into its smallest processable semantic units, outputting tokens. The vehicle's local text processing model predicts the probability of each token appearing in the instruction text within its context. The vehicle calculates the Shannon entropy based on the probabilities of consecutive tokens appearing in the instruction text, as predicted by the local model. This Shannon entropy characterizes the uncertainty of the instruction text and is used as the first evaluation result. In practical applications, the formula for calculating Shannon entropy can be expressed as:

[0104] H=-ΣP(token_i)×logP(token_i)

[0105] In the formula, H represents Shannon entropy; token_i represents the i-th token in the instruction text; and P(token_i) represents the probability of the i-th token appearing in the instruction text as predicted by the vehicle's local text processing model. The larger H is, the higher the uncertainty; the smaller H is, the lower the uncertainty.

[0106] In this embodiment, the instruction text is segmented using a vehicle-local word segmentation processing model to obtain word units in the instruction text. Then, the Shannon entropy is calculated by combining the probability of multiple consecutive word units appearing in the instruction text predicted by the text processing model. The Shannon entropy is used as the first evaluation result to characterize the degree of semantic ambiguity, providing a reliable processing basis for subsequent steps to determine the processing model corresponding to the instruction text.

[0107] In an exemplary embodiment, step S203 above, which evaluates the complexity of the instruction text to obtain a second evaluation result, includes the following: extracting features from the instruction text based on a vehicle-local feature extraction model to obtain feature data of the instruction text, and evaluating the complexity of the feature data to obtain a second evaluation result; the feature data includes length features, time features, fuzzy word features, modal word features, and question features. Specifically, evaluating the complexity of the feature data to obtain the second evaluation result includes the following: obtaining the length evaluation result of the length feature, obtaining the time evaluation result of the time feature, obtaining the fuzzy word evaluation result of the fuzzy word feature, obtaining the modal word evaluation result of the modal word feature, and obtaining the question evaluation result of the question feature; and obtaining the second evaluation result based on at least one of the length evaluation result, time evaluation result, fuzzy word evaluation result, modal word evaluation result, and question evaluation result, and the corresponding feature weights.

[0108] Feature extraction models refer to models that extract multi-dimensional feature data from input text data (such as instruction text).

[0109] For example, the vehicle inputs the instruction text into a locally deployed feature extraction model. The model then performs feature extraction on the instruction text, obtaining the following data: length (reflecting sentence length), time features (reflecting whether the instruction text contains time words, such as "x o'clock," "xx minutes," "morning," "afternoon," etc.), fuzzy word features (reflecting whether the instruction text contains fuzzy pronouns, such as "this," "that," etc.), modal particle features (reflecting modal particle density, such as "ah," "oh," "um," etc.), and question features (reflecting whether the instruction text is a question type, such as "ma?" "how?" etc.). The vehicle evaluates the length score corresponding to the length data to obtain the length evaluation result of the length feature; evaluates the time score corresponding to the time data to obtain the time evaluation result of the time feature; evaluates the fuzzy word score corresponding to the fuzzy word feature to obtain the fuzzy word evaluation result of the fuzzy word feature; evaluates the modal particle score corresponding to the modal particle feature to obtain the modal particle evaluation result of the modal particle feature; and evaluates the question score corresponding to the question feature to obtain the question evaluation result of the question feature.

[0110] The vehicle-side terminal obtains a second evaluation result based on at least one of the following: length evaluation result, time evaluation result, fuzzy word evaluation result, modal word evaluation result, and question evaluation result, along with their corresponding feature weights. For example, the vehicle-side terminal can perform a weighted summation of the length evaluation result, time evaluation result, fuzzy word evaluation result, modal word evaluation result, and question evaluation result based on the feature weights corresponding to the length feature, time feature, fuzzy word feature, modal word feature, and question feature, to obtain the complexity of the instruction text, and use this complexity as the second evaluation result. For example, complexity (i.e., the second evaluation result) = feature weight corresponding to the length feature (e.g., 0.4) * length evaluation result + feature weight corresponding to the time feature (e.g., 0.15) * time evaluation result + feature weight corresponding to the fuzzy word feature (e.g., 0.15) * fuzzy word evaluation result + feature weight corresponding to the modal word feature (e.g., 0.15) * modal word evaluation result + feature weight corresponding to the question feature (e.g., 0.15) * question evaluation result.

[0111] In this embodiment, a multi-dimensional feature data, including length, time, ambiguous words, modal particles, and questions, is extracted from the instruction text using a vehicle-local feature extraction model. Then, a second evaluation result representing the complexity of the text is calculated based on at least one feature data and the feature weight corresponding to the feature data, providing a reliable processing basis for subsequent steps to determine the processing model corresponding to the instruction text.

[0112] In an exemplary embodiment, after obtaining the instruction text corresponding to the vehicle's interaction instruction in step S201, the method further includes: performing action recognition processing on the image to be processed acquired by the vehicle to obtain the action recognition result of the image to be processed; performing keyword processing on the instruction text to obtain the keyword recognition result of the instruction text; responding to at least one of the action recognition result matching the target action and the keyword recognition result matching the target keyword, performing image recognition on the image to be processed based on the vehicle's local first image processing model to obtain the first image recognition result corresponding to the image to be processed; the first image recognition result is used to match with a preset image recognition result to determine whether to send the first image recognition result to the processing model corresponding to the instruction text for processing.

[0113] The images to be processed include interior and exterior images of the vehicle.

[0114] The first image processing model refers to an artificial intelligence model used to identify the image intent represented by the image to be processed.

[0115] For example, in the above steps, in addition to sending the instruction text to the server and / or the vehicle for processing, the in-vehicle and out-of-vehicle images can also be captured periodically or in real time by the DMS (Driver Monitoring System) camera installed in the vehicle and set as images to be processed.

[0116] The vehicle-side performs keyword (visual pronoun) recognition processing on the instruction text to obtain the keyword recognition result of the instruction text, so as to characterize whether the instruction text contains target keywords; among them, target keywords include visual pronouns (such as "this", "that", "look", and "what is it"); for example, the instruction text is "Help me open this".

[0117] The vehicle performs action recognition processing on the image to be processed to obtain the action recognition result of the image to be processed, so as to characterize whether the image to be processed contains a target action (such as a gesture action, a gaze action, etc.); for example, it is recognized that the user in the image to be processed has a gesture action (such as pointing to the window) or a gaze action (such as looking at the window).

[0118] The vehicle can also perform contextual analysis on the command text to obtain contextual information of the command text, so as to characterize whether the command text contains target information; among which, target information includes item information and / or location information; for example, the user mentions descriptions of items or locations in the context, such as "what building I just passed by outside" or "where did the key I was holding go when I got in the car?"

[0119] In response to at least one of the following conditions being met: the keyword recognition result of the instruction text matches the target keyword, the action recognition result of the image to be processed matches the target action, and the context information of the instruction text contains the target information, the vehicle-side performs image recognition on the image to be processed based on the vehicle's local first image processing model to obtain the first image recognition result corresponding to the image to be processed; the first image recognition result is matched with a preset image recognition result to determine whether to send the first image recognition result to the processing model corresponding to the instruction text for processing.

[0120] In this embodiment, by performing action recognition processing on the image to be processed, keyword recognition processing on the instruction text, or context analysis processing on the instruction text, the action recognition results, keyword recognition results, and / or context information can be used to determine whether the image to be processed contains information that can reflect the user's intent (such as target keywords, target actions, target information). Then, if the image to be processed contains information that can reflect the user's intent, the image to be processed is image recognized based on the vehicle's local first image processing model to obtain the first image recognition result corresponding to the image to be processed. This allows the processing model to perform multimodal joint processing on the first image recognition result and the instruction text, thereby improving the accuracy of the processed result.

[0121] In an exemplary embodiment, after obtaining the first image recognition result corresponding to the image to be processed, the method further includes: in response to the first image recognition result matching a preset image recognition result, processing the first image recognition result and the instruction text based on the processing model corresponding to the instruction text to obtain the processing result corresponding to the interaction instruction; in response to the first image recognition result not matching the preset image recognition result, performing image recognition on the image to be processed based on a second image processing model deployed on the server to obtain the second image recognition result corresponding to the image to be processed; and processing the second image recognition result and the instruction text based on the processing model corresponding to the instruction text to obtain the processing result corresponding to the interaction instruction.

[0122] The second image processing model refers to an artificial intelligence model used to identify the image intent represented by the image to be processed.

[0123] The first and second image processing models can be Vision Language Models (VLMs). A Vision Language Model refers to a multimodal generative artificial intelligence system that combines a large language model with a visual encoder to achieve comprehensive processing of images, text, and video. It should be noted that the first image processing model has a larger number of parameters than the second image processing model. The first image processing model can handle most scene content, but it cannot handle intents requiring online searching. In such cases, the image (e.g., the image to be processed) needs to be uploaded to the second image processing model deployed on the server for processing.

[0124] The image recognition result refers to the image intent information represented by the recognized image to be processed. For example, the image intent information could be that the user is pointing at the driver's side window.

[0125] The preset image recognition result refers to the target image intent information that can be processed by the first image processing model deployed locally in advance.

[0126] For example, in response to the matching of the image intent information in the first image recognition result with the target image intent information in the preset image recognition result, indicating that the first image processing model deployed locally on the vehicle is sufficient to handle the intent, the first image recognition result and the instruction text are subjected to multimodal joint processing based on the processing model corresponding to the instruction text to obtain the corresponding processing result.

[0127] In response to a mismatch between the first image recognition result and the preset image recognition result on the vehicle, indicating that the first image processing model deployed locally on the vehicle cannot process the intent (e.g., the intent could be to search for the address corresponding to a place name online), the image to be processed is encrypted to obtain an encrypted image. This encrypted image is then routed to a second image processing model deployed on the server for image recognition processing, yielding a second image recognition result corresponding to the image to be processed. The vehicle then receives the second image recognition result returned by the server; based on the processing model corresponding to the instruction text (including the text processing model locally on the vehicle and / or the text processing model deployed on the server), it performs multimodal joint processing on the second image recognition result and the instruction text to obtain at least one processing result.

[0128] Understandably, the processing model corresponding to the instruction text can use the first or second image recognition result as the context of the instruction text, enabling the processing model to perform multimodal joint processing by combining the first or second image recognition result with the instruction text. For example, suppose the instruction text is "Help me open this." If the language model only analyzes the instruction text, it cannot accurately determine what device "this" specifically refers to. Suppose the vehicle also receives the first or second image recognition result as "The user is pointing at the driver's side window." Then, the language model combines the instruction text and the first or second image recognition result to analyze and obtain "The user is pointing at the driver's side window + help me open this." It is easy to know that the user's true intention is to open the driver's side window, and then the driver's side window is controlled to open.

[0129] In practical applications, a standardized JSON (JavaScript Object Notation) format can be used to route the first or second image recognition result and the instruction text to the processing model corresponding to the instruction text. For example, the routing information includes, but is not limited to: timestamp (e.g., "2026-03-24T10:30:00.000Z"), instruction type, confidence level (e.g., 0.0~1.0), the first or second image recognition result, the image to be processed and the coordinates of the target region in the image to be processed, the instruction text, whether server-side deep processing has been performed, and the urgency level, etc.

[0130] In this embodiment, when the first image recognition result matches the preset image recognition result, the local image recognition result and the instruction text are jointly parsed using a multimodal processing model corresponding to the instruction text. This eliminates the need to upload the image to be processed to the server, significantly reducing the amount of image data transmission, interaction latency, and network resource consumption. When the first image recognition result does not match the preset image recognition result, the second image processing model deployed on the server is invoked to re-perform high-precision image recognition, outputting the second image recognition result. Then, the instruction text matching processing model is used to jointly process and analyze the second image recognition result and the instruction text, outputting the processed result of the instruction text. This approach fully leverages the fast recognition response speed of the first image processing model to handle conventional image interaction scenarios, while also using the high-performance second image processing model on the server to compensate for the insufficient recognition capability of the first image processing model. This achieves collaborative parsing of images and text, effectively improving the recognition accuracy of interactive instructions in complex scenarios.

[0131] In an exemplary embodiment, after processing the instruction text based on the processing model corresponding to the instruction text in step S204, the method further includes: obtaining feedback information returned for the processing result; and updating the preset intent text and / or feature weights associated with complexity assessment based on the feedback information.

[0132] Feedback information refers to the information provided by users regarding the processing results. Feedback information includes positive feedback (such as the processing result being accurate) and negative feedback (such as the processing result being inaccurate).

[0133] For example, the vehicle-side obtains the processing result corresponding to the interactive command returned by the target processing end (including the vehicle-side and / or the server). Feedback options for the processing result (such as "accurate" and "inaccurate" options) can also be displayed on the vehicle's in-vehicle display. The vehicle-side obtains the user's selected option and receives feedback information. If the feedback information is positive, the intent text in the intent word library is added, such as adding interjections and ambiguous words. The weights in the calculation formula for the fusion evaluation result and / or the feature weights in the calculation formula for the complexity evaluation can also be adjusted. If the feedback information is negative, the intent text in the intent word library is modified or deleted. The weights in the calculation formula for the fusion evaluation result and / or the feature weights in the calculation formula for the complexity evaluation can also be modified.

[0134] In practical applications, feedback information can be periodically collected and updated based on the feedback information within a preset period (such as one week, one month, etc.). This update comprehensively updates the preset intention text and / or complexity assessment associated feature weights and / or fusion assessment results associated with the intention word library.

[0135] In this embodiment, after obtaining the processing result corresponding to the interaction command, the preset feature weights associated with the intent text and complexity assessment can be continuously optimized based on the actual feedback information of the processing result, and even the weights associated with the fusion assessment result can be optimized, thereby continuously improving the accuracy of determining the processing model corresponding to the command text, that is, improving the accuracy of routing and distribution, and thus improving the processing accuracy of the command text.

[0136] In one embodiment, such as Figure 4 As shown, another instruction processing method is provided, which can be applied to... Figure 1 Taking the vehicle side as an example, the explanation includes the following steps:

[0137] Step S401: Obtain the instruction text corresponding to the vehicle's interaction instructions.

[0138] Step S402: In response to the matching of the instruction text with the preset intent text, the local text processing model of the vehicle is used as the processing model corresponding to the instruction text.

[0139] Step S403: In response to the mismatch between the instruction text and the preset intent text, perform uncertainty evaluation on the instruction text to obtain a first evaluation result, and perform complexity evaluation on the instruction text to obtain a second evaluation result. Then, fuse the first evaluation result and the second evaluation result to obtain a fused evaluation result.

[0140] Step S404: In response to the fusion evaluation result satisfying the first evaluation condition, the vehicle's local text processing model is used as the processing model corresponding to the instruction text.

[0141] In step S405, in response to the fusion evaluation result satisfying the second evaluation condition, both the local text processing model of the vehicle and the text processing model deployed on the server are used as the processing models corresponding to the instruction text; the number of parameters of the text processing model deployed on the server is greater than the number of parameters of the local text processing model of the vehicle.

[0142] Step S406: In response to the fusion evaluation result meeting the third evaluation condition, the text processing model deployed on the server is used as the processing model corresponding to the instruction text.

[0143] Step S407: Obtain the processing result corresponding to the interactive command, and obtain the feedback information returned for the processing result.

[0144] Step S408: Based on the feedback information, update the preset intent text and / or feature weights associated with complexity assessment.

[0145] The above-described instruction processing method achieves the following beneficial effects: For the instruction text of vehicle interaction commands, it differentiates and matches the corresponding processing model. When the instruction text matches the preset intent text, it calls the vehicle's local text processing model to process the instruction text, improving the processing speed of conventional interaction commands. When the instruction text does not match the preset intent text, it selects one or more suitable processing models by comprehensively evaluating the uncertainty and complexity of the instruction text. This achieves accurate differentiation of ambiguous, complex, and other unconventional interaction commands and assigns appropriate processing models to them, effectively balancing the response speed of conventional instruction text with the processing accuracy of unconventional instruction text. This improves both the processing efficiency and accuracy of vehicle interaction commands.

[0146] To more clearly illustrate the instruction processing method provided in the embodiments of this disclosure, the following specific embodiment will be used to describe the above instruction processing method in detail. For example... Figure 5 As shown, another instruction processing method is provided, which can be applied to... Figure 1 The vehicle-side component specifically includes the following:

[0147] Users issue interactive commands to the vehicle via voice. The vehicle uses a locally deployed speech recognition model to process the commands and obtain the corresponding command text. Based on the pre-set intent text in the intent word library, the vehicle performs precise matching (and fuzzy prefix matching) on ​​the command text. If the command text contains a target word that is exactly the same as the intent text, and / or if the prefix action word of the command text is the same as the prefix action word of the intent text, the command text is routed to the locally deployed text processing model on the vehicle for processing, resulting in the vehicle's processing result.

[0148] Images inside and outside the vehicle are captured by the vehicle's DMS camera and designated as images to be processed. If the image to be processed and / or the instruction text satisfy any one of the image processing conditions (including keyword recognition, action recognition, and context awareness), the image to be processed is routed to the first image processing model on the vehicle for processing, resulting in a first image recognition result. Based on the first image recognition result, it is determined whether the processing model deployed on the vehicle is sufficient to process the image to be processed. If the first image recognition result matches a preset image recognition result, it indicates that the processing model is sufficient to process the image to be processed. The processing model corresponding to the instruction text then processes the first image recognition result and the instruction text to obtain the processing result corresponding to the interaction instruction. If the first image recognition result does not match the preset image recognition result, the image to be processed is encrypted and transmitted to the server. The second image processing model deployed on the server performs image recognition on the image to be processed, resulting in a second image recognition result corresponding to the image to be processed.

[0149] If the command text is found to contain no target word that is exactly the same as the intent text, and if the prefix action word of the command text is found to be different from the prefix action word of the intent text, then the command text is subjected to uncertainty evaluation to obtain a first evaluation result, and the command text is subjected to complexity evaluation to obtain a second evaluation result. The first evaluation result and the second evaluation result are fused to obtain a fused evaluation result, and routing scheduling is performed based on the fused evaluation result to determine the processing model corresponding to the command text: If the fused evaluation result meets the first evaluation condition, such as the fused evaluation result being within the scoring threshold range of [0, 0.3), then the command text and image recognition results (including the first image recognition result and the second image recognition result) are routed to the vehicle terminal. The text processing model on the local end is used to process the data to obtain the processing result on the vehicle end. If the fusion evaluation result meets the second evaluation condition, such as the fusion evaluation result being within the scoring threshold range of [0.3, 0.7), then the instruction text and the above image recognition result are routed to the text processing model deployed locally on the vehicle end for processing to obtain the processing result on the vehicle end, and the instruction text and the above image recognition result are routed to the text processing model deployed on the server end for processing to obtain the processing result on the server end. If the fusion evaluation result meets the third evaluation condition, such as the fusion evaluation result being within the scoring threshold range of [0.7, 1.0], then the instruction text and the above image recognition result are routed to the text processing model deployed on the server end for processing to obtain the processing result on the server end.

[0150] The vehicle displays the processing results from either the vehicle or the server. Based on user feedback regarding the displayed processing results, the intent text in the intent word library, the weights associated with the fusion evaluation results, and / or the feature weights associated with the complexity evaluation are periodically updated iteratively (equivalent to optimizing the routing and scheduling strategy of the processing model corresponding to the instruction text).

[0151] In this embodiment, the following beneficial effects can be achieved: Firstly, based on the preset intent text, instruction text with clear intent can be detected and routed to the local processing model on the vehicle for processing. For instruction text without intent text, the processing models corresponding to instruction text with different uncertainties and complexities are distinguished based on the fusion evaluation results, thus balancing the processing efficiency and accuracy of vehicle interaction instructions. Secondly, by utilizing the image recognition results of the image to be processed, richer contextual information is provided for subsequent processing of instruction text, further improving the accuracy and reliability of the instruction text processing results. Thirdly, the preset intent text and routing scheduling strategy can be continuously optimized based on actual feedback information, thereby continuously improving the accuracy of routing distribution. This improves both the processing efficiency and accuracy of vehicle interaction instructions.

[0152] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0153] Based on the same inventive concept, this application also provides an instruction processing apparatus for implementing the instruction processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more instruction processing apparatus embodiments provided below can be found in the limitations of the instruction processing method described above, and will not be repeated here.

[0154] In one exemplary embodiment, such as Figure 6 As shown, an instruction processing device 600 is provided, including: a text acquisition module 601, a text matching module 602, a routing scheduling module 603, and an instruction processing module 604, wherein:

[0155] The text acquisition module 601 is used to acquire the instruction text corresponding to the vehicle's interaction commands.

[0156] The text matching module 602 is used to match the instruction text with the preset intent text and use the vehicle's local text processing model as the processing model corresponding to the instruction text.

[0157] The routing and scheduling module 603 is used to respond to a mismatch between the instruction text and the preset intent text by performing an uncertainty assessment on the instruction text to obtain a first assessment result, and performing a complexity assessment on the instruction text to obtain a second assessment result, and determining the processing model corresponding to the instruction text based on the first assessment result and the second assessment result.

[0158] The instruction processing module 604 is used to process the instruction text based on the processing model corresponding to the instruction text to obtain the processing result of the instruction text; the processing includes one or more of intent recognition, response text generation and instruction splitting; wherein, response text generation includes performing intent recognition on the instruction text and generating response text based on the intent recognition result; instruction splitting includes performing intent recognition on the instruction text, and if multiple intent recognition results are obtained, the instruction text is split into multiple instructions, with one instruction corresponding to one intent recognition result.

[0159] In one embodiment, the routing and scheduling module 603 is further configured to fuse the first evaluation result and the second evaluation result to obtain a fused evaluation result; in response to the fused evaluation result satisfying the first evaluation condition, the vehicle's local text processing model is used as the processing model corresponding to the instruction text; in response to the fused evaluation result satisfying the second evaluation condition, both the vehicle's local text processing model and the text processing model deployed on the server are used as the processing models corresponding to the instruction text; the number of parameters of the text processing model deployed on the server is greater than the number of parameters of the vehicle's local text processing model; in response to the fused evaluation result satisfying the third evaluation condition, the text processing model deployed on the server is used as the processing model corresponding to the instruction text.

[0160] In one embodiment, the routing scheduling module 603 is further configured to segment the instruction text based on a vehicle-local word segmentation processing model to obtain word units in the instruction text, and to perform uncertainty evaluation on the word units to obtain a first evaluation result. The instruction processing device 600 also includes a first evaluation model, configured to obtain the probability of word units appearing in the instruction text predicted by the vehicle-local text processing model; to obtain Shannon entropy based on the probability of multiple consecutive word units appearing in the instruction text, and to determine the Shannon entropy as the first evaluation result.

[0161] In one embodiment, the routing scheduling module 603 is further configured to extract features from the instruction text based on a vehicle-local feature extraction model to obtain feature data of the instruction text, and to evaluate the complexity of the feature data to obtain a second evaluation result; the feature data includes length features, time features, fuzzy word features, modal word features, and question features. The instruction processing device 600 further includes a second evaluation model, configured to obtain a length evaluation result for the length feature, a time evaluation result for the time feature, a fuzzy word evaluation result for the fuzzy word feature, a modal word evaluation result for the modal word feature, and a question evaluation result for the question feature; and to obtain the second evaluation result based on at least one of the length evaluation result, time evaluation result, fuzzy word evaluation result, modal word evaluation result, and question evaluation result, and the corresponding feature weights.

[0162] In one embodiment, the instruction processing device 600 further includes an image recognition module, configured to perform action recognition on the image to be processed acquired by the vehicle to obtain an action recognition result of the image to be processed; perform keyword recognition on the instruction text to obtain a keyword recognition result of the instruction text; in response to at least one of the action recognition result matching a target action and the keyword recognition result matching a target keyword, perform image recognition on the image to be processed based on a first image processing model local to the vehicle to obtain a first image recognition result corresponding to the image to be processed; the first image recognition result is used to match with a preset image recognition result to determine whether to send the first image recognition result to the processing model corresponding to the instruction text for processing.

[0163] In one embodiment, the instruction processing device 600 further includes a multimodal processing module, configured to: respond to a first image recognition result matching a preset image recognition result; process the first image recognition result and the instruction text based on a processing model corresponding to the instruction text to obtain a processing result corresponding to the interactive instruction; respond to a first image recognition result not matching a preset image recognition result; perform image recognition on the image to be processed based on a second image processing model deployed on the server to obtain a second image recognition result corresponding to the image to be processed; and process the second image recognition result and the instruction text based on a processing model corresponding to the instruction text to obtain a processing result corresponding to the interactive instruction.

[0164] In one embodiment, the instruction processing device 600 further includes an information update module for obtaining feedback information returned in response to the processing result; and updating the preset intent text and / or complexity assessment based on the feedback information.

[0165] Each module in the aforementioned instruction processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the new energy vehicle in hardware form or independent of it, or stored in the memory of the new energy vehicle in software form, so that the processor can call and execute the operations corresponding to each module.

[0166] In one exemplary embodiment, a new energy vehicle is provided, the internal structure of which can be as follows: Figure 7 As shown, the new energy vehicle includes a processor and a memory. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium storing a computer program. When executed by the processor, the computer program implements an instruction processing method.

[0167] Those skilled in the art will understand that Figure 7The structure shown is a block diagram of a part of the structure related to the solution of this application, and does not constitute a limitation on the new energy vehicle (new energy vehicle) on which the solution of this application is applied. A specific new energy vehicle may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0168] In one exemplary embodiment, a new energy vehicle is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0169] In one exemplary embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method embodiments.

[0170] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0171] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0172] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program mentioned can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0173] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0174] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An instruction processing method, characterized in that, The method includes: Obtain the command text corresponding to the vehicle's interaction commands; In response to the instruction text matching a preset intent text, the vehicle's local text processing model is used as the processing model corresponding to the instruction text. In response to a mismatch between the instruction text and the preset intent text, the instruction text is segmented based on the vehicle's local word segmentation processing model to obtain the word units in the instruction text; The probability of the word appearing in the instruction text predicted by the local text processing model of the vehicle is obtained; the Shannon entropy is obtained based on the probability of multiple consecutive words appearing in the instruction text, and the Shannon entropy is determined as the first evaluation result; The complexity of the instruction text is evaluated to obtain a second evaluation result, and the processing model corresponding to the instruction text is determined based on the first evaluation result and the second evaluation result. The instruction text is processed based on the processing model corresponding to the instruction text to obtain the processing result of the instruction text; the processing includes one or more of intent recognition, response text generation and instruction splitting; wherein, the response text generation includes performing intent recognition on the instruction text and generating response text based on the intent recognition result; the instruction splitting includes performing intent recognition on the instruction text, and if multiple intent recognition results are obtained, the instruction text is split into multiple instructions, with one instruction corresponding to one intent recognition result.

2. The method according to claim 1, characterized in that, The step of determining the processing model corresponding to the instruction text based on the first evaluation result and the second evaluation result includes: The first evaluation result and the second evaluation result are merged to obtain a fused evaluation result; In response to the fusion evaluation result satisfying the first evaluation condition, the local text processing model of the vehicle is used as the processing model corresponding to the instruction text; In response to the fusion evaluation result satisfying the second evaluation condition, both the vehicle-local text processing model and the server-deployed text processing model are used as the processing model corresponding to the instruction text; the number of parameters of the server-deployed text processing model is greater than the number of parameters of the vehicle-local text processing model. In response to the fusion evaluation result satisfying the third evaluation condition, the text processing model deployed on the server is used as the processing model corresponding to the instruction text.

3. The method according to claim 1, characterized in that, The process of performing a complexity evaluation on the instruction text to obtain a second evaluation result includes: Based on the vehicle's local feature extraction model, feature extraction is performed on the instruction text to obtain feature data of the instruction text, and the complexity of the feature data is evaluated to obtain the second evaluation result; the feature data includes length features, time features, fuzzy word features, modal word features, and question features; The step of evaluating the complexity of the feature data to obtain the second evaluation result includes: Obtain the length evaluation result of the length feature, the time evaluation result of the time feature, the fuzzy word evaluation result of the fuzzy word feature, the modal word evaluation result of the modal word feature, and the question evaluation result of the question feature; The second evaluation result is obtained based on at least one of the length evaluation result, the time evaluation result, the fuzzy word evaluation result, the modal word evaluation result, and the question evaluation result, as well as the corresponding feature weights.

4. The method according to claim 1, characterized in that, After obtaining the command text corresponding to the vehicle's interaction commands, the following is also included: The image to be processed acquired by the vehicle is subjected to action recognition to obtain the action recognition result of the image to be processed; The instruction text is subjected to keyword recognition to obtain the keyword recognition result of the instruction text; In response to at least one of the action recognition result matching the target action and the keyword recognition result matching the target keyword, the image to be processed is image-recognized based on the vehicle's local first image processing model to obtain a first image recognition result corresponding to the image to be processed; the first image recognition result is used to match with a preset image recognition result to determine whether to send the first image recognition result to the processing model corresponding to the instruction text for processing.

5. The method according to claim 4, characterized in that, After obtaining the first image recognition result corresponding to the image to be processed, the process further includes: In response to the first image recognition result matching a preset image recognition result, the first image recognition result and the instruction text are processed based on the processing model corresponding to the instruction text to obtain the processing result corresponding to the interaction instruction; In response to the mismatch between the first image recognition result and the preset image recognition result, the image to be processed is performed on the image based on the second image processing model deployed on the server to obtain the second image recognition result corresponding to the image to be processed. Furthermore, the second image recognition result and the instruction text are processed based on the processing model corresponding to the instruction text to obtain the processing result corresponding to the interaction instruction.

6. The method according to any one of claims 1 to 5, characterized in that, After processing the instruction text based on the processing model corresponding to the instruction text to obtain the processing result of the instruction text, the method further includes: Obtain feedback information returned regarding the processing result; Based on the feedback information, update the preset intent text and / or the feature weights associated with the complexity assessment.

7. An instruction processing device, characterized in that, The device includes: The text acquisition module is used to acquire the instruction text corresponding to the vehicle's interaction commands; The text matching module is used to respond to the instruction text matching with the intent text containing a preset value, and to use the local text processing model of the vehicle as the processing model corresponding to the instruction text. A routing and scheduling module is used to respond to a mismatch between the instruction text and the preset intent text by segmenting the instruction text into words based on the vehicle's local word segmentation processing model to obtain the word units in the instruction text; obtaining the probability of the word units appearing in the instruction text predicted by the vehicle's local text processing model; obtaining Shannon entropy based on the probability of multiple consecutive word units appearing in the instruction text, and determining the Shannon entropy as a first evaluation result; performing a complexity evaluation on the instruction text to obtain a second evaluation result, and determining the processing model corresponding to the instruction text based on the first evaluation result and the second evaluation result. The instruction processing module is used to process the instruction text based on the processing model corresponding to the instruction text to obtain the processing result of the instruction text; the processing includes one or more of intent recognition, response text generation, and instruction splitting; wherein, the response text generation includes performing intent recognition on the instruction text and generating response text based on the intent recognition result; the instruction splitting includes performing intent recognition on the instruction text, and if multiple intent recognition results are obtained, then splitting the instruction text into multiple instructions, with one instruction corresponding to one intent recognition result.

8. A vehicle comprising a memory and a processor, said memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intention recognition method, device, apparatus, and storage medium

    CN109543190A

  • Intention identification method and device, computer equipment, and computer-readable storage medium

    CN113672696A