Cross-language driving cooperative control method and device, medium and product

By parsing the intent of users' multilingual commands and combining them with local traffic rules and road sign recognition models, the system generates target driving commands using fleet collaborative decision-making. This solves the problem of adapting intelligent driving systems to multilingual environments and cultural differences in cross-border scenarios, achieving accurate driving decisions and compliance, and improving the safety and efficiency of cross-border driving.

CN121291473APending Publication Date: 2026-01-09CHINA MOBILE FINANCIAL TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511399272.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing intelligent driving systems based on large language models lack the comprehensive adaptability to multilingual environments, cross-cultural differences, and dynamic localization rules in cross-border scenarios. This results in insufficient language and cultural adaptation, weak collaborative decision-making capabilities, and a lack of dynamic rule loading, making it unable to meet the needs of cross-border driving.

Method used

By parsing the intent of multilingual commands input by users, a standardized first driving command is generated. This command is then corrected by combining local traffic rules and a multilingual road sign recognition model. Finally, a target driving command is generated using fleet collaborative decision-making, achieving a precise conversion from general commands to scenario-based driving decisions.

Benefits of technology

It improves the accuracy of intent parsing in cross-language intelligent driving, ensures driving compliance, reduces communication latency, and enhances the safety and efficiency of cross-border driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121291473A_ABST
    Figure CN121291473A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a cross-language driving cooperative control method and device, a medium and a product, and relates to the technical field of intelligent driving. The method comprises the following steps: performing intention analysis on a multi-language instruction input by a user to generate a first driving instruction; loading a local traffic rule based on the position of the vehicle, and correcting the first driving instruction into a second driving instruction by using the local traffic rule and a preset multi-language road sign recognition model; and correcting the second driving instruction into a target driving instruction according to shared information sent by a motorcade to which the vehicle belongs. According to the scheme of the invention, through intention analysis, localization correction and motorcade cooperation, accurate conversion from a general instruction to a scenarized driving decision is realized, and the problem that an existing intelligent driving system based on a large language model lacks comprehensive adaptation capability to a multi-language environment, cross-cultural differences and dynamic localization rules is solved. And the driving requirement in a transnational scene cannot be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent driving technology, specifically to a cross-language driving cooperative control method, device, medium, and product. Background Technology

[0002] In existing technologies, intelligent driving systems based on large language models generate bird's-eye view (BEV) features by fusing LiDAR and a monocular camera, and then use a deep learning object detection model to identify obstacles. User voice commands are parsed into navigation tasks by the large language model, divided into multiple sub-tasks such as localization and path planning. Finally, relevant algorithms are used to plan the path, and obstacle avoidance algorithms are combined to generate an attraction field to dynamically adjust the route. However, existing solutions have the following drawbacks: 1) Insufficient language and cultural adaptation: Current solutions primarily focus on single-language environments, failing to address the conflict between parsing multilingual mixed commands and localization rules in cross-border scenarios. 2) Weak collaborative decision-making capabilities: Vehicle-to-everything (V2X) communication is mostly based on standardized protocols, lacking integration of multilingual translation and intent calibration to achieve cross-cultural collaborative driving. 3) Lack of dynamic rule loading: Existing high-precision map systems do not integrate multilingual road sign recognition with dynamic traffic regulation databases, potentially leading to compliance risks. Summary of the Invention

[0003] At least one embodiment of this application provides a cross-language driving cooperative control method, device, medium, and product to address the problem that existing intelligent driving systems based on large language models lack comprehensive adaptability to multilingual environments, cross-cultural differences, and dynamic localization rules, and thus cannot meet driving needs in cross-border scenarios.

[0004] To solve the above-mentioned technical problems, this application is implemented as follows:

[0005] In a first aspect, embodiments of this application provide a cross-language driving cooperative control method, including:

[0006] The intent of the multilingual commands input by the user is parsed to generate the first driving command;

[0007] Based on the vehicle's location, local traffic rules are loaded, and the first driving instruction is corrected into a second driving instruction using the local traffic rules and a preset multilingual road sign recognition model.

[0008] Based on the shared information sent by the fleet to which the vehicle belongs, the second driving instruction is modified into the target driving instruction.

[0009] Optionally, after modifying the second driving instruction to the target driving instruction, the method further includes:

[0010] Acquire users' native language and cultural characteristics;

[0011] Based on the native language and cultural characteristics, perceptual interaction information corresponding to the target driving command is generated.

[0012] Optionally, the multilingual commands obtained from user input are subjected to intent parsing to generate the first driving command, including:

[0013] Perform local and / or online speech recognition on multilingual commands input by users to obtain corresponding multilingual text information;

[0014] The multilingual text information is input into a preset multi-intent recognition model to generate a first driving instruction corresponding to the multilingual instruction.

[0015] Optionally, the multilingual text information is input into a preset multi-intent recognition model to generate a first driving instruction corresponding to the multilingual instruction, including:

[0016] The multilingual text information is input into a preset multi-intent recognition model, the multilingual text information is segmented into words, the segmented sub-word sequence is obtained, the sub-word sequence is added to a preset position code, and mapped into a first vector;

[0017] Using the first vector and the network architecture layer of the multi-intent recognition model, a probability distribution of each intent corresponding to the multilingual text information is generated;

[0018] Based on the preset reinforcement learning model and the probability distribution of each intention, a corresponding first driving instruction is generated.

[0019] Optionally, based on a preset reinforcement learning model and the probability distribution of each intention, a corresponding first driving instruction is generated, including:

[0020] Obtain the encoding of users' historical behavioral characteristics and real-time environmental characteristics;

[0021] The historical behavioral features and the real-time environmental features are encoded and input into the reinforcement learning model, and the reinforcement learning model outputs safety weights.

[0022] The probability distribution of each intent is adjusted according to the security weights.

[0023] Based on the adjusted intent probability distribution, the corresponding first driving instruction is generated.

[0024] Optionally, based on the vehicle's location, local traffic rules are loaded, and using these local traffic rules and a preset multilingual road sign recognition model, the first driving instruction is modified into a second driving instruction, including:

[0025] Based on the obtained vehicle location, load the preset local traffic rules;

[0026] Obtain the user's historical response information to the corresponding local traffic rules;

[0027] Based on the response history information and Bayesian inference, the rule priority of the local traffic rules is adjusted;

[0028] Obtain the road sign image corresponding to the vehicle location and the multilingual prompt text input by the user;

[0029] The road sign image and the multilingual prompt text are input into a preset multilingual road sign recognition model, and the road sign semantic classification result is output.

[0030] Based on the rule priority and the road sign semantic classification result, the first driving instruction is modified into a second driving instruction.

[0031] Optionally, based on shared information sent by the fleet to which the vehicle belongs, the second driving instruction is modified into the target driving instruction, including:

[0032] Receive a shared message that establishes a fleet communication connection with the vehicle;

[0033] Based on multiple shared messages, identify the fleet decision leader who establishes fleet communication connections with the vehicle;

[0034] Determine the target language required by the fleet;

[0035] Send a second driving instruction, translated into the target language, to the fleet decision-making leader;

[0036] Receive the group consensus instruction sent by the fleet decision leader; the group consensus instruction is the instruction with the highest decision weight determined by the fleet decision leader based on multiple candidate instructions to construct a weighted decision.

[0037] Based on the aforementioned group consensus instruction, the second driving instruction is modified into the target driving instruction.

[0038] Secondly, embodiments of this application provide a cross-language driving cooperative control device, comprising:

[0039] The first generation module is used to perform intent parsing on the multilingual instructions obtained from the user input and generate the first driving instruction;

[0040] The first processing module is used to load local traffic rules based on the vehicle's location, and use the local traffic rules and a preset multilingual road sign recognition model to correct the first driving instruction into a second driving instruction.

[0041] The second processing module is used to modify the second driving instruction into the target driving instruction based on the shared information sent by the fleet to which the vehicle belongs.

[0042] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in any of the first aspects.

[0043] Fourthly, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the steps of the method described in any of the first aspects.

[0044] Compared with existing technologies, the cross-language driving cooperative control method, device, medium, and product provided in this application parsed the mixed language commands input by the user, generated a standardized first driving command, loaded local traffic rules based on the vehicle's location, and corrected the command by combining a multilingual road sign recognition model to generate a corrected second driving command. Through fleet collaboration, the command was further corrected based on shared information sent by the fleet to which the vehicle belongs, thus correcting the second driving command into the target driving command. Through intent parsing, localization correction, and fleet collaboration, the accurate conversion from general commands to scenario-based driving decisions is achieved. This solves the problem that existing intelligent driving systems based on large language models lack comprehensive adaptability to multilingual environments, cross-cultural differences, and dynamic localization rules, and cannot meet the driving needs in cross-border scenarios. Attached Figure Description

[0045] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0046] Figure 1 A flowchart illustrating the cross-language driving cooperative control method provided in this application embodiment;

[0047] Figure 2 This is a schematic diagram of the cross-language driving cooperative control system provided in an embodiment of this application.

[0048] Figure 3 This is a schematic diagram of the cross-language driving cooperative control device provided in the embodiments of this application. Detailed Implementation

[0049] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, without limiting the number of objects; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, "A or B" covers three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0050] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc., in the instruction sent. An indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.

[0051] As described in the background section, existing technologies often focus on a single language environment, failing to resolve conflicts between multilingual mixed command parsing and localization rules in cross-border scenarios. Vehicle-to-everything (V2X) communication is mostly based on standardized protocols, without combining multilingual translation and intent calibration to achieve cross-cultural collaborative driving. Existing high-precision map systems do not integrate multilingual road sign recognition with dynamic traffic regulation databases, potentially leading to compliance risks. These issues severely impact the user experience of cross-language intelligent driving. To address at least one of these problems, this application provides a cross-language driving collaborative control method that can reduce or avoid these situations. Through intent parsing, localization correction, and fleet collaboration, it achieves accurate conversion from general commands to scenario-based driving decisions.

[0052] This application provides a cross-language driving cooperative control method and apparatus. The method and apparatus are based on the same concept, and since the principles by which they solve problems are similar, their implementations can be mutually referenced; repeated details will not be repeated.

[0053] Please refer to Figure 1 This application provides a cross-language driving cooperative control method, comprising:

[0054] Step 11: Perform intent parsing on the multilingual instructions obtained from the user input to generate the first driving instruction.

[0055] In this embodiment, multilingual and multi-form user input commands (such as voice and text) are converted into standardized initial driving commands, solving the fundamental problem of cross-language command understanding. User input forms include, but are not limited to: users may input commands via voice, such as a mixed Chinese and English "Turn left at the next intersection," or a combination of text or gestures and voice. The input commands are denoised (voice) and segmented (text). The XLM-RoBERTa model is used to convert multilingual text into a unified semantic vector, bridging the semantic gap caused by language differences. A fine-tuned Transformer model is used to classify the semantic vector into preset driving intent categories, such as 12 core intents including "turn left," "accelerate," "check speed limit," and "overtaking request," improving classification accuracy. The classified intents are mapped to structured first driving commands, such as "Intent: Turn left; Trigger condition: Next intersection; Priority: High," serving as the basis for subsequent corrections.

[0056] For example, if a user enters the French phrase "Puis-jedoublerenface?" ("May I overtake ahead?"), the system will parse it and generate the first driving instruction: "Intent: Request to overtake; Target location: Road ahead; Status: Pending verification".

[0057] Step 12: Load local traffic rules based on the vehicle's location, and use the local traffic rules and a preset multilingual road sign recognition model to correct the first driving instruction into a second driving instruction.

[0058] In this application, step 12 combines the current regional traffic rules and real-time road sign information to modify the initial instruction into an intermediate instruction that complies with local regulations, thus resolving the compliance issues of cross-border driving.

[0059] Local traffic rules are dynamically loaded by obtaining the vehicle's real-time location via GPS, converting it to a projected coordinate system, and matching it with a 1km grid cell ID. Rules for the corresponding region are loaded from a pre-set global national legal database, while temporary rules are obtained via a local traffic API. Rule confidence is updated by retrieving historical user response data; mandatory constraints are set for high-priority rules, while low-priority rules are user-defined. Images of road signs ahead are captured by the vehicle's camera and input into a pre-set multilingual road sign recognition model, such as a dual-branch VLM model. This allows the first driving instruction to be corrected to a second driving instruction using the local traffic rules and the pre-set multilingual road sign recognition model.

[0060] It should be noted that if the first driving instruction is consistent with local rules and road signs, it will be retained and strengthened; if there is a conflict, it will be amended to be a compliant instruction; if there are temporary rules, the temporary rules will be used as the priority for amendment.

[0061] Step 13: Based on the shared information sent by the fleet to which the vehicle belongs, modify the second driving instruction to the target driving instruction.

[0062] This application resolves command conflicts in cross-language fleets by sharing information among multiple vehicles and making group decisions, generating unified final instructions for execution and ensuring safe collaborative driving.

[0063] The three-step process of this application is progressive. Step 11 solves the problem of cross-language understanding of user instructions, step 12 achieves localized and compliant adaptation of instructions, and step 13 completes the unified decision-making of fleet collaboration. Finally, it generates target driving instructions that not only meet user intentions and comply with local rules, but also collaborate with the fleet, which significantly improves the safety and efficiency of cross-language and cross-border driving.

[0064] Optionally, after step 13 above, the method further includes:

[0065] Acquire users' native language and cultural characteristics;

[0066] Based on the native language and cultural characteristics, perceptual interaction information corresponding to the target driving command is generated.

[0067] In this embodiment, contextual adaptation of a multilingual user interface can be provided. This application employs an index ID mapping algorithm to achieve dynamic language switching of interface elements. Based on the user's initial settings, such as selecting native language (French) and cultural region (France) during registration, or historical interaction data, such as frequent use of French commands or receiving traffic updates from France, the system automatically identifies the user's native language (e.g., Chinese, Spanish, German) and cultural characteristics (e.g., language expression habits, cultural taboos, interaction preferences). The native language is directly associated with the language type of the interface text and voice feedback; cultural characteristics include expression style (e.g., French prefers polite sentences, Chinese prefers concise and direct language), symbol understanding (e.g., red indicates warning in most cultures, but may differ in some cultures), and interaction habits such as prompt frequency and information density.

[0068] Based on native language and cultural characteristics, perceptual interaction information corresponding to the target driving command can be generated. This can transform the target driving command into multi-dimensional interactive information (text, image, voice) that conforms to the user's native language and cultural habits, reducing language barriers and improving user acceptance.

[0069] For example, based on an index ID mapping algorithm, instruction keywords (such as "speed warning") are bound to corresponding text in a multilingual library. The corresponding language text is automatically invoked based on the user's native language, ensuring real-time adaptation of interface elements. Combined with semantic segmentation technology, multilingual labels are overlaid onto the display module or system of the vehicle perception interaction, along with culturally compatible icons, reducing reliance on language.

[0070] For example, Chinese user interfaces have a higher information density (integrating multiple types of prompts), while some Western users prefer to display information in bullet points; avoid using taboo symbols in the target culture.

[0071] For example, "Speed ​​Warning" is displayed as "Caution Speed!" in Chinese or "Watch Your Speed!" in Spanish in different language environments. "Con la velocidad!" This application can combine semantic segmentation technology to overlay multilingual labels (such as the German "Stau") onto real-world road condition images, and reduce language dependence through icon-based design (such as red triangles indicating danger). The voice feedback module integrates a cultural adaptation strategy: French users receive polite prompts ("Veuillez ralentir"), while Chinese users prefer concise instructions ("Slow down").

[0072] This step, through a dual design of "language adaptation and cultural customization," makes the interactive information of the target driving instructions more in line with the user's cognitive habits. It not only solves the problem of cross-language communication, but also improves the user experience and the accuracy of instruction execution through cultural details.

[0073] This application, through the deep integration of multimodal artificial intelligence (AI) and edge computing, constructs the first autonomous driving collaborative method that supports cross-language, cross-cultural, and cross-regulatory communication. It solves the problems of language ambiguity, rule conflict, and inefficient collaboration in cross-border scenarios of traditional solutions. The solution of this application improves the accuracy of intent parsing, ensures compliance, and reduces communication latency.

[0074] Optionally, step 11 above includes:

[0075] Perform local and / or online speech recognition on multilingual commands input by users to obtain corresponding multilingual text information;

[0076] The multilingual text information is input into a preset multi-intent recognition model to generate a first driving instruction corresponding to the multilingual instruction.

[0077] In this embodiment, by integrating local and online speech recognition technologies, multilingual commands input by the user are subjected to local and / or online speech recognition to achieve accurate transcription of mixed language and non-standard accent commands. Optionally, speech input preprocessing is performed, including but not limited to: 1) Front-end noise reduction and speech separation: A speech enhancement model based on deep neural networks (SEGAN) is used to reduce noise in the original in-vehicle audio (filtering engine and external noise) and separate the speech of multiple speakers. For example, when multiple people in the vehicle issue commands at the same time, the source of the main command is distinguished to ensure the purity of the speech signal. 2) Local speech recognition (applicable to offline scenarios): This application uses the DeepSpeech model, trained for mixed language commands (such as "turn left ahead into MainStreet" with mixed Chinese and English), to support real-time recognition in offline environments. After fine-tuning on a mixed speech dataset, the model improves the recognition accuracy of mixed Chinese and English commands. 3) Online speech recognition (supplementary scenarios): The speech recognition API is called to cover the recognition needs of minority languages, such as Thai and Arabic, to solve the problem of insufficient support for minority languages ​​by the local model.

[0078] Furthermore, a fusion strategy is implemented. Local and online recognition results are integrated based on confidence-weighted voting. The local model's confidence weight is set to 0.7 to prioritize offline reliability, while the online API's weight is set to 0.3 to supplement accuracy for less common languages. The final output is unified multilingual text information. This multilingual text information is then input into a pre-defined multi-intention recognition model to generate the first driving instruction corresponding to the multilingual commands.

[0079] Optionally, the multilingual text information is input into a preset multi-intent recognition model to generate a first driving instruction corresponding to the multilingual instruction, including:

[0080] The multilingual text information is input into a preset multi-intent recognition model, the multilingual text information is segmented into words, the segmented sub-word sequence is obtained, the sub-word sequence is added to a preset position code, and mapped into a first vector;

[0081] Using the first vector and the network architecture layer of the multi-intent recognition model, a probability distribution of each intent corresponding to the multilingual text information is generated;

[0082] Based on the preset reinforcement learning model and the probability distribution of each intention, a corresponding first driving instruction is generated.

[0083] In this embodiment, a multilingual intent classification model can be used to transform text information into structured driving intents. The core of this model is to achieve accurate mapping of cross-language intents through word segmentation, semantic encoding, and context extraction. The specific process is as follows: The SentencePiece segmenter of the XLM-RoBERTa model is used to convert multilingual text into a sequence of words, solving the segmentation problems of mixed languages ​​and uncommon words. For example, the input German instruction "Biegen Sie an der" Kreuzung links ab” is followed by a participle: ["Biegen","Sie","an","der", ["_Kreuzung"," links"," ab"]. The output intent category is 0 (direction), with a confidence level of 0.92.

[0084] A pre-defined positional encoding (Sinusoidal function) is added to the sub-word sequence to mark word order information (such as the position of "links" (left) affecting the judgment of "turn" intent). This is then mapped to a 768-dimensional first vector through an embedding layer, which is used to unify the semantic space. This first vector is input to the 12-layer Transformer encoder of the multi-intent recognition model to extract contextual semantic features, such as… The combination of "Kreuzung" (next intersection) and "links" (left) reinforces the intention to "turn".

[0085] Furthermore, contextual semantic extraction is performed: features are extracted through a 12-layer Transformer encoder, with the last layer [CLS] labeled vector serving as the input for intent classification.

[0086] Furthermore, the classification head is optimized: the fully connected layer and Softmax output the probabilities of 5 types of intents (turn, accelerate, stop, query, others).

[0087] Furthermore, feature dimensionality reduction is performed: the [CLS] label vector (768 dimensions) extracted by the XLM-RoBERTa model is input into the fully connected layer, and the dimensionality is reduced to 256 dimensions through linear transformation.

[0088] The specific formula used is: h hidden =W fc1 *h CLS +b fc1 W fc1 ∈R 256×768 b fc1 ∈R 256 The dimension is reduced to 256 through linear transformation, where h CLSThe hidden state representing the classification token (CLS token) originates from the output vector of the special token [CLS] used to aggregate sequence information in the Transformer model. It serves as the model's feature input for the classification task; W fc1 This represents the weight matrix of the fully connected layer. The output vector has 256 rows, and the input vector h... CLS The dimension has 768 columns, used to map the 768-dimensional input to a 256-dimensional space; b fc1 h represents the bias vector of the fully connected layer. hidden This represents the output of the fully connected layer.

[0089] Furthermore, intent probability calculation is performed: the second fully connected layer maps the 256-dimensional features to 5 dimensions (corresponding to 5 types of intents). The formula used is: h logits =W fc2 *h hidden +b fc2 Intent probability calculation is performed. Where W fc2 It is a weight matrix of shape R5×256, meaning it has 5 rows and 256 columns, used to weight the hidden layer output h. hidden (Typically a 256-dimensional vector) undergoes a linear transformation. fc2 It is a bias vector of shape R5, containing 5 elements, used to offset the result of the linear transformation. hidden h represents the output vector of the hidden layer. logits This is the output calculated by the fully connected layer, typically used as the input to the softmax activation function in subsequent classification tasks. It has a dimension of 5, corresponding to the original scores for the 5 categories. Wherein, W... fc2 This represents the weight matrix, with dimensions 5×256 (5 rows and 256 columns), used for linear transformation of the hidden layer vectors. fc2 This represents the bias vector, with a dimension of 5, used to adjust the result after the linear transformation. hidden This represents the hidden layer output vector (typically 256-dimensional), which serves as the input to that layer. logits This represents the output (5-dimensional vector) of the fully connected layer, which is generally used as the raw score for classification tasks (without softmax activation).

[0090] Finally, Softmax normalization is performed. The 5-dimensional output is processed using Softmax to obtain the probability distribution for each intent category. Specifically, it is mapped to a 5-dimensional vector through the second fully connected layer, corresponding to the five core intents: "turn, accelerate, stop, query, and others." Softmax normalization is then performed on the 5-dimensional vector to generate the probability distribution for each intent. For example, the output of the above German instruction might be: turn (0.92), accelerate (0.03), stop (0.02), query (0.02), and others (0.01).

[0091] Optionally, based on a preset reinforcement learning model and the probability distribution of each intention, a corresponding first driving instruction is generated, including:

[0092] Obtain the encoding of users' historical behavioral characteristics and real-time environmental characteristics;

[0093] The historical behavioral features and the real-time environmental features are encoded and input into the reinforcement learning model, and the reinforcement learning model outputs safety weights.

[0094] The probability distribution of each intent is adjusted according to the security weights.

[0095] Based on the adjusted intent probability distribution, the corresponding first driving instruction is generated.

[0096] In this embodiment, reinforcement learning is used to dynamically calibrate intent priority and adjust intent probability by combining user behavior and real-time environment to ensure that the generated instructions conform to user habits and scenario requirements. The specific process is as follows: Obtain the user's historical behavior features and real-time environment feature encodings. The user's historical behavior and real-time environment feature encodings can be selected as 10-dimensional vectors.

[0097] The state space is defined as user historical behavior features (such as braking frequency and lane-changing preferences) and real-time context (weather, road conditions), encoded as a 10-dimensional vector. This vector is the source of input layer data for the PPO policy network.

[0098] User behavior characteristics (5 dimensions) include: braking frequency (times / 100 km); average lane change interval time (seconds); voice command response latency (ms); number of speeding incidents (times / 1000 km); and frequency of manual intervention (times / hour).

[0099] Environmental characteristics (5 dimensions) include: weather (rain / snow / sunny, coded as 0-1); real-time traffic conditions (congestion index 0-1); road type (highway / urban / rural, unique hot coding); time (day / night, 0 or 1); visibility (meters, normalized to 0-1).

[0100] The example vector can be represented as: [0.2,15.3,200,2,0.5,0,0.7,1,0,0.8].

[0101] Optionally, this application uses reinforcement learning (PPO algorithm) to optimize intent weights and adjust decision priorities based on user historical behavior. Vehicle control parameters are adjusted based on user actions, and rewards (user satisfaction) are calculated. Then, a PPO policy model is trained by calling the `model.learn` method to optimize intent weights and adjust decision priorities based on user historical behavior. The state vector is input into the PPO (Proximal Policy Optimization) reinforcement learning model.

[0102] Optionally, the decision priority adjustment mechanism defines the action space as follows: safety weight adjustment (0.1–0.9 steps, 0.1 increments) and efficiency weight adjustment (inversely complementary). This safety weight adjustment directly affects the vehicle control logic. The adjustment logic includes: when the safety weight increases, the vehicle control parameters become more conservative (e.g., increasing following distance); when the efficiency weight increases, the path planning becomes more aggressive (e.g., allowing shorter lane change intervals).

[0103] This application also establishes a user satisfaction calculation and reward function. The user satisfaction R calculation formula is: R = α * Safety Score + β * Efficiency Score + γ * Interaction Satisfaction. Here, user satisfaction R is a comprehensive index, and α, β, and γ are the weighting coefficients of each score, used to measure the importance of safety score, efficiency score, and interaction satisfaction in user satisfaction R. Each score is multiplied by its corresponding weight and then summed to obtain the comprehensive result. The safety score is used to represent the number of emergency braking incidents (+0.2 points for each reduction); the efficiency score is used to represent the percentage reduction in travel time (+0.5 points for each 10% reduction); and the interaction satisfaction score is used to represent the voice assistant's rating (1-5 stars, linearly mapped to 0-1 points). Coefficient settings: α = 0.5, β = 0.3, γ = 0.2.

[0104] The policy update formula for the reward function is:

[0105]

[0106] The advantage function A is calculated using GAE (Generalized Advantage Estimation); the KL divergence threshold is set to 0.01 to prevent policy mutations. θ represents the policy network parameters; πθk represents the policy in the k-th iteration; Aπθk(s,a) represents the advantage function, which measures the relative value of action a in state s; λ represents the KL divergence penalty coefficient (default 0.01).

[0107] The architecture of the PPO model includes: the input layer can be a 10-dimensional state vector; the hidden layer can be a 2-layer fully connected network (256→128 nodes, ReLU activation); the output layer can be a safe weight adjustment action (discrete action space, step size ±0.1); the optimizer can be Adam (learning rate 3e-4); and the GAE parameters can be discount factors γ = 0.99 and λ = 0.95.

[0108] The training process of the PPO model includes: Data collection: Interacting with the environment using the current policy πθk to collect trajectory data (st, at, rt). Advantage estimation: Calculating the advantage value At for each state-action pair using GAE (Generalized Advantage Estimation).

[0109] Strategy optimization: 1) Maximize the objective function L(θ) while constraining the KL divergence to not exceed a threshold (e.g., 0.01). 2) Unimproved part: Directly adopt the PPO-Clip algorithm without modifying the core formula.

[0110] User satisfaction feedback and weight optimization linkage mechanism. 1) Reward calculation: Calculate the real-time reward R (weighted sum of safety, efficiency, and interaction scores) based on trip data. 2) Policy gradient update: Update the policy network parameters θ using the PPO algorithm to make future decisions more inclined towards high-reward actions. 3) Intent weight adjustment: The policy network outputs the safety weights safewsafe (selection in the action space), which directly affect vehicle control parameters: Following distance = base distance × (1 + safe) Following distance = base distance × (1 + wsafe). For example, if the user frequently brakes suddenly, the safety score decreases, and the policy network will increase wsafe, thus increasing the following distance.

[0111] Furthermore, by strengthening the safety weights output by the learning model, the initial probability distribution generated by the multi-intent recognition model is dynamically corrected, making the intent priority more aligned with the user's safety preferences and real-time scenario. The highest priority intent is selected from the optimized probability distribution, and combined with scenario details, a structured and executable initial driving instruction is generated, i.e., the corresponding first driving instruction is generated.

[0112] For example, in the original probability distribution, "turning" accounted for 0.92. After adjusting for safety weights, the instruction was refined to "turn left at the next intersection, with a recommended speed of ≤30km / h (safety first)".

[0113] This application achieves accurate conversion of multilingual instructions into standardized first-drive instructions through a complete closed-loop logic of speech recognition transcription, text intent classification, and reinforcement learning calibration. It solves the problem of intent ambiguity in cross-language and cross-scenario situations and improves the accuracy of intent matching.

[0114] Optionally, step 12 above includes:

[0115] Based on the obtained vehicle location, load the preset local traffic rules;

[0116] Obtain the user's historical response information to the corresponding local traffic rules;

[0117] Based on the response history information and Bayesian inference, the rule priority of the local traffic rules is adjusted;

[0118] Obtain the road sign image corresponding to the vehicle location and the multilingual prompt text input by the user;

[0119] The road sign image and the multilingual prompt text are input into a preset multilingual road sign recognition model, and the road sign semantic classification result is output.

[0120] Based on the rule priority and the road sign semantic classification result, the first driving instruction is modified into a second driving instruction.

[0121] In this embodiment of the application, this step is the core process of localized traffic rule adaptation and multilingual road sign integration. By dynamically loading regional rules, adapting to user behavior preferences, and accurately identifying multilingual road signs, the initial driving instruction is modified into a second driving instruction that conforms to local regulations and user habits.

[0122] Specifically, the system obtains the vehicle's real-time location via GPS, converts latitude and longitude to a projected coordinate system using a coordinate converter, and then calculates the corresponding 1km grid cell ID. Next, it loads rules, such as basic rules for the current area from a pre-built global database of traffic regulations from multiple countries, based on the grid cell ID. Examples include the EU's "right turn priority" rule and Japan's "30km / h speed limit" for residential areas. Simultaneously, it connects to the local traffic management department's API to obtain temporary rules, such as "construction detour" and "time-limited one-way street," ensuring the rules' timeliness. For instance, when a vehicle enters Munich, Germany, it loads basic rules such as the default 50km / h speed limit and right-hand traffic priority at intersections without traffic lights, based on the grid cell ID, while also obtaining a temporary notice that a certain section of road is closed from 2:00 PM to 6:00 PM.

[0123] The system retrieves historical information about user responses to local traffic rules from a database recording user interactions with traffic rules in the current area. This information includes, but is not limited to: actions taken to manually correct rules; the degree of compliance with the rules; and the frequency of interactions. This provides a basis for adjusting rule priorities based on user behavior, avoiding user resistance caused by mechanically applying rules.

[0124] Based on user response history, the confidence level of rules is dynamically updated through Bayesian inference to adjust the priority of local traffic rules. The Bayesian inference process for adjusting priority is as follows: Prior probability: The default confidence level of the initially loaded local rules, such as the prior probability P1 (rule valid) = 0.9 for the German highway "130km / h speed limit" rule; Likelihood calculation: Calculate the association probability between "user behavior and rule validity" based on user behavior, such as the likelihood probability P2 (speeding | rule valid) = 0.1 for "user speeding when complying with valid rules" if a user speeds 5 times consecutively; Posterior update: Calculate the posterior probability based on the preset Bayesian formula, prior probability P1, and likelihood probability P2. If the posterior probability decreases significantly, the rule is determined to conflict with user habits; the execution priority of the rule is reduced, such as from "mandatory compliance" to "suggested reference," allowing users to customize rules and marking them as "low confidence rules." For example, after a user repeatedly exceeds the speed limit on a German highway, the priority of the "130km / h speed limit" rule is reduced, and the user-defined 150km / h speed limit is used instead.

[0125] The system acquires road sign images corresponding to the vehicle's location and multilingual prompt text input by the user. It can collect road sign images of the current location through the vehicle's multimodal sensors (such as a forward-facing camera), including traffic signs such as speed limit signs and yield signs. It can also receive query text input by the user, such as "What is this sign?" in 10 languages ​​including Chinese, English, German, and French, as text guidance for model recognition.

[0126] Furthermore, the road sign images and multilingual prompt texts are input into a pre-defined multilingual road sign recognition model, and the road sign semantic classification results are output. This application uses a two-branch visual language model (VLM) to achieve cross-modal road sign recognition. The specific process is as follows:

[0127] Visual branch processing: The road sign image is input into the ResNet-50 backbone network (ImageNet pre-trained) to extract high-level semantic features of 7×7×2048; the YOLOv5 detection head outputs the road sign position or bounding box coordinates, and the DeepLabv3+ segmentation head outputs pixel-level semantic masks to distinguish the road sign from the background, such as excluding tree occlusion.

[0128] Text branching: Multilingual prompt text, such as "What is this sign" or "This is what sign", is input into the XLM-RoBERTa model to generate a 768-dimensional text feature vector; cross-language semantic alignment is achieved by constraining the similarity of text features in different languages ​​through contrastive loss.

[0129] Cross-modal fusion: Using visual features as the query and text features as the key / value pair, the cross-attention weights can be calculated using the following formula:

[0130]

[0131] Where: Q represents the query vector obtained from visual feature extraction; K represents the key vector obtained from text feature extraction; V represents the value vector obtained from text feature extraction; d k is the dimension of the key vector, used to scale the weights to avoid the inner product result being too large due to the high vector dimension, which would affect the gradient stability of the softmax function; the softmax function is used to normalize the calculation result into a probability distribution so that the sum of the attention weights is 1.

[0132] The cross-attention weights calculated using this formula can measure the degree of correlation between visual features and different text features, enabling the effective fusion of visual information and multilingual text information, and providing a basis for road sign semantic classification.

[0133] Attention-weighted visual features are concatenated with original text features and input into a fully connected layer to generate the final classification result. The final classification result will be used for road sign semantic parsing, and will also trigger rule base matching, path modification and vehicle control.

[0134] Furthermore, by combining the adjusted rule priorities with the road sign recognition results, the initial first driving instruction is revised, including: if the road sign semantics are consistent with local rules, then the rule priority is followed; if the road sign semantics conflict with local rules, then the road sign result is prioritized for correction; if the rule priority has been reduced (e.g., user-defined speed limit), then fine-tuning is done in conjunction with the road sign. For example, if the first driving instruction is "continue straight," but a French "Déviation (detour)" road sign is recognized, and local temporary rules indicate construction ahead, then the second driving instruction is revised to "turn right and detour to XX Road."

[0135] This application achieves precise adaptation of driving instructions to local regulations, user habits, and real-time road conditions through a process of dynamic rule loading, user behavior adaptation, multilingual road sign recognition, and instruction fusion and correction, significantly reducing compliance risks for cross-border driving.

[0136] Optionally, step 13 above includes:

[0137] Receive a shared message that establishes a fleet communication connection with the vehicle;

[0138] Based on multiple shared messages, identify the fleet decision leader who establishes fleet communication connections with the vehicle;

[0139] Determine the target language required by the fleet;

[0140] Send a second driving instruction, translated into the target language, to the fleet decision-making leader;

[0141] Receive the group consensus instruction sent by the fleet decision leader; the group consensus instruction is the instruction with the highest decision weight determined by the fleet decision leader based on multiple candidate instructions to construct a weighted decision.

[0142] Based on the aforementioned group consensus instruction, the second driving instruction is modified into the target driving instruction.

[0143] In this embodiment, step 13 is the core process of group decision-making in cross-language driving collaboration. Based on the Vehicle-to-Everything (V2X) communication protocol and the Raft consensus algorithm, it realizes the collaboration and conflict resolution of multilingual commands within the fleet, ultimately generating a unified target driving command. Specifically, the vehicle establishes real-time communication with other vehicles in the fleet via the V2X communication protocol (DSRC or 5G) and receives shared messages from other vehicles. These messages include each vehicle's driving intentions, such as "turn left" or "slow down," and real-time traffic conditions, such as "congestion ahead," and support automatic multilingual translation, such as automatically translating Chinese messages into English, German, etc., ensuring cross-language vehicle understanding. Example: This vehicle (Chinese interface) receives the German message "Voraus wird rechts abgebogen" from a German vehicle, which the system automatically translates as "Turn right ahead."

[0144] A leader election mechanism based on the Raft consensus algorithm elects a fleet decision leader from the fleet. During the election, weights are dynamically adjusted according to vehicle location density. Vehicles in high-density areas (such as densely packed convoy sections) are given increased priority, ensuring the leader can more accurately coordinate surrounding vehicles. The fleet decision leader is responsible for aggregating instructions from all vehicles, handling conflicts, and generating group consensus, avoiding the chaos caused by multiple vehicles making independent decisions.

[0145] Once the target language for the fleet is determined, instructions are uniformly translated into the "majority language" based on the language distribution of fleet members. For example, if most vehicles in a European fleet use English and German, the target language is set to "English and German"; for a Sino-German mixed fleet, it can be set to "Chinese and English". This ensures that all vehicles can understand the leader's decisions and instructions, thus overcoming cross-language communication barriers.

[0146] The vehicle translates its second driving instructions into the target language and sends them to the elected team decision-making leader as candidate instructions for group decision-making.

[0147] Before receiving the group consensus instruction (represented by "Score(c)") sent by the fleet decision leader, the fleet decision leader summarizes the candidate instructions of all vehicles. For example, some vehicle instructions are "turn left" and some are "turn right". Based on the conflict decision weight voting formula, the weight score of each candidate instruction is calculated, and finally the instruction with the highest weight is selected as the group consensus instruction.

[0148] The formula for weighted voting in conflict decision-making is as follows:

[0149]

[0150] Among them, w v This represents the rule weight of the country where vehicle v is located. For example, a vehicle in Germany has a weight of 1.2 due to strict local traffic rules, while a vehicle in China has a weight of 1.0. V represents the set of all vehicles participating in the decision-making process. I(c v =c) represents the indicator function. If the instruction of vehicle v is the candidate instruction c, the value is 1, otherwise it is 0; c represents the candidate instruction.

[0151] For example, suppose out of 10 cars, 6 German cars are instructed to "turn right" and 4 Chinese cars are instructed to "turn left". The score for turning right = 6 × 1.2 (weight for German cars) = 7.2; the score for turning left = 4 × 1.0 (weight for Chinese cars) = 4.0; the consensus instruction is "turn right" (highest weight).

[0152] Based on the group consensus instruction, the second driving instruction is modified into the target driving instruction. This includes: the vehicle receiving the group consensus instruction sent by the leader and modifying its original second driving instruction to this consensus instruction, which becomes the final target driving instruction to be executed. This ensures that all vehicles in the fleet execute unified instructions, avoids the risk of collisions caused by conflicting decisions among multiple vehicles, and achieves collaborative driving across languages ​​and regions.

[0153] This application resolves the command conflict problem in cross-language fleets through a process of "information sharing - leader election - multilingual unification - weighted voting - command correction", and achieves efficient and safe collaborative driving by leveraging an algorithmic group decision-making mechanism.

[0154] Reference Figure 2 As shown in the illustration, this application also provides a cross-language driving cooperative control system, applied to vehicles with cross-language driving cooperation. The system includes a perception layer, a decision layer, and an execution layer. The perception layer includes a multilingual V2X communication module and multimodal sensors; the decision layer includes an intent parsing engine, a local rule base, and a group cooperation algorithm; the execution layer includes a multilingual control interface and an AR-HUD display module.

[0155] The system consists of three layers: the perception layer, which is the information input center and is responsible for collecting multi-source environmental information and cross-vehicle interaction data to provide the original basis for decision-making; the decision layer, which is the brain of the system, performs intent parsing, rule matching and group collaborative decision-making based on the data from the perception layer to generate standardized driving instructions; and the execution layer, which is the action terminal of the system, which transforms the instructions from the decision layer into vehicle control actions and user interaction feedback.

[0156] Optionally, a multilingual V2X communication module can achieve cross-vehicle, cross-language information interaction based on the V2X communication protocol, supporting real-time transmission and automatic translation of driving commands in different languages ​​(such as Chinese "congestion ahead," German, etc.). It receives warning information and road condition updates from other vehicles and translates the vehicle's information into the target vehicle's interface language, ensuring barrier-free information sharing across multinational fleets. The multimodal sensor integrates cameras, LiDAR, millimeter-wave radar, and voice acquisition devices to collect data on the vehicle's surrounding environment and user commands. Visual sensors recognize multilingual road signs (such as Chinese "speed limit 80," French, etc.), traffic lights, pedestrians, etc.; voice devices collect mixed-language user commands; and radar senses the distance and relative speed between the vehicle and obstacles, providing data for safety decisions.

[0157] Optionally, the decision layer performs real-time parsing of users' multilingual voice commands (such as mixed Chinese and English, non-standard accents) and maps them to standardized driving intentions (such as turning, accelerating, stopping). The decision layer's intention parsing engine can process mixed language text, accurately classifying intentions through word segmentation and contextual semantic extraction; combined with reinforcement learning, it dynamically adjusts intention priorities based on users' historical behavior and the real-time environment. The local rule base can pre-set traffic regulations for countries around the world (such as right-turn priority in the EU, and a speed limit of 30km / h in Japanese residential areas), and dynamically load current area rules based on vehicle GPS coordinates. Driving strategies are corrected by combining road sign information identified by multimodal sensors; Bayesian inference is used to adapt to user behavior. In the group collaboration algorithm, a consensus algorithm is used to coordinate path planning and decision-making conflicts among multiple vehicles, enabling collaborative driving of multinational fleets.

[0158] The multilingual control interface in the execution layer interfaces with the vehicle's underlying control system (such as steering, throttle, and brakes), converting standardized commands generated by the decision-making layer into control signals that the vehicle can execute. Additional features include support for multilingual voice feedback, adapting to different cultural expression habits.

[0159] The AR-HUD (Augmented Reality-Head-Up Display) module uses augmented reality (AR) head-up display technology to overlay key information (such as multilingual road sign translations, navigation instructions, and warning prompts) onto the real road conditions on the windshield. It can display localized information and reduce language dependence through icon-based design (such as red triangles indicating danger); it also updates collaborative decision-making results in real time to ensure users intuitively understand driving instructions.

[0160] In summary, the cross-language driving cooperative control system can apply the aforementioned cross-language driving cooperative control methods. Through closed-loop collaboration of "perception-decision-execution", it realizes driving command parsing, local rule adaptation, group collaboration and user interaction in cross-language environments, solving the problems of language barriers, rule conflicts and inefficient collaboration in cross-border driving.

[0161] This application solves the language and cultural barriers in cross-border autonomous driving by embedding cross-language command real-time parsing and intent mapping, combined with a multilingual collaborative driving group decision optimization strategy, applied to a model that integrates localized traffic rules and map data, and optimizing the contextual adaptation function of multilingual user interfaces.

[0162] In summary, existing technologies rely on single-language instruction parsing, which cannot handle mixed-language input (such as mixed Chinese and English instructions) or non-standard accents. This application integrates a multilingual large model to achieve end-to-end parsing of mixed-language instructions; it dynamically calibrates user intent priorities (e.g., adjusting security preference weights) through reinforcement learning; and it employs an intent calibration model based on user's native language habits, significantly improving intent matching accuracy. Specifically, a reinforcement learning intent calibration model is used, and the PPO algorithm optimizes user preference weights.

[0163] This application broadcasts multilingual traffic events (such as congestion and accidents) in real time and automatically translates them into the target vehicle's language (such as Chinese to English). At the same time, it coordinates the behavior of multiple vehicles based on a consensus algorithm and dynamically adjusts path priorities (such as avoiding construction areas marked in multiple languages).

[0164] Existing technologies use static binding of rule bases, which cannot adapt to the regulations of multiple countries in real time. This application uses a dynamic rule base loading engine to achieve real-time matching of regional traffic rule bases based on GPS coordinates; it also uses a visual language model to recognize multilingual road signs and integrate them with map data to correct navigation strategies.

[0165] The various methods of the embodiments of this application have been described above. Apparatus for implementing the above methods will now be provided.

[0166] Please refer to Figure 3 This application also provides a cross-language driving cooperative control device, including:

[0167] The first generation module 31 is used to perform intent parsing on the multilingual instructions obtained from the user input and generate the first driving instruction;

[0168] The first processing module 32 is used to load local traffic rules based on the vehicle's location, and use the local traffic rules and a preset multilingual road sign recognition model to correct the first driving instruction into a second driving instruction.

[0169] The second processing module 33 is used to modify the second driving instruction into the target driving instruction based on the shared information sent by the fleet to which the vehicle belongs.

[0170] Optionally, the cross-language driving cooperative control device in this application embodiment further includes:

[0171] The first acquisition module is used to acquire the user's native language and cultural characteristics;

[0172] The second generation module is used to generate perceptual interaction information corresponding to the target driving command based on the native language and cultural characteristics.

[0173] Optionally, the first generation module 31 described above includes:

[0174] The first acquisition unit is used to perform local and / or online speech recognition on the multilingual instructions input by the user to acquire the corresponding multilingual text information.

[0175] The first generation unit is used to input the multilingual text information into a preset multi-intent recognition model to generate a first driving instruction corresponding to the multilingual instruction.

[0176] Optionally, the first generation unit described above is specifically used for:

[0177] The multilingual text information is input into a preset multi-intent recognition model, the multilingual text information is segmented into words, the segmented sub-word sequence is obtained, the sub-word sequence is added to a preset position code, and mapped into a first vector;

[0178] Using the first vector and the network architecture layer of the multi-intent recognition model, a probability distribution of each intent corresponding to the multilingual text information is generated;

[0179] Based on the preset reinforcement learning model and the probability distribution of each intention, a corresponding first driving instruction is generated.

[0180] Optionally, based on a preset reinforcement learning model and the probability distribution of each intention, a corresponding first driving instruction is generated, including:

[0181] Obtain the encoding of users' historical behavioral characteristics and real-time environmental characteristics;

[0182] The historical behavioral features and the real-time environmental features are encoded and input into the reinforcement learning model, and the reinforcement learning model outputs safety weights.

[0183] The probability distribution of each intent is adjusted according to the security weights.

[0184] Based on the adjusted intent probability distribution, the corresponding first driving instruction is generated.

[0185] Optionally, the first processing module 32 described above includes:

[0186] The first processing unit is used to load preset local traffic rules based on the acquired vehicle location;

[0187] The second acquisition unit is used to acquire the user's historical response information to the corresponding local traffic rules;

[0188] The second processing unit is used to adjust the rule priority of the local traffic rules based on the response history information and Bayesian inference.

[0189] The third acquisition unit is used to acquire the road sign image corresponding to the vehicle location and the multilingual prompt text input by the user;

[0190] The third processing unit is used to input the road sign image and the multilingual prompt text into a preset multilingual road sign recognition model and output the road sign semantic classification result;

[0191] The fourth processing unit is used to modify the first driving instruction into a second driving instruction based on the rule priority and the road sign semantic classification result.

[0192] Optionally, the second processing module 33 described above includes:

[0193] The first receiving unit is used to receive a shared message that establishes a fleet communication connection with the vehicle.

[0194] The first determining unit is configured to determine, based on multiple shared messages, the fleet decision leader who establishes a fleet communication connection with the vehicle;

[0195] The second determining unit is used to determine the target language required by the fleet;

[0196] The fifth processing unit is used to send a second driving instruction, translated into the target language, to the fleet decision-making leader;

[0197] The second receiving unit is used to receive the group consensus instruction sent by the fleet decision leader; the group consensus instruction is the instruction of the fleet decision leader to determine the highest decision weight based on multiple candidate instructions to construct a weighted decision.

[0198] The sixth processing unit is used to modify the second driving instruction into the target driving instruction according to the group consensus instruction.

[0199] It should be noted that the device in this embodiment corresponds to the device applied to the cross-language driving cooperative control method described above. The implementation methods in the above embodiments are all applicable to the embodiments of this device and can achieve the same technical effect. The device provided in this application embodiment can implement all the method steps implemented in the above method embodiments and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiments and the beneficial effects will not be described in detail.

[0200] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described cross-language driving cooperative control method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may include, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0201] This application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the various processes of the above-described cross-language driving cooperative control method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0202] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0203] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0204] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A cross-language driving cooperative control method, characterized in that, include: The intent of the multilingual commands input by the user is parsed to generate the first driving command; Based on the vehicle's location, local traffic rules are loaded, and the first driving instruction is corrected into a second driving instruction using the local traffic rules and a preset multilingual road sign recognition model. Based on the shared information sent by the fleet to which the vehicle belongs, the second driving instruction is modified into the target driving instruction.

2. The method according to claim 1, characterized in that, After modifying the second driving instruction to the target driving instruction, the method further includes: Acquire users' native language and cultural characteristics; Based on the native language and cultural characteristics, perceptual interaction information corresponding to the target driving command is generated.

3. The method according to claim 1, characterized in that, The intent of the multilingual commands obtained from user input is parsed to generate the first driving command, including: Perform local and / or online speech recognition on multilingual commands input by users to obtain corresponding multilingual text information; The multilingual text information is input into a preset multi-intent recognition model to generate a first driving instruction corresponding to the multilingual instruction.

4. The method according to claim 3, characterized in that, The multilingual text information is input into a preset multi-intent recognition model to generate a first driving instruction corresponding to the multilingual instruction, including: The multilingual text information is input into a preset multi-intent recognition model, the multilingual text information is segmented into words, the segmented sub-word sequence is obtained, the sub-word sequence is added to a preset position code, and mapped into a first vector; Using the first vector and the network architecture layer of the multi-intent recognition model, a probability distribution of each intent corresponding to the multilingual text information is generated; Based on the preset reinforcement learning model and the probability distribution of each intention, a corresponding first driving instruction is generated.

5. The method according to claim 4, characterized in that, Based on a preset reinforcement learning model and the probability distribution of each intent, a corresponding first driving instruction is generated, including: Obtain the encoding of users' historical behavioral characteristics and real-time environmental characteristics; The historical behavioral features and the real-time environmental features are encoded and input into the reinforcement learning model, and the reinforcement learning model outputs safety weights. The probability distribution of each intent is adjusted according to the security weights. Based on the adjusted intent probability distribution, the corresponding first driving instruction is generated.

6. The method according to claim 1, characterized in that, Based on the vehicle's location, local traffic rules are loaded. Using these local traffic rules and a preset multilingual road sign recognition model, the first driving instruction is modified into a second driving instruction, including: Based on the obtained vehicle location, load the preset local traffic rules; Obtain the user's historical response information to the corresponding local traffic rules; Based on the response history information and Bayesian inference, the rule priority of the local traffic rules is adjusted; Obtain the road sign image corresponding to the vehicle location and the multilingual prompt text input by the user; The road sign image and the multilingual prompt text are input into a preset multilingual road sign recognition model, and the road sign semantic classification result is output. Based on the rule priority and the road sign semantic classification result, the first driving instruction is modified into a second driving instruction.

7. The method according to claim 1, characterized in that, Based on the shared information sent by the fleet to which the vehicle belongs, the second driving instruction is modified into the target driving instruction, including: Receive a shared message that establishes a fleet communication connection with the vehicle; Based on multiple shared messages, identify the fleet decision leader who establishes fleet communication connections with the vehicle; Determine the target language required by the fleet; Send a second driving instruction, translated into the target language, to the fleet decision-making leader; Receive the group consensus instruction sent by the fleet decision leader; the group consensus instruction is the instruction with the highest decision weight determined by the fleet decision leader based on multiple candidate instructions to construct a weighted decision. Based on the aforementioned group consensus instruction, the second driving instruction is modified into the target driving instruction.

8. A cross-language driving cooperative control device, characterized in that, include: The first generation module is used to perform intent parsing on the multilingual instructions obtained from the user input and generate the first driving instruction; The first processing module is used to load local traffic rules based on the vehicle's location, and use the local traffic rules and a preset multilingual road sign recognition model to correct the first driving instruction into a second driving instruction. The second processing module is used to modify the second driving instruction into the target driving instruction based on the shared information sent by the fleet to which the vehicle belongs.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, Includes computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 7.