Vehicle control method and system based on cloud large model and vehicle situational awareness

By deeply integrating and understanding user commands and contextual data in the cloud, a strategy description is generated and then transformed into device control commands on the vehicle side. This solves the problems of semantic understanding and multimodal contextual fusion in vehicle control solutions, improving the accuracy and adaptability of vehicle control.

CN122239552APending Publication Date: 2026-06-19DONGFENG LIUZHOU MOTOR
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DONGFENG LIUZHOU MOTOR
Filing Date
2026-03-18
Publication Date
2026-06-19

Smart Images

  • Figure CN122239552A_ABST
    Figure CN122239552A_ABST
Patent Text Reader

Abstract

This application discloses a vehicle control method and system based on a cloud-based large model and in-vehicle context awareness, relating to the field of vehicle control technology. The method includes: the vehicle uploading user-generated original commands and context data snapshots to the cloud; the cloud receiving the user-generated original commands and context data snapshots uploaded by the vehicle; using the user-generated original commands and context data snapshots as input to a target policy model to obtain a policy description; the vehicle obtaining a sequence of device control commands based on a user historical preference database, preset device constraints, real-time context data, and the policy description; and controlling the operation of corresponding device actuators in the vehicle based on the sequence of device control commands. This solves the technical problems of traditional solutions where a single terminal side struggles to simultaneously handle complex semantics and multimodal contexts, and where the cloud-generated policy is disconnected from the actual execution environment of the vehicle, leading to poor control adaptability. It improves the accuracy and scenario adaptability of vehicle control in complex dynamic scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle control technology, and in particular to a vehicle control method and system based on cloud-based large model and in-vehicle context perception. Background Technology

[0002] With the rapid development of automotive intelligence and connectivity technologies, user demands for in-vehicle interaction systems have evolved from simple command execution to deep intent understanding and personalized services. In complex driving scenarios, in-vehicle systems not only need to accurately interpret users' ambiguous or polysemous natural language commands, but also need to perceive multimodal contextual information inside and outside the vehicle in real time, and integrate users' long-term usage habits to generate refined control strategies that conform to the current driving environment and equipment capability constraints, thereby achieving a truly intelligent cockpit experience.

[0003] Existing vehicle control solutions mostly employ locally preset rules or finite state machine models, relying on predefined command-action mapping tables for decision-making. This makes it difficult to handle the complex semantics implicit in natural language and personalized user preferences. While some solutions introduce cloud-based semantic parsing capabilities, they remain limited to a shallow understanding of command text and lack deep fusion modeling of multimodal vehicle contextual data (such as driving state, environmental perception, and user behavior sequences). This results in a disconnect between cloud-generated strategies and the actual execution environment on the vehicle, leading to poor adaptability and insufficient control precision in dynamic scenarios.

[0004] Therefore, how to achieve a method that can integrate the semantic understanding capabilities of large cloud models with in-vehicle multimodal contextual awareness information to dynamically generate device control command sequences that adapt to the current driving scenario and user preferences has become an urgent technical problem to be solved. Summary of the Invention

[0005] The main purpose of this application is to provide a vehicle control method and system based on cloud-based large models and in-vehicle situation awareness, aiming to solve the technical problem of how to improve the accuracy of vehicle control response to dynamic driving situations.

[0006] To achieve the above objectives, this application proposes a vehicle control method based on a cloud-based large model and in-vehicle context awareness. The vehicle control method applied to the vehicle includes:

[0007] Upload the user's original commands and a snapshot of the scenario data to the cloud, so that the cloud can use the user's original commands and the snapshot of the scenario data as input to the target policy model to obtain a policy description; Based on the user's historical preference database, preset device constraints, real-time context data, and the strategy description, a sequence of device control instructions is obtained; The corresponding device actuators in the vehicle are controlled based on the sequence of device control commands.

[0008] In one embodiment, obtaining the device control instruction sequence based on the user's historical preference database, preset device constraints, real-time context data, and the policy description includes: Parse the strategy description to obtain a list of target states and a list of constraints; The target policy function is determined based on the user's historical preference database, real-time contextual data, and the target state list; Set constraints for the target strategy function based on preset device constraints and the constraint list; Based on the constraints and the target strategy function, a sequence of device control instructions is obtained.

[0009] In one embodiment, determining the target policy function based on the user's historical preference database, real-time contextual data, and the target state list includes: Based on the target state list, determine the vector of actuators to be controlled; Based on the vector of the actuator to be controlled, query the user's historical preference database to determine the actuator weight vector; A target policy function is generated based on real-time context data, the actuator weight vector, the target state list, and the actuator vector to be controlled.

[0010] In one embodiment, after controlling the corresponding device actuator in the vehicle based on the device control command sequence, the method further includes: The monitoring results are obtained by detecting whether there is manual intervention by the user during the operation of the corresponding equipment actuators in the vehicle. Feedback data is obtained based on the user identifier, the monitoring results, and the policy description identifier corresponding to the device control command sequence; The feedback data is uploaded to the cloud so that the cloud can adjust the target strategy model based on the feedback data.

[0011] To achieve the above objectives, this application proposes a vehicle control method based on a cloud-based large model and in-vehicle context perception. The cloud-based vehicle control method based on a cloud-based large model and in-vehicle context perception includes: Receive user-generated commands and snapshots of contextual data uploaded from the vehicle terminal; The user's original instructions and the scenario data snapshot are used as inputs to the target policy model to obtain a policy description, which is then sent to the vehicle terminal so that the vehicle terminal can obtain a sequence of device control instructions based on the policy description.

[0012] In one embodiment, the user's original instruction and the scenario data snapshot are used as inputs to the target policy model to obtain a policy description, including: Based on the user's original instructions and the context data snapshot, a knowledge retrieval result is obtained by retrieving information from the knowledge graph. The knowledge retrieval results, the user's original instructions, and the scenario data snapshot are processed into text to obtain the model input text. The input text of the model is used as the input of the target policy model, so that the target policy model can reason based on the input text to obtain a policy description, wherein the policy description includes user intent labels, a list of target states, a list of constraints, and key parameters.

[0013] In one embodiment, before obtaining the policy description by taking the original user instruction and the scenario data snapshot as input to the target policy model, the method further includes: Multiple sets of training data samples and multiple sets of optimized data samples are obtained, wherein both the training data samples and the optimized data samples include user original instruction samples, scenario data snapshots and policy description samples; An initial model is trained based on the training data samples described above to obtain the strategy model to be optimized. Based on the optimized data samples and the strategy model to be optimized, a target strategy model is constructed.

[0014] In one embodiment, constructing the target policy model based on each of the optimized data samples and the policy model to be optimized includes: The user's original instruction sample and the scenario data snapshot in each of the optimized data samples are used as input to the strategy model to be optimized, so as to obtain multiple candidate strategy descriptions for each set of optimized data samples. A reward model is constructed based on the preference ranking data described in each of the candidate strategies. Based on the reward model and the strategy model to be optimized, a target strategy model is constructed.

[0015] In one embodiment, after taking the original user command and the scenario data snapshot as input to the target policy model to obtain a policy description and feeding the policy description back to the vehicle terminal so that the vehicle terminal can obtain a sequence of device control commands based on the policy description, the method further includes: Receive feedback data uploaded by the vehicle terminal; By aggregating multiple feedback data with the same user identifier, the preference adjustment direction for the corresponding user can be obtained; The feedback data with different user identifiers are clustered to obtain similar user groups; The target strategy model is optimized based on the stated preference adjustment direction and the stated similar user groups.

[0016] Furthermore, to achieve the above objectives, this application also provides a vehicle control system based on a cloud-based large model and in-vehicle context perception. The vehicle control system based on a cloud-based large model and in-vehicle context perception includes: a vehicle-side and a cloud-based system. The vehicle-side executes the vehicle control method based on a cloud-based large model and in-vehicle context perception applied to the vehicle-side as described above, and the cloud-based system executes the vehicle control method based on a cloud-based large model and in-vehicle context perception applied to the cloud-based system as described above.

[0017] This application involves the vehicle-side uploading original user commands and scenario data snapshots to the cloud. The cloud then uses these commands and snapshots as input to a target policy model to obtain a policy description. Based on a user historical preference database, preset device constraints, real-time scenario data, and the policy description, a sequence of device control commands is generated. The corresponding device actuators in the vehicle are then controlled based on this sequence. The cloud receives the original user commands and scenario data snapshots uploaded by the vehicle-side; uses these commands and snapshots as input to the target policy model to obtain a policy description; and sends the policy description to the vehicle-side, enabling the vehicle-side to generate a sequence of device control commands based on the policy description. This solves the technical problems of traditional solutions where a single terminal side struggles to simultaneously handle complex semantics and multimodal scenarios, and where the cloud-generated policy is disconnected from the vehicle's actual execution environment, leading to poor control adaptability. Compared to existing technologies, this approach achieves effective synergy between the cloud's large-scale model's deep semantic understanding capabilities and the vehicle-side's personalized preferences and real-time device constraints, improving the accuracy and scenario adaptability of vehicle control in complex dynamic scenarios. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram illustrating the process of applying the vehicle control method based on cloud-based large model and in-vehicle context perception in this application to the vehicle end. Figure 2 This is a schematic diagram illustrating the process of applying the vehicle control method based on cloud-based large model and in-vehicle context perception provided in the cloud in Implementation 1 of this application. Figure 3 This is a schematic diagram illustrating the process of applying the vehicle control method based on cloud-based large model and in-vehicle context perception in Embodiment 2 of this application to the cloud. Figure 4 This is a system architecture block diagram of the vehicle control method based on cloud-based large model and vehicle context perception provided in Embodiment 2 of this application.

[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0024] The main solution of this application embodiment is as follows: the vehicle uploads the user's original commands and a snapshot of the scenario data to the cloud, so that the cloud uses the user's original commands and the snapshot of the scenario data as input to the target policy model to obtain a policy description; based on the user's historical preference database, preset device constraints, real-time scenario data, and the policy description, a sequence of device control commands is obtained; and the corresponding device actuators in the vehicle are controlled to operate based on the sequence of device control commands. The cloud receives the user's original commands and the snapshot of the scenario data uploaded by the vehicle; uses the user's original commands and the snapshot of the scenario data as input to the target policy model to obtain a policy description, and sends the policy description to the vehicle, so that the vehicle obtains a sequence of device control commands based on the policy description.

[0025] In this embodiment, for ease of description, the following description will focus on a vehicle control system based on a cloud-based large model and in-vehicle context perception.

[0026] Existing vehicle control solutions mostly employ local preset rules or finite state machine models, relying on predefined command-action mapping tables for decision-making. This makes it difficult to handle the complex semantics implicit in natural language and personalized user preferences. While some solutions introduce cloud-based semantic parsing capabilities, they remain limited to a shallow understanding of command text and lack deep fusion modeling of multimodal vehicle contextual data (such as driving state, environmental perception, and user behavior sequences). This results in a disconnect between cloud-generated strategies and the actual execution environment on the vehicle, leading to poor adaptability and insufficient control precision in dynamic scenarios.

[0027] This application provides a solution that uploads user-generated commands and snapshots of contextual data from the vehicle to the cloud. A cloud-based large-scale model uses these as input to a target policy model to generate a policy description. The vehicle then combines this policy description with a user history preference database, real-time contextual data, and preset device constraints to transform it into a specific sequence of device control commands, ultimately controlling the corresponding actuators. This solves the technical problems of traditional solutions, where a single terminal device struggles to simultaneously handle complex semantics and multimodal scenarios, and where the disconnect between the cloud-generated policy and the actual vehicle execution environment leads to poor control adaptability. Compared to existing technologies, this solution achieves effective synergy between the deep semantic understanding capabilities of the cloud-based large-scale model and the vehicle's personalized preferences and real-time device constraints, improving the accuracy and adaptability of vehicle control in complex dynamic scenarios.

[0028] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, such as a vehicle control system based on a cloud-based large model and in-vehicle context perception. The following description uses a vehicle control system based on a cloud-based large model and in-vehicle context perception as an example to illustrate this embodiment and the subsequent embodiments.

[0029] Based on this, embodiments of this application provide a vehicle control method based on a cloud-based large model and in-vehicle context perception, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the vehicle control method based on cloud-based large model and in-vehicle context awareness applied to the vehicle.

[0030] In this embodiment, the vehicle control method based on cloud-based large model and vehicle context awareness is applied to the vehicle end. The vehicle control method based on cloud-based large model and vehicle context awareness includes steps S10~S30: Step S10: Upload the user's original instructions and scenario data snapshot to the cloud, so that the cloud can use the user's original instructions and scenario data snapshot as input to the target policy model to obtain a policy description.

[0031] It should be noted that the original user command refers to the original interactive command input by the user inside the vehicle, such as voice collected by the microphone array, gesture captured by the in-vehicle camera, and touch operation of the touch screen / physical button; the contextual data snapshot refers to the lightweight structured data package of real-time vehicle status, environmental information, and user status collected by the in-vehicle multimodal contextual perception module at the moment the user issues the command, including but not limited to key contextual information such as vehicle speed, gear position, geographical location, in-vehicle temperature, ambient light, user identity, and fatigue level; the strategy description refers to the high-level abstract intent and goal expressed in a structured domain-specific language, generated by the cloud-based large model after fusing and understanding the user command and contextual data, such as including core intent labels, a list of desired goal states, and related constraints. The goal-policy model refers to the multimodal large language model (MLLM) deployed in the cloud and specially trained. Its core function is to fuse and understand the original user command and contextual data snapshot and perform deep reasoning to ultimately generate a structured abstract strategy description.

[0032] Specifically, when a user issues a command inside the vehicle, the vehicle-side system converts the audio stream into raw command text T via its local ASR module. Simultaneously, the multimodal context awareness module is triggered, collecting key contextual information from various sensors and the vehicle bus. After filtering and anonymization, this information is packaged into a structured contextual snapshot C, containing vehicle status (e.g., speed, gear, location), user status (e.g., user ID, detected fatigue level), and environmental information (e.g., in-vehicle temperature, ambient light, time). The vehicle-side system encrypts the {T, C} data packet and securely uploads it to the cloud service subsystem via the vehicle network, awaiting further understanding and generation by the cloud-based large model.

[0033] For example, during the morning rush hour on a weekday, user Alice gets into the car, fastens her seatbelt, and says, "Okay, perk up, let's get started!" The vehicle is on a congested urban road. At this time, the data collected by the vehicle is instruction T and scenario C (low vehicle speed, traffic congestion, morning time, user ID identified as Alice, DMS detects slight drowsiness).

[0034] The cloud-based big model combines knowledge (morning rush hour = congestion, high pressure; "energize" = refresh the mind and improve efficiency; "combat" = may refer to efficient commuting) to generate an abstract strategy description S1: {intent: "efficient commuting and energy-boosting mode", goals: [improve driving focus, provide efficient navigation information, adjust the atmosphere to be invigorating], constraints: [ensure absolute driving safety and do not affect vehicles behind]}.

[0035] Understandably, user natural language commands are often ambiguous, metaphorical, and polysemous, and the same command should trigger completely different responses in different contexts. If only the limited rule model on the vehicle's local end is used for parsing, it will lead to misjudgment of intent and inaccurate service. Therefore, step S10, which synchronously uploads the original command text and a contextual snapshot reflecting the complete context to a cloud-based large model for fusion and understanding, can avoid policy generation deviations caused by missing contextual information or insufficient semantic understanding capabilities, thereby improving the accuracy of command parsing and the contextual adaptability of policy generation.

[0036] Step S20: Based on the user's historical preference database, preset device constraints, real-time context data, and the strategy description, a sequence of device control instructions is obtained.

[0037] It should be noted that the user history preference database refers to a collection of user history adjustment records stored locally in the vehicle, indexed by user ID, reflecting the personalized setting habits formed by users over a long period of use. For example, for seat position, it stores the final position parameters manually adjusted by the user in different scenarios (driving, resting, watching a movie); for air conditioning temperature, it stores the user's preference setting curves under different outdoor temperatures. Preset equipment constraints refer to the set of constraints describing the capability model of all controllable actuators in the vehicle, i.e., the vehicle equipment capability library, including the adjustable dimensions of each actuator (such as seat fore-aft, height, pitch), the range of each dimension (0-100), accuracy (step value), adjustment speed, and other physical limitations. Preset equipment may also include safety hard constraints, which are mandatory restrictions related to driving safety and occupant safety that must be unconditionally followed under any circumstances. These constraints have the highest priority and can override or modify any output of the strategy optimizer. Real-time contextual data refers to the vehicle status data stream obtained from the multimodal contextual perception module at the current moment, which is more real-time and richer than the uploaded snapshot, including but not limited to dynamic information such as real-time vehicle speed, gear position, in-vehicle temperature, ambient light, user identity, and fatigue level. The device control instruction sequence refers to the list of specific executable instructions generated by the local policy optimizer and arranged according to timing and dependencies. Each instruction includes the executor identifier, action type, and parameters.

[0038] For example, the vehicle-mounted device receives the strategy description S1 sent from the cloud and combines it with user preferences P_Alice (e.g., the user prefers a cooler air conditioner, listening to podcasts and news, and displaying detailed navigation on the HUD), preset device constraints D_cap, and real-time contextual data C_local (severe traffic congestion) to generate a personalized sequence of device control commands. For example, the device may lower the air conditioner temperature by 2°C (from preference) and turn on "Fresh Air Mode," play Alice's favorite "Morning Energy" playlist at a moderate volume, highlight "estimated remaining congestion time" and "next feasible detour," and simultaneously display real-time traffic conditions on the central control screen in a split-screen format; adjust the ambient lighting on the driver's side to a cool tone and slow breathing mode; and ensure that all entertainment information displays do not obstruct the driver's view, and that podcast content is not selected to be overly engaging or story-based.

[0039] Understandably, since the policy descriptions delivered from the cloud are device-independent abstract goals (such as "minimizing light intensity interference"), they do not include specific device control parameters. Furthermore, different users have drastically different personalized preferences for the same abstract goal (e.g., user A prefers a completely dark environment when taking a nap, while user B prefers to retain a faint warm light). Simultaneously, the vehicle's real-time status (such as current speed) imposes safety constraints on certain adjustment actions (e.g., prohibiting fully reclining the seat while driving at high speeds). Directly executing fixed actions based on abstract goals would lead to a lack of personalization or safety hazards. Therefore, step S20 integrates user historical preferences, device physical constraints, hard safety constraints, and real-time contextual data on the in-vehicle side, transforming the abstract policy description into a specific sequence of device control commands. This avoids poor user experience or security risks caused by a lack of personalized adaptation and real-time security verification, thereby improving the personalization accuracy and security of policy execution.

[0040] In one feasible implementation, step S20 may include: parsing the strategy description to obtain a target state list and a constraint list; determining a target strategy function based on a user historical preference database, real-time context data, and the target state list; setting constraints for the target strategy function based on preset device constraints and the constraint list; and obtaining a device control instruction sequence based on the constraints and the target strategy function.

[0041] It should be noted that the target state list refers to the set of desired states parsed from the abstract strategy description issued from the cloud; the constraint list refers to the set of constraints carried in the strategy description; and the target strategy function refers to the satisfaction function constructed for each target state, used to quantify the degree to which the current device parameter configuration satisfies the target state. The closer the function value is to 1, the more satisfied the state is. Constraints refer to the mathematical or logical constraint expressions that transform strategy constraints and device physical limitations into, used to limit the search space of feasible solutions.

[0042] Specifically, the vehicle parses the JSON-formatted strategy description from the cloud, extracting the target state list `goal_states` and the constraint list `constraints`. For each item in `goal_states`, it maps it locally, for example: Regarding goal_states[0](light: intensity -> minimum interference), Query device capability constraints D_cap: This vehicle has "ceiling reading light", "foot ambient light", and "door panel ambient light".

[0043] Querying the user's historical preferences database P_user: user_123 previously manually turned off all overhead lights and left only the foot ambient light at 5% brightness during a "short nap".

[0044] Based on real-time contextual data C_local: the current ambient light is "low".

[0045] Based on the above data, control commands can be generated. At the same time, the target states in the control commands need to be jointly optimized. The joint optimization is accomplished by using the target policy function. For example, goal_states[2] (seat: posture → recline and relax) may conflict with constraints[0] (safety: can quickly restore control). The optimizer needs to find a balance point: adjust the seat to "memory position 1" (this position is the compromise between the driver's seat and the reclining posture set by the user), rather than fully reclining.

[0046] After optimizing the control commands, it is also necessary to consider the dependencies and timing between device actions and generate a sequence of device control commands. For example, "close the sunroof sunshade" should be executed before "adjust the seat to a reclining position" to prevent the user from being stuck by the moving seat.

[0047] Furthermore, determining the target policy function based on the user's historical preference database, real-time context data, and the target state list includes: determining the actuator vector to be controlled based on the target state list; querying the user's historical preference database based on the actuator vector to be controlled to determine the actuator weight vector; and generating the target policy function based on the real-time context data, the actuator weight vector, the target state list, and the actuator vector to be controlled.

[0048] It should be noted that the actuator vector to be controlled refers to the vector representation of all actuators that may need to be invoked to achieve the current target state list. For example, for the target of "minimizing light interference," the actuators involved may include ceiling reading lights, foot ambient lights, and door ambient lights. The actuator weight vector refers to the weight coefficient vector obtained from the user's historical preference database, reflecting the user's preference for different actuators in achieving the current target. For example, if a user tends to turn off all ceiling lights while leaving the foot ambient light dimly lit when taking a nap, then the ceiling light weight is 0 and the foot light weight is 1.

[0049] Specifically, all controllable device parameters are defined as a vector. ,For example This represents the seat back angle. Representing the target temperature of the air conditioner, a satisfaction function is defined based on each actuator parameter in the target state list. Where x represents the device parameter that needs to be adjusted for the actuator to be controlled. P Represents user's historical preferences, C l This represents real-time context data; the closer the function value is to 1, the more satisfied the condition is.

[0050] The optimization objective of the satisfaction function is to maximize the weighted total satisfaction. Therefore, the objective policy function is set as follows: as follows:

[0051] Among them, weight It represents the weight of each target state in the target state list, which can be queried from the user's historical preference database and is determined by the user's preferences.

[0052] It is also necessary to set constraints for the target policy function. Preset device constraints include physical constraints and hard safety constraints. Based on the physical constraints, the range of values ​​for x is set as follows: Based on hard safety constraints, x must meet strict safety conditions, for example, when the first target constraint... When the vehicle speed is less than 5 km / h, the corresponding other objective constraint is required. For the seat angle, the following is required: It must be greater than 20°.

[0053] In this embodiment, by introducing executor weight vectors, user historical preferences are quantified into computable objective function weights, which solves the problem of lack of personalized guidance in the mapping between abstract objectives and specific executors. This achieves accurate embedding of user personalized habits in the process of constructing the objective policy function, thereby improving the matching degree between the policy generation results and the user's true expectations.

[0054] The above are merely feasible implementations of step S20 provided in this embodiment. This embodiment does not specifically limit the specific implementation of step S20.

[0055] Step S30: Control the operation of the corresponding device actuator in the vehicle based on the device control command sequence.

[0056] It should be noted that an actuator refers to a physical or virtual device unit in a vehicle that can be controlled to perform specific actions. Examples include seat motors that support multi-directional electric adjustment, damper actuators for multi-zone independent air conditioning, LED ambient lighting with zoned dimming and color adjustment, multi-speaker audio systems, electric sunroof and sunshade motors, and central control rotating screen drive mechanisms. These actuators receive control commands via vehicle buses (such as CAN, LIN, and Ethernet) and convert them into actual physical actions.

[0057] Understandably, since the device control command sequence has undergone multi-objective optimization by the local policy optimizer and final safety checks by the safety and conflict arbitrator, it contains precise actuator identification, action type, parameter values, and timing dependencies. If the command sequence is generated but not executed, or if execution is not scheduled in sequence, the policy intent will fail or device actions will conflict. Therefore, step S30, which sends the optimized command sequence sequentially to each actuator via the vehicle's standard protocol, avoids policy failures caused by missing execution steps or timing errors, thus ensuring that the abstract policy description is ultimately translated into a user-perceptible actual cockpit experience.

[0058] In one feasible implementation, step S30 may include: monitoring whether there is manual intervention by the user during the operation of the corresponding device actuator in the vehicle, and obtaining monitoring results; obtaining feedback data based on the user identifier, the monitoring results, and the policy description identifier corresponding to the device control command sequence; and uploading the feedback data to the cloud so that the cloud can adjust the target policy model based on the feedback data.

[0059] It should be noted that manual intervention refers to the user's active manual adjustment of a device actuator after the system automatically executes the policy. For example, a user manually increases the brightness of an ambient light from the system-set 5% to 20%, or manually lowers the air conditioner temperature by 2°C. Monitoring results refer to the detection results of user intervention behavior by the feedback recording module, including whether intervention occurred, the device involved, the parameter value after intervention, and the timestamp of the intervention. User identifier refers to the identity information used to uniquely identify the current user, such as the user ID determined by facial recognition or account login. Policy description identifier refers to the unique ID of the policy description executed this time, used to associate feedback data with the specific policy generated in the cloud. Feedback data refers to the structured data record formed after associating user intervention behavior with the original policy, including policy ID, user ID, abstract goal, execution action, user overriding action, and contextual information.

[0060] Specifically, after the sequence of device control commands is executed, the feedback recording module initiates a monitoring process. This module continuously listens for device status change events on the vehicle bus and compares them with the most recently executed policy commands. For example, suppose the system sets the footwell ambient light to 5% brightness according to the "rest mode" policy. After 15 minutes, the user manually adjusts the brightness to 20% via a physical knob or touchscreen. The feedback recording module captures this device status change event, identifies that the change was not triggered by a system policy command, and determines it as a manual intervention. The module then collects relevant information about this intervention: obtaining the user identifier (e.g., "user_123") from the current session, obtaining the policy description identifier (e.g., "strategy_vehicle control based on cloud-based large model and vehicle context awareness") from the execution context, extracting the intervened device ("footwell_lights") and the user-adjusted value (20%) from the intervention event, and correlating it with the original executed value (5%). This information is then encapsulated into feedback data F in JSON format.

[0061] For example, the structure of the feedback data F is as follows: { "strategy_id": "Vehicle control based on cloud-based large-scale models and in-vehicle context awareness", "user_id": "user_123", "abstract_goal": {"domain": "lighting", "target": "minimal_intrusive"}, "executed_action": {"device": "footwell_lights", "value": 5}, "user_override_action": {"device": "footwell_lights", "value": 20}, "context": {"time_since_start": 900} }

[0062] The vehicle-mounted device temporarily stores this type of feedback data locally, and then uploads it to the cloud in batches with encryption when network conditions are suitable (such as when connected to WiFi).

[0063] In this implementation, by introducing a closed-loop mechanism for monitoring and feedback of user manual intervention behavior, the problem of the inability to self-correct the possible deviation between the system strategy and the user's actual preferences is solved. This enables the system to continuously evolve by "understanding you better the more you use it", thereby continuously improving the personalized accuracy of strategy generation and user satisfaction.

[0064] The above are merely feasible implementations of step S30 provided in this embodiment. This embodiment does not specifically limit the specific implementation of step S30.

[0065] This embodiment provides a vehicle control method based on a cloud-based large-scale model and in-vehicle context awareness. The vehicle uploads the user's original command text and a snapshot of the context data to the cloud. The cloud then uses these as input to a target policy model to obtain a policy description. Based on a user historical preference database, preset device constraints, real-time context data, and the policy description, a sequence of device control commands is generated. The corresponding device actuators in the vehicle are then controlled based on this sequence of commands. This method solves the technical problems of existing solutions where a single terminal side struggles to simultaneously handle complex semantic understanding and multimodal context fusion, and where the disconnect between cloud-generated policies and the actual execution environment on the vehicle leads to rigid control decisions and a lack of personalized adaptation. It avoids misjudgment of intent due to insufficient semantic understanding, poor user experience due to lack of personalized adaptation, and execution risks due to failure to consider real-time safety constraints. This achieves effective synergy between the deep semantic understanding capabilities of the cloud-based large-scale model and the personalized preferences, real-time device constraints, and safety rules on the vehicle side, improving the accuracy, safety, and personalized experience of vehicle control in complex dynamic scenarios.

[0066] Based on this, embodiments of this application provide a vehicle control method based on a cloud-based large model and in-vehicle context perception, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the vehicle control method based on cloud-based large model and in-vehicle context perception applied to the cloud.

[0067] In this embodiment, the vehicle control method based on cloud-based large model and vehicle context perception is applied to the cloud, and the vehicle control method based on cloud-based large model and vehicle context perception includes steps S10'~S20': Step S10': Receive the user's original instructions and scenario data snapshot uploaded by the vehicle terminal.

[0068] Specifically, the cloud continuously monitors secure connection requests from a massive number of vehicles. Upon receiving a {T, C} data packet uploaded by a vehicle, it performs protocol parsing, integrity verification, and identity authentication. The cloud extracts the original instruction text string and contextual data snapshot object from the network load and temporarily stores them in the session context of this request. The cloud preprocesses these two types of heterogeneous data: the text portion is normalized, and the key-value pairs in the contextual data snapshot are converted into a unified vector representation or flattened text description. These two are then combined into an input format acceptable to the target policy model, ready to be fed into the target policy model for inference. This processing flow ensures that heterogeneous data from different vehicle models and different sensor configurations can be uniformly understood and processed by the model.

[0069] Understandably, the original instruction text uploaded by the vehicle differs significantly from the contextual data snapshot in terms of data format, semantic space, and units. Furthermore, large models require strictly structured input formats. Without unified preprocessing and format conversion, the model may fail to effectively fuse the two types of information or produce input format errors. Therefore, step S10', through a unified data receiving, verification, and preprocessing process, avoids model inference failures caused by input format mismatches or data quality defects. This provides high-quality input data for subsequent cross-modal semantic understanding, ensuring the accuracy and reliability of policy generation.

[0070] Step S20': The user's original instruction and the scenario data snapshot are used as inputs to the target policy model to obtain a policy description, and the policy description is sent to the vehicle terminal so that the vehicle terminal can obtain a sequence of device control instructions based on the policy description.

[0071] Understandably, user commands are often ambiguous, metaphorical, and polysemous (e.g., "Fight!" might refer to efficient commuting rather than its literal meaning), and the same command should trigger completely different responses in different contexts. Relying solely on simple keyword matching or local rules for parsing would lead to misjudgment of intent and inaccurate service. Therefore, step S20' involves inputting the user's original command and contextual data snapshots into a cloud-based big data model for fusion understanding and deep reasoning, generating a structured abstract strategy description which is then sent to the vehicle. This avoids policy generation biases caused by insufficient semantic understanding or missing contextual information, thereby improving the accuracy of command parsing and the rationality of policy generation. Simultaneously, by outputting a device-independent abstract description, it leaves room for personalized adaptation on the vehicle side.

[0072] In one feasible implementation, step S20' may include: retrieving knowledge from the knowledge graph based on the user's original instruction and the context data snapshot to obtain knowledge retrieval results; textualizing the knowledge retrieval results, the user's original instruction, and the context data snapshot to obtain model input text; and using the model input text as input to the target policy model so that the target policy model can reason based on the model input text to obtain a policy description, wherein the policy description includes user intent tags, a target state list, a constraint list, and key parameters.

[0073] It should be noted that a knowledge graph refers to a structured knowledge base that stores entities, attributes, and relationships related to car-related life, such as triples like "Sunset - Direction: West - Best viewing time: Dusk - Associated mood: Romantic and peaceful", "Driving in the rain - Need to improve visibility - Recommended to turn on wipers / fog lights - Associated risk: Sideslip", and "Infant - Need stable ambient temperature - Sensitive to noise - Associated scenario: Child sleep mode". Knowledge retrieval results refer to the set of relevant entities and relationships retrieved from the knowledge graph based on the user's original instructions and contextual data snapshots, used to enhance the model's common-sense understanding of the scenario. Model input text refers to a text sequence that conforms to the input format of a large model, formed by textually concatenating the knowledge retrieval results, the user's original instructions, and the contextual data snapshots. User intent tags refer to the model's abstract summary of the user's core intent, such as "In-car nap mode", "Efficient commuting refreshment mode", and "Child pick-up and drop-off care mode". The target state list refers to the set of multi-dimensional target states that are expected to be achieved. The constraint list refers to the restrictions that need to be followed during the execution of the strategy. Key parameters refer to supplementary parameters related to the current strategy.

[0074] Specifically, after receiving the user's original command T and contextual data snapshot C uploaded from the vehicle, the cloud-based big data model first triggers the knowledge retrieval module. This module uses key information from T and C (such as location "home_garage", time "night", user status "fatigued") as query conditions to retrieve relevant entities and relationships from common sense and scenario knowledge graphs. For example, for the command "I'm tired and want to relax," combined with "night, garage, P gear, high fatigue" in C, the knowledge graph might return relevant knowledge triples such as "take a nap - need to lower the light - need to maintain air circulation - need a comfortable posture."

[0075] The cloud platform processes the knowledge retrieval results K, the user's original instructions T, and the contextual data snapshot C into text: it converts the structured contextual data snapshot into a natural language description (e.g., "The vehicle is in P gear and stationary, located in the home garage, currently at night, the interior temperature is 24℃, the ambient light is dim, and the user is highly fatigued"), and converts the knowledge retrieval results into knowledge prompt text (e.g., "Common sense tip: In a nap scenario, it is recommended to reduce the light intensity, maintain gentle air circulation, and adjust the seat to a relaxed posture"), and concatenates it with the original instructions to form a complete model input text Prompt.

[0076] After receiving the Prompt, the target policy model (a multimodal large language model) performs autoregressive inference, generating a structured abstract policy description S_abstract for each token. During inference, the model integrates instruction semantics, contextual information, and common-sense knowledge, outputting a JSON structure containing the intent label, the goal state list (goal_states), the constraint list (constraints), and the key parameters. The cloud then distributes this policy description to the corresponding in-vehicle terminal for subsequent personalized adaptation and instantiation.

[0077] For example, the structure of the strategy description S_abstract is as follows: { "intent": "In-car nap mode", "goal_states": [ {"domain": "lighting", "target": "minimal_intrusive", "attribute": "intensity"}, {"domain": "climate", "target": "gentle_airflow", "attribute": "air_mode"}, {"domain": "seat", "target": "reclinated_relax", "attribute": "posture"}, {"domain": "acoustic", "target": "mask_ambient_noise", "attribute": "soundscape"} ], "constraints": [ {"type": "safety", "condition": "user_can_quickly_regain_control"}, {"type": "duration", "condition": "approx_30_min"} ], "parameters": { "user_state": "fatigued" } }

[0078] In this embodiment, by introducing Knowledge Graph Retrieval Enhanced Generation (RAG) technology, external common sense knowledge is integrated into the reasoning process of the large model, which solves the problem of common sense gaps or insufficient scenario understanding that large models may have in vertical domains. This makes the generated strategy descriptions more in line with best practices and common sense cognition in car use. At the same time, through the structured strategy description output, a clear and parsable abstract target framework is provided for personalized adaptation on the vehicle side, thereby improving the rationality, interpretability and cross-vehicle adaptation capability of strategy generation.

[0079] The above are merely feasible implementations of step S20' provided in this embodiment. This embodiment does not specifically limit the specific implementation of step S20'.

[0080] In one feasible implementation, step S20' may include: receiving feedback data uploaded by the vehicle; aggregating multiple feedback data with the same user identifier to obtain the corresponding user's preference adjustment direction; clustering feedback data with different user identifiers to obtain similar user groups; and optimizing the target strategy model based on the preference adjustment direction and the similar user groups.

[0081] It should be noted that the preference adjustment direction refers to the quantitative direction reflecting the evolution trend of a user's personalized preferences, derived from the aggregation and analysis of the user's historical feedback data. For example, "users tend to prefer brighter lights than the default policy when taking a nap" or "users prefer lower air conditioning temperatures when commuting." Similar user groups refer to user groups formed by dividing different users with similar preference characteristics through clustering algorithms. For example, "young female users who like warm ambient lighting" or "male users who prefer strong air conditioning."

[0082] Specifically, the cloud-based model continuous learning and optimization module continuously receives feedback data F from massive amounts of anonymized vehicle uploads. Each piece of feedback data F includes a user identifier (user_id), a strategy description identifier (strategy_id), an abstract goal (abstract_goal), an executed action (executed_action), a user override action (user_override_action), and context information.

[0083] For example, the module aggregates and analyzes feedback data according to user identifiers. Statistical analysis of feedback data generated by the same user (e.g., user_123) during multiple "rest mode" executions reveals that this user manually increases the brightness of the foot ambient light from the system default of 5% to the 15%-20% range in 80% of cases. This statistical pattern is quantified as the user's preferred adjustment direction: "user_123's preference for light brightness in rest scenarios is 10-15 percentage points higher than the general strategy."

[0084] The module performs cluster analysis on feedback data from different users. Using a clustering algorithm, users are grouped based on their preference feature vectors in different scenarios. For example, users who tend to turn up the lights and turn down the air conditioning speed in "Rest Mode" are grouped together to form a user group that "prefers a gentle resting environment"; users who tend to turn down the air conditioning and play fast-paced music in "Commuting Mode" are grouped together to form a user group that "prefers an energetic commute".

[0085] For specific users (e.g., user_123), adjustments to their explicit preferences can be made during model inference by injecting personalized prompts. This allows the generated strategy descriptions to better align with the user's individual needs. For new users or users with sparse feedback data, cold-start optimization can be performed using the preference characteristics of similar user groups. For example, when the system identifies that the profile characteristics of new user_456 (e.g., age, gender, region) highly match those of the "preferring a gentle resting environment" group, the model output can be adjusted based on the common preferences of this group, making its initial strategy more in line with user expectations.

[0086] In this implementation, by establishing a complete closed-loop mechanism from individual feedback to group clustering and then to model optimization, the problems of the system's personalization capabilities being limited to the historical data of a single user and the poor cold start effect for new users are solved. This achieves multi-level optimization of "precise individual adaptation, commonality mining of the group, and continuous global evolution", thereby continuously improving the adaptability and accuracy of the target strategy model for different users and different scenarios.

[0087] The above are merely feasible implementations of step S20' provided in this embodiment. This embodiment does not specifically limit the specific implementation of step S20'.

[0088] This embodiment provides a vehicle control method based on a cloud-based large model and in-vehicle context awareness. The cloud receives original user command text and a snapshot of context data uploaded from the vehicle. The original user command text and the snapshot of context data are used as input to a target policy model to obtain a policy description, which is then sent to the vehicle. Based on the policy description, the vehicle generates a sequence of device control commands. This method solves the technical problems in existing solutions where the cloud can only process single-modal commands and lacks a deep understanding of multimodal vehicle context data, leading to a disconnect between policy generation and real-world scenarios and an inability to provide effective guidance for personalized adaptation on the in-vehicle side. It avoids misjudgment of intent due to insufficient semantic understanding and policy deviation due to missing context information. The method achieves cross-modal fusion understanding of ambiguous user commands and real-time context through a cloud-based large model, and outputs a structured, device-independent abstract policy description, providing a precise and universal framework for personalized execution on the in-vehicle side.

[0089] Based on the first embodiment of this application applied to the cloud, in the second embodiment of this application applied to the cloud, the content that is the same as or similar to the first embodiment applied to the cloud described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the vehicle control method based on cloud-based large model and in-vehicle context perception applied to the cloud.

[0090] In the second embodiment, the vehicle control method based on cloud-based large model and vehicle context awareness is applied to the cloud. Before step S20', the vehicle control method based on cloud-based large model and vehicle context awareness further includes steps S17'~S19': Step S17': Obtain multiple sets of training data samples and multiple sets of optimized data samples, wherein both the training data samples and the optimized data samples include user original instruction samples, scenario data snapshots, and policy description samples.

[0091] It should be noted that the training data samples refer to the manually labeled dataset used in the supervised fine-tuning stage. Each sample contains a user's original instruction T (such as "I'm tired and want to relax"), a corresponding snapshot of the contextual data C (such as a structured description of vehicle status, environmental information, and user status), and a manually labeled standard policy description S_abstract (such as a JSON structure containing intent, goal_states, and constraints). The optimization data samples refer to the dataset used in the reinforcement learning stage based on human feedback. Each sample also contains {T, C} pairs, but does not contain a standard answer. Instead, it is used to generate multiple candidate policies for human preference ranking.

[0092] Specifically, cloud-based model training platforms construct large-scale training datasets through manual or semi-automatic generation. For example, testers are invited to simulate various driving scenarios, record user commands, and collect corresponding contextual data, which is then annotated by expert annotators for each command.<T, C> The optimal strategy description S_abstract, which aligns with the intended annotations, is used to create a supervised fine-tuning dataset ranging from tens of thousands to hundreds of thousands of records. Simultaneously, an optimization dataset is constructed: a portion of the dataset is selected.<T, C> Yes, multiple candidate policy descriptions are generated from the model to be optimized, and then these candidate policies are manually ranked by preference to form a preference ranking dataset for subsequent reward model training.

[0093] Understandably, the target policy model needs to generate high-quality policy descriptions from user instructions and contextual data, a capability that cannot be achieved solely through rule writing but must rely on learning from large-scale, high-quality data. Therefore, step S17' is performed to obtain data containing...<T, C, S_abstract> Training data samples of triples and containing<T, C> Optimizing data samples for preference ranking can avoid model capability defects caused by insufficient or low-quality training data, thereby providing a data foundation for subsequent model training and improving the accuracy and rationality of the strategy description generated by the model.

[0094] Step S18': Train the initial model based on each of the training data samples to obtain the strategy model to be optimized.

[0095] It should be noted that the initial model refers to the basic large language model that has been pre-trained on a large-scale general corpus, possessing basic language understanding and generation capabilities, but has not yet been optimized for the automotive vertical domain; the policy model to be optimized refers to the model that has been supervised fine-tuned, and has learned to generate policy descriptions that meet the format requirements based on user instructions and contextual data, but has not yet been optimized for human preferences.

[0096] Specifically, the cloud-based training platform will use the training data samples obtained in step S17'<T, C, S_abstract> Batch processing is performed, concatenating T and C as the model input and using S_abstract as the target output. Model training employs the standard cross-entropy loss function, calculating the difference between the model-generated sequence and the target labeled sequence at each token position, and updating the model parameters through backpropagation. Loss function... The format is:

[0097] in, It is the length of the target sequence. It is the first One token, This represents all tokens prior to the i-th token. A token represents a basic semantic unit in natural language processing and large language models. It can be understood as the smallest unit of granularity when the model processes text. For example, the sentence "I am tired" may be split into two tokens: ["I", "tired"].

[0098] T represents the original user instruction sample in the training data sample, C represents the scenario data snapshot in the training data sample, and KG represents the policy description sample in the training data sample.

[0099] Understandably, since the initial model was only pre-trained on a general corpus and lacks expertise in understanding vehicle-specific commands, context fusion, and policy generation, directly applying it to real-world scenarios would result in low-quality and non-standardized generated policies. Therefore, step S18' involves supervised fine-tuning of the initial model based on high-quality, manually annotated training data samples. This avoids policy generation biases caused by a lack of domain knowledge, thereby improving the model's adaptability to vehicle-specific scenarios and the structured standardization of generated policies.

[0100] Step S19': Based on the optimized data samples and the strategy model to be optimized, construct the target strategy model.

[0101] It should be noted that the target policy model refers to a multimodal large language model that is finally obtained after supervised fine-tuning and reinforcement learning based on human feedback, and is capable of generating high-quality policy descriptions. It achieves optimal performance in terms of intent understanding accuracy, policy rationality, and user preference matching.

[0102] Specifically, it involves collecting user preference ranking data for multiple policy descriptions generated by the model, and training a reward model based on this preference ranking data. This is used to predict user preference scores; the cloud training platform utilizes optimized data samples and reward models, and further optimizes the policy model to be optimized through reinforcement learning algorithms. .

[0103] Understandably, since the supervised fine-tuning phase only learns and imitates the standard answers labeled by humans, and the standard answers may not be the only optimal solution and are difficult to cover all user preferences, directly using the policy model to be optimized may result in a policy that, while formatted correctly, does not conform to the actual user preferences. Therefore, step S19' introduces reinforcement learning based on human feedback, which avoids the model merely mechanically imitating labeled data without a deep understanding of user preferences, thereby improving the matching degree between the policy description and user expectations and achieving continuous evolutionary capability.

[0104] In one feasible implementation, step S19' may include: using the user's original instruction sample and the scenario data snapshot in each of the optimized data samples as input to the strategy model to be optimized, to obtain multiple candidate strategy descriptions for each set of optimized data samples; constructing a reward model based on the preference ranking data of each of the candidate strategy descriptions; and constructing a target strategy model based on the reward model and the strategy model to be optimized.

[0105] It should be noted that the candidate strategy description refers to the description of the same group.<T, C> The input consists of multiple differentiated policy descriptions generated by the policy model to be optimized through different random sampling temperatures or decoding strategies; preference ranking data refers to the results of human annotators ranking multiple candidate policy descriptions from high to low quality, such as "policy A is better than policy B is better than policy C"; the reward model is an auxiliary model trained based on the preference ranking data, whose input is...<T, C, S> The triplet outputs a predicted human preference score, which is used to guide the optimization direction of the policy model during the reinforcement learning phase.

[0106] Specifically, for each of the optimization datasets<T, C> Sample, strategy model to be optimized N candidate policy descriptions S1...SN are generated by setting different decoding parameters. These candidate policy descriptions are then randomly sorted and presented to human annotators. The annotators rank the policies according to their preferences, considering factors such as rationality, completeness, and matching degree with instructions and context, forming preference ranking data. The training platform uses this preference ranking data to train a reward model. After the reward model is trained, it is used as a reward signal in the reinforcement learning phase to guide the policy model to be optimized. The target policy model is ultimately generated through optimization using algorithms such as PPO.

[0107] For example, the target strategy model The expression is:

[0108] in, This represents the mathematical expression of expectation. Indicates from dataset D Randomly sample a user command text T and corresponding context snapshots C, Indicates based on the current strategy model π θ (i.e., the policy model to be optimized) generates an abstract policy description. S, Overall Expectations E It is in all possible ( T , C Sample and model generation S The average value is calculated, which means calculating the expected value of the expression within the parentheses (reward score minus KL divergence penalty); It is the KL divergence penalty coefficient that controls the degree of deviation from the original model. Indicates the initial model. This represents the strategy model to be optimized.

[0109] In this implementation, by introducing a reward model based on human preferences and reinforcement learning optimization, the problem that the model in the supervised fine-tuning stage only imitates labeled data and cannot understand the deep preferences of users and the differences in policy quality is solved. This enables the target policy model to generate high-quality policy descriptions that are more in line with human expectations and more in line with real-world scenarios, thereby improving user experience and policy acceptance.

[0110] The above are merely feasible implementations of step S19' provided in this embodiment. This embodiment does not specifically limit the specific implementation of step S19'.

[0111] This embodiment provides a vehicle control method based on a cloud-based large model and in-vehicle context awareness. It acquires multiple sets of training data samples and multiple sets of optimized data samples, wherein both the training data samples and the optimized data samples include user original command text samples, context data snapshots, and policy description samples. An initial model is trained based on each of the training data samples to obtain a policy model to be optimized. A target policy model is constructed based on each of the optimized data samples and the policy model to be optimized. This avoids the limitations of model capabilities caused by a single training method and solves the problems of insufficient adaptability of the model in vertical domains and difficulty in aligning with the depth of user preferences. Thus, it achieves comprehensive optimization of the target policy model in terms of intent understanding accuracy, policy generation rationality, and user preference matching.

[0112] For example, to help understand the implementation process of the vehicle control method based on cloud-based large model and vehicle context awareness obtained by combining this embodiment with the above embodiment one, please refer to... Figure 4 , Figure 4A system architecture diagram is provided for a vehicle control method based on a cloud-based large model and in-vehicle context perception. Specifically: The overall workflow begins in the vehicle. After capturing the user's original interaction command (such as "I'm tired, I want to relax"), the vehicle's multimodal context awareness module simultaneously collects real-time data on vehicle status, environmental information, and user status, generating a structured "contextual data snapshot." This snapshot, along with the "original user command text" converted from local speech recognition, is encrypted and securely uploaded to the cloud service subsystem via the vehicle-to-everything (V2X) network. This step corresponds to the process in the solution where the vehicle collects and uploads basic information for cloud-based understanding.

[0113] Upon receiving the data, the cloud decrypts and parses it. The intent understanding and abstract strategy generation model (target strategy model) fuses the input command text with the context snapshot for understanding, and selectively retrieves relevant knowledge from common sense and scenario knowledge graphs to enhance reasoning capabilities, ultimately generating a structured "abstract strategy description." This description is defined in a domain-specific language and includes core intent tags, a list of desired target states, and related constraints, but it does not involve any specific device control details. The generated abstract strategy description is encrypted and sent to the requesting vehicle terminal, thus completing the cloud's deep understanding of the user's complex intent and high-level strategy planning.

[0114] After receiving the policy description from the cloud, the policy conversion and executor module begins operation. The parser within this module first interprets the abstract policy, clarifying its objectives and constraints. The local policy optimizer integrates the user's historical preferences from the local personalized database, real-time contextual data provided by the multimodal contextual awareness module, and preset vehicle equipment capability constraints (including physical limitations and hard safety constraints). Through a multi-objective optimization algorithm, it transforms the abstract objectives into optimal equipment parameter configurations for the current vehicle and user. The control command sequence generator converts these parameters into specific equipment control commands arranged in a time sequence, which are then sent to various actuators, such as seat motors, air conditioning vents, and ambient lighting, via the vehicle execution network to achieve the physical execution of personalized services. After execution, the feedback recording module monitors the user's manual intervention behavior, generates anonymized feedback data, and uploads it to the cloud as needed. This data is then used by the model continuous learning and optimization module for subsequent model iterations, forming a complete data loop.

[0115] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the vehicle control method based on cloud-based large models and in-vehicle context perception. Any simple modifications based on this technical concept are within the scope of protection of this application.

[0116] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A vehicle control method based on a cloud large model and a vehicle situational awareness, characterized in that, The method is applied to the vehicle end, and the method includes: Upload the user's original commands and a snapshot of the scenario data to the cloud, so that the cloud can use the user's original commands and the snapshot of the scenario data as input to the target policy model to obtain a policy description; Based on the user's historical preference database, preset device constraints, real-time context data, and the strategy description, a sequence of device control instructions is obtained; The corresponding device actuators in the vehicle are controlled based on the sequence of device control commands.

2. The method of claim 1, wherein, The device control command sequence obtained based on the user's historical preference database, preset device constraints, real-time context data, and the strategy description includes: Parse the strategy description to obtain a list of target states and a list of constraints; The target policy function is determined based on the user's historical preference database, real-time contextual data, and the target state list; Set constraints for the target strategy function based on preset device constraints and the constraint list; Based on the constraints and the target strategy function, a sequence of device control instructions is obtained.

3. The method of claim 2, wherein, The step of determining the target policy function based on the user's historical preference database, real-time contextual data, and the target state list includes: Based on the target state list, determine the vector of actuators to be controlled; Based on the vector of the actuator to be controlled, query the user's historical preference database to determine the actuator weight vector; A target policy function is generated based on real-time context data, the actuator weight vector, the target state list, and the actuator vector to be controlled.

4. The method of claim 1, wherein, After controlling the corresponding device actuator in the vehicle to operate based on the device control command sequence, the method further includes: The monitoring results are obtained by detecting whether there is manual intervention by the user during the operation of the corresponding equipment actuators in the vehicle. Feedback data is obtained based on the user identifier, the monitoring results, and the policy description identifier corresponding to the device control command sequence; The feedback data is uploaded to the cloud so that the cloud can adjust the target strategy model based on the feedback data.

5. A vehicle control method based on a cloud large model and a vehicle situational awareness, characterized in that, The method is applied in the cloud and includes: Receive user-generated commands and snapshots of contextual data uploaded from the vehicle terminal; The user's original instructions and the scenario data snapshot are used as inputs to the target policy model to obtain a policy description, which is then sent to the vehicle terminal so that the vehicle terminal can obtain a sequence of device control instructions based on the policy description.

6. The method of claim 5, wherein, Using the original user commands and the snapshot of the scenario data as input to the target policy model, a policy description is obtained, including: Based on the user's original instructions and the context data snapshot, a knowledge retrieval result is obtained by retrieving information from the knowledge graph. The knowledge retrieval results, the user's original instructions, and the scenario data snapshot are processed into text to obtain the model input text. The input text of the model is used as the input of the target policy model, so that the target policy model can reason based on the input text to obtain a policy description, wherein the policy description includes user intent labels, a list of target states, a list of constraints, and key parameters.

7. The method of claim 5, wherein, Before obtaining the policy description by taking the user's original instructions and the scenario data snapshot as input to the target policy model, the process further includes: Multiple sets of training data samples and multiple sets of optimized data samples are obtained, wherein both the training data samples and the optimized data samples include user original instruction samples, scenario data snapshots and policy description samples; An initial model is trained based on the training data samples described above to obtain the strategy model to be optimized. Based on the optimized data samples and the strategy model to be optimized, a target strategy model is constructed.

8. The method of claim 7, wherein, The construction of the target policy model based on the optimized data samples and the policy model to be optimized includes: The user's original instruction sample and the scenario data snapshot in each of the optimized data samples are used as input to the strategy model to be optimized, so as to obtain multiple candidate strategy descriptions for each set of optimized data samples. A reward model is constructed based on the preference ranking data described in each of the candidate strategies. Based on the reward model and the strategy model to be optimized, a target strategy model is constructed.

9. The method of claim 5, wherein, After taking the user's original command and the scenario data snapshot as input to the target policy model to obtain a policy description and feeding the policy description back to the vehicle terminal so that the vehicle terminal can obtain a sequence of device control commands based on the policy description, the method further includes: Receive feedback data uploaded by the vehicle terminal; By aggregating multiple feedback data with the same user identifier, the preference adjustment direction for the corresponding user can be obtained; The feedback data with different user identifiers are clustered to obtain similar user groups; The target strategy model is optimized based on the stated preference adjustment direction and the stated similar user groups.

10. A vehicle control system based on cloud large model and vehicle on-board context awareness, characterized in that, The vehicle control system based on cloud-based large model and vehicle context awareness includes: a vehicle end and a cloud end, wherein the vehicle end executes the vehicle control method based on cloud-based large model and vehicle context awareness as described in any one of claims 1 to 4, and the cloud end executes the vehicle control method based on cloud-based large model and vehicle context awareness as described in any one of claims 5 to 9.