Vehicle escape method and device, electronic equipment and vehicle

By acquiring vehicle environment images and voice data, combined with vehicle signal tags, determining vehicle status and user intentions, and generating escape control strategies, the system solves the user's difficulty in maneuvering in trapped vehicles and improves driving experience and safety.

CN120804576APending Publication Date: 2025-10-17CHENGDU GREAT WALL MOTOR R&D CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510894766.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

When the vehicle is trapped, users do not know how to control the vehicle to get out of trouble, resulting in a poor driving experience. The existing system has inaccurate environmental perception, lacks semantic understanding in voice interaction, and lacks the ability to make coordinated judgments on vehicle status.

Method used

By acquiring vehicle environment image data and voice data, combined with vehicle signal labels, the vehicle status, user intention and type of distress are determined. After fusion processing, the data is input into the pre-trained escape processing model to generate the target escape control strategy.

Benefits of technology

The accuracy of the escape control strategy and user experience are improved, user panic is avoided, and driving safety is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804576A_ABST
    Figure CN120804576A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle safety, and provides a vehicle de-trapping method and device, electronic equipment and a vehicle, and the method comprises the steps: obtaining the image data and voice data of an environment where the vehicle is located, and determining a target intention response vector in response to the situation that the vehicle state is determined to be a trapped state according to the environment image data and the voice data; obtaining a vehicle signal label, and determining a target trapped type according to the environment image data, the voice data and the vehicle signal label; fusing the vehicle state, the target intention response vector and the target trapped type to obtain a target fusion vector; and inputting the target fusion vector into a pre-trained de-trapping processing model, and obtaining a target de-trapping control strategy through the de-trapping processing model. Therefore, the user can control the vehicle to get out of trouble according to the target getting-out control strategy, the situation that the user masters the vehicle in a limited degree and does not know how to control the vehicle to get out of trouble when the vehicle is in a trapped scene, and consequently the panic emotion is caused is avoided, and the driving safety and the vehicle using experience of the user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of vehicle safety, in particular to a vehicle escape method and device, an electronic device and a vehicle. BACKGROUND

[0002] With the rapid development of the vehicle technology field, vehicles have become an important means of transportation in people's daily life. In actual use, vehicles may be trapped due to vehicle trapping or bad weather, resulting in a trapped scenario.

[0003] Currently, due to the limited understanding of vehicles, when the vehicle is in a trapped scenario, it is often unclear how to control the vehicle to escape, which affects the driving experience of users. SUMMARY

[0004] Therefore, the purpose of the present disclosure is to provide a vehicle escape method, device, electronic device and vehicle to solve the problem that, due to the limited understanding of vehicles, when the vehicle is in a trapped scenario, it is often unclear how to control the vehicle to escape, which affects the driving experience of users.

[0005] To achieve the above purpose, the first aspect of the present disclosure provides a vehicle escape method, which comprises:

[0006] obtaining environmental image data and voice data of a vehicle, and determining a target intention response vector in response to determining that the vehicle state is a trapped state according to the environmental image data and the voice data;

[0007] obtaining a vehicle signal label, and determining a target trapped type according to the environmental image data, the voice data and the vehicle signal label;

[0008] fusing the vehicle state, the target intention response vector and the target trapped type to obtain a target fusion vector;

[0009] inputting the target fusion vector into a pre-trained escape processing model, and obtaining a target escape control strategy through processing by the escape processing model.

[0010] Specifically, the voice data comprises user voice data and voice interaction data, and the obtaining of the environmental image data and the voice data of the vehicle, and the determination of the target intention response vector in response to the determination that the vehicle state is the trapped state according to the environmental image data and the voice data comprises:

[0011] obtaining environmental image data and user voice data of a vehicle, inputting the environmental image data and the user voice data into a trained first language model, and outputting a vehicle state through processing by the first language model;

[0012] In response to the vehicle state being a trapped state, a first prompt information containing an initial escape control strategy is generated and outputted;

[0013] The voice interaction information corresponding to the first prompt information is received, and a target intention response vector is determined according to the voice interaction information.

[0014] Specifically, the voice interaction information corresponding to the first prompt information is received, and a target intention response vector is determined according to the voice interaction information, including:

[0015] The voice interaction information corresponding to the first prompt information is received, and the vehicle state and the voice interaction information are inputted into a trained second language model;

[0016] The target intention response vector is outputted via the second language model processing.

[0017] Specifically, the vehicle signal label is acquired, and a target trapped type is determined according to the environmental image data, the voice data and the vehicle signal label, including:

[0018] The vehicle signal label is acquired, and a time sequence signal matrix is obtained by time stamping analysis processing on the vehicle signal label;

[0019] The environmental image data, the voice data and the time sequence signal matrix are inputted into a trained third language model, and a target trapped type is outputted via the third language model processing.

[0020] Specifically, the third language model includes an input layer, a hidden layer and an output layer, and the environmental image data, the voice data and the time sequence signal matrix are inputted into a trained third language model via the third language model processing, and a target trapped type is outputted, including:

[0021] The environmental image data, the voice data and the time sequence signal matrix are inputted into the third language model through the input layer;

[0022] An initial event graph is constructed according to the environmental image data, the voice data and the time sequence signal matrix by using the hidden layer;

[0023] The initial event graph is compared with a preset event graph by the hidden layer to determine a target escape scene and a target trapped type corresponding to the initial event graph;

[0024] The target trapped type and the target escape scene are outputted by the output layer.

[0025] Specifically, the vehicle state, the target intention response vector, and the target trapped type are fused to obtain a target fusion vector, including:

[0026] A vehicle state vector corresponding to the vehicle state and a trapped type vector corresponding to the target trapped type are determined.

[0027] The vehicle state vector, the target intention response vector, and the trapped type vector are mapped and aligned to obtain a vehicle state embedding vector, a target intention response embedding vector, and a trapped type embedding vector.

[0028] A target fusion weight is determined, and the vehicle state embedding vector, the target intention response embedding vector, the trapped type embedding vector, and the target fusion weight are input into a trained vector fusion model to output a target fusion vector.

[0029] Specifically, after obtaining the target escape control strategy, further comprising:

[0030] A safety level corresponding to the target escape control strategy is determined.

[0031] In response to the safety level being greater than a preset level threshold, second prompt information containing the target escape control strategy is output; or,

[0032] In response to the safety level being less than the preset level threshold, the target escape control strategy is executed.

[0033] Based on the same inventive concept, a second aspect of the present disclosure provides a vehicle escape device, including:

[0034] A data acquisition module configured to acquire environmental image data and voice data of a vehicle, and determine a target intention response vector in response to determining that the vehicle state is a trapped state according to the environmental image data and the voice data.

[0035] A trapped cause vector determination module configured to acquire a vehicle signal label, and determine a target trapped type according to the environmental image data, the voice data, and the vehicle signal label.

[0036] A fusion vector determination module configured to fuse the vehicle state, the target intention response vector, and the target trapped type to obtain a target fusion vector.

[0037] An escape control strategy determination module configured to input the target fusion vector into a pre-trained escape processing model, and obtain a target escape control strategy via processing of the escape processing model.

[0038] Based on the same inventive concept, a third aspect of the present disclosure provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable by the processor, wherein the processor implements the vehicle escape method as described above when executing the computer program.

[0039] Based on the same inventive concept, a fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the vehicle escape method as described above.

[0040] Based on the same inventive concept, a fifth aspect of the present disclosure provides a vehicle comprising the vehicle escape device of the second aspect, the electronic device of the third aspect, or the storage medium of the fourth aspect.

[0041] As can be seen from the above, the present disclosure provides a vehicle escape method, device, electronic device, and vehicle. The environment image data and voice data of the vehicle are obtained. When it is determined according to the environment image data and the voice data that the vehicle state is a trapped state, it means that the vehicle is in a trapped scenario. A target intention response vector is determined, wherein the target intention response vector represents the current intention of the user, specifically whether the user has the intention to need assistance to escape. A vehicle signal label is obtained. The target trapped type is determined according to the environment image data, the voice data, and the vehicle signal label. The vehicle signal label represents the CAN bus signal of the vehicle. When it is determined that the current vehicle is in a trapped state, the vehicle signal label is considered when determining the trapped type, which makes the determination of the trapped type more accurate. The vehicle state, the target intention response vector, and the target trapped type are fused to obtain a target fusion vector. By fusing the vehicle state, the target intention response vector representing the user's intention, and the determined target trapped type, a target fusion vector is obtained, which improves the accuracy of the determination of the escape control strategy. At the same time, because the target intention response vector is considered when determining the target fusion vector, the user's intention is considered when determining the target escape control strategy based on the target fusion vector, which makes the target escape control strategy more closely related to the user's intention, i.e., the target escape control strategy determined for different users is different, and can better meet the user's needs. The target fusion vector is input into a pre-trained escape processing model, and the target escape control strategy is obtained by processing the escape processing model. This allows the user to control the vehicle to escape according to the target escape control strategy, avoids the situation that the user has limited knowledge of the vehicle and is not clear about how to control the vehicle to escape when the vehicle is in a trapped scenario, which leads to panic, improves the safety of driving and the user's driving experience. BRIEF DESCRIPTION OF DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the present disclosure or the related art, the following will briefly introduce the drawings needed to be used in the embodiments or the related description. Obviously, the drawings in the following description only constitute part of the embodiments of the present disclosure, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0043] Figure 1 Flow chart of the vehicle escape method according to an embodiment of the present disclosure;

[0044] Figure 2 Flow chart of the vehicle escape method according to another embodiment of the present disclosure;

[0045] Figure 3 Structure block diagram of the vehicle escape device according to an embodiment of the present disclosure;

[0046] Figure 4 Structure schematic diagram of the electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0047] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, the following will further describe the present disclosure in detail with specific embodiments and with reference to the drawings.

[0048] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood as the general meanings understood by those skilled in the art to which the present disclosure belongs. The terms "first", "second" and similar terms used in the embodiments of the present disclosure do not represent any order, number or importance, but are only used to distinguish different components. The terms "include" or "contain" and similar terms mean that the elements or objects before the terms cover the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right" and the like only represent relative positional relationships, and when the absolute positions of the described objects change, the relative positional relationships may also change accordingly.

[0049] The following are the explanations of the nouns:

[0050] ECU: ECU (Electronic Control Unit) electronic controller unit, also known as "vehicle computer", whose purpose is to control the driving state of the vehicle and realize various functions. It mainly uses various sensors, bus data acquisition and exchange to judge the vehicle state and the driver's intention and controls the vehicle through actuators.

[0051] CAN bus: Controller Area Network (CAN) is a serial communication protocol bus for real-time applications. It can use twisted pair wires to transmit signals and is one of the most widely used field buses in the world.

[0052] DeepSeek-R1: DeepSeek-R1 is a reasoning model developed by DeepSeek, an AI company under the Fantasy Quantification flag. DeepSeek-R1 uses reinforcement learning for post-training to improve reasoning ability, especially good at complex tasks such as mathematical, code, and natural language reasoning.

[0053] QwQ-VL-Max: QwQ-VL-Max is a QwQ reasoning model that has greatly improved its reasoning ability through reinforcement learning. It is used to understand and analyze user input natural language, as well as multi-modal data such as pictures, audio, and video.

[0054] QwQ-Plus: QwQ-Plus model is a QwQ reasoning model based on Qwen2.5 model training, which has greatly improved the model's reasoning ability through reinforcement learning.

[0055] Yolov8 model: Yolov8 is the latest target detection model of Ultralytics company, based on Yolo series, introducing new features such as new backbone network, decoupling detection head and VFLLoss, providing more efficient and flexible solutions.

[0056] DeepSeek-LLM: General large language model, using a similar architecture design as Llama, and has undergone multiple optimizations.

[0057] DeepSeek-MoE: The first MoE large model in China, with excellent generalization effect.

[0058] With the rapid development of vehicle technology, vehicles have become an important means of transportation in people's daily life. In actual use, vehicles may be trapped due to vehicle trouble or bad weather, resulting in a trapped scenario.

[0059] Currently, due to the limited understanding of vehicles by users, when the vehicle is in a trapped scenario, it is often unclear how to control the vehicle to escape, affecting the user's driving experience.

[0060] At the same time, most current vehicles often cannot accurately identify and respond to complex environments when trapped, which manifests in the following problems:

[0061] The image recognition system cannot understand semantic scenes, and current visual models mainly focus on target detection and image segmentation, and cannot understand higher-level semantics such as getting stuck, wheel spin, road closure, and lack of comprehensive judgment ability of "environmental status + user semantics".

[0062] The voice interaction system does not have scene perception ability, and most vehicle-mounted voice systems only rely on corpus matching, and cannot match user voice with the current escape scene. For example, the user says that I cannot move the car, and the system cannot determine whether it is really stuck or navigation error.

[0063] The existing system lacks the ability to judge the linkage of vehicle regulations and multi-source states, and many systems do not combine CAN bus data such as ESP operation and slip detection, resulting in a lack of system perception of the real state and inability to perform dynamic self-diagnosis.

[0064] Therefore, in view of the current situation that users have limited knowledge of vehicles, when the vehicle is in a trapped scene, they often do not know how to control the vehicle to escape, and the current vehicle in the trapped car, collision and other escape scenes has the problems of inaccurate environmental perception, lack of semantic understanding of voice interaction, and the system cannot link the vehicle state for comprehensive judgment. To solve the above problems, the embodiment provides a vehicle escape method, as shown in Figure 1 The method comprises the following steps:

[0065] Step 101, acquiring environmental image data and voice data of a vehicle, and determining a target intention response vector in response to determining that the vehicle state is a trapped state according to the environmental image data and the voice data.

[0066] In specific implementation, the environmental image data of the vehicle is acquired, wherein the environmental image data represents the image of the environment where the vehicle is located, and the weather data, road data and actual driving state of the vehicle in the current environment can be determined through the environmental image data.

[0067] For example, the weather data of the current environment where the vehicle is located is determined according to the environmental image data, such as heavy rain, snow, heavy fog, etc. The road data of the current environment where the vehicle is located is determined according to the environmental image data, such as congestion, smooth, existence of traffic accident, etc. The actual driving state of the vehicle is determined according to the environmental image data, such as getting stuck, rollover, wheel slip, etc.

[0068] In the embodiment, the environmental image data can be acquired by the camera device arranged on the vehicle, and the environmental image data can be picture data or video data, which is not limited in specific.

[0069] In this embodiment, the voice data is voice data issued by a user in the vehicle, including voice data exchanged between a driver and a passenger, voice data exchanged between the driver and the vehicle machine, voice data when the driver uses an electronic device to make a call, and the like.

[0070] The vehicle state is determined according to the environmental image data and the voice data, wherein the vehicle state indicates whether the vehicle is currently driving normally, and the vehicle state includes a trapped state and a non-trapped state. The trapped state indicates that the vehicle is currently in a trapped scene and cannot normally travel.

[0071] If it is determined according to the environmental image data and the voice data that the vehicle state is in the trapped state, a target intention response vector is determined, wherein the target intention response vector indicates the current intention of the user, specifically whether the user has the intention to need assistance to escape from the trapped state.

[0072] For example, when the vehicle is in the trapped state, some users may have certain driving experience, and they tend to handle the current trapped scene by themselves, that is, no additional assistance escape suggestion is needed from the vehicle machine, or no detailed assistance escape operation is needed, only a general reference strategy is needed, so the target intention response vector corresponds to providing a simple assistance escape control suggestion. Some users may have short driving experience and insufficient driving experience, and they cannot handle it by themselves, so the target intention response vector corresponds to providing a detailed assistance escape control suggestion.

[0073] In step 102, a vehicle signal tag is obtained, and a target trapped type is determined according to the environmental image data, the voice data, and the vehicle signal tag.

[0074] In specific implementation, the vehicle signal tag is obtained, wherein the vehicle signal tag is a state signal obtained by analyzing vehicle log data, and the vehicle log data is structured log data obtained from the whole vehicle ECU and CAN bus.

[0075] The target trapped type is determined according to the obtained environmental image data, voice data, and vehicle signal tag, wherein the target trapped type is used to indicate the type of the current trapped event of the vehicle, that is, which trapped scene it belongs to. For example, the trapped scene is a trapped car, a collision, a vehicle that cannot move, and the like.

[0076] Specifically, the event graph modeling is performed by inferring and analyzing the environmental image data and the voice data. The determined vehicle signal tag is used to assist in determining the accident causal chain, and the event graph, that is, the accident causal chain, is used to determine which trapped scene it belongs to.

[0077] Exemplarily, it is determined according to the environmental image data that the vehicle has a collision behavior, it is identified that the user voice data is unable to move, the vehicle signal label is obtained, it is determined according to the vehicle signal label that the vehicle is in an off state, and then it is inferred according to the environmental image data, the voice data and the vehicle signal label that the vehicle trapped type is that the vehicle is stuck and cannot move.

[0078] In step 103, the vehicle state, the target intention response vector and the target trapped type are fused to obtain a target fusion vector.

[0079] In implementation, a target fusion weight is determined, wherein the target fusion weight is a weight value corresponding to the vehicle state, the target intention response vector and the target trapped type respectively when the vehicle state, the target intention response vector and the target trapped type are fused.

[0080] Based on the determined target fusion weight, the vehicle state, the target intention response vector and the target trapped type are fused, that is, the vehicle state, the target intention response vector and the target trapped type are fused according to the respective weights to obtain a target fusion vector.

[0081] In step 104, the target fusion vector is input into a pre-trained trapped processing model, and a target trapped control strategy is obtained by processing via the trapped processing model.

[0082] In implementation, the target fusion vector obtained by fusion is input into a pre-trained trapped processing model, and a target trapped control strategy is output by processing via the trapped processing model. The target trapped control strategy represents a specific control mode for realizing vehicle trapped in the current vehicle trapped scenario.

[0083] In this embodiment, the training process of the trapped processing model specifically includes:

[0084] In step a, a first training data set and an initial trapped processing model are obtained, wherein the first training data set includes historical fusion vectors and historical trapped control strategies.

[0085] In step b, training data in the first training data set is input into the initial trapped processing model for training, and a first preset training end condition is determined to satisfy, and a trapped processing model is obtained.

[0086] In implementation, a first training data set and an initial trapped processing model are obtained, wherein the first training data set includes historical fusion vectors and historical trapped control strategies. Training data in the first training data set is input into the initial trapped processing model for training, and a first preset training end condition is determined to satisfy, and a trapped processing model is obtained.

[0087] The first preset training end condition comprises at least one of the following: determining that all data in the first training data set is input into the initial escape processing model for training, determining that a loss function of the initial escape processing model converges to a first convergence threshold, or determining that the initial escape processing model is iteratively trained to a first preset iteration number.

[0088] For example, the first preset training end condition is to determine that all data in the first training data set is input into the initial escape processing model for training.

[0089] The first training data set comprises fifty groups of data, and each group of data comprises a historical fusion vector and a historical escape control strategy. The first preset training end condition is to determine that all data in the first training data set is input into the initial escape processing model for training, that is, when the fifty groups of data are input into the initial escape processing model, there is no training data in the first training data set that has not been input into the initial escape processing model, and it is determined that the training of the initial escape processing model is completed, and the escape processing model is obtained.

[0090] For another example, the first preset training end condition is to determine that a loss function of the initial escape processing model converges to a first convergence threshold.

[0091] The training data in the first training data set is input into the initial escape processing model for training, and a training result is output. The loss function is determined according to the training result and the historical fault type, the historical fault processing method and the historical fault cause. The type of the loss function comprises at least one of the following: mean square error loss function, cross-entropy loss function, logarithmic loss function, exponential loss function, square loss function or absolute value loss function. When the loss function converges to the first convergence threshold, it is determined that the first preset training end condition is met, and the escape processing model is obtained.

[0092] For another example, the first preset training end condition is to determine that the initial escape processing model is iteratively trained to a first preset iteration number.

[0093] The training data in the first training data set is input into the initial escape processing model for iterative training, and the iteration number is recorded. When the iteration number is equal to the first preset iteration number, it is determined that the first preset training end condition is met, and the escape processing model is obtained.

[0094] In this embodiment, the training data in the first training data set comprises a historical fusion vector and a historical escape control strategy. The initial escape processing model is trained using the training data in the first training data set until the first preset training end condition is met, and the escape processing model is obtained. The escape processing model is used to output a target escape control strategy corresponding to a target fusion vector, and the determined target escape control strategy is more accurate.

[0095] By the above scheme, the environmental image data and the voice data in which the vehicle is located are obtained. When it is determined that the vehicle state is a trapped state according to the environmental image data and the voice data, it means that the vehicle is in a trapped scenario at this time. The target intention response vector is determined, wherein the target intention response vector represents the current intention of the user, specifically whether the user has the intention to need assistance to escape from the trapped state. The vehicle signal label is obtained. The target trapped type is determined according to the environmental image data, the voice data, and the vehicle signal label. The vehicle signal label represents the vehicle CAN bus signal. When it is determined that the current vehicle is in a trapped state, the vehicle signal label is considered in the determination of the trapped type, so that the determination of the trapped type is more accurate. The vehicle state, the target intention response vector, and the target trapped type are fused to obtain a target fusion vector. By fusing the vehicle state, the target intention response vector representing the user's intention, and the determined target trapped type, the target fusion vector is obtained, so that the accuracy of the determination of the trapped control strategy is improved according to the target fusion vector. At the same time, because the target intention response vector is considered in the determination of the target fusion vector, the user's intention is considered in the determination of the target trapped control strategy according to the target fusion vector, so that the target trapped control strategy is more closely related to the user's intention, that is, the target trapped control strategy determined for different users is different, and the user's demand is better met. The target fusion vector is input into the pre-trained trapped processing model, and the target trapped control strategy is obtained by processing the trapped processing model. The user can control the vehicle to escape from the trapped state according to the target trapped control strategy, avoid the user's limited knowledge of the vehicle, and not know how to control the vehicle to escape from the trapped state when the vehicle is in a trapped scenario, so as to avoid the user's panic emotion, improve the safety of driving and the user's experience of using the vehicle.

[0096] At the same time, in the embodiment, the trapped processing model is a pre-trained model, that is, the trapped processing model is trained by using the first training data set, and the training data in the first training data set includes historical fusion vectors and historical trapped control strategies. The trapped processing model can learn the corresponding historical trapped control strategies according to the historical fusion vectors, and then the trained trapped processing model can be used subsequently. The trapped processing model can analyze the determined target fusion vector, and then the target trapped control strategy can be obtained. The determined target trapped control strategy is more accurate.

[0097] In some embodiments, the voice data includes user voice data and voice interaction data. In step 101, the environmental image data and the voice data in which the vehicle is located are obtained. In response to determining that the vehicle state is a trapped state according to the environmental image data and the voice data, the target intention response vector is determined, specifically including:

[0098] Step 1011, obtain the environmental image data of the vehicle and the user voice data, input the environmental image data and the user voice data into the trained first language model, output the vehicle state through the first language model processing;

[0099] Step 1012, in response to the vehicle state being a trapped state, generating and outputting first prompt information containing an initial escape control strategy;

[0100] Step 1013, receiving the voice interaction information corresponding to the first prompt information, and determining the target intention response vector according to the voice interaction information.

[0101] In specific implementation, the environmental image data of the vehicle and the user voice data are obtained, wherein the user voice data is the natural interaction expectation of the in-vehicle user recognized. For example, the user voice data is "I can't move the car", "Is it trapped" and the like.

[0102] The environmental image data and the user voice data are input into the trained first language model, and the vehicle state is output through the first language model processing.

[0103] In this embodiment, the first language model can be a QwQ-VL-Max model or a Yolov8 model. Taking the QwQ-VL-Max model as an example, the process of analyzing the environmental image data and the user voice data by the QwQ-VL-Max model includes:

[0104] The input of the QwQ-VL-Max model is the environmental image data and the user voice data, the task target is to determine the trapped scene jointly by the image content and the user voice, and the output is whether it belongs to the trapped scene.

[0105] When the first language model determines that the vehicle state is a trapped state, the first prompt information containing the initial escape control strategy is generated and output. The initial escape control strategy can be determined by the first language model according to the environmental image data and the user voice data. That is, when the first language model determines that the vehicle state is a trapped state, the corresponding initial escape control strategy is determined, and the initial escape control strategy is output together with the vehicle state.

[0106] In this embodiment, the prompt mode of the first prompt information includes at least one of the following: voice broadcast, HUD display, instrument display, center screen display, window display, and in-vehicle device linkage. The HUD display is a head-up display. For example, the prompt mode of the first prompt information is the HUD display, and the initial escape control strategy is projected in front of the user.

[0107] Exemplarily, the content of the first prompt information is in the form of text, and the prompt manner is the display of the central control screen. In this case, the central control screen displays the text description corresponding to the initial escape control strategy.

[0108] The voice interaction information corresponding to the first prompt information is received, and a target intention response vector is determined according to the voice interaction information. The voice interaction information is voice interaction dialogue data between the user and the vehicle machine after the user receives the initial escape control strategy.

[0109] Exemplarily, the initial escape control strategy included in the first prompt information is to suggest calling a rescue phone. After the first prompt information is output, the vehicle machine outputs prompt information about whether to attempt to call a rescue phone, and receives feedback voice data of the user. At this time, the prompt information about whether to attempt to call a rescue phone and the feedback voice data of the user are the voice interaction information.

[0110] Through the above scheme, the vehicle state is determined by using the first language model according to the environmental image data and the user voice data. The vehicle state determination is more accurate. Meanwhile, the target intention response vector is determined according to the voice interaction information with the user, that is, the user intention is determined. The determination of the user intention is more accurate.

[0111] In some embodiments, in step 1013, the voice interaction information corresponding to the first prompt information is received, and a target intention response vector is determined according to the voice interaction information. Specifically, the step includes:

[0112] In step 10131, the voice interaction information corresponding to the first prompt information is received, and the vehicle state and the voice interaction information are input into a trained second language model.

[0113] In step 10132, the target intention response vector is output via processing of the second language model.

[0114] In specific implementation, the voice interaction information corresponding to the first prompt information is received, the vehicle state and the voice interaction information are input into a trained second language model, and the target intention response vector is output via processing of the second language model.

[0115] In this embodiment, the second language model is a QwQ-Plus model. The process of using the QwQ-Plus model to process the vehicle state and the voice interaction information to obtain the target intention response vector is as follows:

[0116] The vehicle state and the voice interaction data are input into the QwQ-Plus model. The QwQ-Plus model controls the dialogue rhythm, such as “whether to attempt to call a rescue”, “whether to perform vehicle self-check”, etc. The state confirmation, intention recognition and semantic analysis are performed, and finally the target intention response vector is output.

[0117] In the embodiment, to improve the accuracy of the target intent response vector output by the QwQ-Plus model, during training, the training data is the vehicle enterprise customer service dialogue corpus and the user case of getting rid of the trouble. Meanwhile, the loss function of the QwQ-Plus model is determined in the manner of combining the intent classification loss and the slot filling and extraction loss.

[0118] Through the above scheme, when determining the target intent response vector, the second language model is used for determination, and the output target intent response vector is more accurate, that is, the user's intent recognition is more accurate, and then the subsequent target escape control strategy is more matched with the user's real intent.

[0119] In some embodiments, the vehicle signal label is acquired in step 102, and the target trapped type is determined according to the environment image data, the voice data and the vehicle signal label, specifically comprising:

[0120] In step 1021, the vehicle signal label is acquired, and the time stamp analysis processing is performed on the vehicle signal label to obtain a time sequence signal matrix.

[0121] In step 1022, the environment image data, the voice data and the time sequence signal matrix are input into the trained third language model, and the target trapped type is output via the third language model processing.

[0122] In specific implementation, after the vehicle signal label is acquired, the time stamp of the vehicle signal label is analyzed and organized according to the channel, and the time sequence signal matrix is output. The time sequence signal matrix is a multi-channel time sequence. Exemplarily, the window length corresponding to the time sequence signal matrix is 128 to 256 frames.

[0123] The environment image data, the voice data and the time sequence signal matrix are input into the trained third language model, and the target trapped type is output via the third language model processing.

[0124] In the embodiment, the third language model includes an input layer, a hidden layer and an output layer, and the specific process of outputting the target trapped type by the third language model includes:

[0125] The environment image data, the voice data and the time sequence signal matrix are input into the third language model through the input layer, and the hidden layer is used to construct an initial event graph according to the environment image data, the voice data and the time sequence signal matrix. The initial event graph is compared with a preset event graph through the hidden layer to determine the target escape scene and the target trapped type corresponding to the initial event graph. The target trapped type and the target escape scene are output through the output layer.

[0126] In this embodiment, the third language model is the DeepSeek-R1 model, and may also be a language model of the same type, such as DeepSeek-LLM, DeepSeek-MoE, etc. Taking the DeepSeek-R1 model as an example, the process of using the DeepSeek-R1 model to process the environmental image data, the speech data, and the time series signal matrix is ​​as follows:

[0127] The environmental image data, voice data, and time-series signal matrix are structured to generate structured data. The environmental image data, voice data, and time-series signal matrix are converted into nodes and connected using causal relationships to generate an event graph. For example, tire slippage corresponds to vehicle immobility. Based on the event graph, the specific escape scenario is determined, i.e., the target trapped type is determined.

[0128] In this embodiment, when a vehicle is trapped, there may be multiple possible causes of vehicle failure, i.e., multiple possible types of trapped conditions. Therefore, when the third language model is used to process the environmental image data, the voice data, and the time series signal matrix, and outputs the target trapped condition, it also outputs the priority level corresponding to each trapped condition. Subsequently, when determining the target escape control strategy, it is executed sequentially according to the priority level.

[0129] For example, if the vehicle is currently immobile, the third language model analyzes the environmental image data, the voice data, and the time-series signal matrix to determine that the target distress type includes collision and vehicle entrapment. If the collision priority level is determined to be medium and the vehicle entrapment priority level is determined to be high, then when determining the target control strategy, resolving the entrapment caused by the vehicle entrapment is prioritized over resolving the collision issue.

[0130] Through the above scheme, while determining the target trapped type, the priority level corresponding to each trapped type can also be output at the same time. The priority level corresponds to the severity. Then, when determining the target control strategy, the problems with higher severity are solved first, avoiding solving the problems that cause insufficient vehicle resources at the same time.

[0131] In some embodiments, step 103 fuses the vehicle state, the target intention response vector, and the target trapped type to obtain a target fusion vector, specifically including:

[0132] Step 1031, determining a vehicle state vector corresponding to the vehicle state and a distress type vector corresponding to the target distress type;

[0133] Step 1032, mapping and aligning the vehicle state vector, the target intention response vector and the trapped type vector to obtain a vehicle state embedding vector, a target intention response embedding vector and a trapped type embedding vector;

[0134] Step 1033, determining a target fusion weight, inputting the vehicle state embedding vector, the target intention response embedding vector, the trapped type embedding vector and the target fusion weight into a trained vector fusion model to output a target fusion vector.

[0135] In specific implementation, a vehicle state vector corresponding to a vehicle state is determined, wherein the vehicle state vector is used to reflect different vehicle states, and in this embodiment, the vehicle state vector represents a vector representation of whether the vehicle is in a trapped state determined based on a QwQ-VL-Max model.

[0136] A trapped type vector corresponding to a target trapped type is determined, and in this embodiment, the trapped type vector represents a vector representation of reasons for causing the vehicle to be trapped determined based on a DeepSeek-R1 model.

[0137] The vehicle state vector, the target intention response vector and the trapped type vector are mapped and aligned to obtain a vehicle state embedding vector, a target intention response embedding vector and a trapped type embedding vector.

[0138] Specifically, the vehicle state vector is linearly transformed to map the vehicle state vector to a preset semantic space to obtain a vehicle state embedding vector. The target intention response vector is linearly transformed to map the target intention response vector to the preset semantic space to obtain a target intention response embedding vector. The trapped type vector is linearly transformed to map the trapped type vector to the preset semantic space to obtain a trapped type embedding vector.

[0139] The aligned vehicle state embedding vector, target intention response embedding vector and trapped type embedding vector are further fused through a multi-head attention mechanism, so that the model can identify the deep relationship and causal relationship therebetween. That is, the vehicle state embedding vector, the target intention response embedding vector, the trapped type embedding vector and the target fusion weight are input into a trained vector fusion model to output a target fusion vector.

[0140] In this embodiment, the target fusion weight represents the weight value corresponding to the vehicle state embedding vector, the target intent response embedding vector, and the trapped type embedding vector when the vehicle state embedding vector, the target intent response embedding vector, and the trapped type embedding vector are fused to obtain the target fusion vector. In this embodiment, the target fusion weight can be pre-set, or can be dynamically adjusted according to the accuracy of the vehicle state embedding vector, the target intent response embedding vector, and the trapped type embedding vector, to ensure that the obtained target fusion vector is more consistent with the actual trapped scene of the vehicle and the user intent.

[0141] In this embodiment, the vector fusion model adopts a cross-modal contrast loss function, and is trained through historical paired data to make related samples close in a semantic space. The historical paired data includes a vehicle state vector, an intent response vector, and a trapped type vector. The historical matched paired data is used as a positive sample, and the unmatched data is used as a negative sample, to construct a contrast learning model, that is, the vector fusion model, which can determine the relevance between the three modalities.

[0142] Through the above scheme, the vehicle state vector, the target intent response vector, and the trapped type vector are fused to obtain the target fusion vector, so that the subsequent trapped processing model automatically identifies the relevance between the three modalities when analyzing and processing the target fusion vector, thereby improving the determination accuracy of the trapped control strategy.

[0143] In some embodiments, the safety level corresponding to the target trapped control strategy can be determined, and when the safety risk is high, only a prompt is output instead of a control instruction, thereby avoiding the problem that inaccurate vehicle control threatens driving safety, that is, after obtaining the target trapped control strategy in step 104, the method further includes:

[0144] Step 10A, determining the safety level corresponding to the target trapped control strategy;

[0145] Step 10B, in response to the safety level being greater than a preset level threshold, outputting second prompt information containing the target trapped control strategy; or,

[0146] Step 10C, in response to the safety level being less than the preset level threshold, executing the target trapped control strategy.

[0147] In specific implementation, the safety level corresponding to the target trapped control strategy is determined, specifically, a database can be searched according to the target trapped control strategy to determine the safety level corresponding to the target trapped control strategy. The database stores a corresponding relationship between trapped control strategies and safety levels. The form of the corresponding relationship can include at least one of the following: a relationship table, a functional relationship, a curve relationship, a key-value pair relationship, and a column chart relationship.

[0148] The determined safety level is compared with a preset level threshold. If the safety level is greater than the preset level threshold, at this time, it indicates that the risk of executing the target escape control strategy by the vehicle is relatively high, and a safety accident is likely to occur. Therefore, at this time, only the second prompt information containing the target escape control strategy is output, and the user executes the target escape control strategy according to the prompt of the second prompt information, so as to realize the escape of the vehicle.

[0149] In this embodiment, the content form of the second prompt information includes at least one of the following: voice, text, picture, video and rich text. The second prompt information can be displayed on the vehicle side, or can be displayed on the APP of the user terminal.

[0150] Case one: if displayed on the vehicle side, the prompt mode of the second prompt information includes at least one of the following: voice broadcast, HUD display, instrument display, central screen display, window display and in-vehicle equipment linkage. The HUD display is a head-up display. For example, if the prompt mode of the second prompt information is the HUD display, the second prompt information will be projected in front of the user.

[0151] For example, the content form of the second prompt information is text, and the prompt mode is central screen display. The central screen displays text descriptions of how the user executes the target escape control strategy. For example, it prompts the user to reverse, suggests to slow down, recommends to call a rescue phone, etc.

[0152] Case two: if displayed on the APP of the user terminal, the terminal includes at least one of the following: mobile phone, watch, bracelet, notebook computer or tablet computer. The prompt mode of the second prompt information includes at least one of the following: voice broadcast or pop-up window display.

[0153] For example, the terminal is a mobile phone, the content form of the second prompt information is text, and the prompt mode is pop-up window display. The pop-up window is displayed on the user's mobile phone, and the pop-up window contains text descriptions of how the user executes the target escape control strategy.

[0154] If the safety level is less than the preset level threshold, at this time, it indicates that the target escape control strategy can be directly executed to realize the escape of the vehicle.

[0155] Through the above scheme, by determining the safety level corresponding to the target escape control strategy, and then determining whether to directly execute or only output the prompt information according to the safety level, the problem that the driving risk is increased and there is a safety hazard when directly executing the escape control strategy in a high safety risk is avoided, and the safety of the user driving is improved.

[0156] Based on the same inventive concept, another embodiment of the present disclosure provides a vehicle escape method, which is similar to the above-mentioned vehicle escape method. Figure 2As shown, the method specifically includes:

[0157] Step 201, collecting multi-source vehicle data.

[0158] In specific implementation, the multi-source vehicle data includes image data, user voice and vehicle state data, wherein the image data includes scenes such as vehicle sinking, vehicle rollover, wheel skidding, snow blocking, etc., the voice data is natural interaction corpus such as I can't move the car, is it stuck, etc., and the vehicle state data is CAN bus signal data such as high ESP trigger frequency, skidding detection, long time stop, etc.

[0159] Step 202, using QwQ-VL-Max model for image semantic and language joint recognition, using QwQ-Plus model for multi-round semantic dialogue and instruction understanding, and using DeepSeek-R1 model for semantic structure understanding and event modeling.

[0160] In specific implementation, the input of the QwQ-VL-Max model is environmental image data and user voice data, the task target is to jointly determine the escape scene of image content and user voice, and the output is whether it belongs to the trapped scene and the preliminary semantic suggestion.

[0161] According to the image and user language recognition result, the QwQ-VL-Max outputs the preliminary semantic suggestion (such as "suggesting alarm" and "suggesting trying to reverse"), which can provide the first reaction processing suggestion for the user within milliseconds, so as to win the key time for the user. Before the user clearly expresses the demand, the system intervenes first, shortens the response time, and improves the user's trust.

[0162] The preliminary semantic suggestion will be an important input for the next dialogue module, providing context basis for the multi-round semantic interaction of the QwQ-Plus. The system can automatically generate subsequent guidance according to the "preliminary judgment of wheel skidding": "whether to try to reverse at low speed?", "whether to need rescue?", so as to build a natural and close to user psychological expectation semantic process.

[0163] The QwQ-VL-Max model uses multi-modal cross attention to fuse visual targets and language intentions, can automatically assign semantic labels to visual targets, and generate reasoning process chains. During training, it is fine-tuned based on self-collected escape picture-text pairs and accident picture-question-answer data sets.

[0164] The vehicle state and voice interaction data are input into the QwQ-Plus model, and the QwQ-Plus model controls the dialogue rhythm such as "whether to try to call for rescue", "whether to perform vehicle self-check", etc. The execution state confirmation, intention recognition and semantic analysis are performed, and finally the target intention response vector is output.

[0165] To improve the accuracy of the target intent response vector output by the QwQ-Plus model, during training, the training data is the vehicle enterprise customer service dialogue corpus and the user case of the escape from the predicament. At the same time, the loss function of the QwQ-Plus model is determined by combining the intent classification loss and the slot filling and extraction loss.

[0166] The input of the DeepSeek-R1 model is the determination result output by the QwQ-VL-Max model, the target intent response vector output by the QwQ-Plus model, and the vehicle state data. The DeepSeek-R1 model uses the event graph modeling to determine whether the current state is an emergency escape, and assists in determining the accident causal chain. The output of the DeepSeek-R1 model is an event chain in the form of a graph structure, an accident classification result, and a priority level. For example, according to the impact, the engine stall, and the user's inability to persuade, it is inferred that the vehicle is stuck in a stuck type.

[0167] Step 203, constructing a multi-modal fusion vector.

[0168] In specific implementation, the determination result provided by the QwQ-VL-Max, i.e., the vehicle state vector in the above embodiment, the intent response vector of the QwQ-Plus, and the vehicle stuck type output by the DeepSeek-R1 are input into a unified fusion module to construct a multi-modal vector representation of the escape response.

[0169] Step 204, generating a system response according to the fusion vector.

[0170] In specific implementation, the system response is generated according to the fusion vector, and the system response is the target escape control strategy described in the above embodiment, including prompting, guiding attempts to act, and calling for help. At the same time, in this embodiment, reinforcement learning can be used for training when generating the system response, such as using a simulator or historical success / failure case feedback to optimize the suggestion sequence.

[0171] For example, a simulated escape environment is constructed or real case data is used, such as "successful escape from a stuck car" and "attempting to fail and seeking help". The system tries various response actions (such as prompting to reverse, suggesting to slow down, and recommending to seek help) in different scenarios, and each behavior has corresponding feedback (such as success rate and user satisfaction). Using the reward mechanism to continuously adjust the strategy, the system gradually masters the optimal response path "when to suggest trying to save oneself and when to directly recommend seeking help" in different situations;

[0172] For example, in the case of "wheel slip + voice expression for emergency help", the model will learn to preferentially output "prepare to call for road assistance".

[0173] On the basis of model training, the internal escape rules library of the vehicle enterprise is added, such as "new energy vehicle model is not recommended to frequently reverse escape", "mountainous road conditions are recommended to try to switch to low speed mode" and the like. The rules are hard-coded in the form of "IF-THEN" and inserted into the model post-processing flow, and the training sample generation can also be guided by the rules. Different vehicle models, regions, seasons and the like can be associated with different response strategies to improve the generalization ability and safety guarantee of the model.

[0174] In the embodiment, the escape rule library can be inserted into the vehicle enterprise self-defined strategy. Exemplarily, the recommended solutions are different under different vehicle models and different terrains.

[0175] Step 205, vehicle control strategy assistance and safety confirmation mechanism.

[0176] In specific implementation, when the safety risk is high, the system must output a prompt instead of a control instruction according to the rules, for example, "when the confidence is lower than the threshold, only generate a voice prompt, and do not trigger any control operation".

[0177] Step 206, vehicle private data continuous optimization mechanism.

[0178] In specific implementation, the data source is the user operation behavior log and the rescue success / failure record, the training strategy adopted is incremental learning, and Replay Buffer is adopted to prevent the model from forgetting historical emergency experience. At the same time, a privacy policy is set, all data are locally cached, and after receiving the explicit authorization of the user, the data can be synchronized to the training cloud.

[0179] It should be noted that the method of the embodiment of the disclosure can be executed by a single device, such as a computer or a server. The method of the embodiment can also be applied to a distributed scenario, and completed by multiple devices cooperating with each other. In this distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiment of the disclosure, and the multiple devices can interact with each other to complete the method.

[0180] It should be noted that some embodiments of the disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different than that described above and still achieve the desired result. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or necessary.

[0181] Based on the same inventive concept, the disclosure also provides a vehicle escape device corresponding to the method of any of the above embodiments.

[0182] ReferenceFigure 3 , Figure 3 The vehicle escape device of the embodiment comprises:

[0183] The data acquisition module 301 is configured to acquire environmental image data and voice data of a vehicle, and determine a target intention response vector in response to determining that a vehicle state is a trapped state according to the environmental image data and the voice data.

[0184] The trapped cause vector determination module 302 is configured to acquire vehicle signal labels, and determine a target trapped type according to the environmental image data, the voice data and the vehicle signal labels.

[0185] The fusion vector determination module 303 is configured to fuse the vehicle state, the target intention response vector and the target trapped type to obtain a target fusion vector.

[0186] The escape control strategy determination module 304 is configured to input the target fusion vector into a pre-trained escape processing model, and obtain a target escape control strategy through processing of the escape processing model.

[0187] In some embodiments, the voice data comprises user voice data and voice interaction data, and the data acquisition module 301 is specifically configured to:

[0188] acquire environmental image data and user voice data of a vehicle, input the environmental image data and the user voice data into a trained first language model, process through the first language model, and output a vehicle state;

[0189] generate and output first prompt information containing an initial escape control strategy in response to the vehicle state being a trapped state;

[0190] receive voice interaction information corresponding to the first prompt information, and determine a target intention response vector according to the voice interaction information.

[0191] In some embodiments, the data acquisition module 301 is specifically configured to:

[0192] receive voice interaction information corresponding to the first prompt information, and input the vehicle state and the voice interaction information into a trained second language model;

[0193] process through the second language model, and output a target intention response vector.

[0194] In some embodiments, the trapped cause vector determination module 302 is specifically configured to:

[0195] Obtaining a vehicle signal label, performing timestamp analysis on the vehicle signal label to obtain a time sequence signal matrix;

[0196] Inputting the environmental image data, the voice data, and the time sequence signal matrix into a trained third language model, performing processing via the third language model, and outputting a target trapped type.

[0197] In some embodiments, the third language model includes an input layer, a hidden layer, and an output layer, and the trapped cause vector determination module 302 is specifically configured to:

[0198] Inputting the environmental image data, the voice data, and the time sequence signal matrix into the third language model through the input layer;

[0199] Using the hidden layer to construct an initial event graph according to the environmental image data, the voice data, and the time sequence signal matrix;

[0200] Comparing the initial event graph with a preset event graph through the hidden layer to determine a target escape scene and a target trapped type corresponding to the initial event graph;

[0201] Outputting the target trapped type and the target escape scene through the output layer.

[0202] In some embodiments, the fusion vector determination module 303 is specifically configured to:

[0203] Determine a vehicle state vector corresponding to the vehicle state and a trapped type vector corresponding to the target trapped type; perform mapping and alignment processing on the vehicle state vector, the target intention response vector, and the trapped type vector to obtain a vehicle state embedding vector, a target intention response embedding vector, and a trapped type embedding vector;

[0204] Determine a target fusion weight, input the vehicle state embedding vector, the target intention response embedding vector, the trapped type embedding vector, and the target fusion weight into a trained vector fusion model, and output a target fusion vector.

[0205] In some embodiments, the device further includes an output module, which is specifically configured to:

[0206] Determine a safety level corresponding to the target escape control strategy;

[0207] In response to the safety level being greater than a preset level threshold, output second prompt information containing the target escape control strategy; or,

[0208] In response to the safety level being less than the preset level threshold, execute the target escape control strategy.

[0209] For the convenience of description, the above apparatus is described in various modules in terms of functions. Of course, the functions of the modules can be implemented in one or more software and / or hardware when implementing the present disclosure.

[0210] The apparatus of the above embodiments is used to implement the corresponding vehicle escape method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here again.

[0211] Based on the same inventive concept, the present disclosure also provides an electronic device corresponding to the method of any of the above embodiments, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the vehicle escape method of any of the above embodiments when executing the program.

[0212] Figure 4 A more specific hardware structure of an electronic device provided by the present embodiment is shown, which can include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040 and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030 and the communication interface 1040 are connected to each other through the bus 1050 for communication within the device.

[0213] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, etc., for executing related programs to implement the technical solutions provided by the present embodiment.

[0214] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs, and when the technical solutions provided by the present embodiment are implemented by software or firmware, the related program codes are saved in the memory 1020 and executed by the processor 1010.

[0215] The input / output interface 1030 is configured to connect an input / output module to realize information input and output. The input / output module can be configured in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0216] The communication interface 1040 is configured to connect a communication module (not shown in the figure) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as a USB, a network cable, etc.) or a wireless manner (such as a mobile network, WIFI, Bluetooth, etc.).

[0217] The bus 1050 includes a channel to transmit information between various components (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040) of the device.

[0218] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only include components necessary for implementing the embodiments of the present disclosure, and does not necessarily include all components shown in the figure.

[0219] The electronic device of the above embodiments is used to implement the corresponding vehicle escape method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.

[0220] Based on the same inventive concept, the disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the vehicle escape method according to any of the above embodiments.

[0221] The computer readable medium of the embodiments includes permanent and non-permanent, removable and non-removable media, which can realize information storage by any method or technology. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0222] The storage medium of the above-mentioned embodiments stores computer instructions for causing the computer to execute the vehicle escape method according to any one of the above-mentioned embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0223] Based on the same inventive concept, the present application also provides a vehicle corresponding to the method of any one of the above-mentioned embodiments, comprising the vehicle escape device of the above-mentioned embodiments, the electronic device of the above-mentioned embodiments, the computer readable storage medium of the above-mentioned embodiments, and the vehicle device realizes the vehicle escape method of any one of the above-mentioned embodiments.

[0224] The vehicle of the above-mentioned embodiments is used to realize the vehicle escape method of any one of the above-mentioned embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0225] It can be understood that before using the technical solutions of various embodiments in the present disclosure, the user will be informed of the type, scope of use, use scenario, etc. of the personal information involved by appropriate means, and the authorization of the user will be obtained.

[0226] For example, in response to receiving the user's active request, the user is sent prompt information to explicitly prompt the user that the operation requested to be performed will require the acquisition and use of the user's personal information. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic devices, application programs, servers or storage media that perform the technical solutions of the present disclosure according to the prompt information.

[0227] As an optional but not limited implementation manner, in response to accepting the user's active request, the way of sending prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0228] It can be understood that the above notification and obtaining user authorization process is only illustrative, and does not limit the implementation of the present disclosure, and other ways meeting relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0229] It should be understood by those of ordinary skill in the art that the above discussion of any embodiment is only exemplary and is not intended to imply that the scope of the present disclosure (including claims) is limited to these examples; the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity.

[0230] In addition, in order to simplify the description and discussion, and so as not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections of integrated circuit (IC) chips and other components can or can not be shown in the provided drawings. In addition, the apparatus can be shown in the form of a block diagram in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram apparatus are highly dependent on the platform to be implemented to implement the embodiments of the present disclosure (i.e., these details should be fully within the understanding of those skilled in the art). Where specific details (e.g., circuitry) are set forth in order to describe an illustrative embodiment of the present disclosure, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details or with variations on these specific details. Therefore, these descriptions should be considered as illustrative rather than limiting.

[0231] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications and variations will be apparent to those skilled in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.

[0232] The embodiments of the present disclosure are intended to cover all such alternatives, modifications and variations as falling within the broad scope of the appended claims. Accordingly, any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A method for escaping a vehicle, characterized in that: include: acquiring image data and voice data of an environment in which the vehicle is located, and determining a target intention response vector in response to determining that the vehicle is in a trapped state based on the image data and the voice data; Obtaining a vehicle signal tag, and determining the target trapped type according to the environmental image data, the voice data, and the vehicle signal tag; fusing the vehicle state, the target intention response vector, and the target trapped type to obtain a target fusion vector; The target fusion vector is input into a pre-trained escape processing model, and is processed by the escape processing model to obtain a target escape control strategy.

2. The method according to claim 1, characterized in that The voice data includes user voice data and voice interaction data, The acquiring of the environment image data and the voice data of the vehicle, and determining the vehicle state as a trapped state based on the environment image data and the voice data, and determining the target intention response vector, includes: Obtaining environmental image data and user voice data of the vehicle, inputting the environmental image data and the user voice data into a trained first language model, processing the data through the first language model, and outputting a vehicle state; In response to the vehicle being in a trapped state, generating and outputting first prompt information including an initial escape control strategy; Receive voice interaction information corresponding to the first prompt information, and determine a target intention response vector based on the voice interaction information.

3. The method according to claim 2, characterized in that The receiving voice interaction information corresponding to the first prompt information, and determining a target intention response vector according to the voice interaction information, includes: receiving voice interaction information corresponding to the first prompt information, and inputting the vehicle status and the voice interaction information into a trained second language model; After being processed by the second language model, a target intention response vector is output.

4. The method according to claim 1, wherein The acquiring of the vehicle signal tag and determining the target trapped type according to the environmental image data, the voice data, and the vehicle signal tag includes: Obtaining a vehicle signal tag, performing time stamp analysis on the vehicle signal tag, and obtaining a time series signal matrix; The environmental image data, the voice data, and the time series signal matrix are input into a trained third language model, and processed by the third language model to output the target distress type.

5. The method according to claim 4, characterized in that The third language model includes an input layer, a hidden layer and an output layer. The step of inputting the environmental image data, the voice data, and the time series signal matrix into a trained third language model, processing the data through the third language model, and outputting the target distress type includes: Inputting the environmental image data, the voice data and the time series signal matrix into a third language model through an input layer; Using a hidden layer, constructing an initial event graph based on the environmental image data, the voice data, and the time series signal matrix; Comparing the initial event graph with a preset event graph through a hidden layer to determine the target escape scenario and target trapped type corresponding to the initial event graph; The target trapped type and the target escape scenario are output through the output layer.

6. The method according to claim 1, characterized in that The fusing the vehicle state, the target intention response vector, and the target trapped type to obtain a target fusion vector includes: Determining a vehicle state vector corresponding to the vehicle state and a distress type vector corresponding to the target distress type; Mapping and aligning the vehicle state vector, the target intention response vector, and the trapped type vector to obtain a vehicle state embedding vector, a target intention response embedding vector, and a trapped type embedding vector; Determine a target fusion weight, input the vehicle state embedding vector, the target intention response embedding vector, the trapped type embedding vector and the target fusion weight into a trained vector fusion model, and output a target fusion vector.

7. The method according to claim 1, characterized in that After obtaining the target escape control strategy, it also includes: Determine the safety level corresponding to the target escape control strategy; In response to the safety level being greater than a preset level threshold, outputting a second prompt message including the target escape control strategy; or, In response to the safety level being less than a preset level threshold, executing the target escape control strategy.

8. A vehicle escape device, characterized in that: include: a data acquisition module configured to acquire image data and voice data of an environment in which the vehicle is located, and determine a target intention response vector in response to determining that the vehicle is in a trapped state based on the image data and the voice data; a trapped cause vector determination module, configured to obtain a vehicle signal label and determine a target trapped type based on the environmental image data, the voice data, and the vehicle signal label; a fusion vector determination module configured to fuse the vehicle state, the target intention response vector, and the target trapped type to obtain a target fusion vector; The escape control strategy determination module is configured to input the target fusion vector into a pre-trained escape processing model, and obtain a target escape control strategy through processing by the escape processing model.

9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

10. A vehicle, characterized in that: The vehicle includes the electronic device according to claim 9.