Vehicle control method and device, electronic equipment and storage medium

By acquiring multi-source heterogeneous data to generate synchronized data with a unified time base and converting it into context vectors, the real-time and adaptability issues of multimedia systems in existing technologies are solved, enabling personalized immersive experiences and secure adaptability, thereby improving the user experience.

CN121361473AActive Publication Date: 2026-01-20CHONGQING LANDIAN AUTOMOBILE TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511936564.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-01-20
Estimated Expiration
2045-12-22

AI Technical Summary

Technical Problem

Existing smart cockpit multimedia systems achieve multimedia content linkage with in-vehicle hardware through pre-recording and static binding, resulting in a poor user experience, lack of real-time performance, adaptability, and personalization, and difficulty in providing a personalized immersive experience.

Method used

By acquiring data from multiple sensors and user preferences, synchronized data with a unified time base is generated, converted into a unified numerical feature representation, fused into a context vector, input into the narrative generation engine, generates narrative content, and controls the vehicle at key event points, achieving hardware collaboration with cross-sensory and precise time synchronization.

Benefits of technology

It achieves a personalized and immersive experience, generating narrative content in real time based on vehicle and user status, enhancing the personalization, novelty, and immersion of the user experience, while also possessing safety adaptability to ensure driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121361473A_ABST
    Figure CN121361473A_ABST
Patent Text Reader

Abstract

The invention relates to a vehicle control method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining data generated by a plurality of sensors in a vehicle and pre-stored user preference data to obtain corresponding multi-source heterogeneous data, and carrying out the alignment processing of the multi-source heterogeneous data in a time dimension, generating synchronous data with a unified time reference; converting different types of data in the synchronous data into uniform numeralization feature representation, and fusing all numeralization features to generate a situation vector of a target dimension; inputting the situation vector into a narrative generation engine, generating narrative content and generating an abstract sensory instruction set associated with key event points in the narrative content; and controlling the vehicle to play the narrative content, and controlling the vehicle based on instructions in the abstract sensory instruction set when the key event point in the narrative content is played.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicle control technology, and particularly relates to a vehicle control method and device, an electronic device and a storage medium. BACKGROUND

[0002] The existing intelligent cockpit multimedia system mainly focuses on realizing the functional atomization capability, such as the calling and cooperation of navigation, playing, air conditioning and the like, and usually adopts a pre-recording and static binding manner to realize the linkage of multimedia content and vehicle hardware. Such a mechanism is usually a one-way instruction flow from content to hardware, lacks real-time, adaptability and personalization, and leads to poor user experience. It can be seen that the existing technology completely relies on a pre-prepared content library, and cannot dynamically and instantaneously generate a unique narrative content according to the personalized preferences of a driver and the current unique geographical location, and is difficult to convert a driving journey into a continuous and deeply personalized immersive experience.

[0003] In view of the above technical problems in the prior art, there is currently no effective solution. SUMMARY

[0004] The present application provides a vehicle control method and device, an electronic device and a storage medium to solve the problem that the prior art adopts a pre-recording and static binding manner to realize the linkage of multimedia content and vehicle hardware, leading to poor user experience.

[0005] In a first aspect, the present application provides a vehicle control method, comprising: obtaining data generated from a plurality of sensors in a vehicle and pre-stored user preference data to obtain corresponding multi-source heterogeneous data, and performing alignment processing on the multi-source heterogeneous data in a time dimension to generate synchronized data with a unified time reference, wherein the plurality of sensors at least include a sensor for detecting vehicle data and a sensor for detecting a target object in the vehicle; converting different types of data in the synchronized data into unified numerical feature representation, and fusing each numerical feature to generate a context vector of a target dimension, wherein the context vector represents a vector of a vehicle state and a target object state; inputting the context vector into a narrative generation engine, and generating narrative content and an abstract sensory instruction set associated with a key event point in the narrative content; wherein the instructions in the abstract sensory instruction set are instructions for controlling the vehicle at the key event point in the process of playing the narrative content in the vehicle; controlling the vehicle to play the narrative content, and based on the instructions in the abstract sensory instruction set, controlling the vehicle at the key event point in playing the narrative content.

[0006] In a second aspect, the application provides a control device of a vehicle, comprising: a first processing module configured to obtain data generated by a plurality of sensors in the vehicle and pre-stored user preference data to obtain corresponding multi-source heterogeneous data, and perform alignment processing on the multi-source heterogeneous data in a time dimension to generate synchronized data with a unified time reference, wherein the plurality of sensors at least include a sensor for detecting vehicle data and a sensor for detecting a target object in the vehicle; a second processing module configured to convert different types of data in the synchronized data into unified numerical feature representation, and fuse each numerical feature to generate a context vector of a target dimension, wherein the context vector represents a vector of vehicle state and target object state; a third processing module configured to input the context vector into a narrative generation engine, and generate a narrative content and an abstract sensory instruction set associated with a key event point in the narrative content; wherein the instructions in the abstract sensory instruction set are instructions for controlling the vehicle at the key event point during playing of the narrative content in the vehicle; and a fourth processing module configured to control the vehicle to play the narrative content, and control the vehicle based on the instructions in the abstract sensory instruction set at the key event point in playing the narrative content.

[0007] In a third aspect, the application provides an electronic device, comprising: at least one communication interface; at least one bus connected with the at least one communication interface; at least one processor connected with the at least one bus; and at least one memory connected with the at least one bus, wherein the processor is configured to execute the control method of the vehicle according to the first aspect of the application.

[0008] In a fourth aspect, the application further provides a computer storage medium storing computer executable instructions for executing the control method of the vehicle according to the first aspect of the application.

[0009] Compared with the prior art, the technical scheme provided by the embodiment of the present application has the following advantages: through the method provided by the embodiment of the present application, data generated from multiple sensors in the vehicle and pre-stored user preference data can be obtained to obtain corresponding multi-source heterogeneous data, and the multi-source heterogeneous data is aligned in the time dimension to generate synchronized data with a unified time reference; then, different types of data in the synchronized data are converted into unified numerical feature representation, and each numerical feature is fused to generate a context vector of the target dimension, and then the context vector is input into a narrative generation engine, and a narrative content and an abstract sensory instruction set associated with a key event point in the narrative content are generated; finally, the vehicle plays the narrative content, and when the vehicle plays the key event point in the narrative content, the vehicle is controlled based on the instructions in the abstract sensory instruction set. It can be seen that in the embodiment of the present application, the state associated with the vehicle and the state of the user in the vehicle can be obtained in real time, and then corresponding narrative content is generated according to the real-time states of the two, accompanied by cross-sensory, time-accurate synchronous vehicle hardware cooperative control instructions (instructions in the abstract sensory instruction set), which provides personalized immersive experience for the user. BRIEF DESCRIPTION OF DRAWINGS

[0010] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate an embodiment consistent with the present application and, together with the description, serve to explain the principles of the application.

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced below. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without creative labor.

[0012] One or more embodiments are exemplarily illustrated by pictures in the drawings corresponding thereto, and these exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, unless otherwise specified. The drawings in the drawings do not constitute a proportional limitation.

[0013] Figure 1 A flowchart of a vehicle control method provided by the embodiment of the present application; Figure 2 An optional flowchart of a vehicle control method provided by the embodiment of the present application; Figure 3 A structural schematic diagram of the seat cabin multi-sensory narrative content real-time generation and cooperative control system based on context perception and user portrait provided by the embodiment of the present application; Figure 4A flowchart of the method for real-time generation and collaborative control of multi-sensory narrative content in the cockpit based on context awareness and user portrait is provided for the embodiments of the present application. Figure 5 A structural schematic diagram of the control device of the vehicle is provided for the embodiments of the present application. Figure 6 A structural schematic diagram of the electronic device is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0015] The following disclosure provides many different embodiments, or examples, for implementing different structures of the present application. For the purpose of simplicity, the description of a particular example will not necessarily be repeated in the description of each example. Of course, they are only examples and are not intended to limit the present application. In addition, reference numerals and / or letters can be repeated in different examples. Such repetition is for the purpose of simplicity and clarity, and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0016] To solve the problem that the user experience is poor due to the pre-recording and static binding way to realize the linkage of multimedia content and vehicle hardware in the prior art, the present application provides a control method of a vehicle, as shown in Figure 1 The steps of the method include: Step 101, obtaining data generated from a plurality of sensors in the vehicle and pre-stored user preference data to obtain corresponding multi-source heterogeneous data, and performing alignment processing on the multi-source heterogeneous data in the time dimension to generate synchronized data with a unified time reference, wherein the plurality of sensors at least include a sensor for detecting vehicle data and a sensor for detecting a target object in the vehicle; The target object in the embodiments of the present application refers to a user in the vehicle cabin, such as a driver and / or a passenger. In addition, in the embodiments of the present application, data collection can be started after detecting a specific event, such as a voice wake-up word, an emergency brake signal, a state mutation of the target object, etc. In addition, data collection is periodic data collection, such as a collection period of 500 ms, which can also be adjusted according to actual needs, and the collection period is dynamically adjusted according to the load and scene complexity. For example, the period can be appropriately lengthened (such as 800 ms) when driving on a highway, and shortened (such as 300 ms) in complex urban road conditions.

[0017] The specific collected data can be vehicle environment data, user physiological data, and user behavior data. Further, the vehicle environment data includes current vehicle position data (obtained through a GPS positioning system), light intensity (obtained through a light sensor arranged on the vehicle), weather information (obtained from a server through a weather interface), and vehicle speed (obtained through a vehicle speed sensor arranged on the vehicle), and the like. The user physiological data refers to physiological data of the target object, and specifically can include heart rate (detected through a smart wearable device and then sent to the server and then sent to the vehicle controller), galvanic skin response (detected through a smart wearable device and then sent to the server and then sent to the vehicle controller), facial expression (obtained through a camera arranged in the cabin), seat pressure distribution (obtained through a pressure sensor arranged below the seat), and the like. The user behavior data is also data associated with the target object, and specifically can include interaction events (such as voice instructions and touch operations), and the like. The user preference data can refer to user portrait data, and specifically can be preference labels obtained from the vehicle cloud or the vehicle locally, such as preference for suspense stories and preference for children mode, and the like.

[0018] It should be noted that the more rich the data types in the multi-source heterogeneous data in the embodiments of the present application and the more data under each data type, the richer the subsequent generated scenario vector representing the vehicle state and the target object state, and thus the more in line with the needs of the current vehicle state and the user state for generating an abstract sensory instruction set based on the scenario vector. That is, the multi-source heterogeneous data obtained through synchronous collection by multiple sensors can construct a comprehensive real-time scenario portrait, providing rich and accurate input for subsequent generation of narrative content.

[0019] Further, in the embodiments of the present application, the multi-source heterogeneous data is aligned in the time dimension to solve the problem of inconsistent sampling frequencies of various sensors in the vehicle. Specifically, low-frequency data (such as GPS) can be aligned to the current timestamp using nearest neighbor interpolation; high-frequency continuous data (such as heart rate) can calculate statistical features (mean, variance) within a sliding window; and event data can be judged whether to be included in the current frame according to the timestamp. After alignment in the time dimension, the situation misjudgment caused by data delay can be avoided, and the synchronization accuracy of the narrative content and the hardware linkage in the vehicle can be improved.

[0020] In step 102, different types of data in the synchronized data are converted into unified numerical feature representations, and the numerical features are fused to generate a scenario vector of a target dimension, wherein the scenario vector represents a vector of the vehicle state and the target object state. It can be seen that various data can be converted into a unified numerical representation through this step. Specifically, continuous numerical data can be normalized using Min-Max normalization or logarithmic compression, and categorical data (such as location labels) can be converted into one-hot encoding or low-dimensional embedding vectors. Finally, the structured feature vectors, i.e., scenario vectors, are obtained by concatenating the above-mentioned normalized and encoded data. Through the above normalization and encoding processing, the inference efficiency of the subsequent machine learning model (narrative generation engine) can be improved.

[0021] In addition, the dimension of the scenario vector in the embodiments of the present application can be a high-dimensional scenario vector with a dimension of 256 / 512. The high-dimensional vector can accommodate complex multi-modal information and provide rich semantic input for the subsequent narrative generation engine.

[0022] In step 103, the scenario vector is input into the narrative generation engine, and a narrative content and an abstract sensory instruction set associated with a key event point in the narrative content are generated. The instructions in the abstract sensory instruction set are used to control the vehicle at the key event point during the playing of the narrative content in the vehicle. In the embodiments of the present application, the generated narrative content is generated according to the vehicle state and the user state, i.e., it conforms to the current real-time state. For example, a suspenseful story is generated according to the current real-time state, or a light-hearted and humorous story is generated. The corresponding abstract sensory instruction set can include instructions associated with various senses. For example, if a light-hearted story is needed to relax, the instructions for controlling the light color will control the light in the cabin to gradually change to blue at the beginning of the narrative content, control the vibration intensity in the cabin to be soft waves when introducing characters in the narrative content, and control the sound effect type in the cabin to be natural wind sound when introducing scenery in the narrative content. Moreover, the instructions in the abstract sensory instruction set are matched with the specific situation in the narrative content at the key points during the playing of the narrative content. Therefore, after generating the narrative content, the key event points in the narrative content are determined, such as a key event point at the beginning of the narrative content, a key event point when introducing characters in the narrative content, and a key event point when introducing scenery in the narrative content. Based on these key event points, the corresponding instructions are generated, thereby obtaining the abstract sensory instruction set.

[0023] In step 104, the vehicle is controlled to play the narrative content, and at the key event point in the playing of the narrative content, the vehicle is controlled based on the instructions in the abstract sensory instruction set.

[0024] It can be seen that through the above steps 101 to 104, the data generated from the plurality of sensors in the vehicle and the pre-stored user preference data can be obtained to obtain corresponding multi-source heterogeneous data, and the multi-source heterogeneous data is aligned in the time dimension to generate synchronized data with a unified time reference; then, different types of data in the synchronized data are converted into unified numerical feature representations, and the numerical features are fused to generate a context vector of the target dimension, and then the context vector is input into the narrative generation engine, and a narrative content and an abstract sensory instruction set associated with a key event point in the narrative content are generated; finally, the vehicle plays the narrative content, and at the key event point in the playing of the narrative content, the vehicle is controlled based on the instructions in the abstract sensory instruction set. It can be seen that in the embodiment of the present application, the state associated with the vehicle and the state of the user in the vehicle can be obtained in real time, and then the corresponding narrative content is generated according to the real-time states of the two, accompanied by cross-sensory, time-accurate synchronous vehicle hardware cooperative control instructions (instructions in the abstract sensory instruction set), which provides personalized immersive experience for the user.

[0025] In an optional implementation of the embodiment of the present application, for the manner of aligning the multi-source heterogeneous data in the time dimension involved in the above step 101, further comprising: Step 11, for the first data in the multi-source heterogeneous data with a sampling frequency lower than a first preset threshold, acquiring first target data aligned to the target moment from the first data by using the nearest neighbor value method; Step 12, for the second data in the multi-source heterogeneous data with a sampling frequency higher than the first preset threshold, taking the statistical feature value corresponding to the second data in a preset time window including the target moment as the second target data aligned to the target moment; Step 13, for the image frame in the multi-source heterogeneous data, acquiring frame data at the target moment or third target data near the target moment; Since the user facial expression data collected by the camera, i.e., the image data collected by the camera, is stored as an image frame with a time label in the multi-source heterogeneous data.

[0026] Step 14, for the event type data in the multi-source heterogeneous data, acquiring fourth target data within a preset time range centered on the target moment; Step 15, determining the first target data, the second target data, the third target data and the fourth target data as the data aligned to the target moment in the multi-source heterogeneous data.

[0027] To this end, in a specific example, the first data is GPS data, weather data, etc., and the second data can be user heart rate, illumination data, etc. The event type data can be voice instructions, key instructions, etc. Therefore, the nearest neighbor value method involved in the above step 11 refers to taking the data point closest to the alignment time t, and if the delay exceeds the threshold, it is considered invalid. Taking the GPS data as an example, and the current GPS data recorded within the last one hour, the data at the alignment time t is found by the nearest neighbor value method, such as the GPS data at the time corresponding to the current time minus 15 minutes as the first target data. The sliding window statistical feature extraction method involved in the above step 12 can be: taking a window W (such as 1 second) before and after the alignment time t, calculating the mean, variance, slope, etc. in the window as the representative value at time t. Taking the user heart rate as an example, and the current user heart rate data recorded within the last one hour, then by sliding window method, sliding to the time corresponding to the current time minus 15 minutes, the user heart rate data within 1s (sliding window) before and after the time is obtained, and the obtained user heart rate data is the second target data. The frame data at the target time or the data near the target time involved in the above step 13 refers to, if there is no corresponding frame at time t, taking the nearest two frames before and after for linear interpolation (such as target detection box position), or directly taking the nearest frame; such as the current image frame recorded within the last one hour, the third target data is the image frame at the time corresponding to the current time minus 15 minutes, if there is no image frame at the time, the image frame obtained by linear interpolation according to the two frames before and after the time is the third target data, or the nearest image frame is directly determined as the third target data. The above step 14 refers to matching by timestamp, which can be, if the event occurs within the range of [t Δt, t] (such as Δt=200ms), the event is classified into the current frame, and the trigger time and type are recorded. Taking the voice instruction event as an example, if the voice instruction event corresponding to the time near the time corresponding to the current time minus 15 minutes is obtained and taken as the fourth target data. Since the data types in the multi-source heterogeneous data include various types, after the time sequence consistency is unified, the situation misjudgment caused by data delay can be avoided, and the synchronization accuracy of the narrative content and vehicle control linkage is improved.

[0028] In the optional implementation of the embodiments of the present application, the above step 102 involved in converting different types of data in the synchronization data into unified numerical feature representation and fusing the numerical features to generate the situation vector of the target dimension can further include: Step 21, scaling the continuous numerical data in the synchronization data based on a preset normalization interval to obtain the corresponding numerical feature representation; To this end, in specific examples, continuous numerical data can be heart rate, illumination, vehicle speed, etc.; physical signals with different dimensions and value ranges are mapped to a unified numerical interval (such as [0, 1]), eliminating the influence of dimensions; Specifically, it can be, for example, the preset range [50, 120] bpm of the heart rate, and the measured 72 bpm is normalized, that is, (72-50) / (120-50), and the heart rate is 0.31. Or, the preset range [0, 1000] lux of the light intensity, the measured light intensity is 200 lux, and the normalized processing is 200 / 1000, and the light intensity is 0.2.

[0029] Step 22, converting discrete data in the synchronization data into a one-hot encoding vector or a low-dimensional dense vector, wherein the one-hot encoding vector and the low-dimensional dense vector represent the corresponding data through numerical characteristics; The purpose of this step 22 is to convert non-numerical category information into a numerical vector that the model can understand, while retaining as much semantic information as possible. Specifically, one-hot encoding assigns a unique integer index to each possible category and creates a vector with a length equal to the total number of categories. The vector is 1 at the corresponding index and 0 elsewhere. For example, the current discrete data is music style, which supports [“pop”, “classical”, “electronic”, “suspense music”] four styles. Among them, suspense music corresponds to index 3, which is encoded as vector [0, 0, 0, 1]. In addition, in the embodiments of the present application, each category can be mapped to a low-dimensional dense vector (such as 32 dimensions or 64 dimensions, etc.) through a pre-trained or online learning lookup table. The vector can implicitly capture the semantic similarity between categories.

[0030] Step 23, the continuous numerical data after scaling and the discrete data converted into a one-hot encoding vector or a low-dimensional dense vector are fused by vector splicing or weighted summation with attention weights to generate a context vector of the target dimension.

[0031] It can be seen that all the above processed sub-vectors are spliced according to the predefined order and dimension to form a long and fixed dimension context vector. For example, after the above steps 21 and 22, the following data is obtained: environment sub-vector (light, weather, etc.): 64 dimensions; position semantic sub-vector: 64 dimensions; user physiological sub-vector (heart rate, emotion, etc.): 64 dimensions; user preference embedding sub-vector: 64 dimensions; visual state sub-vector (eye opening degree, gaze direction, etc.): 64 dimensions; event feature sub-vector: 64 dimensions. The total dimension of the context vector after splicing: 384 dimensions. The dimensions of the above sub-vectors can be set according to actual needs. For example, in order to more comprehensively collect data of users and vehicles, data from all aspects needs to be collected, so that the elements of the sub-vectors obtained are more, and the dimensions of the corresponding sub-vectors are higher, and vice versa. Taking the environment sub-vector as an example, the elements in the current environment sub-vector can include: light intensity, temperature, humidity, noise, AQI (air quality index), time, weather, season and week. After time alignment and normalization of each element data, and converting it into a one-hot encoding vector or a low-dimensional dense vector, the environment sub-vector is obtained as follows: 0.72, # light 0.5, -0.25, # temperature and humidity 0.3, 0.7, # noise and AQI 0.97, -0.24, # time encoding 0.9, -0.2, 0.4, 0.1, # weather embedding 0.1, 0.8, -0.3, 0.2, # season embedding 0.97, -0.22, # week encoding ] It can be seen that the dimension of the environment sub-vector is 17. The above is only an example, and the specific dimension can be determined according to the number of required elements. In addition, other sub-vectors are processed in a similar manner, which will not be described here.

[0032] In an optional embodiment of the present application, the manner of inputting the context vector into the narrative generation engine and generating narrative content in step 103 can further include: Step 31, determining a corresponding target template from a plurality of preset narrative templates based on the location label, user preference label and environment and time label in the context vector; Step 32, inputting the context vector as input to drive the target template to generate narrative content associated with the current location and user preference.

[0033] To this end, in specific examples, the geographic / location tags can be, for example, mountain road, business district, specific place (broken bridge). The user profile tags: for example, children present, preference for suspense, emotional state (bored). The environment and time tags: for example, night, rainy day, weekend. In specific application scenarios, the tag content specifically includes historical sites, children, preference for suspense, based on which a child-friendly myth suspense model is selected, and corresponding narrative content related to a children's suspense story is generated based on the child-friendly myth suspense model.

[0034] In an optional implementation of the embodiment of the present application, for the manner of controlling the vehicle based on the instructions in the abstract sensory instruction set involved in the above step 104, further can include: Step 41, dynamically mapping each abstract sensory instruction in the abstract sensory instruction set into an executable control command of the corresponding hardware device in the vehicle cabin; Step 42, sending the executable control command to the corresponding cabin executor to drive multiple hardware devices in the vehicle cabin to output in coordination and present a multi-sensory narrative experience synchronized with the key event point in the narrative content.

[0035] For the steps 41 and 42, taking the narrative content generated above as the narrative content related to a children's suspense story as an example, specifically generating the narrative content of the Legend of White Snake, and generating corresponding abstract sensory instructions based on the key event point of the appearance of the White Snake to create a tense suspense atmosphere, therefore, during the playing of the narrative content, when playing to the appearance of the White Snake, the tactile experience is controlled by multiple instructions in the abstract sensory instruction set: the seat vibrates slightly, the intensity is 70%; the visual experience is generated: the light is pulse red light, the frequency is 1Hz; the auditory experience is generated: the audio is added with a low-frequency thunder sound effect, the sound image is located in front. It can be seen that in the embodiment of the present application, corresponding narrative content can be generated according to the real-time vehicle state and user state, and the multi-sensory narrative experience of the user can also be increased during the playing of the narrative content.

[0036] In the embodiment of the present application, before controlling the vehicle to play the narrative content and at the key event point in the playing of the narrative content, based on the instructions in the abstract sensory instruction set, the vehicle is controlled, in order to ensure the safety of the vehicle, for example Figure 2 As shown, the method of the embodiment of the present application further includes: Step 201, acquiring the current vehicle state and the driver's attention state; Step 202, risk assessment of the abstract sensory instruction set based on the vehicle state and the driver's attention state; In step 203, based on the risk assessment result, the target instruction in the abstract sensory instruction set is adjusted in degradation or selectively shielded to obtain the final abstract sensory instruction set, wherein the target instruction refers to the instruction that currently affects the driving safety in the risk assessment result.

[0037] In the embodiment of the present application, the vehicle state includes the driving state data of the vehicle, such as the driving speed of the vehicle, the use of the vehicle light during driving, the use of the auxiliary driving device, etc. The real-time image data of the driver can be collected through the camera arranged in the cockpit, the face recognition and hand action recognition are performed on the collected user image data, and the recognition result is compared with the preset image to obtain the driver attention state, such as the expression of the user, the degree of eye opening, the state of driving posture, etc. Therefore, for the above steps 201 to 203, the narrative content of the White Snake is generated as an example, and in the process of playing the White Snake, the tactile experience is generated through the abstract sensory instruction control: the seat slightly vibrates with an intensity of 70%; the visual experience is generated: the light is a pulse red light with a frequency of 1 Hz; the auditory experience is generated: the audio is added with a low-frequency thunder sound effect, and the sound image is positioned in front. The risk assessment is obtained based on the vehicle state and the driver attention state, and is used to evaluate how the driving safety is in the current narrative content playing. If the current user is determined to be in a state of fatigue and high-speed driving through the vehicle state and the driver attention state at this time, if the instruction of starting the light control will seriously affect the driving safety, then the risk assessment result at this time is an instruction with a high degree of danger, and therefore the instruction needs to be selectively shielded. As for the seat slight vibration instruction, if the driver is in a state of fatigue at this time when driving at high speed, then the risk assessment result at this time is an instruction with a moderate degree of danger, because the vibration can relieve fatigue at this time, but there is also a certain degree of danger, and the instruction can be adjusted in degradation after the two are combined, that is, the vibration intensity is reduced. As can be seen, in the embodiment of the present application, if the risk assessment result is high degree of danger, the target instruction is selectively shielded, if the risk assessment result is moderate degree of danger, the target instruction can be adjusted in degradation, and if the risk assessment result is light degree, the instruction in the abstract sensory instruction set does not need to be adjusted, and the instruction in the abstract sensory instruction set can be executed normally. It can be seen that through the risk assessment, the multi-sensory narrative experience of the user can be ensured while the driving safety of the vehicle is ensured.

[0038] Further, the degradation adjustment on the target instruction in the abstract sensory instruction set in the embodiment of the present application comprises: reducing the intensity parameter of the target instruction, shortening the duration of the target instruction, or limiting the action range of the target instruction according to a preset strategy, wherein the preset strategy represents that different risk levels of instructions correspond to different degradation adjustment strategies. Wherein, according to the preset strategy, the intensity parameter of the target instruction is reduced, the duration of the target instruction is shortened, or the action range of the target instruction is limited, which comprises: when the risk level is the first level, only the instruction intensity of the non-critical sensory channel is reduced; when the risk level is the second level, all instructions involving physical tactile feedback are closed; when the risk level is the third level, all narrative content output and multi-sensory linkage are interrupted, and a safety prompt mode is switched to, wherein the higher the risk level represents the higher the danger degree of the current instruction execution.

[0039] The present application will be further explained and described in conjunction with the specific implementation manners of the embodiments of the present application. The specific implementation manners provide a cockpit multi-sensory narrative content real-time generation and collaborative control system and method based on situational awareness and user portrait, to solve the problems of content staticization, linkage rigidity and lack of safety adaptability of the existing cockpit multimedia system. As shown in Figure 3 The cockpit multi-sensory narrative content real-time generation and collaborative control system based on situational awareness and user portrait comprises: A collection module is configured to collect vehicle state information and user preference information, such as GPS, camera, seat heart rate, illumination, vehicle speed, and user account preferences. Specifically, data collection can be performed by sensors placed at different positions in the vehicle, such as seat sensors, steering wheel cameras, rear cameras, etc.

[0040] A synchronization module is configured to align the data of different sensors at the same time point, to solve the problem of sampling frequency difference of each sensor. The synchronization module is arranged in the main control computing unit of the vehicle, so that it can receive all sensor inputs.

[0041] A normalization module is configured to convert various numerical values and categories into a unified digital vector (for example, converting heart rate, illumination, and labels into numerical values or small vectors within [0, 1]), to facilitate subsequent processing.

[0042] A live vector construction module is configured to splice all normalized features into a fixed-length vector (such as 256 / 512 dimensions), to represent the complete situation at the current time.

[0043] A narrative generation engine (EGOE) is configured to take the constructed high-dimensional vector as input, select a suitable narrative template or model, generate a story text / voice in real time, and synchronously generate abstract sensory instructions (original parameters), such as “pulse light at climax, rear light vibration, increase wind sound effect”.

[0044] An ASDM (Assessment and Degradation Module) is used to assess whether to allow or require degradation / masking of sensory effects based on vehicle speed, braking, driver attention, etc. to ensure driving safety.

[0045] A mapping and execution module is used to translate the abstract instructions output by the EGOE / ASDM into specific hardware commands (such as mapping "blue pulse" to specific RGB values, frequency, seat motor instructions), and executing on the subsystems (lights, sound, seat).

[0046] Based on this, the specific embodiment provides a cockpit multi-sensory narrative content real-time generation and collaborative control method based on context awareness and user portrait, as shown in the figure, the steps of the method include: Figure 4 As shown in the figure, the steps of the method include: Step 401, synchronously collecting the latest data packets from various sensors, such as GPS, camera frames, heart rate streams, illumination, user interaction events, etc., and event type data, such as voice, buttons, occurring within the last 200 ms window, which can also be collected within the collection period.

[0047] Step 402, time alignment; Among them, the low-frequency data (GPS) is aligned by taking the nearest neighbor sample to the current frame, and for high-frequency continuous data (heart rate, GSR), the mean / variance is taken in [t W, t] to align (W is generally 1-2s), and for camera / video frames, linear interpolation or the nearest frame can be used for alignment, and event type is determined by timestamp whether it belongs to the current frame.

[0048] Step 403, normalization and encoding, in which numerical values are directly normalized by Min-Max or log according to the preset interval; category / label is converted to one-hot or small dimension embedding vector, and then the two are merged to obtain a fixed length vector segment. Step 404, constructing RCV, in which all sub-vectors are concatenated or fused (spliced or weighted sum) according to model requirements, and a high-dimensional vector can be obtained, such as a 256 / 512-dimensional RCV(t).

[0049] Step 405, template selection, in which a narrative template is selected according to the obvious labels (position, preference, occupant role) in RCV (for example, "broken bridge + suspense + child → child-friendly myth suspense template").

[0050] Step 406, narrative and original parameter generation, in which key event points (KEP) are annotated during the generation process. Abstract sensory parameters are output for each KEP: module, effect name, intensity, duration, target seat, priority, trigger time offset, etc.

[0051] Step 407, security check / demotion process; wherein, vehicle dynamics and driver physiological indicators are read: if high risk is detected (e.g. emergency braking, lane deviation, driver pupil dilation), replace or reduce the parameters given by EGOE (e.g. set vibration intensity to 0, change light to weak static color). If an emergency situation in the vehicle is detected (e.g. child crying, someone taking off the seat belt), jump to the pacification or stop strategy.

[0052] Step 408, data packaging, wherein the abstract instructions after security are packaged into a multi-modal instruction set (containing timestamp, priority, rollback strategy) on the timeline.

[0053] Step 409, hardware mapping and execution, i.e. according to the vehicle's hardware capability table, abstract instructions are converted into hardware API calls (e.g. LED.setRGB, Audio.playSound, SeatHaptic.setPattern). Further, the instructions are issued and executed synchronously through the vehicle's bus. Record execution feedback (whether the execution is successful, whether the sensor detects the expected effect) backflow for online calibration.

[0054] Step 410, closed-loop learning / log, i.e. RCV, generated content, ASDM behavior and execution feedback are saved to local or cloud for model fine-tuning and user preference update.

[0055] It can be seen that, through the present application, the content generation is changed from static playing to real-time creation: by introducing the experience generation and optimization engine (EGOE) and taking the real-time context-aware vector (RCV) as input, the system can instantly and automatically generate narrative content that highly fits the current geographical location, environmental state and user preferences. This solves the defects of existing technology that the content is rigid and cannot adapt to changes in the journey, greatly improving the personalization, freshness and uniqueness of the user experience. Moreover, through the present application, the depth and efficiency of the cockpit immersive experience can also be significantly improved: the EGOE directly outputs multi-sensory output raw parameters, and via a hardware instruction dynamic mapper, the emotional climax and key events in the narrative content are synchronously and accurately converted into cross-sensory coordinated actions (including light and shadow, sound effects, somatosensory, etc.). This end-to-end, zero-intermediary generation-to-execution path ensures the time synchronization accuracy and response speed of the immersive linkage, eliminates the delay caused by preset labeling, thereby enhancing the psychological / emotional management value of the system; that is, by taking the user mood curve as a key input of the RCV, the system can avoid generating narrative content and linkage effects that may exacerbate road rage or anxiety (such as through narrative content emotion filtering). This not only improves the user experience, but also gives the cockpit system the added value of assisting emotional management and optimizing the driving experience. In addition, in the present application, a context-adaptive safety control mechanism is also constructed: that is, the system integrates an active safety degradation module that forcibly checks the driver state and vehicle dynamics before executing the narrative instructions. This solves the potential safety hazards in the pursuit of immersion in the prior art, ensuring that all multi-sensory outputs (especially somatosensory vibrations involving physical feedback) can be degraded or canceled in real time and automatically in high-risk driving situations, ensuring driving safety and the driver's attention.

[0056] Corresponding to the above Figure 1 , the embodiment of the present application provides a control device of a vehicle, as shown in Figure 5 , the device comprises: A first processing module 502 is configured to obtain data generated from a plurality of sensors in the vehicle and pre-stored user preference data to obtain corresponding multi-source heterogeneous data, and perform alignment processing on the multi-source heterogeneous data in the time dimension to generate synchronized data with a unified time reference, wherein the plurality of sensors at least include a sensor for detecting vehicle data and a sensor for detecting a target object in the vehicle; A second processing module 504 is configured to convert different types of data in the synchronized data into a unified numerical feature representation, and fuse each numerical feature to generate a context vector of a target dimension, wherein the context vector represents a vector of the vehicle state and the target object state; The third processing module 506 is configured to input the context vector into a narration generation engine, and generate narration content and an abstract sensory instruction set associated with a key event point in the narration content; wherein an instruction in the abstract sensory instruction set is an instruction for controlling the vehicle at the key event point in the process of playing the narration content by the vehicle. The fourth processing module 508 is configured to control the vehicle to play the narration content, and control the vehicle based on the instruction in the abstract sensory instruction set when playing the key event point in the narration content.

[0057] In an optional implementation of the embodiments of the present application, the first processing module in the embodiments of the present application can further include: a first processing unit configured to acquire, from the first data, first target data aligned to the target moment by using a nearest neighbor value method for the first data in the multi-source heterogeneous data with a sampling frequency lower than a first preset threshold; a second processing unit configured to acquire, as second target data aligned to the target moment, a statistical feature value corresponding to the second data in a preset time window including the target moment for the second data in the multi-source heterogeneous data with a sampling frequency higher than the first preset threshold; a third processing unit configured to acquire, as third target data, frame data at the target moment or frame data close to the target moment for the image frames in the multi-source heterogeneous data; a fourth processing unit configured to acquire, as fourth target data, data in a preset time range centered on the target moment for the event type data in the multi-source heterogeneous data; and a fifth processing unit configured to determine the first target data, the second target data, the third target data, and the fourth target data as data in the multi-source heterogeneous data aligned to the target moment.

[0058] In an optional implementation of the embodiments of the present application, the second processing module in the embodiments of the present application can further include: a scaling unit configured to scale continuous numerical data in the synchronization data based on a preset normalization interval to obtain a corresponding numerical feature representation; a conversion unit configured to convert discrete data in the synchronization data into a one-hot encoding vector or a low-dimensional dense vector, wherein the one-hot encoding vector and the low-dimensional dense vector correspond to data through the numerical feature representation; and a sixth processing unit configured to fuse the continuous numerical data after the scaling processing and the discrete data converted into the one-hot encoding vector or the low-dimensional dense vector by using a vector splicing manner or a weighted summation manner with attention weights to generate a context vector of a target dimension.

[0059] In an optional implementation of the embodiments of the present application, the third processing module in the embodiments of the present application can further include: a determination unit configured to determine a corresponding target template from a plurality of preset narration templates based on a location label, a user preference label, and an environment and time label in the context vector; and a seventh processing unit configured to take the context vector as an input to drive the target template to generate narration content associated with a current location and a user preference.

[0060] In an optional implementation of the embodiment of the application, the fourth processing module in the embodiment of the application further can comprise: a mapping unit, configured to dynamically map each abstract sensory instruction in the abstract sensory instruction set into an executable control command of a corresponding hardware device in the vehicle cabin; and an eighth processing unit, configured to send the executable control command to the corresponding cabin executor to drive the multiple hardware devices in the vehicle cabin to perform collaborative output and present the multi-sensory narrative experience synchronized with the key event point in the narrative content.

[0061] In an optional implementation of the embodiment of the application, the device in the embodiment of the application further comprises: an acquisition module, configured to acquire the current vehicle state and the driver attention state before controlling the vehicle to play the narrative content and to control the vehicle based on the instructions in the abstract sensory instruction set when playing the key event point in the narrative content; a fifth processing module, configured to perform risk assessment on the abstract sensory instruction set based on the vehicle state and the driver attention state; and a sixth processing module, configured to perform degradation adjustment or selective shielding on a target instruction in the abstract sensory instruction set based on the risk assessment result to obtain a final abstract sensory instruction set, wherein the target instruction is an instruction that currently affects the driving safety in the risk assessment result.

[0062] In an optional implementation of the embodiment of the application, the sixth processing module in the embodiment of the application comprises: a ninth processing unit, configured to reduce the intensity parameter of the target instruction, shorten the duration of the target instruction, or limit the action range of the target instruction according to a preset strategy, wherein the preset strategy represents that instructions of different risk levels correspond to different degradation adjustment strategies. Wherein, reducing the intensity parameter of the target instruction, shortening the duration of the target instruction, or limiting the action range of the target instruction according to the preset strategy means: when the risk level is the first level, only reducing the instruction intensity of the non-key sensory channel; when the risk level is the second level, closing all instructions involving physical tactile feedback; when the risk level is the third level, interrupting all narrative content output and multi-sensory linkage, and switching to a safety prompt mode, wherein the higher the risk level, the higher the risk degree of the current instruction execution.

[0063] As shown in FIG. 6, Figure 6 The embodiment of the application provides an electronic device, which comprises a processor 611, a communication interface 612, a memory 613 and a communication bus 614, wherein the processor 611, the communication interface 612 and the memory 613 complete mutual communication through the communication bus 614, The memory 613 is used to store a computer program. In an embodiment of the application, the processor 611 is used to execute the program stored in the memory 613, and the vehicle control method provided by any one of the preceding method embodiments is implemented, which also has similar effects, and will not be described here.

[0064] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the control method of the vehicle provided by any one of the foregoing method embodiments.

[0065] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment.

[0066] Through the description of the foregoing embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course can also be implemented by hardware. Based on such understanding, the foregoing technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, or an optical disk, and includes a plurality of instructions to cause a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0067] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises", "comprising", "includes", "including" and "has" are inclusive and therefore specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order in which they are described, unless specifically identified as an order dependent step. It is also to be understood that additional or alternative steps can be employed.

[0068] The above description is merely illustrative of the application and the specific examples, so that those skilled in the art can understand or implement the application. Various modifications to these examples will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the application. Therefore, the application will not be limited to the examples shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A control method of a vehicle, characterized by, The method comprises: obtaining data generated by a plurality of sensors in a vehicle and pre-stored user preference data to obtain corresponding multi-source heterogeneous data, and aligning the multi-source heterogeneous data in a time dimension to generate synchronized data with a unified time reference, wherein the plurality of sensors at least include a sensor for detecting vehicle data and a sensor for detecting a target object in the vehicle; converting different types of data in the synchronized data into unified numerical feature representations, and fusing each numerical feature to generate a context vector of a target dimension, wherein the context vector represents a vector of vehicle state and target object state; inputting the context vector into a narrative generation engine, and generating narrative content and an abstract sensory instruction set associated with a key event point in the narrative content; wherein the instructions in the abstract sensory instruction set are instructions for controlling the vehicle at the key event point during playing of the narrative content in the vehicle; controlling the vehicle to play the narrative content, and controlling the vehicle based on the instructions in the abstract sensory instruction set at the key event point in playing the narrative content.

2. The method of claim 1, wherein, The aligning of the multi-source heterogeneous data in the time dimension comprises: for first data in the multi-source heterogeneous data with a sampling frequency lower than a first preset threshold, acquiring first target data aligned to a target time from the first data by using a nearest neighbor value method; for second data in the multi-source heterogeneous data with a sampling frequency higher than the first preset threshold, acquiring a statistical feature value corresponding to the second data in a preset time window including the target time as second target data aligned to the target time; for image frames in the multi-source heterogeneous data, acquiring frame data at the target time or third target data adjacent to the target time; for event type data in the multi-source heterogeneous data, acquiring fourth target data within a preset time range centered on the target time; determining the first target data, the second target data, the third target data and the fourth target data as data in the multi-source heterogeneous data aligned to the target time.

3. The method of claim 1, wherein, The conversion of different types of data in the synchronized data into unified numerical feature representations, and the fusion of each numerical feature to generate a context vector of a target dimension, comprises: scaling continuous numerical data in the synchronized data based on a pre-set normalization interval to obtain corresponding numerical feature representations; converting discrete data in the synchronized data into one-hot encoding vectors or low-dimensional dense vectors, wherein the one-hot encoding vectors and the low-dimensional dense vectors correspond to data represented by numerical features; fusing the scaled continuous numerical data and the discrete data converted into one-hot encoding vectors or low-dimensional dense vectors by using vector splicing or weighted summation with attention weights to generate a context vector of a target dimension.

4. The method of claim 1, wherein, Inputting the context vector into a narrative generation engine and generating narrative content comprises: determine a corresponding target template from a plurality of preset narrative templates based on the location label, the user preference label, and the environment and time label in the context vector; input the context vector to drive the target template to generate the narrative content associated with the current location and the user preference.

5. The method of claim 1, wherein, control the vehicle based on the instructions in the abstract sensory instruction set, including: dynamically map each abstract sensory instruction in the abstract sensory instruction set into an executable control command for a corresponding hardware device in the vehicle cabin; send the executable control command to a corresponding cabin actuator to drive multiple hardware devices in the vehicle cabin to output in coordination and present a multi-sensory narrative experience synchronized with a key event point in the narrative content.

6. The method of claim 1, wherein, Before controlling the vehicle to play the narrative content and controlling the vehicle based on the instructions in the abstract sensory instruction set when playing the key event point in the narrative content, the method further includes: obtain the current vehicle state and the driver attention state; risk assess the abstract sensory instruction set based on the vehicle state and the driver attention state; based on the risk assessment result, degrade or selectively shield a target instruction in the abstract sensory instruction set to obtain a final abstract sensory instruction set, wherein the target instruction is an instruction that currently affects driving safety in the risk assessment result.

7. The method of claim 6, wherein, degrading the target instruction in the abstract sensory instruction set includes: decrease the intensity parameter of the target instruction, shorten the duration of the target instruction, or limit the action range of the target instruction according to a preset strategy, wherein the preset strategy represents that different risk levels of instructions correspond to different degradation adjustment strategies.

8. The method of claim 7, wherein, decreasing the intensity parameter of the target instruction, shortening the duration of the target instruction, or limiting the action range of the target instruction according to a preset strategy, includes: when the risk level is the first level, only decrease the instruction intensity of non-critical sensory channels; when the risk level is the second level, close all instructions involving physical tactile feedback; when the risk level is the third level, interrupt all narrative content output and multi-sensory linkage, and switch to a safety prompt mode, wherein the higher the risk level, the higher the risk degree of the current instruction execution.

9. A control device of a vehicle characterized by comprising: including: a first processing module configured to obtain data generated from a plurality of sensors in the vehicle and pre-stored user preference data to obtain corresponding multi-source heterogeneous data, and align the multi-source heterogeneous data in a time dimension to generate synchronized data with a unified time reference, wherein the plurality of sensors include at least a sensor for detecting vehicle data and a sensor for detecting a target object in the vehicle; a second processing module configured to convert different types of data in the synchronized data into a unified numerical feature representation, and fuse each numerical feature to generate a context vector of a target dimension, wherein the context vector represents a vector of vehicle state and target object state; The third processing module is configured to input the context vector into a narration generation engine, and generate narration content and an abstract sensory instruction set associated with a key event point in the narration content; wherein the instructions in the abstract sensory instruction set are instructions for controlling the vehicle at the key event point in the process of playing the narration content by the vehicle. The fourth processing module is configured to control the vehicle to play the narration content, and control the vehicle based on the instructions in the abstract sensory instruction set when playing the key event point in the narration content.

10. An electronic device, comprising: The vehicle comprises: a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus; The memory is configured to store a computer program; and the processor is configured to execute the computer program to implement the control method of the vehicle according to any one of claims 1-8.

11. A storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the control method of the vehicle according to any one of claims 1-8.

Citation Information

Patent Citations

  • Feeling estimation device, feeling estimation method, and storage medium

    CN111341349A

  • Controller system and control method

    CN115129023A

  • Vehicle safety control method and device, medium and vehicle

    CN119459579A

  • Vehicle control method, device, equipment and medium

    CN119811396A

  • Multimedia content interaction method, device and equipment based on vehicle cabin and storage medium

    CN120191303A