Vehicle, control method and device thereof and medium
By comprehensively processing the video, audio and environmental data of the target object in the car, outputting the status language description and controlling the vehicle operation, it solves the problem of the existing technology that cannot accurately identify the status of the cared-for object in the car, and achieves higher recognition accuracy and driver risk understanding.
Patent Information
- Application Number
- CN202510575670.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-16
AI Technical Summary
Existing vehicle-mounted monitoring systems are unable to fully utilize multimodal information, resulting in the inability to accurately identify the status of objects requiring care in the vehicle in complex environments, and lack an effective risk feedback mechanism.
By acquiring the video data, audio data and environmental data of the target object, a comprehensive analysis is performed using a preset state detection model, and a state language description of the target object is output. The vehicle is then controlled to perform corresponding operations based on the description, including playing sound and light alarms and natural language interpretation.
It improves the accuracy of identifying the status of target objects, reduces the risk of false alarms and missed alarms, reduces driver anxiety, and improves user satisfaction.
Smart Images

Figure CN120645820A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle technology, and in particular to a vehicle control method, a vehicle control device, a computer-readable storage medium, and a vehicle. Background Art
[0002] With the development of intelligent automotive cockpit technology, the safety of in-car care recipients (such as infants) has become a key concern for major automakers and users. In-vehicle monitoring systems typically rely solely on visual input to monitor the in-car environment, failing to fully utilize multimodal information within the vehicle (such as sound and temperature). This results in significant deficiencies in identifying the status of care recipients in complex environments, making it impossible to accurately determine their health and safety, such as risks of suffocation, falling, and overheating. Furthermore, when a care recipient is in danger, these systems typically only provide sound or light alarms, failing to provide detailed risk feedback to the driver through more explanatory information or natural language. Summary of the Invention
[0003] The present application aims to solve at least one of the technical problems in the related art to a certain extent. To this end, the first purpose of the present application is to propose a vehicle control method, which includes: obtaining video data of a target object, audio data of a target object and environmental data of the environment in which the target object is located; using the video data, audio data and environmental data as inputs of a preset state detection model to output a state language description of the target object; and controlling the vehicle according to the state language description of the target object. The control method of the present application controls the vehicle according to the state language description of the target object output by comprehensive processing of the video data, audio data and environmental data of the target object by a multimodal large model. It can analyze the complex state of the target object in more dimensions, improve the accuracy of target object state recognition, and reduce the risk of false alarms and missed alarms. At the same time, the state and risk of the target object are explained to the driver in detail through natural language, which can reduce the driver's anxiety and enable him to make correct judgments quickly, thereby improving user satisfaction.
[0004] A second objective of the present application is to provide a vehicle control device.
[0005] The third object of this application is to provide a computer-readable storage medium.
[0006] A fourth object of the present application is to provide a vehicle.
[0007] To achieve the above-mentioned objectives, the first embodiment of the present application proposes a vehicle control method, which includes: obtaining video data of a target object, audio data of a target object, and environmental data of the environment in which the target object is located; using the video data, audio data, and environmental data as inputs of a preset state detection model to output a state language description of the target object; and controlling the vehicle according to the state language description of the target object.
[0008] According to one embodiment of the present application, controlling a vehicle according to a state language description of a target object includes: controlling the vehicle to play an audible and visual alarm and / or a state language description of the target object.
[0009] According to one embodiment of the present application, controlling a vehicle according to a state language description of a target object includes: identifying keywords in the state language description of the target object; matching corresponding vehicle control instructions according to the keywords; and controlling the vehicle according to the vehicle control instructions.
[0010] According to one embodiment of the present application, before inputting the video data, audio data and environmental data of the target object into the preset state detection model, the method also includes: performing frame extraction processing on the video data of the target object to determine multiple target object images; and using each target object image, the audio data corresponding to each target object image and the environmental data corresponding to each target object image as input to the preset state detection model.
[0011] According to one embodiment of the present application, the above method also includes: obtaining historical video data of the target object, historical audio data of the target object, and historical environmental data of the environment in which the target object is located; performing frame extraction processing on the historical video data of the target object to determine multiple historical target object images; generating a model training set based on each historical target object image, the historical audio data corresponding to each historical target object image, and the historical environmental data corresponding to each historical target object image; training a preset state detection model based on the model training set to construct a preset state detection model for detecting the state of the target object.
[0012] According to one embodiment of the present application, a model training set is generated based on each historical target object image, the historical audio data corresponding to each historical target object image, and the historical environment data corresponding to each historical target object image, including: performing feature extraction on each historical target object image, the historical audio data corresponding to each historical target object image, and the historical environment data corresponding to each historical target object image to obtain corresponding first feature data; and performing feature fusion processing on the first feature data to generate a model training set.
[0013] According to one embodiment of the present application, a model training set is generated based on each historical target object image, the historical audio data corresponding to each historical target object image, and the historical environment data corresponding to each historical target object image, and also includes: determining the text data corresponding to each historical target object image based on each historical target object image and / or the historical audio data corresponding to each historical target object image; performing feature extraction on each historical target object image, the historical audio data corresponding to each historical target object image, the historical environment data corresponding to each historical target object image, and the text data corresponding to each historical target object image to obtain corresponding second feature data; and performing feature fusion processing on the second feature data to generate a model training set.
[0014] To achieve the above-mentioned purpose, the second embodiment of the present application proposes a vehicle control device, which includes: an acquisition module for acquiring video data of a target object, audio data of a target object, and environmental data of the environment in which the target object is located; an output module for using the video data, audio data, and environmental data as inputs of a preset state detection model to output a language description of the target object state; and a control module for controlling the vehicle according to the language description of the target object state.
[0015] To achieve the above-mentioned objectives, the third embodiment of the present application proposes a computer-readable storage medium on which a computer program is stored, characterized in that the program is executed by a processor to implement the aforementioned vehicle control method.
[0016] To achieve the above-mentioned objectives, the fourth embodiment of the present application proposes a vehicle, including a memory, a processor, and a vehicle control program stored in the memory and runnable on the processor. When the processor executes the vehicle control program, the aforementioned vehicle control method is implemented.
[0017] According to the vehicle and its control method, device and medium of the embodiment of the present application, the video data of the target object, the audio data of the target object and the environmental data of the environment in which the target object is located are obtained; the video data, audio data and environmental data are used as inputs of a preset state detection model to output a state language description of the target object; and the vehicle is controlled according to the state language description of the target object. The control method of the present application controls the vehicle according to the state language description of the target object output by comprehensive processing of the video data, audio data and environmental data of the target object by a multimodal large model. It can analyze the complex state of the target object in more dimensions, improve the accuracy of the target object state recognition, and reduce the risk of false alarms and missed alarms. At the same time, the state and risk of the target object are explained to the driver in detail through natural language, which can reduce the driver's anxiety and enable him to make correct judgments quickly, thereby improving user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flowchart of a vehicle control method according to some embodiments of the present application;
[0019] Figure 2 is a block diagram of a control device for a vehicle according to some embodiments of the present application;
[0020] Figure 3 is a block diagram of a vehicle according to some embodiments of the present application. DETAILED DESCRIPTION
[0021] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0022] The vehicle and its control method, device and medium according to the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0023] In the embodiment of the present application, the target object refers to an object in the car that needs care, such as infants, the elderly, and the disabled.
[0024] Figure 1 FIG. 1 is a flow chart of a vehicle control method according to some embodiments of the present application. Figure 1 The vehicle control method of the embodiment of the present application may include the following steps:
[0025] S110 , obtaining video data of a target object, audio data of a target object, and environmental data of an environment where the target object is located.
[0026] Specifically, the target object's video data can be captured by a camera installed inside the vehicle. The camera is installed at an appropriate angle to ensure that the target object's posture and facial features are clearly captured. For example, a camera installed on the left B-pillar can be used to capture the target object in the left rear seat.
[0027] The audio data of the target object can be collected by a microphone installed in the car, such as collecting the target object's crying, speaking voice and environmental noise.
[0028] Environmental data of the target object's environment includes, but is not limited to, in-vehicle temperature, humidity, and air quality. In-vehicle temperature can be detected and obtained using a temperature sensor installed in the vehicle; in-vehicle humidity can be collected and obtained using a humidity sensor installed in the vehicle; and in-vehicle air quality can be detected and obtained using on-board air quality monitoring equipment, such as inhalable particulate matter concentration, carbon dioxide concentration, and air quality index.
[0029] It should be noted that the specific method for obtaining the video data of the target object, the audio data of the target object, and the environmental data of the environment in which the target object is located is not limited here.
[0030] S120 , using the video data, audio data, and environmental data as inputs to a preset state detection model to output a state language description of the target object.
[0031] Specifically, after the video data, audio data and environmental data are input into the preset state detection model, the features of the video data, audio data and environmental data are mapped into a shared semantic space to realize the fusion of multimodal data. For example, through the multimodal encoder and projector, the features of different modalities are converted into a unified vector representation; then, the preset state detection model performs comprehensive analysis and reasoning on the input multimodal data to determine the state of the target object. For example, the posture and expression of the target object are captured according to the video data to judge the state of the target object; the voice, breathing, environmental noise, etc. of the target object are captured according to the audio data to judge the emotion or health status of the target object; the temperature, humidity and air quality in the car are captured according to the environmental data to judge whether the environment in the car is suitable. The preset state detection model performs comprehensive analysis and reasoning on the above-captured content. For example, after analysis and reasoning, it is determined that the target object is crying because its face is partially blocked by the blanket; finally, the preset state detection model outputs a language description of the state of the target object, such as "The target object's face is partially blocked by the blanket, please check to ensure its safety."
[0032] S130: Control the vehicle according to the state language description of the target object.
[0033] Specifically, after the preset state detection model outputs the state language description of the target object, the vehicle can be controlled to play the state language of the target object, and the vehicle can also be controlled to play sound and light alarms to remind the driver to pay attention to the state of the target object; the vehicle can also be directly controlled by matching corresponding vehicle control instructions according to the state language description of the target object. For example, if the state language description of the target object contains "the temperature in the car is too high", the vehicle control instruction of "controlling the air conditioner to lower the temperature" is matched to avoid distracted driving due to excessive attention to the target object to a certain extent; the vehicle can also be directly controlled by matching corresponding vehicle control instructions according to the state language description of the target object while controlling the vehicle to play the state language and / or sound and light alarms of the target object.
[0034] The control method of the present application controls the vehicle based on the state language description of the target object output by comprehensive processing of the target object's video data, audio data and environmental data of the environment in which it is located by a multimodal large model. It can analyze the complex state of the target object in more dimensions, improve the accuracy of target object state recognition, and reduce the risk of false alarms and missed alarms. At the same time, the state and risk of the target object are explained in detail to the driver through natural language, which can reduce the driver's anxiety and enable him to make correct judgments quickly, thereby improving user satisfaction.
[0035] In some embodiments, controlling the vehicle according to the state language description of the target object includes controlling the vehicle to play an audible and visual alarm and / or the state language description of the target object.
[0036] Specifically, after determining the language description of the target object's status, the vehicle can be controlled to play the language description of the target object's status. The risk level of the target object can be determined based on the language description of the target object's status played by the vehicle, and the vehicle can be controlled to play sound and light alarms based on the risk level. The vehicle can also be controlled to play the language description of the target object's status while playing sound and light alarms to remind the driver to pay attention to the status of the target object.
[0037] For example, the risk levels include the first risk level, the second risk level and the third risk level, where the first risk level has the lowest risk and the third risk level has the highest risk. The sound and light alarms include the first sound and light alarm and the second sound and light alarm, where the intensity of the first sound and light alarm is higher than that of the second sound and light alarm, but this is not a limitation to the present application.
[0038] If the risk level of the target object is the first risk level, only the vehicle is controlled to play the language description of the target object's status; if the risk level of the target object is the second risk level, the vehicle is controlled to play the language description of the target object's status while issuing a first sound and light alarm; if the risk level of the target object is the third risk level, the vehicle is controlled to play the language description of the target object's status while issuing a second sound and light alarm.
[0039] In some embodiments, determining the risk level of a target object based on a state language description of the target object played by a vehicle includes: identifying keywords in the state language description of the target object; determining a risk level score of the state language description of the target object based on the keywords; and determining the risk level of the target object based on the risk level score.
[0040] Specifically, the state language description of the target object can be input into a preset word extraction model to output the keywords in the state language description of the target object, and the risk level score of the keyword can be determined by looking up a two-dimensional relationship mapping table between the keyword and the risk level score, wherein the two-dimensional relationship mapping table includes multiple keywords and the risk level score corresponding to each keyword. After determining the risk level scores of all keywords in the state language description of the target object, a sum operation is performed to obtain the risk level score of the state language description of the target object, and the risk level of the state language description of the target object is determined according to the risk level score range in which the risk level score of the state language description of the target object is located.
[0041] For example, if the risk level score is less than a first preset risk level score threshold, the target object's risk level is determined to be a first risk level; if the risk level score is greater than or equal to the first preset risk level score threshold and less than a second preset risk level score threshold, the target object's risk level is determined to be a second risk level; and if the risk level score is greater than or equal to the second preset risk level score threshold, the target object's risk level is determined to be a third risk level. The first preset risk level score threshold and the second preset risk level score threshold can be calibrated according to actual circumstances and are not specifically limited here.
[0042] In this way, the status and risks of the target object are explained to the driver in detail through natural language, and sound and light alarms are issued, which can reduce the driver's anxiety, enable him to make correct judgments quickly, and ensure the effectiveness of the prompt.
[0043] In some embodiments, controlling a vehicle according to a state language description of a target object includes: identifying keywords in the state language description of the target object; matching corresponding vehicle control instructions according to the keywords; and controlling the vehicle according to the vehicle control instructions.
[0044] Specifically, the state language description of the target object can be input into a preset word extraction model to output the keywords in the state language description of the target object, and the vehicle control instructions corresponding to the state language description of the target object can be determined by searching a two-dimensional relationship mapping table between the keywords and the vehicle control instructions, wherein the two-dimensional relationship mapping table includes multiple keywords and the vehicle control instructions corresponding to each keyword. After determining the vehicle control instructions corresponding to all keywords in the state language description of the target object, the vehicle is controlled according to the vehicle control instructions.
[0045] For example, assuming that the state language description of the target object is "the temperature inside the car is low and the humidity is high, and the target object feels uncomfortable", the language description is input into the preset word extraction model, and the keywords in the output state language description of the target object are "low temperature" and "high humidity". The vehicle control instructions matched according to the keywords are "control the air conditioner to rise 2°C" and "control the dehumidifier to increase one level". According to the above vehicle control instructions, the air conditioner is controlled to rise 2°C and the dehumidifier is controlled to increase one level.
[0046] In this way, the driver can be assisted in watching over the target object, and at the same time, to a certain extent, be prevented from being distracted due to excessive attention to the target object, thereby improving driving safety.
[0047] In some embodiments, before inputting the video data, audio data, and environmental data of the target object into a preset state detection model, the method further includes: performing frame extraction processing on the video data of the target object to determine multiple target object images; and using each target object image, the audio data corresponding to each target object image, and the environmental data corresponding to each target object image as inputs to the preset state detection model.
[0048] Specifically, directly inputting the target object's video data into a preset state detection model may result in processing speed that cannot keep up with the speed of video generation, thus affecting the real-time output of the target object's state language description. Therefore, it is necessary to perform frame extraction on the target object's video data to determine multiple target object images. For example, single-frame target object images can be extracted from the video data at fixed time intervals, or a motion detection algorithm can be used to trigger frame extraction when the target object's posture or surrounding environment changes significantly, thereby reducing unnecessary image analysis and improving response efficiency.
[0049] Accordingly, it is necessary to process audio data and environmental data based on the time points of multiple target object image extractions. For example, the audio data within a preset time range before and after the time point of target object image extraction is used as the audio data corresponding to the target object image, and the environmental data within a preset time range before and after the time point of target object image extraction is used as the environmental data corresponding to the target object image.
[0050] In this way, the output rate of the state language description of the target object can be improved, thereby improving the accuracy of the target object state recognition.
[0051] In some embodiments, the above method also includes: obtaining historical video data of the target object, historical audio data of the target object, and historical environmental data of the environment in which the target object is located; performing frame extraction processing on the historical video data of the target object to determine multiple historical target object images; generating a model training set based on each historical target object image, the historical audio data corresponding to each historical target object image, and the historical environmental data corresponding to each historical target object image; training a preset state detection model based on the model training set to construct a preset state detection model for detecting the state of the target object.
[0052] Specifically, the historical video data is subjected to frame extraction processing, for example, at fixed time intervals, or a motion detection algorithm is used to trigger frame extraction when the posture of the target object or the surrounding environment of the target formation changes significantly, so as to determine multiple historical target object images.
[0053] Correspondingly, historical audio data and historical environmental data are processed based on the time points of multiple historical target object image extractions. For example, historical audio data within a preset time range before and after the time point of historical target object image extraction is used as historical audio data corresponding to the historical target object image, and historical environmental data within a preset time range before and after the time point of historical target object image extraction is used as historical environmental data corresponding to the historical target object image.
[0054] After processing the target object's historical video, audio, and environmental data, multiple historical target object image, audio, and environmental data combinations are generated. Each combination is labeled, including the target object's state. The labeled combinations serve as the model training set.
[0055] Input any input data in the model training set into the preset state detection model for forward propagation to obtain the predicted state of the target object. Calculate the discrimination loss based on the predicted state of the target object and the labeled state of the target object corresponding to the input data to obtain the loss value. Calculate the gradient of the loss function for each parameter through backpropagation, and use a preset optimization algorithm (such as Adam) to update the model parameters of the preset state detection model based on the gradient. Repeat the above steps until the updated preset state detection model meets the preset convergence conditions, such as the loss value reaches the preset loss threshold or the number of iterations reaches the preset iteration threshold, to obtain the preset state detection model for detecting the state of the target object.
[0056] It should be noted that there is no specific restriction on the specific training process.
[0057] In some embodiments, a model training set is generated based on each historical target object image, the historical audio data corresponding to each historical target object image, and the historical environment data corresponding to each historical target object image, including: performing feature extraction on each historical target object image, the historical audio data corresponding to each historical target object image, and the historical environment data corresponding to each historical target object image to obtain corresponding first feature data; and performing feature fusion processing on the first feature data to generate a model training set.
[0058] Specifically, when constructing the model training set, in order to reduce the computational dimension and remove redundant information, it is necessary to perform feature extraction on each historical target object image, the historical audio data corresponding to each historical target object image, and the historical environment data corresponding to each historical target object image. For example, a convolutional neural network is used for feature extraction to obtain the corresponding first feature data.
[0059] The corresponding first feature data are subjected to feature fusion processing, for example, the first image feature, the first audio feature, and the first environmental feature are concatenated into a long vector, or the first image feature, the first audio feature, and the first environmental feature are linearly weighted summed, or the first image feature, the first audio feature, and the first environmental feature are multiplied to obtain a first fusion vector, each first fusion vector is labeled, and the labeling content includes the state of the target object, and the multiple labeled first fusion vectors are used as a model training set.
[0060] In some embodiments, generating a model training set based on each historical target object image, the historical audio data corresponding to each historical target object image, and the historical environment data corresponding to each historical target object image also includes: determining the text data corresponding to each historical target object image based on each historical target object image and / or the historical audio data corresponding to each historical target object image; performing feature extraction on each historical target object image, the historical audio data corresponding to each historical target object image, the historical environment data corresponding to each historical target object image, and the text data corresponding to each historical target object image to obtain corresponding second feature data; and performing feature fusion processing on the second feature data to generate a model training set.
[0061] Specifically, the output of the preset state detection model is a state language description of the target object. In order to improve the semantic understanding ability of the preset state detection model, the text data corresponding to each historical target object image can be determined based on each historical target object image, the historical audio data corresponding to each historical target object image, or each historical target object image and the historical audio data corresponding to each historical target object image. For example, each historical target object image and the historical audio data corresponding to each historical target object image are input into the preset natural language description model, and text data associated with each historical target object image and the historical audio data corresponding to each historical target object image are output, and the text data is also added to the model training set.
[0062] Exemplarily, feature extraction is performed on each historical target object image, the historical audio data corresponding to each historical target object image, the historical environment data corresponding to each historical target object image, and the text data corresponding to each historical target object image, for example, using a convolutional neural network to perform feature extraction to obtain corresponding second feature data.
[0063] The corresponding second feature data are subjected to feature fusion processing, for example, the image second feature, audio second feature, text second feature and environment second feature are concatenated into a long vector, or the image second feature, audio second feature, text second feature and environment second feature are linearly weighted summed, or the image second feature, audio second feature, text second feature and environment second feature are multiplied to obtain a second fusion vector, each second fusion vector is labeled, and the annotation content includes the state of the target object, and the multiple labeled second fusion vectors are used as a model training set.
[0064] In this way, during the training phase of a large multimodal model, adding text data corresponding to each target object image when constructing the model training set can enhance the model's semantic understanding ability, thereby improving the accuracy and coherence of the target object's state language description.
[0065] To sum up, the control method of the present application controls the vehicle based on the state language description of the target object output by the comprehensive processing of the video data, audio data and environmental data of the target object by the multimodal large model. It can analyze the complex state of the target object in more dimensions, improve the accuracy of target object state recognition, and reduce the risk of false alarms and missed alarms; at the same time, explaining the state and risk of the target object to the driver in detail through natural language can reduce the driver's anxiety, enable the driver to make correct judgments quickly, and thus improve user satisfaction; in the training stage of the multimodal large model, when constructing the model training set, adding the text data corresponding to each target object image can enhance the semantic understanding ability of the model, thereby improving the accuracy and coherence of the state language description of the target object.
[0066] Corresponding to the above embodiments, the present application also proposes a vehicle control device.
[0067] Reference Figure 2 The vehicle control device 200 includes: an acquisition module 210, an output module 220 and a control module 230.
[0068] The acquisition module 210 is configured to acquire video data of a target object, audio data of the target object, and environmental data of the target object's environment. The output module 220 is configured to use the video data, audio data, and environmental data as inputs to a preset state detection model to output a verbal description of the target object's state. The control module 230 is configured to control the vehicle based on the verbal description of the target object's state.
[0069] According to one embodiment of the present application, the control module 230 is specifically configured to control the vehicle to play an audible and visual alarm and / or a language description of the state of a target object.
[0070] According to one embodiment of the present application, the control module 230 is specifically configured to identify keywords in a state language description of a target object; match corresponding vehicle control instructions according to the keywords; and control the vehicle according to the vehicle control instructions.
[0071] According to one embodiment of the present application, before the video data, audio data and environmental data of the target object are input into a preset state detection model, the video data of the target object is subjected to frame extraction processing to determine multiple target object images; each target object image, the audio data corresponding to each target object image and the environmental data corresponding to each target object image are used as inputs to the preset state detection model.
[0072] According to one embodiment of the present application, historical video data of the target object, historical audio data of the target object, and historical environmental data of the environment in which the target object is located are obtained; the historical video data of the target object are subjected to frame extraction processing to determine multiple historical target object images; a model training set is generated based on each historical target object image, the historical audio data corresponding to each historical target object image, and the historical environmental data corresponding to each historical target object image; and a preset state detection model is trained based on the model training set to construct a preset state detection model for detecting the state of the target object.
[0073] According to one embodiment of the present application, feature extraction is performed on each historical target object image, the historical audio data corresponding to each historical target object image, and the historical environment data corresponding to each historical target object image to obtain corresponding first feature data; and feature fusion processing is performed on the first feature data to generate a model training set.
[0074] According to one embodiment of the present application, text data corresponding to each historical target object image is determined based on each historical target object image and / or historical audio data corresponding to each historical target object image; feature extraction is performed on each historical target object image, the historical audio data corresponding to each historical target object image, the historical environment data corresponding to each historical target object image, and the text data corresponding to each historical target object image to obtain corresponding second feature data; and feature fusion processing is performed on the second feature data to generate a model training set.
[0075] It should be pointed out that the above-mentioned explanation of the embodiments and beneficial effects of the vehicle control method is also applicable to the vehicle control device of the embodiment of the present application. To avoid redundancy, it will not be elaborated here.
[0076] Corresponding to the above embodiment, the present application also proposes a computer-readable storage medium.
[0077] The computer-readable storage medium of the present application stores a computer program thereon, which implements the aforementioned vehicle control method when executed by a processor.
[0078] It should be pointed out that the above-mentioned explanation of the embodiments and beneficial effects of the vehicle control method is also applicable to the computer-readable storage medium of the embodiments of the present application. To avoid redundancy, they are not elaborated here.
[0079] Corresponding to the above embodiments, the present application also proposes a vehicle.
[0080] See also Figure 3 As shown, the vehicle 300 of the present application includes a memory 310, a processor 320, and a vehicle control program stored in the memory 310 and executable on the processor 320. When the processor executes the vehicle control program, the aforementioned vehicle control method is implemented.
[0081] It should be pointed out that the above-mentioned explanation of the embodiments and beneficial effects of the vehicle control method are also applicable to the vehicles of the embodiments of the present application. To avoid redundancy, they will not be elaborated here.
[0082] It should be noted that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0083] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0084] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0085] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0086] In this application, unless otherwise specified or limited, the terms "installed," "connected," "connect," "fixed," etc. should be understood in a broad sense. For example, they can refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two elements or interaction between two elements, unless otherwise specified. Those skilled in the art will understand the specific meanings of the above terms in this application based on specific circumstances.
[0087] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A vehicle control method, characterized in that: in, The method comprises: Acquire video data of a target object, audio data of the target object, and environmental data of an environment in which the target object is located; Using the video data, the audio data, and the environmental data as inputs to a preset state detection model to output a state language description of a target object; The vehicle is controlled according to the state language description of the target object.
2. The vehicle control method according to claim 1, characterized in that: Controlling the vehicle according to the state language description of the target object includes: The vehicle is controlled to play an audible and visual alarm and / or a language description of the state of the target object.
3. The vehicle control method according to claim 1 or 2, characterized in that: Controlling the vehicle according to the state language description of the target object includes: Identifying keywords in a state language description of the target object; Matching corresponding vehicle control instructions according to the keywords; The vehicle is controlled according to the vehicle control instruction.
4. The vehicle control method according to claim 1, wherein: Before inputting the video data, the audio data, and the environmental data of the target object into the preset state detection model, the method further includes: Performing frame extraction processing on the video data of the target object to determine a plurality of target object images; Each target object image, audio data corresponding to each target object image, and environmental data corresponding to each target object image are used as inputs to the preset state detection model.
5. The vehicle control method according to claim 1, characterized in that: The method further comprises: Acquire historical video data of the target object, historical audio data of the target object, and historical environmental data of the environment in which the target object is located; Performing frame extraction processing on historical video data of the target object to determine a plurality of historical target object images; Generate a model training set based on each of the historical target object images, the historical audio data corresponding to each of the historical target object images, and the historical environment data corresponding to each of the historical target object images; The preset state detection model is trained based on the model training set to construct a preset state detection model for detecting the state of the target object.
6. The vehicle control method according to claim 5, characterized in that: Generating a model training set according to each of the historical target object images, the historical audio data corresponding to each of the historical target object images, and the historical environment data corresponding to each of the historical target object images includes: performing feature extraction on each of the historical target object images, the historical audio data corresponding to each of the historical target object images, and the historical environment data corresponding to each of the historical target object images to obtain corresponding first feature data; Perform feature fusion processing on the first feature data to generate the model training set.
7. The vehicle control method according to claim 5, characterized in that: Generating a model training set according to each of the historical target object images, the historical audio data corresponding to each of the historical target object images, and the historical environment data corresponding to each of the historical target object images further includes: determining text data corresponding to each of the historical target object images based on each of the historical target object images and / or historical audio data corresponding to each of the historical target object images; performing feature extraction on each of the historical target object images, the historical audio data corresponding to each of the historical target object images, the historical environment data corresponding to each of the historical target object images, and the text data corresponding to each of the historical target object images, respectively, to obtain corresponding second feature data; Perform feature fusion processing on the second feature data to generate the model training set.
8. A vehicle control device, characterized in that: The device comprises: An acquisition module, configured to acquire video data of a target object, audio data of the target object, and environmental data of an environment in which the target object is located; an output module, configured to use the video data, the audio data, and the environmental data as inputs to a preset state detection model to output a language description of the state of the target object; A control module is used to control the vehicle according to the language description of the target object state.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the vehicle control method according to any one of claims 1 to 7.
10. A vehicle, characterized in that: The vehicle control method comprises a memory, a processor, and a vehicle control program stored in the memory and executable on the processor. When the processor executes the vehicle control program, the vehicle control method according to any one of claims 1 to 7 is implemented.