Task execution method and device based on vehicle-mounted monitoring image and electronic equipment

By deploying perception modules and multimodal large model technology in the on-board system, intelligent query and processing of on-board driving records is achieved, and the problem of time-consuming and labor-consuming search of video clips in the existing technology is solved, and the user experience is improved.

CN120406796APending Publication Date: 2025-08-01XG TECHNOLOGIES PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510478329.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing on-board dash recorder cannot achieve intelligent query, and users need to manually trace back a large amount of video data to find specific content, which will consume time and effort and affect the driving experience.

Method used

Through the perception module, perceive the vehicle's surrounding environment, generate and store driving records, and use multimodal large-modal model technology and retrieval enhancement generation technology to automatically find and perform related tasks, such as creating or editing videos based on instructions entered by users.

Benefits of technology

It improves the accuracy and efficiency of driving record query, automatically completes user needs, improves driving experience, and reduces manual operation time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120406796A_ABST
    Figure CN120406796A_ABST
Patent Text Reader

Abstract

The invention discloses a task execution method and device based on a vehicle-mounted monitoring image and electronic equipment, and relates to the technical field of intelligent driving. The method comprises the following steps: determining perception information and image data of a vehicle for a surrounding environment; generating at least one driving record according to the perception information of the vehicle for the surrounding environment and the image data, and storing the at least one driving record in a database; and based on an execution instruction input by a user, determining at least one driving record corresponding to the execution instruction and stored in the database, and executing a to-be-executed task corresponding to the execution instruction based on the at least one driving record corresponding to the execution instruction. According to the technical scheme, intelligent searching of the driving image data can be achieved, the searching efficiency is improved, the driving experience of the user is improved, creation work or monitoring work expected by the user can be automatically completed, and a more convenient content generation function and an event reminding function are provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of intelligent driving, and particularly to a task execution method, apparatus, and electronic device based on in-vehicle monitoring images. Background Art

[0002] With the popularization of vehicles, people's travel modes have changed significantly, and driving for tourism has gradually become a popular travel mode. During vehicle driving, devices such as dash cams usually capture the surrounding of the vehicle to obtain corresponding image data for recording driving conditions. After the trip, users can retrieve the image data to view beautiful sceneries, interesting events, etc. encountered during the journey.

[0003] Currently, users need to manually retrieve the image data to find the required image segments and process them. This method has deficiencies in terms of convenience and intelligence, resulting in a poor user experience. Summary of the Invention

[0004] To solve the above technical problems, the present disclosure provides a task execution method, apparatus, and electronic device based on in-vehicle monitoring images, which can achieve intelligent search of driving image data, improve search efficiency, and enhance the user's driving experience.

[0005] In a first aspect of the present disclosure, a task execution method based on in-vehicle monitoring images is provided, including: determining the perception information and image data of the vehicle for the surrounding environment; generating at least one driving record based on the perception information and image data of the vehicle for the surrounding environment, and storing the at least one driving record in a database; based on an execution instruction input by a user, determining at least one driving record stored in the database corresponding to the execution instruction, and executing a to-be-executed task corresponding to the execution instruction based on the at least one driving record corresponding to the execution instruction.

[0006] In a second aspect of the present disclosure, a task execution apparatus based on in-vehicle monitoring images is provided. The apparatus includes: a perception module for determining the perception information and image data for the surrounding environment; a cognitive memory module for generating at least one driving record based on the perception information and image data of the vehicle for the surrounding environment, and storing the at least one driving record in a database; a task execution module for, based on an execution instruction input by a user, determining at least one driving record stored in the database corresponding to the execution instruction, and executing a to-be-executed task corresponding to the execution instruction based on the at least one driving record corresponding to the execution instruction.

[0007] In a third aspect of the present disclosure, a computer-readable storage medium is provided. The storage medium stores a computer program, and the computer program is used to execute the task execution method based on in-vehicle monitoring images provided in the first aspect of the present disclosure.

[0008] The fourth aspect of the present disclosure provides an electronic device, which includes: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the task execution method based on vehicle-mounted monitoring images provided in the first aspect of the present disclosure.

[0009] The fifth aspect of the present disclosure provides a computer program product, which, when executed by an instruction processor in the computer program product, executes the task execution method based on vehicle-mounted monitoring images provided in the first aspect of the present disclosure.

[0010] Based on the task execution method based on vehicle-mounted monitoring images provided by the present disclosure, it is possible to perceive the surrounding environment of the vehicle while the user is driving and form perception information. At least one driving record is generated based on the perception information and the image data of the vehicle and stored in the database, so that each driving record can have a corresponding information record for later query. When the user wants to query related records, the vehicle-mounted system can automatically find the driving record corresponding to the execution instruction based on the execution instruction input by the user, so as to improve the accuracy of the query result matching the user's needs. At the same time, it is also possible to perform the task to be executed corresponding to the execution instruction based on the driving record corresponding to the execution instruction, so as to automatically complete the user's expected creation work or monitoring work, and provide the user with a more convenient content generation function and event reminder function. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 An in-vehicle system architecture provided for an exemplary embodiment of the present disclosure;

[0012] Figure 2 is a flowchart of a method for executing a task based on vehicle-mounted monitoring images provided by an exemplary embodiment of the present disclosure;

[0013] Figure 3 is a flowchart of a method for executing a task based on vehicle-mounted monitoring images provided by another exemplary embodiment of the present disclosure;

[0014] Figure 4 is a flowchart of a method for executing a task based on vehicle-mounted monitoring images provided by another exemplary embodiment of the present disclosure;

[0015] Figure 5 is a first image schematic diagram provided by an exemplary embodiment of the present disclosure;

[0016] Figure 6 is a second image schematic diagram provided by an exemplary embodiment of the present disclosure;

[0017] Figure 7 is a flowchart of a method for executing a task based on vehicle-mounted monitoring images provided by another exemplary embodiment of the present disclosure;

[0018] Figure 8 It is a schematic flowchart of a task execution method based on on-vehicle monitoring images provided by another exemplary embodiment of the present disclosure;

[0019] Figure 9 It is a schematic flowchart of a task execution method based on on-vehicle monitoring images provided by another exemplary embodiment of the present disclosure;

[0020] Figure 10 It is a schematic flowchart of a task execution method based on on-vehicle monitoring images provided by another exemplary embodiment of the present disclosure;

[0021] Figure 11 It is a schematic flowchart of a task execution method based on on-vehicle monitoring images provided by another exemplary embodiment of the present disclosure;

[0022] Figure 12 It is a schematic flowchart of a task execution method based on on-vehicle monitoring images provided by another exemplary embodiment of the present disclosure;

[0023] Figure 13 It is a schematic flowchart of a task execution method based on on-vehicle monitoring images provided by another exemplary embodiment of the present disclosure;

[0024] Figure 14 It is a schematic flowchart of a task execution method based on on-vehicle monitoring images provided by another exemplary embodiment of the present disclosure;

[0025] Figure 15 It is a schematic structural diagram of a task execution device based on on-vehicle monitoring images provided by an exemplary embodiment of the present disclosure;

[0026] Figure 16 It is an interaction schematic diagram between a task execution device based on on-vehicle monitoring images and a data acquisition device provided by an exemplary embodiment of the present disclosure;

[0027] Figure 17 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed Embodiments

[0028] To explain the present disclosure, exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. It should be understood that the present disclosure is not limited by the exemplary embodiments.

[0029] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.

[0030] Application Overview

[0031] With the popularization of vehicles and the improvement of people's living standards, driving for travel has gradually become a popular way to relax and relieve stress. More and more people choose to drive to scenic spots, visit relatives and friends, or go camping in the suburbs to enjoy a free travel experience. During the driving process, when users encounter beautiful scenery or interesting events, they often hope to record these precious moments immediately, so as to record relevant events after the driving ends, or edit and create videos and share them on social platforms.

[0032] Due to driving safety requirements, if users need to record the scenery or events they encounter in a timely manner during the driving of the vehicle, they may need to stop at a safe location and then take pictures of the scenery or events. However, the time consumed in this process may affect the user's itinerary arrangement, and they may also miss the content they want to record due to inability to stop.

[0033] Currently, in-vehicle driving recorders are often configured in vehicles. The driving recorder can record the video of the vehicle during driving in real time and save it for the user to review later. In other words, the driving recorder can record the events and scenery experienced during driving. However, during the driving process, to ensure driving safety, users need to search for the content they want in the video recorded by the driving recorder after the driving ends.

[0034] Since the existing driving recorders cannot implement the intelligent query function, after the driving ends, if users need to search for video clips with specific content from the video data stored in the driving recorder, they often need to compare the video content one by one to find the desired video clips. This will cause users to spend a lot of time and energy searching back, especially during long-distance travel, where there is a large amount of video data collected by the driving recorder, and it is very time-consuming to search for video clips with specific content, resulting in users being unable to find the content they want in a timely and accurate manner, affecting the user's driving experience. After users find the video clips with specific content, they often need to edit and create the video clips to obtain satisfactory video clips. When users edit and create the video clips, they will also need to consume a lot of time and energy.

[0035] To solve the above technical problems, embodiments of the present disclosure provide a task execution method based on in-vehicle monitoring images, which can perceive the surrounding environment of the vehicle and form perception information. At least one driving record is generated based on the perception information of the vehicle for the surrounding environment and the image data of the vehicle, and stored in the database, so that each driving record can have a corresponding information record for easy later query and use. When the user wants to query relevant records, the in-vehicle system can automatically find the driving record corresponding to the execution instruction based on the execution instruction input by the user, and perform the to-be-executed task corresponding to the execution instruction based on the driving record corresponding to the execution instruction, so as to automatically complete the query or creation work desired by the user, thereby being able to provide accurate query results and intelligent creation means for the user and improving the user's driving experience.

[0036] Exemplary system

[0037] Figure 1 The in-vehicle system architecture provided by an exemplary embodiment of the present disclosure. As Figure 1 shown, the in-vehicle system in the present disclosure may include: a data acquisition device 110, a processor 120, and a memory 130. The data acquisition device 110 can acquire the environmental data of the vehicle, and the processor 120 can process the data acquired by the data acquisition device 110 to obtain a driving record. The memory 130 can store the data acquired by the data acquisition device 110 and the data processed by the processor 120 for easy query by the user.

[0038] Exemplarily, the data acquisition device 110 may include various sensors installed on the vehicle. The in-vehicle sensors can provide corresponding monitoring and feedback capabilities to the vehicle to improve driving safety and comfort. The in-vehicle sensors may include:

[0039] In-vehicle cameras (such as dash cams): used to acquire image data during driving or parking. The in-vehicle cameras can perceive the visible environment of the vehicle, including the situation of the occupants in the vehicle, traffic signs outside the vehicle, pedestrians, and vehicles, etc., and can record during driving. Among them, the image data may include image data and video data. The image data includes static two-dimensional images, usually single-frame images, which do not involve the time dimension, but the image data can have a corresponding shooting time. The video data includes a dynamic image composed of a series of consecutive images (frames). The video data usually contains time information and displays the image sequence at a certain frame rate to present a motion effect.

[0040] In-vehicle microphone: Used to collect sound data inside and outside the vehicle. The in-vehicle microphone can perceive the sound environment where the vehicle is located, including the voices of the vehicle occupants, so as to facilitate the recognition of the voice commands of the vehicle occupants to complete relevant interaction operations. The in-vehicle microphone can also perceive road noise outside the vehicle, etc., so as to realize auxiliary driving functions such as active noise reduction.

[0041] Acceleration sensor: Used to collect the speed and acceleration data of the vehicle. The acceleration sensor can detect the speed change of the vehicle, so as to distinguish whether the vehicle is in a driving state or a parked state. During driving, the acceleration sensor can also provide safety and driving assistance functions for the driver based on the magnitude of the vehicle's acceleration and deceleration.

[0042] Radar sensor: Used to collect the distance data between the vehicle and other objects. The radar sensor can detect obstacles around the vehicle when the driver is reversing or following a vehicle, and provide safety warnings. The radar sensor can also adjust its own speed according to the speed of the vehicle in front to maintain a safe distance to assist the vehicle's adaptive cruise function.

[0043] Ultrasonic sensor: Used to collect the distance data between the vehicle and other objects. The ultrasonic sensor can perceive the obstacle information around the vehicle at low speeds. When the driver is reversing, it can provide real-time distance feedback to avoid collisions.

[0044] Facial recognition sensor: Used to collect the facial data of the driver and occupants inside the vehicle. The facial recognition sensor can recognize the facial features (such as facial expressions) of the vehicle driver or occupants, so as to determine the emotional state and identity verification of the driver or occupants according to the facial features.

[0045] Infrared sensor: Used to collect environmental data in a night environment. The infrared sensor can detect the environment around the vehicle under low light conditions and identify pedestrians and animals, etc. It can also be used to detect the body temperature and heat sources of the vehicle occupants.

[0046] Seat pressure sensor: Used to collect seat pressure data. The seat pressure sensor can detect the pressure situation on the seat, so as to assist in judging the number of vehicle occupants and remind the occupants to wear seat belts, etc.

[0047] Temperature sensor: Used to collect the temperature data inside and outside the vehicle. The temperature sensor can monitor the interior temperature and engine temperature to adjust the air conditioning system and engine working state. It can also be used to monitor the external environmental temperature to provide temperature tips for the occupants.

[0048] Humidity Sensor: Used to collect humidity data inside and outside the vehicle. The humidity sensor can measure the humidity levels inside and outside the vehicle to assist in adjusting the air conditioning system, avoid an overly humid environment inside the vehicle, thereby improving passenger comfort, and reducing the formation of frost and fog.

[0049] It should be noted here that the vehicle-mounted sensors may also include an air quality sensor, an eye tracking sensor, etc. The above settings of the sensors are only examples, and the specific types of vehicle-mounted sensors are not limited in the embodiments of the present disclosure.

[0050] Exemplarily, the data acquisition device 110 may include, in addition to various vehicle-mounted sensors, a Global Navigation Satellite System (GNSS) for collecting the position data of the vehicle. Among them, GNSS may include the Global Positioning System (GPS), the Beidou Navigation Satellite System (BDS), the Galileo positioning system, etc., and the embodiments of the present disclosure do not limit this.

[0051] Exemplarily, the processor 120 can calculate and process the data collected by the data acquisition device 110. In the embodiments of the present disclosure, the processor 120 may include a general-purpose processor 120, such as a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), etc., or may also be an Application-Specific Integrated Circuit (ASIC) designed for specific tasks, such as deep learning tasks, autonomous driving tasks, etc.

[0052] In some embodiments, multiple processors 120 may be provided inside the vehicle. The multiple processors 120 may be independent devices from each other or may be integrated devices. For example, the multiple processors 120 may be integrated on one or more System on Chip (SoC). For example, the SoC may include one or more Central Processing Units (CPUs) and / or one or more Graphics Processing Units (GPUs), etc.

[0053] Exemplarily, the memory 130 may also be used to cache or store the intermediate data or result data generated during the operation of the processor 120, as well as to store system files, application files, data files (such as the image data, sound data, speed data, position data, and temperature data collected by the data acquisition device 110), etc.

[0054] Exemplary method

[0055] Figure 2 is a schematic flowchart of a task execution method based on in-vehicle monitoring images provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to vehicles and electronic devices, such as Figure 2 As shown, the task execution method based on in-vehicle monitoring images includes the following steps:

[0056] Step 210, determine the perception information and image data of the vehicle regarding the surrounding environment.

[0057] Among them, the perception information of the vehicle regarding the surrounding environment may include weather information, location information, road condition information, vehicle start information, and in-vehicle personnel status information of the vehicle's surrounding environment, etc.

[0058] Exemplarily, the perception information of the vehicle regarding the surrounding environment can be obtained by a perception module deployed on the vehicle through preliminary analysis and understanding of various data collected by a data acquisition device, and stored in a memory.

[0059] In some embodiments, the perception module refers to a basic system responsible for collecting and processing various environmental data of the vehicle. The perception module can integrate data from multiple sensors in the vehicle, process and analyze the collected data, including image processing, signal processing, and data fusion, etc., to extract useful information. The perception module can form perception information about the vehicle's surrounding environment based on the processed data.

[0060] Exemplarily, the perception module can generate more accurate and comprehensive perception information through multi-modal data fusion technology. Multi-modal fusion technology refers to combining data from different sources or different types to generate more comprehensive and accurate information. Multi-modal data can include various types of data, such as image data, audio data, text data, etc. Multi-modal fusion technology can fuse these data from different sensors to obtain richer perception information. For example, in a complex environment, an in-vehicle camera may be limited by lighting conditions, while lidar can provide more stable distance measurements. Through multi-modal fusion, it is possible to better identify and track surrounding objects. In this way, the use of multi-modal fusion technology can improve the accuracy and robustness of environmental perception.

[0061] In some embodiments, when storing the perception information in the memory, relevant time information can also be attached, such as the time of collecting relevant data or the time of perceiving relevant data, so as to facilitate determining the time when relevant events occur according to the time information in the perception information later.

[0062] In some embodiments, the image data of the vehicle with respect to the surrounding environment may be the image data captured by an in-vehicle camera. When the in-vehicle camera is installed inside the vehicle, the image data may be the image data inside the vehicle, such as the image data captured of the occupant state. When the in-vehicle camera is installed outside the vehicle, the image data may be the image data outside the vehicle, such as the image data captured of the road conditions. At the same time, the image data outside the vehicle may include the image data collected by an in-vehicle driving recorder during the driving process and the parking process of the vehicle.

[0063] It should be noted that the in-vehicle camera may include at least one of the cameras installed inside the vehicle and the cameras installed outside the vehicle, which is not limited in this embodiment.

[0064] Step 220, generate at least one driving record according to the perception information and image data of the vehicle with respect to the surrounding environment, and store the at least one driving record in a database.

[0065] Exemplarily, each driving record may include corresponding text description information and image data information. In other words, each driving record will not only include a video record, but also include some text description information explaining the content of the video record, so as to facilitate subsequent searching and use by users.

[0066] In some embodiments, the text description information corresponding to each driving record may be analyzed by a cognitive memory module deployed in the vehicle based on the image data and perception information of the vehicle.

[0067] Exemplarily, the cognitive memory module refers to a key module in an intelligent system for managing, storing, and retrieving information. In an intelligent driving system or other intelligent systems, the cognitive memory module can record and regularly update the perception information perceived by the perception module stored in the database to maintain the timeliness and accuracy of the information, so as to quickly respond to and adapt to environmental changes. The cognitive memory module can also understand and infer the context information of the current environment to obtain more comprehensive environmental information.

[0068] In some embodiments, the cognitive memory module can comprehensively capture and record the visual content in the environment through multi-modal large model perception technology, and regularly generate a driving record including the perception results, timestamps, and relevant video clips and store it in the database. Among them, the multi-modal large model perception technology refers to the technology that uses large-scale models (such as deep learning models) to comprehensively analyze and understand by combining multiple data modalities. The multi-modal large model can not only identify and classify information, but also extract context knowledge, fuse various types of information, so as to form a more comprehensive understanding and decision-making ability. In this way, the cognitive memory module can understand the environment more comprehensively through the multi-modal large model. For example, by combining image data and audio data, the cognitive memory module can better identify traffic signals, pedestrians, and potential dangers.

[0069] Exemplarily, when the cognitive memory module performs perception based on the multi-modal large model perception technology, the perception results can be presented in the form of text descriptions, so as to generate text description information for relevant video clips, which is convenient for subsequent searching of relevant videos according to the text description information. The cognitive memory module can also manage the perception information generated by the perception module and the driving record system, so as to generate corresponding text description information and driving records based on the perception information and video records, and store the driving records in the database to provide support for subsequent decision-making.

[0070] Step 230, based on the execution instruction input by the user, determine at least one driving record stored in the database corresponding to the execution instruction, and execute the task to be executed corresponding to the execution instruction based on the at least one driving record corresponding to the execution instruction.

[0071] Among them, the execution instruction input by the user can be input when the vehicle starts, or can be input during the driving process of the vehicle, or can also be input after the vehicle driving ends. The present disclosure embodiment does not limit the input time of the execution instruction.

[0072] In some embodiments, the execution instruction input by the user can be recognized and extracted from the audio data collected by the microphone. Or, the execution instruction input by the user can be obtained by recognizing the text input by the user. Or, the execution instruction input by the user can be obtained by the user clicking on the relevant function control on the touch screen of the user interface in the vehicle. Compared with the system where the user can only manually input instructions, multiple input methods for execution instructions such as voice recognition, text recognition, and click triggering can provide a more flexible input experience for the user. At the same time, when the user is not convenient to manually input, voice recognition can also be used, which can improve the input efficiency. In addition, when the user wants to input an execution instruction during the driving process of the vehicle, voice recognition can also improve the driving safety of the user input, so as to avoid affecting the driving safety, thereby improving the user's driving experience.

[0073] Exemplarily, the execution instruction input by the user may include relevant information of the driving record to be searched and the task to be executed. Among them, the task to be executed refers to the task that needs to process the found driving record, such as creation, editing, and transmission, etc.

[0074] In some embodiments, the recognition of the execution instruction input by the user may be implemented based on a task execution module deployed in the vehicle. The task execution module refers to an important module in the intelligent system for processing user input and generating responses. The task execution module can make decisions based on the user input, the current perception information, and the relevant data stored in the database to respond to the user's needs.

[0075] Exemplarily, the task execution module can receive various forms of input from the user, such as voice, text, and touch screen instructions, etc., so that the user can choose a comfortable way to express their needs. Combining the user input, the current perception information, and the driving records stored in the database, an understanding information of the user input is formed. The task execution module can also make corresponding decisions based on the understanding information, such as answering the user's questions, executing specific tasks, etc. For example, when the task execution module receives that the user needs to search for relevant driving records, it can execute the task of searching for driving records. At the same time, it can also process the found driving records based on the user's instructions.

[0076] In some embodiments, the task execution module can use retrieval-augmented generation technology to query the driving records in the database. Retrieval-augmented generation technology is a method that combines information retrieval and text generation, which can improve the response quality and accuracy of the system to user queries.

[0077] Exemplarily, when retrieving the driving records, the retrieval-augmented generation technology can retrieve the corresponding text description information of each driving record based on the relevant information of the driving record to be searched in the execution instruction input by the user. When retrieving the corresponding text description information of each driving record, the matching degree can be determined based on keyword matching, semantic analysis, etc., so as to be able to query the driving records that the user wants.

[0078] It should be noted here that the driving records stored in the database can be one segment or multiple segments. And the driving records found by the task execution module based on the execution instruction input by the user can be one segment or multiple segments, which is not limited in the embodiments of the present disclosure.

[0079] In some embodiments, after the task execution module queries the driving records that the user wants, it can determine the task to be executed included in the execution instruction based on a natural language model, and execute the corresponding tasks on the found driving records, such as tasks of creation, editing, and transmission, etc.

[0080] In the task execution method based on in-vehicle monitoring images provided by the embodiments of the present disclosure, the environment around the vehicle can be perceived to form perception information. At least one driving record is generated based on the perception information of the vehicle regarding the surrounding environment and the image data of the vehicle, and stored in the database. Each driving record may include corresponding text description information generated based on the perception information and the image data, so as to facilitate querying and using the driving record based on the text description information in the later stage. When the user wants to query relevant records, the in-vehicle system can automatically find the driving record corresponding to the execution instruction based on the execution instruction input by the user, and perform the to-be-executed task corresponding to the execution instruction based on the driving record corresponding to the execution instruction, so as to automatically complete the query or creation work desired by the user, thereby being able to provide accurate query results and intelligent creation means for the user and improving the user's driving experience.

[0081] Figure 3 It is a schematic flowchart of the task execution method based on in-vehicle monitoring images provided by another exemplary embodiment of the present disclosure. As Figure 3 shown, based on the embodiment shown above Figure 2 shown, step 210 may include the following steps:

[0082] Step 211, according to the preset environment perception conditions, perceive the environment data collected by the vehicle to obtain the perception result of the environment perception conditions.

[0083] In some embodiments, when the perception module perceives the data collected by the data collection device of the vehicle based on the multi-modal fusion technology, it can perceive functions such as the start of the vehicle, the boarding of the driver or passengers, face recognition, and voice wake-up, thereby generating comprehensive and accurate environment perception results, and saving these results with timestamps to the database.

[0084] Exemplarily, the environment data collected by the vehicle may be relevant data collected by the data collection device on the vehicle. For example, the data that the data collection device of the vehicle can collect may include, but is not limited to, image data, sound data, speed and acceleration data of the vehicle, distance data, facial data of the driver and passengers, night environment data, seat pressure data, position data, temperature data, and humidity data, etc. It should be noted that the embodiments of the present disclosure do not limit the environment data collected by the vehicle.

[0085] In some embodiments, when the perception module performs environment perception, it can perform perception based on the preset environment perception conditions. Among them, the environment perception conditions may include requirements regarding time, weather, location, road conditions, and the state of the people in the vehicle, etc. in the perception of the environment data.

[0086] It should be noted that the environmental perception conditions can be adjusted according to the actual situation. This embodiment only provides an example and is not specifically limited.

[0087] Exemplarily, the time data can obtain the current driving time based on the global satellite positioning system in the vehicle, or can determine the current time according to the clock inside the vehicle, or can also determine the current time and the events corresponding to each piece of image data based on the timestamp information in the driving recorder. In this embodiment, the specific perception method of the time data is not limited.

[0088] Exemplarily, the weather data can perceive the real-time weather information on the driving route by accessing the weather API based on the vehicle's mobile network, or can be perceived by in-vehicle meteorological sensors (such as humidity sensors, temperature sensors, and rain sensors, etc.), or can be perceived by navigation devices and traffic management systems, etc. In this embodiment, the specific perception method of the weather data is not limited.

[0089] Exemplarily, the position data can perceive the current position based on the global satellite positioning system, or can determine the position of the vehicle by calculating the driving distance of the vehicle according to the in-vehicle acceleration sensor. In this embodiment, the specific perception method of the position data is not limited.

[0090] Exemplarily, the road condition data can be perceived by the driving recorder, or can be perceived based on the in-vehicle camera, or can be perceived based on in-vehicle radar sensors, ultrasonic sensors, etc. In this embodiment, the specific perception method of the road condition data is not limited.

[0091] Exemplarily, the in-vehicle personnel status data can be perceived based on the in-vehicle camera, or can be perceived by the face recognition sensor, or can be perceived by the seat pressure sensor. In this embodiment, the specific perception method of the in-vehicle personnel status data is not limited.

[0092] It should be noted that the above examples are only used as a means of perceiving relevant data, and the specific perception methods are not limited in this embodiment.

[0093] In addition, when the perception module perceives the data collected by the above vehicle according to the preset environmental perception conditions, for the same data, it may be necessary to combine different categories of data for perception determination, or it may only be necessary to perform perception determination on one category of data, which is not limited in this embodiment.

[0094] In some embodiments, after the perception module performs perception according to the preset environmental perception conditions, corresponding perception results can be obtained, so as to facilitate subsequent collation of the vehicle's environmental information based on the perception results.

[0095] Exemplarily, in the storage system, the perception result can be saved in the form of text, or in the form of code, or in the form of a table. In this embodiment, the specific storage form of the perception result is not limited.

[0096] Step 212: Based on the perception result, generate a first text, which is used to record the perception information.

[0097] Exemplarily, when the preset environmental perception conditions include time, location, weather, road conditions, and the state of the vehicle occupants, the generated first text can be:

[0098] "Time: Daytime;

[0099] Location: On the bridge;

[0100] Weather: Sunny;

[0101] Road conditions: Good;

[0102] State of vehicle occupants: Good."

[0103] Among them, the order of the relevant content recorded in the above first text can be adjusted, which is not limited in this embodiment.

[0104] It should be noted that the road conditions can include good, congested, etc., and the state of the vehicle occupants can include good, sleepy, etc. states, which are not limited in this embodiment.

[0105] Figure 4 is a schematic flowchart of a task execution method based on in-vehicle monitoring images provided by another exemplary embodiment of the present disclosure. As Figure 4 shown, on the basis of the above Figure 2 shown embodiment, step 220 can include the following steps:

[0106] Step 221: Identify the visible content in the image data of the vehicle.

[0107] Exemplarily, after the perception module obtains the first text recording the perception information based on the preset environmental perception conditions, the cognitive memory module can perform further perception based on the first text and the image data of the vehicle to obtain more comprehensive perception information. Compared with the perception module, the cognitive memory module can perceive the relevant detailed content (including the scene content inside and outside the vehicle) in the image data collected by the vehicle, so as to be able to perceive the environment where the vehicle is located and has experienced more comprehensively.

[0108] In some embodiments, the cognitive memory module can first identify the visible content in the image data and perceive the relevant features displayed in the visible content.

[0109] Exemplarily, when the cognitive memory module recognizes the visual content in the image data, it can preferentially obtain a single-frame image and perceive the single-frame image based on the multi-modal large model technology.

[0110] Figure 5 It is the first image schematic diagram provided by an exemplary embodiment of the present disclosure.

[0111] As Figure 5 shown, exemplarily, when the cognitive memory module perceives Figure 5 based on the multi-modal large model technology, it can recognize the sky, clouds, suspension bridge, stay cables, main tower and bridge deck in the picture. At the same time, it can also recognize the vehicles, street lights on the bridge and the buildings in the distance.

[0112] Figure 6 It is the second image schematic diagram provided by an exemplary embodiment of the present disclosure.

[0113] As Figure 6 shown, exemplarily, when the cognitive memory module perceives Figure 6 based on the multi-modal large model technology, it can recognize the bridge, sculpture and vehicle in the picture.

[0114] It should be noted that the Figure 5 and Figure 6 images provided by the embodiments of the present disclosure can be single-frame images in video data or separate picture data, which is not limited in this embodiment.

[0115] Exemplarily, when the cognitive memory module recognizes the visual content in the image data, it can also directly recognize based on the dynamic video, and is not limited to only recognizing single-frame images.

[0116] Step 222, obtaining a second text based on the visual content and the first text recording the perception information.

[0117] Exemplarily, after the cognitive memory module recognizes the visual content in the image data, it will extract the first text recording the perception information perceived by the perception module again, and organize and fuse the recognized visual content and the first text to obtain the second text.

[0118] Step 223, generating at least one driving record based on the second text and the image data corresponding to the second text.

[0119] Exemplarily, after obtaining the second text, the cognitive memory module will generate at least one driving record by using the second text and the image data corresponding to the second text. In other words, each driving record can include not only the image data collected by the driving recorder, but also the second text used to describe the segment of image data.

[0120] In some embodiments, the video data collected by the driving recorder may contain content for a relatively long period of time. When processing video data with a long duration, a whole segment of video data with a long duration can be sliced into multiple segments of video data with shorter durations, so that each segment of video data with a shorter duration can be recognized and processed. In this way, not only can the processing efficiency of the video data be improved, but also the recognition accuracy of the visible content in the video data can be enhanced.

[0121] Exemplarily, when the cognitive memory module recognizes the visible content in the video data, it can be recognized during the driving process of the vehicle and form multiple consecutive segments of driving records to record the environmental information passed by the vehicle during the driving process. Or, when the cognitive memory module recognizes the visible content in the video data, it can also be recognized during the parking process of the vehicle and form multiple consecutive segments of driving records to monitor the environmental information of the vehicle during the parking process.

[0122] Step 224, store at least one segment of driving record in the database.

[0123] Exemplarily, after obtaining the driving record, the cognitive memory module can store each segment of the driving record in the database to facilitate the management of each segment of the driving record.

[0124] In some embodiments, the retention period of the driving records in the database can be freely set by the user. For example, the user can set the retention period of each segment of the driving record to one week or one month. Based on the retention period, the cognitive memory module can delete and update the driving records in the database to avoid too much content stored in the database, which may affect the storage of subsequent driving records.

[0125] Figure 7 is a schematic flowchart of a task execution method based on on-vehicle monitoring video provided by another exemplary embodiment of the present disclosure. As Figure 7 shown, based on the above Figure 4 shown embodiment, step 222 may include the following steps:

[0126] Step 2221, extract features from the visible content.

[0127] Exemplarily, after recognizing the visible content in the video data, features can be extracted from the visible content. For example, after recognizing a cloud, features such as the color and shape of the cloud can be extracted; after recognizing a vehicle, features such as the type, quantity, and driving position of the vehicle can be extracted; after recognizing a building, features such as the shape and height of the building can be extracted; after recognizing a person, features such as the gestures and expressions of the person can be extracted. In this way, subsequent descriptions of the visible content can be based on the features of the visible content.

[0128] In some embodiments, when extracting features from visual content, deep learning models such as convolutional neural networks can be used for extraction, or other models or means can also be adopted for extraction, which is not limited in this embodiment.

[0129] Step 2222: Generate a third text corresponding to the visual content based on the extracted features.

[0130] Exemplarily, after extracting the features of the visual content, the visual content can be described based on the extracted features.

[0131] Such as Figure 5 shown, a third text corresponding to the visual content in Figure 5 can be generated based on the extracted features:

[0132] "This picture shows a sunny day with a bright blue sky dotted with a few white clouds. In the center of the picture is a modern suspension bridge with beautiful white stay cables connecting the tall main towers and the wide bridge deck. Vehicles are moving smoothly on the bridge, including sedans, SUVs, and an off-road vehicle towing a motorhome. The street lamp poles by the roadside are neatly arranged to ensure safe driving at night. In the distance, the outline of the city can be seen with high-rise buildings standing in rows, forming a busy urban landscape."

[0133] Such as Figure 6 shown, a third text corresponding to the visual content in Figure 6 can be generated based on the extracted features:

[0134] "This picture shows the famous landmark in Nanjing, Jiangsu Province, China - Nanjing Yangtze River Bridge. At the two entrances of the bridge, there are two large sculptures standing respectively, and each sculpture is composed of multiple statues, presenting a magnificent feeling. The bridge spans the Yangtze River, and at both ends, there are majestic sculptures and red towers standing. The bridge deck is wide, and several vehicles are driving on it. On a sunny day, under the blue sky and white clouds, the bridge looks even more spectacular."

[0135] In the above examples of the third text, the third text can not only describe the visual content in the image data, but also describe the visual content based on the features of the visual content, such as "a bright blue sky", "modern suspension bridge", "beautiful white stay cables", "tall main towers and wide bridge deck", "majestic sculptures and red towers", etc., so that the description of the third text for the image data is more accurate, facilitating subsequent search for relevant content.

[0136] Step 2223: Obtain a second text based on the third text and the first text recording perception information.

[0137] Exemplarily, after obtaining the third text and the first text, the first text and the third text can be fused to obtain the second text.

[0138] In some embodiments, after the cognitive memory module recognizes the visual content based on the picture as Figure 5 shown, and combines the first text and the third text, the following second text can be obtained:

[0139] "Time: Daytime;

[0140] Location: On the bridge;

[0141] Weather: Sunny;

[0142] Road condition: Good;

[0143] Status of people in the car: Good;

[0144] This picture shows a sunny day with a bright blue sky dotted with a few white clouds. In the center of the picture is a modern suspension bridge with beautiful white stay cables connecting the tall main towers and the wide bridge deck. Vehicles are moving smoothly on the bridge, including sedans, SUVs, and an off-road vehicle towing a recreational vehicle. The street lamp poles by the roadside are neatly arranged to ensure safe driving at night. In the distance, the outline of the city can be seen with high-rise buildings standing in rows, forming a busy urban landscape."

[0145] In some embodiments, after the cognitive memory module recognizes the visual content based on the picture as Figure 6 shown, and combines the first text and the third text, the following second text can be obtained:

[0146] "Weather: Sunny;

[0147] Time: Daytime;

[0148] Location: Nanjing Yangtze River Bridge;

[0149] Road condition: Good;

[0150] This picture shows the famous landmark in Nanjing, Jiangsu Province, China - Nanjing Yangtze River Bridge. At the two entrances of the bridge, there are two large sculptures standing respectively, and each sculpture is composed of multiple statues, presenting a magnificent feeling. The bridge spans the Yangtze River, and at both ends, there are magnificent sculptures and red towers standing. The bridge deck is wide, and several vehicles are moving on it. The weather is sunny, and under the blue sky and white clouds, the bridge looks even more spectacular."

[0151] Exemplarily, referring to the above two examples of the second text, it can be seen that compared with the first text, the content in the second text is a descriptive text obtained by organizing and summarizing the recognized visual content.

[0152] In some embodiments, when organizing and summarizing the first text and the third text, the cognitive memory module can generate the second text based on a natural language model, so as to be able to describe the environment where the vehicle is located by using the second text.

[0153] It should be noted that the second text can be a direct combination of the first text and the third text, forming a second text with two paragraphs of text content (such as the above example). Or, the second text can also be a fusion of the content of the first text and the third text, forming a second text with a whole paragraph of text content. In this embodiment, the fusion method of the first text and the third text is not limited.

[0154] Figure 8 It is a schematic flowchart of a task execution method based on in-vehicle monitoring images provided by another exemplary embodiment of the present disclosure. As Figure 8 shown, based on the above Figure 2 shown embodiment, step 230 may include the following steps: [[ID=,13]]

[0155] Step 231, retrieve at least one driving record stored in the database based on the execution instruction input by the user to obtain a target driving record.

[0156] Exemplarily, the execution instruction input by the user can be input at any time, and the content input by the user can be language or text. Compared with an intelligent shooting system that requires the user to trigger by pressing a button or trigger at a fixed time, the user can avoid the lag caused by button triggering and fixed-time triggering, as well as the situation of being difficult to capture the user's interest points. In the embodiments of the present disclosure, the user can record and share the interesting journey scenes at any time, so as to avoid missing beautiful scenery and the situation of being easy to forget afterwards.

[0157] In addition, compared with an in-vehicle travel note system based on the user's travel judgment, in the embodiments of the present disclosure, it is no longer necessary to judge whether the user is in a travel state according to set conditions, but can respond to the input execution instruction at any time, so as to improve the user experience fluency and meet the user's personalized needs.

[0158] In some embodiments, the task execution module may include an interaction decision module, and the interaction decision module can process the execution instruction input by the user to determine the driving record that the user wants to find.

[0159] In some examples, the interaction decision module can identify the execution instruction input by the user based on a natural language model, and can retrieve at least one driving record stored in the database to obtain content that matches the execution instruction, so as to obtain a target driving record.

[0160] In one example, for instance, when the execution instruction input by the user is "Produce a documentary containing driving anecdotes and scenery and including the Yangtze River Bridge", the interaction decision module will retrieve driving records including the Yangtze River Bridge in the database. When there are many driving records including the Yangtze River Bridge, retrieval can be performed based on driving anecdotes and scenery to obtain a driving record that can be closer to the execution instruction input by the user as the target driving record.

[0161] In another example, for instance, when the execution instruction input by the user is "Help me pay attention to scenic spots / summer resorts / fishing places suitable for taking children to play", the interaction decision module will retrieve driving records including content such as scenic spots, summer resorts, and fishing places suitable for taking children to play in the database to obtain a driving record that can be close to the execution instruction input by the user as the target driving record.

[0162] Step 232: Determine the task to be executed based on the execution instruction input by the user.

[0163] Exemplarily, the execution instruction input by the user usually includes the task to be executed in addition to the main content of the driving record to be searched for.

[0164] In one example, for instance, when the execution instruction input by the user is "Produce a documentary containing driving anecdotes and scenery and including the Yangtze River Bridge", the interaction decision module will determine that the task to be executed in the execution instruction input by the user includes producing a documentary based on the natural language model.

[0165] In another example, for instance, when the execution instruction input by the user is "Help me pay attention to scenic spots / summer resorts / fishing places suitable for taking children to play", the interaction decision module will determine that the task to be executed in the execution instruction input by the user includes monitoring the specified places based on the natural language model.

[0166] Step 233: Process the target driving record based on the task to be executed.

[0167] Exemplarily, after obtaining the target driving record and determining the task to be executed, the target driving record can be processed based on the task to be executed.

[0168] In one example, if the task to be executed includes producing a video, after obtaining the target driving record, the target driving record can be edited, captioned, filtered, and special effects added, etc., and the video clip desired by the user can be generated and then sent to the user.

[0169] In another example, if the task to be executed includes monitoring relevant places, then after obtaining the target driving record, the target driving record can be transmitted to the user for the user to consult.

[0170] Figure 9It is a schematic flowchart of a task execution method based on in-vehicle monitoring images provided by another exemplary embodiment of the present disclosure. As Figure 9 shown, based on the above Figure 8 shown embodiment, step 231 may include the following steps:

[0171] Step 2311, convert the execution instruction input by the user into a description text.

[0172] Exemplarily, when retrieving driving records based on the execution instruction input by the user, the interaction decision module may understand and organize the execution instruction input by the user through voice based on speech recognition technology to obtain a description text. Alternatively, the interaction decision module may understand and organize the execution instruction input by the user through text based on a natural language model to obtain a description text. In this way, when retrieving the driving records in the database subsequently, the retrieval can be performed based on the content in the description text.

[0173] Step 2312, retrieve at least one driving record stored in the database based on the description text to obtain a target driving record.

[0174] Exemplarily, when retrieving the driving records in the database, the retrieval may be performed based on the description text to obtain a target driving record.

[0175] In some embodiments, when retrieving the driving records in the database, retrieval enhanced generation technology may be used to find the target driving record that best matches the execution instruction through the description text, so as to achieve efficient retrieval and accurately respond to the user's needs.

[0176] Figure 10 It is a schematic flowchart of a task execution method based on in-vehicle monitoring images provided by another exemplary embodiment of the present disclosure. As Figure 10 shown, based on the above Figure 9 shown embodiment, step 2312 may include the following steps:

[0177] Step a, determine the similarity value corresponding to each driving record based on the similarity between the description text and the second text corresponding to at least one driving record stored in the database.

[0178] Exemplarily, when using retrieval enhanced generation technology for retrieval, after obtaining the description text corresponding to the execution instruction input by the user, the description text may be compared with the second text corresponding to each driving record stored in the database based on a natural language model, and the similarity value corresponding to each driving record may be obtained based on the similarity between the description text and the second text.

[0179] In some embodiments, when the natural language model compares the description text with the second text, it can cut a longer second text into multiple text segments. Then, based on the correspondence between the description text and each text segment, the similarity is determined to obtain the similarity value corresponding to each driving record.

[0180] Step b: Determine at least one driving record with a similarity value corresponding to the driving record greater than or equal to the similarity threshold as at least one preselected driving record.

[0181] Exemplarily, after obtaining the similarity value corresponding to each driving record, the driving record with a similarity value greater than or equal to the similarity threshold can be determined as the preselected driving record. Among them, the preselected driving record can be one segment or multiple segments.

[0182] In one example, if the similarity threshold is 70%, then all driving records with a similarity value corresponding to the driving record greater than or equal to 70% can be used as preselected driving records. For example, if the similarity value 1 corresponding to driving record 1 is 60%, the similarity value 2 corresponding to driving record 2 is 70%, the similarity value 3 corresponding to driving record 3 is 68%, the similarity value 4 corresponding to driving record 4 is 80%, and the similarity value 5 corresponding to driving record 5 is 90%, then similarity value 2, similarity value 4, and similarity value 5 are all greater than or equal to the similarity threshold. At this time, driving record 2, driving record 4, and driving record 5 can be used as preselected driving records. In this way, through the similarity threshold, the driving records in the database can be preliminarily screened, thereby narrowing the range of retrieved driving records.

[0183] Step c: Use the preselected driving record with the largest similarity value among at least one preselected driving record as the target driving record.

[0184] Exemplarily, after obtaining the preselected driving record, the similarity values corresponding to each preselected driving record can be compared with each other, and the preselected driving record with the largest corresponding similarity value is used as the target driving record.

[0185] For example, when driving record 2, driving record 4, and driving record 5 in the above step c are used as preselected driving records, the similarity value 5 corresponding to driving record 5 is 90%, which is the largest corresponding similarity value among the preselected driving records. At this time, driving record 5 can be used as the target driving record.

[0186] It should be noted that in other embodiments, the target driving record can also be multiple segments. For example, the similarity values corresponding to multiple preselected driving records are the same, or after obtaining the preselected driving record, another similarity threshold can be used for screening to obtain multiple target driving records. In the embodiments of the present disclosure, the specific number of target driving records is not limited.

[0187] Figure 11 is a schematic flowchart of a task execution method based on in-vehicle monitoring images provided by another exemplary embodiment of the present disclosure. As Figure 11 shown, based on the above Figure 10 shown embodiment, the following steps may further be included before step a:

[0188] Step d, determining the time information in the description text.

[0189] Exemplarily, the time information in the description text is the time corresponding to the driving record desired by the user. For example, when the description text is: "Produce a documentary containing the interesting driving events and scenery today, including the Yangtze River Bridge passed by in the evening, use a soft filter, and edit a video of about one minute." Among them, "today's" and "evening" in the description text are both time information. In this way, the approximate time corresponding to the driving record to be searched can be determined based on the description text.

[0190] Step e, matching the time information with the time stamps corresponding to each driving record.

[0191] In some embodiments, each driving record usually has a corresponding time stamp. This time stamp is used to represent the time when the driving record is formed. When retrieving the driving record, the driving record can be initially screened based on the time information and the time stamp.

[0192] Exemplarily, after obtaining the time information in the description text, the time information can be matched with the time stamps corresponding to each driving record, so that the driving records meeting the time requirements can be obtained.

[0193] In one example, if the description text includes "noon", then when retrieving the driving record, all driving records with time stamps from 11 o'clock to 13 o'clock can be pre-screened and further screened among the driving records meeting the time information. In this way, the retrieval time can be greatly saved and the response speed can be improved.

[0194] Step f, based on the matching of the time information with the time stamps of at least one driving record, determining the similarity between the description text and at least one driving record.

[0195] Exemplarily, when the time stamp corresponding to the driving record matches the time information, the similarity between the driving record and the description text can be determined, and thus the next screening can be performed based on the similarity.

[0196] Figure 12 is a schematic flowchart of a task execution method based on in-vehicle monitoring images provided by another exemplary embodiment of the present disclosure. As Figure 12 shown, based on the above Figure 8Based on the illustrated embodiments, step 233 may include the following steps:

[0197] Step 2331, determine that the task to be executed includes a monitoring task, and obtain monitoring information based on the target driving record.

[0198] Exemplarily, the task to be executed may include a monitoring task to monitor events of interest during the user's driving or parking process, thereby enhancing the user's driving experience.

[0199] In some embodiments, the task execution module further includes an event monitoring module. The event monitoring module can interact with the interaction decision module and perform work related to the monitoring task. When the interaction decision module determines the task to be executed based on the execution instruction, if the task to be executed includes a monitoring task, the event monitoring module can obtain monitoring information based on the target driving record.

[0200] Exemplarily, the monitoring task may include a predefined monitoring task and a custom monitoring task.

[0201] In some embodiments, the predefined monitoring task may include a predefined driving monitoring task and a predefined parking monitoring task. The predefined driving monitoring task may include event monitoring tasks that occur during driving, such as hard braking monitoring task, cutting in monitoring task, and sharp steering wheel turning monitoring task, etc. The predefined parking monitoring task may include event monitoring tasks that occur during parking, such as scratching monitoring task, door opening kill monitoring task, and collision monitoring task, etc.

[0202] Exemplarily, the predefined monitoring task may be some monitoring tasks set in advance in the vehicle system, and the user can determine which monitoring tasks to enable and which to disable.

[0203] It should be noted that the predefined monitoring task may also include some other event monitoring tasks, which will not be elaborated one by one in this embodiment.

[0204] In some embodiments, the custom monitoring task may include user-defined driving events, such as recording good places for taking children, paying attention to fishing places, and recording places suitable for taking children for walks, etc. The custom monitoring task can improve the flexibility of the user's event monitoring, and at the same time can accurately obtain the user's interest points, so as to provide a more accurate and comprehensive monitoring experience for the user.

[0205] It should be noted that the custom monitoring task may also include some other event monitoring tasks, which will not be elaborated one by one in this embodiment.

[0206] Exemplarily, the custom monitoring task may be a monitoring task added by the user himself to pay attention to relevant events during driving.

[0207] In some embodiments, the monitoring task can be that after the user gets in the vehicle, the system automatically asks whether to start the monitoring task, or it can also be input by the user at any time during driving, or it can also be that after the user parks the vehicle, the system automatically asks whether to start the monitoring task. In this embodiment, the input time of the monitoring task is not limited.

[0208] In some embodiments, the monitoring information can include the detailed information of the monitoring event trigger, such as the event type (predefined driving monitoring event, predefined parking monitoring event, and custom driving monitoring event), occurrence time, location, and video clip, etc.

[0209] Exemplarily, the monitoring information can be extracted by the event monitoring module from the second text corresponding to the driving record, or it can also be obtained by re-sensing based on the driving record. In this embodiment, the acquisition method of the monitoring information is not limited.

[0210] Step 2332, send the monitoring information to the user.

[0211] Exemplarily, after obtaining the monitoring information, it can be sent to the electronic device or email bound by the user for the user to view later. In this way, the user does not need to wait in the vehicle for the result after parking, which can avoid affecting the user's itinerary and improve the user's driving experience.

[0212] Figure 13 It is a schematic flowchart of a task execution method based on in-vehicle monitoring images provided by another exemplary embodiment of the present disclosure. As Figure 13 shown, on the basis of the above Figure 12 shown embodiment, before step 231, the following steps can also be included:

[0213] Step 2341, determine the sensitivity threshold of at least one monitoring event corresponding to the monitoring task.

[0214] Exemplarily, when the task to be executed includes a monitoring task, the sensitivity threshold can be determined based on at least one monitoring event corresponding to the monitoring task.

[0215] In some embodiments, the custom monitoring task and predefined monitoring task provided in the above step 2331 can include multiple monitoring events, and each monitoring event can correspond to a different sensitivity threshold. Among them, the lower the sensitivity threshold, the higher the priority of the monitoring event, and the higher the sensitivity threshold, the lower the priority of the monitoring event. Among them, the priority of each monitoring event can represent the user's interest, that is, for the events that the user is more interested in, the priority is higher, and for the events that the user is not interested in, the priority is lower. In this way, when monitoring, the sensitivity threshold corresponding to each monitoring event can be determined first to determine the monitoring events that need to be focused on based on the sensitivity threshold.

[0216] Step 2342: Determine the similarity threshold corresponding to each monitoring event based on the sensitivity threshold.

[0217] Exemplarily, when retrieving driving records, the similarity threshold corresponding to each monitoring event can be determined based on the sensitivity threshold.

[0218] In some embodiments, when retrieving relevant monitoring events, the lower the sensitivity threshold corresponding to the monitoring event, the lower the corresponding retrieval similarity threshold. In this way, during the retrieval process, driving records with a certain similarity to the monitoring event can be pushed to the user, facilitating the user's selection.

[0219] Alternatively, when retrieving relevant monitoring events, the higher the sensitivity threshold corresponding to the monitoring event, the higher the corresponding retrieval similarity threshold. In this way, during the retrieval process, driving records with a relatively high similarity to the monitoring event can be pushed to the user to save the user's selection time.

[0220] Exemplarily, the sensitivity thresholds corresponding to different monitoring events can be adjusted based on the user's needs. During different driving processes of the user, the same monitoring event can correspond to different sensitivity thresholds. In this embodiment, the sensitivity thresholds of each monitoring event are not limited.

[0221] Figure 14 It is a schematic flowchart of a task execution method based on in-vehicle monitoring images provided by another exemplary embodiment of the present disclosure. As Figure 14 shown, based on the above Figure 8 shown embodiment, step 233 may include the following steps:

[0222] Step 2333: Determine that the task to be executed includes a travel note creation task, and create a travel note based on the target driving record.

[0223] Exemplarily, the task to be executed may further include a travel note creation task to provide the user with a travel note video, facilitating the user's sharing on social platforms and enhancing the user's driving experience.

[0224] In some embodiments, the task execution module further includes a personalized creation module. The personalized creation module can interact with the interaction decision module and perform work related to video creation. When the interaction decision module determines the task to be executed based on the execution instruction, if the task to be executed includes a travel note creation task, the personalized creation module can create a travel note based on the target driving record.

[0225] Exemplarily, the travelogue creation task may include video editing, video synthesis, and adding special effects (such as subtitles, background music, and filters, etc.). The personalized creation module can automatically cut and splice the retrieved target driving records based on large model technology and perform creation in combination with the user's description to complete the travelogue creation task.

[0226] In some embodiments, when the execution instruction input by the user is [specific instruction], the corresponding task to be executed is the travelogue creation task. At the same time, the requirements in the travelogue creation task include producing a documentary containing today's driving anecdotes and scenery, and the documentary needs to include the Yangtze River Bridge passed by in the evening and use a soft filter to obtain a video with a time length of about one minute.

[0227] In addition, the personalized creation module also supports Artificial Intelligence Generated Content (AIGC). For example, it can generate cartoon videos or other creative video content according to the user's description, making the generated videos more unique and creative.

[0228] In some embodiments, after obtaining the target driving record, it can be determined whether a travelogue creation task is included. If not, the retrieved target driving record can be directly sent to the user. If a travelogue creation task is included, the target driving record can be created.

[0229] Exemplarily, when creating a travelogue, when the personalized creation module is creating the target driving record, it can also obtain the expressions of the vehicle occupants in the car according to the user's description, such as surprised expressions or excited expressions, and add them to the travelogue video to increase the personalization and interest of the travelogue.

[0230] In some embodiments, the specific creation process of the travelogue can be determined according to the description text generated by the execution instruction input by the user. For example, when the execution instruction is: "Produce a documentary containing today's driving anecdotes and scenery, which should include the Yangtze River Bridge passed by in the evening, use a soft filter, and edit a video of about one minute." After the interactive decision-making module retrieves the target driving record with the Yangtze River Bridge, when creating the travelogue, it can focus on retaining the scenery of the Yangtze River Bridge. At the same time, it can add the surprised expressions of the occupants seeing the Yangtze River Bridge and use a soft filter. In addition, based on "driving anecdotes", some interesting background music and subtitles can also be added to obtain a documentary containing driving anecdotes and scenery with a time length of about one minute to meet the user's needs.

[0231] It should be noted that the specific creation process of the travelogue can be determined according to the user's needs and is not limited in this embodiment.

[0232] Step 2334, send the travelogue to the user.

[0233] Exemplarily, after the travelogue is created, it can be sent to the electronic device or email bound to the user, or played on the central control display screen of the vehicle for the user to view and receive conveniently.

[0234] In some embodiments, after sending the travelogue to the user, it can be determined whether the user is satisfied with the travelogue. If the user is satisfied, the travelogue creation task ends. If the user is not satisfied, the user can be reminded to re-enter the execution instruction and retrieve and generate again until the user is satisfied, thereby improving the user's driving experience.

[0235] In the task execution method based on in-vehicle monitoring images provided by the embodiments of the present disclosure, a multimodal large model technology with highly accurate natural language understanding and dialogue capabilities can be deployed in the in-vehicle system. In this way, the user can describe their needs in any natural language, and the in-vehicle system can accurately understand and respond, thereby improving the fluency of interaction with the user and enabling the user to more conveniently use the in-vehicle system to create travelogues for trips. At the same time, the task execution method based on in-vehicle monitoring images provided by the embodiments of the present disclosure can also sense the data collected by the in-vehicle camera and related sensors, and store the second text obtained based on the sensing result in the database. When the user issues a request, the in-vehicle system can use the retrieval-enhanced generation technology to accurately match the target driving record from the database. Further, in the task execution method based on in-vehicle monitoring images provided by the embodiments of the present disclosure, the user can also customize the monitoring event, and the in-vehicle system can timely remind the user when the monitored event occurs. Compared with the traditional in-vehicle system that relies on preset rules and fixed detection modes, the in-vehicle system of the present disclosure can more flexibly adapt to the personalized needs of users and provide more effective event monitoring services. In addition, the task execution method based on in-vehicle monitoring images provided by the embodiments of the present disclosure can also determine whether video editing is required and complete the corresponding operations. Compared with the system directly providing multiple videos for the user to select, the technical solution of the present disclosure can more accurately respond to the user's needs, improve the user experience, and improve the response efficiency.

[0236] Exemplary device

[0237] Figure 15 is a schematic structural diagram of a task execution device based on in-vehicle monitoring images provided by an exemplary embodiment of the present disclosure.

[0238] Figure 16 is an interaction schematic diagram of a task execution device based on in-vehicle monitoring images provided by an exemplary embodiment of the present disclosure and a data acquisition device.

[0239] Such as Figure 15 and Figure 16As shown, the task execution device 1500 based on in-vehicle monitoring images can interact with the data collection device 110 to obtain the environmental data collected by the data collection device 110.

[0240] Exemplarily, the task execution device 1500 based on in-vehicle monitoring images includes a perception module 1510, a cognitive memory module 1520, and a task execution module 1530.

[0241] Among them, the perception module 1510 is used to determine the perception information and image data of the vehicle for the surrounding environment. The cognitive memory module 1520 is used to generate at least one driving record according to the perception information and image data of the vehicle for the surrounding environment, and store at least one driving record in the database. The task execution module 1530 is used to determine at least one driving record stored in the database corresponding to the execution instruction based on the execution instruction input by the user, and execute the to-be-executed task corresponding to the execution instruction based on at least one driving record corresponding to the execution instruction.

[0242] In some embodiments, the perception module 1510 is further used to: perceive the environmental data collected by the vehicle according to the preset environmental perception conditions to obtain a perception result that meets the environmental perception conditions. Based on the perception result, generate a first text, and the first text is used to record the perception information.

[0243] In some embodiments, the cognitive memory module 1520 is further used to: identify the visible content in the image data of the vehicle; obtain a second text based on the visible content and the first text recording the perception information; generate at least one driving record based on the second text and the image data corresponding to the second text; store at least one driving record in the database.

[0244] In some embodiments, the cognitive memory module 1520 is further used to: extract features from the visible content; generate a third text corresponding to the visible content based on the extracted features; obtain a second text based on the third text and the first text recording the perception information.

[0245] In some embodiments, the task execution module 1530 is further used to: retrieve at least one driving record stored in the database based on the execution instruction input by the user to obtain a target driving record; determine the to-be-executed task based on the execution instruction input by the user; process the target driving record based on the to-be-executed task.

[0246] In some embodiments, the task execution module 1530 includes an interaction decision module 1531, and the interaction decision module 1531 is used to: convert the execution instruction input by the user into a description text. Retrieve at least one driving record stored in the database based on the description text to obtain a target driving record.

[0247] In some embodiments, the interaction decision module 1531 is further configured to: determine a similarity value corresponding to each driving record based on the similarity between the description text and at least one piece of second text corresponding to the driving records stored in the database. Determine at least one preselected driving record from among the driving records for which the similarity value is greater than or equal to a similarity threshold. Use the preselected driving record with the greatest similarity value among the at least one preselected driving records as the target driving record.

[0248] In some embodiments, the interaction decision module 1531 is further configured to: determine the time information in the description text. Match the time information with the time stamps corresponding to each driving record. Based on the matching of the time information with the time stamps of at least one driving record, determine the similarity between the description text and at least one driving record.

[0249] In some embodiments, the interaction decision module 1531 is further configured to: determine that the task to be executed includes a monitoring task. The task execution module 1530 further includes an event monitoring module 1532, and the event monitoring module 1532 is configured to: obtain monitoring information based on the target driving record and send the monitoring information to the user.

[0250] In some embodiments, the event monitoring module 1532 is further configured to: determine a sensitivity threshold for at least one monitoring event corresponding to the monitoring task. Determine a similarity threshold corresponding to each monitoring event based on the sensitivity threshold.

[0251] In some embodiments, the interaction decision module 1531 is further configured to: determine that the task to be executed includes a travelogue creation task. The task execution module 1530 further includes a personalized creation module 1533, and the personalized creation module 1533 is configured to: create a travelogue from the target driving record and send the travelogue to the user.

[0252] For the beneficial technical effects corresponding to the exemplary embodiments of this device, reference may be made to the corresponding beneficial technical effects in the above exemplary method section, and details are not elaborated herein.

[0253] Exemplary electronic device

[0254] Figure 17 is a structural diagram of an electronic device provided for an embodiment of the present disclosure.

[0255] As Figure 17 shown, the electronic device 1700 includes at least one processor 1701 and a memory 1702.

[0256] The processor 1701 may be a central processing unit (CPU) or other form of processing unit having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1700 to perform desired functions.

[0257] The memory 1702 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1701 may run the one or more computer program instructions to implement the task execution method based on in-vehicle monitoring images and / or other desired functions of various embodiments of the present disclosure described above.

[0258] In one example, the electronic device 1700 may further include: an input device 1703 and an output device 1704, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0259] The input device 1703 may further include, for example, a keyboard, a mouse, and so on.

[0260] The output device 1704 may output various information to the outside, and it may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.

[0261] Of course, for simplicity, Figure 17 only some of the components related to the present disclosure in the electronic device 1700 are shown, and components such as a bus, an input / output interface, and so on are omitted. In addition, according to specific application scenarios, the electronic device 1700 may further include any other appropriate components.

[0262] Exemplary Computer Program Product and Computer Readable Storage Medium

[0263] In addition to the above methods and devices, embodiments of the present disclosure may further provide a computer program product, including computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the task execution method based on in-vehicle monitoring images of various embodiments of the present disclosure described in the above "Exemplary Method" section.

[0264] A computer program product may write program code for performing the operations of the embodiments of the present disclosure in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0265] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the method for performing tasks based on in-vehicle monitoring images of various embodiments of the present disclosure described in the above "Exemplary Method" section.

[0266] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium, for example but not limited to, includes systems, devices or components of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0267] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that they are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for the purpose of illustration and facilitating understanding, rather than limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0268] Those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure is also intended to include these changes and modifications.

Claims

1. A task execution method based on vehicle-mounted monitoring images, wherein, The method includes: Determining the perception information and image data of the vehicle with respect to the surrounding environment; Generating at least one driving record based on the perception information and the image data of the vehicle with respect to the surrounding environment, and storing at least one of the driving records in a database; Based on an execution instruction input by a user, determining at least one of the driving records stored in the database corresponding to the execution instruction, and executing a to-be-executed task corresponding to the execution instruction based on at least one of the driving records corresponding to the execution instruction.

2. The method according to claim 1, wherein, The determining the perception information of the vehicle with respect to the surrounding environment includes: Perceiving the environmental data collected by the vehicle according to preset environmental perception conditions to obtain a perception result that meets the environmental perception conditions; Generating a first text based on the perception result, where the first text is used to record the perception information.

3. The method according to claim 2, wherein The generating at least one driving record based on the perception information and the image data of the vehicle with respect to the surrounding environment, and storing at least one of the driving records in a database includes: Identifying visible content in the image data of the vehicle; Obtaining a second text based on the visible content and the first text recording the perception information; Generating at least one of the driving records based on the second text and the image data corresponding to the second text; Storing at least one of the driving records in the database.

4. According to the method of claim 3, the obtaining a second text based on the visible content and the first text recording the perception information includes: Extracting features from the visible content; Generating a third text corresponding to the visible content based on the extracted features; Obtaining the second text based on the third text and the first text recording the perception information.

5. The method according to any one of claims 1-4, wherein, The based on an execution instruction input by a user, determining at least one of the driving records stored in the database corresponding to the execution instruction, and executing a to-be-executed task corresponding to the execution instruction based on at least one of the driving records corresponding to the execution instruction includes: Retrieving at least one of the driving records stored in the database based on the execution instruction input by the user to obtain a target driving record; Determining the to-be-executed task based on the execution instruction input by the user; Processing the target driving record based on the to-be-executed task.

6. The method according to claim 5, wherein The retrieving at least one of the driving records stored in the database based on the execution instruction input by the user to obtain a target driving record includes: Converting the execution instruction input by the user into a description text; Retrieving at least one of the driving records stored in the database based on the description text to obtain the target driving record.

7. The method according to claim 6, wherein, The retrieving at least one of the driving records stored in the database based on the description text to obtain the target driving record includes: Determining a similarity value corresponding to each driving record based on the similarity between the description text and the second text corresponding to at least one of the driving records stored in the database; Determine at least one segment of the driving records corresponding to the driving record with a similarity value greater than or equal to the similarity threshold as at least one preselected driving record; Use the segment of the preselected driving record with the largest similarity value among at least one segment of the preselected driving records as the target driving record.

8. The method according to claim 7, wherein, Before determining the similarity value corresponding to each segment of the driving record based on the similarity between the description text and the second text corresponding to at least one segment of the driving records stored in the database, it further includes: Determine the time information in the description text; Match the time information with the time stamps corresponding to each segment of the driving records; Based on the matching of the time information with the time stamps of at least one segment of the driving records, determine the similarity between the description text and the second text corresponding to at least one segment of the driving records.

9. The method according to claim 5, wherein The processing of the target driving record based on the to-be-executed task includes: Determine that the to-be-executed task includes a monitoring task, and obtain monitoring information based on the target driving record; Send the monitoring information to the user.

10. The method according to claim 9, wherein, Before retrieving at least one segment of the driving records stored in the database based on the execution instruction input by the user to obtain the target driving record, it further includes: Determine the sensitivity threshold of at least one monitoring event corresponding to the monitoring task; Determine the similarity threshold corresponding to each monitoring event based on the sensitivity threshold.

11. The method according to claim 5, wherein, The processing of the target driving record based on the to-be-executed task includes: Determine that the to-be-executed task includes a travelogue creation task, create a travelogue based on the target driving record; Send the travelogue to the user.

12. A task execution device based on in-vehicle monitoring images, wherein, The device includes: A perception module, configured to determine the perception information and image data of the vehicle for the surrounding environment; A cognitive memory module, configured to generate at least one segment of driving records based on the perception information and image data of the vehicle for the surrounding environment, and store at least one segment of the driving records in a database; A task execution module, configured to determine at least one segment of the driving records stored in the database corresponding to the execution instruction based on the execution instruction input by the user, and execute the to-be-executed task corresponding to the execution instruction based on at least one segment of the driving records corresponding to the execution instruction.

13. A computer-readable storage medium, the storage medium stores a computer program, and the computer program is used to execute the task execution method based on in-vehicle monitoring images according to any one of claims 1-11 above.

14. An electronic device, including: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the task execution method based on in-vehicle monitoring images according to any one of claims 1-11 above.