Dangerous event content description generation method and device, medium and equipment
By acquiring video footage of the vehicle's surroundings in the intelligent vehicle's sentry mode and using a visual language model for secondary verification, a detailed description of the dangerous event is generated. This solves the problem of users not being able to understand the content of the recorded video, improving user experience and vehicle safety.
Patent Information
- Application Number
- CN202511053318.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-11
AI Technical Summary
In the sentry mode of a smart car, the recorded videos of dangerous events lack content prompts, making it impossible for users to view them in a targeted manner, which affects user experience and safety.
By acquiring target videos around the vehicle, a visual language model is used to reconfirm the hazard level of the dangerous event, generating a detailed description of the dangerous event, including specific content and level, to provide users with accurate event information.
It improves the speed and accuracy of users' understanding of dangerous events, ensures vehicle safety, reduces power consumption, and optimizes resource utilization.
Smart Images

Figure CN120932004A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of intelligent vehicle technology, specifically to a method, apparatus, medium, and device for generating descriptions of hazardous events. Background Technology
[0002] Sentry mode is a vehicle security monitoring system in intelligent vehicles. It is mainly used to provide environmental monitoring, anti-theft protection and accident recording functions when the vehicle is parked. In Sentry mode, the surrounding environment is continuously monitored through sensors such as cameras.
[0003] In related technologies, when a vehicle is in sentry mode, a video recording function is activated if a dangerous event is detected. The recorded video is saved locally and a notification is sent to the owner. However, because the recorded video does not provide any content information, users do not specifically review the video, leaving them unaware of the dangerous event and severely impacting the user experience.
[0004] Therefore, there is an urgent need for a sentinel-mode method for generating dangerous event descriptions to improve the user experience. Summary of the Invention
[0005] To address the aforementioned technical issues, this disclosure provides a method, apparatus, medium, and device for generating descriptions of hazardous events, enabling users to quickly understand the specific circumstances of vehicle hazardous events and react promptly, thereby improving user experience and contributing to ensuring vehicle safety.
[0006] In one aspect, a method for generating content descriptions of dangerous events in Sentinel mode is provided, including:
[0007] In response to the detection of a hazardous event of a target hazard level around the vehicle, acquire video of the target around the vehicle;
[0008] Based on the target video, the danger level of the dangerous event is reassessed;
[0009] In response to the redefined hazard level being consistent with the target hazard level, a hazard event content description is generated for the hazard event.
[0010] In another aspect, an apparatus for generating content descriptions of dangerous events in Sentinel mode is provided, comprising:
[0011] The acquisition module is used to acquire target video around the vehicle in response to detecting a dangerous event of a target hazard level around the vehicle;
[0012] The determination module is used to redetermine the danger level of the dangerous event based on the target video;
[0013] The processing module is used to generate a hazard event content description for the hazard event in response to the re-determined hazard level being consistent with the target hazard level.
[0014] In another aspect, the embodiments provide a computer program product that, when an instruction processor in the computer program product is executed, performs the method for generating dangerous event content descriptions for sentinel mode as proposed in the first aspect of the present disclosure.
[0015] In another aspect, an electronic device is proposed, comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the method for generating dangerous event content descriptions for sentinel mode as described in the first aspect above.
[0016] The technical solution provided in this disclosure, in response to the detection of a dangerous event of a target hazard level around a vehicle, acquires target video of the area around the vehicle and redetermines the hazard level of the dangerous event based on the target video. Therefore, in response to the redetermined hazard level matching the target hazard level, a hazard event description is generated. This solution, after initially detecting a dangerous event of a target hazard level (e.g., a high hazard level) around a vehicle, can perform a secondary confirmation of the hazard level based on the acquired target video of the area around the vehicle to determine whether the hazard level matches the target hazard level. If both confirmations indicate a high-risk event, a hazard event description is generated. Therefore, when a serious dangerous event occurs, the user can quickly understand the specific circumstances of the dangerous event based on the generated hazard event description and react promptly, thus contributing to ensuring vehicle safety.
[0017] Moreover, since the hazard level of a dangerous event can be reconfirmed, the authenticity of the hazard level of the initially detected dangerous event can be verified to ensure a more accurate hazard level. Thus, by reconfirming the dangerous event, the detection accuracy of the hazard level of the dangerous event is improved. Attached Figure Description
[0018] Figure 1 This is a system architecture diagram of a sentinel system provided in an exemplary embodiment of this disclosure.
[0019] Figure 2 This disclosure provides an exemplary embodiment of a method for generating content descriptions of dangerous events in sentinel mode.
[0020] Figure 3This is another exemplary embodiment of the present disclosure, which provides a method for generating dangerous event content descriptions for Sentinel mode.
[0021] Figure 4 This is yet another exemplary embodiment of the present disclosure, providing a method for generating content descriptions of dangerous events in Sentinel mode.
[0022] Figure 5 This is a method for generating dangerous event content descriptions for Sentinel mode, provided in yet another exemplary embodiment of this disclosure.
[0023] Figure 6 This is a method for generating dangerous event content descriptions for Sentinel mode, provided in yet another exemplary embodiment of this disclosure.
[0024] Figure 7 This is a generation apparatus for describing dangerous events in sentinel mode, provided in yet another exemplary embodiment of this disclosure.
[0025] Figure 8 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation
[0026] To explain this disclosure, exemplary embodiments of the disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the disclosure, and not all of them. It should be understood that the disclosure is not limited to exemplary embodiments.
[0027] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0028] Application Overview
[0029] In related technologies, in the vehicle's sentry mode, sensors such as cameras and lidar monitor the environment around the vehicle. If abnormal vibrations, others approaching the vehicle, or other potential dangers (such as scratches or theft) are detected, the sentry system will trigger an alarm and control the camera to record video and push it to the vehicle owner.
[0030] However, due to power consumption limitations, the current Sentinel mode employs a low-complexity algorithm that cannot understand the content of dangerous events. Consequently, it does not perform any further processing on the recorded video. Therefore, users only know that a dangerous event has occurred in the vehicle, but they have no idea about the specific details of the event, thus preventing them from making an appropriate response.
[0031] Based on the aforementioned technical problems, the method for generating hazard event content descriptions in sentry mode provided in this disclosure, in response to detecting a hazard event of a target hazard level around a vehicle, can acquire target video around the vehicle and redetermine the hazard level of the hazard event based on the target video. Therefore, in response to the redetermined hazard level being consistent with the target hazard level, a hazard event content description is generated. The solution of this disclosure, after initially detecting a hazard event of a target hazard level (e.g., a high hazard level) around a vehicle, can perform a secondary confirmation of the hazard level based on the acquired target video around the vehicle to determine whether the hazard level of the hazard event is the target hazard level. If the hazard event is confirmed as a high-risk event in both instances, a hazard event content description can be generated. Therefore, when a relatively serious hazard event occurs to a vehicle, the user can quickly understand the specific circumstances of the hazard event based on the generated hazard event content description and react in a timely manner, thus helping to ensure vehicle safety.
[0032] Moreover, since the hazard level of a dangerous event can be reconfirmed, the authenticity of the hazard level of the initially detected dangerous event can be verified to ensure a more accurate hazard level. Thus, by reconfirming the dangerous event, the detection accuracy of the hazard level of the dangerous event is improved.
[0033] Exemplary System
[0034] Figure 1 This is a system architecture diagram of a sentinel system provided in an exemplary embodiment of this disclosure. The sentinel system may include a training phase and an inference phase.
[0035] In some examples, such as Figure 1 As shown, during the inference phase, image-text pairs for Sentinel mode are retrieved from the open-source image-text dataset to obtain image-text data for Sentinel mode at high, medium, and low risk levels. Based on the image-text data of Sentinel mode, the basic model of Visual-Language Model (VLM) is trained to obtain the trained VLM. The trained VLM can then be deployed on vehicles for application to Sentinel mode, i.e., the trained VLM serves as a dedicated visual-language model for Sentinel mode.
[0036] In some embodiments, such as Figure 1 As shown, during the inference phase, dangerous event targets can be detected based on images of the vehicle's surroundings captured by the camera. If a high-risk event is detected, the target video recorded by the camera can be acquired and processed based on the trained VLM to achieve secondary confirmation of the high-risk event.
[0037] One possible scenario: If the redefined hazard level matches the high-hazard level, a hazard event description can be generated. This description can include the specific details of the hazard and its hazard level. Thus, the event details can be obtained through the hazard event description.
[0038] Subsequently, based on the description of the dangerous event, a notification SMS can be generated and sent to the vehicle owner's terminal device to notify the owner to check the SMS for event details; alternatively, a power-on reminder can be generated and sent to the vehicle's in-vehicle terminal so that the owner can intuitively view the reminder and understand the event details when starting the in-vehicle terminal; or, the recorded target videos can be categorized so that the owner can selectively view the videos locally to understand the event details, avoiding the accumulation of a large number of unread videos locally.
[0039] Another possible scenario: If the redefined hazard level is inconsistent with the high-hazard level, meaning the redefined hazard level is a different level than high-hazard (e.g., low-hazard or medium-hazard), the hazard level of the event can be corrected based on the redefined hazard level. For example, the hazard level can be adjusted to a lower hazard level in a timely manner, thus ensuring the accuracy of hazard level detection. Furthermore, the next step can be determined based on the adjusted lower hazard level. For instance, the corrected result can be used to classify videos so that users can selectively view locally stored videos, or the corresponding video resources can be deleted, effectively avoiding waste of storage resources.
[0040] The technical solution provided in this disclosure, after initially detecting a dangerous event of a target danger level (e.g., high danger level) around a vehicle, can perform a secondary confirmation of the danger level of the dangerous event based on the acquired target video around the vehicle to determine whether the danger level of the dangerous event is the target danger level. If the dangerous event is confirmed to be a high danger event in both instances, a description of the dangerous event content can be generated. Therefore, when a relatively serious dangerous event occurs to the vehicle, the user can quickly understand the specific situation of the dangerous event based on the generated description of the dangerous event content and thus react in a timely manner, which helps to ensure vehicle safety.
[0041] Moreover, since the hazard level of a dangerous event can be reconfirmed, the authenticity of the hazard level of the initially detected dangerous event can be verified to ensure a more accurate hazard level. Thus, by reconfirming the dangerous event, the detection accuracy of the hazard level of the dangerous event is improved.
[0042] In addition, since VLM is only triggered when a high-risk event is detected, and other risk levels (such as low risk levels) will not trigger VLM, power consumption can be reduced on the one hand, while ensuring the accuracy of the detection of the risk level of the dangerous event. On the other hand, the powerful understanding of VLM enables the generation of high-precision event descriptions that are easy for users to understand, so that users can respond in a timely and accurate manner.
[0043] Exemplary methods
[0044] Figure 2 This is a flowchart illustrating a method for generating a dangerous event content description in sentinel mode, provided by an exemplary embodiment of this disclosure.
[0045] This embodiment can be applied to electronic devices, such as... Figure 2 As shown, it includes the following steps:
[0046] Step 201: In response to detecting a dangerous event of a target hazard level around the vehicle, acquire target video around the vehicle.
[0047] In some embodiments, the hazard level of a hazardous event can be divided into multiple hazard levels based on the severity of the hazardous event; for example, the multiple hazard levels may include a high hazard level, a medium hazard level, and a low hazard level; or, for another example, the multiple hazard levels may include a high hazard level and a low hazard level. The division of hazard levels can be determined according to the actual usage, and this disclosure does not limit this.
[0048] In some examples, the target hazard level described above may include at least one hazard level.
[0049] For example, the target hazard level is high hazard level, or the target hazard level is medium hazard level; or, for another example, the target hazard level includes both high hazard level and medium hazard level.
[0050] In the vehicle's sentry mode, cameras installed on the outside of the vehicle collect environmental data around the vehicle in real time to detect whether dangerous events are occurring around the vehicle and the level of danger of such events. When a dangerous event is detected, the camera can be controlled to record video, thus obtaining the target video in real time.
[0051] For a description of the detection of a dangerous event of a target hazard level around a vehicle, please refer to the detailed description in the following embodiments, which will not be repeated here.
[0052] Step 202: Based on the target video, redetermine the hazard level of the hazardous event.
[0053] In some examples, the target video can be preprocessed to obtain keyframes, and then feature recognition can be performed on the keyframes to obtain image recognition results. The image recognition results may include information such as the category, distance, and behavior of the target object. The assessment criteria for the danger level are predefined, and then the danger level of the dangerous event is re-determined based on the image recognition results and in combination with the assessment criteria for the danger level. For details, please refer to the detailed description in the following embodiments. This disclosure will not repeat the details in the embodiments.
[0054] Step 203: In response to the redefined hazard level being consistent with the target hazard level, generate a hazard event content description for the hazard event.
[0055] In some examples, the above description of a hazardous event is used to show specific details of the hazardous event; that is, the description may include the specific content of the hazardous event. Of course, the above description may also include the hazard level or other information.
[0056] For example, taking a target hazard level as high hazard level as an example. When a high hazard level dangerous event "a passerby is smashing a vehicle window" is detected around the vehicle, the target video recorded by the camera can be acquired in real time; based on the target video, the hazard level of the dangerous event is re-determined, and if the re-determined hazard level is still high hazard level, a hazard event content description "a passerby is smashing a vehicle window" is generated.
[0057] The method for generating hazard event content descriptions in sentry mode provided in this disclosure, in response to detecting a hazard event of a target hazard level around a vehicle, acquires target video of the area around the vehicle and redetermines the hazard level of the hazard event based on the target video. Therefore, in response to the redetermined hazard level matching the target hazard level, a hazard event content description is generated. This disclosure, after initially detecting a hazard event of a target hazard level (e.g., high hazard level) around a vehicle, can perform a secondary confirmation of the hazard level based on the acquired target video of the area around the vehicle to determine whether the hazard level matches the target hazard level. If both confirmations indicate a high-risk event, a hazard event content description can be generated. Therefore, when a serious hazard event occurs, the user can quickly understand the specific circumstances of the hazard event based on the generated hazard event content description and react promptly, thus contributing to ensuring vehicle safety.
[0058] Moreover, since the hazard level of a dangerous event can be reconfirmed, the authenticity of the hazard level of the initially detected dangerous event can be verified to ensure a more accurate hazard level. Thus, by reconfirming the dangerous event, the detection accuracy of the hazard level of the dangerous event is improved.
[0059] In some embodiments, such as Figure 3 As shown, step 202 above may include the following steps:
[0060] Step 2021: Obtain the preset first prompt text.
[0061] In some examples, the aforementioned first prompt text can be a prompt text, which is used to guide the visual language model to process the target video and output the danger level of the dangerous event; the pre-set first prompt text can be stored locally or on a cloud server, so the first prompt text can be retrieved from local storage or from a cloud server.
[0062] For example, the first prompt text could be "Please determine whether the behavior in the current video will cause damage to the vehicle? Or, please determine whether the vehicle in the current video has been damaged or destroyed? If yes, output a high danger level; if no, output a low danger level."; or, the first prompt text could be "Please determine the distance between the target object in the current video and the vehicle? If the distance is 0, output a high danger level; if the distance is less than 20cm, output a medium danger level; if the distance is greater than 20cm, output a low danger level." The first prompt text could also be other prompt words to guide the visual language model to identify other information about the target objects around the vehicle from the target video (e.g., quantity, dwell time, or duration of contact, etc.), which is not limited in this embodiment of the present disclosure.
[0063] Step 2022: Using the first prompt text, guide the visual language model to process the target video and obtain the redefined danger level.
[0064] Since the first prompt text can determine the operation that the model needs to perform on the video, define the specific judgment rules for the danger level, and specify the output results of the model, the first prompt text can be used to guide the visual language model to perform the corresponding recognition operation on the target video, and then, based on the specific judgment rules for the danger level, determine the danger level of the dangerous event in the target video and output the danger level.
[0065] In some examples, the aforementioned visual language model is a multimodal artificial intelligence model that combines computer vision and natural language processing. By simultaneously learning the semantic relationships between images (visual data) and text (language data), it can understand the visual content (such as objects, scenes, and actions) in images and align it with natural language (such as descriptions, questions, and instructions), ultimately achieving bidirectional "image-text" interaction. The aforementioned visual language model can be a basic Visual-Language Model (VLM) or a pre-trained visual language model; this disclosure does not limit the specific implementation.
[0066] In some embodiments, when the visual language model is a pre-trained visual language model, before step 2022, the method for generating content descriptions of dangerous events in sentinel mode provided in this disclosure may further include: obtaining an initial visual language model; obtaining multiple image-text pairs for sentinel mode from a training dataset; wherein the multiple image-text pairs include image-text pairs with different danger levels; and training the visual language model based on the multiple image-text pairs to obtain a trained visual language model.
[0067] In some examples, the initial visual language model described above is a VLM model based on; if a pre-downloaded initial visual language model is stored locally, it can be obtained directly from the local storage; or, if a pre-downloaded initial visual language model is not stored locally, it can be downloaded directly from an open-source model library.
[0068] In some examples, during the initial training phase of the visual language model, multiple image-text pairs for Sentinel mode can be retrieved from an open-source image-text dataset (i.e., the training dataset). Each image-text pair corresponds to a text description that defines the image's content and hazard level. During retrieval, image-text pairs of different hazard levels (high, medium, and low) can be obtained from the training dataset. These hazard-level pairs can then be used to train the visual language model, resulting in a trained model. This improves the trained model's visual ability to perceive images in Sentinel scenes, thereby enhancing the accuracy of hazard level assessment.
[0069] For example, the criteria used to assess different risk levels may include: a high risk level when there is a destructive action or a destructive result; a medium risk level when a specific target object exists within a preset distance around the vehicle; and a low risk level when a target object exists beyond a preset distance around the vehicle.
[0070] The technical solution provided in this disclosure has two advantages. First, it enhances the visual language model's understanding of sentinel scenes by training it on multiple image-text pairs based on sentinel mode, thereby improving its subsequent understanding of target videos. Second, since the preset first prompt text can predefine the quantitative indicators and grading rules for the hazard level, when guiding the visual language model to process the target video using the first prompt text, the hazard level of the dangerous event can be re-determined according to a unified standard, thus avoiding errors caused by inconsistent human judgment standards in traditional assessments. Therefore, through model fine-tuning and standardized prompts, the accuracy of hazard level assessment can be improved.
[0071] In some embodiments, such as Figure 4 As shown, step 203 above can specifically include the following steps:
[0072] Step 2031: In response to the redefined hazard level being consistent with the target hazard level, obtain the preset second prompt text.
[0073] The second prompt text is used to indicate the generation style of the description of the dangerous event.
[0074] In some examples, the second prompt text can be a prompt text; the pre-set second prompt text can be stored locally or on a cloud server, so the second prompt text can be retrieved from local storage or from a cloud server.
[0075] For example, the second prompt text can define the style and content of the description of the task objective and the hazardous event. For instance, the content of the second prompt text could be as follows:
[0076]
Task Objective
[0077] 1. Based on the content of the video, select the hazard level that is closest to the video from the hazard level list [high risk, medium risk, low risk, unknown].
[0078] 2. If the danger level is high, please select a suitable label from the list of target types and corresponding target behaviors, and describe the event content in no more than 20 words.
[0079] The target object types include: [pedestrians, cyclists, motorcyclists, tricycles, four-wheeled vehicles, and unknown].
[0080] The target's behaviors include: Pedestrians: [touching vehicles with their hands, striking their own vehicles with their bare hands, kicking their own vehicles, striking their own vehicles with objects, maliciously scratching their own vehicles, approaching their own vehicles, approaching and then moving away from their own vehicles, peeking into their own vehicles, taking photos of their own vehicles, loitering around their own vehicles, passing by normally, other]; Two-wheeled vehicles: [colliding with their own vehicles, striking their own vehicles with their bare hands, striking their own vehicles with objects, suspected of scraping their own vehicles, touching their own vehicles with their hands, approaching their own vehicles, approaching and then moving away from their own vehicles, taking photos of their own vehicles, loitering around their own vehicles, passing by normally, other]; Three-wheeled vehicles: [colliding with their own vehicles, suspected of scraping their own vehicles, approaching their own vehicles, approaching and then moving away from their own vehicles, passing by normally, other]; Four-wheeled vehicles: [opening their own vehicles and hitting their own vehicles, suspected of scraping their own vehicles, getting off and touching their own vehicles, passing by normally, other]
[0081] Example Output
[0082] [Danger Level: "High Danger Level", specifically "Pedestrians approach vehicles, peer into the vehicle through windows, and touch the glass, suspected of reconnaissance activity"]
[0083] Step 2032: Using the second prompt text, guide the visual language model to process the target video and generate a description of the dangerous event.
[0084] The description of a hazardous event includes the specific details of the hazardous event and its reclassified hazard level.
[0085] Since the second prompt text can determine the task that the model needs to perform on the video, define the type and behavior of the target object, and specify the output results and style of the model, the visual language model can be guided to perform the corresponding task on the target video based on the obtained target object type and behavior information, and then determine the content included in the description of the dangerous event (e.g., including specific content and the danger level after the danger event is redefined) according to the specified output results and style of the model, and output the description of the dangerous event content.
[0086] In some examples, the aforementioned visual language model can be a basic visual language model (VLM) or a pre-trained visual language model, and this disclosure does not limit this. When the aforementioned visual language model is a pre-trained visual language model, the basic visual language model can be trained with reference to the description in the above embodiments, and this disclosure will not repeat the details.
[0087] The technical solution provided in this disclosure, when the re-determined danger level is consistent with the target danger level, can guide the visual language model to process the target video and generate a description of the danger event content by using a preset second prompt text. This makes the generated description of the danger event content more standardized and comprehensive, thereby making it easier for users to understand the details of the danger event based on the description of the danger event content, and thus enabling users to make more accurate responses through judgment.
[0088] In some embodiments, the method for generating content descriptions of dangerous events in sentinel mode provided in this disclosure may further include the following steps: in response to a discrepancy between the redefined danger level and the target danger level, correcting the danger level of the dangerous event based on the redefined danger level.
[0089] In some examples, the redefined hazard level of a hazardous event differs from the target hazard level, meaning the redefined hazard level is a different level than the target hazard level. For example, if the target hazard level is high, the redefined hazard level may be medium or low.
[0090] If the risk level of a hazard event is redefined as a risk level other than the target risk level, the risk level of the hazard event will be corrected to the redefined risk level, and no hazard event description will be generated.
[0091] For example, taking a target hazard level as high hazard level as an example. When a high hazard level dangerous event is initially detected around the vehicle, target video of the vehicle's surroundings recorded by a camera can be acquired; based on the target video, the hazard level of the dangerous event is re-determined to low hazard level; since the re-determined hazard level is inconsistent with the high hazard level, the hazard level of the dangerous event can be corrected to low hazard level.
[0092] In other embodiments, after correcting the danger level of a dangerous event based on the redefined danger level, the video can be categorized based on the corrected danger level. Alternatively, corresponding video resources can be deleted, or some cameras can be turned off. Thus, by categorizing the video, dangerous events are presented in a categorized manner, making it easier for users to view videos selectively, thereby avoiding the storage of a large number of unread videos locally. Deleting video resources with lower danger levels can save local storage resources, and turning off some cameras can reduce the power consumption of Sentry Mode.
[0093] Based on the above embodiments, since the redefined hazard level is inconsistent with the target hazard level, the hazard level of the dangerous event can be corrected based on the redefined hazard level. Therefore, based on the hazard level correction mechanism, the actual hazard level of the dangerous event can be accurately identified. This can improve the detection accuracy of the hazard level of the dangerous event on the one hand, and effectively avoid resource waste by eliminating the need for vehicles to respond promptly to dangerous events that do not meet the target hazard level on the other hand.
[0094] like Figure 5 As shown above, in the above Figure 2 Based on the illustrated embodiment, prior to step 201 above, the method for generating dangerous event content descriptions for Sentinel mode provided in this disclosure may further include the following steps:
[0095] Step 204: Obtain the first image of the area around the vehicle.
[0096] In some examples, image information about the vehicle's surroundings is acquired in real time using an image sensor to obtain a first image. This image sensor can be one or more cameras located outside the vehicle, and the number of first images can be one or more. For example, the image sensor can be a wide-angle camera with a 180-degree or greater field of view to ensure a comprehensive view of the vehicle's surroundings.
[0097] Step 205: Based on the first image, determine the hazard level of the hazardous event.
[0098] In some embodiments, image processing algorithms can be used to perform image recognition on the first image to extract image features of multiple objects included in the first image. Then, the image features of the multiple objects are compared with the image features of a preset target object. If at least one object matches the image features of the target object, it can be determined that a dangerous event exists around the vehicle. Furthermore, the type, location, and quantity of at least one object are determined, and the danger level of the dangerous event is determined based on this information. For details, please refer to the detailed description in the following embodiments; the embodiments disclosed herein will not be repeated here.
[0099] In some examples, the preset target object may include at least one of the following: pedestrians, vehicles, foreign objects, animals, etc.
[0100] In some examples, if a target object is identified from the first image, an alarm can be triggered on the vehicle.
[0101] In some embodiments, determining the danger level of a dangerous event based on a first image may specifically include: performing target detection on the first image based on a preset neural network model to obtain target detection results; wherein the target detection results include at least one of the following: the distance between the vehicle and the target object, the number of target objects, and the time the target object stays around the vehicle; and determining the danger level of the dangerous event based on the target detection results.
[0102] The distance between the vehicle and the target object is used to indicate the positional relationship between the vehicle and the target object; the number of the target objects can be the total number of all target objects, or the number of target objects of different types; the dwell time of the target object around the vehicle is the duration between the moment the target object appears within the preset range of the vehicle and the moment the target object leaves the preset range of the vehicle.
[0103] In some examples, the aforementioned neural network model is a pre-trained model for image recognition, such as a Convolutional Neural Network (CNN). Based on the neural network model, specific target objects (e.g., pedestrians, animals, or vehicles) can be identified and located from a first image, and the number of target objects, the distance between the target objects and the vehicles, and the appearance and departure of target objects around the vehicles can be output. Furthermore, the dwell time of the target objects can be determined based on their appearance and departure times. In addition, the neural network model can also output information such as the type and size of the target objects. For a detailed description of the neural network model's image processing, please refer to related technologies; this disclosure will not elaborate further.
[0104] In some embodiments, after obtaining the target detection results, the danger level of the dangerous event can be determined based on the target detection results through preset rules or machine learning models.
[0105] In some examples, the above-mentioned determination of the hazard level of a dangerous event based on the target detection results may specifically include: if the target detection result matches the high-risk level, then the high-risk level is determined as the hazard level of the dangerous event; or, if the target detection result matches the medium-risk level, then the medium-risk level is determined as the hazard level of the dangerous event; if the target detection result matches the low-risk level, then the low-risk level is determined as the hazard level of the dangerous event.
[0106] For example, if the target detection result shows that the distance between the target object and the vehicle is 0, that is, the target object is found to be hitting the vehicle, the target detection result matches the high-risk level and is determined as the danger level of the dangerous event; if the target detection result shows that the distance between the target object and the vehicle is less than or equal to a preset distance (e.g., 20cm), the target detection result matches the medium-risk level and is determined as the danger level of the dangerous event; if the target detection result shows that the distance between the target object and the vehicle is greater than the preset distance, the target detection result matches the low-risk level and is determined as the danger level of the dangerous event.
[0107] It should be noted that the closer the target is to the vehicle, the greater the danger to the vehicle, and therefore the higher the danger level of the incident. The more target objects around the vehicle, the greater the potential danger to the vehicle (e.g., three or more suspicious individuals gathered together may intend to damage the vehicle), and therefore the higher the danger level of the incident. The longer the target lingers around the vehicle, the more suspicious their intentions, and therefore the higher the danger level of the incident. Furthermore, the danger level of the incident can also be determined by the type of target object; that is, certain specific objects (such as armed individuals) pose a high-risk hazard to the vehicle.
[0108] The technical solution provided in this disclosure is based on a neural network model to perform target detection on the first image around the vehicle, and outputs multi-dimensional quantitative indicators such as the distance, number, and dwell time of target objects around the vehicle. Then, combined with the multi-dimensional quantitative indicators, the danger level of the dangerous event is assessed, thereby improving the accuracy of the danger level assessment. Based on the accurate danger level, a more appropriate response can be made subsequently.
[0109] In the above Figure 2 On the basis of, such as Figure 6 As shown, after step 203 above, the method for generating dangerous event content descriptions for Sentinel mode provided in this embodiment of the disclosure may further include the following steps:
[0110] Step 206: Based on the description of the hazardous event, perform the target operation.
[0111] The target operation includes at least one of the following: generating and sending target information to the target device, classifying the target video by hazard level; the target information is obtained based on the description of the hazard event content; the target device is the vehicle's in-vehicle terminal or the vehicle owner's mobile device.
[0112] In some embodiments, since the target information is obtained based on the description of the hazard event content, the description of the hazard event content can be directly determined as the target information, that is, the target information is the same as the description of the hazard event content, or the key information in the description of the hazard event content can be determined as the target information, that is, the target information is the key information in the description of the hazard event content.
[0113] In one example, consider a dangerous event described as "a suspicious person circled the car twice and made scratching motions." Based on this description, a target message is generated stating, "A suspicious person circled the car twice and made scratching motions; be highly vigilant." This message is then sent to the car owner's mobile phone to alert them. The owner can then view this message, understand the specifics of the dangerous event, and take timely action (e.g., call the police or change parking locations).
[0114] In another example, let's consider a dangerous event described as "At 10:19 PM last night, there was a suspected car scratching incident." Based on this description, target information "At 10:19 PM last night, there was a suspected car scratching incident" is generated and sent to the vehicle's onboard terminal to alert the owner. Once the owner powers on the vehicle, the onboard terminal displays this target information. In this way, the owner can view the target information, understand the specific details of the dangerous event, and react promptly.
[0115] In another example, let's consider a dangerous event described as "Suspicious individuals vandalized cars last night." Based on this description, the target video is categorized as a high-risk event—car vandalism. This allows users to selectively view the video based on the category tag to learn more about the dangerous event.
[0116] The technical solution provided in this disclosure has two advantages. First, since target information can be sent to the vehicle terminal or the owner's mobile device based on the description of the dangerous event, the owner can receive risk warnings immediately. Even if the user is not on-site, they can still understand the specific content of the dangerous event based on the target information and react in a timely manner, preventing the dangerous event from escalating due to information delays. Second, since the target video is classified into dangerous levels based on the description of the dangerous event, the user can view the target video in a targeted and intuitive way to understand the specific details of the dangerous event, thereby preventing a large number of unread videos from accumulating locally, thus improving the user experience.
[0117] Exemplary device
[0118] Figure 7This is a schematic diagram of a device for generating hazardous event content descriptions in sentinel mode, provided as an exemplary embodiment of the present disclosure. The device can be installed in electronic devices such as terminal devices and servers, or on objects such as vehicles, to execute the hazardous event content description generation method of any of the above embodiments of the present disclosure.
[0119] like Figure 7 As shown, the above-mentioned device 300 may include:
[0120] The first acquisition module 301 can be used to acquire target video around the vehicle in response to detecting a dangerous event of target danger level around the vehicle;
[0121] The first processing module 302 can be used to redetermine the danger level of the dangerous event based on the target video;
[0122] The second processing module 303 can be used to generate a hazard event content description of the hazard event in response to the re-determined hazard level being consistent with the target hazard level.
[0123] In one possible implementation, the first processing module 302 may be specifically used to: obtain a preset first prompt text;
[0124] Using the first prompt text, the visual language model is guided to process the target video to obtain the redefined danger level.
[0125] In one possible implementation, the second processing module 303 may be specifically used to: in response to the re-determined danger level being consistent with the target danger level, obtain a preset second prompt text; the second prompt text is used to prompt the generation style of the description of the danger event content;
[0126] Using the second prompt text, the visual language model is guided to process the target video and generate a description of the dangerous event.
[0127] The description of the hazardous event includes the specific content of the hazardous event and the redefined hazard level of the hazardous event.
[0128] In one possible implementation, the above-mentioned apparatus may further include:
[0129] The second acquisition module can be used to acquire a first image of the area surrounding the vehicle;
[0130] The first determining module can be used to determine the danger level of the dangerous event based on the first image.
[0131] In one possible implementation, the first determining module 301 can be specifically used for:
[0132] The first image is subjected to target detection based on a preset neural network model to obtain target detection results; wherein, the target detection results include at least one of the following: the distance between the vehicle and the target object, the number of target objects, and the time the target object stays around the vehicle;
[0133] Based on the target detection results, the danger level of the dangerous event is determined.
[0134] In one possible implementation, the above-mentioned apparatus may further include:
[0135] The third acquisition module can be used to acquire the initial visual language model;
[0136] The fourth acquisition module can be used to acquire multiple image-text pairs for the sentinel mode from the training dataset; wherein the multiple image-text pairs include image-text pairs with different danger levels;
[0137] The model training module can be used to train the visual language model based on the multiple image texts to obtain the trained visual language model.
[0138] In one possible implementation, the above-mentioned apparatus may further include:
[0139] The execution operation module can be used to perform target operations based on the description of the dangerous event content.
[0140] The target operation includes at least one of the following: generating and sending target information to the target device, and classifying the target video by risk level;
[0141] The target information is obtained based on the description of the dangerous event; the target device is the vehicle's in-vehicle terminal or the vehicle owner's mobile device.
[0142] In one possible implementation, the above-mentioned apparatus may further include:
[0143] The hazard level correction module can be used to correct the hazard level of the hazardous event based on the redefined hazard level in response to a discrepancy between the redefined hazard level and the target hazard level.
[0144] The beneficial technical effects corresponding to the exemplary embodiments of this device can be found in the corresponding beneficial technical effects of the exemplary method section above, and will not be repeated here.
[0145] Exemplary electronic devices
[0146] Figure 8 A structural diagram of an electronic device provided in an embodiment of this disclosure includes at least one processor 111 and a memory 112.
[0147] The processor 111 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 11 to perform desired functions.
[0148] The memory 112 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 111 may execute one or more computer program instructions to implement the methods for generating hazard event content descriptions in sentinel mode and / or other desired functions of the various embodiments of this disclosure described above.
[0149] In one example, the electronic device 11 may also include an input device 113 and an output device 114, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0150] The input device 113 may include various sensors, including but not limited to: a distance sensor for detecting the distance between a target object and the vehicle; an image sensor for acquiring information about the vehicle's surrounding environment. In some examples, the input device may also include a pressure sensor for detecting seat pressure to determine the presence and location of passengers; a temperature sensor for monitoring the temperature inside the cabin; a humidity sensor for monitoring the humidity inside the cabin to assist in regulating the in-vehicle environment; an air quality sensor for monitoring in-vehicle air quality, such as carbon dioxide and volatile organic compounds (VOCs); a light sensor for detecting the intensity of light inside and outside the vehicle; an acceleration sensor for detecting changes in the vehicle's acceleration; a distance sensor for detecting the distance between the vehicle and other objects; a touchscreen sensor for interaction with the vehicle's infotainment system; biometric sensors, such as fingerprint recognition and facial recognition; a heart rate monitor for monitoring the driver's heart rate; a sound sensor for voice recognition and interaction to enable voice control; a seat sensor for monitoring seat usage, such as whether the seat is occupied and the passenger's body size; and wireless communication sensors, such as Bluetooth and Wi-Fi, for connecting to smart devices to achieve data transmission and remote control. In addition to the examples given above, the input device may include more or fewer sensors, which will not be elaborated here.
[0151] The output device 114 can output various information or signals to other hardware or devices, which may include displays, car audio systems, seats, windows, steering wheels, communication networks, and their connected remote output devices. The displays may include multiple different displays such as a driver's side display, a passenger side display, and a rear-seat display. The car audio system may include multiple speakers located in different positions within the vehicle cabin, and each display or speaker can operate independently.
[0152] Of course, for the sake of simplicity, Figure 8 Only some of the components of the electronic device 11 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 11 may include any other suitable components depending on the specific application.
[0153] Exemplary computer program products and computer-readable storage media
[0154] In addition to the methods and apparatus described above, embodiments of this disclosure may also provide a computer program product, including computer program instructions that, when executed by a processor, cause the processor to perform steps in the methods for generating dangerous event content descriptions in sentinel mode according to the various embodiments of this disclosure described in the "Exemplary Methods" section above.
[0155] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. These programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0156] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform steps in the methods for generating dangerous event content descriptions in sentinel mode according to various embodiments of this disclosure described in the "Exemplary Methods" section above.
[0157] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may include, but is not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0158] The basic principles of this disclosure have been described above with reference to specific embodiments. However, the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0159] Various modifications and variations can be made to this disclosure without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, this disclosure is also intended to include such modifications and variations.
Claims
1. A method for generating content descriptions of dangerous events in Sentinel mode, comprising: In response to the detection of a hazardous event of a target hazard level around the vehicle, acquire video of the target around the vehicle; Based on the target video, the danger level of the dangerous event is reassessed; In response to the redefined hazard level being consistent with the target hazard level, a hazard event content description is generated for the hazard event.
2. The method according to claim 1, wherein, The step of redetermining the danger level of the dangerous event based on the target video includes: Get the preset first prompt text; Using the first prompt text, the visual language model is guided to process the target video to obtain the redefined danger level.
3. The method according to claim 2, wherein, The response, in response to the redefined hazard level being consistent with the target hazard level, generates a hazard event content description, including: In response to the redefined hazard level being consistent with the target hazard level, a preset second prompt text is obtained; the second prompt text is used to indicate the generation style of the description of the hazard event. Using the second prompt text, the visual language model is guided to process the target video and generate a description of the dangerous event. The description of the hazardous event includes the specific content of the hazardous event and the redefined hazard level of the hazardous event.
4. The method of claim 1, wherein before acquiring target video around the vehicle in response to detecting a hazardous event of a target hazard level occurring around the vehicle, the method further comprises: Acquire a first image of the area surrounding the vehicle; Based on the first image, the danger level of the dangerous event is determined.
5. The method according to claim 4, wherein, Determining the danger level of the hazardous event based on the first image includes: The first image is subjected to target detection based on a preset neural network model to obtain target detection results; wherein, the target detection results include at least one of the following: the distance between the vehicle and the target object, the number of target objects, and the time the target object stays around the vehicle; Based on the target detection results, the danger level of the dangerous event is determined.
6. The method according to claim 2, wherein, Before using the first prompt text to guide the visual language model to process the target video and obtain the redefined danger level, the method further includes: Obtain the initial visual language model; Multiple image-text pairs for Sentinel mode are obtained from the training dataset; wherein the multiple image-text pairs include image-text pairs with different danger levels; The visual language model is trained based on the multiple image-text pairs to obtain the trained visual language model.
7. The method according to claim 1, after generating a hazard event content description of the hazard event in response to the re-determined hazard level being consistent with the target hazard level, the method may further include: Based on the description of the hazardous event, perform the target operation; The target operation includes at least one of the following: generating and sending target information to the target device, and classifying the target video by risk level; The target information is obtained based on the description of the dangerous event; the target device is the vehicle's in-vehicle terminal or the vehicle owner's mobile device.
8. The method according to claim 1, further comprising: In response to a discrepancy between the redefined hazard level and the target hazard level, the hazard level of the hazardous event is revised based on the redefined hazard level.
9. An apparatus for generating content descriptions of dangerous events in Sentinel mode, comprising: The first acquisition module is used to acquire target video around the vehicle in response to detecting a dangerous event of target danger level around the vehicle; The first processing module is used to redetermine the danger level of the dangerous event based on the target video; The second processing module is used to generate a hazard event content description of the hazard event in response to the re-determined hazard level being consistent with the target hazard level.
10. A computer-readable storage medium storing a computer program for performing the method for generating a dangerous event content description for sentinel mode as described in any one of claims 1-8.
11. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method for generating dangerous event content descriptions for sentinel mode as described in any one of claims 1-8.