Vehicle monitoring method and device and storage medium
By continuously collecting images from vehicle cameras and judging dangerous events based on multiple frames, the problem of false alarms and wasted storage space in the traditional sentry mode is solved, and more accurate video recording and storage management is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG GEELY HLDG GRP CO LTD
- Filing Date
- 2026-01-22
- Publication Date
- 2026-04-17
AI Technical Summary
Traditional sentinel mode relies on sensor triggering, which is prone to false alarms, resulting in a large amount of useless video consuming storage space.
The vehicle's target camera continuously collects monitoring images, identifies target objects within a preset range, and determines event description information based on multiple frames of images, recording video only when a dangerous event is confirmed.
This avoids false alarms, reduces invalid video recordings, and saves storage space.
Smart Images

Figure CN121887941A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle control technology, specifically to a vehicle monitoring method, device, and storage medium. Background Technology
[0002] With the development of vehicle intelligence, Sentry Mode has become a standard feature of most vehicles, providing intelligent security functions for vehicles when parked.
[0003] Currently, the traditional Sentry Mode typically uses vibration sensors to detect when the vehicle vibrates or moves, then activates the vehicle's camera to record video of the area around the vehicle. Users can then review the video to confirm if any dangerous situations have occurred.
[0004] However, the traditional sentinel mode relies solely on sensor triggers, which can easily lead to false alarms and result in a large amount of useless video consuming storage space. Summary of the Invention
[0005] In view of this, this application aims to provide a vehicle monitoring method, device and storage medium to solve the problem that the sentry mode in traditional related technologies can only rely on sensor triggering, which is prone to false alarms and causes a large amount of useless video to occupy storage space.
[0006] The first aspect of this application provides a vehicle monitoring method, comprising: The target camera device controlling the vehicle continuously collects monitoring images around the vehicle, and identifies whether there are target objects within a preset range of the vehicle based on the monitoring images collected by the target camera device during the collection process; In response to the detection of a target object within a preset range of the vehicle, multiple target images containing the target object are obtained from the monitoring images captured by the target camera device; Based on the multi-frame target images, determine the event description information; In response to the event description information indicating a dangerous event, the target camera device is controlled to record video.
[0007] In one possible implementation of this application, determining event description information based on the multi-frame target images includes: inputting the multi-frame target images and preset prompt words into a visual language model to obtain event description information; and / or, identifying whether there is a target object within a preset range of the vehicle based on the monitoring images collected by the target camera device includes: identifying whether there is a target object within a preset range of the vehicle based on the monitoring images collected by the target camera device using a target detection model.
[0008] In one possible implementation of this application, the target camera device is one of at least two basic camera devices deployed on a vehicle; controlling the target camera device to record video includes controlling the at least two basic camera devices to record video.
[0009] In one possible implementation of this application, the method further includes: summarizing the event description information to obtain event summary information in response to the event description information indicating a dangerous event; and / or: triggering the vehicle's warning module to issue a warning prompt in response to the event description information indicating a dangerous event.
[0010] In one possible implementation of this application, the method further includes: pushing the event summary information to a user terminal connected to the vehicle.
[0011] In one possible implementation of this application, the target camera device is one of at least two basic camera devices deployed on a vehicle; the method further includes: in response to the event description information indicating a dangerous event, synchronizing the dangerous event with the state machines corresponding to each of the other camera devices, wherein the other camera devices are those other than the target camera device among the at least two basic camera devices; receiving event description information detected by the state machines corresponding to each of the other camera devices in response to the dangerous event; summarizing all the event description information, and obtaining event summary information based on the summarized event description information.
[0012] In one possible implementation of this application, the step of summarizing all event description information and obtaining event summary information based on the summarized event description information includes: summarizing all event description information to obtain summarized event description information; when the number of event description information meets the triggering condition, or when a set time has elapsed after determining that the event description information belongs to a dangerous event, triggering the generation of event summary information based on the summarized event description information.
[0013] In one possible implementation of this application, before determining the event description information based on the multi-frame target images, the method further includes: determining that the continuous occurrence duration of the target object reaches a preset duration threshold, wherein the continuous occurrence duration is determined by multi-frame target images containing the target object.
[0014] A second aspect of this application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the at least one processor to perform a vehicle monitoring method as described in the first aspect and possible implementations thereof.
[0015] A third aspect of this application provides a computer storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement a vehicle monitoring method as described in the first aspect and possible implementations thereof.
[0016] The fourth aspect of this application provides a computer program product comprising: a computer program that, when executed by a processor, implements a vehicle monitoring method as described in the first aspect and possible implementations thereof.
[0017] The vehicle monitoring method, equipment, and storage medium provided in this application continuously collect monitoring images around the vehicle using the vehicle's target camera device and identify whether there is a target object within a preset range of the vehicle based on the monitoring images. When a target object is detected in the monitoring image and is located within a preset range near the vehicle, multiple frames of target images of the target object are obtained from the monitoring image, and event description information is obtained based on the multiple frames of target images. When it is confirmed that the event description information belongs to a dangerous event, the vehicle's target camera device is controlled to record video. This avoids the false triggering and false alarms of the traditional sentry mode and reduces the recording of invalid video, thereby saving storage space. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this application, the drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram illustrating an application scenario for vehicle monitoring provided in an embodiment of this application.
[0020] Figure 2 Flowchart of the vehicle monitoring method provided in the embodiments of this application Figure 1 .
[0021] Figure 3 This is an example diagram illustrating the event summary information provided in an embodiment of this application.
[0022] Figure 4 Flowchart of the vehicle monitoring method provided in the embodiments of this application Figure 2 .
[0023] Figure 5 This is a schematic diagram of the vehicle monitoring device provided in an embodiment of this application.
[0024] Figure 6This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0026] In vehicle monitoring scenarios, the vehicle's sentry mode has a crucial impact on vehicle security. To address the issue that traditional sentry modes rely solely on sensor triggers, which are prone to false triggers and alarms, resulting in a large amount of useless video consuming storage space, this invention provides the following technical concept: The vehicle's target camera continuously collects monitoring images around the vehicle. When a target object is detected in the monitoring image and is located within a preset range near the vehicle, event description information is obtained based on multiple frames of the target object's appearance. If the event description information confirms that the event is a dangerous event, the vehicle's target camera is controlled to record video. This avoids the false triggers and alarms of traditional sentry modes while reducing the recording of invalid video, thus saving storage space.
[0027] Figure 1 This is a schematic diagram illustrating an application scenario for vehicle monitoring provided in an embodiment of this application. (Reference) Figure 1 The scenario includes a vehicle 10. The vehicle 10 includes a vehicle controller 101, one or more camera devices 102, and a local storage unit 103.
[0028] The vehicle controller 101 can be an Electronic Control Unit (ECU) or an onboard processor. The onboard processor can be a Graphics Processing Unit (GPU). The camera device 102 can be a single AVM camera or multiple AVM cameras. In one application scenario, four AVM cameras are used, installed on the four sides of the vehicle, specifically the front, rear, left, and right sides. It should be noted that the vehicle controller 101 maintains a state machine for each camera device 102. This state machine controls the corresponding camera device 102 to continuously acquire monitoring images around the vehicle. When a target object appears in the detected monitoring images, multiple frames of the target object are acquired for analysis of the large model to determine if a dangerous event has occurred, and video recording is performed. The recorded video is stored in the vehicle's local storage unit 103, which can be an SD (Secure Digital) card.
[0029] Exemplary methods Figure 2 Flowchart of the vehicle monitoring method provided in the embodiments of this application Figure 1 The execution entity in this embodiment can be... Figure 1 The vehicle controller in the illustrated embodiment. (As shown) Figure 2 As shown, the method includes: S201: Control the target camera device of the vehicle to continuously collect monitoring images around the vehicle, and identify whether there is a target object within a preset range of the vehicle based on the monitoring images collected by the target camera device during the collection process.
[0030] In the embodiments of this application, when the vehicle is turned off and enters sentry mode, the target camera device controlling the vehicle collects a preset number of monitoring images around the vehicle at preset intervals.
[0031] The target camera device controlling the vehicle captures a preset number of monitoring images around the vehicle every second. For example, the preset number could be 9 frames.
[0032] For example, the target camera device is an AVM camera, which is controlled to acquire 9 frames of monitoring images around the vehicle per second.
[0033] Specifically, by using the state machine corresponding to the target camera device, a target recognition model is used to detect whether there is a target object in the monitoring image collected by the target camera device, and whether the target object appears within the preset range of the vehicle.
[0034] In this embodiment, the target object can be a person or an object. The person can be someone from various professions, such as a food delivery person or a courier. The object can be a vehicle, including but not limited to two-wheeled vehicles, three-wheeled vehicles, and four-wheeled vehicles.
[0035] For example, the preset range of the vehicle can be a distance of 1 meter from the vehicle.
[0036] S202: In response to the detection of a target object within a preset range of the vehicle, acquire multiple frames of target images containing the target object from the monitoring images captured by the target camera device.
[0037] In the embodiments of this application, when it is identified that the monitoring image collected by the target camera device contains a target object and the target object appears within a preset range of the vehicle, it indicates that the target object poses a potential danger. Then, multiple frames of target images containing the target object are obtained from the collected monitoring images.
[0038] In the embodiments of this application, the collected monitoring images are obtained by calling back the monitoring images collected by the target camera device, and multiple frames of target images containing the target object are obtained from the collected monitoring images.
[0039] The multi-frame target images containing the target object can be multi-frame target images of the target object appearing continuously for a preset duration. For example, the multi-frame target images of the preset duration can be target images captured within any 2 seconds of the target object appearing continuously.
[0040] S203: Determine event description information based on multiple frames of target images.
[0041] In the embodiments of this application, event description information can be obtained by analyzing multiple frames of target images using a large visual language model.
[0042] The Visual Language Model (VLM) is trained using a large number of collected target object images and preset prompts.
[0043] S204: In response to an event description message indicating a dangerous event, control the target camera device to record video.
[0044] In the embodiments of this application, the event description information is classified by keywords, and the event description information is classified into dangerous events or non-dangerous events.
[0045] For example, if the event description information is "at a certain time - a minor collision - a food delivery driver - a two-wheeled vehicle (at a certain time, a food delivery driver riding a two-wheeled vehicle had a minor collision with our vehicle)," and contains the keyword "minor collision," then it is classified as a dangerous event.
[0046] For example, if the event description information is "at a certain time - passing by - courier - two-wheeled vehicle (at a certain time, a courier rode a two-wheeled vehicle and passed by our vehicle)," and contains the keyword "passing by," then it is determined to be classified as a non-dangerous event.
[0047] In the embodiments of this application, when the event description information indicates a dangerous event, the basic sentry mode of the vehicle is activated, and the target camera device of the vehicle is controlled to record video by activating the basic sentry mode.
[0048] The video is recorded in high definition and then stored in the vehicle's local storage unit.
[0049] In embodiments of this application, in response to an event description indicating a non-dangerous event, video recording is not performed for the non-dangerous event.
[0050] As described above, the vehicle's target camera continuously collects monitoring images around the vehicle and identifies whether there is a target object within a preset range of the vehicle based on the monitoring images. When a target object is detected in the monitoring image and is located within a preset range near the vehicle, multiple target images of the target object are obtained from the monitoring image. Based on the multiple target images, event description information is obtained. When it is confirmed that the event description information belongs to a dangerous event, the vehicle's target camera is controlled to record video. This avoids the false triggering and false alarm of the traditional sentry mode and reduces the recording of invalid video, thus saving storage space.
[0051] In one embodiment of this application, determining event description information based on multiple frames of target images in step S203 above may specifically include: Input multiple target images and preset prompts into the visual language model to obtain event description information.
[0052] In the embodiments of this application, multiple target images and preset prompt words are fused to obtain fused prompt words, and the fused prompt words are input into a large visual language model to obtain event description information.
[0053] In this embodiment, multiple target images are encoded according to a preset format and then fused with preset prompt words to obtain fused prompt words.
[0054] The default format can be base64.
[0055] In this embodiment, the Visual Language Model (VLM) is trained using a large number of collected target object images and preset prompts.
[0056] For example, an example of a fusion prompt word is as follows: "You are the AI sentinel of the car, and you need to analyze the following car camera data. {Task Description - Actions to be identified and output results}{img_base64}".
[0057] In this embodiment, the event description information includes, but is not limited to: time information, action type, personnel type, vehicle type, and a description of the event process.
[0058] For example, an example of an event description message is as follows: "At some point - a minor collision - a delivery driver - a two-wheeled vehicle (At some point, a delivery driver riding a two-wheeled vehicle had a minor collision with our vehicle)."
[0059] As can be seen from the above description, by using a large visual language model to process multiple target images and preset prompts, accurate event description information can be obtained, thereby more accurately determining whether the event description information belongs to a dangerous event.
[0060] In one embodiment of this application, the step S201 above, which involves detecting whether there is a target object within a preset range of the vehicle based on the monitoring image acquired by the target camera device, specifically includes: The target detection model identifies whether there is a target object within a preset range of the vehicle based on the monitoring images collected by the target camera equipment.
[0061] In the examples of this application, the target recognition model can be a YOLO model or a Faster R-CNN model.
[0062] In this embodiment, the target objects to be detected are registered with the target recognition model through a pre-deployed VPS (Vision Perception Service). For example, people, two-wheeled vehicles, three-wheeled vehicles, and four-wheeled vehicles are registered with the target recognition model through the VPS.
[0063] As can be seen from the above description, the accuracy of target object detection can be improved by using a target recognition model.
[0064] In one embodiment of this application, the target camera device is one of at least two basic camera devices deployed on a vehicle; correspondingly, in step S204, controlling the target camera device to perform video recording includes: Control at least two basic camera devices to record video.
[0065] In this embodiment, there may be four basic camera devices, each installed on one of the four sides of the vehicle. When any of the target camera devices determines that an event description indicates a dangerous event, multiple camera devices are controlled to simultaneously record video.
[0066] As can be seen from the above description, comprehensive video recording can be achieved by controlling multiple basic camera devices to record video simultaneously.
[0067] In one embodiment of this application, the above-mentioned at least two basic camera devices are deployed at different locations on the vehicle.
[0068] Specifically, multiple camera devices can be installed at different locations on the side of the vehicle. For example, there are four camera devices, installed on the front, rear, left, and right sides of the vehicle respectively.
[0069] In this example, multiple cameras located at different positions on the vehicle enable comprehensive monitoring of the area around the vehicle, further ensuring vehicle safety.
[0070] In one embodiment of this application, the method further includes the following after step S204: S205: In response to the event description information indicating a hazardous event, the event description information is summarized to obtain event summary information.
[0071] In the embodiments of this application, in response to the event description information indicating a dangerous event, the visual language big data model is triggered to summarize the event description information to obtain event summary information.
[0072] In the embodiments of this application, event description information is input into the visual language big model so that the visual language big model can obtain event summary information of the dangerous event based on the event description information and the context.
[0073] In embodiments of this application, the event summary information includes the event sequence and a summary. The event sequence includes a summary description of key scenarios and / or corresponding images of key targets; and / or, the summary includes event conclusions and risk warnings.
[0074] refer to Figure 3 , Figure 3 This is an example diagram illustrating the event summary information provided in an embodiment of this application. Figure 3 In this example, the key target image consists of two frames. The scene description summary for the first key target image is: someone approaches the vehicle at 15:30. The scene description summary for the second key target image is: someone leans against the vehicle at 15:40. The event conclusion and risk warning are: someone is suspected of scratching the vehicle, with a medium risk.
[0075] As can be seen from the above description, the event summary information is obtained by summarizing the event description information, rather than just recording videos, so that users can be informed of dangerous events in a timely manner.
[0076] In one embodiment of this application, after step S205, the method further includes: S206: Push event summary information to the user terminal connected to the vehicle.
[0077] In this embodiment, the event summary information is sent to the user terminal connected to the vehicle via the vehicle's onboard communication module.
[0078] Among them, the communication module can be a mobile communication module, such as a 4G or 5G mobile communication module.
[0079] In another embodiment, the event summary information can also be pushed to the vehicle's infotainment system. Specifically, the event summary information is sent to the vehicle's infotainment system via the vehicle's data bus. The vehicle's data bus can be CAN (Controller Area Network, serial communication bus).
[0080] As can be seen from the above description, by pushing event summary information to user terminals, users can obtain event summary information of dangerous events in a timely manner, which improves processing efficiency compared to existing technologies where users can only view recorded videos.
[0081] refer to Figure 4 , Figure 4 Flowchart of the vehicle monitoring method provided in the embodiments of this application Figure 2 In this embodiment, the target camera device is one of at least two basic camera devices deployed on the vehicle, and the method further includes: S401: In response to the event description information indicating a dangerous event, synchronize the dangerous event with the state machine corresponding to each of the other camera devices, wherein the other camera devices are camera devices other than the target camera device among at least two basic camera devices.
[0082] In this embodiment, the state machine of each camera device synchronously executes steps S201 to S204 as described above, and obtains event description information corresponding to each camera device. When any camera device determines that the event description information is a dangerous event, it synchronizes the state machine corresponding to all other camera devices with this dangerous event.
[0083] In this embodiment, synchronizing the dangerous event with the state machines corresponding to all other camera devices can be achieved by synchronizing the event identifier (session ID) of the dangerous event with the state machines corresponding to all other camera devices.
[0084] S402: Receive event description information from the state machines of other camera devices for the dangerous events they detect.
[0085] In this embodiment, when any target camera determines that the event description information is a dangerous event, the state machines corresponding to all other camera devices are synchronized so that each state machine summarizes the event description information within the same time period under the dangerous event.
[0086] In this embodiment, all event description information within the same time period is appended with the event identifier (session ID) of the dangerous event.
[0087] For example, the time interval can be 15 minutes or 30 minutes.
[0088] S403: Summarize all event description information and obtain event summary information based on the summarized event description information.
[0089] Specifically, all event description information is summarized and then input into the visual language big data model to obtain event summary information.
[0090] In this embodiment, all event description information is summarized to obtain summarized event description information.
[0091] For example, an example of a summarized event description is as follows: "{At a certain moment, a deliveryman on a two-wheeled vehicle grazed our vehicle; at a certain moment, the deliveryman approached the vehicle; at a certain moment, the deliveryman moved away from the vehicle}."
[0092] In this embodiment, the summarized event description information is input into the visual language big model, so that the visual language big model can obtain the event summary information of the dangerous event based on the summarized event description information and combined with the context content.
[0093] As can be seen from the above description, after any target camera detects a dangerous event, the event summary information of each camera is summarized to obtain a summary event description information, which avoids the problem of repeated reporting and video recording of the same dangerous event by each camera.
[0094] In one embodiment of this application, step S403 above, which involves summarizing all event description information and obtaining event summary information based on the summarized event description information, may specifically include: S431. Summarize all the event description information to obtain the summarized event description information.
[0095] S432. When the amount of event description information meets the triggering condition, or when the event description information is determined to be a dangerous event and a set time has elapsed, the event summary information is obtained based on the summarized event description information.
[0096] Specifically, the summarized event description information is input into the visual language big model to obtain event summary information.
[0097] In this embodiment, the quantity meeting the triggering condition can be that the number of event description messages reaches 4.
[0098] In this embodiment, the set duration can be 15 minutes or 30 minutes.
[0099] As can be seen from the above description, when the number of summarized event description information meets the triggering condition or the set duration triggering condition, the summarized event description information is summarized to obtain event summary information, thus avoiding the problem of frequent triggering of event description information summary, which leads to reduced processing efficiency.
[0100] In one embodiment of this application, before step S203 described above, the following step is further included: The duration of continuous appearance of the target object is determined by a preset duration threshold, which is determined by multiple frames of target images containing the target object.
[0101] Specifically, the duration of continuous appearance of the target object is determined by using multiple frames of target images in which the target object appears; when the duration of continuous appearance of the target object reaches a preset duration threshold, step "S203" is triggered.
[0102] In this embodiment, the preset duration threshold can be set according to requirements.
[0103] For example, the preset duration threshold is 2 seconds.
[0104] For example, if the duration of the continuous appearance of the target object is less than 2 seconds, it indicates that the target object only passes by briefly, and the execution of step "S203" is not triggered; if the duration of the continuous appearance of the target object is greater than or equal to 2 seconds, it indicates that the target object does not pass by briefly, and the execution of step "S203" is triggered.
[0105] As can be seen from the above description, if the duration of continuous occurrence of the target object is less than the set duration threshold, the action of obtaining event description information will not be triggered, thus avoiding frequent and invalid analysis of event description information.
[0106] In one embodiment of this application, after obtaining event description information based on multiple target images in step S203, the method further includes: S207: In response to the event description information indicating a dangerous event, the vehicle's warning module is triggered to issue a warning.
[0107] In this embodiment, triggering the vehicle's warning module to issue a warning can be achieved by either lighting up the vehicle's infotainment display or by triggering the vehicle's infotainment system to play a voice message or alarm tone through the vehicle's speaker.
[0108] As can be seen from the above description, when it is confirmed that a dangerous event has occurred against the vehicle by a target object, the vehicle's warning module will issue a warning to alert the target object, so as to prevent the dangerous situation from continuing or escalating, thus achieving the effect of protecting the vehicle.
[0109] Exemplary device Figure 5 This is a schematic diagram of the vehicle monitoring device provided in an embodiment of this application. Figure 5 As shown, the vehicle monitoring device includes: an image detection unit 501, an image extraction unit 502, an event description unit 503, and a recording control unit 504.
[0110] The image detection unit 501 is used to control the vehicle's target camera to continuously collect monitoring images around the vehicle, and to identify whether there is a target object within a preset range of the vehicle based on the monitoring images collected by the target camera during the collection process.
[0111] The image extraction unit 502 is used to obtain multiple frames of target images containing the target object from the monitoring images collected by the target camera device in response to the recognition that a target object is detected within a preset range of the vehicle.
[0112] The event description unit 503 is used to determine event description information based on the multi-frame target images.
[0113] The recording control unit 504 is used to control the target camera device to record video in response to the event description information indicating a dangerous event.
[0114] In one or more embodiments of this application, the event description unit 503 is specifically used to input the multi-frame target images and preset prompt words into a visual language big model to obtain event description information; the image detection unit 501 is specifically used to identify whether there is a target object within a preset range of the vehicle based on the monitoring images collected by the target camera device through a target detection model.
[0115] In one or more embodiments of this application, the target camera device is a camera device among at least two basic camera devices deployed on a vehicle; the recording control unit 504 is used to control the at least two basic camera devices to perform video recording.
[0116] In one or more embodiments of this application, the at least two basic camera devices are deployed at different locations on the vehicle.
[0117] In one or more embodiments of this application, the event description unit 503 is further configured to summarize the event description information in response to the event description information indicating a dangerous event, and obtain event summary information.
[0118] In one or more embodiments of this application, the event summary information includes the event process and a summary content.
[0119] In one or more embodiments of this application, the event process includes a summary of key scene descriptions and / or corresponding key target images; and / or, the summary includes event conclusions and risk warnings.
[0120] In one or more embodiments of this application, the device further includes: an information push unit 505 (see reference 505). Figure 5 This is used to push the event summary information to the user terminal connected to the vehicle.
[0121] In one or more embodiments of this application, the event description unit 503 is further configured to: in response to the event description information indicating a dangerous event, synchronize the dangerous event with the state machines corresponding to each other camera device, wherein the other camera devices are camera devices other than the target camera device among the at least two basic camera devices; receive event description information detected by the state machines corresponding to each other camera device for the dangerous event; summarize all the event description information, and obtain event summary information based on the summarized event description information.
[0122] In one or more embodiments of this application, the event description unit 503 is further specifically used to: summarize all event description information to obtain summarized event description information; when the number of event description information meets the triggering condition, or when a set time has elapsed after determining that the event description information belongs to a dangerous event, trigger the generation of event summary information based on the summarized event description information.
[0123] In one or more embodiments of this application, the event description unit 503 is further configured to: determine that the continuous occurrence duration of the target object reaches a preset duration threshold, wherein the continuous occurrence duration is determined by multiple frames of target images containing the target object.
[0124] In one or more embodiments of this application, the device further includes: a warning notification unit 506 (see reference 506). Figure 5 In response to the event description information indicating a dangerous event, the vehicle's warning module is triggered to issue a warning.
[0125] The apparatus provided in this application embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.
[0126] Exemplary electronic devices Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application. Figure 6As shown, the electronic device of this embodiment includes a processor 601 and a memory 602.
[0127] The memory 602 stores computer execution instructions; the processor 601 executes the computer execution instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiments. For details, please refer to the relevant descriptions in the foregoing method embodiments.
[0128] Alternatively, the memory 602 can be either standalone or integrated with the processor 601.
[0129] When the memory 602 is set up independently, the electronic device also includes a bus 603 for connecting the memory 602 and the processor 601.
[0130] Exemplary vehicles, media, and products This application embodiment also provides a vehicle, which includes: a vehicle body, the vehicle body being configured with the above-described vehicle controller and camera equipment, wherein the vehicle controller is used to execute the above-described vehicle monitoring method.
[0131] This application also provides a computer storage medium storing computer execution instructions. When the processor executes the computer execution instructions, the above-described vehicle monitoring method is implemented.
[0132] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the vehicle monitoring method described above.
[0133] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0134] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.
[0135] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit composed of the above modules can be implemented in hardware or in the form of hardware plus software functional units.
[0136] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0137] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules within the processor.
[0138] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk or optical disc, etc.
[0139] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0140] The aforementioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0141] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. Both the processor and the storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components in an electronic device or host device.
[0142] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A vehicle monitoring method, characterized in that, include: The target camera device controlling the vehicle continuously collects monitoring images around the vehicle, and identifies whether there are target objects within a preset range of the vehicle based on the monitoring images collected by the target camera device during the collection process; In response to the detection of a target object within a preset range of the vehicle, multiple target images containing the target object are obtained from the monitoring images captured by the target camera device; Based on the multi-frame target images, determine the event description information; In response to the event description information indicating a dangerous event, the target camera device is controlled to record video.
2. The method of claim 1, wherein, The determination of event description information based on the multi-frame target images includes: The multi-frame target images and preset prompts are input into the visual language model to obtain event description information. And / or, The method of identifying whether there is a target object within a preset range of the vehicle based on the monitoring images acquired by the target camera device includes: The target detection model identifies whether there is a target object within a preset range of the vehicle based on the monitoring images collected by the target camera equipment.
3. The method of claim 1, wherein, The target camera device is one of at least two basic camera devices deployed on the vehicle; The control of the target camera device to perform video recording includes: Control the at least two basic camera devices to record video.
4. The method of claim 1, wherein, Also includes: In response to the event description information indicating a dangerous event, the event description information is summarized to obtain event summary information; and / or; In response to the event description information indicating a dangerous event, the vehicle's warning module is triggered to issue a warning.
5. The method of claim 4, wherein, Also includes: The event summary information is pushed to the user terminal connected to the vehicle.
6. The method of claim 1, wherein, The target camera device is one of at least two basic camera devices deployed on the vehicle; The method further includes: In response to the event description information indicating a dangerous event, the state machine corresponding to each of the other camera devices is synchronized with the dangerous event, wherein the other camera devices are camera devices other than the target camera device among the at least two basic camera devices; Receive event description information detected by the state machines of other camera devices in response to the dangerous event; All event descriptions are compiled, and event summary information is obtained based on the compiled event descriptions.
7. The method of claim 6, wherein, The process involves summarizing all event descriptions and obtaining event summary information based on the summarized event descriptions, including: All event descriptions are summarized to obtain the summarized event descriptions. When the amount of event description information meets the triggering condition, or when a set time has elapsed after determining that the event description information belongs to a dangerous event, the event summary information is obtained based on the summarized event description information.
8. The method of claim 1, wherein, Before determining the event description information based on the multi-frame target images, the process also includes: The duration of continuous occurrence of the target object is determined to reach a preset duration threshold, and the duration of continuous occurrence is determined by multiple frames of target images containing the target object.
9. An electronic device, comprising: include: At least one processor; The system also includes a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the at least one processor to perform the vehicle monitoring method according to any one of claims 1 to 8.
10. A computer storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the vehicle monitoring method as described in any one of claims 1 to 8.