Event prompting method and device based on camera and storage medium

By identifying the perception status of target objects in camera video data and selecting appropriate visual or auditory cues, the energy consumption and noise pollution problems of existing cameras when detecting target events are solved, achieving better security cues and power consumption control.

CN121963410APending Publication Date: 2026-05-01HANGZHOU HUACHENG SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU HUACHENG SOFTWARE TECH CO LTD
Filing Date
2025-12-04
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

When existing cameras detect a target event, they usually immediately activate audio and visual alerts, resulting in unnecessary energy consumption and environmental noise pollution, making them unsuitable for all application scenarios.

Method used

By identifying the perception state of target objects in camera video data, their visual and auditory perception states can be determined, thereby selecting appropriate visual or auditory cueing strategies for cueing processing and avoiding unnecessary energy consumption and noise pollution.

Benefits of technology

It enables the selection of appropriate prompting strategies based on actual application scenarios, issuing reasonable and effective event prompts, and reducing unnecessary energy consumption and noise pollution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963410A_ABST
    Figure CN121963410A_ABST
Patent Text Reader

Abstract

The invention discloses an event prompting method and device based on a camera and a storage medium, and the method comprises the steps: obtaining the video data of a target event in response to the triggering of the detected target event; identifying a target object in the video data to obtain a sensing state of the target object; determining a target prompt strategy from preset prompt strategies according to the sensing state; and performing prompt processing on the target object according to the target prompt strategy. According to the scheme, a reasonable and effective event prompt can be given out when the target event is detected.
Need to check novelty before this filing date? Find Prior Art

Description

Camera-based event notification methods, devices, and storage media Technical Field

[0001] This application relates to the field of target detection technology, and in particular to a camera-based event notification method, device, and storage medium. Background Technology

[0002] With the development of technologies such as target detection and security, cameras are now commonly seen in almost every aspect of life to achieve functions such as intelligent detection and audio-visual alerts.

[0003] For example, an AOV (Always-On Video) camera is an image acquisition device that uses intelligent low-power technology. Its core principle is to operate in low-power mode to ensure battery life when no target event occurs, and switch to normal mode when a target event is detected. Therefore, an AOV camera can trigger an audio-visual alert when a target event is detected.

[0004] However, existing methods typically activate audio-visual devices to issue alerts immediately upon detecting a target event. In reality, this alert mode is not suitable for all application scenarios and can also generate unnecessary energy consumption and environmental noise and light pollution. Summary of the Invention

[0005] This application provides at least one camera-based event notification method, apparatus, device, and computer-readable storage medium.

[0006] The first aspect of this application provides a camera-based event prompting method, comprising: in response to detecting a target event trigger, acquiring video data of the target event; identifying a target object in the video data and obtaining the perception state of the target object; determining a target prompting strategy from a preset prompting strategy based on the perception state; and performing prompting processing on the target object according to the target prompting strategy.

[0007] In one embodiment, identifying a target object in the video data and obtaining the perception state of the target object includes: identifying a target object in the video data and obtaining a target identification result; and determining the perception state of the target object based on the target identification result and the image features of the video data.

[0008] In one embodiment, the perception state includes a visual perception state and / or an auditory perception state, the preset prompting strategy includes a visual prompting strategy and an auditory prompting strategy, and the step of determining a target prompting strategy from the preset prompting strategies based on the perception state includes: in response to the visual perception state being restricted and the auditory perception state being unrestricted, determining the auditory prompting strategy as the target prompting strategy; and in response to the visual perception state being unrestricted and the auditory perception state being restricted, determining the visual prompting strategy as the target prompting strategy.

[0009] In one embodiment, after processing the target object according to the target prompting strategy, the method further includes: obtaining the behavioral intent of the target object; determining whether the target prompting strategy is effective based on the behavioral intent; if the target prompting strategy is ineffective, performing prompting escalation processing on the target prompting strategy to obtain an escalated prompting strategy; and processing the target object according to the escalated prompting strategy.

[0010] In one embodiment, determining the target prompting strategy from the preset prompting strategy based on the perception state includes: acquiring the behavioral intent of the target object; and determining the target prompting strategy from the preset prompting strategy based on the perception state of the target object and the behavioral intent.

[0011] In one embodiment, after acquiring video data of the target event in response to the detection of the target event, the method further includes: identifying a target object in the video data and obtaining a confidence level of the target object; and in response to the confidence level being less than a confidence level threshold, acquiring a fixed prompting strategy to provide prompts to the target object.

[0012] In one embodiment, after acquiring video data of the target event in response to detecting the target event trigger, the method further includes: identifying candidate objects in the video data, obtaining the number and type of the candidate objects; and determining the target object from each candidate object according to the object type if the number of objects is multiple.

[0013] In one embodiment, determining the target object from candidate objects based on the object type includes: obtaining the behavioral intent of each candidate object; determining the object priority corresponding to each candidate object based on the object type and the behavioral intent; and determining the target object from each candidate object based on the object priority.

[0014] A second aspect of this application provides a camera-based event prompting device, comprising: an event detection module for acquiring video data of the target event in response to the detection of a target event; an object recognition module for identifying a target object in the video data and obtaining the perception state of the target object; a strategy determination module for determining a target prompting strategy from a preset prompting strategy based on the perception state; and a prompting module for prompting the target object according to the target prompting strategy.

[0015] A third aspect of this application provides an electronic device, including a memory and a processor, wherein the processor is configured to execute program instructions stored in the memory to implement the aforementioned camera-based event notification method.

[0016] The fourth aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a processor, implement the aforementioned camera-based event prompting method.

[0017] The above scheme, upon detecting a target event, acquires video data collected by the camera in response to the event. By identifying the target object in the video data, the perceptual state of the target object can be obtained, which may include visual and / or auditory perceptual states. Analyzing the perceptual state of the target object can determine its more receptive perception mode to external information (cues), thus allowing the determination of a target cues strategy from preset cues strategies based on the perceptual state. The target cues strategy is then applied to the target object for cues processing. This allows for the selection of a matching target cues strategy based on the target object's receptive perception mode to provide cues, unlike traditional methods that apply a uniform cues strategy regardless of the type of target event or the type of target object. The method in this application selects an appropriate cues strategy based on the actual application scenario, enabling the issuance of reasonable and effective event cues.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0020] Figure 1 is a flowchart illustrating an exemplary embodiment of the camera-based event notification method of this application; Figure 2 is a block diagram illustrating a camera-based event notification device in an exemplary embodiment of this application; Figure 3 is a structural schematic diagram of an embodiment of the electronic device of this application; Figure 4 is a structural schematic diagram of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0021] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0022] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0023] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0024] To facilitate understanding, one of the applicable scenarios of this application will be illustrated by example.

[0025] With the development of technologies such as target detection and security, cameras are now commonly seen in almost every aspect of life to achieve functions such as intelligent detection and audio-visual alerts.

[0026] For example, an AOV (Always-On Video) camera is an image acquisition device that uses intelligent low-power technology. Its core principle is to operate in low-power mode to ensure battery life when no target event occurs, and switch to normal mode when a target event is detected. Therefore, an AOV camera can trigger an audio-visual alert when a target event is detected.

[0027] However, existing methods typically activate audio-visual devices to issue alerts immediately upon detecting a target event (such as tripwire detection or perimeter detection). In reality, this alert mode is not suitable for all application scenarios and can also generate unnecessary energy consumption and environmental noise and light pollution.

[0028] The event notification method of this application can achieve intelligent identification and adaptive hierarchical response of target objects in target events through software algorithm innovation without changing or with minimal changes to existing hardware, so as to achieve a better balance between security notification effect and power consumption control.

[0029] It should be noted that the event notification method of this application can be applied to image acquisition devices, such as cameras in normal power consumption continuous recording mode and cameras in low power consumption AOV mode. Alternatively, it can also be applied to data processing devices that have a communication connection with the camera (e.g., the camera transmits the acquired data to the data processing device for analysis and processing, and then the data processing device returns the corresponding control commands to the camera). To demonstrate the power consumption control advantages of the method of this application, the following description mainly uses an AOV camera as an example.

[0030] Please refer to Figure 1, which is a flowchart illustrating an exemplary embodiment of the camera-based event notification method of this application. Specifically, it may include the following steps: Step S110, in response to detecting a target event trigger, acquiring video data of the target event.

[0031] In some application scenarios, the camera can detect whether the target event has been triggered, or other detection devices that have a communication connection with the camera can detect whether the target event has been triggered; this is not limited here.

[0032] Other detection devices may include, but are not limited to, radar, photoelectric sensors, radio frequency sensors, and microphone arrays. After detecting a target event, these other detection devices can send a trigger signal indicating "target event triggered" to the image acquisition device; details will not be elaborated here.

[0033] For example, a microwave radar module (such as the LD2411S) can be installed on an AOV camera. In AOV standby mode, only the microwave radar module operates at microampere-level power consumption. Only when the radar detects motion within a preset range is the image sensor and main processor of the AOV camera powered on and put into normal operation.

[0034] For example, a microphone array can be installed on an AOV camera. The system continuously monitors ambient sound. When the microphone array detects a preset target sound event (such as the sound of breaking glass or a loud impact), it can trigger the image sensor and main processor in the AOV camera to power on and begin normal operation. Alternatively, even if the target object is not subsequently identified from the video data, a corresponding preset prompting strategy can be applied based on the target sound event.

[0035] For example, for an AOV camera, after determining that a target event has been triggered, it can acquire video data of the target event through its image sensor to obtain a low frame rate AOV video stream; or it can switch from AOV mode to continuous recording mode and acquire a continuous high-definition video stream.

[0036] In another example, for a continuous recording camera, after determining that a target event has been triggered, the video segment corresponding to the occurrence of the target event can be extracted from the continuously recorded high-definition video stream to obtain the video data of the target event.

[0037] Step S120: Identify the target object in the video data and obtain the perception state of the target object.

[0038] The target object can refer to a person or other animal; there is no limitation here.

[0039] Perceptual state refers to the state of a target object when it perceives external information. Perception here can be understood as awareness, feeling, attention, and perception.

[0040] For example, common perceptions can include, but are not limited to, visual perception, auditory perception, and tactile perception. Visual perception refers to acquiring visual information (such as light information or text information); auditory perception refers to acquiring auditory information (such as sound information).

[0041] Therefore, the event prompting method of this application can also mainly revolve around generating light and / or sound to achieve visual and auditory prompts, which will not be elaborated here.

[0042] Step S130: Determine the target prompting strategy from the preset prompting strategies based on the perception state.

[0043] Based on the steps outlined above, the perception state can be mainly divided into limited perception and unlimited perception.

[0044] It is understandable that if a certain perception is restricted, it means that the target object cannot obtain the corresponding perceptual information from this perceptual pathway, or can only obtain a small amount of perceptual information.

[0045] Therefore, this application starts by analyzing the perceptual state of the target object, determines the prompting method that is relatively easy for the target object to receive information based on the perceptual state of the target object, and then determines the appropriate target prompting strategy from the preset prompting strategies that include multiple prompting methods.

[0046] Step S140: Provide prompts to the target object according to the target prompting strategy.

[0047] Based on the steps described above, after determining the target prompting strategy, the image acquisition device, or a prompting device that is connected to the image acquisition device, or a smart device that is connected to the image acquisition device, can be controlled to provide prompts to the target object.

[0048] The methods for handling prompts may include, but are not limited to, sound prompts (auditory prompts), light prompts (visual prompts), and sound-light prompts (auditory prompts + visual prompts), etc., and are not limited here.

[0049] For example, if a target object is prompted through an image acquisition device, the image acquisition device needs to be equipped with a sound playback element (such as a speaker) and / or a light emission element (such as a flash, red and blue light, spotlight) in an integrated or separate manner.

[0050] The auditory cues may include, but are not limited to, preset ringtones or voice messages, or voice content generated based on target events and / or target objects in the actual application scenario.

[0051] Similarly, visual cues may include, but are not limited to, preset light flashing patterns, or the ability to connect to a display panel and display preset text content, or text content generated based on target events and / or target objects in the actual application scenario.

[0052] In another example, if the target object is prompted by a prompting device that has a communication connection with the image acquisition device, the prompting device may include, but is not limited to, a speaker, an audio device, a light board, a display screen, a flashlight, etc., without limitation here.

[0053] As another example, the smart device that has a communication connection with the image acquisition device may include, but is not limited to, vehicles equipped with V2X (Vehicle to Everything) vehicle networking technology, and / or some wearable devices, mobile terminals, etc.

[0054] For example, image acquisition devices can be connected to intelligent vehicle systems. When a target object is detected inside an intelligent car with vehicle networking capabilities, the device's perception state can be determined to be "more receptive to prompts from the vehicle's screen and / or audio system." In this case, the cross-device prompting strategy in the preset prompting strategy can be selected as the target prompting strategy to provide prompts to the target object.

[0055] Cross-device prompting strategy refers to the prompting strategy implemented based on communication connections that span between the image acquisition device and other smart devices.

[0056] Specifically, communication technologies such as V2X can be used to send alerts (such as "Pedestrian crossing, please be careful") directly to the vehicle's central control display and / or in-vehicle audio system. Compared to alert strategies using cameras or supplementary lighting, cross-device alert strategies may be more accurate and effective in some scenarios.

[0057] For example, when visual recognition detects that a target is wearing a smartwatch or smart glasses (based solely on device shape recognition without involving personal data), it can be determined that the tactile sensing of the smartwatch or the near-eye display of the smart glasses are more effective sensing methods. Therefore, when device linkage is supported, cross-device prompting strategies such as vibration or visual cues can be implemented.

[0058] As can be seen, this application can acquire video data collected by the camera in response to the target event when the target event is detected. By identifying the target object in the video data, the perceptual state of the target object can be obtained. These perceptual states can include visual and / or auditory perceptual states. Analyzing the perceptual state of the target object can determine its perception mode that is more likely to receive external information (cue information). Therefore, a target cue strategy can be determined from the preset cue strategies based on the perceptual state. Then, the target cue strategy is used to cue the target object. This allows for the selection of a matching target cue strategy for the target object's perception mode that is more likely to receive external information, rather than using a uniform cue strategy regardless of the type of target event or the type of target object in traditional methods. The method of this application selects an appropriate cue strategy based on the actual application scenario, enabling the issuance of reasonable and effective event cuees.

[0059] Based on the above embodiments, this embodiment provides an example of an image acquisition device capable of executing the event notification method of this application.

[0060] At the hardware level, image acquisition devices may include, but are not limited to, image sensors used to acquire AOV video streams.

[0061] The processor (CPU / MCU) serves as the core of the event notification system's computation.

[0062] Memory is used to store program instructions and data that implement related functions.

[0063] The prompting unit may include at least one of a white light, a red and blue light, a speaker, etc., for emitting light and / or sound to form visual and / or auditory prompts.

[0064] At the software level, the processor of the image acquisition device can run a target object recognition and confidence assessment module, which can identify target objects in the received video data and output the corresponding confidence scores.

[0065] Perception State Analysis Module: This module can analyze the perception state of a target object.

[0066] Adaptive strategy decision engine: It contains a preset strategy library and a strategy arbitrator, which can generate adaptive response instructions based on the target object's perception state and / or other relevant information, and provide corresponding prompts to the target object.

[0067] The strategy effect self-learning module records relevant data of the target event and the effect of prompting the target object, and can optimize the corresponding prompting strategy by learning from this data.

[0068] For example, an image acquisition device can continuously acquire video data of a target event through its image sensor, or after a target event is detected. Then, a lightweight AI recognition model (such as quantized YOLOv5s, MobileNetV3-SSD) deployed on a processor (main processor or coprocessor) can be used to analyze the video data (video frames) in real time, identify the target object in the video data, and output its bounding box and confidence score.

[0069] This confidence level can serve as an important basis for subsequent decision-making strategies. It is understandable that when the confidence level is lower than the preset confidence threshold (such as 0.7), it indicates that the current image (video frame) quality is poor or the target object is blurry, that is, the reliability of the recognition result is low.

[0070] Based on the above embodiments, this embodiment describes step S120. Specifically, the method for identifying the target object in the video data and obtaining the perception state of the target object in step S120 may include the following steps S121 to S122.

[0071] Step S121: Identify the target object in the video data and obtain the target identification result.

[0072] It should be noted that the method for identifying the target object can refer to one or more related methods in this technical field, which will not be elaborated here.

[0073] The target recognition result refers to the relevant information of the identified target object. This relevant information may include, but is not limited to, the target object's bounding box and confidence score.

[0074] Step S122: Determine the perception state of the target object based on the target recognition result and the image features of the video data.

[0075] Image features may include, but are not limited to, pixel features, structural features, and object features in video frames.

[0076] By comprehensively analyzing the target recognition results and image features, the associated information of the target object in the image (such as environmental information, behavioral information, appearance information, etc. related to the target object) can be determined. Based on this associated information, the perceptual state of the target object can then be further determined. Specific methods can refer to relevant technologies in this field, such as neural network prediction, prior knowledge judgment, and image perspective relationship analysis, which will not be elaborated here.

[0077] For example, in a traffic detection scenario, when the structural feature of the target object being surrounded by the window frame of a motor vehicle is identified, or when the prior knowledge of the driver's cab pose is used to determine that a driver exists in the cab, it can be determined that the target object is located inside the motor vehicle.

[0078] Therefore, the auditory perception state of the target object can be determined by analyzing the state of the vehicle's windows (open or closed).

[0079] Under normal circumstances, people inside a car are more likely to receive sound information when the windows are open (equivalent to unrestricted auditory perception); while people inside a car are less likely to receive sound information when the windows are closed (equivalent to restricted auditory perception).

[0080] The methods for determining the status of vehicle windows can include, but are not limited to, image recognition methods relevant to this technical field, which will not be elaborated here. When determining the status of vehicle windows, the determination can be based on captured video frames, or by extracting a region of interest (such as the vehicle body area or the window area) from the video frames before making a determination.

[0081] Optionally, after extracting the region of interest from the video frame and obtaining the image of interest, image interpolation processing (such as bilinear interpolation, bicubic interpolation, etc.) can be performed on the image of interest to ensure the clarity and image details of the image of interest as much as possible, and avoid loss of resolution due to image cropping, which would affect the accuracy of subsequent processing.

[0082] In addition, the auditory perception state of the target object can be determined by analyzing the vehicle's speed (visual speed measurement can be achieved by analyzing the displacement of consecutive frames).

[0083] Under normal circumstances, if a vehicle is traveling at high speed (high wind noise and tire noise), the occupants will have difficulty receiving sound information (equivalent to limited auditory perception); if the vehicle is traveling at low speed (low wind noise and tire noise), the occupants will have easier access to sound information (equivalent to unrestricted auditory perception).

[0084] Alternatively, one can combine the window status and vehicle speed in their judgment. For example, if the window is determined to be open, the vehicle speed can be analyzed to determine whether the driver's hearing perception is impaired; this will not be elaborated upon here.

[0085] It should be noted that when there are multiple factors that may affect the perception state of the target object in the current application scenario, the influence of these factors can be quantified numerically, and then the degree of influence of the target can be determined by summation (or weighted summation), and then the perception state can be determined based on the degree of influence of the target.

[0086] In addition to being simply divided into "perceptual limitation" and "perceptual unlimitation", the perception state can also be divided into more types as needed according to the actual application scenario, such as classification by level or by value, which will not be elaborated here.

[0087] For example, referring to the previous examples, when a vehicle's window is open and it is traveling at low speed, the driver's perception is determined to be "unrestricted." When a vehicle's window is closed and it is traveling at low speed, the driver's perception is determined to be "slightly restricted." When a vehicle's window is open and it is traveling at high speed, the driver's perception is determined to be "moderately restricted." When a vehicle's window is closed and it is traveling at high speed, the driver's perception is determined to be "severely restricted."

[0088] In another example, in a street detection scenario, if a pedestrian is identified as wearing headphones, it can be determined that the pedestrian's auditory perception is limited. Similarly, if a pedestrian is identified as wearing tinted glasses, it can be determined that the pedestrian's visual perception is limited. Furthermore, if a pedestrian is identified as holding an umbrella, it can be determined that the pedestrian's visual perception is limited. Additionally, visual perception can also be determined by analyzing whether the pedestrian is holding a guide cane and / or leading a guide dog.

[0089] As another example, in street detection scenarios, street types in video data can also be analyzed. Street types can include normal pedestrian paths and tactile paving. If the same target object is detected walking on tactile paving in multiple consecutive frames, it can be determined that the target object's visual perception is impaired.

[0090] Optionally, to improve the reliability of the method in this application, the judgment of information such as the type and state of the target object should be based on the analysis results of multiple consecutive frames (e.g., 3-5 frames). In consecutive frames, if the identification results of the target object's type and state are consistently confirmed, it can be considered a valid judgment; otherwise, it is cached in a pending queue for continuous observation. If the prompting strategy has multiple levels, a degradation strategy will be triggered when prompting such unreliable identification results.

[0091] Based on the above embodiments, this embodiment describes step S130. The perception state includes a visual perception state and / or an auditory perception state, and the preset prompting strategy includes a visual prompting strategy and an auditory prompting strategy. Specifically, the method for determining the target prompting strategy from the preset prompting strategies based on the perception state in step S130 may include the following steps S131 to S132.

[0092] In step S131, in response to the visual perception state being restricted and the auditory perception state being unrestricted, the auditory cueing strategy is determined as the target cueing strategy.

[0093] Referring to the foregoing embodiments, if the visual perception state of the target object is detected as restricted and the auditory perception state is detected as unrestricted, it can be determined that the target object is more sensitive to auditory cues.

[0094] Therefore, in this type of perceptual state, the auditory cueing strategy is identified as the target cueing strategy.

[0095] In step S132, in response to the visual perception state being unrestricted and the auditory perception state being restricted, the visual cueing strategy is determined as the target cueing strategy.

[0096] Similarly, if the visual perception state of the target object is detected as unrestricted and the auditory perception state is restricted, it can be determined that the target object is more sensitive to visual cues.

[0097] Therefore, in this type of perceptual state, the visual cueing strategy is determined as the target cueing strategy.

[0098] In addition, if neither visual perception nor auditory perception is restricted, the audio-visual cue strategy (simultaneously emitting light and sound cues) can be determined as the target cue strategy; or either the auditory cue strategy or the visual cue strategy can be chosen as the target cue strategy.

[0099] If both visual and auditory perception are limited, then an audio-visual cue strategy can be determined as the target cue strategy to increase the probability that the target object receives the cue information, or to enable other targets in the current scene to receive the cue information and notice the target object's behavior.

[0100] The audiovisual cue strategy where neither visual nor auditory perception is restricted differs from the audiovisual cue strategy where both visual and auditory perception are restricted. Specifically, the difference may lie in the intensity of the cue, which will not be elaborated upon here.

[0101] Based on the above embodiments, this embodiment describes the method after step S140. Specifically, after the method of providing hints to the target object according to the target hint strategy in step S140, the method may further include steps S150 to S180.

[0102] It should be noted that the target prompting strategy implemented in the aforementioned embodiments is determined based on the perceptual state of the target object. In reality, the analysis of the perceptual state may not be 100% accurate, and the target object may still not respond accordingly after receiving the prompting information.

[0103] Therefore, in order to ensure the effectiveness (or deterrent effect) of the prompts to the target, after processing the prompts according to the target prompt strategy, it is possible to continue to analyze the target's behavioral intentions to decide whether to adjust the prompt strategy.

[0104] Step S150: Obtain the behavioral intent of the target object.

[0105] In this context, behavioral intent can also be equated to behavioral information and behavioral characteristics in the aforementioned embodiments.

[0106] Behavioral intentions can be varied, such as approaching, leaving, or lingering. Furthermore, each type of behavioral intention can be further subdivided into different degrees of behavior, such as approaching at high speed, approaching at low speed, leaving at high speed, or leaving at low speed, which will not be elaborated upon here.

[0107] Methods for obtaining behavioral intent can refer to methods related to behavior analysis and intent analysis in this technical field, which will not be elaborated here.

[0108] For example, this application may employ distance analysis technology based on visual features. This distance analysis technology can be used to analyze the distance between a target object and an image acquisition device, the distance between a target image and a specified area, the distance between a target object and a specified object, etc., and is not limited thereto. In the embodiments of this application, the distance between the target object and the image acquisition device is mainly used as an example for illustration.

[0109] For example, distance can be determined using principles such as monocular vision ranging or by analyzing changes in the pixel proportion of a target object in an image. When a target object is in consecutive frames, if the size of its bounding box (position box) in the image significantly increases, it is determined to be near the image acquisition device (close); if the size of its bounding box decreases or remains unchanged, it is determined to be far away or stationary. One or more distance thresholds (e.g., far > 10 meters, medium 5-10 meters, near < 5 meters) can be set to provide tiered alerts for the target object at different distances.

[0110] Another example is that by analyzing the trajectory information of the target object, it can be determined whether its behavioral intention is to pass by, linger, or move directly toward the designated area.

[0111] Step S160: Determine whether the target cue strategy is effective based on the behavioral intent.

[0112] Among them, one or more methods for judging whether the target prompting strategy is effective based on behavioral intent can be set as needed according to the actual application scenario, and no limitation is made here.

[0113] For example, the effectiveness of a target prompting strategy can be determined based on the direction of movement of the target object (or its distance from a specified area). For instance, setting the target object to stop moving can indicate that the target prompting strategy is effective; setting the target object to move away from a specified area can indicate that the target prompting strategy is effective; and setting the target object to move closer to a specified area can indicate that the target prompting strategy is ineffective.

[0114] Another example is that the effectiveness of a target prompting strategy can be determined based on the target object's movement speed. For instance, slowing down the target object's movement speed can indicate that the target prompting strategy is effective; or increasing or maintaining the target object's movement speed can indicate that the target prompting strategy is ineffective.

[0115] As another example, the direction of movement and the speed of movement can be combined to determine whether the target prompting strategy is effective, which will not be elaborated here.

[0116] Step S170: If the target prompting strategy is invalid, the target prompting strategy is upgraded to obtain an upgraded prompting strategy.

[0117] Based on the foregoing embodiments, this application can classify different levels of prompting strategies according to different levels of behavior. The method of classifying the levels can be binary or multi-level, and is not limited here. For ease of explanation, this application mainly uses low-level prompting strategies and high-level prompting strategies in the examples.

[0118] Among them, low-level prompt strategies have lower prompt intensity, while high-level prompt strategies have higher prompt intensity. The difference in prompt intensity can mainly be reflected in the different sound intensities and light intensities produced by the prompting devices.

[0119] Understandably, low-level alert strategies typically use lower-decibel sound alerts and lower-brightness light alerts, while high-level alert strategies use higher-decibel sound alerts and higher-brightness light alerts.

[0120] For example, a low-level prompting strategy can be used when the target object approaches at a low speed; a high-level prompting strategy can be used when the target object approaches at a high speed.

[0121] However, if, when the target object is approaching at a low speed, a low-level cueing strategy is applied, and the detected behavioral intent of the target object indicates that it continues to approach or accelerates, then the low-level cueing strategy can be considered ineffective. Therefore, the low-level cueing strategy can be escalated to obtain an escalated cueing strategy.

[0122] The upgrade processing can be to increase the decibel level of the sound cues and / or the brightness of the light cues in the low-level cues strategy; or it can be to replace the low-level cues strategy with a high-level cues strategy, which is not limited here.

[0123] Step S180: Provide prompts to the target object according to the upgrade prompt strategy.

[0124] Based on the foregoing steps, once the upgrade prompt strategy is obtained, the method described in the foregoing embodiments can be used to provide prompts to the target object according to the upgrade prompt strategy.

[0125] Based on the above embodiments, this embodiment describes step S130. Specifically, the method for determining the target prompting strategy from the preset prompting strategies according to the perception state in step S130 may include the following steps S133 to S134.

[0126] Step S133: Obtain the behavioral intent of the target object.

[0127] The method for obtaining behavioral intent can be referred to the description of the foregoing embodiments, and will not be repeated here.

[0128] Step S134: Determine the target prompting strategy from the preset prompting strategies based on the target object's perception state and behavioral intention.

[0129] In conjunction with the foregoing embodiments, this application can not only determine the target prompting strategy from the preset prompting strategy based on the target object's perception state, but also comprehensively analyze the target object's perception state and the target object's behavioral intention, and then determine the target prompting strategy from the preset prompting strategy.

[0130] For example, perceptual state can be used to determine the cueing method (such as visual and / or auditory cues), while behavioral intention can be used to determine the cue intensity (such as low-brightness or high-brightness cue, low-decibel or high-decibel cue, etc.). Therefore, a cueing strategy (targeted cueing strategy) that is reasonable in method and appropriate in intensity can be determined based on perceptual state and behavioral intention.

[0131] Based on the above embodiments, this embodiment describes step S140. Specifically, the method for providing hints to the target object according to the target hint strategy in step S140 may include the following steps S141 to S143.

[0132] In conjunction with the foregoing embodiments, when identifying target objects in video data, the confidence level of the target object (identification confidence level) can also be obtained.

[0133] Confidence level reflects the reliability of identifying the target object. If the confidence level is low, it may indicate that the video data is of poor quality or the target object is blurry. In other words, it can also reflect that the reliability of subsequent processing will be low to a certain extent.

[0134] Therefore, in this application, the selected target prompting strategy can be adjusted according to the confidence level of the target object recognition; or when the recognition confidence level is less than the confidence level threshold, the differentiated strategy is not considered, and a preset standard prompting strategy (such as keeping both sound and light prompts at a medium prompting intensity) is selected to ensure basic prompting capability when perception is unreliable.

[0135] Step S141: Obtain the recognition confidence of the target object in the video data.

[0136] The method for obtaining the identification confidence level can be referred to the description of the foregoing embodiments, and will not be repeated here.

[0137] Step S142: In response to the recognition confidence level being less than the confidence level threshold, the target prompting strategy is downgraded to obtain a downgraded prompting strategy.

[0138] For example, if a target object is detected approaching at high speed, a high-level prompting strategy is generally selected as the target prompting strategy. However, if the confidence level of the detected target object is less than the confidence threshold, the high-level prompting strategy needs to be downgraded.

[0139] The method for downgrading can be illustrated by referring to the description of upgrading in the foregoing embodiments. Downgrading may include, but is not limited to: adjusting the intensity of sound cues and / or light cues in the target cue strategy (high-level cue strategy); or selecting a low-level cue strategy as the downgrading cue strategy, which will not be elaborated here.

[0140] For example, when the identification confidence level is less than a confidence threshold, the intensity of the downgrade process can be determined based on the confidence gap between the identification confidence level and the confidence threshold.

[0141] For example, using the confidence gap as a downgrade weight, the initial light intensity and / or initial sound intensity of the original target cue strategy are multiplied by the downgrade weight to obtain the light intensity and / or sound intensity that need to be reduced. Then, by subtracting the light intensity that needs to be reduced from the initial light intensity, and / or by subtracting the sound intensity that needs to be reduced from the initial sound intensity, the downgraded light intensity and / or downgraded sound intensity in the downgraded cue strategy can be obtained.

[0142] In addition, during the downgrade process, besides the initial light intensity and / or initial sound intensity of the target cue strategy, other cue dimensions such as flash frequency and light color saturation of the target cue strategy can also be downgraded in the same way. This will not be elaborated here.

[0143] For example, multiple confidence gap intervals can be set, and different confidence gap intervals can correspond to different degradation intensities (e.g., a reduction of 25%, 50%, 75%, etc.). The corresponding degradation intensity is selected according to the confidence gap interval in which the confidence gap is located, and the initial light intensity and / or initial sound intensity and other cue dimensions of the original target cue strategy are downgraded.

[0144] For example, different confidence gap ranges can correspond to different low-level cue strategies. The corresponding low-level cue strategy can be selected as the downgrade cue strategy and implemented directly based on the confidence gap range it falls within; this will not be elaborated upon here.

[0145] Step S143: Provide prompts to the target object according to the downgrade prompting strategy.

[0146] Therefore, once the downgrade suggestion strategy is obtained, suggestions can be made to the target object according to the downgrade suggestion strategy.

[0147] Based on the above embodiments, this embodiment describes the method after step S110.

[0148] It is understood that the foregoing embodiments provide a method for dynamically adjusting a target suggestion strategy. This illustrates that after determining the target suggestion strategy, and when the confidence level is less than a confidence threshold, the target suggestion strategy can be downgraded based on the confidence level of the target object.

[0149] This embodiment provides a method for determining a logically fixed prompting strategy, which mainly describes the method of first comparing confidence level and confidence level threshold, and then deciding whether to directly select a fixed prompting strategy or select the corresponding target prompting strategy from preset prompting strategies based on the confidence level and confidence level threshold.

[0150] Specifically, after the method for acquiring video data of the target event in response to the detection of the target event trigger in step S110, it may further include at least the following steps S111 to S114.

[0151] Step S111: Identify the target object in the video data and obtain the identification confidence level of the target object.

[0152] Step S112: In response to the recognition confidence level being less than the confidence level threshold, a fixed prompting strategy is obtained to provide prompting to the target object.

[0153] Step S113: In response to the recognition confidence level being greater than or equal to the confidence threshold, the perception state of the target object is obtained.

[0154] Step S114: Determine the target prompting strategy from the preset prompting strategies based on the perception state and perform prompting processing on the target object.

[0155] Among them, the fixed prompt strategy is essentially preset, and the preset prompt strategy may or may not include the fixed prompt strategy (equivalent to the fixed prompt strategy being one of the preset prompt strategies, or not being one of the preset prompt strategies).

[0156] In this embodiment, the confidence level of the target object can be obtained before acquiring its perception state. The fixed prompting strategy is a prompting strategy that can be directly selected when the confidence level is less than the confidence threshold. This allows us to ignore complex differentiated strategies and directly enable a standard safety response mode (e.g., simultaneously activating a medium-level prompting strategy, including medium-intensity light and sound) to ensure basic prompting capabilities when perception is unreliable.

[0157] When the confidence level is greater than or equal to the confidence level threshold, the same principle applies as described in the previous embodiments (such as the example description of steps S120 to S140) to obtain the perception state of the target object; determine the target prompting strategy from the preset prompting strategy based on the perception state; and perform prompting processing on the target object according to the target prompting strategy, which will not be elaborated here.

[0158] Therefore, when the confidence level is less than the confidence threshold, a fixed prompting strategy is directly used. Compared with the method in the previous embodiment, the intermediate process of downgrading based on the target prompting strategy can be ignored, resulting in higher prompting efficiency and more stable prompting effect.

[0159] Based on the above embodiments, this embodiment describes an example scenario of determining a target prompting strategy from preset prompting strategies.

[0160] Taking an application scenario where the confidence level is greater than or equal to the confidence threshold and the target object is a single target as an example, the adaptive policy decision engine can find the corresponding target prompting policy from the preset prompting policies contained in the preset policy library shown in Table 1 below.

[0161]

[0162] Based on the above embodiments, this embodiment describes the method after step S110. Specifically, after the method of acquiring video data of the target event in response to the detection of the target event trigger in step S110, the method may further include the following steps S210 to S220.

[0163] In conjunction with the foregoing embodiments, in the actual application scenarios of the method of this application, multiple target events may occur, or multiple objects may be detected (referred to as candidate objects here for ease of explanation). Therefore, when multiple candidate objects appear in video data, one object can be selected from the candidate objects as the target object and the corresponding target prompting strategy can be determined.

[0164] Step S210: Identify candidate objects in the video data and obtain the number and type of candidate objects.

[0165] For example, the method for identifying candidate objects can be similarly described with reference to the foregoing embodiments, and will not be repeated here.

[0166] After identifying the candidate objects, we can also obtain relevant information such as the type, perceptual state, behavioral intention, and confidence level of each candidate object.

[0167] Step S220: In response to the number of objects being multiple, the target object is determined from each candidate object based on the object type.

[0168] There can be one or more methods to determine the target object from the candidate objects, and no limit is set here.

[0169] For example, the highest priority object can be selected from different object types according to their pre-set priorities; or the object corresponding to the target object type that needs to be focused on can be selected from different object types; or the object that needs to be focused on can be selected according to the behavioral intentions of different objects, etc.

[0170] For example, priorities can be pre-set for different object types. Then, based on the detected object types, priorities can be sorted, and the candidate object with the highest priority after sorting can be identified as the target object.

[0171] Another example is that if there are multiple candidate objects of the same type, their priorities can be pre-set according to their behavioral intentions. For example, high-speed approach > low-speed approach > low-speed departure > high-speed departure. Thus, the priority can be sorted according to the detected behavioral intentions of each object, and the candidate object with the highest priority after sorting can be determined as the target object.

[0172] If there is only one candidate object, it can be directly used as the target object and the method provided in the previous embodiments can be executed in the same way.

[0173] Based on the above embodiments, this embodiment describes step S220. Specifically, the method for determining the target object from each candidate object according to the object type in step S220 may include the following steps S221 to S223.

[0174] Step S221: Obtain the behavioral intent of each candidate object.

[0175] The method for obtaining behavioral intent can be referred to the description of the foregoing embodiments, and will not be repeated here.

[0176] Step S222: Determine the object priority corresponding to each candidate object based on the object type and behavioral intent.

[0177] The method for determining the target object from candidate objects based on priority in the foregoing embodiments can be exemplified and will not be elaborated here.

[0178] Building upon the aforementioned embodiments, this embodiment combines the object type and behavioral intent judgment methods, and then prioritizes them, which will not be elaborated upon here. For example, "a person approaching quickly" > "a car approaching slowly" > "a pet lingering," etc.

[0179] Step S223: Determine the target object from the candidate objects based on the object priority.

[0180] After obtaining the object priority of each candidate object, the target object can be determined from the candidate objects according to the object priority.

[0181] Furthermore, the policy arbitrator can also select the prompting policy corresponding to the highest priority object (target object) as the target prompting policy to be executed.

[0182] Optionally, in certain application scenarios, to achieve more effective collaborative cues for multiple targets, the system can perform cue strategy fusion. For example, the hardware resources for visual cues can prioritize serving the highest priority target and strictly follow its corresponding cue strategy (such as flashing red and blue lights).

[0183] The hardware resources for auditory cues can be intelligently allocated. For example, the currently broadcast voice content can take into account multiple targets (such as prompting the highest priority target first, then the next highest priority target, and so on, in descending order of priority); or different cues can be assigned to different targets, etc. There are no restrictions here.

[0184] For secondary priority targets, if their alerting strategy does not conflict with the primary alerting strategy for the highest priority target, some of their alerting content can be integrated into the primary strategy. For example, while maintaining the light alerting mode for the highest priority target, the flashing frequency can be fine-tuned to indicate the presence of secondary priority targets.

[0185] This mechanism ensures that, with limited hardware resources, the system can respond decisively to the highest priority targets while taking into account the overall information coverage of the scene as much as possible.

[0186] Based on the above embodiments, this embodiment describes the self-learning module in this application.

[0187] For example, the self-learning module records key data for each target event, forming a log. This includes: the target object type, the target object's perception state, the corresponding prompting strategy, the strategy parameters of the prompting strategy (such as volume, flash frequency, flash radius, etc.), and the object's response time (i.e., the time the prompting strategy takes effect on the target object, such as the time from the start of prompting processing to when the target object leaves, or the time from the start of prompting processing to when the target object slows down).

[0188] The self-learning module can periodically (e.g., weekly) or after a certain number of target events have been accumulated, run a lightweight online learning algorithm (such as the Thompson sampling algorithm, which makes decisions by sampling from a probability distribution, ensuring the utilization of known good options while maintaining the exploration of potential better options, etc., without limiting the optimization strategy parameters here).

[0189] This dynamic programming algorithm analyzes historical logs and fine-tunes strategy parameters to optimize the "object response time" for specific application scenarios (such as "passengers in the vehicle - flashing red and blue lights").

[0190] For example, if the current red and blue light flashing frequency (e.g., 1Hz) is found to be ineffective, the system may try a slightly higher frequency (e.g., 1.5Hz or 2Hz) in the next similar scenario and observe the effect, thereby achieving slow and safe parameter adaptation rather than directly changing the decision logic.

[0191] It should be further noted that the execution entity of the camera-based event notification method can be a camera-based event notification device. For example, the camera-based event notification method can be executed by a terminal device, a server, or other processing devices. The terminal device can be a user equipment (UE), computer, mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. In some possible implementations, the camera-based event notification method can be implemented by a processor calling computer-readable instructions stored in memory.

[0192] Figure 2 is a block diagram illustrating a camera-based event notification device according to an exemplary embodiment of this application. As shown in Figure 2, the exemplary camera-based event notification device 200 includes: an event detection module 210, an object recognition module 220, a strategy determination module 230, and a notification module 240. Specifically, the event detection module 210 is used to acquire video data of the target event in response to the detection of a target event.

[0193] The object recognition module 220 is used to identify target objects in video data and obtain the perception state of the target objects.

[0194] The strategy determination module 230 is used to determine the target prompting strategy from the preset prompting strategies based on the perception state.

[0195] The prompting module 240 is used to provide prompts to the target object according to the target prompting strategy.

[0196] In this exemplary camera-based event notification device, when a target event is detected, video data collected by the camera in response to the target event can be acquired. By identifying the target object in the video data, the perceptual state of the target object can be obtained. These perceptual states may include visual and / or auditory perceptual states. Analyzing the perceptual state of the target object can determine its perception mode that more easily receives external information (notification information). Therefore, a target notification strategy can be determined from preset notification strategies based on the perceptual state. Then, the target notification strategy is used to provide notification to the target object. This allows for the selection of a matching target notification strategy based on the target object's perception mode that it easily receives external information, rather than using a uniform notification strategy regardless of the type of target event or the type of target object in traditional methods. The method of this application selects an appropriate notification strategy according to the actual application scenario, enabling the issuance of reasonable and effective event notifications.

[0197] It should be noted that the apparatus and method provided in the above embodiments belong to the same concept, and the specific ways in which each module and unit performs operations have been described in detail in the method embodiments, and will not be repeated here. In practical applications, the apparatus provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above, and this is not a limitation.

[0198] The functions of each module can be found in the implementation example of the camera-based event prompting method, which will not be repeated here.

[0199] Please refer to Figure 3, which is a schematic diagram of the structure of an embodiment of the electronic device of this application. The electronic device 100 includes a memory 101 and a processor 102. The processor 102 is used to execute program instructions stored in the memory 101 to implement the steps in any of the above-described embodiments of the camera-based event prompting method. In a specific implementation scenario, the electronic device 100 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 100 may also include mobile devices such as laptops and tablets, which are not limited here.

[0200] Specifically, processor 102 controls itself and memory 101 to implement the steps in any of the above-described camera-based event notification method embodiments. Processor 102 may also be referred to as a CPU (Central Processing Unit). Processor 102 may be an integrated circuit chip with signal processing capabilities. Processor 102 may also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor. Furthermore, processor 102 may be implemented using integrated circuit chips.

[0201] In this exemplary electronic device, when a target event is detected, video data collected by a camera in response to the target event can be acquired. By identifying the target object in the video data, the perceptual state of the target object can be obtained. These perceptual states may include visual and / or auditory perceptual states. Analyzing the perceptual state of the target object can determine its perception mode that more easily receives external information (cue information). Therefore, a target cue strategy can be determined from preset cue strategies based on the perceptual state. Then, the target cue strategy is used to provide cue processing to the target object. This allows for the selection of a matching target cue strategy based on the target object's perception mode that easily receives external information, rather than using a uniform cue strategy regardless of the type of target event or the type of target object in traditional methods. The method of this application selects an appropriate cue strategy according to the actual application scenario, enabling the issuance of reasonable and effective event cueing.

[0202] Please refer to Figure 4, which is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. The computer-readable storage medium 110 stores program instructions 111 that can be executed by a processor. The program instructions 111 are used to implement the steps in any of the above-described embodiments of the camera-based event notification method.

[0203] In this exemplary storage medium, by running the program instructions within the storage medium, video data captured by the camera in response to the target event can be acquired when the target event is detected. By identifying the target object in the video data, the perceptual state of the target object can be obtained. These perceptual states may include visual and / or auditory perceptual states. Analyzing the perceptual state of the target object can determine its perception mode that more easily receives external information (cue information). Therefore, a target cue strategy can be determined from preset cue strategies based on the perceptual state. Then, the target cue strategy is used to cue the target object. This allows for the selection of a matching target cue strategy based on the target object's perception mode that easily receives external information, rather than using a uniform cue strategy regardless of the type of target event or the type of target object in traditional methods. The method of this application selects an appropriate cue strategy according to the actual application scenario, enabling the issuance of reasonable and effective event cuees.

[0204] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0205] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0206] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0207] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A camera-based event notification method, characterized in that, The method includes: in response to detecting a target event trigger, acquiring video data of the target event; identifying a target object in the video data and obtaining the perception state of the target object; determining a target prompting strategy from a preset prompting strategy based on the perception state; and performing prompting processing on the target object according to the target prompting strategy.

2. The method according to claim 1, characterized in that, The step of identifying the target object in the video data and obtaining the perception state of the target object includes: identifying the target object in the video data and obtaining a target identification result; and determining the perception state of the target object based on the target identification result and the image features of the video data.

3. The method according to claim 1, characterized in that, The perception state includes a visual perception state and / or an auditory perception state, the preset prompting strategy includes a visual prompting strategy and an auditory prompting strategy, and the step of determining a target prompting strategy from the preset prompting strategies based on the perception state includes: in response to the visual perception state being restricted and the auditory perception state being unrestricted, determining the auditory prompting strategy as the target prompting strategy; and in response to the visual perception state being unrestricted and the auditory perception state being restricted, determining the visual prompting strategy as the target prompting strategy.

4. The method according to claim 1, characterized in that, After the target object is prompted according to the target prompting strategy, the method further includes: obtaining the behavioral intent of the target object; determining whether the target prompting strategy is effective based on the behavioral intent; if the target prompting strategy is ineffective, performing prompting escalation processing on the target prompting strategy to obtain an escalated prompting strategy; and performing prompting processing on the target object according to the escalated prompting strategy.

5. The method according to claim 1, characterized in that, The step of determining the target prompting strategy from the preset prompting strategies based on the perceived state includes: obtaining the behavioral intent of the target object; and determining the target prompting strategy from the preset prompting strategies based on the perceived state of the target object and the behavioral intent.

6. The method according to claim 1, characterized in that, After acquiring video data of the target event in response to the detection of the target event, the method further includes: identifying a target object in the video data and obtaining the identification confidence level of the target object; and in response to the identification confidence level being less than a confidence level threshold, acquiring a fixed prompting strategy to provide prompting to the target object.

7. The method according to claim 1, characterized in that, After acquiring video data of the target event in response to the detection of the target event, the method further includes: identifying candidate objects in the video data, obtaining the number and type of the candidate objects; and determining the target object from each candidate object according to the object type if the number of objects is multiple.

8. The method according to claim 7, characterized in that, The step of determining the target object from each candidate object based on the object type includes: obtaining the behavioral intent of each candidate object; determining the object priority corresponding to each candidate object based on the object type and the behavioral intent; and determining the target object from each candidate object based on the object priority.

9. An electronic device, characterized in that, The method includes a memory and a processor, the processor being configured to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the method described in any one of claims 1 to 8.