A safety warning method and related apparatus
Patent Information
- Application Number
- CN202510019005.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2026-07-07
Smart Images

Figure CN122347822A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a security early warning method and related apparatus. Background Technology
[0002] Action recognition of objects is crucial for determining their state and whether they face safety risks. For example, an object can be someone requiring care, such as an infant, the elderly, or a person with disabilities.
[0003] Because these types of objects cannot independently assess potential safety risks during their activities, they are prone to security problems. Recognizing their movements can help determine their situation, enabling timely warnings when security risks arise, thus better protecting their safety. In related technologies, information collection devices can be installed in locations where these objects frequently appear, allowing relevant personnel to view the data collected at any time to monitor the objects' situation.
[0004] However, this method, which relies on manual inspection, is prone to oversights and cannot detect security problems in a timely manner, which is not conducive to the security protection of the target. Summary of the Invention
[0005] To address the aforementioned technical issues, this application provides a security early warning method and related apparatus, which facilitates the timely detection of potential security problems, reduces oversights, and provides better security protection for target objects.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] On the one hand, embodiments of this application provide a security early warning method, the method comprising:
[0008] Obtain the video stream corresponding to the detection scene targeting the object;
[0009] If the first action is determined when the j-th image frame included in the video stream is detected, the target object included in the detection scene is obtained, and the speed parameter corresponding to the target object performing the first action is obtained. The speed parameter is used to indicate the movement speed of the target part from the first position to the second position. The target part is the part of the target object that performs the first action. The first position refers to the position of the target part in the detection scene at the acquisition time of the i-th image frame. The i-th image frame is the image frame included in the video stream used to determine the start of the first action. The second position refers to the position of the target part in the detection scene at the acquisition time of the j-th image frame. The acquisition time of the i-th image frame is earlier than the acquisition time of the j-th image frame.
[0010] Based on the speed parameter, the second position, and the motion trajectory from the first position to the second position, a safety risk prediction is performed on the target object to obtain a risk prediction result. The risk prediction result is used to indicate whether there is a safety risk to the target object when it continues to move on the predicted trajectory with the speed parameter. The predicted trajectory refers to the trajectory of the target part continuing to perform the first action from the second position.
[0011] If the risk prediction results indicate that the target object poses a security risk, a security warning will be issued.
[0012] In another aspect, embodiments of this application provide a safety early warning device, the device comprising an acquisition unit, a prediction unit, and an early warning unit:
[0013] The acquisition unit is used to acquire the video stream corresponding to the detection scene of the target object;
[0014] The acquisition unit is further configured to determine that a target object included in the detection scene has performed a first action when the j-th image frame included in the video stream is detected, and to acquire a speed parameter corresponding to the target object performing the first action. The speed parameter is used to indicate the movement speed of the target part from a first position to a second position. The target part is the part of the target object that performs the first action. The first position refers to the position of the target part in the detection scene at the acquisition time of the i-th image frame. The i-th image frame is an image frame included in the video stream used to determine the start of the first action. The second position refers to the position of the target part in the detection scene at the acquisition time of the j-th image frame. The acquisition time of the i-th image frame is earlier than the acquisition time of the j-th image frame.
[0015] The prediction unit is used to predict the safety risk of the target object based on the speed parameter, the second position, and the movement trajectory from the first position to the second position, and to obtain a risk prediction result. The risk prediction result is used to indicate whether there is a safety risk to the target object when it continues to move on the predicted trajectory with the speed parameter. The predicted trajectory refers to the trajectory of the target part continuing to perform the first action from the second position.
[0016] The early warning unit is used to issue a safety warning if the risk prediction result indicates that the target object has a safety risk.
[0017] On the other hand, embodiments of this application provide a computer device, the computer device including a processor and a memory:
[0018] The memory is used to store computer programs and to transfer the computer programs to the processor;
[0019] The processor is configured to execute the method described in any of the foregoing aspects according to instructions in the computer program.
[0020] On the other hand, embodiments of this application provide a computer-readable storage medium for storing a computer program, which, when run by a computer device, causes the computer device to perform the methods described in any of the foregoing aspects.
[0021] On the other hand, embodiments of this application provide a computer program product, including a computer program that, when run on a computer device, causes the computer device to perform the methods described in any of the foregoing aspects.
[0022] As can be seen from the above technical solution, for the detection scenario of the target object, the corresponding video stream is acquired. If the j-th image frame included in the video stream is detected, it is determined that the target object has performed a first action. Then, the velocity parameter corresponding to the target object performing the first action is acquired. The velocity parameter is used to indicate the target part of the target object performing the first action and the speed of movement from the first position to the second position. Typically, after performing a certain action, the target object may continue to perform that action, which may lead to safety risks. However, since the target object itself cannot judge the potential safety risks, a safety risk prediction is made for the target object based on the velocity parameter, the second position, and the movement trajectory from the first position to the second position. Because the velocity parameter reflects the movement of the target object performing the first action, it can indicate whether the target object will face safety risks if it continues to move at this speed. Similarly, the movement trajectory reflects the trend of the first action, i.e., it can indicate which direction the target object might continue to move in, and therefore, it can also indicate whether the target object will face safety risks if it continues to move. Therefore, the predicted risk result can be used to indicate whether there is a safety risk for the target object when it continues to move along the predicted trajectory with the velocity parameter. Correspondingly, if the risk prediction results indicate that there is a safety risk to the target object, a safety warning can be issued to remind relevant personnel to take action and protect the target object.
[0023] As can be seen, this application combines the target object's own movement speed and the trend of its first action to determine whether the target object will face safety risks as it continues to move, and then issues an early warning when risks are detected, thereby improving the safety protection of the target object. Furthermore, the prediction is based on the current state of the target object's first action, making it a targeted prediction for that specific target object. Since different target objects may have different speed parameters and movement trajectories for the same first action, combining the current state of the target object's first action can more accurately predict the situations the target object may face, thus enabling more accurate safety risk prediction and early warning. In practical applications, video streams can be acquired by information acquisition devices; therefore, compared to manual review, automatic prediction and early warning based on video streams are more conducive to timely detection of potential safety issues, reducing oversights, and better protecting the target object. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a schematic diagram illustrating an application scenario of a security early warning method provided in an embodiment of this application;
[0026] Figure 2 A flowchart illustrating a security early warning method provided in this application embodiment;
[0027] Figure 3 This is a schematic diagram of an infant's skeletal structure provided in an embodiment of this application;
[0028] Figure 4 A schematic diagram of a dynamic identification process provided in an embodiment of this application;
[0029] Figure 5 A schematic diagram of the structure of a video stream acquisition module provided in an embodiment of this application;
[0030] Figure 6 This application provides a schematic diagram of the application architecture of a security early warning system.
[0031] Figure 7 A schematic diagram of an early warning center interface provided in an embodiment of this application;
[0032] Figure 8A schematic diagram of the logical architecture for a security early warning system provided in an embodiment of this application;
[0033] Figure 9 This application provides a schematic diagram of the architecture of a network service system.
[0034] Figure 10 A structural diagram of a safety early warning device provided in an embodiment of this application;
[0035] Figure 11 A structural diagram of a terminal provided in an embodiment of this application;
[0036] Figure 12 This is a structural diagram of a server provided in an embodiment of this application. Detailed Implementation
[0037] The embodiments of this application will now be described with reference to the accompanying drawings.
[0038] Typically, information collection devices are installed in scenarios where objects need to be monitored (e.g., to determine whether an object faces security risks). The data collected by these devices reflects the specific situation in the current scenario, thus reflecting the object's situation in that scenario and assisting in determining whether the object is safe.
[0039] In related technologies, relevant personnel (such as caregivers of the subject) can view the data collected by the information collection device at any time to check the subject's condition and determine whether the subject faces any security risks. However, this method relies on manual review, which is prone to oversights and cannot detect security problems in a timely manner, thus hindering the subject's safety protection.
[0040] To address this, this application provides a safety warning method and related apparatus. For a target object detection scenario, a corresponding video stream is acquired. This video stream reflects the actual situation in the detection scenario and can therefore be used to detect the target object. The target object refers to the object to be detected. Specifically, when a target object is detected to have performed a first action, the speed parameters corresponding to the first action are acquired, and a safety risk prediction is made based on the speed parameters and the corresponding motion trajectory. Because the speed parameters reflect the movement of the target object performing the first action, they indicate whether the target object will face safety risks if it continues to move at this speed. The motion trajectory reflects the trend of the first action, i.e., which direction it might continue to move in, and therefore also indicates whether the target object will face safety risks if it continues to move. Therefore, the predicted risk result can be used to indicate whether there is a safety risk to the target object if it continues to move along the predicted trajectory at the speed parameters. If a safety risk is determined to exist for the target object, a safety warning is issued to alert relevant personnel for action and to protect the target object.
[0041] In practical applications, video streams can be acquired using the aforementioned information collection devices. Therefore, compared to manual review, automatic prediction and early warning based on video streams are more effective in promptly identifying potential security issues. Eliminating reliance on real-time manual review reduces oversights caused by human error, thus enabling better security protection of target objects.
[0042] Furthermore, this application makes predictions based on the current target object's first action (such as speed parameters and trajectory), thus it is a targeted prediction for that target object. Since different target objects may have different speed parameters and trajectories when performing the same first action, combining the current target object's first action can more accurately predict the situation that the target object may face, thereby enabling more accurate safety risk prediction and safety warning.
[0043] The security warning method provided in this application can be implemented using a computer device, which can be a terminal or a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Terminals include, but are not limited to, smartphones, computers, smart voice interaction devices, smart home appliances, and in-vehicle terminals. Terminals and servers can be directly or indirectly connected via wired or wireless communication, and this application does not impose any limitations on this connection.
[0044] This application's embodiments can be specifically applied to various detection scenarios targeting specific objects. The target object refers to the object to be detected, specifically a person requiring care. Typically, due to their unique characteristics, such individuals may be unable to independently assess potential safety risks, potentially injuring themselves during activities. Therefore, detection is often necessary to determine whether these individuals face safety risks. For example, these individuals may have limited cognitive abilities regarding their environment, making it impossible for them to judge whether their activities will result in self-harm. Alternatively, they may have limited control over their own activities (such as limb movements), making them prone to self-harm during activities. Specifically, care recipients may include infants, the elderly, and people with disabilities.
[0045] It is understandable that different target objects may frequently appear in different locations, thus the detection scenarios will also differ. For example, in places where target objects may be present, such as hospitals, healthcare facilities (e.g., pediatric and geriatric care facilities), bedrooms, and living rooms, there is a need for detection in order to better protect the target objects. Therefore, these locations can be considered as the aforementioned detection scenarios for target objects, and this application can be applied to provide better security protection for target objects in these locations. Specifically, information acquisition equipment can be installed in the detection scenario to collect video streams. These video streams can be used to detect the situation of the target objects in the detection scenario and predict whether there are any security risks associated with the target objects, thereby providing security protection for the target objects.
[0046] It should be noted that in the specific embodiments of this application, the execution of the security warning method may involve user information and other related data. When the above embodiments of this application are applied to specific products or technologies, it is necessary to obtain separate consent or separate permission from the user, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0047] Figure 1 This illustration shows an application scenario of the security early warning method provided in the embodiments of this application. Figure 1 In the scenario shown, server 100 is used as an example of the aforementioned computer device, and the target object is the aforementioned infant or toddler, for illustration:
[0048] For detection scenarios involving infants and young children, in order to better protect the safety of these children, the system can monitor their condition and provide timely warnings when safety risks are present. To this end, server 100 first acquires the video stream corresponding to the detection scenario for infants and young children. This video stream can be used to indicate the scene of the detection scenario, reflecting the actual situation within it, and thus can be used for subsequent detection and warnings. Typically, the video stream consists of image frames, each of which can be used to indicate the specific situation of the detection scenario at a particular moment.
[0049] In practical applications, once the video stream is acquired, subsequent detection can be performed based on it, such as detecting the image frames included in the video stream to determine the infant's condition. Specifically, if the infant is determined to have performed the first action when the j-th image frame in the video stream is detected, the server 100 can obtain the speed parameters corresponding to the infant's first action. It is understood that an action is a continuous process, typically requiring the detection of multiple image frames before determining whether an action has been performed. In this application, the infant's first action is determined to have been performed when the j-th image frame is detected; therefore, the j-th image frame can represent the situation when the first action is determined to have been performed.
[0050] The first action can refer to an action that may pose a danger to infants and young children. Specifically, the first action can be a dangerous action for infants and young children, such as reaching for something (which may result in knocking over and injuring themselves), crawling (which may result in falling off the bed), rolling over (which may result in falling off the bed), or bringing something close to their face (which may result in putting it in their mouth). For ease of understanding, in... Figure 1 In the example, taking the first action as holding something close to the face, the dashed box shows a sitting infant holding a toy close to his face.
[0051] Furthermore, the velocity parameter can be used to indicate the speed at which the target body part moves from the first position to the second position. The target body part is the part of the infant or toddler that performs the first movement, such as... Figure 1 The infant's right hand and right arm shown can be the target body part. Here, the first position refers to the location of the target body part in the detection scene at the acquisition time of the i-th image frame. The i-th image frame is the image frame included in the video stream used to determine the start of the first action, and the acquisition time of the i-th image frame is earlier than the acquisition time of the j-th image frame. That is, by detecting the image frames included in the video stream, it is identified that the infant began performing the first action at the i-th image frame and that the first action was performed at the j-th image frame. Correspondingly, the second position refers to the location of the target body part in the detection scene at the acquisition time of the j-th image frame.
[0052] Typically, infants and toddlers may continue performing a certain action, potentially leading to safety risks, as they cannot assess these risks themselves. Therefore, server 100 can predict these risks based on speed parameters, a second position, and the movement trajectory from the first position to the second position, thus obtaining a risk prediction result. Since the speed parameter reflects the infant's initial movement, it indicates whether continuing at that speed would pose a safety risk. Similarly, the movement trajectory reflects the trend of the initial action, indicating the likely direction of future movement, and thus also indicating potential safety risks. Therefore, the predicted risk result can be used to indicate whether there is a safety risk for the infant continuing to move along the predicted trajectory at the predicted speed.
[0053] Afterwards, follow-up actions can be taken based on the specific circumstances of the risk prediction results. For example, if the risk prediction results indicate that there is a safety risk to the infant, the server 100 can issue a safety warning, such as broadcasting a safety warning message or flashing a safety warning light, to remind relevant personnel (such as the person in charge of the infant's care) to take action and protect the infant's safety.
[0054] For example, targeting Figure 1 The dashed box indicates a predicted trajectory where the right hand continues to move closer to the face, especially towards the mouth, potentially leading to the toy being put in the mouth, which could be dangerous for infants and toddlers. Therefore, regarding... Figure 1 For example, the risk prediction results obtained can indicate that there are safety risks to infants and young children, and that safety warnings are needed to remind relevant personnel to pay attention and take timely action to avoid dangerous events as much as possible.
[0055] Figure 2 A flowchart illustrating a security warning method provided in this application embodiment, using a server as an example of the aforementioned computer device, shows the method comprising S201-S204:
[0056] S201: Obtain the video stream corresponding to the detection scene of the target object.
[0057] The target group refers to the individual being monitored, specifically those requiring care. Often, due to their unique characteristics, these individuals may be unable to independently assess potential safety risks, potentially leading to self-harm during activities. Therefore, monitoring is frequently necessary to determine if they face safety risks. For example, they may have limited cognitive abilities regarding their environment, making it difficult for them to judge whether activities within that environment will cause injury. Alternatively, they may have limited control over their own movements (such as limb movements), making them prone to self-injury during activities. Specific examples of individuals requiring care include infants, the elderly, and people with disabilities.
[0058] Furthermore, the detection scenario can refer to a scenario where there is a need to detect the situation of the target object, such as the aforementioned places where the target object may be, such as hospitals, health centers, bedrooms, living rooms, etc. Also, the video stream refers to a media file that can be processed as a stable and continuous stream during the transmission of video data (such as subsequent detection, recognition, and early warning). In this application, video data can refer to media files obtained by capturing video of the detection scenario, and the video stream can refer to media files transmitted in real time during the detection process. Specifically, it can be used to indicate the scene of the detection scenario, thereby reflecting the real situation in the detection scenario. Therefore, the video stream can be acquired for subsequent detection, etc.
[0059] In practical applications, the detection requirements for the target object detection scenario are performed in real time in order to promptly identify and address potential dangers to the target object. Therefore, the acquired video data is continuously transmitted as a video stream to facilitate real-time detection.
[0060] It should be noted that this application does not impose any limitations on the method of acquiring the video stream. For ease of understanding, the embodiments of this application provide the following examples:
[0061] Typically, video streams can be acquired using information acquisition devices installed in the detection scenario, such as video recording devices, specifically network cameras (IP cameras). It is understood that information acquisition devices are terminal devices. To improve the efficiency of real-time detection, a server is usually used as the aforementioned computer device to execute the implementation methods provided in this application. Accordingly, a connection (such as a network connection) can be established between the information acquisition device and the server to transmit the video stream acquired in real-time by the information acquisition device to the server for subsequent processing. Of course, in some scenarios, a terminal (such as the aforementioned computer) can also be used as a computer device to execute the implementation methods provided in this application, in which case the information acquisition device can establish a connection with the terminal. This can be flexibly configured according to the actual detection scenario and actual detection needs, and this application does not impose any limitations.
[0062] S202: If the j-th image frame included in the video stream is detected, it is determined that the target object included in the detection scene has performed the first action, and the speed parameter corresponding to the target object performing the first action is obtained.
[0063] In practical applications, video streams can include image frames. Each image frame can be used to indicate the acquisition time of that image frame and the specific situation corresponding to the detection scene. The acquisition time can indicate the moment when the image frame was acquired. Typically, by detecting the image frames included in the video stream, the situation of target objects in the detection scene can be identified. Since target objects cannot independently determine whether their actions will pose a danger to themselves, in this application, image frames can be detected to identify whether the target object has performed certain actions, and then predict whether the target object faces a safety risk.
[0064] Specifically, if the target object is determined to have performed a first action when the j-th image frame in the video stream is detected, the server can obtain the speed parameter corresponding to the target object's performance of the first action. Here, the first action can refer to an action that may pose a danger to the target object; that is, the first action can be a dangerous action targeting the target object.
[0065] Understandably, different targets may present different dangerous actions. To better understand this, the following example illustrates the first action:
[0066] As an example, taking infants and toddlers as the target audience, the following example is provided for the first action:
[0067] For example, when reaching for things, infants and toddlers cannot judge whether the object they are reaching for will hurt them. They might knock over things and injure themselves, or the object might be dangerous. Similarly, crawling and rolling over are dangerous actions. Since infants and toddlers who haven't learned to walk primarily play in their beds, these actions could cause them to fall off and face danger. Furthermore, infants often put things in their mouths. If the objects are unsafe, this can also pose a danger. For example, toys can obstruct breathing and cause injury if put in the mouth.
[0068] As yet another example, taking the target audience as elderly people, the following example is provided for the first action:
[0069] For example, when reaching for something, older adults often have weaker control over their limbs and may accidentally bump into something dangerous (like object B) while trying to reach for object A, thus injuring themselves. Similarly, when running quickly, due to weaker control and reaction time to danger, they may fall or collide with obstacles, causing injury.
[0070] As another example, taking the target group as people with disabilities, the following example is provided for the first action:
[0071] For example, the action of putting something into one's mouth. In practical applications, people with disabilities may not be able to judge whether the things they pick up are safe or whether they are edible, so this action may cause danger.
[0072] It is understandable that an action is a continuous process, typically involving the detection of multiple image frames before determining whether an action has been performed. In this application, when performing frame-by-frame detection on the image frames included in the video stream, it is determined that the first action begins to be performed at the i-th image frame, that the first action has been performed at the j-th image frame, and that the target part has moved from a first position to a second position. The target part can be the part of the target object where the first action is performed. Here, j and i are both positive integers.
[0073] For example, targeting Figure 1 In the example, the action of raising the right hand and bringing it close to the face cannot be determined when the infant performs this action during the frame-by-frame detection of the video stream. It is not until the j-th image frame is detected that the infant performs this action, and it was started in the i-th image frame.
[0074] Here, the first position can refer to the position of the target part in the detection scene at the acquisition time of the i-th image frame, and the second position can refer to the position of the target part in the detection scene at the acquisition time of the j-th image frame. Correspondingly, the obtained velocity parameter can be used to indicate the movement speed of the target part from the first position to the second position.
[0075] In summary, the acquisition time of the i-th image frame is earlier than that of the j-th image frame. The i-th image frame is included in the video stream and is used to determine the start of the first action, representing the situation at which the first action is determined to begin. Similarly, the j-th image frame is included in the video stream and is used to determine that the first action has been performed, representing the situation at which the first action has been performed. Accordingly, the j-th image frame can indicate the start time of subsequent processing (such as acquiring speed parameters and performing security risk prediction).
[0076] In practical applications, detection is performed in real time, and the video stream is continuously transmitted to the server. Therefore, when the j-th image frame is detected, if it is determined that the first action has been performed, subsequent processing begins. Conversely, if it is not determined that the target object has performed the first action, subsequent image frames can be detected.
[0077] It should also be noted that this application does not impose any limitations on the method for determining the j-th image frame as the starting point for subsequent processing, the method for obtaining speed parameters, or the method for determining the target location. For ease of understanding, this application will provide examples and illustrations below:
[0078] (i) Regarding the method of determining the j-th image frame as the starting point for subsequent processing, the embodiments of this application provide the following examples:
[0079] Subsequent processing is for safety risk prediction, enabling timely warnings in case of danger and protecting infants and young children. Understandably, different initial actions may pose varying degrees of danger to infants and young children, thus affecting the urgency of safety risk prediction and warning. Therefore, one possible implementation is to determine the j-th image frame based on the risk level of the initial action, allowing for more flexible detection.
[0080] In practice, the video stream is inspected frame by frame. It's possible that when a particular image frame (e.g., the third image frame) is detected, the infant's first action is identified. Correspondingly:
[0081] In one example, if the risk level of this first action is high, that is, it may pose a greater danger to infants and young children, then the corresponding urgency level is high. Therefore, the third image frame can be directly used as the j-th image frame mentioned above to indicate the start of subsequent processing.
[0082] In another example, if the risk level of this first action is low, meaning it poses little danger to the infant or toddler, and thus has a low urgency, then subsequent risk prediction can be temporarily suspended, and subsequent image frames can be detected instead. If, when a subsequent image frame (such as the 6th image frame) is detected, it is recognized that the infant or toddler is still performing this first action, then this image frame (i.e., the 6th image frame) can be used as the aforementioned j-th image frame to indicate the start of subsequent processing.
[0083] Correspondingly, the risk level of an action can also be used to indicate its priority or the sensitivity of detection. For example, for a high-risk first action, the warning priority is higher and the detection sensitivity is also higher; once detected, the warning can be issued immediately. For a low-risk first action, the warning priority is lower and the detection sensitivity is also lower; the warning can be issued only if it is detected continuously.
[0084] In practical applications, target groups of different ages (such as infants and toddlers of different ages) will have different abilities to cope with safety risks. Therefore, in one possible implementation, the characteristics of the target group (such as age characteristics) can be combined to implement flexible safety protection for target groups of different ages.
[0085] For ease of understanding, taking infants and toddlers as an example, in specific implementation, for infants under 1 year old, the image frame in which the first action is detected can be designated as the j-th image frame, indicating the start of subsequent processing to achieve timely risk prediction and early warning. For infants aged 1-3 years, for the first action with a lower risk level, the image frame in which the infant continues to perform the first action can be designated as the j-th image frame.
[0086] (II) Regarding the method of obtaining speed parameters, the embodiments of this application provide the following examples for illustration:
[0087] Understandably, the velocity parameter indicates the speed at which the target part moves from the first position to the second position, reflecting the speed of movement during this process. Accordingly, in one possible implementation, the speed can be determined based on the length of the path traveled and the time taken. In practice, firstly, the length of the trajectory from the first position to the second position can be determined, and the difference between the acquisition times of the i-th and j-th image frames can be defined as the motion duration. Then, the velocity parameter can be determined based on the trajectory length and the motion duration.
[0088] The motion trajectory indicates the path a target part takes to move from a first position to a second position within the detection scene. Correspondingly, the trajectory length indicates the actual path length taken during this movement. The motion duration indicates the time taken for the target part to move from the first position to the second position. Therefore, based on the trajectory length and duration, the speed of movement can be comprehensively evaluated, and a speed parameter that more intuitively reflects the speed of movement can be determined for subsequent use.
[0089] (III) Regarding the method for determining the target location, the embodiments of this application provide the following examples:
[0090] It is understandable that different actions require different body parts to be mobilized. Therefore, the target body part may be associated with the first action, and the corresponding target body part may be different for different first actions. For ease of understanding, the embodiments of this application provide the following examples:
[0091] For example, in Figure 1 In the example, the target area could be the infant's right hand and right arm. In other actions, such as crawling, infants may use their hands, elbows, feet, knees, or even their face (such as the chin area), so multiple areas can be identified as target areas.
[0092] For situations requiring the manipulation of multiple parts to perform a single action, since the relative positions of the various parts of the target object remain largely unchanged, one approach is to use one of the multiple parts as the target part. This simplifies the determination of its location, trajectory, etc., and improves detection efficiency. Typically, such actions involving multiple parts alter the overall state of the target object (e.g., changing its location, posture, etc.). Therefore, another approach is to use the entire target object as the moving object. Accordingly, the position of the target object at the start of the first action can be defined as the first position, and the position at the end of the first action can be defined as the second position.
[0093] S203: Based on the speed parameters, the second position, and the trajectory of the movement from the first position to the second position, perform a safety risk prediction on the target object and obtain the risk prediction result.
[0094] Because the velocity parameter reflects the movement of the target object as it performs the first action, it can indicate whether the target object will face safety risks if it continues to move at this velocity. Similarly, the movement trajectory reflects the trend of the first action, indicating the likely direction of continued movement, and thus also indicating whether the target object will face safety risks if it continues to move. Therefore, the server performs safety risk prediction based on the velocity parameter, the second position, and the movement trajectory. The resulting risk prediction can be used to indicate whether there is a safety risk for the target object if it continues to move along the predicted trajectory at the velocity parameter. The predicted trajectory refers to the path the target part takes as it continues to perform the first action from the second position; that is, the predicted trajectory can indicate the likely direction of continued movement.
[0095] For example, in Figure 1 In the example, the predicted trajectory can indicate that the right hand will continue to move towards the mouth. Furthermore, the risk prediction result can be used to indicate whether there is a safety risk to the infant or toddler as they continue to raise their right hand towards the mouth at the current speed (from the right leg position indicated in the i-th image frame to the chin position indicated in the j-th image frame). Figure 1 In the example, the child is holding a toy with sharp edges in their right hand. If the toy is accidentally put in their mouth, it could pose a serious danger to the infant. Therefore, the risk prediction results can be used to indicate that there is a safety risk to the infant in the current situation.
[0096] S204: If the risk prediction results indicate that there is a safety risk to the target object, a safety warning shall be issued.
[0097] After obtaining the risk prediction results, adaptive post-processing can be performed based on the specific circumstances of the predictions. For example, if the risk prediction results indicate that a target object poses a security risk, the server can issue a security alert to remind relevant personnel to take appropriate security protection measures for that target object.
[0098] Correspondingly, if the risk prediction result indicates that the target object does not pose a security risk, then detection can continue. In practice, continued detection can refer to continuing to detect each image frame in the subsequently acquired video stream to determine whether the target object has performed the first action, and thus predict whether the target object poses a security risk. Therefore, this is a method of real-time detection based on real-time transmitted video streams.
[0099] It should be noted that this application does not impose any limitations on the method of conducting security alerts. For ease of understanding, the embodiments of this application provide the following examples:
[0100] As an example, when a terminal is used as the aforementioned computer device to execute the embodiments provided in this application, the terminal can control the reminder module within the terminal to generate danger warning information when issuing a security alert, so as to inform relevant personnel (such as the person in charge of the care of the target object) that the target object may be in danger and needs to be checked in time. For example, the reminder module can be a sound module within the terminal, which can broadcast danger warning voice messages. Alternatively, the reminder module can be a warning light module within the terminal, which can emit danger warning signals by controlling the warning light module to flash, light up, etc.
[0101] As another example, when a server is used as the aforementioned computer device to execute the embodiments provided in this application, the server can generate a danger warning message and send it to the corresponding terminal (such as the smartphone of the person in charge of caring for the target). Alternatively, the server can send the danger warning message to an information acquisition device used to capture video streams, thereby conveying the danger warning message to relevant personnel through the information acquisition device. Alternatively, the information acquisition device can also be equipped with the aforementioned reminder module, allowing for more direct reminders to relevant personnel.
[0102] As can be seen from the above technical solution, for the detection scenario of the target object, the corresponding video stream is acquired. If the j-th image frame included in the video stream is detected, it is determined that the target object has performed a first action. Then, the velocity parameter corresponding to the target object performing the first action is acquired. The velocity parameter is used to indicate the target part of the target object performing the first action and the speed of movement from the first position to the second position. Typically, after performing a certain action, the target object may continue to perform that action, which may lead to safety risks. However, since the target object itself cannot judge the potential safety risks, a safety risk prediction is made for the target object based on the velocity parameter, the second position, and the movement trajectory from the first position to the second position. Because the velocity parameter reflects the movement of the target object performing the first action, it can indicate whether the target object will face safety risks if it continues to move at this speed. Similarly, the movement trajectory reflects the trend of the first action, i.e., it can indicate which direction the target object might continue to move in, and therefore, it can also indicate whether the target object will face safety risks if it continues to move. Therefore, the predicted risk result can be used to indicate whether there is a safety risk for the target object when it continues to move along the predicted trajectory with the velocity parameter. Correspondingly, if the risk prediction results indicate that there is a safety risk to the target object, a safety warning can be issued to remind relevant personnel to take action and protect the target object.
[0103] As can be seen, this application combines the target object's own movement speed and the trend of its first action to determine whether the target object will face safety risks as it continues to move, and then issues an early warning when risks are detected, thereby improving the safety protection of the target object. Furthermore, the prediction is based on the current state of the target object's first action, making it a targeted prediction for that specific target object. Since different target objects may have different speed parameters and movement trajectories for the same first action, combining the current state of the target object's first action can more accurately predict the situations the target object may face, thus enabling more accurate safety risk prediction and early warning. In practical applications, video streams can be acquired by information acquisition devices; therefore, compared to manual review, automatic prediction and early warning based on video streams are more conducive to timely detection of potential safety issues, reducing oversights, and better protecting the target object.
[0104] The above embodiments illustrate the safety warning method for a target object provided in this application. It should also be noted that this application does not limit the methods for predicting safety risks based on speed parameters, a second position, and a motion trajectory (i.e., the specific implementation of S203 mentioned above), nor does it limit the methods for detecting image frames in the video stream to determine if the target object has performed a first action. For ease of understanding, this application will be described in detail through the following embodiments:
[0105] (i) Regarding how to predict safety risks based on speed parameters, second position, and motion trajectory, the embodiments of this application provide the following examples:
[0106] It is understandable that the potential dangers faced by a target object may differ depending on the first action performed, and the potential dangers may also differ depending on the movement of the first action performed. Therefore, in one possible implementation, the existence of a safety risk for the target object can be determined based on the specific details of the first action performed by the target object (such as the type of action, the degree of risk, etc.) and the movement of the target object performing the first action (such as the speed of movement).
[0107] To facilitate understanding, the embodiments of this application provide the following examples:
[0108] It is understandable that the target object has poor control ability, so it may accidentally injure itself during the movement, especially when the movement speed is relatively fast. For example, crawling too fast may twist the hand or elbow, and running too fast (such as infants who have just learned to run or elderly people with mobility difficulties) may cause them to fall.
[0109] Therefore, as an example, if the speed parameter exceeds the motion speed threshold corresponding to the first action, the risk prediction result indicates that the target object faces a safety risk. The motion speed threshold can be used to indicate the safe motion speed limit for the first action. In practical applications, the safe motion speed limit for the current target object when performing the current first action can be determined comprehensively based on the specific circumstances of the first action (such as the type of action, the degree of risk, etc.) and the characteristics of the target object (such as age, etc.). Based on this, it is possible to quickly detect situations where the target object may face danger due to excessive speed, and subsequently issue warnings to avoid such dangers.
[0110] During movement, the target object cannot autonomously determine whether the detection scene is safe, such as whether objects in the detection scene corresponding to its movement path will harm it. Therefore, as another example, if obstacles exist in the detection scene corresponding to the predicted trajectory, the risk prediction result indicates that the target object faces a safety risk. The detection scene corresponding to the predicted trajectory can indicate which parts of the detection scene the target object might pass through as it continues to move. Obstacles can refer to objects that may pose a danger to the target object, such as stools or kettles in the detection scene. Based on this, it is possible to quickly detect the potential danger of the target object colliding with objects in the detection scene and being injured, and subsequently issue warnings to prevent such dangers from occurring.
[0111] In practical applications, the moving parts of a target object may have an object attached to them (such as a toy in its hand), which could pose a danger to the target object if it continues to move. Therefore, as another example, if a target part has a target object attached to it, and the target object will cause harm to the target object as it continues to move along the predicted trajectory with a velocity parameter, the risk prediction result indicates that there is a safety risk to the target object. Here, the target object can refer to external objects attached to the target part (such as toys, unclean debris, plastic bags, etc.). Based on this, it is possible to quickly detect whether the target object will be injured by its own attached target object as it continues to move, and when it is determined that the target object may cause harm to the target object (such as being stuffed in the mouth and affecting breathing, cutting itself, or tripping over a plastic bag on its head or feet), predictions can be made to avoid such dangers as much as possible.
[0112] In response to this situation, during implementation, the target object can be identified, and its specific characteristics can be considered to determine whether it will damage the target object. For example, such as... Figure 1 The scenario depicting a toy being held in the hand and potentially put in the mouth confirms a safety risk for the infant. Understandably, this is in response to... Figure 1If the action shown indicates that the infant is eating and is holding food (the food in the mealtime can be considered as food that will not harm the infant), then when it is determined that the infant has brought his right hand close to his face and may continue to bring it close to his mouth, it can be considered that there is no safety risk.
[0113] Because the target object cannot independently assess dangerous situations, a corresponding safe activity range is usually determined for it, allowing it to move within a relatively safe area. Understandably, if the target object exceeds this safe activity range, it may still face other dangers. Therefore, as another example, if the target object continues to move along the predicted trajectory with a speed parameter, and its position in the detection scene exceeds the safe area within a preset time, it indicates that the target object will leave the safe area shortly after continuing to move, potentially facing other dangers. Thus, the risk prediction result indicates that the target object faces a safety risk. Here, the safe area refers to the target object's safe activity range within the detection scene, and the preset time is used to determine whether it will quickly leave the safe area if it continues to move. Based on this, it is possible to detect situations where the target object may exceed the safe area and face danger, and issue warnings to prevent such dangers from occurring.
[0114] It should be noted that this application does not impose any limitations on how the preset duration is set. In practical applications, it can be flexibly set according to the actual situation. For ease of understanding, the following examples are provided in the embodiments of this application:
[0115] In one example, the settings can be based on the characteristics of the first action (such as action type, risk level, etc.). For instance, the higher the risk level of the first action, the shorter the preset duration can be set. In another example, the settings can be based on the characteristics of the target object (such as age, disability status, etc.). For example, the younger the target object, such as infants and toddlers, the weaker their ability to cope with danger, so a shorter preset duration can be set. Yet another example can also be based on the detection scenario. Based on this, it is possible to more flexibly predict safety risks in specific detection scenarios and for the current target object, so as to take timely and appropriate measures when potential safety risks exist.
[0116] It should also be noted that this application does not impose any limitations on how to determine the safe zone or how to determine whether one will leave the safe zone soon. For ease of understanding, the embodiments of this application provide the following examples:
[0117] First, regarding how to determine the safe area, this application provides the following example:
[0118] Understandably, the area that can support the target's activities will differ in different testing scenarios. For example, in a hospital setting, the area that can support the target's activities might be the bed or the room itself. In a home setting, the area could be the bed or the living room. Furthermore, target subjects of different ages have varying behavioral abilities, and their corresponding activity ranges will also differ.
[0119] Therefore, in one possible implementation, the safe zone can be determined based on the characteristics of the detection scenario and the characteristics of the target object. The characteristics of the detection scenario indicate the specific situation of the current detection scenario, while the characteristics of the target object indicate the specific situation of the target object itself. Based on this, by combining these two aspects, a safe zone that is suitable for the current detection scenario and supports the activities of the current target object can be determined, making it more flexible and adaptable to different detection needs.
[0120] In practical applications, it also allows relevant personnel (such as the caregiver of the target individual) to customize the safe zone. For example, in a hospital testing scenario, medical staff can customize the safe zone for the target individual (e.g., define it as the area of the medical staff's room where the target individual is located). Similarly, in a home testing scenario, family members of the target individual can customize the safe zone for the target individual (e.g., define it as a specific area in the living room), and so on.
[0121] Secondly, regarding how to determine whether one will soon leave the safe zone, the embodiments of this application provide the following examples:
[0122] Generally, "target leaving the safe zone" refers to a tendency for the target to leave the safe zone entirely. Therefore, it can be considered that the first action that changes the target's location may lead to the target facing the risk of leaving the safe zone. For better understanding, this application uses this as an example:
[0123] If the first action is a position-changing action, this action can be used to change the position of the target object in the detection scene. In other words, the first action is an action that changes the position of the target object, such as rolling over, crawling, walking forward, or running. Accordingly, in the specific implementation of S203 mentioned above:
[0124] First, the server can determine the safe zone corresponding to the target object in the detection scene. Next, the server can perform a safety risk prediction on the target object based on the speed parameter, the second position, the trajectory of movement from the first position to the second position, and the safe zone, obtaining a risk prediction result. In this example, the obtained risk prediction result can be used to indicate whether the target object's position in the detection scene will exceed the safe zone within a preset time period when it continues to move along the predicted trajectory at the speed parameter. That is, it determines whether the target object will quickly leave the corresponding safe zone under the current position change action.
[0125] For example, the safe zone can be used to represent the bed area, and the position change action can be a rolling over action. If the second position is already relatively close to the edge of the bed, it is predicted that if the target object continues to roll over one or two more times, it may fall off the bed. Therefore, it can be assumed that the target object will leave the safe zone quickly, thus determining that there is a safety risk.
[0126] Based on this, during the target object's activities, it is possible to predict potential safety risks related to leaving the safe area, and avoid related dangers (such as falling off the bed, climbing out of the safe activity area and colliding with other obstacles in the detection scene).
[0127] The above embodiments illustrate an implementation method for predicting safety risks by combining the specific circumstances of the first action with the movement of the target object. It is understood that since prediction analyzes the situation during continued movement to determine whether there is a safety risk, and the pose of the target object at the current moment (i.e., the acquisition time of the j-th image frame) can reflect how the target object might move next, it can assist in safety risk prediction. Correspondingly, in another possible implementation, safety risk prediction can also be performed by combining the pose of the target object at the current moment.
[0128] In practical implementation, the server can determine the pose parameters of the target object based on the j-th image frame. These pose parameters indicate the pose of the target object at the acquisition time of the j-th image frame. That is, the pose parameters reflect the pose of the target object's body in the detection scene at the acquisition time of the j-th image frame (specifically, the pose of each part of the target object). Correspondingly, the server can predict the safety risks of the target object based on the velocity parameters, pose parameters, the second position, and the motion trajectory from the first position to the second position, thus obtaining the risk prediction result.
[0129] Because the pose of the target object at the acquisition time of the j-th image frame (i.e. the pose at the current time) can reflect how the target object may move next, combining the pose for safety risk prediction is beneficial to more accurately analyze the possible movement of the target object next, thus obtaining more accurate risk prediction results.
[0130] (ii) Regarding how to detect image frames included in a video stream to determine if a target object has performed a first action, the embodiments of this application provide the following examples:
[0131] Example 1:
[0132] It is understandable that if a target object performs an action, this may cause the target object's posture to change. Therefore, in one possible implementation, when performing frame-by-frame detection of image frames in a video stream, the posture of the target object can be identified, thereby analyzing the changes in the target object's posture over time, and thus determining whether the target object has performed an action.
[0133] In practical applications, to improve recognition efficiency, a target object pose recognition model can be used to identify the pose of the target object in image frames. Specifically, the server can input each image frame from the video stream into the target object pose recognition model, which then outputs a pose recognition image for each frame. This pose recognition image indicates the pose of the target object within the image frame, specifically the pose of the target object at the corresponding acquisition time. The service can then perform action recognition on the target object based on the obtained pose recognition images and their corresponding acquisition times, yielding the action recognition result.
[0134] Understandably, a video stream typically contains multiple image frames. By identifying each frame sequentially according to its capture time, the pose of the target object at each capture time can be determined. Furthermore, by sorting these frames according to their capture times, it's possible to analyze whether the target object has performed its first action over time. Therefore, the action recognition results can be used to indicate whether the first action began in the i-th image frame and to confirm that the first action was performed in the j-th image frame.
[0135] Based on this, by identifying the posture of the target object at each acquisition moment, it is possible to detect whether the target object has performed its first action. In practical applications, when the target object performs certain actions (such as crawling, rolling over, etc.), it mobilizes multiple parts of its body. Therefore, this method can better analyze the posture of each part of the target object, and thus analyze the movement of the target object more comprehensively and accurately from the overall posture of the target object, thereby improving the detection accuracy.
[0136] It is understandable that the aforementioned pose parameters (i.e., the pose of the target object at the acquisition time of the j-th image frame) can also be determined by using a target object pose recognition model to determine the pose recognition image corresponding to the j-th image frame, and then determining the pose parameters based on the corresponding pose recognition image. For example, the corresponding pose recognition image can be directly used as the pose parameters, or the pose indicated in the pose recognition image can be directly described using descriptive text.
[0137] It should be noted that this application does not impose any limitations on the method for determining the target object pose recognition model. For ease of understanding, the following examples are provided in the embodiments of this application:
[0138] In one possible implementation, an existing posture recognition model can be directly reused as the aforementioned target object posture recognition model. This existing model can refer to a model with human posture recognition capabilities, which helps reduce costs. For example, this approach can be used to detect the postures of adults such as the elderly and disabled individuals.
[0139] It is understandable that the shapes and proportions of different body parts in infants and young children differ significantly from those in adults. Therefore, to obtain a posture recognition model with strong infant posture recognition capabilities, existing posture recognition models can be trained using sample images of infant skeletal structures. These existing posture recognition models can refer to models capable of recognizing human postures. Based on this, directly reusing existing posture recognition models helps reduce the cost of obtaining a target object posture recognition model with infant posture recognition capabilities. Furthermore, using sample images of infant skeletal structures for model training allows for fine-tuning of existing posture recognition models, making them suitable for recognizing infant postures. For ease of understanding, the following examples are provided in this application:
[0140] In practical applications, based on the infant's skeletal structure and the proportions of the head, body, and limbs, features such as ellipse fitting information and the projection histogram along the ellipse axis can be used to distinguish the infant's posture from important limb parts (such as the head, hands, elbows, legs, and feet). For an example, see [link to relevant documentation]. Figure 3The diagram illustrates the skeletal structure of an infant. In infancy, the head is relatively large and the limbs are relatively rounded. Since existing posture recognition models inherently possess human posture recognition capabilities, during the fine-tuning model training phase, the aforementioned sample images can include positive sample images. These positive sample images contain images of infants and are labeled with their skeletal structure. The labels on the positive sample images indicate the infant's posture and important limb parts, enabling the posture recognition model to learn relevant knowledge about infant posture. Of course, negative sample images, i.e., sample images excluding infants, can also be added to prevent overfitting of the posture recognition model.
[0141] In practical implementation, the pose recognition model can employ the OpenPose algorithm to extract features from the input sample image, identifying potentially important limb parts of the infant or toddler and the connections between these parts, ultimately outputting a corresponding pose recognition prediction image. In practical applications, the OpenPose algorithm can extract features from the input sample image to obtain a corresponding feature map. Then, confidence and correlation information can be extracted from the feature map. Confidence indicates the reliability of a particular part in the feature map representing a specific infant or toddler part, while correlation information indicates the connections between these parts.
[0142] The pose recognition prediction image can be used to indicate the infant's pose in the sample image as predicted by the pose recognition model. Then, based on the difference between the pose recognition prediction image and the sample label, the pose recognition model can be trained to learn relevant knowledge about infant poses (such as the shape of each limb), ultimately resulting in an infant pose recognition model with strong ability to recognize infant poses. Correspondingly, when applied to the recognition of image frames in a detection scene, it helps to obtain more accurate infant poses, thus enabling more accurate analysis of whether the infant has performed any actions.
[0143] Example 2:
[0144] In practical applications, the frame difference method (i.e., inter-frame difference method) identifies moving targets by comparing the pixel differences between two adjacent image frames. Therefore, in one possible implementation, the frame difference method can be used to process the image frames included in the video stream to determine the moving targets whose positions have changed during the detection process, and then determine whether it is the moving target object.
[0145] In practice, the server can perform frame difference processing on the k-th and (k+1)-th image frames in the video stream to obtain a frame difference image. This frame difference image can be used to indicate moving targets whose positions change within the detection scene from the k-th to the (k+1)-th image frame, where k is a positive integer. Correspondingly, if the moving target is determined to include a target object, the server can perform action recognition on the target object based on the obtained frame difference image to obtain an action recognition result. This action recognition result can be used to indicate whether the first action was initiated at the i-th image frame and whether the first action was confirmed to have been implemented at the j-th image frame.
[0146] Based on this, by performing frame difference processing on two adjacent image frames, moving targets in the detection scene during the detection process can be identified, thereby determining whether the target object has performed the first action. Since changes in lighting in the detection scene will affect the acquired image frames, and the frame difference method is not sensitive to changes in lighting, this method can take into account the differences in lighting in different detection scenes or the differences in lighting at different times in the same detection scene, thus making it more applicable.
[0147] To better understand this, the embodiments of this application also provide the following examples. Specifically, the aforementioned frame difference processing can be implemented using the following formula:
[0148]
[0149] In the above formula, (r, q) can be used to indicate the coordinates of a point in an image frame, f k (r, q) can be used to indicate the pixel value of that point in the k-th image frame, f k+1 (r, q) can be used to indicate the pixel value of that point in the (k+1)th image frame, |f k+1 (r, q)-f k (r, q) can be used to indicate the change in pixel value at a point from the k-th image frame to the (k+1)-th image frame, T h It can be used to indicate the pixel difference threshold. If the difference between the pixel values of a point in two adjacent image frames is less than or equal to the pixel difference threshold, it is considered that the real object corresponding to that point in the detection scene has not changed, so it can be set to 0; otherwise, it is considered to have changed, so it is set to 1. Accordingly, d(r, q) can be used to indicate the frame difference processing result corresponding to that point.
[0150] Based on this, the resulting frame difference image can be a binary image, where the regions identified by points with a value of 1 can be considered as moving targets whose positions have changed in the detected scene. Therefore, subsequent analysis can be performed based on the obtained frame difference image.
[0151] The above embodiments have provided a detailed description of the safety warning method for a target object provided in this application. In summary, frame-by-frame detection of image frames in a video stream enables dynamic recognition of the target object's actions. Based on the results of this dynamic recognition, it is determined whether the target object has performed a certain action. Furthermore, after determining that a first action of a certain dangerous type has been performed, safety risk prediction and warnings can be issued when danger is detected, thereby providing safety protection for the target object.
[0152] For ease of understanding, this application embodiment takes the aforementioned infants and young children as the target object and provides the following... Figure 4 The diagram shown illustrates a dynamic recognition process, specifically:
[0153] For detection scenarios involving infants and toddlers, a video stream can be acquired (see S401). Then, image frames can be determined (see S402), for example, by extracting frames at regular intervals. Next, the obtained image frames can be input into an infant pose recognition model (see S403). The model then identifies each image frame individually to obtain the corresponding pose, such as pose i corresponding to the i-th image frame and pose j corresponding to the j-th image frame (see S404). Here, the infant pose recognition model can refer to the aforementioned target object pose recognition model with infant pose recognition capabilities.
[0154] Subsequently, based on the results of frame-by-frame recognition—that is, the infant's posture over time—it can be comprehensively determined whether the infant has performed the first action (see S405). Based on this, the video stream can be extracted into multiple image frames, and recognition can be performed sequentially according to the acquisition time of each image frame. This decomposes the process of the infant performing an action into multiple sub-processes, identifies each sub-process, and then integrates the recognition results of multiple sub-processes according to the order of acquisition time to obtain the final action recognition result (indicating whether the infant has performed the first action). This achieves dynamic action recognition. Compared to related technologies that rely solely on a static image to determine whether an action has been performed, this method offers higher accuracy.
[0155] Next, a safety risk prediction can be performed on the infants and young children (see S406), which will yield the aforementioned risk prediction results. Then, based on the risk prediction results, it can be determined whether there is a safety risk to the infants and young children (see S407). If a safety risk is determined to exist, a safety warning will be issued (see S408).
[0156] Accordingly, if it is determined that the infant has not performed the first action or that there is no safety risk to the infant, the detection can continue, that is, the acquisition of video streams can continue, thereby achieving continuous detection. It is understood that, in the specific implementation of the above steps, please refer to the relevant descriptions in the foregoing embodiments, which will not be repeated here.
[0157] To better understand, embodiments of this application also provide, as follows: Figure 5 The diagram shows the structure of a video stream acquisition module, which mainly includes a network camera, a video gateway, and a video access service. Specifically:
[0158] A network camera (IP camera), which can be the aforementioned information acquisition device, can be used to capture corresponding video streams in the detection scene. Based on this, the physical space corresponding to the detection scene can be digitized to obtain a video stream for subsequent use.
[0159] A video gateway can include a device status detection process (for detecting the status of network cameras), a video capture process (for detecting the video capture status), a message receiving process (for receiving messages, such as network camera fault messages), a gateway software development kit (for users to process messages), and a gateway master-slave relationship (for managing the master-slave relationship of multiple gateways). For example, multiple gateways can be an integration platform as a service (IPaaS), an Internet of Things gateway (IoT), etc.
[0160] Video access services can retrieve video streams through video gateways and transmit them to servers and terminals (such as mobile devices and computers running security alert systems). In practical applications, video access services can be built on platforms such as WeLink and IoT.
[0161] In addition, an image data conversion model can be included to convert the format of the acquired video stream, so that different video stream formats can be retrieved on demand. For example, it can retrieve multiple bitstream formats such as IoT, GB, and Onvif. In practical applications, relevant GB protocols can be used for broadcasting to retrieve video streams that meet the requirements.
[0162] The above embodiments illustrate the safety warning method provided in this application. It is understood that infants and young children, besides being unable to independently assess safety risks, may also be unable to express their needs verbally, such as infants who have not yet learned to speak, infants who have just learned to speak with weak language skills, or people with aphasia. Therefore, this application also provides a needs analysis method for the target audience, specifically:
[0163] For example, infants and toddlers often use body language (such as making various gestures) to express their needs. Similarly, people with aphasia and the elderly with limited language skills also use body language (such as making various gestures) to express their needs. Therefore, one possible approach is to analyze the needs of the target audience based on the actions they perform.
[0164] In practical implementation, the server can identify the actions performed by the target object during the detection process based on the image frames included in the video stream. Specifically, as shown in the figure, the actions of the target object can be determined using methods such as frame difference analysis and pose recognition. Correspondingly, if the video stream indicates that the target object performs multiple secondary actions during the detection process, the server can perform a requirements analysis on the target object based on these secondary actions to obtain the requirements analysis results. Furthermore, the server can generate requirements prompt information based on the requirements analysis results. This requirement prompt information can be used to indicate the requirements of the target object as indicated by the requirements analysis results, so as to promptly inform relevant personnel that the target object may have a certain requirement.
[0165] The second action can refer to any action performed by the target subject during the detection process, such as crying, laughing, patting their upper limbs, kicking their legs, shaking their head, and the aforementioned actions like rolling over, crawling, and putting something in their mouth. In other words, the second action can include the aforementioned first action. Based on this, the target subject's needs can be comprehensively analyzed according to their actions, that is, action analysis can explain why the target subject performs these actions, thereby assisting relevant personnel in better caring for the target subject.
[0166] It should be noted that this application does not limit the methods for performing requirements analysis on the target object based on multiple second actions. For ease of understanding, the embodiments of this application provide the following examples:
[0167] As an example, we can analyze the combination of multiple simultaneous secondary actions to understand why the target individual performs these actions, thereby clarifying their potential needs. For instance, if the target individual is an infant, and the secondary actions include crying and sucking their fingers, it may indicate that the infant is hungry. Therefore, the needs analysis results indicate that the infant is hungry and needs to eat. Similarly, if the secondary actions include crying and keeping their bottom still while patting their upper limbs, it may indicate that the infant needs to clean their bottom.
[0168] Based on this, compared with the method of judging a single action, this application integrates multiple actions to analyze the possible needs of the target object, which is conducive to obtaining more accurate needs analysis results, and thus better guiding relevant personnel to care for the target object.
[0169] In practical applications, the specific circumstances under which a target individual performs a certain action will vary depending on the situation. For example, the manner of crying will differ depending on whether the individual is hungry or feeling unwell. Therefore, in one possible implementation, when it is determined that an infant is crying, the characteristics of the infant's crying can be considered to analyze the infant's needs.
[0170] In practice, if the target is an infant or toddler, and if the multiple second actions include the action of the target crying, the server can obtain the crying sound parameters corresponding to the infant or toddler during the detection process, and perform a demand analysis on the infant or toddler based on the multiple second actions and the crying sound parameters to obtain the demand analysis results.
[0171] Among these parameters, crying sound parameters can be used to indicate the specific details of an infant's crying action, such as whether the cry is loud, sharp, or continuous. Since infants' cries vary depending on their state, combining crying sound parameters helps to more accurately analyze an infant's potential needs, thus obtaining more accurate needs analysis results.
[0172] For example, if the crying parameters indicate a sharp cry, accompanied by waving hands and feet, throwing objects, etc., the infant may be dissatisfied with their environment or afraid. In this case, caregivers can be reminded to soothe the infant. Conversely, if the crying parameters indicate a loud cry, accompanied by thumb-sucking, the infant may be hungry. Caregivers can be reminded to help the infant eat.
[0173] Understandably, the facial color of the target object may vary depending on its state. For example, when feeling unwell, its face may appear pale or yellowish. Conversely, when crying intensely, its face may appear red or purplish. Therefore, another possible implementation approach is to analyze the target object's needs by considering its facial color.
[0174] In practice, the server can obtain the facial color parameters of the target object during the detection process, and perform a requirement analysis on the target object based on various secondary actions and facial color parameters to obtain the requirement analysis results.
[0175] Among these parameters, facial color can be used to indicate the color of the target subject's facial skin. Normally, the target subject's facial skin appears rosy in a healthy state. However, in other abnormal states (such as crying, physical discomfort, etc.), the facial skin may appear a different color. Therefore, combining facial color parameters helps to more accurately analyze the target subject's potential needs, thus obtaining more accurate needs analysis results.
[0176] For example, if facial color parameters indicate that the face is flushed or purplish, and the person is crying loudly, it may be due to prolonged crying, possibly affecting their breathing. Caregivers should be reminded to soothe the person. Furthermore, continue to observe the facial color after the crying stops. If the aforementioned symptoms persist, a professional examination should be conducted promptly.
[0177] The above embodiments have provided a detailed description of the requirement analysis method for the target object provided in this application. It can be understood that the above requirement analysis method uses the actions performed by the target object to perform requirement analysis, and can therefore be regarded as a secondary processing of the action identification results.
[0178] Furthermore, to better protect the target object, intelligent devices can be used to detect its status, achieving more comprehensive security. Additionally, the presence of strangers in the detection scene can be detected to prevent unauthorized personnel from negatively impacting the target object. For ease of understanding, the embodiments of this application will be illustrated through the following examples:
[0179] On the one hand, regarding the detection of target objects using smart devices, the following example is provided:
[0180] In practical applications, smart devices can be smart bracelets, smart ankle bracelets, etc., which detect the status of a target by having the target wear the smart device. For example, this could include detection of the target's vital signs, location detection, early warning alerts, password unlocking, etc. Specifically:
[0181] Vital signs detection can be used to detect the physical characteristics parameters of a target subject (such as body temperature), thereby detecting the target subject's physical health status and enabling timely detection of any physical discomfort (such as abnormal physical signs such as high temperature).
[0182] Location detection can be used to detect the location of a target object, enabling quick determination of its location. It can also be used to determine if a target object has left a safe area, and to issue a warning directly when the location detection indicates that the target object has left the safe area.
[0183] Early warning and alerts can be used to directly warn people of potential safety risks through smart devices, such as through sound and light warnings.
[0184] Password unlocking can be used to bind information about the target's family members. For example, a family member's fingerprint can be used as the lock password for a smart device to prevent misidentification.
[0185] Based on this, and by combining the relevant settings of smart devices, the security of the target object can be further guaranteed.
[0186] On the other hand, the following example illustrates how to perform stranger detection:
[0187] In one possible implementation, the presence of a stranger in the detection scene can be identified based on the image frames included in the acquired video stream. Specifically, if the image frame contains a target object, the facial features of the target object are acquired and compared with the facial features of objects stored in a database. If the comparison result indicates that the target object does not belong to any stored object, it means that a person not in the database has appeared in the detection scene, and therefore a stranger is considered to be present, triggering an illegal object warning. Conversely, if the comparison result indicates that the target object belongs to any stored object, no warning is needed. Based on this, the system can promptly detect and issue warnings for strangers entering the target object's location, preventing unauthorized individuals from negatively impacting the target object.
[0188] The objects to be identified can refer to people detected in the detection scene. In practical applications, neural network algorithms (such as target recognition algorithms) can be used to determine the objects to be identified in the image frame.
[0189] Furthermore, the objects stored in the object database can refer to legitimate individuals, such as medical staff and family members. In practical applications, the object database can be built as needed. For example, in a hospital detection scenario, information on relevant medical staff can be statistically analyzed to establish a long-term valid medical staff database, and information on family members can be statistically analyzed to establish a temporary valid family member database. In the temporary valid family member database, an expiration period is set for each family member (e.g., based on the length of time the target object corresponding to that family member will stay in the hospital), and a binding relationship is established between each family member and their corresponding target object. This allows for timely notification of the target object's family members when a security risk is detected.
[0190] It should also be noted that this application does not limit the method for determining whether a candidate object belongs to an object stored in the object library based on the comparison results. In practical applications, a feature similarity m (where m is a value between 0 and 1) can be set as a comparison threshold. If the maximum similarity between the facial features of the candidate object and the facial features of an object stored in the object library is less than the comparison threshold, then the candidate object is considered not to belong to an object stored in the object library, i.e., the candidate object is not in the library and an alert is issued. Conversely, if the similarity between the facial features of the candidate object and the facial features of an object stored in the object library is greater than or equal to the comparison threshold, then the candidate object is considered to be that object stored in the object library, i.e., the candidate object is in the library, and no alert is required.
[0191] The security early warning method provided in this application has been described in detail through the above embodiments. In practical applications, in order to facilitate relevant personnel to view the status of target objects at any time and to promptly convey relevant notifications (such as security early warning notifications, illegal object warning notifications, and requirements analysis results notifications) to relevant personnel, this application also provides a security early warning system that can run on terminal devices such as mobile terminals and computers. To this end, this application also provides... Figure 6 The diagram shown illustrates the application architecture of a security early warning system. Specifically:
[0192] Server 601 can be used as the aforementioned computer device to execute the security warning methods provided in the foregoing embodiments of this application. Information acquisition device 602 can be installed in scenarios where data needs to be collected, such as the aforementioned target object detection scenario, thereby enabling it to acquire the aforementioned video stream. In practical applications, a connection (such as a network connection) can be established between information acquisition device 602 and server 601 so that server 601 can obtain the acquired video stream from information acquisition device 602 and perform subsequent processing.
[0193] Furthermore, the terminal 603 (such as a computer, smartphone, etc.) can be used to run the aforementioned security early warning system so that relevant personnel (such as medical staff, family members, etc.) can check the status of the target at any time and receive relevant notifications.
[0194] For ease of understanding, Figure 6 The example, using a hospital testing scenario, also provides a sample system interface for a safety warning system, specifically:
[0195] The security early warning system can include functional modules such as historical playback, detection settings, and early warning center. The historical playback module allows users to view historical videos, the early warning center module allows users to view early warning details, and the detection settings module allows users to set detection-related data (such as security zone settings, early warning settings, etc.).
[0196] For ease of understanding, Figure 6 The system interface of the security early warning system shown is primarily illustrated using the interface after selecting the detection settings module. The information acquisition device list displays the information acquisition devices connected to the security early warning system and their corresponding device identifiers. For example, in a hospital detection scenario, multiple information acquisition devices might be installed in the room where the target object is located, as well as in related corridors and stairwells. The device identifiers clearly indicate which area the information acquisition device is collecting video streams from, allowing users to easily view the information.
[0197] In practical applications, if relevant personnel select the information collection device listed as "Room 703" from the information collection device list, the video stream from Room 703 can be displayed accordingly. Figure 6 The image frame corresponding to the video stream in room 703 is shown. In practical applications, it can also support relevant personnel to zoom in on the image frame to observe the details of specific points. For example, it supports manually selecting the point of interest and double-clicking to zoom in with one click. Furthermore, when an alert is issued, it can also support linking information acquisition devices (such as PTZ cameras) to zoom in on abnormal areas.
[0198] In addition, in the security zone settings, corresponding security zones can be set for the detection scenario corresponding to room 703, such as the entrance / exit area and the care area where the target object is located. Furthermore, in the alert settings, relevant detection parameters can be set for the detection scenario corresponding to room 703. For example, the detection time can be set to t1-t2, the alert method when an alert is issued (e.g., selecting a pop-up alert), the risk level (e.g., selecting a high-risk level), and the duration of time a stranger stays (e.g., setting it to aaa seconds).
[0199] And, for Figure 6 The example of the early warning center module, using the detection of an unauthorized stranger as an example, also provides, in this embodiment of the application, an early warning system module. Figure 7 The diagram shown illustrates an interface for an early warning center, which can display early warning information, operation logs, and processing actions. Specifically:
[0200] The warning information may include: warning type (e.g., a stranger warning, indicating that the current warning refers to a stranger entering the detection scene), personnel identification (e.g., a stranger may be unknown, or a person in the database may be a person's ID number), warning device (e.g., the device identification of the information collection device, indicating which information collection device the warning was based on), warning floor (e.g., the seventh floor, indicating the location where the warning occurred), warning number (e.g., 01001), warning time (e.g., xxxxx, indicating the time when the warning occurred), risk level, and warning status (e.g., pending, processed).
[0201] The processing options can include ignoring controls and processing controls, so that relevant personnel can handle this alert.
[0202] Furthermore, the operation log can be used to record the handling of the alert. For example, when a relevant person handles the alert through a processing control at a certain time, an operation log can be generated stating "At xxx time, xxx performed the processing." for subsequent review of the handling process.
[0203] Based on this, relevant personnel (such as medical staff, family members, etc.) can view the panoramic view of the detection scene corresponding to the target object and the specific situation of the warning at any time through the security warning system on a computer or mobile device. They can also receive warnings, such as pop-up reminders after setting them up, when a warning occurs.
[0204] Understandable Figure 6 and Figure 7 This is merely an example and is not intended to impose any limitations. In practical applications, the functional modules and interfaces of the safety warning system can be developed and designed as needed, and corresponding warning information can be displayed according to different warning situations.
[0205] The safety warning method provided in this application has been described in detail through the above embodiments. To gain a more comprehensive understanding of this application, this embodiment takes infants and young children as the target object as an example, and also provides... Figure 8 The diagram shown illustrates a logical architecture for a security early warning system, which mainly includes a detection system and a security early warning system. Specifically:
[0206] On the one hand, the detection system can perform detection on various parts of the video stream, specifically including:
[0207] The infant feature fitting module can assess the skeletal structure of infants and fit their limb parts to achieve posture recognition. For example, this module can determine the aforementioned infant posture recognition model and the posture recognition image corresponding to each image frame. Because infants' limbs are shorter and often curled up, their limb modeling differs significantly from that of adults. Therefore, infant feature fitting enables the modeling of infants' limbs, improving the accuracy of posture recognition.
[0208] The motion recognition module can identify the actions performed by infants and toddlers to determine whether an action has been performed, the corresponding motion parameters (such as the aforementioned trajectory length, duration, and speed parameters), and the infant's location, and based on this, predict safety risks. Thus, through in-depth analysis of actions, it can predict the infant's subsequent movement trajectory to detect safety risks in advance. For example, whether the infant is approaching the boundary of a safe area or leaving the safe area.
[0209] The needs analysis module can analyze the needs of infants and toddlers, specifically by identifying various secondary actions performed by the infants, facial color parameters, and crying sound parameters, and then conducting needs analysis based on this information. This helps to promptly identify the essential needs of infants and toddlers (such as hunger, physical discomfort, etc.).
[0210] The stranger detection module can detect the presence of strangers in a scene. For example, it can identify potential objects in the scene and compare their features with those of objects stored in a database to determine if the potential object is a registered person. This helps prevent strangers from approaching infants and toddlers and causing them any negative impact.
[0211] The intelligent device detection module can monitor the condition of infants and young children using intelligent devices. For example, it can perform vital sign detection (such as body temperature), anti-loss detection (such as determining the location of the infant or young child, whether they have left the safe area, etc.), and determine whether there are any abnormalities (such as abnormal body temperature, etc.) based on this.
[0212] Based on this, the detection system can assess potential risks to infants and young children from various perspectives. When a safety risk is identified, an early warning can be issued. Examples include the aforementioned safety warnings, demand alerts, and warnings about unauthorized access. In practical applications, when an early warning occurs, it can be pushed to the safety warning system so that relevant personnel can view and address it promptly.
[0213] On the other hand, the security early warning system can support relevant personnel to view videos and handle relevant warnings, specifically including:
[0214] Video viewing supports historical playback and real-time viewing, making it convenient for relevant personnel to view the actual situation in the testing scenario as needed.
[0215] The detection settings allow relevant personnel to configure detection-related parameters, such as safe zone settings and alert settings. For details, please refer to [link to relevant documentation]. Figure 6 Examples and explanations are provided.
[0216] The early warning center allows relevant personnel to view early warning information and take appropriate actions. For details, please refer to [link / reference needed]. Figure 7 Examples and explanations are provided.
[0217] Furthermore, embodiments of this application also provide, as follows: Figure 9 The diagram shown illustrates the architecture of a network service system, which mainly includes a Web module, a middleware module, and an algorithm cluster module. Specifically:
[0218] The web module can perform task management, provide routing services, and interface services. In practical applications, the aforementioned security alert system can be developed based on the web module. Correspondingly, users can configure and manage detection tasks through the security alert system. Furthermore, the routing and interface services can be used to enable interaction between the web module and related services in the middleware module to achieve the detection tasks described in this application.
[0219] Middleware modules can include distributed coordination services and message queues. The distributed coordination service, for example, could be Zookeeper, which can distribute tasks from the web module to the algorithm cluster module, allowing the algorithm cluster module to call relevant modules to monitor the tasks. The message queue can be used to cache various messages, such as messages received from the web module corresponding to user-defined detection parameters through the security alert system, and messages received from the algorithm cluster module corresponding to issuing alerts.
[0220] The algorithm cluster module may include a video gateway, a central processing unit (CPU) module, a graphics processing unit (GPU) module, an algorithm module, an early warning module, a configuration module, a proxy module, and a storage module. The video gateway can be used to pull video streams; see [link to relevant documentation] for details. Figure 5Example. The Central Processing Unit (CPU) module manages algorithm tasks, while the Graphics Processing Unit (GPU) module determines the image frames included in the video stream (i.e., frame segmentation) and performs detection on these frames. The algorithm module manages various algorithms (such as the aforementioned object pose recognition model), analyzes for anomalies to determine whether to issue warnings, etc. The configuration module allows configuration of relevant algorithms via configuration files. Finally, the proxy module acts as a communication proxy between the algorithm cluster module and external services, ensuring operational reliability.
[0221] In practical applications, Figure 9 The network service system shown can be constructed using Spatial Pyramid Pooling (SPP), a high-efficiency, robust, and general-purpose network server framework. SPP provides various basic functions, such as logging, alert statistics, and content allocation, and also offers an Application Programming Interface (API) to facilitate convenient development by business personnel. For example, a corresponding SPP plugin can be developed for the security alert method provided in this application. Based on this, convenient on-demand development can be achieved to meet different detection needs.
[0222] It should be noted that, based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods.
[0223] based on Figure 2 Corresponding to the security warning method provided in the embodiments, this application also provides a security warning device 1000, which includes an acquisition unit 1001, a prediction unit 1002, and a warning unit 1003:
[0224] The acquisition unit 1001 is used to acquire the video stream corresponding to the detection scene of the target object;
[0225] The acquisition unit 1001 is further configured to determine that a target object included in the detection scene has performed a first action when the j-th image frame included in the video stream is detected, and to acquire a speed parameter corresponding to the target object performing the first action. The speed parameter is used to indicate the movement speed of the target part from a first position to a second position. The target part is the part of the target object that performs the first action. The first position refers to the position of the target part in the detection scene at the acquisition time of the i-th image frame. The i-th image frame is an image frame included in the video stream used to determine the start of the first action. The second position refers to the position of the target part in the detection scene at the acquisition time of the j-th image frame. The acquisition time of the i-th image frame is earlier than the acquisition time of the j-th image frame.
[0226] The prediction unit 1002 is used to predict the safety risk of the target object based on the speed parameter, the second position, and the movement trajectory from the first position to the second position, and to obtain a risk prediction result. The risk prediction result is used to indicate whether there is a safety risk to the target object when it continues to move on the predicted trajectory with the speed parameter. The predicted trajectory refers to the trajectory of the target part continuing to perform the first action from the second position.
[0227] The early warning unit 1003 is used to issue a safety warning if the risk prediction result indicates that the target object has a safety risk.
[0228] In one possible implementation, if the first action is a position transformation action, the position transformation action is used to change the position of the target object in the detection scene, and the prediction unit is further used to:
[0229] Determine the safe zone corresponding to the target object in the detection scenario;
[0230] Based on the speed parameter, the second position, the trajectory of movement from the first position to the second position, and the safe area, a safety risk prediction is performed on the target object to obtain the risk prediction result. The risk prediction result is used to indicate whether the position of the target object in the detection scenario exceeds the safe area within a preset time period when it continues to move on the predicted trajectory at the speed parameter.
[0231] In one possible implementation, the prediction unit is further configured to:
[0232] If the speed parameter exceeds the motion speed threshold corresponding to the first action, the risk prediction result indicates that the target object has a safety risk.
[0233] If there are obstacles in the detection scene corresponding to the predicted trajectory, the risk prediction result indicates that the target object has a safety risk.
[0234] If the target location has a target object, and the target object will cause damage to the target object when it continues to move on the predicted trajectory with the speed parameter, the risk prediction result indicates that there is a safety risk to the target object;
[0235] If the target object continues to move on the predicted trajectory at the speed parameter, and its position in the detection scenario exceeds the safe zone within a preset time period, the risk prediction result indicates that the target object faces a safety risk. The safe zone refers to the safe activity range of the target object in the detection scenario.
[0236] In one possible implementation, the prediction unit is further configured to:
[0237] The safe zone is determined based on the characteristics of the detection scene and the characteristics of the target object.
[0238] In one possible implementation, the acquisition unit is further configured to:
[0239] Based on the j-th image frame, the pose parameters corresponding to the target object are determined, and the pose parameters are used to indicate the pose of the target object at the acquisition time of the j-th image frame;
[0240] The prediction unit is further configured to predict the safety risks of the target object based on the speed parameters, the attitude parameters, the second position, and the motion trajectory from the first position to the second position, and obtain the risk prediction result.
[0241] In one possible implementation, the prediction unit is further configured to:
[0242] If it is determined from the video stream that the target object has multiple second actions during the detection process, a demand analysis is performed on the target object based on the multiple second actions to obtain the demand analysis results;
[0243] Based on the requirements analysis results, requirements prompt information is generated, which is used to indicate the requirements of the target object indicated by the requirements analysis results.
[0244] In one possible implementation, the prediction unit is further configured to:
[0245] Obtain the facial color parameters of the target object during the detection process;
[0246] Based on the various second actions and the facial color parameters, a requirements analysis is performed on the target object to obtain the requirements analysis results.
[0247] In one possible implementation, if the target object is an infant, and if the plurality of second actions includes the action of the target object crying, the prediction unit is further configured to:
[0248] Obtain the crying parameters of the infant during the detection process;
[0249] Based on the various second actions and the crying parameters, a needs analysis is performed on the infant to obtain the needs analysis results.
[0250] In one possible implementation, the prediction unit is further configured to:
[0251] Frame difference processing is performed on the k-th image frame and the (k+1)-th image frame included in the video stream to obtain a frame difference image. The frame difference image is used to indicate the moving target whose position changes in the detection scene during the process from the k-th image frame to the (k+1)-th image frame.
[0252] If the moving target is determined to include the target object, the target object is subjected to action recognition based on the obtained frame difference image to obtain an action recognition result. The action recognition result is used to indicate whether the first action was started at the i-th image frame and whether the first action was determined to have been performed at the j-th image frame.
[0253] In one possible implementation, the prediction unit is further configured to:
[0254] Each image frame included in the video stream is input into the target object pose recognition model, and the target object pose recognition model outputs a pose recognition image corresponding to each image frame. The pose recognition image is used to indicate the pose of the target object included in the image frame.
[0255] Based on the obtained posture recognition image, the target object is subjected to action recognition according to the acquisition time corresponding to the obtained posture recognition image, and the action recognition result is used to indicate whether the first action is started at the i-th image frame and whether the first action is determined to be implemented at the j-th image frame.
[0256] In one possible implementation, the acquisition unit is further configured to:
[0257] The length of the motion trajectory from the first position to the second position is determined, and the difference between the acquisition time of the i-th image frame and the j-th image frame is determined as the motion duration;
[0258] The speed parameters are determined based on the length of the motion trajectory and the duration of the motion.
[0259] As can be seen from the above technical solution, for the detection scenario of the target object, the corresponding video stream is acquired. If the j-th image frame included in the video stream is detected, it is determined that the target object has performed a first action. Then, the velocity parameter corresponding to the target object performing the first action is acquired. The velocity parameter is used to indicate the target part of the target object performing the first action and the speed of movement from the first position to the second position. Typically, after performing a certain action, the target object may continue to perform that action, which may lead to safety risks. However, since the target object itself cannot judge the potential safety risks, a safety risk prediction is made for the target object based on the velocity parameter, the second position, and the movement trajectory from the first position to the second position. Because the velocity parameter reflects the movement of the target object performing the first action, it can indicate whether the target object will face safety risks if it continues to move at this speed. Similarly, the movement trajectory reflects the trend of the first action, i.e., it can indicate which direction the target object might continue to move in, and therefore, it can also indicate whether the target object will face safety risks if it continues to move. Therefore, the predicted risk result can be used to indicate whether there is a safety risk for the target object when it continues to move along the predicted trajectory with the velocity parameter. Correspondingly, if the risk prediction results indicate that there is a safety risk to the target object, a safety warning can be issued to remind relevant personnel to take action and protect the target object.
[0260] As can be seen, this application combines the target object's own movement speed and the trend of its first action to determine whether the target object will face safety risks as it continues to move, and then issues an early warning when risks are detected, thereby improving the safety protection of the target object. Furthermore, the prediction is based on the current state of the target object's first action, making it a targeted prediction for that specific target object. Since different target objects may have different speed parameters and movement trajectories for the same first action, combining the current state of the target object's first action can more accurately predict the situations the target object may face, thus enabling more accurate safety risk prediction and early warning. In practical applications, video streams can be acquired by information acquisition devices; therefore, compared to manual review, automatic prediction and early warning based on video streams are more conducive to timely detection of potential safety issues, reducing oversights, and better protecting the target object.
[0261] This application also provides a computer device, which can be a terminal, taking a smartphone as an example:
[0262] Figure 11The diagram shown is a block diagram of a portion of the structure of a smartphone provided in an embodiment of this application. (Reference) Figure 11 The smartphone includes components such as a radio frequency (RF) circuit 1110, a memory 1120, an input unit 1130, a display unit 1140, a sensor 1150, an audio circuit 1160, a Wi-Fi module 1170, a processor 1180, and a power supply 1190. The input unit 1130 may include a touch panel 1131 and other input devices 1132, the display unit 1140 may include a display panel 1141, and the audio circuit 1160 may include a speaker 1161 and a microphone 1162. Those skilled in the art will understand that... Figure 11 The smartphone structure shown does not constitute a limitation on smartphones and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0263] The memory 1120 can be used to store software programs and modules. The processor 1180 executes various functions and data processing of the smartphone by running the software programs and modules stored in the memory 1120. The memory 1120 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the smartphone (such as audio data, phonebook, etc.). In addition, the memory 1120 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0264] The processor 1180 is the control center of the smartphone, connecting various parts of the smartphone via various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 1120 and calling data stored in the memory 1120. Optionally, the processor 1180 may include one or more processing units; preferably, the processor 1180 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1180.
[0265] In this embodiment, the steps performed by the processor 1180 in the smartphone can be based on Figure 11 The structure shown is implemented.
[0266] The computer device provided in this application embodiment can also be a server. Please refer to [link / reference]. Figure 12 As shown, Figure 12 This is a structural diagram of the server 1200 provided in this application embodiment. The server 1200 can vary significantly due to different configurations or performance. It may include one or more processors, such as a central processing unit (CPU) 1222, and a memory 1232, and one or more storage media 1230 (e.g., one or more mass storage devices) for storing application programs 1242 or data 1244. The memory 1232 and storage media 1230 can be temporary or persistent storage. The program stored in the storage media 1230 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the CPU 1222 may be configured to communicate with the storage media 1230 and execute the series of instruction operations in the storage media 1230 on the server 1200.
[0267] Server 1200 may also include one or more power supplies 1226, one or more wired or wireless network interfaces 1250, one or more input / output interfaces 1258, and / or one or more operating systems 1241, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.
[0268] In this embodiment, the central processing unit 1222 in server 1200 can perform the following steps:
[0269] Obtain the video stream corresponding to the detection scene targeting the object;
[0270] If the first action is determined when the j-th image frame included in the video stream is detected, the target object included in the detection scene is obtained, and the speed parameter corresponding to the target object performing the first action is obtained. The speed parameter is used to indicate the movement speed of the target part from the first position to the second position. The target part is the part of the target object that performs the first action. The first position refers to the position of the target part in the detection scene at the acquisition time of the i-th image frame. The i-th image frame is the image frame included in the video stream used to determine the start of the first action. The second position refers to the position of the target part in the detection scene at the acquisition time of the j-th image frame. The acquisition time of the i-th image frame is earlier than the acquisition time of the j-th image frame.
[0271] Based on the speed parameter, the second position, and the motion trajectory from the first position to the second position, a safety risk prediction is performed on the target object to obtain a risk prediction result. The risk prediction result is used to indicate whether there is a safety risk to the target object when it continues to move on the predicted trajectory with the speed parameter. The predicted trajectory refers to the trajectory of the target part continuing to perform the first action from the second position.
[0272] If the risk prediction results indicate that the target object poses a security risk, a security warning will be issued.
[0273] According to one aspect of this application, a computer-readable storage medium is provided for storing a computer program that, when executed by a computer device, causes the computer device to perform the security warning method described in the foregoing embodiments.
[0274] According to one aspect of this application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in various optional implementations of the above embodiments.
[0275] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.
[0276] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0277] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0278] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0279] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0280] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0281] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0282] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A safety early warning method, characterized in that, The method includes: Obtain the video stream corresponding to the detection scene targeting the object; If the first action is determined when the j-th image frame included in the video stream is detected, the target object included in the detection scene is obtained, and the speed parameter corresponding to the target object performing the first action is obtained. The speed parameter is used to indicate the movement speed of the target part from the first position to the second position. The target part is the part of the target object that performs the first action. The first position refers to the position of the target part in the detection scene at the acquisition time of the i-th image frame. The i-th image frame is the image frame included in the video stream used to determine the start of the first action. The second position refers to the position of the target part in the detection scene at the acquisition time of the j-th image frame. The acquisition time of the i-th image frame is earlier than the acquisition time of the j-th image frame. Based on the speed parameter, the second position, and the motion trajectory from the first position to the second position, a safety risk prediction is performed on the target object to obtain a risk prediction result. The risk prediction result is used to indicate whether there is a safety risk to the target object when it continues to move on the predicted trajectory with the speed parameter. The predicted trajectory refers to the trajectory of the target part continuing to perform the first action from the second position. If the risk prediction results indicate that the target object poses a security risk, a security warning will be issued.
2. The method according to claim 1, characterized in that, If the first action is a position transformation action, the position transformation action is used to change the position of the target object in the detection scene. The step of predicting the safety risk of the target object based on the velocity parameter, the second position, and the trajectory of the movement from the first position to the second position, to obtain a risk prediction result, includes: Determine the safe zone corresponding to the target object in the detection scenario; Based on the speed parameter, the second position, the trajectory of movement from the first position to the second position, and the safe area, a safety risk prediction is performed on the target object to obtain the risk prediction result. The risk prediction result is used to indicate whether the position of the target object in the detection scenario exceeds the safe area within a preset time period when it continues to move on the predicted trajectory at the speed parameter.
3. The method according to claim 1, characterized in that, The method further includes at least one of the following: If the speed parameter exceeds the motion speed threshold corresponding to the first action, the risk prediction result indicates that the target object has a safety risk. If there are obstacles in the detection scene corresponding to the predicted trajectory, the risk prediction result indicates that the target object has a safety risk. If the target location has a target object, and the target object will cause damage to the target object when it continues to move on the predicted trajectory with the speed parameter, the risk prediction result indicates that there is a safety risk to the target object; If the target object continues to move on the predicted trajectory at the speed parameter, and its position in the detection scenario exceeds the safe zone within a preset time period, the risk prediction result indicates that the target object faces a safety risk. The safe zone refers to the safe activity range of the target object in the detection scenario.
4. The method according to claim 2 or 3, characterized in that, The method further includes: The safe zone is determined based on the characteristics of the detection scene and the characteristics of the target object.
5. The method according to claim 1, characterized in that, The method further includes: Based on the j-th image frame, the pose parameters corresponding to the target object are determined, and the pose parameters are used to indicate the pose of the target object at the acquisition time of the j-th image frame; The step of predicting the safety risk of the target object based on the speed parameter, the second position, and the trajectory of movement from the first position to the second position, and obtaining the risk prediction result, includes: Based on the speed parameters, the attitude parameters, the second position, and the motion trajectory from the first position to the second position, a safety risk prediction is performed on the target object to obtain the risk prediction result.
6. The method according to claim 1, characterized in that, The method further includes: If it is determined from the video stream that the target object has multiple second actions during the detection process, a demand analysis is performed on the target object based on the multiple second actions to obtain the demand analysis results; Based on the requirements analysis results, requirements prompt information is generated, which is used to indicate the requirements of the target object indicated by the requirements analysis results.
7. The method according to claim 6, characterized in that, The step of performing a requirements analysis on the target object based on the multiple second actions to obtain the requirements analysis results includes: Obtain the facial color parameters of the target object during the detection process; Based on the various second actions and the facial color parameters, a requirements analysis is performed on the target object to obtain the requirements analysis results.
8. The method according to claim 6, characterized in that, If the target object is an infant or toddler, and if the plurality of second actions includes the action of the infant or toddler crying, the step of performing a requirements analysis on the target object based on the plurality of second actions to obtain the requirements analysis results includes: Obtain the crying parameters of the infant during the detection process; Based on the various second actions and the crying parameters, a needs analysis is performed on the infant to obtain the needs analysis results.
9. The method according to claim 1, characterized in that, The method further includes: Frame difference processing is performed on the k-th image frame and the (k+1)-th image frame included in the video stream to obtain a frame difference image. The frame difference image is used to indicate the moving target whose position changes in the detection scene during the process from the k-th image frame to the (k+1)-th image frame. If the moving target is determined to include the target object, the target object is subjected to action recognition based on the obtained frame difference image to obtain an action recognition result. The action recognition result is used to indicate whether the first action was started at the i-th image frame and whether the first action was determined to have been performed at the j-th image frame.
10. The method according to claim 1, characterized in that, The method further includes: Each image frame included in the video stream is input into the target object pose recognition model, and the target object pose recognition model outputs a pose recognition image corresponding to each image frame. The pose recognition image is used to indicate the pose of the target object included in the image frame. Based on the obtained posture recognition image, the target object is subjected to action recognition according to the acquisition time corresponding to the obtained posture recognition image, and the action recognition result is used to indicate whether the first action is started at the i-th image frame and whether the first action is determined to be implemented at the j-th image frame.
11. The method according to claim 1, characterized in that, The step of obtaining the speed parameters corresponding to the target object performing the first action includes: The length of the motion trajectory from the first position to the second position is determined, and the difference between the acquisition time of the i-th image frame and the j-th image frame is determined as the motion duration; The speed parameters are determined based on the length of the motion trajectory and the duration of the motion.
12. A safety early warning device, characterized in that, The device includes an acquisition unit, a prediction unit, and an early warning unit: The acquisition unit is used to acquire the video stream corresponding to the detection scene of the target object; The acquisition unit is further configured to determine that a target object included in the detection scene has performed a first action when the j-th image frame included in the video stream is detected, and to acquire a speed parameter corresponding to the target object performing the first action. The speed parameter is used to indicate the movement speed of the target part from a first position to a second position. The target part is the part of the target object that performs the first action. The first position refers to the position of the target part in the detection scene at the acquisition time of the i-th image frame. The i-th image frame is an image frame included in the video stream used to determine the start of the first action. The second position refers to the position of the target part in the detection scene at the acquisition time of the j-th image frame. The acquisition time of the i-th image frame is earlier than the acquisition time of the j-th image frame. The prediction unit is used to predict the safety risk of the target object based on the speed parameter, the second position, and the movement trajectory from the first position to the second position, and to obtain a risk prediction result. The risk prediction result is used to indicate whether there is a safety risk to the target object when it continues to move on the predicted trajectory with the speed parameter. The predicted trajectory refers to the trajectory of the target part continuing to perform the first action from the second position. The early warning unit is used to issue a safety warning if the risk prediction result indicates that the target object has a safety risk.
13. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store computer programs and to transfer the computer programs to the processor; The processor is configured to execute the method according to any one of claims 1-11 according to instructions in the computer program.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which, when executed by a computer device, causes the computer device to perform the method according to any one of claims 1-11.
15. A computer program product, comprising a computer program, characterized in that, When it is run on a computer device, it causes the computer device to perform the method according to any one of claims 1-11.