Object state recognition method and apparatus, computer device, and storage medium

By acquiring the background and target frames of the target video for image segmentation and generating interactive information groups, the problem of low accuracy in object state recognition in IoT technology is solved, and higher recognition accuracy is achieved.

WO2026011984A1PCT designated stage Publication Date: 2026-01-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/096755
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-11
Filing Date
2025-05-23
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing IoT technologies suffer from low accuracy in identifying object states, especially when it is difficult to distinguish object states with similar kinematic signals.

Method used

By acquiring the background and target frames of the target video, image segmentation is performed to generate an interactive information set, and the state of the target object is determined based on the interactive information set. Specific steps include acquiring the target video, determining the background and target frames, performing image segmentation, generating the interactive information set, and using an interactive model to identify the state of the target object.

Benefits of technology

It improves the accuracy of object state recognition, effectively distinguishes object states that are easily confused due to the same action, and enhances the accuracy of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025096755_15012026_PF_FP_ABST
    Figure CN2025096755_15012026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to an object state recognition method and apparatus, a computer device, a storage medium, and a computer program product. The method comprises: acquiring a target video (step S202); acquiring a background frame corresponding to the target video, and determining, from the target video, a target frame corresponding to a target action of a target object (step S204); performing image segmentation on the background frame and the target frame to obtain object regions respectively corresponding to objects in the background frame and a target region corresponding to the target object in the target frame (step S206); generating interaction information groups on the basis of position information of the target region and the object regions, the objects corresponding to the object regions, and the target action, so as to obtain an interaction information group between the target object and each object (step S208); and on the basis of each interaction information group, determining a target object state of the target object corresponding to the target frame (step S210).
Need to check novelty before this filing date? Find Prior Art

Description

Object state recognition method, apparatus, computer equipment and storage medium

[0001] This application claims priority to Chinese Patent Application No. 2024109290569, filed on July 11, 2024, entitled “Method, Apparatus, Computer Equipment and Storage Medium for Object Status Recognition”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer technology, and in particular to an object state recognition method, apparatus, computer equipment, storage medium, and computer program product. Background Technology

[0003] With the development of computer technology, the Internet of Things (IoT) technology has emerged. The IoT refers to connecting any object to a network through information sensing devices and according to agreed protocols. Objects exchange and communicate information through information transmission media to achieve functions such as intelligent identification, positioning, tracking, and monitoring.

[0004] Currently, user status can be identified through IoT technology. This is mainly achieved by using sensors such as accelerometers, gyroscopes, and magnetometers in wearable devices to measure the device's motion changes and direction of movement in space, thereby determining the wearer's posture and direction of movement, and ultimately identifying the wearer's object status. However, some object states exhibit similarities in their kinematic signal representations, making them difficult to distinguish, and current identification methods suffer from low accuracy. Summary of the Invention

[0005] This application provides an object state recognition method, apparatus, computer device, computer-readable storage medium, and computer program product.

[0006] This application provides an object state recognition method, executed by a computer device, including:

[0007] Acquire the target video;

[0008] Obtain the background frame corresponding to the target video, and determine the target frame corresponding to the target action of the target object from the target video;

[0009] Image segmentation is performed on the background frame and the target frame to obtain the object regions corresponding to each object in the background frame and the target regions corresponding to the target objects in the target frames.

[0010] Based on the location information of the target region and the object region, the object corresponding to the object region, and the target action, an interaction information group is generated to obtain the interaction information groups between the target object and each of the objects respectively.

[0011] Based on each group of interactive information, the target object state corresponding to the target frame is determined.

[0012] This application also provides an object state recognition device, including:

[0013] The video acquisition module is used to acquire the target video;

[0014] An image determination module is used to acquire the background frame corresponding to the target video and determine the target frame corresponding to the target action of the target object from the target video;

[0015] An image segmentation module is used to segment the background frame and the target frame to obtain the object regions corresponding to each object in the background frame and the target regions corresponding to the target objects in the target frame.

[0016] The information group generation module is used to generate interactive information groups based on the location information of the target area and the object area, the object corresponding to the object area, and the target action, so as to obtain the interactive information groups between the target object and each of the objects respectively;

[0017] The state recognition module is used to determine the state of the target object in the target frame based on each group of interactive information.

[0018] This application also provides a computer device, including a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps described in the object state recognition method above.

[0019] This application also provides a computer-readable storage medium storing computer-readable instructions thereon, which, when executed by a processor, implement the steps described in the object state recognition method above.

[0020] This application also provides a computer program product, including computer-readable instructions that, when executed by a processor, implement the steps described in the object state recognition method above.

[0021] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features, objects, and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 is an application environment diagram of an object state recognition method in one embodiment;

[0024] Figure 2 is a flowchart illustrating an object state recognition method in one embodiment;

[0025] Figure 3 is a schematic diagram of determining the background frame from the target video in one embodiment;

[0026] Figure 4 is a schematic diagram of the process of determining the target frame from the target video in one embodiment;

[0027] Figure 5 is a schematic diagram of identifying the target frame in one embodiment;

[0028] Figure 6 is a schematic diagram of interactive analysis based on image segmentation results in one embodiment;

[0029] Figure 7 is a flowchart illustrating the process of determining the state of a target object from the state of a reference object in one embodiment;

[0030] Figure 8 is a flowchart illustrating the process of determining the state of a target object through an interaction model in one embodiment;

[0031] Figure 9 is a schematic diagram of a fall detection method in one embodiment;

[0032] Figure 10 is a schematic diagram of a fall detection method in another embodiment;

[0033] Figure 11 is a structural block diagram of an object state recognition device in one embodiment;

[0034] Figure 12 is an internal structure diagram of a computer device in one embodiment;

[0035] Figure 13 is an internal structural diagram of a computer device in another embodiment. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0037] The object state recognition method provided in this application embodiment can be applied to the application environment shown in Figure 1. The terminal 102 communicates with the server 104 via a network. A data storage system can store the data that the server 104 needs to process. The data storage system can be set up independently, integrated into the server 104, or placed in the cloud or on other devices. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented using a standalone server, a server cluster composed of multiple servers, or a cloud server.

[0038] It is understood that both the terminal and the server can be used independently to execute the object state recognition method provided in the embodiments of this application. The terminal and the server can also be used collaboratively to execute the object state recognition method provided in the embodiments of this application.

[0039] For example, the server acquires the target video and the corresponding background frame, and determines the target frame corresponding to the target action of the target object from the target video. The server performs image segmentation on the background frame and the target frame, obtaining the object region corresponding to each object in the background frame and the target region corresponding to the target object in the target frame. Based on the position information of the target region and the object region, the object corresponding to the object region, and the target action, the server generates interaction information groups, obtaining the interaction information groups between the target object and each object, and based on each interaction information group, determines the target object state corresponding to the target frame.

[0040] In one embodiment, as shown in Figure 2, an object state recognition method is provided. Taking the application of this method to a computer device as an example, the computer device can be a terminal or a server. It is understood that this method can be executed independently by the terminal or server, or it can be implemented through interaction between the terminal and the server. Wherein:

[0041] Step S202: Obtain the target video.

[0042] Here, the target video refers to the video to be analyzed. It can be understood that the target video can be a pre-captured video, such as a security recording used for a demonstration. Alternatively, the target video can be a video captured in real-time, such as a real-time monitoring video.

[0043] Specifically, computer devices can acquire target videos locally or from other devices, analyze the target videos, and identify the object state of target objects in the target videos.

[0044] Step S204: Obtain the background frame corresponding to the target video, and determine the target frame corresponding to the target action of the target object from the target video.

[0045] In this context, the background frame is a background image used to represent the static or relatively stable background portion of a scene. Background frames primarily provide a sense of space and atmosphere, helping to identify the environment in which the scene is situated. The background frame corresponding to the target video is an image representing the background of the scene within the target video. It can be understood that the background frame corresponding to the target video can be a pre-captured background image. For example, if the target video is a recording of room A, an image of room A without any living creatures, captured by the same camera before recording begins, can be used as the background frame for the target video. Alternatively, the background frame can be a background image determined from the target video itself; for example, a video frame without living creatures can be obtained from the target video as the background frame.

[0046] A target frame is an image representing a target object in a target action. In the target frame, the target object is performing the target action. The target object is the object to be analyzed, and its specific configuration can be set as needed. The target object is typically a living being; for example, it could be a person, an animal, etc. The target action is the action to be analyzed, and its specific configuration can be set as needed. For example, the target action could be a falling action, an eating action, etc. For instance, if the target object is a person and the target action is a falling action, the person could be depicted lying down in the target frame.

[0047] It is understandable that the background frame does not include the target object, while the target frame does include the target object.

[0048] Specifically, after determining the target video, the computer device can acquire the background frame corresponding to the target video, determine the target frame corresponding to the target action of the target object from the target video, and identify the target object state corresponding to the target frame based on the background frame and the target frame.

[0049] For the target frame, the computer device can perform target object recognition on the video frames in the target video, further perform target action recognition on the video frames where the target object is identified, and use the video frames where the target action is identified as the target frame.

[0050] In one embodiment, target object recognition is performed on video frames of the target video, and video frames in which the target object is identified are used as candidate frames; each candidate frame within the target time period is obtained to form a candidate frame sequence; the candidate frame sequence is input into the action recognition model to obtain the predicted action corresponding to the candidate frame sequence; when the predicted action is the target action, the target frame corresponding to the target action of the target object is determined from the candidate frame sequence.

[0051] Step S206: Perform image segmentation on the background frame and the target frame to obtain the object regions corresponding to each object in the background frame and the target regions corresponding to the target objects in the target frame.

[0052] Image segmentation refers to dividing an image into several different regions, where pixels within the same region have similar features, and the features of different regions are significantly different.

[0053] Image segmentation is performed on the background frame to obtain the object regions corresponding to each object in the background frame. The object region refers to the area where the object is located in the image, used to represent the object's position. For example, if the background frame is an image of room A, image segmentation of the background frame can obtain the regions where each facility in room A is located, allowing us to understand the distribution of facilities in room A. Image segmentation is performed on the target frame to obtain the target region corresponding to the target object in the target frame. The target region refers to the area where the target object is located in the image, used to represent the target object's position.

[0054] It is understandable that the same image segmentation algorithm can be used for the background frame and the target frame, or different image segmentation algorithms can be used.

[0055] Specifically, computer equipment can perform image segmentation on a background frame to obtain the object regions corresponding to each object in the background frame, and can perform image segmentation on a target frame to obtain the target region corresponding to the target object in the target frame. For example, the background frame can be input into an image segmentation model, and the image segmentation model can output the image segmentation result of the background frame. The image segmentation result can be that the object is outlined in the background frame using a bounding box. The bounding box can be a regular border or an irregular border, and the image region outlined by the bounding box corresponding to the object is the object region corresponding to the object. As another example, the target frame can be input into a target object recognition model, and the target object recognition model can output the target region corresponding to the target object in the target frame. The target object recognition model only cares about the region where the target object is located.

[0056] Step S208: Based on the position information of the target area and the object area, the object corresponding to the object area, and the target action, generate an interaction information group to obtain the interaction information groups between the target object and each object.

[0057] Each object corresponds to an interaction information group. The interaction information group includes relevant information about the object in the background frame and relevant information about the target object in the target frame, used to analyze the interaction relationship between the target object and the object. The object information includes at least one of the following: the object identifier corresponding to the object, and the position information of the object region. The object identifier is used to identify the object; for example, the object identifier can be the object name, object number, etc. The position information of the object region indicates the location of the object region in the image. The target object information includes at least one of the following: the action identifier corresponding to the target object's target action, the position information of the target region, and the object identifier corresponding to the target object. The action identifier is used to identify the action; for example, the action identifier can be the action name, action number, etc. The position information of the target region indicates the location of the target region in the image. The object identifier is used to identify the object; for example, the object identifier can be the object number.

[0058] Specifically, for any object in the background frame, the computer device can establish an interaction information group based on the positional information of the object's corresponding object region, the positional information of the object and the target region, and the target action. For example, the interaction information group may include the overlap between the object region and the target region, the object's identifier, and the target action's action identifier; or it may include the positional information of the object region, the positional information of the target region, the object's identifier, and the target action's action identifier; and so on. Ultimately, the computer device can obtain the interaction information groups between the target object in the target frame and each object in the background frame.

[0059] Step S210: Based on each group of interactive information, determine the target object state corresponding to the target frame.

[0060] In this context, object state refers to the state an object is in due to the execution of an action. It can be understood that the same behavior can correspond to different object states in different situations. For example, when a person performs the action of falling down, they may be in a resting state or in a fallen state. When the falling action overlaps significantly with furniture such as a bed, sofa, or recliner, the person is usually in a resting state; however, when the falling action overlaps significantly with the floor, the person is usually in a fallen state.

[0061] Specifically, computer devices can analyze the interaction relationships between target objects and objects based on various interaction information groups to determine the target object's state in the target frame. For example, the target frame and interaction information groups are input into a trained interaction model, which outputs the target object's state in the target frame. The trained interaction model is obtained through supervised training of the interaction model to be trained, based on training frames, training information groups, and training labels corresponding to the training frames. Another example is determining the overlap between the target region and each object region based on each interaction information group. Objects corresponding to object regions with overlap greater than a threshold are identified as target objects. The target object's state in the target frame is then determined based on the target object and its actions.

[0062] In the above object state recognition method, the following steps are taken: First, a target video is acquired. Then, the background frame corresponding to the target video is acquired, and the target frame corresponding to the target object's action is determined from the target video. Image segmentation is performed on the background frame and the target frame to obtain the object regions corresponding to each object in the background frame and the target region corresponding to the target object in the target frame. Based on the positional information of the target region and object region, the object corresponding to the object region, and the target action, interaction information groups are generated to obtain the interaction information groups between the target object and each object. Based on each interaction information group, the target object state corresponding to the target frame is determined. In this way, the target action of the target object in the target frame is analyzed based on the background frame and the target frame to identify the target object state corresponding to the target frame. The background frame can provide information about the object, and the target frame can provide information about the target object. Establishing interaction information groups based on the object information and the target object information, and analyzing the interaction relationship between the target object and the object through these interaction information groups, can effectively improve the accuracy of object state recognition. Considering the interaction relationship between the target object and the environment when recognizing object state can effectively distinguish object states that are easily confused due to the same action, thus improving the accuracy of object state recognition.

[0063] In one embodiment, obtaining the background frame corresponding to the target video includes:

[0064] Target object recognition is performed on video frames in the target video. If the target object is not recognized after a certain number of consecutive video frames, the last video frame in the consecutive target video frames is taken as the background frame.

[0065] Target object recognition refers to identifying whether a target object exists in an image. For example, if the target object is a person, various human detection algorithms can be used for human identification. It can be understood that the number of targets can be set according to actual needs. The number of targets can be a pre-set number, or it can be a number that changes flexibly depending on the scene.

[0066] Specifically, computer equipment can determine background frames from target videos to improve their accuracy. The equipment can sequentially perform target object recognition on video frames within the target video to determine if a target object exists within each frame. The equipment can count the number of consecutive video frames in which no target object is detected. If no target object is detected for a certain number of consecutive video frames, the last video frame from those consecutive frames is selected as the background frame, ensuring that the background frame is the most up-to-date background image.

[0067] It is understandable that after determining the background frame, the computer device can clear the statistical results and start counting the number of consecutive video frames in which the target object is not detected again. When it is found that no target object is detected in a consecutive number of video frames, the last video frame is taken from the newly found consecutive number of video frames as the new background frame, so as to further ensure that the background frame is always the latest background image.

[0068] In a specific application, referring to Figure 3, human detection is performed frame by frame on the target video to determine whether a human body (i.e., N) has been detected. object (Whether it equals 0), count the number of consecutive frames N where no human body is detected. empty When T consecutive e If no target object is detected in any frame, then the current frame f will be... i The image is updated to the background frame f bg When performing human detection, object detection models such as YOLO series object detectors, SSD (Single Shot MultiBox Detector), CornerNet, and CenterNet can be used, but are not limited to those that can be employed.

[0069] In the above embodiments, target object identification is performed on the video frames in the target video. If the target object is not identified after a certain number of consecutive video frames, the last video frame from the consecutive number of video frames is taken as the background frame. This ensures that the latest and most accurate video frame representing the background in the target video is used as the background frame, which helps to improve the accuracy of object state identification.

[0070] In one embodiment, the number of targets and the recognition accuracy of the target object recognition algorithm are negatively correlated.

[0071] Among them, the target object recognition algorithm refers to the algorithm used to identify target objects. The recognition accuracy of the target object recognition algorithm refers to the accuracy of the recognition of the target object.

[0072] Specifically, the number of targets can be dynamically changed, specifically determined by the target object recognition algorithm. The number of targets is related to the accuracy of the target object recognition algorithm; the higher the accuracy, the smaller the number of targets. It can be understood that if no target object is identified in a consecutive number of video frames, taking the last frame as the background frame can effectively prevent the background frame from containing the target object. If the accuracy of the target object recognition algorithm is poor, the accuracy of the background frame can be improved by increasing the number of targets. When the accuracy of the target object recognition algorithm is poor, it is necessary to identify no target object in more consecutive video frames before determining a background frame.

[0073] In one embodiment, the target object identification algorithm is determined based on the protection priority of the target object. The higher the protection priority of the target object, the more accurate the target object identification algorithm is used. It can be understood that the higher the importance of the target object, the higher its protection priority. For example, the protection priority of a person is higher than that of a cat or dog. Similarly, the older a person is, the higher their protection priority; and the protection priority of a patient is higher than that of an ordinary person.

[0074] Furthermore, object identification can be performed first using a general object recognition algorithm, with the first identified live object serving as the target object. Then, a target object recognition algorithm can be determined based on the protection priority of the target object, and subsequent video frames can be used to identify the target object, thereby improving the accuracy of target object identification. For example, a liveness detection algorithm can be used first to identify whether a live object exists in a video frame. If a human body is identified in the video frame first, a human body recognition algorithm can then be used to accurately identify whether a human body exists in subsequent video frames, thus improving the accuracy of human body recognition.

[0075] In the above embodiments, the number of targets and the recognition accuracy of the target object recognition algorithm are negatively correlated. When the recognition accuracy of the target object recognition algorithm is poor, it is necessary to continuously identify more video frames without target objects before determining the background frame, which can ensure the accuracy of the background frame.

[0076] In one embodiment, as shown in Figure 4, determining the target frame corresponding to the target action of the target object from the target video includes:

[0077] Step S402: Perform target object recognition on the video frames of the target video, and use the video frames with the identified target objects as candidate frames.

[0078] Specifically, computer equipment can perform target object recognition on video frames of a target video, use video frames that identify the target object as candidate frames, and then further perform target action recognition on the candidate frames, using the candidate frames that identify the target action as the target frames.

[0079] In one embodiment, in a home monitoring or caregiving scenario, target object identification is performed on video frames of the target video, and video frames that identify a single target object are selected as candidate frames. It is understood that in a home monitoring or caregiving scenario, the target object is usually a person. If there is more than one target object, and someone experiences an abnormal condition, other people present can provide assistance without the need for automatic alarms.

[0080] Step S404: Obtain the candidate frames within the target time period to form a candidate frame sequence.

[0081] The execution of an action by a target object causes a change in its posture. Therefore, by detecting these posture changes, the corresponding action can be identified. Since posture changes are time-dependent, multiple video frames within a certain time frame can be analyzed to identify whether the target action has occurred. For example, falling down represents a change from a standing to a lying position, and sitting down represents a change from a standing to a sitting position. The target time period is the time segment to be analyzed. The target time period can be set as needed. The length of the target time period is usually greater than or equal to the execution time of an action. For example, the target time period could be 1 second, and the target object's performance of the target action can be determined from candidate frames spanning 1 second. The candidate frame sequence includes multiple candidate frames within the target time period.

[0082] Specifically, computer equipment can assemble candidate frames within a target time period into a candidate frame sequence in chronological order, and then perform target action recognition on the candidate frame sequence to determine the target frame from the candidate frame sequence containing the target action.

[0083] Step S406: Target object skeleton point recognition is performed on the candidate frames in the candidate frame sequence to obtain the skeleton point data corresponding to the candidate frames, forming a skeleton point data sequence corresponding to the candidate frame sequence.

[0084] Target object skeleton point recognition refers to identifying the skeleton points of a target object. For example, if the target object is a person, various human skeleton point recognition algorithms can be used for human skeleton point recognition. Target object skeleton point recognition is performed on a video frame to obtain the skeleton point data corresponding to that video frame. The skeleton point data sequence corresponding to the candidate frame sequence includes the skeleton point data corresponding to each candidate frame in the candidate frame sequence. The skeleton point data includes the type and location of the skeleton points.

[0085] Specifically, for any candidate frame in the candidate frame sequence, the computer device can perform target object skeleton point recognition on the candidate frame to obtain the skeleton point data corresponding to the candidate frame. The computer device can obtain the skeleton point data corresponding to each candidate frame in the candidate frame sequence and combine the skeleton point data corresponding to each candidate frame to form a skeleton point data sequence corresponding to the candidate frame sequence. For example, the skeleton point data can be directly combined to obtain the skeleton point data sequence, or the skeleton point data of each type can be arranged in chronological order to obtain the time-series data of skeleton points corresponding to each type. The time corresponding to the skeleton point data can be the time corresponding to the candidate frame to which the skeleton point data belongs. The skeleton point data sequence can reflect the action characteristics of the candidate frame sequence, which helps to identify the action corresponding to the candidate frame sequence.

[0086] Step S408: Input the skeleton point data sequence into the multi-frame action recognition model to obtain the predicted action corresponding to the candidate frame sequence.

[0087] The multi-frame action recognition model is an artificial intelligence model used to identify actions based on multiple frames of images. The input data for the multi-frame action recognition model includes a sequence of skeleton points corresponding to the image sequence, and the output data includes predicted actions. The predicted actions are the actions identified by the model.

[0088] Specifically, computer devices can use artificial intelligence models to identify target actions. The computer device acquires a multi-frame action recognition model, inputs the skeleton point data sequence corresponding to the candidate frame sequence into the multi-frame action recognition model, the multi-frame action recognition model processes the input data, and outputs the predicted action corresponding to the candidate frame sequence.

[0089] It is understandable that a multi-frame action recognition model can be obtained through supervised training based on the skeleton point data sequence corresponding to the training image sequence and the action training labels. The training image sequence is an image sequence of known actions. The action training labels corresponding to the training image sequence are the known actions reflected in the training image sequence.

[0090] Step S410: When the predicted action is the target action, determine the target frame corresponding to the target action of the target object from the candidate frame sequence.

[0091] Specifically, when the predicted action output by the multi-frame action recognition model is the target action, the computer device can determine the target frame corresponding to the target action of the target object from the candidate frame sequence. For example, the last frame in the candidate frame sequence can be selected as the target frame. Another example is to randomly select a frame from the candidate frame sequence as the target frame.

[0092] In the above embodiments, target object recognition is performed on the video frames of the target video. The video frames in which the target object is identified are used as candidate frames. A candidate frame sequence is formed by acquiring all candidate frames within the target time period. Target action recognition is then performed on the candidate frame sequence, and the target frame is determined from the candidate frame sequence in which the target action is identified. This comprehensive consideration of candidate frames within a certain time period to identify whether a target action occurs effectively improves the accuracy of target action recognition. Target object skeleton point recognition is performed on the candidate frames in the candidate frame sequence to obtain the skeleton point data corresponding to the candidate frames. This skeleton point data sequence is formed and input into a multi-frame action recognition model to obtain the predicted action corresponding to the candidate frame sequence. When the predicted action is the target action, the target frame corresponding to the target action of the target object is determined from the candidate frame sequence. Thus, the skeleton point data sequence corresponding to the candidate frame sequence is extracted first. The skeleton point data sequence reflects action features. The action corresponding to the candidate frame sequence is identified by processing the skeleton point data sequence through a multi-frame action recognition model. The multi-frame action recognition model has powerful data processing capabilities, which can further improve the accuracy of target action recognition.

[0093] In one embodiment, determining the target frame corresponding to the target action of the target object from the candidate frame sequence includes:

[0094] Each candidate frame in the candidate frame sequence is input into the trained single-frame action recognition model to obtain the action recognition result corresponding to each candidate frame. The action recognition result includes the number of target objects and actions in the candidate frames. The action recognition result includes candidate frames of a single target object and target action, which are the target frames corresponding to the target actions of the target objects.

[0095] The single-frame action recognition model is an artificial intelligence model used to identify actions based on a single frame of an image. The input data for the single-frame action recognition model includes a single-frame image, and the output data includes the number of target objects and their actions within the single-frame image.

[0096] Specifically, when performing target action recognition, the computer device can first perform pre-recognition using a multi-frame action recognition model to initially filter out video frames that are obviously unrelated to the target action. When the target action is pre-recognized, it can then perform further recognition using a single-frame action recognition model to finally obtain the accurate target frame.

[0097] The computer device inputs the skeleton point data sequence corresponding to the candidate frame sequence into a multi-frame action recognition model to obtain the predicted action corresponding to the candidate frame sequence. It is understandable that, due to the influence of various factors such as computing resources, lighting, and angle, the initial screening results obtained through the multi-frame action recognition model may contain some misidentified actions, or even some data containing multiple target objects. Therefore, a single-frame action recognition model is further used for target action recognition. When the predicted action corresponding to the candidate frame sequence is a target action, each candidate frame in the candidate frame sequence is then input into the trained single-frame action recognition model to obtain the action recognition result corresponding to each candidate frame. The action recognition result corresponding to the candidate frame includes the number of target objects and their actions within the candidate frame. For example, the action recognition result could be the location of the target object within the candidate frame using a bounding box and the annotation of the target object's action. The computer device obtains candidate frames from the candidate frame sequence that contain action recognition results for a single target object and its target action, which are then used as the target frames corresponding to the target action of the target object.

[0098] In the above embodiments, the single-frame action recognition model can not only identify the number of target objects, but also identify target actions. It also achieves the functions of filtering multi-object scenes and avoiding misidentification of actions, thereby improving the recognition accuracy of effective target frames.

[0099] In a specific application, referring to Figure 5, video frames with a time span of 1 second are acquired and input into a multi-frame action recognition model in chronological order. If the multi-frame action recognition model outputs a target action, each video frame is then input into a single-frame action recognition model. If the multi-frame action recognition model outputs a target action, it can be considered that there are no target frames among these video frames. The single-frame action recognition model can identify the number of target objects and their actions within a video frame. If the single-frame action recognition model outputs a single target object and its action for a particular video frame, that video frame is considered the target frame.

[0100] In one embodiment, the object state recognition method further includes:

[0101] Obtain the first training image and the corresponding target object location training label and target object action training label; input the first training image into the single-frame action recognition model to be trained to obtain the target object location prediction label and target object action prediction label corresponding to the first training image; based on the differences between the target object location training label and the target object location prediction label, and the differences between the target object action training label and the target object action prediction label, adjust the model parameters of the single-frame action recognition model to be trained until the first convergence condition is met, and obtain the trained single-frame action recognition model.

[0102] The first training image is used to train the single-frame action recognition model. The target object position training label corresponding to the first training image indicates the accurate position information of the target object in the first training image. The target object action training label corresponding to the first training image indicates the accurate action information of the target object in the first training image. It can be understood that the first training image may contain no target object, only a single target object, or multiple target objects. For example, if the target object is a person, the first training image may contain no person, only one person, or multiple people; the multiple people in the first training image can be of the same type or different types. If the target object is a dog, the first training image may contain no dog, only one dog, or multiple dogs; and so on. Each target object in the first training image has a corresponding target object action training label. The target object action training label can be the target action or other actions. For example, the target object action training label can be sitting, standing, lying down, etc.

[0103] The first training image is input into the single-frame action recognition model to be trained, resulting in the target object location prediction label and the target object action prediction label corresponding to the first training image. The target object location prediction label indicates the predicted position information of the target object in the first training image. The target object action prediction label indicates the predicted action information of the target object in the first training image.

[0104] The first convergence criterion is used to determine whether a single-frame action recognition model has reached convergence. The first convergence criterion can be set according to actual needs. It can be at least one of the following: model loss is less than a loss threshold, the number of model iterations is greater than a threshold for the number of iterations, or the rate of change of model loss is less than a threshold for the rate of change of loss.

[0105] Specifically, the computer device can acquire the first training image and the target object location training label and target object action training label corresponding to the first training image, and perform supervised training on the single-frame action recognition model based on the first training image and the training label corresponding to the first training image.

[0106] The computer device can input the first training image into the single-frame action recognition model to be trained. After data processing by the single-frame action recognition model, the model outputs the predicted target object location label and the predicted target object action label corresponding to the first training image. Then, the computer device calculates the model loss based on the differences between the training and predicted target object location labels and the differences between the training and predicted target object action labels. The model loss is then backpropagated to adjust the model parameters of the single-frame action recognition model to be trained until the first convergence condition is met, resulting in the trained single-frame action recognition model.

[0107] It is understandable that there can be multiple first training images. Multiple first training images can be acquired for training in each round. In one round, the model sub-loss is calculated based on the differences between the target object location training label and the target object location prediction label, and the differences between the target object action training label and the target object action prediction label, corresponding to the same first training image. The model sub-loss for each first training image is obtained, and the model loss is derived based on these sub-losses. The single-frame action recognition model to be trained is then adjusted based on the model loss, and the adjusted single-frame action recognition model is used as the single-frame action recognition model to be trained in the next round.

[0108] In the above embodiments, the single-frame action recognition model is trained in a supervised manner based on the first training image and the target object location training label and the target object action training label corresponding to the first training image, which enables the trained single-frame action recognition model to have the ability to recognize target objects and actions.

[0109] In a specific application, referring to Figure 6, the process of determining the state of the target object is illustrated using the human body as the target object and the falling action as the target action. Background frames and target frames are determined from the target video. Image semantic segmentation is performed on both the background and target frames to obtain the object regions corresponding to each object in the background frame and the human body region corresponding to the human body in the target frame. Interaction information groups are established based on the object regions and human body regions. An interaction information group includes the action identifier of the falling action (i.e., P falling in Figure 6), the object identifier of the object in the background frame (i.e., floor O1 or sofa O2 in Figure 6), and the overlap area between the object region and the human body region (i.e., E1 or E2 in Figure 6). The target object state in the target frame is determined based on each interaction information group.

[0110] Image semantic segmentation can be performed using deep neural network models such as Light-Weight RefineNet and SAM (Segment Anything Model), but not limited to those used.

[0111] In one embodiment, as shown in Figure 7, determining the target object state in the target frame based on each group of interactive information includes:

[0112] Step S702: Obtain the state of each reference object corresponding to the target action, and obtain the reference object corresponding to each reference object state.

[0113] In this context, the target object can have at least one reference object state during the target action. A reference object state is an object state that the target object can be in during the target action. For example, for the action of falling down, the reference object state corresponding to the target object can be either a resting state or a fallen state; for the action of sitting down, the reference object state corresponding to the target object can be either a resting state or a social state; for the action of running, the reference object state corresponding to the target object can be either a danger-avoiding state or an exercise state; and so on.

[0114] Each state of a reference object has a corresponding reference object. A reference object is an object related to the state of the reference object. For example, if the target object is in a resting state, the reference object could be a bed, sofa, chair, etc.; if the target object is in a fallen state, the reference object could be a floor, steps, etc.; if the target object is in an exercise state, the reference object could be a treadmill, sandbag, etc.; and so on.

[0115] Specifically, the computer device can acquire the state of each reference object corresponding to the target action, and determine the target object state corresponding to the target frame from the state of each reference object.

[0116] Step S704: Take the reference object in the background frame as the target object, and obtain the reference overlap threshold corresponding to the state of the reference object corresponding to the target object as the target overlap threshold.

[0117] The reference object state has a corresponding reference overlap threshold. The reference overlap threshold refers to the condition that the overlap between the target object and the reference object in the reference object state must meet. If the overlap between the target object and the reference object during the target action is greater than the reference overlap threshold, then the target object is determined to be in the reference object state corresponding to that reference object during the target action. For example, for the action of falling down, if the overlap between the target object and objects such as beds, sofas, and chairs is greater than a certain threshold, then the target object is determined to be in a resting state; if the overlap between the target object and objects such as floors and steps is greater than a certain threshold, then the target object is determined to be in a falling state. For the action of running, if the overlap between the target object and objects such as treadmills and running tracks is greater than a certain threshold, then the target object is determined to be in an exercising state. The reference overlap threshold corresponding to the reference object state can be obtained by statistically analyzing a large number of known images where the target object is in the reference object state. For example, calculate the overlap between the target region corresponding to the target object in a known image and the object region corresponding to the reference object; calculate the average overlap of the regions corresponding to each known image as a reference overlap threshold; obtain the median overlap of the regions corresponding to each known image as a reference overlap threshold; and so on.

[0118] The target object refers to a reference object contained in the background frame. For example, if the target action is a falling action, the target object could be an object in the background frame such as a bed, sofa, chair, floor, or steps. The target overlap threshold refers to the reference overlap threshold corresponding to the state of the reference object corresponding to the target object.

[0119] Specifically, the computer device can obtain the reference objects corresponding to each reference object state, take the reference objects existing in the background frame as the target objects, and determine the target object state of the target object in the target frame from the reference object states corresponding to the target objects.

[0120] Step S706: For the interaction information group corresponding to the target object, when the overlap between the object region corresponding to the target object and the target region is greater than the target overlap threshold, the reference object state corresponding to the target object is taken as the target object state corresponding to the target object in the target frame.

[0121] Among them, the degree of overlap between the object region and the target region refers to the degree of overlap between the object region and the target region, which can be the overlapping area, overlapping ratio, etc.

[0122] Specifically, the computer device obtains the reference overlap threshold corresponding to the reference object state corresponding to the target object as the target overlap threshold, determines the region overlap between the object region corresponding to the target object and the target region corresponding to the target object based on the interaction information group corresponding to the target object, and if the region overlap is greater than the target overlap threshold, the reference object state corresponding to the target object is taken as the target object state corresponding to the target object in the target frame.

[0123] It is understandable that if there are multiple reference objects for a given reference object state, and the overlap between the target region and the object region corresponding to any one of the reference objects is greater than the corresponding reference overlap threshold, then the reference object state corresponding to the reference object will be taken as the target object state.

[0124] It is understandable that if the regional overlap is less than or equal to the target overlap threshold, the target object's state in the target frame can be further determined through other methods. For example, the target object's state in the target frame can be determined using a trained interaction model. The target frame and each interaction information group are input into the trained interaction model, which outputs the target object's state in the target frame. In this way, first determining whether the reference object's state is matched using relevant information about the reference object's state can improve the efficiency of object state recognition. If the reference object's state is not matched, the object state can then be identified through the interaction model, which can improve the accuracy of object state recognition and compensate for the insufficiency of the reference object's state. Another example is sending the target frame to a management terminal, which displays the target frame to the administrator so that the administrator can determine the target object's state in the target frame based on the target frame. If the regional overlap is less than or equal to the target overlap threshold, the target object's state in the target frame can also be determined to be an unknown state.

[0125] In the above embodiments, the states of each reference object corresponding to the target action are obtained. Each reference object state has a corresponding reference object and a reference overlap threshold. By determining the region overlap between the object region corresponding to the reference object in the background frame and the target region corresponding to the target object in the target frame, and comparing the region overlap with the corresponding reference overlap threshold, the target object state corresponding to the target frame can be quickly determined. Furthermore, each reference object state has its own corresponding reference overlap threshold, which helps to further improve the accuracy of object state recognition.

[0126] In one embodiment, as shown in Figure 8, determining the target object state in the target frame based on each group of interactive information includes:

[0127] Step S802: Input the target frame and each group of interaction information into the trained interaction model.

[0128] The interaction model is an artificial intelligence model used to determine the state of an object based on an image and a set of interaction information. The input data for the interaction model includes an image corresponding to a specific action of the object, and a set of interaction information determined based on that image and the corresponding background image. The output data includes the object state identified by the model.

[0129] Specifically, computer devices can identify the state of a target object using artificial intelligence models. The computer device acquires a trained interaction model, inputs the target frame corresponding to the target action of the target object, and the various interaction information groups determined based on the background frame and the target frame into the trained interaction model. After data processing by the interaction model, the interaction model outputs the state of the target object corresponding to the target frame.

[0130] Step S804: Determine the overlapping regions in the target frame based on each interactive information group.

[0131] Step S806: Extract image features of the overlapping region to obtain the first interaction feature; obtain the second interaction feature based on the object and target action in the interaction information group corresponding to the overlapping region; and extract image features of the target frame to obtain the third interaction feature.

[0132] Step S808: Based on the first interaction feature, the second interaction feature, and the third interaction feature, a comprehensive interaction feature is obtained.

[0133] Step S810: Output the target object state corresponding to the target frame based on the comprehensive interaction features.

[0134] The overlapping region refers to the portion that overlaps between the object region corresponding to an object in the background frame and the target region corresponding to the target object in the target frame. The first interaction feature is the image feature of the overlapping region, containing information about the overlapping region. The second interaction feature is determined based on the object and target actions in the interaction information group corresponding to the overlapping region, containing information about the object and target actions corresponding to the overlapping region. The third interaction feature is the image feature of the target frame, containing information about the target frame. The comprehensive interaction feature is a feature obtained based on the first, second, and third interaction features, containing the most important information among the three interaction features.

[0135] Specifically, the computer device inputs the target frame and various interaction information groups into a trained interaction model. In the interaction model, overlapping regions are determined in the target frame based on each interaction information group, resulting in at least one overlapping region. Image features of the overlapping regions are extracted to obtain the first interaction feature, and image features of the target frame are extracted to obtain the third interaction feature. For example, image features can be extracted using convolutional layers. In the interaction model, a second interaction feature is generated based on the object and target action in the interaction information group corresponding to the overlapping region. For example, features of the object and target action are extracted, and the extracted features are concatenated to obtain the second interaction feature. In the interaction model, a comprehensive interaction feature is obtained based on the first, second, and third interaction features. For example, the comprehensive interaction feature can be obtained by fusing the first, second, and third interaction features using fully connected layers; attention features can be applied to each pair of the first, second, and third interaction features to obtain multiple attention features, and these multiple attention features can be fused to obtain the comprehensive interaction feature. Attention mechanisms can be used to extract the parts that influence each other between features; and so on. In the interaction model, the target object state in the target frame is determined based on comprehensive interaction features. The interaction model ultimately outputs the target object state in the target frame.

[0136] In the above embodiments, the target frame and each interaction information group are input into the trained interaction model. The interaction model has powerful data processing capabilities and can quickly output accurate target object states. Furthermore, the interaction information groups can provide interaction information between the target object and the object, helping the interaction model understand the target frame. As a result, the target object states output by the interaction model have even higher accuracy.

[0137] In one embodiment, the object state recognition method further includes:

[0138] Obtain the second training image corresponding to the training action of the training object and the training object state and training information group corresponding to the second training image; the training object state is the object state of the training object corresponding to the second training image, and the training information group is determined based on the position information of the area where the training object is located in the second training image, the position information and object information of the area where the object is located in the background image corresponding to the second training image, and the training action; input the second training image and training information group into the interaction model to be trained to obtain the predicted object state corresponding to the second training image; based on the difference between the training object state and the predicted object state, adjust the model parameters of the interaction model to be trained until the second convergence condition is met to obtain the trained interaction model.

[0139] The second training image is used to train the interaction model. The second training image corresponds to the training action of the training object. The training object state corresponding to the second training image indicates the accurate object state of the training object in the corresponding training image. The training information group is an interaction information group determined based on the image segmentation results of the second training image and the image segmentation results of the background image corresponding to the second training image. The method for determining the training information group can refer to the aforementioned method for determining the interaction information group.

[0140] The second training image and the corresponding training information set are input into the trained interaction model to obtain the predicted object state corresponding to the second training image. The predicted object state is the object state of the training object predicted by the model corresponding to the second training image.

[0141] The second convergence condition is used to determine whether the interaction model has reached convergence. The second convergence condition can be set according to actual needs. It can be at least one of the following: model loss is less than a loss threshold, the number of model iterations is greater than a threshold for the number of iterations, or the rate of change of model loss is less than a threshold for the rate of change of loss.

[0142] It is understandable that the training object and the target object can be the same object or different objects. The training action and the target action can be the same action or different actions. The first training image and the second training image can be the same training image or different training images.

[0143] Specifically, the computer device can acquire the second training image corresponding to the training action of the training object and the state of the training object corresponding to the second training image, as well as the training information group. Based on this data, the interaction model is trained in a supervised manner. The computer device inputs the second training image and the training information group into the interaction model to be trained. After data processing by the interaction model, the interaction model outputs the predicted state of the object corresponding to the second training image. Then, the model loss is calculated based on the difference between the training object state and the predicted object state. The model loss is backpropagated to adjust the model parameters of the interaction model to be trained until the second convergence condition is met, thus obtaining the trained interaction model.

[0144] It is understandable that there can be multiple second training images. Multiple second training images can be obtained in each round for training.

[0145] In the above embodiments, supervised training of the interaction model is performed based on the second training image and the training object state and training information group corresponding to the second training image, which enables the trained interaction model to have the ability to recognize object state.

[0146] In one embodiment, the target video is a video stream obtained by real-time acquisition of images of the living environment of the target object. The object state recognition method further includes:

[0147] When the target object is in an abnormal state, an alarm message for the abnormal object state is sent to the management terminal corresponding to the living environment.

[0148] The target video is a video stream obtained by real-time acquisition of images of the target's living environment. For example, the target video could be real-time monitoring video acquired by a surveillance camera installed in the target's home; the target video could be real-time monitoring video acquired by a surveillance camera installed in a nursing home or hospital where the target resides; or the target video could be real-time monitoring video acquired by a surveillance camera installed in a place the target frequently visits.

[0149] An abnormal object state refers to an unusual object state. Abnormal object states can be set as needed. For example, a dangerous object state can be defined as an abnormal object state; an object state that differs from its usual state can be defined as an abnormal object state; and so on. For instance, if a target object's life is relatively regular, and its object state is generally stable and predictable throughout the day, then if the target object's current object state at a certain time period differs from its historical object states for that time period, then the target object's current object state for that time period can be determined as an abnormal object state.

[0150] The management terminal corresponding to the living environment is a management terminal for the target object, used to manage the object's status in the living environment. For example, if the target object's living environment is its home environment, the management terminal can be the target object's family members, guardians, etc.; if the target object's living environment is the environment where the target object is located in a nursing home or hospital, the management terminal can be the staff of the nursing home or hospital, or the target object's family members, guardians, etc.

[0151] Alarm messages for abnormal object states are used to alert that the target object is in an abnormal state. Alarm messages can be in at least one media format, such as text, image, video, or audio. Alarm messages can be displayed via SMS, applications, telephone, and other methods.

[0152] Specifically, the target video is a video stream obtained by real-time acquisition of images of the target object's living environment. In other words, the target video is a continuously generated video stream. The computer device can acquire the target video and decode video frames from it to perform object state recognition. If a target frame corresponding to the target object's target action appears, the target object's state at the corresponding target frame is determined based on the target frame and the background frame. Further identification is then made to determine whether the target object's state is abnormal. If the target object's state is abnormal, the computer device can send an alarm message for the abnormal object state to the management terminal corresponding to the living environment. This allows the management terminal to promptly become aware that the target object is in an abnormal state and take appropriate protective measures.

[0153] In the above embodiments, when the target object is in an abnormal state, an alarm message for the abnormal object state is sent to the management terminal corresponding to the living environment, which helps to restore the target object to a normal state in a timely manner and ensures the safety of the target object.

[0154] In one embodiment, the target video is a video stream obtained by real-time image acquisition of the living environment by a target device placed in the living environment of the target object. The object state recognition method further includes:

[0155] When the target object is in an abnormal state, an alarm message for the abnormal object state is sent to the target device. The target device displays an assistance request confirmation message for the abnormal object state, and in response to the agreement to accept assistance event in the assistance request confirmation message, determines the target assistance method for the target object, and sends a collaboration message for the target object to the collaborating device corresponding to the target assistance method.

[0156] Assistance request confirmation messages are used to determine whether a target object needs assistance. Assistance request confirmation messages can also be called assistance request inquiry messages. If the target object needs help, it interacts with the target device, causing the target device to send assistance messages to other devices. For example, an assistance request confirmation message could be a voice message asking if the target object needs help; the target object can then interact with the target device to indicate whether assistance is needed. Another example is a notification message asking if the target object needs help; the target object can then perform actions on the target device to indicate whether assistance is needed.

[0157] An "Agree to Assistance" event is an event that indicates agreement to accept assistance and confirms the need for assistance. This event can be generated based on a user's trigger action. For example, a trigger action in response to a confirmation message for assistance requests triggers an "Agree to Assistance" event; displaying an "Agree" control triggers an event in response to a trigger action on that control. An "Agree to Assistance" event can also be generated based on a user's contactless actions. Contactless actions are operations that convey information to the target device in a non-contact manner. For example, an assistance request confirmation message could be a voice message asking if the target needs help; the target can then send a voice message indicating agreement to receive assistance. If the target device receives the voice message from the target, an "Agree to Assistance" event is generated. Of course, an "Agree to Assistance" event can also be automatically triggered. For example, if the user does not refuse assistance within a certain time, an "Agree to Assistance" event is automatically triggered.

[0158] The target assistance method refers to the specific way to help the target object when the target object is in an abnormal state. The target assistance method can be a preset collaboration method, such as any one of emergency contact assistance, professional organization assistance, or robot assistance. Alternatively, the target assistance method can be selected from candidate assistance methods. The assistance request confirmation message can further include multiple candidate assistance methods for the target object to choose from. If the target object needs assistance, it can select a candidate assistance method, which will then be the target assistance method. The target object can select the target assistance method from the candidate methods through contact or contactless methods. Of course, the target assistance method can also be a custom assistance method. The target object can provide a custom assistance method to the target device. For example, the target object can provide a friend's contact information to the target device, allowing the target device to send an assistance message to the friend's friend's device based on that contact information.

[0159] The target device sends an assistance message to the assistance device corresponding to the target assistance method. The assistance message for the target object includes the target object's location information and status. For example, candidate assistance methods include emergency contact assistance, professional organization assistance, and robot assistance. If the target object selects emergency contact assistance, the assistance device can be the communication device corresponding to the target object's emergency contact. If the target object selects professional organization assistance, the assistance device can be the communication device corresponding to the professional organization. If the target object selects robot assistance, the assistance device is a robot within a certain range of the living environment. The target device sends an assistance message to the robot, allowing the robot to locate the target object based on its location information and take appropriate measures to assist the target object based on its status.

[0160] In this way, when an anomaly is detected in the target object, timely inquiry into whether the target object needs assistance via a request confirmation message can effectively help the target object return to normal as soon as possible.

[0161] In one embodiment, the target video is a video stream obtained by real-time acquisition of images of the living environment of the target object. The object state recognition method further includes:

[0162] When the target object is in an abnormal state, obtain the set of adjacent frames of the target frame from the target video; identify the target object in the adjacent frames of the adjacent frame set, and use the adjacent frames of the identified target object as auxiliary frames; determine the auxiliary object from the auxiliary frames; the auxiliary object and the target object in the target frame are different objects of the same type; send an alarm message for the abnormal object state to the terminal of the object corresponding to the auxiliary object.

[0163] The set of adjacent frames of the target frame refers to video frames in the target video whose acquisition time distance from the target frame is less than a time distance threshold. In other words, the set of adjacent frames of the target frame includes several video frames in the target video that precede and follow the target frame.

[0164] A secondary frame is a video frame selected from the set of adjacent frames of the target frame. A secondary object refers to another target object that exists in the secondary frame, but is different from the target object exhibiting an abnormal object state. For example, if the target object is a person, and the abnormal object state in the target frame is Zhang San, the secondary object could be Li Si appearing in the secondary frame. It can be understood that when an abnormal object state is detected in the target object in the target frame, no abnormal object state is detected in the secondary object.

[0165] The object terminal corresponding to the auxiliary object refers to the communication device corresponding to the auxiliary object. For example, the object terminal corresponding to the auxiliary object can be the auxiliary object's mobile phone, smartwatch, etc.

[0166] Specifically, when the target object is in an abnormal state in the target frame, the computer device can further determine the auxiliary object from the target video and send an alarm message for the abnormal object state to the terminal corresponding to the auxiliary object, so that the auxiliary object can take timely measures to help the target object.

[0167] When the target object is in an abnormal state, a set of adjacent frames is obtained from the target video. For example, the set of adjacent frames is composed of video frames within 5 minutes of the target frame. Then, target object identification is performed on the adjacent frames in the adjacent frame set. The adjacent frames of the identified target object are used as auxiliary frames, and auxiliary objects are determined from the auxiliary frames. For example, the adjacent frame of the identified target object that is closest to the target frame is used as an auxiliary frame. From the auxiliary frames, at least one other target object different from the target object in the target frame that exhibits an abnormal state is selected as the auxiliary object. Finally, an alarm message for the abnormal object state is sent to the terminal corresponding to the auxiliary object.

[0168] For example, if the target object is a person, and user A in the target frame is in an abnormal state, if user B appears in the set of adjacent frames of the target frame, then user B is used as an auxiliary object, and an alarm message for user A is sent to the communication device corresponding to user B, so that user B can be notified of the abnormality of user A in a timely manner through the communication device.

[0169] In the above embodiments, when the target object is in an abnormal object state, an auxiliary object is determined from the set of adjacent frames of the target frame, and an alarm message for the abnormal object state is sent to the object terminal corresponding to the auxiliary object. The auxiliary object is usually close to the target object, which helps to quickly help the target object restore to a normal object state and ensure the safety of the target object.

[0170] In a specific embodiment, the method of this application can be applied to fall detection in a home setting, effectively improving the accuracy of fall recognition. Referring to Figures 9 and 10, the server acquires the video stream captured by a camera in the home setting and decodes it into images. First, human detection (i.e., human body detection) is performed on the images. Once a target person is detected, the action recognition module identifies the action performed by the target person in the image frame. Simultaneously, based on the human body detection results of each image, the background frames in the video stream can also be identified (i.e., environmental context detection). It can be understood that the image in which the target person is identified to be falling is used as the foreground frame (i.e., the target frame). Image semantic segmentation is performed on both the foreground and background frames. In the foreground frame, the focus is mainly on the pixel-level semantic segmentation region of the target person, while in the background frame, the focus is on the segmentation results of objects overlapping or connected to the target person's region. The overlap rate between the target person and objects in the surrounding environment is calculated based on the image semantic segmentation results. The overlap rate, the falling action, and the objects together constitute an interaction triple (i.e., an interaction information group). The fall recognition result of the target person is determined based on the interaction triple. The interaction triple can be input into an interaction model, and the interaction model outputs the final fall recognition result.

[0171] If the fall detection result indicates that the target person has fallen, the server can return the fall detection result to the front end. The front end can then perform further operations such as image display and abnormal behavior alerts based on the fall detection result. For example, the front end can display a foreground frame and circle the fallen target person in the foreground frame.

[0172] It's understandable that the target audience can be set as needed. For example, the target audience could be the elderly or children.

[0173] It is understandable that with the increase in the elderly population, there is a need for more advanced home devices to monitor the activities of the elderly in order to ensure their safety. The method of this application can more effectively distinguish the falling actions of the elderly in different situations in areas such as bedrooms and living rooms, when furniture such as beds and sofas that are strongly associated with falling actions are present, and distinguish whether the elderly are in a resting state or have fallen.

[0174] Of course, this method can be applied not only to home safety monitoring, but also to scenarios such as patient assistance and care in medical / elderly care facilities. This method can also be used to detect conditions other than falls.

[0175] In one specific embodiment, the method of this application can be applied to a cloud-edge collaborative system for fall detection. The cloud-edge collaborative system is a fusion of cloud computing and edge computing, aiming to achieve efficient data processing and transmission. The cloud-edge collaborative system includes cloud devices and edge devices. Edge devices, i.e., intelligent camera devices, are used to capture real-time video, decode it into images, and perform preliminary analysis on the local device, including human detection, human skeleton point detection, and action recognition based on skeleton point temporal data, to obtain interaction triples. Based on the interaction triples, a preliminary judgment result is determined as to whether the target person has fallen. For example, if the overlap rate in the interaction triples is greater than a corresponding threshold, the preliminary judgment result is that the target person has fallen. When the preliminary judgment result indicates that the target person has fallen, the locally processed result, the current frame (i.e., the foreground frame), and the interaction triples are uploaded to the cloud device. The cloud device performs secondary analysis based on the data sent from the edge device to obtain the final judgment result. For example, the foreground frame and the interaction triples are input into an interaction model, and the interaction model outputs the final judgment result. When the final judgment is that the target person has fallen, the cloud device can push the final judgment result (such as fall alarm and related images) to the user terminal or monitoring center according to the business logic. The cloud device can also automatically dial the emergency contact number of the target person.

[0176] In this way, efficient and accurate fall detection is achieved through the collaborative work of edge devices and cloud devices. The edge devices make an initial judgment and filter the data. Only when the initial judgment indicates that the target person has fallen is the relevant data sent to the cloud devices, which then make a second judgment to obtain the final judgment result.

[0177] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0178] Based on the same inventive concept, this application also provides an object state recognition device for implementing the object state recognition method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more object state recognition device embodiments provided below can be found in the limitations of the object state recognition method described above, and will not be repeated here.

[0179] In one embodiment, as shown in FIG11, an object state recognition device is provided, including: a video acquisition module 1102, an image determination module 1104, an image segmentation module 1106, an information group generation module 1108, and a state recognition module 1110, wherein:

[0180] The video acquisition module 1102 is used to acquire the target video.

[0181] The image determination module 1104 is used to acquire the background frame corresponding to the target video and determine the target frame corresponding to the target action of the target object from the target video.

[0182] The image segmentation module 1106 is used to segment the background frame and the target frame to obtain the object regions corresponding to each object in the background frame and the target regions corresponding to the target objects in the target frame.

[0183] The information group generation module 1108 is used to generate interactive information groups based on the location information of the target area and the object area, the object corresponding to the object area, and the target action, so as to obtain the interactive information groups between the target object and each object.

[0184] The state recognition module 1110 is used to determine the state of the target object in the target frame based on each group of interactive information.

[0185] In one embodiment, the image determination module 1104 is further configured to:

[0186] Target object recognition is performed on video frames in the target video. If the target object is not recognized after a certain number of consecutive video frames, the last video frame in the consecutive target video frames is taken as the background frame.

[0187] In one embodiment, the image determination module 1104 is further configured to:

[0188] Based on a general object recognition algorithm, object recognition is performed on video frames in the target video, and the first live object identified is taken as the target object.

[0189] Target object identification algorithm based on protection priority of target object;

[0190] Based on the target object recognition algorithm, target objects are identified in video frames of the target video.

[0191] In one embodiment, the number of targets and the recognition accuracy of the target object recognition algorithm are negatively correlated.

[0192] In one embodiment, the image determination module 1104 is further configured to:

[0193] Target object recognition is performed on the video frames of the target video, and the video frames in which the target object is identified are used as candidate frames;

[0194] The candidate frames within the target time period are obtained to form a candidate frame sequence.

[0195] Target object skeleton point recognition is performed on the candidate frames in the candidate frame sequence to obtain the skeleton point data corresponding to the candidate frames, and the skeleton point data sequence corresponding to the candidate frame sequence is formed.

[0196] Input the skeleton point data sequence into the multi-frame action recognition model to obtain the predicted action corresponding to the candidate frame sequence;

[0197] When the predicted action is the target action, the target frame corresponding to the target action of the target object is determined from the candidate frame sequence.

[0198] In one embodiment, the image determination module 1104 is further configured to:

[0199] Each candidate frame in the candidate frame sequence is input into the trained single-frame action recognition model to obtain the action recognition result corresponding to each candidate frame; the action recognition result includes the number of target objects and actions in the candidate frame;

[0200] The action recognition results include candidate frames of a single target object and the target action, which are the target frames corresponding to the target action of the target object.

[0201] In one embodiment, the object state recognition device is further used for:

[0202] Obtain the first training image and the corresponding target object location training label and target object action training label;

[0203] The first training image is input into the single-frame action recognition model to be trained to obtain the target object location prediction label and the target object action prediction label corresponding to the first training image.

[0204] Based on the differences between the training labels and predicted labels of the target object's location, and between the training labels and predicted labels of the target object's actions, the model parameters of the single-frame action recognition model to be trained are adjusted until the first convergence condition is met, thus obtaining the trained single-frame action recognition model.

[0205] In one embodiment, the state recognition module 1110 is further configured to:

[0206] Get the state of each reference object corresponding to the target action, and get the reference object corresponding to each reference object state;

[0207] The reference object in the background frame is taken as the target object, and the reference overlap threshold corresponding to the state of the reference object corresponding to the target object is obtained as the target overlap threshold.

[0208] For the interactive information group corresponding to the target object, when the overlap between the object region corresponding to the target object and the target region is greater than the target overlap threshold, the reference object state corresponding to the target object is used as the target object state corresponding to the target object in the target frame.

[0209] In one embodiment, the state recognition module 1110 is further configured to:

[0210] When the overlap between the object region corresponding to the target object and the target region is less than or equal to the target overlap threshold, the target frame and each interaction information group are input into the trained interaction model to obtain the target object state corresponding to the target frame.

[0211] In one embodiment, the state recognition module 1110 is further configured to:

[0212] Input the target frame and each group of interaction information into the trained interaction model;

[0213] The overlapping region is determined in the target frame based on each group of interactive information.

[0214] Image features of the overlapping region are extracted to obtain the first interaction feature. Based on the object and target action in the interaction information group corresponding to the overlapping region, the second interaction feature is obtained. Image features of the target frame are extracted to obtain the third interaction feature.

[0215] Based on the first interaction feature, the second interaction feature, and the third interaction feature, a comprehensive interaction feature is obtained;

[0216] Output the target object state in the target frame based on comprehensive interactive features.

[0217] In one embodiment, the object state recognition device is further used for:

[0218] Obtain the second training image corresponding to the training action of the training object and the state of the training object corresponding to the second training image, and the training information group; the state of the training object is the object state of the training object corresponding to the second training image, and the training information group is determined based on the position information of the area where the training object is located in the second training image, the position information and object information of the area where the object is located in the background image corresponding to the second training image, and the training action;

[0219] Input the second training image and training information group into the interactive model to be trained to obtain the predicted object state corresponding to the second training image;

[0220] Based on the difference between the state of the training object and the state of the predicted object, the model parameters of the interaction model to be trained are adjusted until the second convergence condition is met, and the trained interaction model is obtained.

[0221] In one embodiment, the target video is a video stream obtained by real-time acquisition of images of the living environment of the target object. The object state recognition device is also used for:

[0222] When the target object is in an abnormal state, an alarm message for the abnormal object state is sent to the management terminal corresponding to the living environment.

[0223] In one embodiment, the target video is a video stream obtained by real-time acquisition of images of the living environment of the target object. The object state recognition device is also used for:

[0224] When the target object is in an abnormal state, obtain the set of adjacent frames of the target frame from the target video;

[0225] Target object identification is performed on adjacent frames in the adjacent frame set, and the adjacent frames where the target object is identified are used as auxiliary frames;

[0226] The auxiliary object is determined from the auxiliary frame; the auxiliary object and the target object in the target frame are different objects of the same type.

[0227] Send an alarm message for the abnormal object status to the corresponding object terminal of the auxiliary object.

[0228] In one embodiment, the target video is a video stream obtained by a target device positioned in the living environment of the target object, capturing images of the living environment in real time. The object state recognition device is also used for:

[0229] When the target object is in an abnormal state, an alarm message for the abnormal object state is sent to the target device. The target device displays an assistance request confirmation message for the abnormal object state, and in response to the agreement to accept assistance event in the assistance request confirmation message, determines the target assistance method for the target object, and sends a collaboration message for the target object to the collaborating device corresponding to the target assistance method.

[0230] In one embodiment, the assistance request confirmation message includes candidate assistance methods, and the target device is further configured to determine a target assistance method in response to a selection operation for a candidate assistance method.

[0231] In one embodiment, when the target assistance method is robot assistance, the assisting device is a robot whose distance from the living environment is within a preset threshold.

[0232] The aforementioned object state recognition device analyzes the target action of a target object within a target frame based on a background frame and a target frame to identify the target object's state corresponding to that frame. The background frame provides information about the object, while the target frame provides information about the target object. An interaction information group is established based on the object's information and the target object's information. Analyzing the interaction relationship between the target object and the target object through this interaction information group effectively improves the accuracy of object state recognition. Considering the interaction relationship between the target object and the environment when recognizing object states effectively distinguishes object states that are easily confused due to the same action, thus improving the accuracy of object state recognition.

[0233] Each module in the aforementioned object state recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0234] In one embodiment, a computer device, which may be a server, is provided, and its internal structure can be as shown in Figure 12. The computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the non-volatile storage media. The database of the computer device stores various models, reference information, and other data. The I / O interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communication with external terminals via a network connection. When the computer-readable instructions are executed by the processor, they implement an object state recognition method.

[0235] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as shown in Figure 13. The computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer-readable instructions. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer-readable instructions are executed by the processor, they implement an object state recognition method. The display unit of the computer device is used to form a visually visible image and may be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0236] Those skilled in the art will understand that the structures shown in Figures 12 and 13 are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements.

[0237] In one embodiment, a computer device is also provided, including a memory and a processor, the memory storing computer-readable instructions, the processor executing the computer-readable instructions to implement the steps in the above method embodiments.

[0238] In one embodiment, a computer-readable storage medium is provided storing computer-readable instructions that, when executed by a processor, implement the steps in the above method embodiments.

[0239] In one embodiment, a computer program product is provided, the computer program product including computer-readable instructions that, when executed by a processor, implement the steps in the above method embodiments.

[0240] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0241] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a non-volatile computer-readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0242] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0243] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for object state recognition, characterized in that, Performed by a computer device, the method includes: Acquire the target video; Obtain the background frame corresponding to the target video, and determine the target frame corresponding to the target action of the target object from the target video; Image segmentation is performed on the background frame and the target frame to obtain the object regions corresponding to each object in the background frame and the target regions corresponding to the target objects in the target frames. Based on the location information of the target region and the object region, the object corresponding to the object region, and the target action, an interaction information group is generated to obtain the interaction information groups between the target object and each of the objects; and Based on each group of interactive information, the target object state corresponding to the target frame is determined.

2. The method according to claim 1, characterized in that, The step of obtaining the background frame corresponding to the target video includes: Target object identification is performed on the video frames in the target video. If no target object is identified for a consecutive number of video frames, the last video frame from the consecutive number of video frames is taken as the background frame.

3. The method according to claim 2, characterized in that, The step of identifying target objects in video frames of the target video includes: Based on a general object recognition algorithm, object recognition is performed on video frames in the target video, and the first live object identified is taken as the target object. The target object identification algorithm is determined based on the protection priority of the target object; and Based on the target object recognition algorithm, target object recognition is performed on video frames in the target video.

4. The method according to claim 2 or 3, characterized in that, The number of targets and the recognition accuracy of the target object recognition algorithm are negatively correlated.

5. The method according to any one of claims 1 to 3, characterized in that, Determining the target frame corresponding to the target action of the target object from the target video includes: Target object identification is performed on the video frames of the target video, and the video frames in which the target object is identified are used as candidate frames; The candidate frames within the target time period are obtained to form a candidate frame sequence. Target object skeleton point recognition is performed on the candidate frames in the candidate frame sequence to obtain the skeleton point data corresponding to the candidate frames, and the skeleton point data sequence corresponding to the candidate frame sequence is formed. The skeleton point data sequence is input into a multi-frame action recognition model to obtain the predicted action corresponding to the candidate frame sequence; and When the predicted action is the target action, the target frame corresponding to the target action of the target object is determined from the candidate frame sequence.

6. The method according to claim 5, characterized in that, Determining the target frame corresponding to the target action of the target object from the candidate frame sequence includes: Each candidate frame in the candidate frame sequence is input into the trained single-frame action recognition model to obtain the action recognition result corresponding to each candidate frame; the action recognition result includes the number of target objects and their actions in the candidate frame; and The action recognition result includes a single target object and candidate frames of the target action, which are the target frames corresponding to the target action of the target object.

7. The method according to claim 6, characterized in that, The method further includes: Obtain the first training image and the target object location training label and target object action training label corresponding to the first training image; The first training image is input into the single-frame action recognition model to be trained to obtain the target object location prediction label and the target object action prediction label corresponding to the first training image; and Based on the differences between the target object location training label and the target object location prediction label, and the differences between the target object action training label and the target object action prediction label, the model parameters of the single-frame action recognition model to be trained are adjusted until the first convergence condition is met, thereby obtaining the trained single-frame action recognition model.

8. The method according to any one of claims 1 to 7, characterized in that, Determining the target object state corresponding to the target frame based on each group of interactive information includes: Obtain the state of each reference object corresponding to the target action of the target object, and obtain the reference object corresponding to each reference object state respectively; Using a reference object within the background frame as the target object, the reference overlap threshold corresponding to the state of the reference object corresponding to the target object is obtained as the target overlap threshold; and For the interactive information group corresponding to the target object, when the overlap between the object region corresponding to the target object and the target region is greater than the target overlap threshold, the reference object state corresponding to the target object is taken as the target object state of the target object in the target frame.

9. The method according to claim 8, characterized in that, The method further includes: When the overlap between the object region corresponding to the target object and the target region is less than or equal to the target overlap threshold, the target frame and each interaction information group are input into the trained interaction model to obtain the target object state corresponding to the target frame.

10. The method according to any one of claims 1 to 7, characterized in that, Determining the target object state corresponding to the target frame based on each group of interactive information includes: The target frame and each group of interactive information are input into the trained interactive model; Based on the aforementioned interactive information groups, overlapping regions are determined in the target frame, respectively. Extract image features from the overlapping region to obtain a first interaction feature; obtain a second interaction feature based on the object and target action in the interaction information group corresponding to the overlapping region; and extract image features from the target frame to obtain a third interaction feature. Based on the first interaction feature, the second interaction feature, and the third interaction feature, a comprehensive interaction feature is obtained; and Based on the comprehensive interactive features, the target object state corresponding to the target frame is output.

11. The method according to claim 10, characterized in that, The method further includes: The training object obtains a second training image corresponding to the training action of the training object, and a training object state and training information group corresponding to the second training image; the training object state is the object state of the training object in the second training image, and the training information group is determined based on the location information of the area where the training object is located in the second training image, the location information and object information of the area where the object is located in the background image corresponding to the second training image, and the training action. The second training image and the training information set are input into the interaction model to be trained to obtain the predicted object state corresponding to the second training image; and Based on the difference between the state of the training object and the state of the predicted object, the model parameters of the interaction model to be trained are adjusted until the second convergence condition is met, and the trained interaction model is obtained.

12. The method according to any one of claims 1 to 11, characterized in that, The target video is a video stream obtained by real-time acquisition of images of the living environment of the target object, and the method further includes: When the target object is in an abnormal state, an alarm message for the abnormal object state is sent to the management terminal corresponding to the living environment.

13. The method according to any one of claims 1 to 11, characterized in that, The target video is a video stream obtained by real-time acquisition of images of the living environment of the target object, and the method further includes: When the target object is in an abnormal object state, obtain the set of adjacent frames of the target frame from the target video; Target object identification is performed on adjacent frames in the adjacent frame set, and adjacent frames with identified target objects are used as auxiliary frames; An auxiliary object is determined from the auxiliary frame; the auxiliary object and the target object in the target frame are different objects of the same type; and Send an alarm message for the abnormal object status to the object terminal corresponding to the auxiliary object.

14. The method according to any one of claims 1 to 11, characterized in that, The target video is a video stream obtained by real-time image acquisition of the living environment by a target device set up in the living environment of the target object. The method further includes: When the target object is in an abnormal state, an alarm message for the abnormal object state is sent to the target device; the target device is used to display an assistance request confirmation message for the abnormal object state, and in response to the agreement to accept assistance event in the assistance request confirmation message, determine the target assistance method for the target object, and send a collaboration message for the target object to the collaborating device corresponding to the target assistance method.

15. The method according to claim 14, characterized in that, The assistance request confirmation message includes candidate assistance methods, and the target device is further configured to determine the target assistance method in response to a selection operation for the candidate assistance methods.

16. The method according to claim 14, characterized in that, When the target assistance method is robot assistance, the assistance device is a robot whose distance from the living environment is within a preset threshold.

17. An object state recognition device, characterized in that, The device includes: The video acquisition module is used to acquire the target video; An image determination module is used to acquire the background frame corresponding to the target video and determine the target frame corresponding to the target action of the target object from the target video; An image segmentation module is used to segment the background frame and the target frame to obtain the object regions corresponding to each object in the background frame and the target regions corresponding to the target objects in the target frame. The information group generation module is used to generate interaction information groups based on the location information of the target region and the object region, the object corresponding to the object region, and the target action, thereby obtaining interaction information groups between the target object and each of the objects; and The state recognition module is used to determine the state of the target object in the target frame based on each group of interactive information.

18. A computer device comprising a memory and a processor, the memory storing computer-readable instructions, characterized in that, When the processor executes the computer-readable instructions, it implements the steps of the method according to any one of claims 1 to 16.

19. A computer-readable storage medium storing computer-readable instructions thereon, characterized in that, When the computer-readable instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 16.

20. A computer program product comprising computer-readable instructions, characterized in that, When executed by a processor, the computer-readable instructions implement the steps of the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Falling behavior identification method, equipment and system

    CN112949417A

  • Video-based pedestrian static state identification method, electronic equipment and storage medium

    CN113313007A

  • Human body target detection method and system in conference polling based on dynamic and static detection combination

    CN114299138A

  • Systems and methods for responding to a vehicle parked on shoulder of the road

    US20200202714A1