Information processing method and device, equipment, medium and program product
The infrared camera collects infrared speckle array calibration information, performs depth calculation and posture recognition, which solves the problem of user information exposure and realizes safety monitoring and rapid response in the scenario where the user is alone.
Patent Information
- Application Number
- CN202510422828.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
AI Technical Summary
In the scene where the user is alone, the existing fall recognition and call system is fully monitored through visible light cameras, resulting in the user's face, body parts images or non-public areas of activity images being directly exposed, and information exposure cannot be avoided.
An infrared camera is used to collect infrared speckle array calibration information, obtain multi-frame depth maps through depth calculation, perform cluster detection and posture recognition of moving objects, determine the user's motion state, and reduce information exposure.
In the case of reducing user information exposure, accurately monitor the user's movement status, promptly respond to user's abnormal behavior, and ensure user's safety.
Smart Images

Figure CN119942652A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and in particular, relates to an information processing method, device, equipment, medium and program product. Background Art
[0002] When the user is alone, if an accident such as falling or fainting occurs, the user may face safety risks because he or she cannot get up or seek help on his or her own. In order to ensure the safety of the user when alone, a fall recognition and help calling system can be developed to trigger a rescue mechanism when an accident occurs to the user.
[0003] The relevant technology involves using a visible light camera to capture video images, processing the video images, identifying the user's fall condition, and then triggering a rescue mechanism when the user falls.
[0004] However, the above solution requires all-round monitoring of the user through a visible light camera. The images collected by the visible light camera are a true reflection of the user's environment and may contain rich details. As a result, it is inevitable that the user's face and body parts will be directly exposed, or the user's activities in non-public areas, and other information that the user does not want to be exposed. Summary of the invention
[0005] The embodiments of the present application provide an information processing method, apparatus, device, computer storage medium, and computer program product, which can accurately monitor the user's motion status while reducing the exposure of user information.
[0006] In a first aspect, an embodiment of the present application provides an information processing method, including: When it is detected that the user is in the first area, collecting infrared speckle array calibration information of the first area in the first target time period through an infrared camera; Performing depth calculation on the infrared speckle array calibration information to obtain a multi-frame first depth map corresponding to a first target time period; Perform moving object cluster detection according to the first depth map of multiple frames to obtain a moving object cluster detection result; Determine the human body region in the first depth map of each frame according to the moving object cluster detection result; Performing posture recognition on the human body area in the first depth map of each frame respectively to obtain the posture recognition result of the first depth map of each frame; The motion state of the user in the first target time period is determined according to the posture recognition result of the first depth map of each frame.
[0007] In an optional implementation, after determining the motion state of the user in the first target period according to the gesture recognition results of the first depth map of each frame, the method further includes: When the user's motion state in the first target period is the target state, prompt information is sent to the target client.
[0008] In an optional implementation, performing moving object cluster detection according to multiple frames of first depth images to obtain a moving object cluster detection result includes: Performing ground detection and background detection on the first depth map of each frame respectively to obtain detection results corresponding to the first depth map of each frame; According to the detection results corresponding to the first depth map of each frame, the ground area and the background area in the first depth map of each frame are removed to obtain a second depth map corresponding to the first depth map of each frame, where the second depth map includes the foreground area; Perform moving object cluster detection on the second depth map corresponding to the first depth map of each frame to obtain a moving object cluster detection result.
[0009] In an optional implementation, performing moving object cluster detection on the second depth map corresponding to the first depth map of each frame to obtain a moving object cluster detection result includes: For the second depth map corresponding to the first depth map of each frame, respectively: dividing the voxels in the second depth map into a plurality of groups according to a preset clustering algorithm; performing connectivity analysis on the plurality of groups to obtain connectivity analysis results; and dividing the plurality of groups into a plurality of components according to the connectivity analysis results; For each component, determining a motion trajectory of each component based on position information of the component in a second depth map corresponding to the first depth map of each frame; According to the motion trajectories of the various components, multiple components are matched to obtain matching results; According to the matching results, multiple components are divided to obtain multiple moving object clusters.
[0010] In an optional implementation, posture recognition is performed on the human body region in the first depth map of each frame to obtain the posture recognition result of the first depth map of each frame, including performing the following steps for the human body region in the first depth map of each frame: Identify key points of the human body area to obtain spatial coordinates of multiple target key points, where the multiple target key points include key points corresponding to multiple human body parts; According to the spatial coordinates of multiple target key points, the center of mass coordinates and head turning information are extracted; Posture recognition is performed based on the center of mass coordinates and head turning information to obtain a posture recognition result.
[0011] In an optional embodiment, the method further includes: When it is detected that the user is in the second area, a video frame queue of the second area in the second target time period is collected by a visible light camera; Performing human posture recognition on each video frame in the video frame queue respectively to obtain posture recognition results of each video frame; The motion state of the user in the second target time period is determined according to the gesture recognition result of each video frame in the video frame queue.
[0012] In a second aspect, an embodiment of the present application provides an information processing device, including: A collection module, configured to collect infrared speckle array calibration information of the first area in a first target period through an infrared camera when detecting that the user is in the first area; A calculation module, used to perform depth calculation on the infrared speckle array calibration information to obtain a multi-frame first depth map corresponding to a first target time period; A detection module, used to perform moving object cluster detection based on multiple frames of first depth images to obtain a moving object cluster detection result; A determination module, used to determine a human body area in the first depth map of each frame according to the moving object cluster detection result; A recognition module, used to perform posture recognition on the human body area in the first depth map of each frame respectively, to obtain the posture recognition result of the first depth map of each frame; The determination module is further used to determine the motion state of the user in the first target time period according to the posture recognition results of the first depth map of each frame.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, the device comprising: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the information processing method as described in any optional implementation manner of the first aspect of the present application.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, an information processing method as in any optional implementation manner of the first aspect of the present application is implemented.
[0015] In a fifth aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes an information processing method as described in any optional implementation of the first aspect of the present application.
[0016] The information processing method, device, equipment, computer storage medium and computer program product of the embodiment of the present application can collect infrared speckle array calibration information of the first area in the first target period through an infrared camera when the user is detected to be in the first area. The first area can be a non-public area. The infrared speckle array calibration information mainly reflects the temperature distribution and thermal radiation characteristics of the object, rather than the specific shape or color of the object, thereby avoiding direct exposure of the user's face, body part image, or the user's activity screen in the non-public area and other information that the user does not want to expose. Then, the infrared speckle array calibration information can be depth calculated to obtain a multi-frame first depth map corresponding to the first target period. According to the multi-frame first depth map, moving object cluster detection is performed to obtain a moving object cluster detection result. According to the moving object cluster detection result, the human body area in each frame of the first depth map is determined. The human body area in each frame of the first depth map is respectively subjected to posture recognition to obtain the posture recognition result of each frame of the first depth map. According to the posture recognition result of each frame of the first depth map, the user's motion state in the first target period is determined. In this way, the user's motion state can be accurately monitored while reducing the exposure of user information, which is conducive to rapid response when the user has abnormal behavior and protects the user's safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 It is a schematic diagram of the architecture of an information processing system provided by an embodiment of the present application; Figure 2 is a flowchart of an information processing method provided by another embodiment of the present application; Figure 3 is a structural diagram of an information processing device provided by yet another embodiment of the present application; Figure 4 It is a structural diagram of an electronic device provided in yet another embodiment of the present application. DETAILED DESCRIPTION
[0019] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.
[0020] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0021] In the fall recognition and rescue solution that uses visible light cameras to monitor users in all directions, the images collected by the visible light cameras are a true reflection of the user's environment and may contain rich details. As a result, it is inevitable that the user's face and body parts will be directly exposed, or the user's activities in non-public areas, and other information that the user does not want to be exposed.
[0022] In view of this, the embodiments of the present application propose an information processing method, apparatus, device, computer storage medium and computer program product.
[0023] The following is an introduction to the information processing method provided by the embodiment of the present application through specific embodiments and their application scenarios in conjunction with the accompanying drawings. The information processing method provided by the embodiment of the present application, the device that executes the sending may be an information processing device, or a partial module in the information processing device for executing the information processing method. In the embodiment of the present application, the information processing method provided by the embodiment of the present application is described in detail by taking the information processing device executing the information processing method as an example.
[0024] It should be noted that in the technical solutions of the embodiments of the present application, the collection of information will be reasonably used within the legal scope.
[0025] Taking a fall alarm system as an example, the information processing method of the embodiment of the present application can be based on the following Figure 1The architecture shown is implemented.
[0026] like Figure 1 As shown, the monitoring module can be used to monitor public areas. The human body contour recognition algorithm module can be used to identify the human body contour features of users in public areas, focusing on monitoring abnormal behaviors such as falls that endanger the health of users. If abnormal behaviors such as falls are identified, the system can promptly transmit the video screenshots, warning information and location information of the falls to the client of the community staff so that the community staff can call for help in time. The infrared detection module can be used to monitor non-public areas. The infrared speckle array calibration information of non-public areas can be obtained through the infrared detection module, instead of directly obtaining the optical image of non-public areas. Based on the infrared speckle array calibration information obtained by the infrared detection module, the human body posture recognition algorithm module uses machine vision combined with artificial intelligence neural network algorithm to perform posture recognition, and recognize human standing, lying, sitting, falling and other actions. In addition, taking elderly users as an example, according to the posture recognition results, the time required for users to carry out daily activities can also be recorded, and the data can be uploaded to the cloud and synchronized to the client of the community staff, so that the community staff can visit at any time and call for help from the community staff in time for abnormal falls.
[0027] The following is combined with Figure 2 The information processing method provided in the embodiments of the present application is described in detail.
[0028] Figure 2 FIG. 1 is a flow chart showing an information processing method provided by an embodiment of the present application. Figure 2 As shown, the information processing method may specifically include the following steps S110 to S160.
[0029] S110: When it is detected that the user is in the first area, infrared speckle array calibration information of the first area in the first target period is collected by an infrared camera.
[0030] In step S110, the user may include any object that needs to be monitored, for example, the elderly, children or other objects with safety detection requirements who are alone at home. The first area may be a non-public area. The first target period may include any time period before the current moment, for example, it may be 1s before the current moment, 0.5s before the current moment, 6h before the current moment, 8h before the current moment, etc. The appropriate duration may be selected according to the real-time requirements of the status monitoring, which is not limited here. Before monitoring the user, an infrared camera may be set in the first area with the authorization of the user. When it is detected that the user is in the first area, the infrared speckle array may be automatically calibrated in the space of the first area through the infrared camera to obtain the infrared speckle array calibration information of the first area in the first target period.
[0031] S120, performing depth calculation on the infrared speckle array calibration information to obtain a plurality of frames of first depth maps corresponding to a first target time period.
[0032] In step S120, the depth calculation may be implemented by methods known in the art, which are not limited herein. For example, the distance information corresponding to the infrared speckle array calibration information may be calculated by a complementary metal oxide semiconductor (CMOS) depth chip to obtain a first depth map.
[0033] S130, performing moving object cluster detection according to the multiple frames of first depth images to obtain a moving object cluster detection result.
[0034] In step S130, the moving object cluster detection result may include dividing the moving objects in the first depth map into a plurality of moving object clusters. Moving objects belonging to the same cluster may belong to the same human body or object.
[0035] S140, determining a human body region in the first depth map of each frame according to the moving object cluster detection result.
[0036] In step S140, determining the human body area in the first depth map of each frame can be achieved through a variety of detection strategies, for example, it can be achieved using a detection model based on deep learning. Exemplarily, a detection model based on deep learning can be used to detect human body areas from images, and the detection results can be optimized by a non-maximum suppression method (NMS). Due to different human body postures, multiple detection frames may appear in the same part, which not only increases the complexity of subsequent processing, but also may lead to incorrect associations. By applying NMS, this redundancy can be effectively reduced, and an optimal bounding box can be selected for each detected human body area. In this way, a reliable foundation can be provided for subsequent tasks such as human posture estimation and behavior recognition.
[0037] S150 , performing posture recognition on the human body region in the first depth map of each frame respectively, to obtain a posture recognition result of the first depth map of each frame.
[0038] S160: Determine the motion state of the user in the first target time period according to the gesture recognition result of the first depth map of each frame.
[0039] In step S160, the motion state may include but is not limited to normal states such as lying, standing, sitting, walking, and abnormal states such as falling.
[0040] The information processing method, device, equipment, computer storage medium and computer program product of the embodiment of the present application can collect infrared speckle array calibration information of the first area in the first target period through an infrared camera when the user is detected to be in the first area. The first area can be a non-public area. The infrared speckle array calibration information mainly reflects the temperature distribution and thermal radiation characteristics of the object, rather than the specific shape or color of the object, thereby avoiding direct exposure of the user's face, body part image, or the user's activity screen in the non-public area and other information that the user does not want to expose. Then, the infrared speckle array calibration information can be depth calculated to obtain a multi-frame first depth map corresponding to the first target period. According to the multi-frame first depth map, moving object cluster detection is performed to obtain a moving object cluster detection result. According to the moving object cluster detection result, the human body area in each frame of the first depth map is determined. The human body area in each frame of the first depth map is respectively subjected to posture recognition to obtain the posture recognition result of each frame of the first depth map. According to the posture recognition result of each frame of the first depth map, the user's motion state in the first target period is determined. In this way, the user's motion state can be accurately monitored while reducing the exposure of user information, which is conducive to rapid response when the user has abnormal behavior and protects the user's safety.
[0041] In one embodiment, after determining the motion state of the user in the first target period according to the gesture recognition results of the first depth map of each frame, the method may further include: When the user's motion state in the first target time period is the target state, first prompt information is sent to the target client.
[0042] In the above implementation, the target state may include an abnormal state such as falling. The target client may include a client of a third party such as a community worker or a relative of the user. The first prompt information may be used to prompt the third party that the user is in an abnormal state, so that the third party can provide assistance to the user in a timely manner. This is conducive to timely response to the abnormal situation of the user, thereby protecting the safety of the user in a timely manner.
[0043] In one embodiment, when the user's motion state in the first target period remains in the preset state, the duration of the user in the preset state can be continuously recorded after the first target period until the user is no longer in the preset state, and the duration of the user in the preset state is sent to the third-party client.
[0044] The above-mentioned preset state may include a lying state. When the user maintains the preset state, it may be considered that the user is likely to be in a sleeping state. By recording the duration of the user in the preset state, the user's sleep time can be fed back to a third party, such as a community member. In this way, on the one hand, it is conducive to the third party's return visit, and on the other hand, it is conducive to reflecting the user's physical and mental health status through the user's sleep time, thereby facilitating attention to the user's health status and providing timely help to the user.
[0045] In one embodiment, performing moving object cluster detection according to multiple frames of first depth images to obtain a moving object cluster detection result may specifically include: The ground detection and the background detection are performed on the first depth map of each frame respectively to obtain the detection result corresponding to the first depth map of each frame.
[0046] According to the detection results corresponding to the first depth map of each frame, the ground area and the background area in the first depth map of each frame are removed to obtain a second depth map corresponding to the first depth map of each frame, where the second depth map includes the foreground area.
[0047] Perform moving object cluster detection on the second depth map corresponding to the first depth map of each frame to obtain a moving object cluster detection result.
[0048] In the above implementation, ground detection and background detection can be implemented by using pre-established ground model and background model.
[0049] In one example, after setting up the infrared camera, a depth map of the first area can be collected, which provides distance information from each pixel to the camera plane. The height data of each point in the first area can be directly obtained through the depth map. Then the depth map can be preprocessed, including denoising, smoothing and other steps to reduce the impact of noise on subsequent processing. Through the depth information of each point, a height threshold can be set to extract candidate areas that may belong to the ground. Specifically, the area with a height lower than the height threshold can be considered as a candidate area of the ground because the candidate area is closer to the ground than other objects. Then, the candidate area can be further screened in combination with texture features. Specifically, the texture features can be obtained by calculating the grayscale co-occurrence matrix of the candidate area, and matched with the predefined ground texture model, so as to further exclude those areas whose texture features do not match the ground. Subsequently, the remaining candidate areas can be verified by plane fitting. The remaining candidate areas are fitted into a plane by the least squares method or other optimization algorithms. If the fitting result is good, the remaining candidate area is confirmed as a ground area, and a mathematical model of the ground is obtained. Similarly, a background model can be constructed in combination with texture features. In one example, the background model may also be constructed by using an initial frame or multi-frame averaging method or a machine learning method.
[0050] In the above implementation, the second depth map corresponding to the first depth map of each frame can be obtained by the following steps: Based on the ground model, a mask image of the same size as the depth map is generated, in which the ground area is marked with a specific value, and the non-ground area remains the same or is marked with another value. The mask is applied to the depth map, and the area marked as the ground is removed to obtain a depth map containing only non-ground objects. The difference between the depth map of the current frame and the background model is calculated to identify those areas that are significantly different from the background. Based on the difference detection results, the foreground object is extracted, and a depth map or mask image containing only the foreground is generated to obtain a second depth map.
[0051] According to the above implementation, by performing foreground extraction on the first depth map of each frame, the ground and background in the first depth map of each frame can be removed, thereby helping to reduce the impact of the ground and background and obtain a more accurate moving object cluster detection result. In this way, by determining the human body area in the first depth map of each frame based on the moving object cluster detection result, the accuracy of human body area determination can be improved, thereby improving the accuracy of posture recognition. In this way, it is helpful to accurately identify the abnormal situation of the user and reduce the safety risk of the user.
[0052] In one embodiment, performing moving object cluster detection on the second depth map corresponding to the first depth map of each frame to obtain a moving object cluster detection result may specifically include: For the second depth map corresponding to the first depth map of each frame, respectively perform: divide the voxels in the second depth map into multiple groups according to a preset clustering algorithm. Perform connectivity analysis on the multiple groups to obtain connectivity analysis results. According to the connectivity analysis results, divide the multiple groups into multiple components.
[0053] For each component, the motion trajectory of each component is determined based on the position information of the component in the second depth map corresponding to the first depth map of each frame.
[0054] According to the motion trajectory of each component, multiple components are matched to obtain matching results.
[0055] According to the matching results, multiple components are divided to obtain multiple moving object clusters.
[0056] In the above implementation, the preset clustering algorithm may include but is not limited to one or more of a distance-based clustering algorithm, a density-based clustering algorithm, or a geometric feature-based clustering algorithm. Through the clustering algorithm, the voxels in the second depth map are divided into a plurality of groups, each of which may represent a potential connected component.
[0057] In the above implementation, the connectivity analysis may include performing connectivity analysis on multiple groups through spatial adjacency or geometric continuity. Through the connectivity analysis, it is possible to determine which groups among the multiple groups are truly connected, and merge the connected groups into complete components. For each component, features such as shape, size, and position may be further extracted for subsequent classification and recognition tasks.
[0058] In the above-mentioned implementation, the motion trajectory of each component is determined based on the position information of the component in the second depth map corresponding to the first depth map of each frame. Specifically, it may include: for consecutive multi-frame second depth maps, using feature matching or optical flow method and other technologies to track the position changes of the same component or feature point between different frames, and based on the results of inter-frame matching, calculating the motion vector or trajectory of the component, thereby determining the motion trajectory of each component.
[0059] In the above embodiment, matching multiple components according to the motion trajectories of the components may include: determining at least two components with similar motion patterns or trajectories as mutually matching components. According to the matching results, the components with similar motion patterns or trajectories may be classified into the same moving object cluster.
[0060] According to the above implementation, multiple components are extracted from the second depth map through clustering algorithm and connectivity analysis, and then the multiple components are divided into moving object clusters according to the motion trajectory of each component. This is conducive to accurately detecting and marking human body parts in the depth map, thereby improving the accuracy of human body area detection, and further helping to improve the accuracy of posture recognition.
[0061] In one embodiment, gesture recognition is performed on the human body region in the first depth map of each frame to obtain the gesture recognition result of the first depth map of each frame. Specifically, the following steps may be performed for the human body region in the first depth map of each frame: Key points of the human body region are identified to obtain spatial coordinates of multiple target key points, where the multiple target key points include key points corresponding to multiple human body parts.
[0062] According to the spatial coordinates of multiple target key points, the center of mass coordinates and head turning information are extracted.
[0063] Posture recognition is performed based on the center of mass coordinates and head turning information to obtain a posture recognition result.
[0064] In the above implementation, the target key points may include key points such as the head, shoulders, and elbows. The target key points can be identified by a posture estimation model. Multiple target key points can be associated to form a human skeleton structure. Then, based on the spatial positions of the multiple target key points, the relative positions between the target key points can be determined, and then combined with the prior knowledge of the human body structure, such as the relative proportions between the various parts of the human body and other constraints, these target key points can be associated using graph structure modeling and graph theory algorithms, and the human body area can be divided into multiple parts that conform to human anatomy and movement laws, such as the head, upper limbs, lower limbs, etc., and a unique identifier is assigned to each part. Then, based on the division results, the center of mass coordinates (x arg ,y arg ,z) and the human head turning information o, forming a four-dimensional coordinate (x arg ,y arg ,z,o) feature information, where z can represent the depth of the centroid position.
[0065] As an example, the x, y coordinates of the centroid (x arg ,y arg ), which can be calculated by the following formulas 1 and 2.
[0066] Formula 1 Formula 2 In equations 1 and 2, p, q can represent the moments of the human body contour; m p,q can represent the origin moment; I(x,y) can represent the pixel value at point (x,y); m 00 It can represent the zero-order moment, which is the sum of all pixel values; (m 10 ,m 01 ) can represent the distribution of the first-order moment image in the horizontal and vertical directions.
[0067] As an example, P1(x1, y1) and P2(x2, y2) can represent the positions of the left eye and the right eye respectively, and the head turning information o can be calculated by the following formula 3.
[0068] Formula 3 In one example, posture recognition is performed based on the centroid coordinates and the head turning information to obtain a posture recognition result, which may specifically include: converting the four-dimensional feature vector (x arg ,y arg ,z,o) is input into a three-layer backpropagation neural network (BP). For the output of layer n, the output of layer n+1 can be expressed as follows:
[0069] Formula 4 In Formula 4, f can represent an activation function, which can be a sigmoid function; y i It can represent the output activation value of the i-th neuron in the n+1-th layer; It can represent the connection weight from the jth neuron in the nth layer to the ith neuron in the n+1th layer; x j It can represent the output activation value of the jth neuron from the previous layer (i.e., the nth layer); It can represent the bias term; i, j can represent the connection relationship between neurons in different layers.
[0070] In one example, the output neuron may use a sigmoid function as an activation function, which can be expressed by Formula 5.
[0071] Formula 5 In Formula 5, f can represent the simoid function; z i It can represent the total input received by the neuron.
[0072] According to the above implementation, key points of the human body region can be identified, and the centroid coordinates and head turning information can be extracted. Posture recognition is performed based on the centroid coordinates and head turning information to obtain a posture recognition result. In this way, the user's posture can be accurately identified.
[0073] It is understandable that the above-mentioned three-layer BP neural network can be trained by methods known in the art. For example, a depth map and posture labels containing a human body area can be used as training samples, and the four-dimensional feature coordinates corresponding to the depth map can be input into a preset three-layer BP neural network to obtain a posture classification result output by the three-layer BP neural network. The three-layer BP neural network is iteratively trained according to the posture classification results and posture labels until the preset training stop conditions are met to obtain a trained three-layer BP neural network. Exemplarily, the hidden layer output during the training process can be expressed by the following formula 6. The output layer output can be expressed by the following formula 7.
[0074] Formula 6 In formula 6, H j can represent the output activation value of the jth neuron; g can represent the activation function; w ij It can represent the connection weight from the i-th neuron in the previous layer to the j-th neuron in the current layer; x i It can represent the output activation value of the i-th neuron from the previous layer; a j Can represent the bias term; n can represent the number of neurons in the previous layer.
[0075] Formula 7 In formula 7, O k can represent the total input (also called linear combination or weighted sum) of the kth neuron in the output layer; l can represent the number of neurons in the previous layer (e.g., hidden layer); j can represent the index of each neuron in the previous layer; k can represent the kth neuron; H j It can represent the activation value of the jth neuron in the previous layer (such as the hidden layer); w jk It can represent the connection weight from the jth neuron in the previous layer to the kth neuron in the output layer; b k It can represent the bias term of the kth neuron.
[0076] The error E can be calculated by equation 8.
[0077] Formula 8 In formula 8, m can represent the number of neurons in the output layer, that is, the dimension of the model output; Y k It can represent the actual target value or label corresponding to the kth output unit; k It can represent the predicted value of the kth output unit. During the iterative training process, the stochastic gradient descent method can be used to update the weights.
[0078] In one embodiment, the method may further include: When it is detected that the user is in the second area, a video frame queue of the second area in the second target time period is collected by a visible light camera.
[0079] Human body posture recognition is performed on each video frame in the video frame queue respectively to obtain posture recognition results of each video frame.
[0080] The motion state of the user in the second target time period is determined according to the gesture recognition result of each video frame in the video frame queue.
[0081] In the above implementation, the second area may include a public area, for example, a living room or a public activity area in a community. The visible light camera may include a high-definition camera.
[0082] The second target period has a similar meaning to the first target period, and will not be described in detail here. As an example, the visible light camera can capture a video stream at a speed of 30 frames per second and store the most recent 30 frames in a circular queue. The second target period can be the most recent second from the current moment, and the video frame queue can include 30 video frames.
[0083] According to the above implementation, when the user is in the second area, i.e., the public area, the video frame queue can be directly collected through the visible light camera, the video frame can be recognized, and the user's motion state can be determined based on the gesture recognition result of each video frame. In this way, when the user is in a public area, the user's motion state can be monitored directly by monitoring the area where the user is located. In this way, differentiated monitoring of public areas and non-public areas can be achieved, and different technical means can be used for effective monitoring and data analysis according to different monitoring areas, which is conducive to improving the efficiency and flexibility of monitoring.
[0084] In one example, performing human posture recognition on video frames may include performing human posture recognition through a deep learning algorithm.
[0085] For example, OpenPose can be used to detect human postures in images and generate human posture graphs. Specifically, the original color image is used as input, and the key point positions and two-dimensional skeleton graphs of each person in the image are used as output. First, the feedforward network can predict a set of two-dimensional confidence maps S of feature points and a set of human body part association vector fields L. The association vector fields can be used to encode the degree of association between various parts of the human body. In , each key point has A confidence maps, where A is a positive integer. In , each limb part has B associated vector fields, where B is a positive integer. A position L in the set encodes a two-dimensional vector.
[0086] Next, the confidence map and association field can be parsed by a greedy algorithm to output the two-dimensional key points of all people in the image and generate a human posture map. First, the confidence map can be traversed to locate the potential key point positions by finding the local maximum points in the confidence map of each key point, and the NMS technology can be applied to eliminate redundant detections to ensure the uniqueness of each key point. By evaluating the direction and strength of the part affinity field vector (PAF) vector field and the path integral, the greedy algorithm starts from a key point and greedily selects the next key point along the most likely connection direction until the termination condition is met. Among them, the path integral can represent the confidence of the connection between two points. Combining the path integral technology for connection can enhance the accuracy and robustness of the connection. After all key points and the connections between them are determined, the human body region can be assembled into individual human postures according to the human anatomical structure based on the connection results. When dealing with key points that may overlap or conflict, they can be parsed based on prior knowledge of the human anatomical structure or additional constraints to ensure the accuracy and consistency of each individual posture. After parsing the obtained human posture, the human posture can be displayed on the original image in a graphical manner. Specifically, the detected key points can be drawn on the image, and adjacent key points can be connected using lines of different colors to form an intuitive and easy-to-understand human posture diagram.
[0087] Next, the feature information of the video frames in the spatial and temporal domains can be extracted through a three-dimensional (3D) convolutional neural network (CNN) model to classify the human posture in the video.
[0088] Before classifying the human posture in the video through the 3D-CNN model, the 3D-CNN model needs to be trained, and the model training can be achieved by methods known in the art. Exemplarily, the richness and representativeness of the training set can be ensured by collecting and annotating surveillance video data containing various human postures. Subsequently, the video data is preprocessed, including cropping, scaling, and normalization, to adapt to the input requirements of the 3D-CNN model and reduce the computational burden. For the 3D-CNN model, a network structure containing multiple 3D convolutional layers can be constructed, and the convolutional layers can simultaneously capture the spatial and temporal features in the video, thereby effectively extracting feature maps across time dimensions. In order to reduce the dimension of the feature map and enhance the robustness of the model, a 3D pooling layer can be inserted between the convolutional layers. In addition, the expression ability and training efficiency of the model can be further improved by introducing nonlinear activation functions and fully connected layers. During the training process, the cross entropy loss function can be used to evaluate the difference between the model prediction and the actual label, and the model parameters can be updated by the back propagation algorithm and the Adaptive Moment Estimation (Adam) optimizer. Through iterative training, the performance of the model is continuously optimized until a satisfactory classification accuracy is achieved on the validation set. After training, the 3D-CNN model can be deployed in the actual monitoring system to efficiently and accurately classify human postures in real-time video streams.
[0089] For example, during the model training process, video data can be used as input and features can be extracted in the convolution layer. The three-dimensional convolution in the convolution layer is achieved by convolution of the convolution kernel with a cube composed of multiple video frame queues. Each convolution kernel can only extract one local feature. In order to generate multiple local feature maps and enrich feature information, there are generally multiple convolution kernels in the convolution layer. Among them, the value with coordinates (x, y, z) in the jth feature map in the i-th layer It can be calculated by formula 9.
[0090] Formula 9 In Formula 9, tanh() can represent a hyperbolic tangent function; P, Q, and R can represent the height, width, and depth of the convolution kernel, respectively; b can represent the bias of the feature map; ω can represent the connection weight between the convolution kernel and the mth feature map; and p, q, and r can represent the position of the convolution kernel.
[0091] After the video frame queue is calculated by the convolution layer, the data volume increases. The maximum downsampling can be used to reduce the connection between the convolution layers and thus reduce the scale of the feature map to reduce the difficulty of training. The calculation formula of the maximum downsampling can be shown in Formula 10.
[0092] Formula 10 In formula 10, It can represent the output at the (x, y, z) position after the pooling operation; μ can be the three-dimensional input vector of the downsampling layer; s, t, r can be the sampling steps in three directions; i, j, k can represent the relative offset in three directions, which is used to traverse all elements in the maximum downsampling window.
[0093] After the model training is completed, each video frame can be preprocessed and human body detection can be performed, and posture analysis can be performed through the 3D-CNN model. By comparing the changes in human body posture in the head and tail frames of the queue, it can be determined whether a fall has occurred based on the set height change and tilt angle.
[0094] In one example, after determining the motion state of the user in the second target period according to the gesture recognition result of each video frame in the video frame queue, the method may further include: When the user's motion state in the second target time period is the target state, second prompt information is sent to the target client.
[0095] The second prompt information can be used to remind third-party personnel that the user is in an abnormal state, so that the third-party personnel can provide timely assistance to the user. In this way, it is helpful to respond to the user's abnormal situation in a timely manner, thereby protecting the user's safety in a timely manner. In one example, once it is confirmed that the user has sent an abnormal event such as a fall, the system will immediately save the key frame corresponding to the moment the user fell, and send the second prompt information containing information such as time, location and user identity to the third-party client in real time through customized communication software, which is conducive to the third-party personnel to respond quickly and ensure the safety of the user.
[0096] Based on the same inventive concept as the information processing method, an embodiment of the present application also provides an information processing device.
[0097] like Figure 3 As shown, the information processing device 200 may include a first acquisition module 201 , a calculation module 202 , a detection module 203 , a first determination module 204 and a first identification module 205 .
[0098] The first acquisition module 201 is used to acquire infrared speckle array calibration information of the first area in a first target period through an infrared camera when it is detected that the user is in the first area.
[0099] The calculation module 202 is used to perform depth calculation on the infrared speckle array calibration information to obtain a plurality of first depth maps corresponding to the first target time period.
[0100] The detection module 203 is used to perform moving object cluster detection according to the multiple frames of the first depth map to obtain a moving object cluster detection result.
[0101] The first determination module 204 is used to determine the human body area in the first depth map of each frame according to the moving object cluster detection result.
[0102] The first recognition module 205 is used to perform posture recognition on the human body area in the first depth map of each frame to obtain a posture recognition result of the first depth map of each frame.
[0103] The first determination module 204 is further configured to determine the motion state of the user in the first target time period according to the gesture recognition results of the first depth images of each frame.
[0104] The information processing device of the embodiment of the present application can collect infrared speckle array calibration information of the first area in the first target period through an infrared camera when detecting that the user is in the first area. The first area can be a non-public area. The infrared speckle array calibration information mainly reflects the temperature distribution and thermal radiation characteristics of the object, rather than the specific shape or color of the object, thereby avoiding direct exposure of the user's face, body part image, or the user's activity screen in the non-public area and other information that the user does not want to expose. Then, the infrared speckle array calibration information can be depth calculated to obtain a multi-frame first depth map corresponding to the first target period. According to the multi-frame first depth map, moving object cluster detection is performed to obtain a moving object cluster detection result. According to the moving object cluster detection result, the human body area in each frame of the first depth map is determined. The human body area in each frame of the first depth map is respectively subjected to posture recognition to obtain a posture recognition result of each frame of the first depth map. According to the posture recognition result of each frame of the first depth map, the user's motion state in the first target period is determined. In this way, the user's motion state can be accurately monitored while reducing the exposure of user information, which is conducive to rapid response when the user has abnormal behavior and protects the user's safety.
[0105] In one embodiment, the apparatus may further include: The sending module is used to determine the motion state of the user in the first target period according to the posture recognition result of the first depth map of each frame, and then send prompt information to the target client when the motion state of the user in the first target period is the target state.
[0106] In one embodiment, the detection module is used to perform moving object cluster detection according to the multiple frames of the first depth map to obtain the moving object cluster detection result, which may specifically include: The detection submodule is used to perform ground detection and background detection on the first depth map of each frame respectively, and obtain the detection result corresponding to the first depth map of each frame.
[0107] The removal submodule is used to remove the ground area and the background area in the first depth map of each frame according to the detection result corresponding to the first depth map of each frame, and obtain the second depth map corresponding to the first depth map of each frame, and the second depth map includes the foreground area.
[0108] The detection submodule is further used to perform moving object cluster detection on the second depth map corresponding to the first depth map of each frame to obtain a moving object cluster detection result.
[0109] In one embodiment, the detection submodule is used to perform moving object cluster detection on the second depth map corresponding to the first depth map of each frame to obtain a moving object cluster detection result, which may specifically include: The execution unit is used to respectively execute, for each second depth map corresponding to the first depth map of each frame: dividing the voxels in the second depth map into a plurality of groups according to a preset clustering algorithm. Performing connectivity analysis on the plurality of groups to obtain connectivity analysis results. According to the connectivity analysis results, dividing the plurality of groups into a plurality of components.
[0110] The determination unit is used to determine the motion trajectory of each component based on the position information of the component in the second depth map corresponding to the first depth map of each frame.
[0111] The matching unit is used to match multiple components according to the motion trajectory of each component to obtain a matching result.
[0112] The division unit is used to divide the multiple components according to the matching results to obtain multiple moving object clusters.
[0113] In one embodiment, the first identification module may be specifically used for: Key points of the human body region are identified to obtain spatial coordinates of multiple target key points, where the multiple target key points include key points corresponding to multiple human body parts.
[0114] According to the spatial coordinates of multiple target key points, the center of mass coordinates and head turning information are extracted.
[0115] Posture recognition is performed based on the center of mass coordinates and head turning information to obtain a posture recognition result.
[0116] In one embodiment, the apparatus may further include: The second acquisition module is used to acquire a video frame queue of the second area in a second target time period through a visible light camera when it is detected that the user is in the second area.
[0117] The second recognition module is used to perform human posture recognition on each video frame in the video frame queue to obtain posture recognition results of each video frame.
[0118] The second determination module is used to determine the motion state of the user in the second target time period according to the gesture recognition result of each video frame in the video frame queue.
[0119] The information processing device provided in the embodiment of the present application can realize Figure 2 To avoid repetition, the various processes implemented by the method embodiment are not described here.
[0120] Figure 4 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.
[0121] The electronic device may include a processor 301 and a memory 302 storing computer program instructions.
[0122] Specifically, the processor 301 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0123] The memory 302 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 302 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 302 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 302 is a non-volatile solid-state memory.
[0124] The memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.
[0125] The processor 301 implements any one of the information processing methods in the above embodiments by reading and executing computer program instructions stored in the memory 302 .
[0126] As an example, the electronic device may further include a communication interface 303 and a bus 310. Figure 4 As shown, the processor 301, the memory 302, and the communication interface 303 are connected via a bus 310 and communicate with each other.
[0127] The communication interface 303 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0128] The bus 310 includes hardware, software or both, coupling the components of the information processing device to each other. For example and not limitation, the bus may include an accelerated graphics port (AGP) or other graphics bus, an enhanced industry standard architecture (EISA) bus, a front-side bus (FSB), a hypertransport (HT) interconnect, an industry standard architecture (ISA) bus, an infinite bandwidth interconnect, a low pin count (LPC) bus, a memory bus, a microchannel architecture (MCA) bus, a peripheral component interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a serial advanced technology attachment (SATA) bus, a video electronics standard association local (VLB) bus or other suitable bus or a combination of two or more of these. Where appropriate, the bus 310 may include one or more buses. Although the present application embodiment describes and illustrates a specific bus, the present application considers any suitable bus or interconnect.
[0129] The electronic device can execute the information processing method in the embodiment of the present application, thereby realizing the combination Figure 2 and Figure 3 Described information processing method and device.
[0130] In addition, in combination with the information processing method in the above embodiments, the embodiments of the present application may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any one of the information processing methods in the above embodiments is implemented.
[0131] An embodiment of the present application also provides a computer program product, including a computer program, which implements any one of the information processing methods in the above embodiments when the computer program is processed and executed.
[0132] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0133] The functional blocks shown in the structural block diagram described above can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier. "Machine-readable medium" may include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0134] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0135] Aspects of the present disclosure are described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0136] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.
Claims
1. An information processing method, characterized in that: include: When it is detected that the user is in the first area, collecting infrared speckle array calibration information of the first area in the first target time period through an infrared camera; Performing depth calculation on the infrared speckle array calibration information to obtain a plurality of first depth maps corresponding to the first target time period; Perform moving object cluster detection according to the multiple frames of first depth images to obtain a moving object cluster detection result; Determine a human body region in a first depth map of each frame according to the moving object cluster detection result; Performing posture recognition on the human body area in the first depth map of each frame respectively to obtain the posture recognition result of the first depth map of each frame; The motion state of the user in the first target time period is determined according to the posture recognition results of the first depth images of each frame.
2. The method according to claim 1, characterized in that After determining the motion state of the user in the first target period according to the gesture recognition results of the first depth map of each frame, the method further includes: When the user's motion state in the first target time period is the target state, prompt information is sent to the target client.
3. The method according to claim 1 or 2, characterized in that: The performing moving object cluster detection according to the multiple frames of the first depth map to obtain a moving object cluster detection result includes: Performing ground detection and background detection on the first depth map of each frame respectively to obtain detection results corresponding to the first depth map of each frame; According to the detection results corresponding to the first depth map of each frame, the ground area and the background area in the first depth map of each frame are removed to obtain a second depth map corresponding to the first depth map of each frame, wherein the second depth map includes a foreground area; Perform moving object cluster detection on the second depth map corresponding to the first depth map of each frame to obtain a moving object cluster detection result.
4. The method according to claim 1 or 2, characterized in that: The performing moving object cluster detection on the second depth map corresponding to the first depth map of each frame to obtain a moving object cluster detection result includes: For the second depth map corresponding to the first depth map of each frame, respectively: dividing the voxels in the second depth map into a plurality of groups according to a preset clustering algorithm; performing connectivity analysis on the plurality of groups to obtain connectivity analysis results; and dividing the plurality of groups into a plurality of components according to the connectivity analysis results; For each component, determining a motion trajectory of each component based on position information of the component in a second depth map corresponding to the first depth map of each frame; Matching the multiple components according to the motion trajectory of each component to obtain a matching result; According to the matching results, the multiple components are divided to obtain multiple moving object clusters.
5. The method according to claim 1 or 2, characterized in that: The step of performing posture recognition on the human body region in the first depth map of each frame to obtain the posture recognition result of the first depth map of each frame includes performing the following steps on the human body region in the first depth map of each frame: Performing key point recognition on the human body region to obtain spatial coordinates of a plurality of target key points, wherein the plurality of target key points include key points corresponding to a plurality of human body parts; Extracting centroid coordinates and head turning information according to the spatial coordinates of the multiple target key points; Perform posture recognition according to the center of mass coordinates and the head turning information to obtain a posture recognition result.
6. The method according to claim 1 or 2, characterized in that: The method further comprises: When it is detected that the user is in the second area, collecting a video frame queue of the second area in the second target time period through a visible light camera; Performing human posture recognition on each video frame in the video frame queue respectively to obtain posture recognition results of each video frame; The motion state of the user in the second target time period is determined according to the gesture recognition result of each video frame in the video frame queue.
7. An information processing device, characterized in that: include: A collection module, configured to collect infrared speckle array calibration information of the first area in a first target period through an infrared camera when detecting that the user is in the first area; A calculation module, configured to perform depth calculation on the infrared speckle array calibration information to obtain a plurality of first depth maps corresponding to the first target time period; A detection module, configured to perform moving object cluster detection according to the plurality of first depth maps to obtain a moving object cluster detection result; A determination module, configured to determine a human body region in the first depth map of each frame according to the moving object cluster detection result; A recognition module, used to perform posture recognition on the human body area in the first depth map of each frame respectively, to obtain the posture recognition result of the first depth map of each frame; The determination module is further used to determine the motion state of the user in the first target time period according to the posture recognition results of the first depth map of each frame.
8. An electronic device, characterized in that: The device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the information processing method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the information processing method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the information processing method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Body action identification method and system based on depth image induction
CN102831380A
Posture identification method and device based on near-infrared TOF camera depth information
CN104463146A
Sitting posture information generation method and device, terminal equipment and storage medium
CN112712053A
Sitting posture geometric parameter detection method based on an RGBD image
CN113065532A
Table lamp capable of recognizing sitting posture of child and conducting automatic voice correction
CN113989936A