Object recognition method and system, electronic equipment and storage medium

By simulating human visual perception and cognitive reasoning, and utilizing deep learning models and human keypoint detection technology, the search radius is dynamically adjusted to solve the occlusion problem in object recognition and pose estimation, providing more accurate image analysis.

CN120877259APending Publication Date: 2025-10-31BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410535928.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-29
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies are susceptible to interference from occluded objects in complex scenes, pose estimation accuracy decreases in multi-person scenarios, lack cognitive understanding of object functionality, and are insufficiently adaptable to scale changes in dynamic scenes.

Method used

By simulating human visual perception and cognitive reasoning, the system uses a deep learning model to identify target objects, combines human keypoint detection technology to determine posture, and searches for target objects centered on human keypoints, dynamically adjusting the search radius to adapt to visual effects, and inferring the relationship between the human body and objects.

Benefits of technology

It provides more accurate and comprehensive image analysis in complex scenes, identifies obvious objects and infers situations where visual information is incomplete, solves occlusion problems, and adapts to the visual phenomenon of objects appearing larger when closer and smaller when farther away.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877259A_ABST
    Figure CN120877259A_ABST
Patent Text Reader

Abstract

The invention discloses an object recognition method and system, electronic equipment and a storage medium, and belongs to the technical field of computer vision. The object identification method comprises the following steps: identifying target objects in an input image by using a deep learning model, and marking all visible target objects; identifying key points of a human body in the input image by adopting a human body key point detection technology; determining the posture of each human body based on the key points; and according to the key points and postures of each human body and the position information of all visible target objects, reasoning the relationship between the human body and the target object in the input image. According to the method, the limitation of a traditional computer vision technology in object recognition and human body posture analysis is solved by simulating the visual perception and cognitive reasoning ability of human beings, the visibility of the object in the image is considered, the relation between the human body and the object is analyzed from the perspective of functionality, and more accurate and more comprehensive image analysis can be provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer vision technology, and in particular relates to an object recognition method, system, electronic device and storage medium. Background Technology

[0002] In computer vision, object recognition, especially chair recognition in images and videos, typically relies on deep learning models such as convolutional neural networks (CNNs). These models are trained to recognize objects like chairs in various environments. Pose estimation techniques focus on detecting and tracking key points of the human body, such as the head, shoulders, and hips, to analyze human posture.

[0003] Despite the progress made in their respective fields, object recognition technology cannot effectively handle occlusion problems in complex scenes. The recognition results are easily affected by occluded objects. When pose estimation technology is applied in multi-person scenes, especially when parts of the human body are occluded, the accuracy will decrease. Summary of the Invention

[0004] This application aims to address at least one of the technical problems existing in related technologies. To this end, this application proposes an object recognition method, system, electronic device, and storage medium that, by simulating human visual perception and cognitive reasoning abilities, overcomes the limitations of traditional computer vision technology in object recognition and human posture analysis. It not only considers the visibility of objects in the image but also analyzes the relationship between the human body and objects from a functional perspective, providing more accurate and comprehensive image analysis.

[0005] In a first aspect, this application provides an object recognition method, the method comprising:

[0006] The deep learning model is used to identify target objects in the input image and mark all visible target objects.

[0007] Human key point detection technology is used to identify key points of the human body in the input image;

[0008] Based on the aforementioned key points, the posture of each human body is determined;

[0009] Based on the key points and posture of each human body and the positional information of all visible target objects, the relationship between the human body and target objects in the input image is inferred.

[0010] The object recognition method of this application identifies target objects in an input image using a deep learning model, marking all visible target objects. It then employs human keypoint detection technology to identify key points of the human body in the input image. Based on these key points, the posture of each human body in the input image is determined. Finally, based on the key points, posture, and position information of all visible target objects, the relationship between the human body and target objects in the input image is inferred. By simulating human visual perception and cognitive reasoning abilities, this method overcomes the limitations of traditional computer vision technology in object recognition and human posture analysis. It not only considers the visibility of objects in the image but also analyzes the relationship between the human body and objects from a functional perspective. When processing images, it more closely approximates human perception and cognitive processes, not only recognizing obvious objects but also inferring situations where visual information is incomplete, thus providing more accurate and comprehensive image analysis results.

[0011] According to one embodiment of this application, determining the posture of each human body based on the key points includes:

[0012] The posture of each human body is determined by the relative position and angle between key points detected in each human body; the posture includes standing or sitting posture.

[0013] According to one embodiment of this application, the step of inferring the relationship between the human body and target objects in the image based on the key points, posture, and position information of each human body and all visible target objects includes:

[0014] Using the hip bone key point of each human body in a sitting posture as the center, search among all visible target objects with a first preset radius, and match the target object that is closest to the human body. If a match is found, the matched target object is marked as the assigned target object.

[0015] If no match is found, the search continues among all visible target objects with a second preset radius. If no target object is still found, it is determined that the target object matching the corresponding human body is occluded, and the second preset radius is equal to a preset multiple of the first preset radius.

[0016] The object recognition method provided in this application uses the hip bone key point of each human body in a sitting posture as the center, and searches among all visible target objects with a first preset radius to match the target object closest to the human body. If no match is found, the search continues among all visible target objects with a second preset radius. If no match is found, it is determined that the target object matched for the corresponding human body is occluded. When processing images, it can more closely resemble the human perception and cognition process. It can not only identify obvious objects, but also infer situations where visual information is incomplete, thereby solving the occlusion problem in complex scenes and providing more accurate and comprehensive image analysis results.

[0017] According to one embodiment of this application, the first preset radius includes: half the distance from the shoulder key point to the hip key point of the corresponding human body.

[0018] This application embodiment sets the first preset radius to half the distance from the shoulder key point to the hip key point of the corresponding human body, so that the size of the search area can be dynamically adjusted according to the distance from the shoulder to the hip of different human bodies, thereby adapting to the visual effect of near objects appearing larger and distant objects appearing smaller in the image, making the analysis of the image more in line with the human perception and cognitive process.

[0019] According to one embodiment of this application, the method further includes:

[0020] Count the number of all visible target objects and the number of occluded target objects;

[0021] The total number of all target objects in the input image is obtained based on the number of all visible target objects and the number of occluded target objects.

[0022] The object recognition method provided in this application not only considers the visibility of target objects in the image, but also analyzes the relationship between the human body and the target object from a functional perspective to infer and recognize the target object, and then infers the total number of target objects in the image, thus solving the occlusion problem in complex scenes.

[0023] According to one embodiment of this application, the target object is a chair.

[0024] Secondly, this application provides an object recognition system, which includes:

[0025] The object recognition unit is used to identify target objects in the input image using a deep learning model and to mark all visible target objects.

[0026] A key point recognition unit is used to identify key points of the human body in the input image using human key point detection technology.

[0027] A posture classification unit is used to determine the posture of each human body based on the key points;

[0028] The reasoning unit is used to infer the relationship between the human body and target objects based on the key points, posture, and positional information of each human body and all visible target objects.

[0029] The object recognition device provided in this application uses a deep learning model to identify target objects in an input image, marking all visible target objects. It employs human keypoint detection technology to identify key points of the human body in the input image. Based on these key points, it determines the posture of each human body in the input image. Finally, based on the key points, posture, and position information of all visible target objects, it infers the relationship between the human body and target objects in the input image. By simulating human visual perception and cognitive reasoning abilities, it overcomes the limitations of traditional computer vision technology in object recognition and human posture analysis. It not only considers the visibility of objects in the image but also analyzes the relationship between the human body and objects from a functional perspective. When processing images, it more closely approximates human perception and cognitive processes, not only recognizing obvious objects but also inferring situations where visual information is incomplete, thereby providing more accurate and comprehensive image analysis results.

[0030] According to one embodiment of this application, the posture classification unit is used for:

[0031] The posture of each human body is determined by the relative position and angle between key points detected in each human body; the posture includes standing or sitting posture.

[0032] According to one embodiment of this application, the inference unit is used for:

[0033] Using the hip bone key point of each human body in a sitting posture as the center, search among all visible target objects with a first preset radius, and match the target object that is closest to the human body. If a match is found, the matched target object is marked as the assigned target object.

[0034] If no match is found, the search continues among all visible target objects with a second preset radius. If no target object is still found, it is determined that the target object matching the corresponding human body is occluded, and the second preset radius is equal to a preset multiple of the first preset radius.

[0035] The object recognition device provided in this application uses the hip bone key point of each human body in a sitting posture as the center and searches among all visible target objects with a first preset radius to match the target object closest to the human body. If no match is found, the search continues among all visible target objects with a second preset radius. If no match is found, it is determined that the target object matched for the corresponding human body is occluded. When processing images, it can be closer to the human perception and cognition process. It can not only identify obvious objects, but also infer the situation when visual information is incomplete, thereby providing more accurate and comprehensive image analysis results.

[0036] According to one embodiment of this application, the first preset radius includes: half the distance from the shoulder key point to the hip key point of the corresponding human body.

[0037] This application embodiment sets the first preset radius to half the distance from the shoulder key point to the hip key point of the corresponding human body, so that the size of the search area can be dynamically adjusted according to the distance from the shoulder to the hip of different human bodies, thereby adapting to the visual effect of near objects appearing larger and distant objects appearing smaller in the image, making the analysis of the image more in line with the human perception and cognitive process.

[0038] According to one embodiment of this application, the system further includes a statistical unit, the statistical unit being used for:

[0039] Count the number of all visible target objects and the number of occluded target objects;

[0040] The total number of all target objects in the input image is obtained based on the number of all visible target objects and the number of occluded target objects.

[0041] The object recognition device provided in this application not only considers the visibility of target objects in the image, but also analyzes the relationship between the human body and the target object from a functional perspective to infer and recognize the target object, and then infers the total number of target objects in the image, thus solving the occlusion problem in complex scenes.

[0042] According to one embodiment of this application, the target object is a chair.

[0043] Thirdly, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the object recognition method as described in the first aspect above.

[0044] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the object recognition method as described in the first aspect above.

[0045] Fifthly, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the object recognition method as described in the first aspect.

[0046] In a sixth aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the object recognition method as described in the first aspect above.

[0047] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0048] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0049] Figure 1 This is one of the flowcharts illustrating the object recognition method provided in the embodiments of this application;

[0050] Figure 2 A second schematic flowchart illustrating the object recognition method provided in this application embodiment;

[0051] Figure 3 A schematic diagram of the input image provided in the embodiments of this application;

[0052] Figure 4 This is a schematic diagram illustrating the identification and marking of all visible chairs in an input image, as provided in an embodiment of this application.

[0053] Figure 5 A schematic diagram of key points of the human body in an input image identified according to an embodiment of this application;

[0054] Figure 6 A schematic diagram of human sitting posture detection provided in an embodiment of this application;

[0055] Figure 7 A schematic diagram of human standing posture detection provided in an embodiment of this application;

[0056] Figure 8 A schematic diagram illustrating the matching of a human body and a chair, provided for an embodiment of this application;

[0057] Figure 9 This is a schematic diagram of the structure of the object recognition device provided in the embodiments of this application;

[0058] Figure 10 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0059] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0060] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0061] In computer vision, object recognition, especially chair recognition in images and videos, typically relies on deep learning models such as convolutional neural networks (CNNs). These models are trained to recognize objects like chairs in various environments. Pose estimation focuses on detecting and tracking key points of the human body, such as the head, shoulders, and hips, to analyze posture. Current market products and research largely concentrate on single object recognition or pose estimation, and are primarily used in motion capture, augmented reality, and health monitoring.

[0062] Existing invention patents frequently explore technologies that combine object recognition with human posture estimation. For example, some patents propose using keypoint detection to improve the accuracy of object recognition, or using object recognition to assist in more accurate posture estimation. These patents typically focus on specific application scenarios, such as intelligent monitoring and interaction design. They improve recognition accuracy and computational efficiency through innovative algorithms and model architectures.

[0063] Despite advancements in their respective fields, several common shortcomings remain in object recognition technologies. First, the occlusion problem in complex scenes remains unresolved, making recognition results susceptible to interference from occluded objects. Second, pose estimation techniques often suffer from decreased accuracy in multi-person scenarios, particularly when human figures are partially occluded. Furthermore, these technologies lack a cognitive understanding of object functionality, failing to infer and identify objects from a functional perspective, such as whether a chair is in use. Finally, existing systems are inefficient at handling scale variations in images, especially in dynamic scenes where they are ill-suited to the visual phenomenon of objects appearing larger when closer and smaller when farther away.

[0064] To at least address one of the technical problems existing in related technologies, embodiments of this application provide an object recognition method, system, electronic device, and storage medium. The object recognition method, system, electronic device, and storage medium provided in this application will be described in detail below with reference to the accompanying drawings and specific embodiments and application scenarios.

[0065] Among them, the object recognition method can be applied to the terminal, and can be executed by the hardware or software in the terminal.

[0066] The object recognition method provided in this application embodiment can be executed by an electronic device or a functional module or entity in an electronic device that can implement the object recognition method. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras, and wearable devices. The object recognition method provided in this application embodiment will be described below using an electronic device as the execution subject.

[0067] Figure 1 This is one of the flowcharts illustrating the object recognition method provided in this application. Figure 1 As shown, the object recognition method includes steps 110, 120, 130 and 140.

[0068] Step 110: Use a deep learning model to identify target objects in the input image and mark all visible target objects;

[0069] Deep learning models, based on artificial neural networks (ANNs), are machine learning techniques that simulate the human brain's ability to perform high-level abstract processing of data. They typically have multiple hidden layers (i.e., deep structures) to learn complex feature representations of input data. Deep learning models have achieved great success in the field of machine learning and are widely used in various fields such as image recognition, speech recognition, and natural language processing.

[0070] Deep learning models can be trained to learn patterns and features in data, enabling prediction and classification for various tasks. Common models in deep learning include Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM) networks, and Transformers.

[0071] Specifically, deep learning models can be categorized as follows:

[0072] Supervised neural networks require labeled data during training; that is, both input and expected output data are provided simultaneously. The model learns the relationships between these data pairs to make predictions. Common supervised networks include Convolutional Neural Networks (CNNs) and fully connected networks. Training supervised neural networks typically employs backpropagation to optimize model parameters, allowing the model to continuously adjust and minimize the error between the predicted output and the actual label.

[0073] Recurrent Neural Networks (RNNs) are capable of processing sequential data, such as time series or natural language text. They are characterized by circular connections within the network, which allow for the transmission of state information, making the network sensitive to both preceding and following elements in the sequence.

[0074] Unsupervised learning neural networks: Unsupervised learning does not require labeled data. The model extracts features or generates data by learning the structure and distribution of the data itself, such as autoencoders and generative adversarial networks (GANs).

[0075] This application embodiment uses a deep learning model to identify target objects in the input image and marks all visible target objects.

[0076] The input image can be one or more frames of continuous or discontinuous images from a video captured by a camera, and the input image must contain at least a person and a target object.

[0077] Specifically, by using a deep learning model to identify target objects in the input image, all visible target objects are marked, including:

[0078] By using a pre-trained target object recognition model, target objects are identified in the input image to obtain the location information of all visible target objects in the input image.

[0079] Step 120: Use human key point detection technology to identify key points of the human body in the input image;

[0080] Human keypoint detection technology is used to locate key parts of the human body in an input image, such as the head, shoulders, elbows, wrists, and hips. This technology has wide applications in computer vision and artificial intelligence, including motion capture, pose recognition, and human tracking. Generally, the human keypoint detection process includes the following steps:

[0081] Data preparation: Collect and prepare labeled human pose datasets, ensuring that the datasets contain human images of various poses and angles, and that the key points of the human body are labeled.

[0082] Model selection: Choose a suitable deep learning model for human keypoint detection. Commonly used models include Hourglass network, OpenPose, PoseNet, etc. These models are usually based on convolutional neural networks or recurrent neural networks with temporal modeling capabilities.

[0083] Model training: The selected deep learning model is trained using the prepared dataset to accurately identify human keypoints in images. During training, model parameters need to be adjusted to minimize the loss function, typically using backpropagation for optimization.

[0084] Model evaluation: The trained model is evaluated using an independent validation dataset to check its performance on unseen data. Evaluation metrics may include the accuracy of keypoint localization, robustness, etc.

[0085] Inference and Application: The trained model is applied to real image data to detect key points of the human body, locate key parts of the human body, and output the corresponding results.

[0086] Step 130: Based on the key points, determine the posture of each human body;

[0087] In some embodiments, the pose classification model obtained through training can be used to classify the key points of the human body in the input image to obtain the pose of each human body.

[0088] The pose classification model is pre-trained using the following method:

[0089] Construct a sample set, which includes multiple images labeled with key points of the human body and corresponding pose labels;

[0090] The original pose classification model is trained using the sample set. Training is stopped when the training termination condition is met, and the pose classification model is obtained.

[0091] In some embodiments, step 130 includes:

[0092] The posture of each human body is determined by the relative position and angle between key points detected in each human body; the posture includes standing or sitting posture.

[0093] That is, the posture of each human body can be determined based on predefined posture standards by detecting the relative positions and angles between key points of each human body.

[0094] Posture includes standing or sitting posture.

[0095] Predefined posture standards refer to the predefined relative positions and angles between key points of the human body in a standing posture, and the relative positions and angles between key points of the human body in a sitting posture.

[0096] For example, for each key point of the human body detected in the input image, the relative positions and angles between the key points of the human body are compared with the relative positions and angles between the key points of the human body in the standing state and the relative positions and angles between the key points of the human body in the sitting state. If the similarity is higher than that in the standing state, the human body is determined to be in a standing state; if the similarity is higher than that in the sitting state, the human body is determined to be in a sitting state.

[0097] Step 140: Based on the key points, postures, and positional information of each human body and all visible target objects, infer the relationship between the human body and target objects in the input image.

[0098] After obtaining the position information of all visible target objects in the input image in step 110, obtaining the key points of the human body in the input image in step 120, and determining the posture of each human body in the input image in step 130, the relationship between the human body and the target objects in the input image is inferred based on the key points, posture, and position information of all visible target objects of each human body. This allows for the analysis of the relationship between the human body and objects from a functional perspective, thereby enabling the inference of situations where visual information is incomplete and providing more accurate and comprehensive image analysis results.

[0099] The object recognition method provided in this application utilizes a deep learning model to identify target objects in an input image, marking all visible target objects. It then employs human keypoint detection technology to identify key points of the human body in the input image. Based on these key points, the posture of each human body in the input image is determined. Finally, based on the key points, posture, and position information of all visible target objects, the relationship between the human body and target objects in the input image is inferred. By simulating human visual perception and cognitive reasoning abilities, this method overcomes the limitations of traditional computer vision technology in object recognition and human posture analysis. It not only considers the visibility of objects in the image but also analyzes the relationship between the human body and objects from a functional perspective. When processing images, it more closely approximates human perception and cognitive processes, not only recognizing obvious objects but also inferring situations where visual information is incomplete, thereby providing more accurate and comprehensive image analysis results.

[0100] In some embodiments, step 140, based on the key points, posture, and positional information of each human body and all visible target objects, infers the relationship between the human body and target objects in the input image, including:

[0101] Using the hip bone key point of each human body in a sitting posture as the center, search among all visible target objects with a first preset radius, and match the target object that is closest to the human body. If a match is found, the matched target object is marked as the assigned target object.

[0102] If no match is found, the search continues among all visible target objects with a second preset radius. If no target object is still found, it is determined that the target object matching the corresponding human body is occluded.

[0103] This application embodiment further processes the human body in a sitting posture. Specifically, it searches among all visible target objects with a first preset radius, using the hip bone key point of each human body in a sitting posture as the center, and matches the target object that is closest to the human body.

[0104] For a human body in a seated posture, the search is conducted among all visible target objects with the hip bone as the center and a first preset radius. If the target object that is closest to the human body can be matched, the matched target object is marked as an assigned target object, which means that the matched target object is used by the human body and can be attributed to the human body.

[0105] For a human body in a seated posture, the search is performed on all visible target objects with the hip bone as the center and a first preset radius. If no target object closest to the human body can be found, the search range is expanded. Specifically, the search continues with a second preset radius among all visible target objects. If no target object is found, it means that the target object used by the human body is occluded, and the target object matched by the corresponding human body is determined to be occluded.

[0106] The second preset radius is a preset multiple of the first preset radius. For example, the second preset radius = 1.5 * the first preset radius.

[0107] The object recognition method provided in this application uses the hip bone key point of each human body in a sitting posture as the center, and searches among all visible target objects with a first preset radius to match the target object closest to the human body. If no match is found, the search continues among all visible target objects with a second preset radius. If no match is found, it is determined that the target object matched for the corresponding human body is occluded. When processing images, it can more closely resemble the human perception and cognition process. It can not only identify obvious objects, but also infer situations where visual information is incomplete, thereby solving the occlusion problem in complex scenes and providing more accurate and comprehensive image analysis results.

[0108] In some embodiments, the first preset radius includes half the distance from the shoulder key point to the hip key point of the corresponding human body.

[0109] This application embodiment sets the first preset radius to half the distance from the shoulder key point to the hip key point of the corresponding human body, so that the size of the search area can be dynamically adjusted according to the distance from the shoulder to the hip of different human bodies, thereby adapting to the visual effect of near objects appearing larger and distant objects appearing smaller in the image, making the analysis of the image more in line with the human perception and cognitive process.

[0110] In some embodiments, the method further includes:

[0111] Count the number of all visible target objects and the number of occluded target objects;

[0112] The total number of all target objects in the input image is obtained based on the number of all visible target objects and the number of occluded target objects.

[0113] Based on the description in the foregoing embodiments, after inferring the relationship between the human body and target objects in the input image based on the key points, posture, and position information of each human body and all visible target objects, it is possible to determine whether there are any occluded target objects, and then the number of occluded target objects can be counted.

[0114] By using a deep learning model to identify target objects in the input image, all visible target objects are marked, and thus the total number of all visible target objects can be obtained.

[0115] The total number of all objects in the input image can be obtained by adding the number of all visible objects to the number of occluded objects.

[0116] The object recognition method provided in this application not only considers the visibility of target objects in the image, but also analyzes the relationship between the human body and the target object from a functional perspective to infer and recognize the target object, and then infers the total number of target objects in the image, thus solving the occlusion problem in complex scenes.

[0117] In some embodiments, the target object is a chair.

[0118] Figure 2 This is a second schematic flowchart illustrating the object recognition method provided in an embodiment of this application. Figure 2 As shown, the object recognition method includes: using a deep learning model to identify chairs in the input image and marking all visible chairs; using human key point detection technology to identify key points of the human body in the input image; determining the posture of each human body based on the key points; and inferring the relationship between the human body and chairs in the input image based on the key points, posture, and position information of all visible target objects of each human body.

[0119] Figure 3 This is a schematic diagram of the input image provided in an embodiment of this application. Figure 4 This is a schematic diagram illustrating the identification and marking of all visible chairs in an input image, as provided in an embodiment of this application. Figure 5 This is a schematic diagram of key points of the human body in the input image identified according to an embodiment of this application. Figure 6 This is a schematic diagram of human sitting posture detection provided in an embodiment of this application. Figure 7 This is a schematic diagram of human posture detection provided in an embodiment of this application. Figure 8 This is a schematic diagram illustrating the matching of a human body and a chair, provided as an embodiment of this application. Figures 3-8 This allows for a further understanding of the object recognition method provided in the embodiments of this application.

[0120] The object recognition method provided in this application uses a deep learning model to identify chairs in an input image, marking all visible chairs. It then employs human keypoint detection technology to identify key points of the human body in the input image. Based on these key points, the posture of each human body in the input image is determined. Finally, based on the key points, posture, and position information of all visible chairs, the relationship between the human body and chairs in the input image is inferred. By simulating human visual perception and cognitive reasoning abilities, this method overcomes the limitations of traditional computer vision technology in object recognition and human posture analysis. It not only considers the visibility of chairs in the image but also analyzes the relationship between the human body and chairs from a functional perspective. When processing images, it more closely approximates human perception and cognitive processes, not only recognizing obvious objects but also inferring situations where visual information is incomplete, thereby providing more accurate and comprehensive image analysis results.

[0121] The object recognition method provided in this application can be executed by an object recognition device. This application uses an object recognition device executing the object recognition method as an example to illustrate the object recognition device provided in this application.

[0122] This application also provides an object recognition device.

[0123] Figure 9 This is a schematic diagram of the structure of the object recognition device provided in an embodiment of this application. Figure 9 As shown, the object recognition device includes: an object recognition unit 310, a key point recognition unit 320, a posture classification unit 330, and a reasoning unit 340.

[0124] The object recognition unit 310 is used to identify target objects in the input image using a deep learning model and mark all visible target objects.

[0125] Key point recognition unit 320 is used to identify key points of the human body in the input image using human key point detection technology;

[0126] The posture classification unit 330 is used to determine the posture of each human body based on the key points.

[0127] The reasoning unit 340 is used to infer the relationship between the human body and the target objects based on the key points, posture, and position information of each human body and all visible target objects.

[0128] The object recognition device provided in this application uses a deep learning model to identify target objects in an input image, marking all visible target objects. It employs human keypoint detection technology to identify key points of the human body in the input image. Based on these key points, it determines the posture of each human body in the input image. Finally, based on the key points, posture, and position information of all visible target objects, it infers the relationship between the human body and target objects in the input image. By simulating human visual perception and cognitive reasoning abilities, it overcomes the limitations of traditional computer vision technology in object recognition and human posture analysis. It not only considers the visibility of objects in the image but also analyzes the relationship between the human body and objects from a functional perspective. When processing images, it more closely approximates human perception and cognitive processes, not only recognizing obvious objects but also inferring situations where visual information is incomplete, thereby providing more accurate and comprehensive image analysis results.

[0129] In some embodiments, the attitude classification unit is used for:

[0130] The posture of each human body is determined by the relative position and angle between key points detected in each human body; the posture includes standing or sitting posture.

[0131] In some embodiments, the inference unit is used for:

[0132] Using the hip bone key point of each human body in a sitting posture as the center, search among all visible target objects with a first preset radius, and match the target object that is closest to the human body. If a match is found, the matched target object is marked as the assigned target object.

[0133] If no match is found, the search continues among all visible target objects with a second preset radius. If no target object is still found, it is determined that the target object matching the corresponding human body is occluded, and the second preset radius is equal to a preset multiple of the first preset radius.

[0134] The object recognition device provided in this application uses the hip bone key point of each human body in a sitting posture as the center and searches among all visible target objects with a first preset radius to match the target object closest to the human body. If no match is found, the search continues among all visible target objects with a second preset radius. If no match is found, it is determined that the target object matched for the corresponding human body is occluded. When processing images, it can be closer to the human perception and cognition process. It can not only identify obvious objects, but also infer the situation when visual information is incomplete, thereby providing more accurate and comprehensive image analysis results.

[0135] In some embodiments, the first preset radius includes half the distance from the shoulder key point to the hip key point of the corresponding human body.

[0136] This application embodiment sets the first preset radius to half the distance from the shoulder key point to the hip key point of the corresponding human body, so that the size of the search area can be dynamically adjusted according to the distance from the shoulder to the hip of different human bodies, thereby adapting to the visual effect of near objects appearing larger and distant objects appearing smaller in the image, making the analysis of the image more in line with the human perception and cognitive process.

[0137] In some embodiments, the system further includes a statistical unit, the statistical unit being used for:

[0138] Count the number of all visible target objects and the number of occluded target objects;

[0139] The total number of all target objects in the input image is obtained based on the number of all visible target objects and the number of occluded target objects.

[0140] The object recognition device provided in this application not only considers the visibility of target objects in the image, but also analyzes the relationship between the human body and the target object from a functional perspective to infer and recognize the target object, and then infers the total number of target objects in the image, thus solving the occlusion problem in complex scenes.

[0141] In some embodiments, the target object is a chair.

[0142] The object recognition device provided in this application uses a deep learning model to identify chairs in an input image, marking all visible chairs. It then employs human keypoint detection technology to identify key points of the human body in the input image. Based on these key points, it determines the posture of each human body in the input image. Finally, based on the key points, posture, and position information of all visible chairs, it infers the relationship between the human body and chairs in the input image. By simulating human visual perception and cognitive reasoning abilities, it overcomes the limitations of traditional computer vision technology in object recognition and human posture analysis. It not only considers the visibility of chairs in the image but also analyzes the relationship between the human body and chairs from a functional perspective. When processing images, it more closely approximates human perception and cognitive processes, not only recognizing obvious objects but also inferring situations where visual information is incomplete, thereby providing more accurate and comprehensive image analysis results.

[0143] The object recognition device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the specific device.

[0144] The object recognition device in this application embodiment can be a device with an operating system. This operating system can be a Microsoft (Windows) operating system, an Android operating system, an iOS operating system, or other possible operating systems; this application embodiment does not specifically limit the specific operating system.

[0145] The object recognition device provided in this application embodiment can achieve... Figures 1 to 2 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0146] In some embodiments, such as Figure 10 As shown, this application embodiment also provides an electronic device 400, including a processor 401, a memory 402, and a computer program stored in the memory 402 and executable on the processor 401. When the program is executed by the processor 401, it implements the various processes of the above-described object recognition method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0147] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0148] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described object recognition method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0149] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0150] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described object recognition method.

[0151] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0152] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described object recognition method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0153] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0154] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0155] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0156] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0157] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0158] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. An object recognition method, characterized in that, include: The deep learning model is used to identify target objects in the input image and mark all visible target objects. Human key point detection technology is used to identify key points of the human body in the input image; Based on the aforementioned key points, the posture of each human body is determined; Based on the key points and posture of each human body and the positional information of all visible target objects, the relationship between the human body and target objects in the input image is inferred.

2. The object recognition method according to claim 1, characterized in that, Determining the posture of each human body based on the key points includes: The posture of each human body is determined by the relative position and angle between key points detected in each human body; the posture includes standing or sitting posture.

3. The object recognition method according to claim 1, characterized in that, The step of inferring the relationship between the human body and target objects in the image based on the key points, posture, and position information of each human body and all visible target objects includes: Using the hip bone key point of each human body in a sitting posture as the center, search among all visible target objects with a first preset radius, and match the target object that is closest to the human body. If a match is found, the matched target object is marked as the assigned target object. If no match is found, the search continues among all visible target objects with a second preset radius. If no target object is still found, it is determined that the target object matching the corresponding human body is occluded, and the second preset radius is equal to a preset multiple of the first preset radius.

4. The object recognition method according to claim 3, characterized in that, The first preset radius includes half the distance from the shoulder key point to the hip key point of the corresponding human body.

5. The object recognition method according to claim 3 or 4, characterized in that, The method further includes: Count the number of all visible target objects and the number of occluded target objects; The total number of all target objects in the input image is obtained based on the number of all visible target objects and the number of occluded target objects.

6. The object recognition method according to any one of claims 1 to 4, characterized in that, The target object is a chair.

7. An object recognition system, characterized in that, include: The object recognition unit is used to identify target objects in the input image using a deep learning model and to mark all visible target objects. A key point recognition unit is used to identify key points of the human body in the input image using human key point detection technology. A posture classification unit is used to determine the posture of each human body based on the key points; The reasoning unit is used to infer the relationship between the human body and target objects based on the key points, posture, and positional information of each human body and all visible target objects.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the object recognition method as described in any one of claims 1-6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the object recognition method as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the object recognition method as described in any one of claims 1-6.