Object recognition method, device, computer-readable storage medium and processor
By using multiple detection boxes to detect target parts and associating the state recognition results with object recognition, the problem of low object recognition accuracy is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202011528573.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-22
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2040-12-22
AI Technical Summary
In existing technologies, when multiple objects move and cluster together, the key points of the skeleton are not accurately located, resulting in low object recognition accuracy.
By determining multiple detection boxes for the target object, the target parts are detected based on the detection boxes to obtain multiple state recognition results. The first state recognition result is then associated with the second state recognition result to achieve accurate recognition.
It improves the accuracy of object recognition, ensuring accurate association and identification of target objects.
Smart Images

Figure CN114663933B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computer, in particular to an object recognition method and device, computer readable storage medium and processor. BACKGROUND
[0002] At present, when an object is recognized, a method of recognizing skeleton key points is usually used to recognize the behavior of the object.
[0003] The above method can realize the recognition of the object, but when multiple objects frequently move and gather together, the key points cannot be positioned accurately, thereby seriously affecting the accuracy of recognizing the object.
[0004] At present, no effective solution has been proposed to solve the problem of low accuracy of recognizing the object. SUMMARY
[0005] Embodiments of the present application provide an object recognition method and device, computer readable storage medium and processor to at least solve the technical problem of low accuracy of recognizing the object.
[0006] According to an aspect of an embodiment of the present application, an object recognition method is provided. The method can include: determining a plurality of detection boxes of a target object, wherein the detection boxes are used to detect target parts corresponding to a state of the target object; detecting the corresponding target parts based on the plurality of detection boxes respectively, and recognizing the state of the target object based on the obtained detection results to obtain a plurality of state recognition results, wherein the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object; and associating the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0007] According to another aspect of the embodiments of the present application, another object recognition method is also provided. The method can include: a teaching management system displaying a plurality of detection boxes of a target object on a display interface, wherein the detection boxes are used to detect target parts corresponding to a state of the target object; the teaching management system receiving a recognition request and displaying a plurality of state recognition results on the display interface, wherein the plurality of state recognition results are obtained by the teaching management system based on the plurality of detection boxes detecting the state of the target object respectively, and the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object; and the teaching management system displaying a correlation result on the display interface, wherein the correlation result is a result of correlating the first state recognition result and a second state recognition result, and the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0008] According to another aspect of the embodiments of the present application, another object recognition method is also provided. The method can include: displaying a plurality of detection boxes of a target object on an operation interface, wherein the detection boxes are used to detect target parts corresponding to a state of the target object; sensing a recognition operation instruction on the operation interface and displaying a plurality of state recognition results, wherein the plurality of state recognition results are obtained by detecting the state of the target object based on the plurality of detection boxes respectively, and the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object; and displaying a correlation result on the operation interface, wherein the correlation result is a result of correlating the first state recognition result and a second state recognition result, and the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0009] According to another aspect of the embodiments of the present application, another object recognition method is also provided. The method can include: a front-end client uploading a plurality of detection boxes of a target object, wherein the detection boxes are used to detect target parts corresponding to a state of the target object; the front-end client transmitting the plurality of detection boxes to a back-end server; the front-end client receiving a plurality of state recognition results returned by the back-end server, wherein the plurality of state recognition results are obtained by the back-end server based on the plurality of detection boxes detecting the state of the target object respectively, and the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object; and the front-end client correlating the first state recognition result and a second state recognition result to obtain a correlation result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0010] According to another aspect of the embodiments of the present application, there is also provided an object recognition device for recognizing an object. The device can include: a determination unit configured to determine a plurality of detection boxes of a target object, wherein the detection boxes are configured to detect target parts corresponding to a state of the target object; a first recognition unit configured to detect the target parts based on the plurality of detection boxes respectively, and recognize the state of the target object based on the obtained detection results to obtain a plurality of state recognition results, wherein the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is configured to identify the target object; and a first association unit configured to associate the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0011] According to another aspect of the embodiments of the present application, there is also provided another object recognition device for recognizing an object. The device can include: a first display unit configured to cause a teaching management system to display a plurality of detection boxes of a target object on a display interface, wherein the detection boxes are configured to detect target parts corresponding to a state of the target object; a second recognition unit configured to cause the teaching management system to receive a recognition request, and display a plurality of state recognition results on the display interface, wherein the plurality of state recognition results are obtained by causing the teaching management system to recognize the state of the target object based on the plurality of detection boxes respectively, and the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is configured to identify the target object; and a second display unit configured to cause the teaching management system to display an association result on the display interface, wherein the association result is a result of associating the first state recognition result with a second state recognition result, and the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0012] According to another aspect of the embodiments of the present application, there is also provided another object recognition device for recognizing an object. The device can include: a third display unit configured to display a plurality of detection boxes of a target object on an operation interface, wherein the detection boxes are configured to detect target parts corresponding to a state of the target object; a fourth display unit configured to receive a recognition operation instruction on the operation interface, and display a plurality of state recognition results, wherein the plurality of state recognition results are obtained by recognizing the state of the target object based on the plurality of detection boxes respectively, and the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is configured to identify the target object; and a fifth display unit configured to display an association result on the operation interface, wherein the association result is a result of associating the first state recognition result with a second state recognition result, and the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0013] According to another aspect of the embodiments of the present application, there is also provided another object recognition device for recognizing an object. The device can include: an uploading unit configured to cause a front-end client to upload a plurality of bounding boxes of a target object, wherein the bounding boxes are used to detect target parts corresponding to a state of the target object; a transmitting unit configured to cause the front-end client to transmit the plurality of bounding boxes to a back-end server; a receiving unit configured to cause the front-end client to receive a plurality of state recognition results returned by the back-end server, wherein the plurality of state recognition results are obtained by the back-end server based on detection of the plurality of bounding boxes on the state of the target object, and the plurality of state recognition results include a first state recognition result of the target object, the first state recognition result being used to identify the target object; and a second associating unit configured to cause the front-end client to associate the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any one of the plurality of state recognition results other than the first state recognition result.
[0014] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium. The computer-readable storage medium includes a stored program, wherein the program, when executed by a processor, controls a device in which the computer-readable storage medium is located to perform the object recognition method according to the embodiments of the present application.
[0015] According to another aspect of the embodiments of the present application, there is also provided a processor. The processor is configured to execute a program, wherein the program, when executed by the processor, performs the object recognition method according to the embodiments of the present application.
[0016] According to another aspect of the embodiments of the present application, there is also provided an object recognition system, which includes: a processor; and a memory connected to the processor, configured to provide the processor with instructions for processing the following processing steps: determining a plurality of bounding boxes of a target object, wherein the bounding boxes are used to detect target parts corresponding to a state of the target object; detecting the corresponding target parts based on the plurality of bounding boxes respectively, and recognizing the state of the target object based on the obtained detection results to obtain a plurality of state recognition results, wherein the plurality of state recognition results include a first state recognition result of the target object, the first state recognition result being used to identify the target object; and associating the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any one of the plurality of state recognition results other than the first state recognition result.
[0017] In the embodiment of the present application, a plurality of detection boxes of the target object are determined, wherein the detection boxes are used for detecting target parts corresponding to the state of the target object; the corresponding target parts are detected based on the plurality of detection boxes respectively, and the state of the target object is recognized based on the obtained detection results to obtain a plurality of state recognition results, wherein the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used for identifying the target object; the first state recognition result is associated with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result except the first state recognition result in the plurality of state recognition results. That is, the target parts corresponding to the target object are detected by the plurality of detection boxes, the first state recognition result used for identifying the target object is included in the plurality of state recognition results, and the first state recognition result is associated with the second state recognition result in the plurality of state recognition results, so that the purpose of accurately associating the second state recognition result with the target object is achieved, the technical problem of low accuracy of object recognition is solved, and the technical effect of improving the accuracy of object recognition is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings, which are included to provide a further understanding of the present application and constitute a part of this application, illustrate embodiments of the present application and together with the description serve to explain the present application. In the drawings:
[0019] Figure 1 FIG. 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing an object recognition method according to an embodiment of the present application;
[0020] Figure 2 FIG. 2 is a flowchart of an object recognition method according to an embodiment of the present application;
[0021] Figure 3 FIG. 3 is a flowchart of another object recognition method according to an embodiment of the present application;
[0022] Figure 4 FIG. 4 is a flowchart of another object recognition method according to an embodiment of the present application;
[0023] Figure 5 FIG. 5 is a flowchart of another object recognition method according to an embodiment of the present application;
[0024] Figure 6 FIG. 6 is a schematic diagram of object recognition according to an embodiment of the present application;
[0025] Figure 7A FIG. 7 is a scene schematic diagram of object recognition according to an embodiment of the present application;
[0026] Figure 7B is another object recognition scenario according to an embodiment of the present application;
[0027] Figure 8 is a schematic diagram of an object recognition device according to an embodiment of the present application;
[0028] Figure 9 is another schematic diagram of an object recognition device according to an embodiment of the present application;
[0029] Figure 10 is another schematic diagram of an object recognition device according to an embodiment of the present application;
[0030] Figure 11 is another schematic diagram of an object recognition device according to an embodiment of the present application;
[0031] Figure 12 is a structural block diagram of a computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to make the personnel in the technical field better understand the present application scheme, the technical scheme in the embodiment of the present application will be described clearly and completely in combination with the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, not all. Based on the embodiment in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the scope of protection of the present application.
[0033] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0034] First, some nouns or terms appearing in the description of the embodiments of the present application are applicable to the following explanations:
[0035] Intelligent recognition, image recognition using artificial intelligence (Artificial Intelligence, AI) technology;
[0036] behavior recognition, recognizing the action of the object, such as, raising hands, standing, playing a mobile phone, etc.;
[0037] expression recognition, recognizing the expression of the object, such as, happy, angry, etc.
[0038] Embodiment 1
[0039] According to the embodiment of the present application, an embodiment of an object recognition method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from here.
[0040] The method embodiment provided by the embodiment of the present application can be executed in a mobile terminal, a computer terminal or similar computing device. Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing an object recognition method according to the embodiment of the present application. As shown in Figure 1 , the computer terminal 10 (or mobile device 10) can include one or more processors 102 (the processor 102 can include but not limited to a microprocessor MCU or a programmable logic device FPGA processing device), a memory 104 for storing data, and a transmission module 106 for communication function. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 the structure shown is only schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can include more or less components than Figure 1 shown, or have a different configuration than Figure 1 shown.
[0041] It should be noted that the one or more processors 102 and / or other data processing circuits described above can be referred to as "data processing circuits" herein. The data processing circuit can be embodied in whole or in part as software, hardware, firmware or any combination thereof. In addition, the data processing circuit can be a single independent processing module, or any one of the other elements combined into the computer terminal 10 (or mobile device) in whole or in part. As referred to in the embodiments of the present application, the data processing circuit as a kind of processor control (for example, the selection of variable resistance terminal path connected with the interface).
[0042] The memory 104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the object recognition method in the embodiments of the present application. The processor 102 executes various functional applications and data processing, i.e., implements the object recognition method described above, by running the software programs and modules stored in the memory 104. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include memories remotely arranged with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0043] The transmission device 106 is used to receive or send data via a network. The specific examples of the network include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network interface controller (NIC) which can be connected to other network devices through a base station so as to be able to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module which is used to communicate with the Internet in a wireless manner.
[0044] The display can be, for example, a touch screen type liquid crystal display (LCD) which can enable a user to interact with a user interface of the computer terminal 10 (or a mobile device).
[0045] It should be noted that, in some optional embodiments, the above-mentioned Figure 1 The computer device (or mobile device) shown can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or a combination of both hardware and software elements. It should be noted that, Figure 1 is merely one example of a particular implementation and is intended to illustrate the types of components that can be present in the above-described computer device (or mobile device).
[0046] In Figure 1 the operating environment shown, the present application provides an object recognition method as shown in Figure 2 It should be noted that the object recognition method of this embodiment can be executed by the mobile terminal of the embodiment shown in Figure 1
[0047] Figure 2 is a flowchart of an object recognition method according to an embodiment of the present application. As shown in Figure 2 As shown, the method can include the following steps:
[0048] In step S202, a plurality of detection boxes of the target object are determined, wherein the detection boxes are used to detect target parts corresponding to the state of the target object.
[0049] In the technical scheme provided by the above step S202 of the present application, the target object is an object to be recognized, and an image of the target object can be obtained, and a plurality of detection boxes of the target object are determined in the image, wherein the image can be a video frame collected by an image collection device, such as a video frame collected by a plurality of cameras arranged in a certain area. Each detection box corresponds to a part of the target object, and the detection box is used to detect a target part corresponding to the state of the target object. For example, the target part is the face of the target object, which corresponds to the expression state of the target object, so the face can be detected through the detection box of the face. For another example, the target part is the head of the target object, which corresponds to the identity state of the target object, so the head can be detected through the detection box of the head. For another example, the target part is the body of the target object, which corresponds to the behavior state of the target object, so the body can be detected through the detection box of the body.
[0050] The embodiment can simultaneously track a plurality of target parts corresponding to the target object through a plurality of detection boxes in real time.
[0051] It should be noted that the above target part of the embodiment can be any part that can be used to determine the state of the target object, and correspondingly, the above detection box can be any detection box corresponding to the target part of the target object, which is not limited here.
[0052] In step S204, the corresponding target parts are detected based on the plurality of detection boxes respectively, and the state of the target object is recognized based on the obtained detection results to obtain a plurality of state recognition results.
[0053] In the technical scheme provided by the above step S204 of the present application, after the plurality of detection boxes of the target object are determined, the corresponding target parts can be detected based on the plurality of detection boxes respectively, and the state of the target object is recognized based on the obtained detection results to obtain a plurality of state recognition results, wherein the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object.
[0054] In this embodiment, the corresponding target part is detected based on the plurality of detection boxes, which can be simultaneously detected based on the plurality of detection boxes, and then the detection result corresponding to the target part is obtained. The state corresponding to the target part of the target object is identified through the detection result, and the state identification result corresponding to the target part is obtained. A plurality of state identification results can be obtained.
[0055] In this embodiment, the plurality of state identification results include a first state identification result, which is an identity state identification result for identifying the target object. That is, the target part of the target object is identified through the detection box, and the identity of the target object can be identified.
[0056] In step S206, the first state identification result is associated with the second state identification result to obtain an association result, wherein the second state identification result is any state identification result in the plurality of state identification results except the first state identification result.
[0057] In the technical solution provided by step S206 of the present application, in addition to the first state identification result, the plurality of state identification results also include a second state identification result. The second state identification result can be a state identification result obtained by detecting other parts (such as the face, body, etc.) of the target object through the detection box.
[0058] In order to correspond the second state identification result to the target object, and the first state identification result of the target object is used to identify the target object, the first state identification result is associated with the second state identification result in this embodiment to obtain an association result. That is, the second state identification result can be associated with the target object, so as to accurately identify the target object.
[0059] Through the steps S202 to S206, the multiple bounding boxes of the target object are determined, wherein the bounding boxes are used to detect the target parts corresponding to the state of the target object; the corresponding target parts are detected based on the multiple bounding boxes respectively, and the state of the target object is recognized based on the obtained detection results to obtain multiple state recognition results, wherein the multiple state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object; and the first state recognition result is associated with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the multiple state recognition results except the first state recognition result. That is, the target parts corresponding to the target object are detected through the multiple bounding boxes, the first state recognition result used to identify the target object is included in the multiple state recognition results, and the first state recognition result is associated with the second state recognition result in the multiple state recognition results, so that the purpose of accurately associating the second state recognition result with the target object is achieved, the technical problem of low accuracy of object recognition is solved, and the technical effect of improving the accuracy of object recognition is achieved.
[0060] The above method of the embodiment will be further introduced below.
[0061] As an optional implementation, after the multiple bounding boxes of the target object are determined in the step S202, the method further includes: binding the multiple bounding boxes to obtain a binding result, wherein the binding result includes the multiple bounding boxes, and the multiple bounding boxes correspond to the multiple state recognition results one by one.
[0062] In the embodiment, the multiple bounding boxes can be bound together through a bounding box binding algorithm to obtain a binding result, the binding result can include the multiple bounding boxes, and the multiple bounding boxes can have a one-to-one correspondence with the multiple state recognition results, so that the corresponding target parts are detected based on the multiple bounding boxes respectively, and the state of the target object is recognized based on the obtained detection results to obtain the corresponding multiple state recognition results.
[0063] As an optional implementation, in the step S204, the corresponding target parts are detected based on the multiple bounding boxes respectively, including: in the binding result, images of regions where the multiple bounding boxes are located are extracted respectively, and the corresponding target parts are detected based on the images.
[0064] In the embodiment, when the corresponding target parts are detected based on the multiple bounding boxes respectively, the images of the regions where the multiple bounding boxes are located can be determined in the above binding result, and then extracted, and the corresponding target parts are detected through a corresponding recognition algorithm, wherein the recognition algorithm can be an AI algorithm.
[0065] As an optional implementation, step S204, respectively based on the plurality of detection frames, the corresponding target part is detected, and based on the obtained detection result, the state of the target object is recognized, and a plurality of state recognition results are obtained, including: based on the first detection frame, the corresponding first target part is detected, and based on the obtained first detection result, the identity state of the target object is recognized, and the first state recognition result is obtained; based on the second detection frame, the corresponding second target part is detected, and based on the obtained second detection result, the expression state of the target object is recognized, and the expression recognition result is obtained, wherein the second state recognition result includes the expression recognition result; and / or based on the third detection frame, the corresponding third target part is detected, and based on the obtained third detection result, the behavior state of the target object is recognized, and the behavior recognition result is obtained, wherein the second state recognition result includes the behavior recognition result.
[0066] In this embodiment, the identification of the target object includes the identification of the identity of the target object, and the plurality of detection frames includes a first detection frame corresponding to a first target part. The first target part can be detected based on the first detection frame to obtain a first detection result, and then the identity state of the target object is identified through the first detection result, that is, the identity of the target object is confirmed through the first detection result, so as to obtain the first state recognition result.
[0067] In this embodiment, the identification of the target object also includes the identification of the expression of the target object, and the plurality of detection frames includes a second detection frame corresponding to a second target part. The second target part can be detected based on the second detection frame to obtain a second detection result, and then the expression state of the target object is identified through the second detection result, so as to obtain the expression recognition result. The second state recognition result includes the expression recognition result.
[0068] In this embodiment, the identification of the target object also includes the identification of the behavior of the target object, and the plurality of detection frames includes a third detection frame corresponding to a third target part. The third target part can be detected based on the third detection frame to obtain a third detection result, and then the behavior state of the target object is identified through the third detection result, so as to obtain the behavior recognition result. The second state recognition result includes the expression recognition result.
[0069] As an optional implementation, based on the first detection frame, the corresponding first target part is detected, and based on the obtained first detection result, the identity state of the target object is recognized, and the first state recognition result is obtained, including: based on the first detection frame, the picture of the first target part is detected, and the posture information of the first target part is recognized from the obtained picture; based on the posture information, the identity state is recognized, and the first state recognition result is obtained.
[0070] In this embodiment, when the corresponding first target part is detected based on the first detection box, and the identity state of the target object is recognized based on the obtained first detection result to obtain the first state recognition result, the first target part can be detected by using the first detection box. The picture is a single image used for identity recognition of the target object, and the posture information of the first target part can be recognized therefrom. The posture information of the first target part can be estimated by using the first detection box, so as to obtain the posture information, and then the identity state of the target object is recognized based on the posture information to obtain the first state recognition result for identifying the identity of the target object.
[0071] As an optional implementation, the identity state is recognized based on the posture information to obtain the first state recognition result, including: positioning the key points corresponding to the posture information to obtain a positioning result; and recognizing the identity state based on the positioning result to obtain the first state recognition result.
[0072] In this embodiment, when the identity state is recognized based on the posture information to obtain the first state recognition result, the key points can be positioned. The key points corresponding to the posture information can be determined, and then positioned to obtain the positioning result of the key points, and then the identity state is recognized based on the positioning result to obtain the first state recognition result.
[0073] As an optional implementation, the first target part is the head of the target object, the identity state is recognized based on the positioning result to obtain the first state recognition result, including: recognizing the face of the head based on the positioning result to obtain the first state recognition result.
[0074] In this embodiment, the first target part can be specifically the head of the target object, and the first detection box can also be referred to as a head box. When the identity state is recognized based on the positioning result to obtain the first state recognition result, the face quality can be evaluated based on the positioning result to obtain an evaluation result, and then the identity of the target object is recognized based on the evaluation result to obtain the first state recognition result for identifying the identity of the target object.
[0075] As an optional implementation, the corresponding second target part is detected based on the second detection box, and the expression state of the target object is recognized based on the obtained second detection result to obtain an expression recognition result, including: detecting a first video sequence of the second target part based on the second detection box, and performing alignment processing on the second target part based on the obtained first video sequence to obtain an alignment result; and recognizing the expression state based on the alignment result to obtain the expression recognition result.
[0076] In this embodiment, when the corresponding second target part is detected based on the second detection frame, and the expression state of the target object is recognized based on the obtained second detection result to obtain an expression recognition result, the expression state of the target object can be recognized in a video sequence-based manner. Alternatively, the embodiment detects the second target part based on the first video sequence of the second detection frame, and the first video sequence can be a buffered face sequence. The second target part can be aligned based on the first video sequence to obtain an alignment result, and the expression state of the target object is recognized based on the alignment result to obtain the expression recognition result of the target object.
[0077] As an optional implementation, the second target part is the face of the target object, and the face of the second target part is aligned based on the obtained first video sequence to obtain an alignment result.
[0078] In this embodiment, the second target part can be the face of the target object, and the second detection frame can be a face frame, so that the embodiment can perform face alignment based on the obtained face sequence to obtain an alignment result.
[0079] As an optional implementation, the corresponding second target part is detected based on the second detection frame, and the expression state of the target object is recognized based on the obtained second detection result to obtain an expression recognition result, including: detecting the first video sequence of the second target part based on the second detection frame, and extracting the first spatiotemporal feature of the second target part from the first video sequence; and recognizing the expression state based on the first spatiotemporal feature to obtain the expression recognition result.
[0080] This embodiment can also recognize the expression state of the target object based on the spatiotemporal feature of the video sequence. The embodiment can detect the first video sequence of the second target part based on the second detection frame, the first video sequence includes the first spatiotemporal feature of the second target part, and the first spatiotemporal feature can be extracted from the first video sequence. Then, the expression state of the target object is recognized based on the first spatiotemporal feature to obtain the expression recognition result of the target object.
[0081] As an optional implementation, the second target part is the face of the target object, and the expression state is recognized based on the first spatiotemporal feature to obtain the expression recognition result, including: determining the expression composed of the plurality of first spatiotemporal features of the face as the expression recognition result.
[0082] In this embodiment, the second target part can be a face of the target object, and the second detection frame can be a face frame, so that when the expression state is identified based on the first spatiotemporal features, a plurality of first spatiotemporal features of the face of the target object can be determined, and then an expression composed of the plurality of first spatiotemporal features is determined as an expression recognition result for finally identifying the expression state of the target object.
[0083] As an optional implementation, the third target part is detected based on the third detection frame, and a behavior state of the target object is identified based on the obtained third detection result to obtain a behavior recognition result, including: a second video sequence of the third target part is detected based on the third detection frame, and a second spatiotemporal feature of the third target part is extracted from the second video sequence; and the behavior state of the target object is identified based on the second spatiotemporal feature to obtain the behavior recognition result.
[0084] In this embodiment, when the third target part is detected based on the third detection frame, and the behavior state of the target object is identified based on the obtained third detection result to obtain the behavior recognition result, the spatiotemporal feature can be extracted based on the video sequence to identify the behavior state of the target object. Alternatively, in this embodiment, a second video sequence of the third target part is detected based on the third detection frame, the second video sequence can be a buffered full-body extended sequence, a second spatiotemporal feature of the third target part can be extracted from the second video sequence, and then the behavior state of the target object is identified based on the second spatiotemporal feature to obtain the behavior recognition result of the target object.
[0085] As an optional implementation, the third target part is the body of the target object, and the behavior state of the target object is identified based on the second spatiotemporal feature to obtain the behavior recognition result, including: a behavior composed of a plurality of second spatiotemporal features of the body is determined as the behavior recognition result.
[0086] In this embodiment, the third target part can be the body of the target object, so that the third detection frame can be a full-body frame, so that when the behavior state of the target object is identified based on the second spatiotemporal feature to obtain the behavior recognition result, a plurality of second spatiotemporal features of the body of the target object can be determined, and then a behavior composed of the plurality of second spatiotemporal features is determined as a behavior recognition result for finally identifying the behavior state of the target object.
[0087] As an optional implementation, after the first state recognition result and the second state recognition result are associated to obtain an association result in step S206, the method further includes: receiving and responding to an information extraction instruction to output target information, wherein the target information is information of the target object including the recognition result.
[0088] In this embodiment, after the first state recognition result is associated with the second state recognition result to obtain an association result, the recognition result can be extracted according to a requirement, an information extraction instruction can be received, the information extraction instruction can be an instruction triggered by a user according to a requirement of the user, a target information meeting the requirement of the user is output in response to the information extraction instruction, the target information is information of a target object including the recognition result of the merchant, and the target information can be an event video (life behavior, etc.), an abstract video (highlight moment of the object, etc.), live broadcast / playback (remote watching, etc.), information summary (statistics of learning and life in a preset time period, etc.), which is not limited specifically here.
[0089] The embodiment of the application further provides another object recognition method applied to a teaching scene.
[0090] Figure 3 is a flowchart of another object recognition method according to an embodiment of the application. As shown in Figure 3 , the method can include the following steps:
[0091] In step S302, the teaching management system displays a plurality of detection boxes of a target object on a display interface, wherein the detection boxes are used to detect target parts corresponding to a state of the target object.
[0092] In the technical solution provided in step S302 of the application, the target object is an object to be recognized in a teaching scene, such as a student, a teacher, etc. in an adult education scene. The embodiment can obtain an image of the target object, and determine a plurality of detection boxes of the target object in the image, wherein the image can be a video frame collected by an image collection device, such as a plurality of cameras arranged in a teaching area. The camera can be a general camera, and the number of cameras to be arranged can be determined according to the specific size and arrangement of the teaching area to ensure full coverage of the teaching area.
[0093] Each detection box of the embodiment corresponds to a part of the target object, and the detection box is used to detect a target part corresponding to a state of the target object, and then the plurality of detection boxes are displayed on the display interface of the teaching management system.
[0094] In step S304, the teaching management system receives a recognition request and displays a plurality of state recognition results on the display interface.
[0095] In the technical solution provided in step S304 of the application, the plurality of state recognition results are obtained by the teaching management system based on the detection of the plurality of detection boxes on the state of the target object, and the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object.
[0096] In this embodiment, the identification request described above is a request for identifying the target object, and the teaching management system can detect the corresponding target part based on the plurality of detection boxes, which can be simultaneously detecting the corresponding target part based on the plurality of detection boxes, and then obtaining the detection result corresponding to the target part. The state corresponding to the target part of the target object is identified through the detection result, and the state identification result corresponding to the target part is obtained. The state identification result corresponds to the detection box one by one, so that the plurality of detection boxes correspond to a plurality of state identification results, and the plurality of state identification results are displayed on the display interface.
[0097] In this embodiment, the plurality of state identification results displayed on the display interface include a first state identification result, which is an identity state identification result for identifying the target object. That is, the target part of the target object is identified through the detection box, and the identity of the target object can be identified.
[0098] Step S306, the teaching management system displays the association result on the display interface, wherein the association result is the result of associating the first state identification result and the second state identification result, and the second state identification result is any state identification result in the plurality of state identification results except the first state identification result.
[0099] In the technical solution provided by the above step S306 of the present application, in addition to the first state identification result, the plurality of state identification results also include the second state identification result, which can be a state identification result obtained by detecting other parts (such as face, body, etc.) of the target object through the detection box.
[0100] In this embodiment, in order to correspond the second state identification result to the target object, and the first state identification result of the target object is used to identify the target object, the teaching management system of this embodiment can associate the first state identification result with the second state identification result, and display the association result on the display interface. That is, the second state identification result can be associated with the target object, so as to accurately identify the target object.
[0101] The embodiment of the present application also provides another object identification method from the perspective of human-computer interaction.
[0102] Figure 4 is a flowchart of another object identification method according to an embodiment of the present application. As shown in Figure 4 , the method can include the following steps:
[0103] Step S402, display a plurality of detection boxes of a target object on an operation interface, wherein the detection box is used to detect a target part corresponding to a state of the target object.
[0104] In the technical solution provided in the step S402 of the present application, the target object is an object to be identified, an image of the target object can be acquired, a plurality of detection boxes of the target object are determined in the image, and the plurality of detection boxes of the target object are displayed on the operation interface.
[0105] In step S404, the identification operation instruction is sensed on the operation interface, and a plurality of state identification results are displayed.
[0106] In the technical solution provided in the step S404 of the present application, the plurality of state identification results are obtained by identifying the state of the target object based on the plurality of detection boxes respectively, the plurality of state identification results include a first state identification result of the target object, and the first state identification result is used to identify the target object.
[0107] In this embodiment, the identification operation instruction can be an operation instruction triggered by a user on the operation interface, the identification operation instruction can be sensed on the operation interface, a corresponding target part can be detected based on the plurality of detection boxes, and a detection result corresponding to the target part is obtained, the state corresponding to the target part of the target object is identified through the detection result, and a state identification result corresponding to the target part is obtained, the state identification result corresponds to the detection box one by one, so that the plurality of detection boxes correspond to the plurality of state identification results, and the plurality of state identification results obtained above are displayed on the operation interface.
[0108] In this embodiment, the plurality of state identification results displayed on the operation interface include the first state identification result, the first state identification result is an identity state identification result used to identify the target object, that is, the target part of the target object is identified through the detection box, and the identity of the target object can be identified.
[0109] In step S406, the association result is displayed on the operation interface, the association result is a result of associating the first state identification result with a second state identification result, and the second state identification result is any state identification result in the plurality of state identification results except the first state identification result.
[0110] In the technical solution provided in the step S406 of the present application, in order to correspond the second state identification result to the target object, the first state identification result of the target object is used to identify the target object, the first state identification result is associated with the second state identification result in this embodiment, the association result is obtained, and the association result is displayed on the operation interface, that is, the second state identification result can be associated with the target object, so as to accurately identify the target object.
[0111] The embodiment of the present application further provides another object identification method.
[0112] Figure 5 is a flowchart of another object recognition method according to an embodiment of the present application. As shown in the figure, the method can include the following steps: Figure 5
[0113] In step S502, the front-end client uploads a plurality of bounding boxes of the target object, wherein the bounding boxes are used to detect target parts corresponding to the state of the target object.
[0114] In the technical solution provided by the above step S502 of the present application, the target object is the object to be recognized, and the front-end client can obtain an image of the target object, determine a plurality of bounding boxes of the target object in the image, and then start uploading the plurality of bounding boxes of the target object.
[0115] In step S504, the front-end client transmits the plurality of bounding boxes to the back-end server.
[0116] In the technical solution provided by the above step S504 of the present application, a communication connection is established between the front-end client and the back-end server, and the plurality of bounding boxes can be transmitted to the back-end server.
[0117] In step S506, the front-end client receives a plurality of state recognition results returned by the back-end server.
[0118] In the technical solution provided by the above step S506 of the present application, the plurality of state recognition results are obtained by the back-end server based on the plurality of bounding boxes detecting the state of the target object, and the plurality of state recognition results include a first state recognition result of the target object, which is used to identify the target object.
[0119] In this embodiment, after the front-end client transmits the plurality of bounding boxes to the back-end server, the back-end server can detect the corresponding target parts based on the plurality of bounding boxes, and then obtain the detection results corresponding to the target parts. The state corresponding to the target part of the target object is recognized through the detection results, and the state recognition result corresponding to the target part is obtained. The state recognition result corresponds to the bounding box one by one, so that the plurality of bounding boxes correspond to the plurality of state recognition results, and the plurality of state recognition results are issued to the front-end client.
[0120] In this embodiment, the plurality of state recognition results obtained by the front-end client include the first state recognition result, which is the identity state recognition result used to identify the target object. That is, the target part of the target object is recognized through the bounding box, and the identity of the target object can be recognized.
[0121] In step S508, the front-end client associates the first state recognition result with the second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0122] In the technical solution provided in step S508 of the present application, after the front-end client receives the plurality of state recognition results returned by the background server, the first state recognition result can be associated with the second state recognition result to obtain an association result. In order to associate the second state recognition result with the target object, the first state recognition result of the target object is used to identify the target object. Therefore, the front-end client associates the first state recognition result with the second state recognition result to obtain an association result, that is, the second state recognition result can be associated with the target object, thereby achieving the purpose of accurately identifying the target object.
[0123] It should be noted that the above object recognition method of the embodiment can be applied to a teaching management (such as adult education) scene, and can also be applied to a farm management scene, an industrial park management scene, and other scenes that require object recognition. Here, it will not be illustrated one by one.
[0124] In related technologies, human skeleton key points are usually used to identify the behavior of an object. However, in the case where objects often move and gather together, the key points may not be accurately positioned, which seriously affects the accuracy of identifying the object. In another related technology, a single image or image sequence can be used to identify the behavior of a target object based on a detected single object. However, this method cannot accurately associate the behavior recognition result of an object with the corresponding object when there is occlusion in the image, thereby causing the technical problem of low accuracy in identifying the object.
[0125] To solve the above problems, the embodiment can automatically identify the behavior state, expression state, and identity state of an object without human intervention, and accurately associate the behavior state and expression state with the identity state of the corresponding target object, that is, accurately associate the recognition result with the corresponding target object. This solves the technical problem of low accuracy in identifying the object and achieves the technical effect of improving the accuracy of identifying the object.
[0126] Embodiment 2
[0127] The technical solutions of the present application will be illustrated below in conjunction with preferred embodiments.
[0128] In a related technology, human skeleton key points can be used to identify the behavior of an object, but in the teaching management and industrial park management scenarios, the target objects often move and gather together, which can cause inaccurate key point positioning and seriously affect the recognition accuracy; in another related technology, after detecting a single object, a single image or image sequence method can be used to identify the behavior, and when there is occlusion, the behavior result cannot be accurately associated with the corresponding object.
[0129] To solve the above problems, a set of efficient and blind area coverage object recognition method can be constructed by using multiple ordinary cameras, big data, cloud computing and AI recognition algorithm, which can automatically identify the behavior, expression and non-contact identity confirmation of students and teachers without human intervention, and accurately associate the recognition result with the corresponding object.
[0130] Figure 6 is a schematic diagram of object recognition according to an embodiment of the application. As shown in Figure 6 The acquisition device can use a camera to obtain the video frame of the object, for example, through camera 1... camera N to obtain the video frame of the object, and the number of cameras to be arranged can be determined according to the specific size and arrangement of the classroom to ensure full coverage of the classroom.
[0131] Since the corresponding area images are processed separately during the identification of expressions and behaviors, it is difficult to match the object's identity in the case of image occlusion. The embodiment uses a bounding algorithm to bind the face, head and body detection boxes of the object together through an algorithm to obtain a binding result, and then extracts the image of the face detection box corresponding area from the binding result, and sends it to the expression recognition algorithm to identify the expression of the object. The image corresponding to the body detection box can also be extracted from the binding result and sent to the behavior recognition algorithm to identify the behavior of the object.
[0132] Optionally, after the face, head and body detection boxes of the object are bound together through an algorithm to obtain a binding result, the embodiment can perform face frame tracking, cache face sequences, and then perform face alignment based on the cached face sequences, and identify the expression of the object through the face alignment result; optionally, the embodiment can perform head frame tracking to estimate the posture of the head, and then perform key point positioning based on the results of face frame tracking and posture estimation to obtain a positioning result, and then perform face quality evaluation based on the positioning result, and identify the identity of the object according to the obtained evaluation result; optionally, the embodiment can perform full-body frame tracking, cache full-body expansion sequences, extract spatio-temporal features from the full-body expansion sequences, and then use the spatio-temporal features to identify the behavior of the object.
[0133] From the above, the embodiment can simultaneously adopt real-time tracking on the face, body and head of the object, record the expression recognition result and behavior recognition result associated with the object and the identity, so as to ensure the accuracy of the recognition result.
[0134] In addition, in order to improve the accuracy of object recognition, the embodiment can extract the spatio-temporal features of the behavior based on the full-body frame to recognize the behavior of the object, extract the spatio-temporal features of the expression based on the face frame to recognize the expression of the object, and perform identity recognition on the detected head of the object based on a single image.
[0135] The embodiment can mainly realize the behavior, expression and identity recognition of the object, and can extract the obtained recognition result according to the requirement, such as remote live broadcast / playback (remote viewing, etc.), information summary (statistics of learning and life situation in a preset time period, etc.), event video (life behavior, etc.), summary video (highlights of the object, etc.) and the like.
[0136] The computer language of the embodiment can adopt C / C++, and the hardware can adopt a common camera.
[0137] Figure 7A is a scene schematic diagram of object recognition according to an embodiment of the application. As shown in Figure 7A The embodiment can capture the video frame of the target object through the camera, determine the plurality of detection frames of the target object through the video frame, and then input the plurality of detection frames to the computing device, wherein the detection frame is used to detect the target part corresponding to the state of the target object. After the computing device obtains the plurality of detection frames, the corresponding target part can be detected based on the plurality of detection frames respectively, and the state of the target object is recognized based on the obtained detection result to obtain a plurality of state recognition results, wherein the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object. The computing device can associate the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result. After the computing device obtains the association result, the above association result can be displayed on the display interface.
[0138] Figure 7B is another scene schematic diagram of object recognition according to an embodiment of the application. As shown in Figure 7BAs shown, a plurality of detection boxes of the target object can be added to the operation interface, the plurality of detection boxes of the target object are displayed on the operation interface, wherein the detection box is used for detecting a target part corresponding to a state of the target object; a recognition operation instruction is sensed on the operation interface, and a plurality of state recognition results are displayed, wherein the plurality of state recognition results are obtained by recognizing the state of the target object based on the plurality of detection boxes respectively, the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used for identifying the target object; and an association result is displayed on the display interface, wherein the association result is a result of associating the first state recognition result with a second state recognition result, and the second state recognition result is any state recognition result except the first state recognition result in the plurality of state recognition results.
[0139] The embodiment can automatically identify the behavior state, expression state and identity state of the object without interference, and accurately associate the behavior state, expression state and the identity state of the corresponding object, that is, accurately associate the recognition result and the corresponding target, thereby solving the technical problem of low accuracy of object recognition and achieving the technical effect of improving the accuracy of object recognition.
[0140] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0141] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and a general hardware platform, and of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk), and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device) to execute the methods described in the various embodiments of the present application.
[0142] Embodiment 3
[0143] According to the embodiments of the present application, a method for implementing the above-mentioned Figure 2 The object recognition device for implementing the object recognition method shown.
[0144] Figure 8 is a schematic diagram of an object recognition device according to an embodiment of the present application. As shown in the figure, the object recognition device 80 can include a determination unit 81, a first recognition unit 82 and a first association unit 83. Figure 8
[0145] The determination unit 81 is configured to determine a plurality of detection boxes of a target object, wherein the detection boxes are used to detect target parts corresponding to a state of the target object.
[0146] The first recognition unit 82 is configured to detect the corresponding target parts based on the plurality of detection boxes respectively, and identify the state of the target object based on the obtained detection results to obtain a plurality of state recognition results, wherein the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object.
[0147] The first association unit 83 is configured to associate the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0148] It should be noted that the determination unit 81, the first recognition unit 82 and the first association unit 83 correspond to steps S202 to S206 of the embodiment 1 respectively, and the three units have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in the above embodiment 1. It should be noted that the above units can run in the computer terminal 10 provided in the embodiment 1 as a part of the device.
[0149] According to the embodiments of the present application, an object recognition device for implementing the above-mentioned object recognition method is also provided. Figure 3
[0150] Figure 9 is a schematic diagram of another object recognition device according to an embodiment of the present application. As shown in the figure, the object recognition device 90 can include a first display unit 91, a second recognition unit 92 and a second display unit 93. Figure 9
[0151] The first display unit 91 is configured to make a teaching management system display a plurality of detection boxes of a target object on a display interface, wherein the detection boxes are used to detect target parts corresponding to a state of the target object.
[0152] The second identification unit 92 is configured to cause the teaching management system to receive an identification request and display a plurality of state identification results on the display interface, wherein the plurality of state identification results are obtained by identifying the state of the target object based on the plurality of detection frames respectively, and the plurality of state identification results include a first state identification result of the target object, and the first state identification result is used to identify the target object.
[0153] The second display unit 93 is configured to cause the teaching management system to display an association result on the display interface, wherein the association result is obtained by associating the first state identification result and a second state identification result, and the second state identification result is any state identification result in the plurality of state identification results except the first state identification result.
[0154] It should be noted that the first display unit 91, the second identification unit 92, and the second display unit 93 correspond to steps S302 to S306 of Embodiment 1 respectively, and the three units have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can run in the computer terminal 10 provided in Embodiment 1 as a part of the device.
[0155] According to the embodiments of the present application, a computer terminal 10 for implementing the above-mentioned object identification method is also provided. Figure 4 An object identification device for implementing the above-mentioned object identification method is also provided.
[0156] Figure 10 is a schematic diagram of another object identification device according to the embodiments of the present application. As shown in Figure 10 The object identification device 100 can include a third display unit 101, a fourth display unit 102, and a fifth display unit 103.
[0157] The third display unit 101 is configured to display a plurality of detection frames of the target object on the operation interface, wherein the detection frames are used to detect the target part corresponding to the state of the target object.
[0158] The fourth display unit 102 is configured to sense an identification operation instruction on the operation interface and display a plurality of state identification results, wherein the plurality of state identification results are obtained by identifying the state of the target object based on the plurality of detection frames respectively, and the plurality of state identification results include a first state identification result of the target object, and the first state identification result is used to identify the target object.
[0159] The fifth display unit 103 is configured to display an association result on the operation interface, wherein the association result is obtained by associating the first state identification result and a second state identification result, and the second state identification result is any state identification result in the plurality of state identification results except the first state identification result.
[0160] It should be noted that the third display unit 101, the fourth display unit 102 and the fifth display unit 103 correspond to steps S402 to S406 of Embodiment 1 respectively, and the three units have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be run in the computer terminal 10 provided in Embodiment 1 as part of the device.
[0161] According to an embodiment of the present application, a computer terminal for implementing the object recognition method is also provided. Figure 5 An object recognition device for implementing the object recognition method shown in the figure is provided.
[0162] Figure 11 is a schematic diagram of another object recognition device according to an embodiment of the present application. As shown in the figure, the object recognition device 110 can include an uploading unit 111, a transmission unit 112, a receiving unit 113 and a second association unit 114. Figure 11 The uploading unit 111 is configured to enable the front-end client to upload a plurality of detection boxes of a target object, wherein the detection boxes are used to detect target parts corresponding to the state of the target object.
[0163] The transmission unit 112 is configured to enable the front-end client to transmit the plurality of detection boxes to the back-end server.
[0164] The receiving unit 113 is configured to enable the front-end client to receive a plurality of state recognition results returned by the back-end server, wherein the plurality of state recognition results are obtained by the back-end server based on the plurality of detection boxes detecting the state of the target object respectively, and the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object.
[0165] The second association unit 114 is configured to enable the front-end client to associate the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0166] It should be noted that the uploading unit 111, the transmission unit 112, the receiving unit 113 and the second association unit 114 correspond to steps S502 to S506 of Embodiment 1 respectively, and the four units have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be run in the computer terminal 10 provided in Embodiment 1 as part of the device.
[0167]
[0168] In the object recognition device of the embodiment, the target part corresponding to the target object is detected by multiple detection boxes, and the multiple state recognition results include the first state recognition result for identifying the target object, and the first state recognition result is associated with the second state recognition result in the multiple state recognition results, so that the second state recognition result is accurately associated with the target object, the technical problem of low accuracy of object recognition is solved, and the technical effect of improving the accuracy of object recognition is achieved.
[0169] Embodiment 4
[0170] The embodiment of the present application can provide an object recognition system, which can include a computer terminal, which can be any computer terminal device in a computer terminal group. Alternatively, in the embodiment, the computer terminal can be replaced by a mobile terminal or other terminal device.
[0171] Alternatively, in the embodiment, the computer terminal can be located in at least one network device of multiple network devices of a computer network.
[0172] In the embodiment, the computer terminal can execute the program code of the following steps of the object recognition method: determining multiple detection boxes of a target object, wherein the detection box is used to detect a target part corresponding to the state of the target object; detecting the corresponding target part based on the multiple detection boxes respectively, and identifying the state of the target object based on the obtained detection result, to obtain multiple state recognition results, wherein the multiple state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object; and associating the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the multiple state recognition results except the first state recognition result.
[0173] Alternatively, Figure 12 is a structural block diagram of a computer terminal according to an embodiment of the present application. As shown in Figure 12 , the computer terminal A can include one or more (only one is shown in the figure) processors 122, a memory 124, and a transmission device 126.
[0174] The transmission device is configured to transmit a plurality of bounding boxes of the target object. The memory is configured to store software programs and modules, such as program instructions / modules corresponding to the object recognition method and device. The processor executes various functions and data processing by running the software programs and modules stored in the memory, thereby implementing the object recognition method described above. The memory can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include remotely disposed memories relative to the processor, which can be connected to the computer terminal A through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0175] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: determining a plurality of bounding boxes of the target object, wherein the bounding boxes are used to detect target parts corresponding to the state of the target object; detecting the corresponding target parts based on the plurality of bounding boxes respectively, and identifying the state of the target object based on the obtained detection results to obtain a plurality of state recognition results, wherein the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object; and associating the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0176] Optionally, the processor can further execute program codes of the following steps: after determining the plurality of bounding boxes of the target object, binding the plurality of bounding boxes to obtain a binding result, wherein the binding result includes the plurality of bounding boxes, and the plurality of bounding boxes correspond to the plurality of state recognition results one by one.
[0177] Optionally, the processor can further execute program codes of the following steps: in the binding result, extracting images of regions where the plurality of bounding boxes are located respectively, and detecting the corresponding target parts based on the images.
[0178] Optionally, the processor can further execute program codes of the following steps: detecting the first target part based on the first detection frame, and identifying the identity state of the target object based on the obtained first detection result to obtain a first state identification result; detecting the second target part based on the second detection frame, and identifying the expression state of the target object based on the obtained second detection result to obtain an expression identification result, wherein the second state identification result comprises the expression identification result; and / or detecting the third target part based on the third detection frame, and identifying the behavior state of the target object based on the obtained third detection result to obtain a behavior identification result, wherein the second state identification result comprises the behavior identification result.
[0179] Optionally, the processor can further execute program codes of the following steps: detecting a picture of the first target part based on the first detection frame, and identifying the posture information of the first target part from the obtained picture; identifying the identity state based on the posture information to obtain the first state identification result.
[0180] Optionally, the processor can further execute program codes of the following steps: positioning the key points corresponding to the posture information to obtain a positioning result; identifying the identity state based on the positioning result to obtain the first state identification result.
[0181] Optionally, the processor can further execute program codes of the following steps: identifying the face of the head based on the positioning result to obtain the first state identification result.
[0182] Optionally, the processor can further execute program codes of the following steps: detecting a first video sequence of the second target part based on the second detection frame, and performing alignment processing on the second target part based on the obtained first video sequence to obtain an alignment result; identifying the expression state based on the alignment result to obtain the expression identification result.
[0183] Optionally, the processor can further execute program codes of the following steps: performing face alignment on the face based on the obtained first video sequence to obtain the alignment result.
[0184] Optionally, the processor can further execute program codes of the following steps: detecting a first video sequence of the second target part based on the second detection frame, and extracting a first spatio-temporal feature of the second target part from the first video sequence; identifying the expression state based on the first spatio-temporal feature to obtain the expression identification result.
[0185] Optionally, the processor can further execute program codes of the following steps: determining the expression composed of the plurality of first spatio-temporal features of the face as the expression identification result.
[0186] Optionally, the processor can further execute program codes of the following steps: detecting a second video sequence of the third target part based on the third detection frame, and extracting a second spatio-temporal feature of the third target part from the second video sequence; and identifying a behavior state of the target object based on the second spatio-temporal feature to obtain a behavior recognition result.
[0187] Optionally, the processor can further execute program codes of the following steps: determining a behavior composed of the plurality of second spatio-temporal features of the body as the behavior recognition result.
[0188] Optionally, the processor can further execute program codes of the following steps: after associating the first state recognition result with the second state recognition result to obtain an association result, receiving and responding to an information extraction instruction to output target information, wherein the target information is information of the target object including the recognition result.
[0189] As an optional example, the processor can call information and application programs stored in the memory through the transmission device to execute the following steps: the teaching management system displays a plurality of detection frames of the target object on a display interface, wherein the detection frames are used to detect target parts corresponding to states of the target object; the teaching management system receives an identification request and displays a plurality of state recognition results on the display interface, wherein the plurality of state recognition results are obtained by the teaching management system based on a plurality of detection frames detecting states of the target object respectively, and the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object; and the teaching management system displays an association result on the display interface, wherein the association result is a result of associating the first state recognition result with a second state recognition result, and the second state recognition result is any state recognition result other than the first state recognition result in the plurality of state recognition results.
[0190] As an optional example, the processor can call information and application programs stored in the memory through the transmission device to execute the following steps: displaying a plurality of detection frames of the target object on an operation interface, wherein the detection frames are used to detect target parts corresponding to states of the target object; sensing an identification operation instruction on the operation interface and displaying a plurality of state recognition results, wherein the plurality of state recognition results are obtained by detecting states of the target object based on a plurality of detection frames respectively, and the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object; and displaying an association result on the operation interface, wherein the association result is a result of associating the first state recognition result with a second state recognition result, and the second state recognition result is any state recognition result other than the first state recognition result in the plurality of state recognition results.
[0191] As an optional example, the processor can call the information and application stored in the memory through the transmission device to perform the following steps: the front-end client uploads a plurality of detection boxes of the target object, wherein the detection box is used for detecting a target part corresponding to a state of the target object; the front-end client transmits the plurality of detection boxes to the back-end server; the front-end client receives a plurality of state recognition results returned by the back-end server, wherein the plurality of state recognition results are obtained by the back-end server based on the plurality of detection boxes detecting the state of the target object respectively, and the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used for identifying the target object; and the front-end client associates the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0192] By adopting the embodiment of the present application, an object recognition scheme is provided. A plurality of detection boxes of a target object are determined, wherein the detection box is used for detecting a target part corresponding to a state of the target object; the corresponding target part is detected based on the plurality of detection boxes respectively, and the state of the target object is recognized based on the obtained detection result to obtain a plurality of state recognition results, wherein the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used for identifying the target object; and the first state recognition result is associated with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result. That is, the embodiment detects the target part corresponding to the target object through the plurality of detection boxes, and the first state recognition result used for identifying the target object is included in the plurality of state recognition results, and the first state recognition result is associated with the second state recognition result in the plurality of state recognition results, so that the purpose of accurately associating the second state recognition result with the target object is achieved, the technical problem of low accuracy of object recognition is solved, and the technical effect of improving the accuracy of object recognition is achieved.
[0193] Those skilled in the art can understand that Figure 12 The structure shown is only schematic, and the computer terminal A can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a mobile Internet device (MID), a PAD, and the like. Figure 12 It does not limit the structure of the computer terminal A. For example, the computer terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 12 The computer terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 12 The computer terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure.
[0194] Those skilled in the art can understand that all or part of the steps in the above-mentioned various methods of the embodiments can be completed by instructing the terminal device related hardware through a program, and the program can be stored in a computer readable storage medium, which can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0195] Embodiment 5
[0196] The embodiments of the present application also provide a computer readable storage medium. Optionally, in the present embodiment, the above-mentioned computer readable storage medium can be used to save the program code executed by the object recognition method provided in the above-mentioned embodiment 1.
[0197] Optionally, in the present embodiment, the above-mentioned computer readable storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0198] Optionally, in the present embodiment, the computer readable storage medium is configured to store program code for performing the following steps: determining a plurality of detection boxes of a target object, wherein the detection boxes are used to detect target parts corresponding to the state of the target object; detecting the corresponding target parts based on the plurality of detection boxes respectively, and identifying the state of the target object based on the obtained detection results to obtain a plurality of state recognition results, wherein the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object; and associating the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0199] Optionally, the computer readable storage medium is further configured to store program code for performing the following steps: after determining the plurality of detection boxes of the target object, binding the plurality of detection boxes to obtain a binding result, wherein the binding result includes the plurality of detection boxes, and the plurality of detection boxes correspond to the plurality of state recognition results one by one.
[0200] Optionally, the computer readable storage medium is further configured to store program code for performing the following steps: in the binding result, extracting images of the regions where the plurality of detection boxes are located respectively, and detecting the corresponding target parts based on the images.
[0201] Optionally, the computer readable storage medium is further configured to store program code for performing the following steps: detecting the first target part based on the first detection frame, and identifying the identity state of the target object based on the obtained first detection result to obtain a first state identification result; detecting the second target part based on the second detection frame, and identifying the expression state of the target object based on the obtained second detection result to obtain an expression identification result, wherein the second state identification result comprises the expression identification result; and / or detecting the third target part based on the third detection frame, and identifying the behavior state of the target object based on the obtained third detection result to obtain a behavior identification result, wherein the second state identification result comprises the behavior identification result.
[0202] Optionally, the computer readable storage medium is further configured to store program code for performing the following steps: detecting a picture of the first target part based on the first detection frame, and identifying the posture information of the first target part from the obtained picture; identifying the identity state based on the posture information to obtain the first state identification result.
[0203] Optionally, the computer readable storage medium is further configured to store program code for performing the following steps: positioning the key points corresponding to the posture information to obtain a positioning result; identifying the identity state based on the positioning result to obtain the first state identification result.
[0204] Optionally, the computer readable storage medium is further configured to store program code for performing the following steps: identifying the face of the head based on the positioning result to obtain the first state identification result.
[0205] Optionally, the computer readable storage medium is further configured to store program code for performing the following steps: detecting a first video sequence of the second target part based on the second detection frame, and performing alignment processing on the second target part based on the obtained first video sequence to obtain an alignment result; identifying the expression state based on the alignment result to obtain the expression identification result.
[0206] Optionally, the computer readable storage medium is further configured to store program code for performing the following steps: performing face alignment on the face based on the obtained first video sequence to obtain the alignment result.
[0207] Optionally, the computer readable storage medium is further configured to store program code for performing the following steps: detecting a first video sequence of the second target part based on the second detection frame, and extracting a first spatio-temporal feature of the second target part from the first video sequence; identifying the expression state based on the first spatio-temporal feature to obtain the expression identification result.
[0208] Optionally, the computer readable storage medium is further configured to store program code for performing the following step: determining the expression composed of the plurality of first spatio-temporal features of the face as the expression recognition result.
[0209] Optionally, the computer readable storage medium is further configured to store program code for performing the following step: detecting a second video sequence of the third target part based on the third detection frame, and extracting second spatio-temporal features of the third target part from the second video sequence; identifying the behavior state of the target object based on the second spatio-temporal features to obtain a behavior recognition result.
[0210] Optionally, the computer readable storage medium is further configured to store program code for performing the following step: determining the behavior composed of the plurality of second spatio-temporal features of the body as the behavior recognition result.
[0211] Optionally, the computer readable storage medium is further configured to store program code for performing the following step: after associating the first state recognition result with the second state recognition result to obtain an association result, receiving and responding to an information extraction instruction to output target information, wherein the target information is information of the target object including the recognition result.
[0212] As an optional example, the computer readable storage medium is configured to store program code for performing the following steps: the teaching management system displays a plurality of detection frames of the target object on a display interface, wherein the detection frames are used to detect target parts corresponding to the state of the target object; the teaching management system receives an identification request and displays a plurality of state recognition results on the display interface, wherein the plurality of state recognition results are obtained by the teaching management system identifying the state of the target object based on the plurality of detection frames respectively, and the plurality of state recognition results include a first state recognition result of the target object, the first state recognition result being used to identify the target object; the teaching management system displays an association result on the display interface, wherein the association result is a result of associating the first state recognition result with a second state recognition result, and the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result.
[0213] The computer readable storage medium is configured to store program code for performing the following steps: displaying a plurality of detection boxes of a target object on an operation interface, wherein the detection boxes are used for detecting target parts corresponding to a state of the target object; sensing an identification operation instruction on the operation interface and displaying a plurality of state identification results, wherein the plurality of state identification results are obtained by identifying the state of the target object based on the plurality of detection boxes respectively, the plurality of state identification results include a first state identification result of the target object, and the first state identification result is used for identifying the target object; and displaying an association result on the operation interface, wherein the association result is a result of associating the first state identification result with a second state identification result, and the second state identification result is any state identification result except the first state identification result in the plurality of state identification results.
[0214] The computer readable storage medium is configured to store program code for performing the following steps: uploading, by a front-end client, a plurality of detection boxes of a target object, wherein the detection boxes are used for detecting target parts corresponding to a state of the target object; transmitting, by the front-end client, the plurality of detection boxes to a back-end server; receiving, by the front-end client, a plurality of state identification results returned by the back-end server, wherein the plurality of state identification results are obtained by identifying the state of the target object based on the plurality of detection boxes by the back-end server respectively, the plurality of state identification results include a first state identification result of the target object, and the first state identification result is used for identifying the target object; and associating, by the front-end client, the first state identification result with a second state identification result to obtain an association result, wherein the second state identification result is any state identification result except the first state identification result in the plurality of state identification results.
[0215] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0216] In the above-mentioned embodiments of the present application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0217] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented by other ways. Among them, the above-mentioned device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division way, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between units or modules, which can be electrical or other forms.
[0218] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0219] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0220] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the part of the prior art that essentially contributes or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various program code storage media.
[0221] The above is only the preferred embodiment of the present application. It should be noted that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. A method of recognizing an object, characterized by, The method comprises: determining a plurality of bounding boxes of a target object, wherein the bounding boxes are used to detect target parts corresponding to a state of the target object; detecting the corresponding target parts based on the plurality of bounding boxes respectively, and identifying the state of the target object based on the obtained detection results to obtain a plurality of state identification results, wherein the plurality of state identification results comprise a first state identification result of the target object, and the first state identification result is used to identify the target object; associating the first state identification result with a second state identification result to obtain an association result, wherein the second state identification result is any state identification result in the plurality of state identification results except the first state identification result; wherein detecting the corresponding target parts based on the plurality of bounding boxes respectively, and identifying the state of the target object based on the obtained detection results to obtain a plurality of state identification results comprises: performing pose estimation on a head of the target object based on a bounding box corresponding to the head to obtain pose information, positioning key points corresponding to the pose information to obtain a positioning result, performing face quality evaluation based on the positioning result to obtain an evaluation result, and identifying an identity state of the target object based on the evaluation result to obtain the first state identification result; detecting a video series of a face of the target object based on a bounding box corresponding to the face, wherein the video series of the face is a buffered face sequence, extracting spatio-temporal features of the face from the video series of the face to identify an expression state of the target object to obtain the second state identification result, and / or detecting a video sequence of a body of the target object based on a bounding box corresponding to the body, wherein the video sequence of the body is a buffered full-body expansion sequence, extracting spatio-temporal features of the body from the video sequence of the body to identify a behavior state of the target object to obtain the second state identification result.
2. The method of claim 1, wherein, After determining the plurality of bounding boxes of the target object, the method further comprises: binding the plurality of bounding boxes to obtain a binding result, wherein the binding result comprises the plurality of bounding boxes, and the plurality of bounding boxes correspond to the plurality of state identification results one by one.
3. The method of claim 2, wherein, detecting the corresponding target parts based on the plurality of bounding boxes respectively comprises: extracting images of regions where the plurality of bounding boxes are located in the binding result respectively, and detecting the corresponding target parts based on the images.
4. The method of claim 1, wherein, detecting the corresponding target parts based on the plurality of bounding boxes respectively, and identifying the state of the target object based on the obtained detection results to obtain a plurality of state identification results comprises: detecting a first target part based on a first bounding box, and identifying an identity state of the target object based on a first detection result obtained by the detection to obtain the first state identification result; detect a corresponding second target part based on the second detection frame, and identify an expression state of the target object based on a second detection result obtained, to obtain an expression recognition result, wherein the second state recognition result comprises the expression recognition result; and / or detect a corresponding third target part based on a third detection frame, and identify a behavior state of the target object based on a third detection result obtained, to obtain a behavior recognition result, wherein the second state recognition result comprises the behavior recognition result.
5. The method of claim 4, wherein, The first target part is a head of the target object, and the head is posture-estimated based on a detection frame corresponding to the head, to obtain posture information, comprising: detecting a picture of the first target part based on the first detection frame, and performing posture estimation on the first target part based on the first detection frame in the obtained picture, to obtain posture information of the first target part.
6. The method of claim 1, wherein, identify the identity state based on the evaluation result, to obtain the first state recognition result, comprising: identifying the face of the head based on the evaluation result, to obtain the first state recognition result.
7. The method of claim 4, wherein, detect a corresponding second target part based on a second detection frame, and identify an expression state of the target object based on a second detection result obtained, to obtain an expression recognition result, comprising: detecting a first video sequence of the second target part based on the second detection frame, and performing alignment processing on the second target part based on the first video sequence obtained, to obtain an alignment result; identify the expression state based on the alignment result, to obtain the expression recognition result.
8. The method of claim 7, wherein, The second target part is a face of the target object, and the second target part is aligned based on the first video sequence obtained, to obtain an alignment result, comprising: performing face alignment on the face based on the first video sequence obtained, to obtain the alignment result.
9. The method of claim 4, wherein, detect a corresponding second target part based on a second detection frame, and identify an expression state of the target object based on a second detection result obtained, to obtain an expression recognition result, comprising: detecting a first video sequence of the second target part based on the second detection frame, and extracting a first spatio-temporal feature of the second target part from the first video sequence; identify the expression state based on the first spatio-temporal feature, to obtain the expression recognition result.
10. The method of claim 9, wherein, The second target part is a face of the target object, and the expression state is identified based on the first spatio-temporal feature, to obtain the expression recognition result, comprising: determining an expression composed of a plurality of the first spatio-temporal features of the face as the expression recognition result.
11. The method of claim 4, wherein, detect a corresponding third target part based on a third detection frame, and identify a behavior state of the target object based on a third detection result obtained, to obtain a behavior recognition result, comprising: detecting a second video sequence of the third target part based on the third detection frame, and extracting a second spatio-temporal feature of the third target part from the second video sequence; Identify a behavior state of the target object based on the second spatiotemporal features, to obtain a behavior recognition result.
12. The method of claim 11, wherein, The third target part is a body of the target object, and the behavior state of the target object is identified based on the second spatiotemporal features to obtain the behavior recognition result, including: The behavior composed of the plurality of second spatiotemporal features of the body is determined as the behavior recognition result.
13. The method according to any one of claims 1 to 12, characterized in that, After the first state recognition result and the second state recognition result are associated to obtain an association result, the method further includes: Receiving and responding to an information extraction instruction, and outputting target information, wherein the target information is information of the target object including the recognition result.
14. A method of recognizing an object, characterized by, Including: The teaching management system displays a plurality of detection boxes of a target object on a display interface, wherein the detection boxes are used to detect target parts corresponding to states of the target object; The teaching management system receives a recognition request and displays a plurality of state recognition results on the display interface, wherein the plurality of state recognition results are obtained by the teaching management system based on detection of the plurality of detection boxes on the states of the target object, and the plurality of state recognition results include a first state recognition result of the target object, which is used to identify the target object; The teaching management system displays an association result on the display interface, wherein the association result is a result of associating the first state recognition result with a second state recognition result, and the second state recognition result is any state recognition result other than the first state recognition result in the plurality of state recognition results; The method further includes: performing pose estimation on the head of the target object based on a detection box corresponding to the head to obtain pose information; positioning key points corresponding to the pose information to obtain a positioning result; performing face quality evaluation based on the positioning result to obtain an evaluation result; and identifying an identity state of the target object based on the evaluation result to obtain the first state recognition result. Based on a detection box corresponding to the face of the target object, a video series of the face is detected, wherein the video series of the face is a cached face sequence; spatiotemporal features of the face are extracted from the video series of the face to identify an expression state of the target object to obtain the second state recognition result; and / or, based on a detection box corresponding to the body of the target object, a video sequence of the body is detected, wherein the video sequence of the body is a cached full-body expansion sequence; spatiotemporal features of the body are extracted from the video sequence of the body to identify a behavior state of the target object to obtain the second state recognition result.
15. A method of recognizing an object, characterized by, Including: Display a plurality of detection boxes of a target object on an operation interface, wherein the detection boxes are used to detect target parts corresponding to states of the target object; The operation interface is used to sense an operation instruction and display a plurality of state recognition results, wherein the plurality of state recognition results are obtained by recognizing the state of the target object based on the plurality of detection frames respectively, and the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object; An association result is displayed on the operation interface, wherein the association result is obtained by associating the first state recognition result with a second state recognition result, and the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result; The method further comprises: performing pose estimation on the head of the target object based on the detection frame corresponding to the head to obtain pose information; positioning the key points corresponding to the pose information to obtain a positioning result; performing face quality evaluation based on the positioning result to obtain an evaluation result; and recognizing the identity state of the target object based on the evaluation result to obtain the first state recognition result; Based on the detection frame corresponding to the face of the target object, a video series of the face is detected, wherein the video series of the face is a buffered face sequence; the spatio-temporal features of the face are extracted from the video series of the face to recognize the expression state of the target object to obtain the second state recognition result; and / or, based on the detection frame corresponding to the body of the target object, a video sequence of the body is detected, wherein the video sequence of the body is a buffered full-body expansion sequence; the spatio-temporal features of the body are extracted from the video sequence of the body to recognize the behavior state of the target object to obtain the second state recognition result.
16. A method of recognizing an object, characterized by, The method comprises: A front-end client uploads a plurality of detection frames of a target object, wherein the detection frames are used to detect target parts corresponding to the state of the target object; The front-end client transmits the plurality of detection frames to a back-end server; The front-end client receives a plurality of state recognition results returned by the back-end server, wherein the plurality of state recognition results are obtained by recognizing the state of the target object based on the plurality of detection frames respectively by the back-end server, and the plurality of state recognition results include a first state recognition result of the target object, and the first state recognition result is used to identify the target object; The front-end client associates the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result; The method further comprises: performing pose estimation on the head of the target object based on the detection frame corresponding to the head to obtain pose information; positioning the key points corresponding to the pose information to obtain a positioning result; performing face quality evaluation based on the positioning result to obtain an evaluation result; and recognizing the identity state of the target object based on the evaluation result to obtain the first state recognition result; detect a video series of the face based on the detection frame corresponding to the face of the target object, wherein the video series of the face is a buffered face sequence; extract spatio-temporal features of the face from the video series of the face, identify an expression state of the target object, and obtain the second state identification result; and / or detect a video sequence of the body based on the detection frame corresponding to the body of the target object, wherein the video sequence of the body is a buffered full-body expansion sequence; extract spatio-temporal features of the body from the video sequence of the body, identify a behavior state of the target object, and obtain the second state identification result.
17. An identification device of an object, characterized by Comprise: A determination unit is configured to determine a plurality of detection frames of a target object, wherein the detection frames are used to detect target parts corresponding to states of the target object; A first identification unit is configured to detect the corresponding target parts based on the plurality of detection frames respectively, and identify the states of the target object based on the obtained detection results, and obtain a plurality of state identification results, wherein the plurality of state identification results comprise a first state identification result of the target object, and the first state identification result is used to identify the target object; A first association unit is configured to associate the first state identification result with a second state identification result to obtain an association result, wherein the second state identification result is any state identification result except the first state identification result in the plurality of state identification results; The first identification unit is further configured to: perform pose estimation on a head of the target object by using a detection frame corresponding to the head to obtain pose information; position key points corresponding to the pose information to obtain a positioning result; perform face quality evaluation based on the positioning result to obtain an evaluation result; identify an identity state of the target object based on the evaluation result to obtain the first state identification result; Detect a video series of the face based on the detection frame corresponding to the face of the target object, wherein the video series of the face is a buffered face sequence; extract spatio-temporal features of the face from the video series of the face, identify an expression state of the target object, and obtain the second state identification result; and / or detect a video sequence of the body based on the detection frame corresponding to the body of the target object, wherein the video sequence of the body is a buffered full-body expansion sequence; extract spatio-temporal features of the body from the video sequence of the body, identify a behavior state of the target object, and obtain the second state identification result.
18. An apparatus for recognizing an object, characterized by comprising: Comprise: A first display unit is configured to enable a teaching management system to display a plurality of detection frames of a target object on a display interface, wherein the detection frames are used to detect target parts corresponding to states of the target object; The second identification unit is configured to cause the teaching management system to receive an identification request and display a plurality of state identification results on the display interface, wherein the plurality of state identification results are obtained by identifying the state of the target object based on the plurality of detection boxes respectively, the plurality of state identification results comprise a first state identification result of the target object, and the first state identification result is used to identify the target object. The second display unit is configured to cause the teaching management system to display an association result on the display interface, wherein the association result is an association result between the first state identification result and a second state identification result, and the second state identification result is any state identification result other than the first state identification result in the plurality of state identification results. The device is further configured to: perform pose estimation on the head of the target object by using the detection box corresponding to the head to obtain pose information; position the key points corresponding to the pose information to obtain a positioning result; perform face quality evaluation based on the positioning result to obtain an evaluation result; and identify the identity state of the target object based on the evaluation result to obtain the first state identification result. Based on the detection box corresponding to the face of the target object, a video series of the face is detected, wherein the video series of the face is a buffered face sequence; spatiotemporal features of the face are extracted from the video series of the face to identify the expression state of the target object to obtain the second state identification result; and / or, based on the detection box corresponding to the body of the target object, a video sequence of the body is detected, wherein the video sequence of the body is a buffered full-body expansion sequence; spatiotemporal features of the body are extracted from the video sequence of the body to identify the behavior state of the target object to obtain the second state identification result.
19. An apparatus for recognizing an object, characterized by comprising: The third display unit is configured to display a plurality of detection boxes of a target object on an operation interface, wherein the detection boxes are used to detect target parts corresponding to the state of the target object. The fourth display unit is configured to receive an identification operation instruction on the operation interface and display a plurality of state identification results, wherein the plurality of state identification results are obtained by identifying the state of the target object based on the plurality of detection boxes respectively, the plurality of state identification results comprise a first state identification result of the target object, and the first state identification result is used to identify the target object. The fifth display unit is configured to display an association result on the operation interface, wherein the association result is an association result between the first state identification result and a second state identification result, and the second state identification result is any state identification result other than the first state identification result in the plurality of state identification results. The device is further configured to: perform pose estimation on the head of the target object by using the detection frame corresponding to the head to obtain pose information; position the key points corresponding to the pose information to obtain a positioning result; perform face quality evaluation based on the positioning result to obtain an evaluation result; and identify the identity state of the target object based on the evaluation result to obtain the first state recognition result. Detect a video series of the face of the target object based on the detection frame corresponding to the face, wherein the video series of the face is a buffered face sequence; extract the spatiotemporal features of the face from the video series of the face to identify the expression state of the target object to obtain the second state recognition result; and / or detect a video sequence of the body of the target object based on the detection frame corresponding to the body, wherein the video sequence of the body is a buffered full-body expansion sequence; extract the spatiotemporal features of the body from the video sequence of the body to identify the behavior state of the target object to obtain the second state recognition result.
20. An apparatus for recognizing an object, characterized by comprising: The device comprises: an uploading unit configured to enable a front-end client to upload a plurality of detection frames of a target object, wherein the detection frames are used to detect target parts corresponding to states of the target object; a transmission unit configured to enable the front-end client to transmit the plurality of detection frames to a back-end server; a receiving unit configured to enable the front-end client to receive a plurality of state recognition results returned by the back-end server, wherein the plurality of state recognition results are obtained by the back-end server based on detection of the plurality of detection frames on the states of the target object, and the plurality of state recognition results comprise a first state recognition result of the target object, which is used to identify the target object; a second association unit configured to enable the front-end client to associate the first state recognition result with a second state recognition result to obtain an association result, wherein the second state recognition result is any state recognition result in the plurality of state recognition results except the first state recognition result; The device is further configured to: perform pose estimation on the head of the target object by using the detection frame corresponding to the head to obtain pose information; position the key points corresponding to the pose information to obtain a positioning result; perform face quality evaluation based on the positioning result to obtain an evaluation result; and identify the identity state of the target object based on the evaluation result to obtain the first state recognition result. detect a video series of the face based on the bounding box corresponding to the face of the target object, wherein the video series of the face is a buffered face sequence; extract spatio-temporal features of the face from the video series of the face, and identify an expression state of the target object to obtain the second state identification result; and / or detect a video sequence of the body based on the bounding box corresponding to the body of the target object, wherein the video sequence of the body is a buffered full-body expansion sequence; extract spatio-temporal features of the body from the video sequence of the body, and identify a behavior state of the target object to obtain the second state identification result.
21. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a stored program, wherein the program, when executed by a processor, controls a device in which the computer readable storage medium is located to perform the method of any one of claims 1 to 16.
22. A processor, comprising: The processor is configured to execute a program, wherein the program, when executed by the processor, performs the method of any one of claims 1 to 16.
23. An identification system of objects, characterized in that Comprise: A processor; A memory connected to the processor, configured to provide the processor with instructions for processing the following processing steps: determining a plurality of bounding boxes of a target object, wherein the bounding boxes are used to detect target parts corresponding to a state of the target object; detecting the corresponding target parts based on the plurality of bounding boxes respectively, and identifying the state of the target object based on the obtained detection results to obtain a plurality of state identification results, wherein the plurality of state identification results comprise a first state identification result of the target object, and the first state identification result is used to identify the target object; and associating the first state identification result with a second state identification result to obtain an association result, wherein the second state identification result is any state identification result in the plurality of state identification results except the first state identification result; The memory is further configured to provide the processor with instructions for processing the following processing steps: performing pose estimation on the head of the target object based on a bounding box corresponding to the head to obtain pose information; positioning key points corresponding to the pose information to obtain a positioning result; performing face quality evaluation based on the positioning result to obtain an evaluation result; and identifying an identity state of the target object based on the evaluation result to obtain the first state identification result; detect a video series of the face based on the bounding box corresponding to the face of the target object, wherein the video series of the face is a buffered face sequence; extract spatio-temporal features of the face from the video series of the face, and identify an expression state of the target object to obtain the second state identification result; and / or detect a video sequence of the body based on the bounding box corresponding to the body of the target object, wherein the video sequence of the body is a buffered full-body expansion sequence; extract spatio-temporal features of the body from the video sequence of the body, and identify a behavior state of the target object to obtain the second state identification result.
Citation Information
Patent Citations
Video processing method and device, electronic device and storage medium
CN110675433A
Video data processing method and device and storage medium
CN111586466A