A method, apparatus and readable medium for controlling a distributed display device
By performing target detection and feature recognition on the image of the operating area, the unique operator's gesture posture can be determined, solving the problems of cumbersome operation and misoperation of traditional display devices, and realizing precise control of distributed display devices.
Patent Information
- Application Number
- CN202411411943.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-10-11
AI Technical Summary
Traditional display device control methods are cumbersome and not intuitive enough. In situations with multiple people, non-operators are prone to misoperation, and the accuracy and robustness are insufficient.
By performing target detection on the current frame image of the operation area, a human body detection box is obtained. Combined with facial feature information and skeletal recognition, the gesture posture of the unique operator is determined, thereby achieving precise control.
This avoids accidental touches in situations with multiple users, enables precise operation of distributed display devices, and improves the accuracy and robustness of control.
Smart Images

Figure CN119166089B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic information technology, and in particular to a method, apparatus, and readable medium for controlling distributed display devices. Background Technology
[0002] With the rapid development of information technology, distributed display devices have been widely used in public places, commercial displays, home entertainment, and many other fields. However, traditional control methods for display devices, such as physical buttons and remote controls, are cumbersome and not intuitive. In recent years, although vision-based control methods have emerged, when multiple people are present, misoperation by non-operators often occurs, and their accuracy and robustness still need to be improved. Summary of the Invention
[0003] To address at least one of the aforementioned technical problems, this disclosure proposes a method for controlling a distributed display device in a first aspect, comprising: performing target detection on a current frame image of an operating area to obtain human body detection boxes; comparing a first distance and a first threshold between each human body detection box in the current frame image and a pre-defined box of the operator in the current frame image; calculating the ratio of the intersection and union of the area of the human body detection box and the area of the pre-defined box when the first distance is less than the first threshold; when there is one human body detection box with the largest ratio, determining the human body detection box with the largest ratio as the human body detection box of the operator in the current frame image; when there is more than one human body detection box with the largest ratio, comparing the length and width of multiple human body detection boxes with the largest ratios, determining the human body detection box with the largest length and width as the human body detection box of the operator in the current frame image; determining the operator's left or right hand and obtaining the operator's gesture posture based on the human body detection box of the operator in the current frame image and the operator's facial feature information; and responding to the operation corresponding to the gesture posture through the display device.
[0004] Preferably, it further includes: acquiring the previous few frames of images of the operation area, and obtaining the pre-frame bounding box of the operator in the current frame image through the trajectory of the operator in the previous few frames of images of the operation area.
[0005] Preferably, skeletal recognition is performed on the current frame image where the operator's human detection bounding box has been determined, to obtain multiple skeletal node information and corresponding multiple skeletal bounding boxes; the magnitudes of a second distance and a second threshold are compared, where the second distance is the distance between each skeletal bounding box in the current frame image and the determined operator's human detection bounding box in the current frame image; the ratio of the intersection and union of the area of the skeletal bounding box and the area of the human detection bounding box when the second distance is less than the second threshold is calculated; when there is only one skeletal bounding box with the largest ratio, the skeletal bounding box with the largest ratio is determined as the operator's skeletal bounding box in the current frame image; when there is more than one skeletal bounding box with the largest ratio, the length and width of multiple skeletal bounding boxes with the largest ratios are compared, and the skeletal bounding box with the largest length and width is determined as the operator's skeletal bounding box in the current frame image.
[0006] Preferably, the skeletal node corresponding to the skeletal bounding box of the operator in the current frame image is determined as the skeletal node of the operator in the current frame image; based on the skeletal node of the operator in the current frame image, the skeletal face bounding box of the operator in the current frame image is obtained.
[0007] Preferably, face detection is performed on the current frame image where the operator's human body detection bounding box has been determined, obtaining multiple face bounding boxes; the magnitude of a third distance and a third threshold are compared, where the third distance is the distance between each face bounding box in the current frame image and the operator's skeletal face bounding box in the current frame image; the ratio of the intersection and union of the areas of the face bounding boxes and the skeletal face bounding boxes when the third distance is less than the third threshold is calculated; when there is only one face bounding box with the largest ratio, the face bounding box with the largest ratio is determined as the operator's face bounding box in the current frame image; when there is more than one face bounding box with the largest ratio, the length and width of multiple face bounding boxes with the largest ratios are compared, and the face bounding box with the largest length and width is determined as the operator's face bounding box in the current frame image; feature extraction is performed on the operator's face bounding box in the current frame image to obtain the operator's facial features.
[0008] Preferably, multiple face features and corresponding multiple face bounding boxes of the current frame image are obtained through face detection and recognition. Multiple face features and corresponding multiple face bounding boxes are selected based on the operator's face features. The size of a fourth distance and a fourth threshold are compared. The fourth distance is the distance between each selected face bounding box and a pre-defined box. The ratio of the intersection and union of the area of the face bounding box and the area of the pre-defined box when the fourth distance is less than the fourth threshold is calculated. When there is only one face bounding box with the largest ratio, the face bounding box with the largest ratio is determined as the operator's face bounding box in the current frame image. When there is more than one face bounding box with the largest ratio, the length and width of multiple face bounding boxes with the largest ratios are compared, and the face bounding box with the largest length and width is determined as the operator's face bounding box in the current frame image.
[0009] Preferably, skeletal recognition is performed on the current frame image where the operator's face bounding box is determined, to obtain multiple skeletal node information and corresponding multiple skeletal bounding boxes; according to preset conditions, the fifth distance and the fifth threshold between each skeletal bounding box in the current frame image and the determined operator's face bounding box in the current frame image are compared, and the ratio of the intersection and union of the area of the skeletal bounding box and the area of the face bounding box when the fifth distance is less than the fifth threshold is calculated; when there is only one skeletal bounding box with the largest ratio, the skeletal bounding box with the largest ratio is determined as the operator's skeletal bounding box in the current frame image; when there is more than one skeletal bounding box with the largest ratio, the length and width of multiple skeletal bounding boxes with the largest ratios are compared, and the skeletal bounding box with the largest length and width is determined as the operator's skeletal bounding box in the current frame image.
[0010] Preferably, based on the operator's skeletal bounding box in the current frame image, the skeletal face bounding box of the operator in the current frame image is obtained; when the center distance between the operator's face bounding box in the current frame image and the operator's skeletal face bounding box in the current frame image is less than a threshold or the intersection area is greater than a threshold, it is determined whether the relationship between the operator's face bounding box in the current frame image and the operator's skeletal face bounding box meets an adjustable preset condition; if it does, the current frame image is determined as a valid frame image; hand skeleton recognition is performed on the valid frame image to obtain multiple hand node information; based on the hand node information of the valid frame and the skeletal node information corresponding to the current operator's skeletal bounding box, the hand node information of the current operator's left or right hand is determined; based on the hand node information of the current operator's left or right hand, the gesture of the current operator's left or right hand is determined.
[0011] The second aspect of this disclosure includes: an acquisition module for performing target detection on the current frame image of the operation area to obtain a human body detection box; a judgment module for comparing a first distance and a first threshold between each human body detection box in the current frame image and a pre-defined box of the operator in the current frame image, calculating the ratio of the intersection and union of the area of the human body detection box and the area of the pre-defined box when the first distance is less than the first threshold, determining the human body detection box with the largest ratio as the human body detection box of the operator in the current frame image when there is only one human body detection box with the largest ratio, and comparing the length and width of multiple human body detection boxes with the largest ratios to determine the human body detection box of the operator with the largest length and width as the human body detection box of the operator in the current frame image; and a determination module for determining the operator's left or right hand and obtaining the operator's gesture posture based on the human body detection box of the operator in the current frame image and the operator's facial feature information, and responding to the operation corresponding to the gesture posture through a display device.
[0012] In a third aspect, this disclosure provides a computer-readable medium storing a computer program that is loaded and executed by a processing module to implement the steps of any of the methods described above.
[0013] Some technical advantages of this disclosure are: a method for controlling a distributed display device, comprising:
[0014] The system performs target detection on the current frame image of the operation area to obtain human detection boxes. It compares the first distance and a first threshold between each human detection box in the current frame image and a pre-defined box containing the operator in the current frame image. It calculates the ratio of the intersection to the union of the areas of the human detection boxes and the pre-defined boxes when the first distance is less than the first threshold. If there is only one human detection box with the largest ratio, it is determined as the operator's human detection box in the current frame image. If there are more than one human detection box with the largest ratio, the length and width of multiple human detection boxes with the largest ratios are compared, and the human detection box with the largest length and width is determined as the operator's human detection box in the current frame image. Based on the operator's human detection box in the current frame image and the operator's facial feature information, it determines whether the operator has their left or right hand and obtains the operator's gesture posture. The display device responds to the operation corresponding to the gesture posture. This ensures that a unique operator uses only their left or right hand to operate the display device, avoiding accidental touches in multi-person scenarios. Ultimately, it enables precise control of distributed display devices. Attached Figure Description
[0015] To better understand the technical solutions of this disclosure, reference can be made to the following accompanying drawings, which are used to assist in the illustration of the prior art or embodiments. These drawings selectively illustrate the products or methods involved in the prior art or some embodiments of this disclosure. The basic information of these drawings is as follows:
[0016] Figure 1 This is a flowchart of one embodiment of a method for controlling a distributed display device according to this application. Detailed Implementation
[0017] The following will further describe the technical means or effects involved in this disclosure. Obviously, the provided embodiments (or implementation methods) are only some of the implementation methods covered by this disclosure, and not all of them. Based on the embodiments in this disclosure and the explicit or implicit descriptions in the figures and text, all other embodiments that can be obtained by those skilled in the art without creative effort will be within the scope of protection claimed in this disclosure.
[0018] Existing methods for controlling display devices, such as volume control, are primarily single-handed. Therefore, when both hands are present, it's impossible to determine which hand is controlling the volume, leading to difficulty adjusting the volume or accidental volume adjustment. In other situations, such as when multiple people are present, other people's hands might also be identified and used to control the volume, resulting in unauthorized volume control and a poor user experience. Furthermore, if someone is standing behind the operator, existing technologies can easily misidentify them as the operator.
[0019] This application discloses a method, apparatus, and readable medium for controlling a distributed display device. The method for controlling a distributed display device provided in this application can be applied to various industries and other scenarios requiring rapid and precise human-computer interaction through controlling a large screen on a display device. Examples include emergency command and dispatch centers, public security command and dispatch centers, traffic command and dispatch centers, energy command and dispatch centers, and smart city command and dispatch centers. By controlling the large screen, the dispatch system can be controlled, such as switching distributed signal sources or taking over the mouse within the signal source, thereby allowing arbitrary operations on the content within the signal source. As the central brain of command and dispatch control, the command center plays a crucial role in social governance and people's livelihood development, requiring high accuracy and speed in operation. Therefore, the method for controlling the large screen in this disclosure does not require any complex control equipment or wearable sensors. It can quickly and accurately take over and control the command center's large screen using only the behavior recognition of a living person. Through simple air gesture operations, it efficiently achieves rapid and precise interactive operations such as signal input, switching, and scaling between the person and the large screen content. Of course, the method disclosed herein can also be applied to general application scenarios with less demanding requirements, such as performing page turning operations in PPT.
[0020] An exemplary system architecture for controlling a distributed display device, or an apparatus for controlling a distributed display device, can be applied using the methods and apparatus disclosed herein. The system architecture may include a camera with a pan-tilt-zoom (PTZ) sensor, a server, and a display device. For example, the display device is a large screen. The PTZ camera is connected to the server via a serial cable and a USB cable, and the server is then connected to the display device via a network cable. Image information captured by the camera is transmitted to the server via the USB cable. The server processes, analyzes, and makes decisions based on the received image information, generating information or commands. The server may also send the information to a distributed scheduling and image management platform via the network port. The distributed scheduling and image management platform receives the information and displays corresponding operation feedback on the large screen.
[0021] Display devices generally require large screens, multiple colors, high brightness, and high resolution. For example, a display device is a large-screen display, referring to the large screen in a direct-view color television or a rear-projection television; typically, the diagonal size of the screen is over 40 inches. The display surface of a large-screen display can be flat or curved. Large-screen display devices can also be tiled, and there are no further restrictions.
[0022] In this embodiment, the camera with a gimbal is located directly above the large screen. The gimbal is the device that supports the camera.
[0023] The target user can interact with the server via a camera using gestures, and then the server interacts with the large screen to achieve gesture-based control of the large screen. The server can be a single server, a server cluster consisting of several servers, or a cloud computing center. The server can provide various services to the display device. For different applications on the display device, the server can be considered a backend server providing corresponding network services. Therefore, the method disclosed in this application can be considered to be primarily executed by the server side.
[0024] In one embodiment of this disclosure, an application scenario includes a large screen, a pan-tilt-zoom camera positioned above the large screen, and an operable area in front of the large screen. The operable area is roughly a ring-shaped region. The operator can control the large screen within this operable area. If outside the operable area, for example, if too far from the large screen, gesture recognition may fail, potentially leading to control errors. If too close to the large screen, the operator may not be able to observe the entire content, hindering screen operation. In this embodiment, the large screen is 10 meters wide, and the operable area is a ring-shaped region ranging from 3 to 12 meters from the large screen.
[0025] like Figure 1 The diagram illustrates a flowchart of an embodiment of a method for controlling a distributed display device according to the present disclosure. The method includes the following steps:
[0026] S10: Perform target detection on the current frame image of the operation area to obtain the human detection box;
[0027] S20: Compare the first distance and the first threshold between each human detection box in the current frame image and the pre-made box of the operator in the current frame image. Calculate the ratio of the intersection and union of the area of the human detection box and the area of the pre-made box when the first distance is less than the first threshold. When there is only one human detection box with the largest ratio, determine the human detection box with the largest ratio as the human detection box of the operator in the current frame image. When there is more than one human detection box with the largest ratio, compare the length and width of multiple human detection boxes with the largest ratio, and determine the human detection box with the largest length and width as the human detection box of the operator in the current frame image.
[0028] S30: Based on the human body detection box in the current frame image and the operator's facial feature information, determine whether the operator has left or right hand and obtain the operator's gesture posture, and respond to the operation corresponding to the gesture posture through the display device.
[0029] The camera captures images of the operating area and transmits them to the server. The camera acquires the first few frames of the operating area, and the operator's trajectory in these frames is used to obtain a pre-defined bounding box for the operator in the current frame. The current frame of the operating area is then acquired, and object detection is performed to obtain human detection boxes. If no person is detected in the current frame, detection ends, the next frame is acquired, and the object detection process is repeated. When a person is detected within the pre-defined bounding box in the current frame, human detection boxes for all persons located in the operating area are obtained.
[0030] S20: Compare the first distance and the first threshold between each human detection box in the current frame image and the pre-made box of the operator in the current frame image. Calculate the ratio of the intersection and union of the area of the human detection box and the area of the pre-made box when the first distance is less than the first threshold. When there is only one human detection box with the largest ratio, determine the human detection box with the largest ratio as the human detection box of the operator in the current frame image. When there is more than one human detection box with the largest ratio, compare the length and width of multiple human detection boxes with the largest ratio, and determine the human detection box with the largest length and width as the human detection box of the operator in the current frame image.
[0031] In this embodiment, the first distance refers to the center distance between each detected human body bounding box and the pre-defined bounding box of the operator in the current frame image. The first threshold can be preset, and its value can be obtained based on the relationship between the maximum and minimum center distances. Human body bounding boxes whose center distance to the pre-defined bounding box of the operator in the current frame image is within the first threshold range are retained, while human body bounding boxes whose center distance to the pre-defined bounding box of the operator in the current frame image is not within the first threshold range are discarded.
[0032] Then, the human detection boxes are further filtered based on the intersection and union ratio of the areas of the human detection boxes and the pre-made boxes. The ratio of the intersection and union of the areas of the human detection boxes and the pre-made boxes when the first distance is less than the first threshold is calculated, and the human detection boxes with the largest ratio are retained.
[0033] If the human detection bounding boxes obtained in the previous step are still complex values, further filtering is performed to retain the human detection bounding boxes with the largest length and width. This is because the length, width, and sides of the human detection bounding box reflect the distance relationship between different people in front of the camera and the camera; the larger the value, the closer they are to the camera. This avoids misidentifying other people standing behind the operator who are closer to the center of the pre-defined box as the operator. Moreover, both the length and width are simultaneously at their maximum or minimum. That is, the length and width of the human detection bounding box of the operator in front are both greater than the length and width of the human detection bounding boxes of others behind. This is especially important in current real-world scenarios where misidentification of operators frequently occurs. For example, when someone stands behind the controller and is closer to the exact center of the pre-defined box than the controller, the wrong person may be selected. This is particularly true when using the DIOU Loss loss function, which involves bounding box filtering in the field of object detection. The distance relationship in DIOU Loss is the square of the center distance between the ground truth bounding box and the detection bounding box divided by the square of the diagonal of the outermost bounding box surrounding the two boxes.
[0034] S30: Based on the human body detection box in the current frame image and the operator's facial feature information, determine whether the operator has left or right hand and obtain the operator's gesture posture, and respond to the operation corresponding to the gesture posture through the display device.
[0035] For the current frame image with the determined human body detection bounding box of the operator, skeletal recognition is performed to obtain multiple skeletal node information and corresponding multiple skeletal bounding boxes. The size of a second distance and a second threshold are compared. The second distance is the distance between each skeletal bounding box in the current frame image and the determined human body detection bounding box of the operator in the current frame image. The second distance can be the center distance between the two. Skeletal bounding boxes with a second distance greater than the second threshold are discarded, and skeletal bounding boxes with a second distance less than the second threshold are retained. The ratio of the intersection and union of the area of the skeletal bounding boxes with a second distance less than the second threshold and the area of the operator's human body detection bounding box is calculated. If there is one skeletal bounding box with the largest ratio, it is determined as the operator's skeletal bounding box in the current frame image. If there are more than one skeletal bounding box with the largest ratio, the length and width of the multiple skeletal bounding boxes with the largest ratios are compared, and the skeletal bounding box with the largest length and width is determined as the operator's skeletal bounding box in the current frame image. "Maximum length and width" refers to either the maximum length or the maximum width; generally, both are simultaneously maximum. Because the operator is closest to the camera, the width and length of the operator's human body detection bounding box are both maximum.
[0036] Each person's skeletal bounding box has corresponding skeletal nodes. The skeletal nodes corresponding to the operator's bounding box in the current frame are determined as the operator's skeletal nodes in the current frame. Based on the operator's skeletal nodes in the current frame, the skeletal face bounding box of the operator in the current frame is obtained. Skeletal nodes include eye nodes, nose nodes, etc., which allow for the calculation of the approximate skeletal face bounding box.
[0037] Face detection is performed on the current frame image with the operator's human body detection bounding box determined, obtaining multiple face bounding boxes. The magnitude of the third distance and the third threshold are compared. The third distance is the distance between each face bounding box in the current frame image and the operator's skeletal face bounding box in the current frame image. The third distance can be the center distance between the two boxes. Face bounding boxes with a third distance greater than the third threshold are discarded. The ratio of the intersection and union of the area of the face bounding box and the area of the skeletal face bounding box with a third distance less than the third threshold is calculated. When there is one face bounding box with the largest ratio, it is determined as the operator's face bounding box in the current frame image. When there is more than one face bounding box with the largest ratio, the length and width of the multiple face bounding boxes with the largest ratios are compared, and the face bounding box with the largest length and width is determined as the operator's face bounding box in the current frame image. Feature extraction is performed on the operator's face bounding box in the current frame image to obtain the operator's facial features.
[0038] After obtaining the operator's facial features as described above, the next step is to determine the unique hand gesture of the operator in any frame of the image. Multiple facial features and corresponding multiple bounding boxes are obtained from the current frame image through face detection and recognition. A subset of these multiple facial features and corresponding bounding boxes are then selected based on the operator's facial features. For example, this can be done by filtering based on the similarity between each detected facial feature and the previously determined operator's facial features. Bounding boxes corresponding to facial features that do not meet the similarity criteria are discarded, while those corresponding to facial features that do meet the similarity criteria are retained. The fourth distance and the fourth threshold are compared. The fourth distance is the distance between each selected face bounding box and the pre-defined bounding box. The ratio of the intersection and union of the areas of the face bounding boxes and the pre-defined bounding boxes when the fourth distance is less than the fourth threshold is calculated. When there is only one face bounding box with the largest ratio, the face bounding box with the largest ratio is determined as the operator's face bounding box in the current frame image. When there is more than one face bounding box with the largest ratio, the length and width of the multiple face bounding boxes with the largest ratios are compared, and the face bounding box with the largest length and width is determined as the operator's face bounding box in the current frame image.
[0039] For the current frame image where the operator's face bounding box has been determined, skeletal recognition is performed to obtain multiple skeletal node information and corresponding multiple skeletal bounding boxes. Based on preset conditions, the fifth distance and a fifth threshold are compared between each skeletal bounding box in the current frame image and the determined operator's face bounding box in the current frame image. The ratio of the intersection and union of the areas of the skeletal bounding boxes and the face bounding boxes when the fifth distance is less than the fifth threshold is calculated. When there is only one skeletal bounding box with the largest ratio, it is determined as the operator's skeletal bounding box in the current frame image. When there are more than one skeletal bounding box with the largest ratio, the length and width of the multiple skeletal bounding boxes with the largest ratios are compared, and the skeletal bounding box with the largest length and width is determined as the operator's skeletal bounding box in the current frame image. The preset conditions can be adjusted according to the on-site conditions of the distributed system and are not limited to the center distance between two bounding boxes being less than a threshold or the intersection area of two bounding boxes being greater than a threshold.
[0040] Based on the operator's skeletal bounding box in the current frame image, obtain the operator's skeletal face bounding box in the current frame image; the calculation method for obtaining the skeletal face bounding box is the same as described above. When the center distance between the operator's face bounding box in the current frame image and the operator's skeletal face bounding box in the current frame image is less than a threshold or the intersection area is greater than a threshold, determine whether the relationship between the operator's face bounding box in the current frame image and the operator's skeletal face bounding box meets an adjustable preset condition. If it does, the current frame image is determined as a valid frame image. Perform hand skeleton recognition on the valid frame image to obtain multiple hand node information. Based on the hand node information of the valid frame and the skeletal node information corresponding to the current operator's skeletal bounding box, determine the left or right hand node information of the current operator. Based on the left or right hand node information of the current operator, determine the gesture of the current operator's left or right hand. Obtain the operation mode corresponding to the gesture posture and respond to the operation corresponding to the gesture posture through the display device. For example, perform operations such as turning pages in a PPT.
[0041] This disclosure provides, in a second aspect, an apparatus for controlling a distributed display device, comprising:
[0042] The acquisition module is used to perform target detection on the current frame image of the operation area to obtain human detection boxes; the judgment module is used to compare the size of a first distance and a first threshold between each human detection box in the current frame image and the pre-defined box of the operator in the current frame image, calculate the ratio of the intersection and union of the area of the human detection box and the area of the pre-defined box when the first distance is less than the first threshold, and determine the human detection box with the largest ratio as the human detection box of the operator in the current frame image when there is only one human detection box with the largest ratio. When there is more than one human detection box with the largest ratio, compare the length and width of multiple human detection boxes with the largest ratio and determine the human detection box with the largest length and width as the human detection box of the operator in the current frame image; the determination module is used to determine the operator's left or right hand and obtain the operator's gesture posture based on the human detection box of the operator in the current frame image and the operator's facial feature information, and respond to the operation corresponding to the gesture posture through the display device.
[0043] In a third aspect, this disclosure provides a computer-readable medium storing a computer program, which is loaded and executed by a processing module to implement the steps of the acquisition method. Those skilled in the art will understand that all or part of the steps in the embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable medium, which may include various media capable of storing program code, such as flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0044] Within the scope of knowledge and ability of those skilled in the art, the various embodiments or technical features mentioned herein can be combined with each other as other optional embodiments without conflict. These finite number of optional embodiments, which are not listed one by one and are formed by combining a finite number of technical features, still fall within the scope of the technology disclosed herein and are also derived by those skilled in the art from the accompanying drawings and the foregoing.
[0045] In addition, the descriptions of most embodiments are based on different focuses. For further understanding of the parts not described in detail, reasonable inference can be made by referring to the relevant content of the prior art, other relevant descriptions in this document, or the inventive intent.
[0046] To reiterate, the embodiments listed above are typical and preferred embodiments of this disclosure, and are only used to describe and explain the technical solutions of this disclosure in detail to facilitate the reader's understanding. They are not intended to limit the scope or application of the protection claimed in this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure to obtain technical solutions should be covered within the scope of protection claimed in this disclosure.
Claims
1. A method for controlling a distributed display device, characterized in that, include: Perform target detection on the current frame image of the operation area to obtain the human detection box; Compare the first distance and the first threshold between each human detection box in the current frame image and the pre-made box of the operator in the current frame image. Calculate the ratio of the intersection and union of the areas of the human detection boxes and the pre-made boxes when the first distance is less than the first threshold. When there is only one human detection box with the largest ratio, determine the human detection box with the largest ratio as the human detection box of the operator in the current frame image. When there is more than one human detection box with the largest ratio, compare the length and width of multiple human detection boxes with the largest ratio, and determine the human detection box with the largest length and width as the human detection box of the operator in the current frame image. Based on the operator's human detection bounding box and facial feature information in the current frame image, determine whether the operator has their left or right hand and obtain the operator's gesture posture. This includes: obtaining multiple facial features and corresponding multiple face bounding boxes in the current frame image through face detection and recognition; filtering out multiple facial features and corresponding multiple face bounding boxes based on the operator's facial features; comparing the magnitude of a fourth distance and a fourth threshold, where the fourth distance is the distance between each filtered face bounding box and a pre-defined box; calculating the ratio of the intersection and union of the area of the face bounding box and the area of the pre-defined box when the fourth distance is less than the fourth threshold; when there is only one face bounding box with the largest ratio, determine that the face bounding box with the largest ratio is the operator's face bounding box in the current frame image; when there are more than one face bounding box with the largest ratio, compare the length and width of multiple face bounding boxes with the largest ratios, and determine that the face bounding box with the largest length and width is the operator's face bounding box in the current frame image. The face bounding box in the previous frame image is used to obtain the skeletal face bounding box of the operator in the current frame image. If the center distance between the operator's face bounding box and the operator's skeletal face bounding box in the current frame image is less than a threshold or the intersection area is greater than a threshold, it is determined whether the relationship between the operator's face bounding box and the operator's skeletal face bounding box in the current frame image meets an adjustable preset condition. If it does, the current frame image is determined to be a valid frame image. Hand skeleton recognition is performed on the valid frame image to obtain multiple hand node information. Based on the hand node information of the valid frame and the skeletal node information corresponding to the current operator's skeletal bounding box, the hand node information of the current operator's left or right hand is determined. Based on the hand node information of the current operator's left or right hand, the gesture of the current operator's left or right hand is determined, and the operation corresponding to the gesture posture is responded to through the display device.
2. The method according to claim 1, characterized in that, Also includes: Acquire the previous few frames of the operation area, and obtain the pre-defined bounding box of the operator in the current frame image by using the operator's trajectory in the previous few frames of the operation area.
3. The method according to claim 1, characterized in that, Perform skeletal recognition on the current frame image of the human body detection box that has determined the operator, and obtain multiple skeletal node information and corresponding multiple skeletal boxes; Compare the magnitudes of the second distance and the second threshold. The second distance is the distance between each skeleton bounding box in the current frame image and the human detection box of the determined operator in the current frame image. Calculate the ratio of the intersection to the union of the areas of the skeleton bounding boxes and the human detection boxes when the second distance is less than the second threshold. When there is only one skeleton with the largest ratio, the skeleton with the largest ratio is determined to be the skeleton of the operator in the current frame image. When there is more than one bone box with the largest ratio, compare the length and width of the multiple bone boxes with the largest ratio, and determine the bone box with the largest length and width as the operator's bone box in the current frame image.
4. The method according to claim 3, characterized in that, The skeletal node corresponding to the operator's bounding box in the current frame image is determined as the operator's skeletal node in the current frame image; Based on the operator's skeletal nodes in the current frame image, obtain the operator's skeletal face bounding box in the current frame image.
5. The method according to claim 4, characterized in that, Perform face detection on the current frame image where the human body detection bounding box of the operator has been determined, and obtain multiple face bounding boxes; Compare the third distance with the third threshold. The third distance is the distance between each face bounding box in the current frame image and the operator's skeletal face bounding box in the current frame image. Calculate the ratio of the intersection to the union of the area of the face bounding box and the area of the skeletal face bounding box when the third distance is less than the third threshold. When there is only one face bounding box with the largest ratio, the face bounding box with the largest ratio is determined to be the operator's face bounding box in the current frame image; When there is more than one face bounding box with the largest ratio, compare the length and width of the multiple face bounding boxes with the largest ratio, and determine the face bounding box with the largest length and width as the operator's face bounding box in the current frame image; Feature extraction is performed on the operator's face bounding box in the current frame image to obtain the operator's facial features.
6. The method according to claim 1, characterized in that, Perform skeletal recognition on the current frame image of the face bounding box of the operator to obtain multiple skeletal node information and corresponding multiple skeletal bounding boxes; Based on preset conditions, compare the fifth distance and the fifth threshold between each skeleton bounding box in the current frame image and the face bounding box of the determined operator in the current frame image, and calculate the ratio of the intersection and union of the area of the skeleton bounding box and the area of the face bounding box when the fifth distance is less than the fifth threshold. When there is only one skeleton with the largest ratio, the skeleton with the largest ratio is determined to be the skeleton of the operator in the current frame image. When there is more than one bone box with the largest ratio, compare the length and width of the multiple bone boxes with the largest ratio, and determine the bone box with the largest length and width as the operator's bone box in the current frame image.
7. A device for controlling a distributed display device, characterized in that, include: The acquisition module is used to perform target detection on the current frame image of the operation area and obtain the human detection box; The judgment module is used to compare the size of the first distance and the first threshold between each human detection box in the current frame image and the pre-made box of the operator in the current frame image, calculate the ratio of the intersection and union of the area of the human detection box and the area of the pre-made box when the first distance is less than the first threshold, and determine the human detection box with the largest ratio as the human detection box of the operator in the current frame image when there is only one human detection box with the largest ratio. When there is more than one human detection box with the largest ratio, compare the length and width of multiple human detection boxes with the largest ratio and determine the human detection box with the largest length and width as the human detection box of the operator in the current frame image. The determination module is used to determine the operator's left or right hand and obtain the operator's gesture posture based on the human detection bounding box and the operator's facial feature information in the current frame image. This includes: obtaining multiple facial features and corresponding multiple face bounding boxes in the current frame image through face detection and recognition; filtering out multiple facial features and corresponding multiple face bounding boxes based on the operator's facial features; comparing the magnitude of a fourth distance and a fourth threshold, where the fourth distance is the distance between each filtered face bounding box and a pre-defined bounding box; calculating the ratio of the intersection and union of the areas of face bounding boxes and pre-defined bounding boxes when the fourth distance is less than the fourth threshold; determining the face bounding box with the largest ratio as the operator's face bounding box in the current frame image when there is only one face bounding box with the largest ratio; and comparing the length and width of multiple face bounding boxes with the largest ratios to determine the face bounding box with the largest length and width as the operator's face bounding box. The author obtains the operator's skeletal face bounding box in the current frame image based on the operator's skeletal face bounding box in the current frame image. If the center distance between the operator's face bounding box and the operator's skeletal face bounding box in the current frame image is less than a threshold or the intersection area is greater than a threshold, it is determined whether the relationship between the operator's face bounding box and the operator's skeletal face bounding box in the current frame image meets an adjustable preset condition. If it does, the current frame image is determined as a valid frame image. Hand skeleton recognition is performed on the valid frame image to obtain multiple hand node information. Based on the hand node information of the valid frame and the skeletal node information corresponding to the current operator's skeletal bounding box, the left or right hand node information of the current operator is determined. Based on the left or right hand node information of the current operator, the gesture of the current operator's left or right hand is determined, and the operation corresponding to the gesture posture is responded to through the display device.
8. A computer-readable medium, characterized in that: The computer-readable medium stores a computer program, which is loaded and executed by a processing module to implement the steps of the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Behavior recognition method and device, electronic equipment and computer storage medium
CN116129522A
Human body posture key point identification method and device for auxiliary rehabilitation exercise
CN118570838A