A precise and fast light chasing method

By combining visual methods with facial recognition and position prediction, the problem of accurate and rapid tracking of target figures in stage lighting systems has been solved, achieving a fast and accurate light following effect.

CN119445630BActive Publication Date: 2025-11-18GUANGZHOU HAOYANG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411545214.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-11-18
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

In existing stage lighting systems, there are problems such as low precision in manual control when the light follows people on the stage, inaccurate positioning or jitter caused by wearing tags, slow machine vision recognition speed and easy loss of the target when it is blocked. There is a lack of accurate and fast tracking methods.

Method used

Using a purely visual approach, combining facial recognition and location prediction, the system continuously captures live footage with a camera, detects target individuals and determines their location confidence, and then uses Kalman filtering and Hungarian algorithms for location matching to achieve accurate and rapid light tracking.

Benefits of technology

It enables the rapid and accurate locking of target figures in the stage lighting system, balancing the speed of position tracking with the accuracy of facial recognition, avoiding slow light tracking speeds caused by obstructions or poor lighting, and saving computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119445630B_ABST
    Figure CN119445630B_ABST
Patent Text Reader

Abstract

The application discloses a precise and fast light tracking method, which comprises the following steps: finding a target person from a live picture through face recognition, judging the confidence of the real-time position of the target person, marking the next frame for face recognition tracking of the target person when the confidence is lower than a first preset value, matching the predicted position of the target person with the real-time positions of all persons in the next frame if the face recognition fails or is not marked, taking the corresponding person matched as the target person, and then entering the tracking of the target person in the next frame, taking the predicted position of the target person as the real-time position when no corresponding person is matched, and then performing face recognition in the next frame to track the real-time position of the target person by light. The method mainly tracks the position, combines the judgment of the confidence of the real-time position of the current frame, decides whether to perform face recognition in the next live picture, and takes into account the rapidity of position tracking and the precision of face recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of stage lighting technology, and more specifically, to a precise and rapid method for tracking light. Background Technology

[0002] In current stage lighting systems, there are generally three ways for lights to follow people on stage: One method requires the lighting technician to manually control the light's movement to illuminate a specific person. Since the lights are usually far from the stage, manual tracking accuracy is greatly affected by human factors and is prone to jitter. Another common method involves stage actors wearing special tags, such as UWB positioning tags or infrared marker tags. However, this can be inconvenient for actors in certain situations, and the data transmission of the tags can be affected by interference from the human body or the surrounding environment, leading to inaccurate positioning or jitter. A third method uses machine vision to identify and track the target human body. However, machine vision recognition is slow, and the target is easily lost when obscured. Therefore, a precise and fast light-tracking method is needed. Summary of the Invention

[0003] To overcome at least one of the defects described in the prior art, this invention provides a precise and rapid light-tracking method that uses pure vision to quickly and accurately lock onto a target person.

[0004] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a precise and rapid light-tracking method, comprising the following steps:

[0005] S1. Continuously monitor the scene captured by the camera until a target person with specific facial features is identified, and obtain the real-time location and motion information of the target person.

[0006] S2. The light from the lamp tracks the real-time position of the target person and determines whether the confidence level of the real-time position is less than the first preset value. If so, it is marked that face recognition needs to be performed in the next frame.

[0007] S3. Based on the target person's real-time position and motion information, obtain the predicted position of the target person in the next frame;

[0008] S4. The camera captures the next frame of the scene and obtains the real-time position and face position of all the people in the next frame of the scene;

[0009] S5. When face recognition is required in the next frame, the faces of all people in the scene in the next frame of step S4 are recognized. If the target person is recognized, the real-time position of the target person in this scene is updated and marked as not requiring face recognition in the next frame. Then, the process jumps to step S2. If the target person is not recognized, the process jumps to step S6. If face recognition is marked as not requiring face recognition in the next frame in step S2, the process jumps directly to step S6.

[0010] S6. Match the predicted position of the target person in the next frame in step S3 with the real-time positions of all people in the scene in the next frame in step S4. If the match is successful, the matched person is considered to be the target person. Update the real-time position of the target person in this scene and jump to step S2. If the match is unsuccessful, the predicted position of the target person in step S3 is considered to be the real-time position of the target person in the next frame of the scene, and the motion information remains unchanged. The confidence of the real-time position is considered to be less than the first preset value. Then jump to step S2.

[0011] The precise and rapid light-tracking method first locates the target person in the scene using facial recognition. Then, it determines the confidence level of the target person's real-time position. If the confidence level is lower than a first preset value, it marks the next frame of the scene as requiring facial recognition-assisted tracking of the target person. If facial recognition fails or is not marked, it directly matches the predicted position of the target person in the next frame with the real-time positions of all people in the next frame of the scene. When a matching person is found, it is taken as the target person, and the real-time position and confidence level of the target person are updated. Then, it proceeds to the next frame for target person tracking. When no matching person is found, the predicted position of the target person in the next frame is taken as the real-time position of the target person in the next frame of the scene, and it is assumed that the motion information remains unchanged. Then, it performs facial recognition in the next frame to track the target person, while the light from the lamp tracks the real-time position of the target person.

[0012] This method primarily uses position tracking and predicts the target person's position in the next frame. It also combines the confidence level of the current frame's real-time position to determine whether to perform facial recognition in the next frame to assist in tracking the target person. This approach balances the speed of position tracking with the accuracy of facial recognition, thus enabling precise and rapid light tracking.

[0013] Furthermore, in step S6, the predicted position of the target person in the next frame from step S3 is matched with the real-time positions of all people in the scene in the next frame from step S4. If the number of consecutive matching failures exceeds a second preset value, the target person is considered lost, and the process returns to step S1. Thus, in step S6, the target person is not immediately considered lost after a matching failure, but rather confirmed multiple times before being considered lost. Complex calculations are then performed to locate the target person again using facial recognition, saving computational power, accelerating the light-tracking speed, and avoiding the slow light-tracking speed caused by occasional occlusion or poor lighting conditions leading to failure to recognize the target person before immediately performing facial recognition.

[0014] Furthermore, an initialization step precedes step S1. In this initialization step, specific facial features are obtained by inputting a photo or manually selecting a target person from the scene. If specific facial features are obtained by inputting a photo, the process jumps to step S1. If specific facial features are obtained by manually selecting a target person from the scene, it is considered that the target person was directly identified in step S1, and the target person's real-time location and motion information are obtained. Then, the process jumps to step S2. Obtaining specific facial features by inputting a photo requires comparing the specific facial features with the facial information of all people in the scene to find the target person. Obtaining specific facial features by manually selecting a target person from the scene is equivalent to directly providing the target person's real-time location and motion information. The comparison process between the specific facial features and the facial information of all people in the scene is completed by the human brain.

[0015] Furthermore, the order of steps S3 and S4 can be freely interchanged. The contents of steps S3 and S4 can be completed independently without interfering with each other, so their order can be changed. However, the jumps between steps S3 and S4 in the entire method also need to be modified accordingly.

[0016] Furthermore, in step S3, when determining the predicted position of the target in the next frame based on the target's real-time position and motion information, a Kalman filter is used. The Kalman filter is a recursive state estimation algorithm that provides the best state estimation result by continuously estimating and updating the system's state. This involves iteratively repeating the process of "making a prediction - updating the predicted value to the optimal value based on the measured value." It is a mature technology suitable for tracking targets with stable motion.

[0017] Furthermore, the real-time position is determined based on the identified person bounding box, and the confidence level of the real-time position is the confidence level of the person bounding box. Any point within the person bounding box can be selected as the real-time position; determining the confidence level of the real-time position is equivalent to determining the confidence level of the person bounding box, thereby confirming whether it is occluded and the reliability of the detection result.

[0018] Furthermore, the "person frame" refers to the "head frame." A smaller head frame allows for fewer poses, resulting in more accurate detection and less susceptibility to jitter.

[0019] Furthermore, in steps S1 and S4, when detecting people, the backbone network is used to extract features from the scene and form a multi-scale feature pyramid. Each feature layer at each scale outputs five branches: person bounding box confidence, person bounding box position, face confidence, face position, and at least three facial key points. Each feature layer corresponds to several anchor points to adapt to different proportions of target people. Then, the person bounding box confidence is filtered, retaining all outputs greater than a third preset value. Next, the person bounding box position is combined with anchor point information to calculate the real-time position of the corresponding person in the scene, and non-maximum suppression is applied to obtain the final real-time positions of all people. Then, the face confidence of all detected people is filtered, retaining all outputs greater than a fourth preset value. Finally, the face position in the scene is calculated using the face position and at least three facial key points combined with anchor point information. This scheme simultaneously detects person bounding boxes and faces, using the person bounding box detection results as real-time positions and face features as ReID information, thus greatly improving the detection speed.

[0020] Furthermore, the camera is mounted on the lamp head of the lamp that emits a beam of light and rotates with the lamp head. This allows the camera to rotate to locate the target person, and the synchronized rotation ensures that once the camera finds the target person, the light from the lamp shines on that person.

[0021] Furthermore, when the light from the lamp tracks the real-time position of the target person, the target person's head is positioned at the center of the camera's field of view. This ensures that no matter which direction the target person moves, they will not disappear from the field of view in a short period of time, facilitating tracking.

[0022] Furthermore, in step S1, as the camera continuously captures the scene, it follows the rotation of the light head to locate the target person. This ensures the camera can find the target person regardless of their position on the stage, making it more intelligent.

[0023] Furthermore, in step S6, when matching the predicted position of the target character in the next frame from step S3 with the real-time positions of all characters in the scene in the next frame from step S4, the Hungarian algorithm is used for matching. The IoU between the predicted position of the target character and the real-time positions of all characters is calculated, and the Hungarian algorithm is used for matching based on the confidence level of the real-time position and the IoU. This technique is mature and provides accurate matching.

[0024] Furthermore, in steps S1 and S5, when identifying a target person with specific facial features, it is necessary to compare the specific facial features with the facial information of all persons, and select the person with the highest similarity that is greater than the fifth preset value as the target person; otherwise, it is considered that no target person has been found. This is to ensure that the probability of the selected person being the target person is maximized and that it basically meets the requirements. Attached Figure Description

[0025] Figure 1 This is a flowchart illustrating the precise and rapid light-tracking method of the present invention.

[0026] Figure 2 This is a schematic diagram of the structure of a system that applies the precise and rapid light-tracking method of this invention.

[0027] Figure 3 This is a schematic diagram of the process for detecting human subjects according to the present invention.

[0028] Figure 4 This is a schematic diagram of the process of facial recognition of a target person according to the present invention. Detailed Implementation

[0029] The accompanying drawings are for illustrative purposes only and should not be construed as limiting this patent. To better illustrate this embodiment, some components in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings. The positional relationships described in the drawings are for illustrative purposes only and should not be construed as limiting this patent.

[0030] like Figures 1 to 2 This invention provides a precise and rapid light-tracking method, comprising the following steps:

[0031] S1. Continuously monitor the scene captured by the camera until a target person with specific facial features is identified, and obtain the real-time location and motion information of the target person.

[0032] S2. The light from the lamp tracks the real-time position of the target person and determines whether the confidence level of the real-time position is less than the first preset value. If so, it is marked that face recognition needs to be performed in the next frame.

[0033] S3. Based on the target person's real-time position and motion information, obtain the predicted position of the target person in the next frame;

[0034] S4. The camera captures the next frame of the scene and obtains the real-time position and face position of all the people in the next frame of the scene;

[0035] S5. When face recognition is required in the next frame, the faces of all people in the scene in the next frame of step S4 are recognized. If the target person is recognized, the real-time position of the target person in this scene is updated and marked as not requiring face recognition in the next frame. Then, the process jumps to step S2. If the target person is not recognized, the process jumps to step S6. If face recognition is marked as not requiring face recognition in the next frame in step S2, the process jumps directly to step S6.

[0036] S6. Match the predicted position of the target person in the next frame in step S3 with the real-time positions of all people in the scene in the next frame in step S4. If the match is successful, the matched person is considered to be the target person. Update the real-time position of the target person in this scene and jump to step S2. If the match is unsuccessful, the predicted position of the target person in step S3 is considered to be the real-time position of the target person in the next frame of the scene, and the motion information remains unchanged. The confidence of the real-time position is considered to be less than the first preset value. Then jump to step S2.

[0037] The precise and rapid light-tracking method first locates the target person in the scene using facial recognition. Then, it determines the confidence level of the target person's real-time position. If the confidence level is lower than a first preset value, it marks the next frame of the scene as requiring facial recognition-assisted tracking of the target person. If facial recognition fails or is not marked, it directly matches the predicted position of the target person in the next frame with the real-time positions of all people in the next frame of the scene. When a matching person is found, it is taken as the target person, and the real-time position and confidence level of the target person are updated. Then, it proceeds to the next frame for target person tracking. When no matching person is found, the predicted position of the target person in the next frame is taken as the real-time position of the target person in the next frame of the scene, and it is assumed that the motion information remains unchanged. Then, it performs facial recognition in the next frame to track the target person, while the light from the lamp tracks the real-time position of the target person.

[0038] This method primarily uses position tracking and predicts the target person's position in the next frame. It also combines the confidence level of the current frame's real-time position to determine whether to perform facial recognition in the next frame to assist in tracking the target person. This approach balances the speed of position tracking with the accuracy of facial recognition, thus enabling precise and rapid light tracking.

[0039] It should be noted that as long as face recognition is marked as required in the next frame in step S2, face recognition will be required in the next frame after jumping to step S2, regardless of whether the confidence level of the target person's real-time position is greater than or equal to the first preset value in subsequent steps, unless the mark indicating that face recognition is required in the next frame is removed in subsequent steps.

[0040] In a preferred embodiment of the present invention, in step S6, the number of consecutive matching failures when matching the predicted position of the target person in the next frame of step S3 with the real-time positions of all people in the scene in the next frame of step S4 is counted. When the number of consecutive matching failures exceeds a second preset value, the target person is considered lost, and the process returns to step S1. Thus, in step S6, the target person is not immediately considered lost after a matching failure, but is only considered lost after multiple confirmations. Complex calculations are then performed to find the target person again through facial recognition, saving computing power, accelerating the light-tracking speed, and avoiding the slow light-tracking speed caused by immediately performing facial recognition before the target person is identified due to occasional obstruction or poor lighting.

[0041] The second preset value can be 2, 3, 4, 5, 6, 7, 8, 9 or 10, preferably 5.

[0042] In a preferred embodiment of the present invention, an initialization step is included before step S1. In the initialization step, specific facial features are obtained by inputting a photo or manually selecting a target person from the scene. When specific facial features are obtained by inputting a photo, the process jumps to step S1. When specific facial features are obtained by manually selecting a target person from the scene, it is considered that the target person has been directly identified in step S1, and the real-time location and motion information of the target person are obtained. Then, the process jumps to step S2. Obtaining specific facial features by inputting a photo requires comparing the specific facial features with the facial information of all people in the scene to find the target person. However, obtaining specific facial features by manually selecting a target person from the scene is equivalent to directly providing the real-time location and motion information of the target person. The process of comparing the specific facial features with the facial information of all people in the scene is completed by the human brain.

[0043] In this application, specific facial features are preferably obtained by inputting a photo, which is then sent to the lighting fixture via a control server.

[0044] In a preferred embodiment of the present invention, the order of steps S3 and S4 can be arbitrarily changed. The contents of steps S3 and S4 can be completed independently without interfering with each other, so their order can be changed. However, the jumps between steps S3 and S4 in the entire method also need to be modified accordingly.

[0045] In a preferred embodiment of the present invention, in step S3, when obtaining the predicted position of the target person in the next frame based on the real-time position and motion information of the target person, Kalman filtering is used. The Kalman filter is a recursive state estimation algorithm that provides the best state estimation result by continuously estimating and updating the system's state. This involves iteratively repeating the process of "making a prediction - updating the predicted value to the optimal value based on the measured value." This technology is mature and suitable for tracking targets with stable motion. In this application, obtaining the predicted position of the target person in the next frame based on the real-time position and motion information of the target person is equivalent to "making a prediction." Updating the real-time position of the target person based on the real-time position of the target person identified by face recognition in the next frame, or by matching the predicted position with the real-time positions of all other people to obtain the corresponding real-time position, or by forcibly assuming that the predicted position of the target person in step S3 is the real-time position of the target person in the next frame of the scene, is equivalent to "updating the predicted value to the optimal value based on the measured value."

[0046] In a preferred embodiment of the present invention, the real-time position is determined based on the identified person frame, and the confidence level of the real-time position is the confidence level of the person frame. Any point within the person frame can be selected as the real-time position; determining the confidence level of the real-time position is equivalent to determining the confidence level of the person frame, thereby confirming whether it is occluded and assessing the reliability of the detection result.

[0047] The character frame can be a body frame or a head frame. In a preferred embodiment of the present invention, the character frame refers to a head frame. Compared to the body frame, the head frame is smaller, has fewer poses, is more accurate in detection, and is less prone to jitter.

[0048] like Figure 3In a preferred embodiment of the present invention, when detecting people in steps S1 and S4, a backbone network is used to extract features from the scene and form a multi-scale feature pyramid. Each feature layer at each scale outputs five branches: person bounding box confidence, person bounding box position, face confidence, face position, and at least three facial key points. Each feature layer corresponds to several anchor points to adapt to different proportions of target people. Then, the person bounding box confidence is filtered, retaining all outputs greater than a third preset value. Next, the real-time position of the corresponding person in the scene is calculated using the person bounding box position combined with anchor point information, and non-maximum suppression is performed to obtain the final real-time positions of all people. Then, the face confidence corresponding to all detected people is filtered, retaining all outputs greater than a fourth preset value. Finally, the face position in the scene is calculated using the face position and at least three facial key points combined with anchor point information. Compared to the prior art, which requires first detecting people in the scene and then detecting faces in all person bounding boxes, this method is significantly more efficient. This solution detects both person bounding boxes and faces simultaneously. It uses the person bounding box detection results as the real-time location and the face features as ReID information, thus greatly improving the detection speed.

[0049] Five branches predict whether feature points have detected a person, person location, face location, face confidence, face location, and at least three facial key points, respectively.

[0050] Preferably, the number of facial key points is 5.

[0051] In this application, the backbone network is a ResNet50 network, and then FPN is used to form a multi-scale feature pyramid.

[0052] After obtaining the real-time location of all people in the scene, the system filters out those whose faces are visible to the camera (some people may be facing away from the camera and their faces may not be visible) and obtains their facial positions. When performing facial recognition, only the corresponding area needs to be recognized.

[0053] The aforementioned person detection model still requires training. The training data includes labels for people, faces, and at least three key points. The model training adopts a similar strategy to RetinaNet. The FocalLoss loss function is used for person bounding box confidence and face confidence, while the smooth-L1 loss function is used for person bounding box position, face position, and at least three key points. It should be noted that when the person bounding box is a head bounding box, the ground truth assignment for both the person bounding box position and the face position is based on the center of the head bounding box.

[0054] In a preferred embodiment of the present invention, the camera is mounted on the lamp head of the lamp for emitting light beams and rotates with the lamp head. This allows the camera to rotate to locate the target person, and the synchronized rotation ensures that once the camera finds the target person, the light from the lamp shines on that person.

[0055] It should be noted that if the camera is mounted on the lamp head of the lamp for emitting light beams and rotates with the lamp head, then when calculating the predicted position of the target person, the displacement of the camera following the lamp head between two frames needs to be considered.

[0056] In a preferred embodiment of the present invention, when the light from the lamp tracks the real-time position of the target person, the target person's head is positioned at the center of the camera's field of view. This ensures that no matter which direction the target person moves, they will not disappear from the field of view in a short time, facilitating tracking.

[0057] Of course, the target person's head can also be appropriately offset relative to the center point of the camera's field of view to ensure that the light beam illuminates the target person's face or feet.

[0058] In a preferred embodiment of the present invention, during step S1, as the camera continuously captures the scene, it follows the rotation of the light head to locate the target person. This ensures that the camera can find the target person regardless of their position on the stage, making it more intelligent.

[0059] In a preferred embodiment of the present invention, in step S6, when matching the predicted position of the target person in the next frame from step S3 with the real-time positions of all people in the scene in the next frame from step S4, the Hungarian algorithm is used for matching. The IoU between the predicted position of the target person and the real-time positions of all people is calculated, and the Hungarian algorithm is used for matching based on the confidence level of the real-time position and the IoU. This technique is mature and provides accurate matching.

[0060] like Figure 4 In a preferred embodiment of the present invention, in steps S1 and S5, when identifying a target person with specific facial features, it is necessary to compare the specific facial features with the facial information of all persons, and select the person with the highest similarity and greater than a fifth preset value as the target person; otherwise, it is considered that no target person has been found. This is to ensure that the selected person has the highest probability of being the target person and basically meets the requirements. If the similarity is less than or equal to the fifth preset value, it is considered that no target person has been identified.

[0061] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A precise and rapid light-tracking method, characterized in that, Includes the following steps: S1. Continuously monitor the scene captured by the camera until a target person with specific facial features is identified, and obtain the real-time location and motion information of the target person. S2. The light from the lamp tracks the real-time position of the target person and determines whether the confidence level of the real-time position is less than the first preset value. If so, it is marked that face recognition needs to be performed in the next frame. S3. Based on the target person's real-time position and motion information, obtain the predicted position of the target person in the next frame; S4. The camera captures the next frame of the scene and obtains the real-time position and face position of all the people in the next frame of the scene; S5. When face recognition is required in the next frame, the faces of all people in the scene in the next frame of step S4 are recognized. If the target person is recognized, the real-time position of the target person in this scene is updated and marked as not requiring face recognition in the next frame. Then, the process jumps to step S2. If the target person is not recognized, the process jumps to step S6. If it is marked in step S2 that face recognition is not needed in the next frame, proceed directly to step S6; S6. Match the predicted position of the target person in the next frame in step S3 with the real-time positions of all people in the scene in the next frame in step S4. If the match is successful, the matched person is considered to be the target person. Update the real-time position of the target person in this scene and jump to step S2. If the match is unsuccessful, the predicted position of the target person in step S3 is considered to be the real-time position of the target person in the next frame of the scene, and the motion information remains unchanged. The confidence of the real-time position is considered to be less than the first preset value. Then jump to step S2. The camera is mounted on the lamp head of the lamp for emitting light beams and rotates with the lamp head; When the light from the lamp tracks the real-time position of the target person, the head of the target person is placed at the center of the camera's field of view. In step S1, as the camera continues to capture images of the scene, it will follow the rotation of the light head to locate the target person.

2. The precise and rapid light-tracking method according to claim 1, characterized in that, In step S6, the number of consecutive matching failures when matching the predicted position of the target person in the next frame of step S3 with the real-time positions of all people in the scene in the next frame of step S4 is counted. When the number of consecutive matching failures exceeds the second preset value, the target person is considered lost and the process returns to step S1.

3. The precise and rapid light-tracking method according to claim 1, characterized in that, Before step S1, there is an initialization step. In the initialization step, specific facial features are obtained by inputting a photo or manually selecting a target person from the scene. When specific facial features are obtained by inputting a photo, the process jumps to step S1. When specific facial features are obtained by manually selecting a target person from the scene, it is considered that the target person has been directly identified in step S1, and the real-time location and motion information of the target person are obtained. Then, the process jumps to step S2.

4. The precise and rapid light-tracking method according to claim 1, characterized in that, The order of steps S3 and S4 can be changed at will.

5. The precise and rapid light-tracking method according to claim 1, characterized in that, In step S3, when the predicted position of the target person in the next frame is obtained based on the real-time position and motion information of the target person, Kalman filtering is used.

6. The precise and rapid light-tracking method according to claim 1, characterized in that, The real-time location is determined based on the identified person frame, and the confidence level of the real-time location is the confidence level of the person frame.

7. The precise and rapid light-tracking method according to claim 6, characterized in that, The "character frame" refers to the frame around a person's head.

8. The precise and rapid light-tracking method according to claim 6 or 7, characterized in that, In steps S1 and S4, when detecting people, the backbone network is used to extract features from the scene and form a multi-scale feature pyramid. Each feature layer at each scale outputs five branches: person bounding box confidence, person bounding box position, face confidence, face position, and at least three facial key points. Each feature layer corresponds to several anchor points to adapt to different proportions of target people. Then, the person bounding box confidence is filtered, and all outputs greater than the third preset value are retained. Then, the person bounding box position is combined with the anchor point information to calculate the real-time position of the corresponding person in the scene, and non-maximum suppression is performed to obtain the final real-time position of all people. Then, the face confidence corresponding to all detected people is filtered, and all outputs greater than the fourth preset value are retained. Finally, the face position in the scene is calculated by combining the face position and at least three facial key points with the anchor point information.

9. The precise and rapid light-tracking method according to claim 1, characterized in that, In step S6, when matching the predicted position of the target person in the next frame of step S3 with the real-time positions of all people in the scene in the next frame of step S4, the Hungarian algorithm is used for matching.

10. The precise and rapid light-tracking method according to claim 1, characterized in that, In steps S1 and S5, when identifying a target person with specific facial features, it is necessary to compare the specific facial features with the facial information of all people, and select the person with the highest similarity and greater than the fifth preset value as the target person; otherwise, it is considered that no target person has been found.

Citation Information

Patent Citations

  • Method and device for finding and tracking pairs of eyes

    CN101861118A

  • Automatic stage light tracking system and method based on limb motions

    CN108198221A