Face coding method, system and equipment based on area tracking and dynamic interpolation and medium
By using a face masking method based on region tracking and dynamic interpolation, and leveraging the YOLO-Face model and detection result queue, the masking area is dynamically adjusted, solving the problem of facial privacy leakage caused by worker operation, hand occlusion, and fabric interference in factory workshops, and realizing real-time monitoring and automated analysis.
Patent Information
- Application Number
- CN202510917691.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies cannot effectively solve the problem of facial privacy leaks caused by workers looking down, covering their faces with their hands, being interfered with by fabric, and moving quickly, especially in video collection and analysis where there is a risk of missed detection.
A face masking method based on region tracking and dynamic interpolation is adopted. By acquiring the video stream of a sewing machine camera, face detection is performed using the YOLO-Face model. The detection result queue is maintained, and the maximum interpolation box is obtained using dynamic interpolation and region tracking strategies for masking.
It effectively avoids missing codes when the face detection model fails to detect, dynamically adjusts the code-blocking area, and solves the problem of face privacy leakage caused by worker operation, hand covering and cloth interference, realizing real-time monitoring and automated analysis of the factory workshop.
Smart Images

Figure CN120746818A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image processing technology and relates to a face coding method, system, device and medium based on region tracking and dynamic interpolation. Background Art
[0002] As a vital global livelihood industry, the apparel manufacturing industry boasts a massive market and employs a large number of people. In modern apparel factories, industrial engineering (IE) engineers often need to conduct in-depth analysis of sewing machine operators' operations to improve efficiency, optimize processes, and achieve refined cost management. Engineers film videos of each worker performing a specific process on-site and use these videos to perform process breakdown and motion studies, provide operational guidance and standardization, and measure standard time (ST). However, current video acquisition and analysis methods have significant drawbacks: Manually filming videos of each worker and each process is time-consuming and labor-intensive, making it difficult to meet the continuous monitoring and analysis requirements of large-scale production environments. This workload is high and inefficient. Furthermore, manual filming relies heavily on manual intervention and cannot meet the factory's needs for real-time monitoring, immediate feedback, and automated analysis.
[0003] To address the bottleneck of video capture, integrating intelligent components with automatic video capture capabilities into sewing equipment has become an inevitable trend in the industry. These components can automatically and continuously record video during worker operations, significantly freeing up IE engineers' productivity and laying the foundation for automated collection, storage, and analysis of production data. However, due to employee privacy concerns, especially in overseas markets, facial blurring is required in recorded videos.
[0004] With the rise of AI in recent years, using AI technology for face coding has become a more reliable solution. Directly using a target detection frame for face coding can easily cause the frame to shrink due to target movement or detection frame jitter, exposing the edges of the face. Furthermore, the target detection model is prone to missed detections in complex backgrounds, poor lighting, partial face occlusion (such as by the machine or hand when looking down), rapid movement, or at specific angles. Once a missed detection occurs, the face in that frame is completely exposed, posing a serious risk of privacy leakage. This missed detection problem is an inherent limitation that cannot be solved by simply improving the stability of the detection frame. Summary of the Invention
[0005] The present application provides a face coding method, system, device and medium based on area tracking and dynamic interpolation, which is used to solve the problem of facial privacy leakage caused by missing codes due to reasons such as workers lowering their heads, hand occlusion, fabric interference and rapid movement, which cannot be solved in the existing technology.
[0006] In the first aspect, the present application provides a face coding method based on region tracking and dynamic interpolation, the method comprising: obtaining a video stream captured by a camera arranged on a sewing machine; performing face detection based on the video stream and a target detection model to obtain a face detection result; the face detection result includes a face detection frame of a frame image in the video stream; maintaining a detection result queue with a set number of frame images based on the face detection result; the detection result queue is used to store the image frames and face detection results of the latest set number of frame images; processing is performed based on the detection result queue, dynamic interpolation strategy and region tracking strategy to obtain the maximum interpolation frame of each frame image; and coding each face in the detection result queue based on the maximum interpolation frame of each face.
[0007] In an implementation of the first aspect, the detection result queue is implemented using a ring buffer, and the oldest frame face detection result is automatically eliminated each time a new frame is written.
[0008] In an implementation of the first aspect, processing is performed based on the detection result queue, dynamic interpolation strategy, and area tracking strategy to obtain the maximum interpolation frame of each frame image, including: performing face matching based on the detection result queue and dynamic interpolation strategy to obtain matching results; and obtaining the maximum interpolation frame of each face based on the matching results.
[0009] In an implementation of the first aspect, face matching is performed based on the detection result queue and the dynamic interpolation strategy, and obtaining the matching result includes: performing pairwise IOU calculations on the detection frames in different indexes based on the detection result queue; if the IOU value of the current detection frame and the comparison detection frame exceeds a preset threshold, the matching result is determined to be that the two correspond to the same face; if the IOU value of the current detection frame and the comparison detection frame does not exceed the preset threshold, the matching result is determined to be that the current detection frame is a newly added face.
[0010] In an implementation of the first aspect, obtaining a maximum interpolation frame for each frame image based on the matching results includes: obtaining a detection frame for each face based on the matching results; and generating a maximum interpolation frame covering all detection frames of each face based on the detection frame of each face.
[0011] In an implementation of the first aspect, the method further includes: for a frame image after a set frame image in a video stream, dividing the current frame image into regions and obtaining divided region information; the divided region information includes key regions and non-key regions; performing face tracking processing based on the key regions, the region tracking strategy and the maximum interpolation frame of each face in the detection result queue, judging whether there is face missing in the key region, and obtaining the maximum interpolation frame of the key region based on the judgment result; performing dynamic interpolation based on the non-key region and the dynamic interpolation strategy to obtain the maximum interpolation frame of the non-key region.
[0012] In an implementation of the first aspect, determining whether a face is missing in the key area, and obtaining a maximum interpolation frame of the key area based on the determination result includes: performing detection frame tracking using a Kalman filter based on the maximum interpolation frame of the key area and each face in the detection result queue to obtain a prediction frame of the key area; performing detection frame matching using a Hungarian matching algorithm based on the prediction frame of the key area to obtain a face loss result of the key area; if the face loss result of the key area is that the face is not lost, obtaining the maximum interpolation frame of the key area based on the detection frame of the key area and the prediction frame of the key area; if the face loss result of the key area is that the face is lost, dynamically interpolating the key area based on the dynamic interpolation strategy to obtain a maximum interpolation frame of the newly added face in the key area.
[0013] In the second aspect, the present application provides a face coding system based on region tracking and dynamic interpolation, the system comprising: an acquisition module, configured to acquire a video stream captured by a camera arranged on a sewing machine; a face detection module, configured to perform face detection based on the video stream and a target detection model, and obtain a face detection result; the face detection result includes a face detection frame of a frame image in the video stream; a queue maintenance module, configured to maintain a detection result queue with a set number of frame images based on the face detection result; the detection result queue is used to store the image frames and face detection results of the latest set number of frame images; a processing module, configured to perform processing based on the detection result queue, dynamic interpolation strategy and region tracking strategy, and obtain the maximum interpolation frame of each face in the detection result queue; and a coding module, configured to code each face in the detection result queue based on the maximum interpolation frame of each face.
[0014] In a third aspect, the present application provides an electronic device, comprising: a memory configured to store a computer program; and a processor configured to run the computer program to implement the face coding method based on region tracking and dynamic interpolation as described above.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned face coding method based on region tracking and dynamic interpolation.
[0016] As described above, the face coding method, system, device, and medium based on region tracking and dynamic interpolation described in this application have the following beneficial effects:
[0017] The present application obtains a video stream captured by a camera installed on a sewing machine; performs face detection based on the video stream and a target detection model to obtain a face detection result; the face detection result includes a face detection frame of a frame image in the video stream; maintains a detection result queue with a set number of frame images based on the face detection result; the detection result queue is used to store the image frames and face detection results of the latest set number of frame images; processes based on the detection result queue, dynamic interpolation strategy and regional tracking strategy to obtain the maximum interpolation frame of each face in the detection result queue; and codes each face in the detection result queue based on the maximum interpolation frame of each face. The present application utilizes a dynamic interpolation strategy and a regional tracking strategy to immediately trigger interpolation when the face detection model fails to detect, thereby avoiding missing codes, and can dynamically adjust the coding area based on interpolation, which can effectively solve the problem of facial privacy leakage in sewing factory workshops caused by workers lowering their heads, hand occlusion, fabric interference and rapid movement. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Shown is a hardware scenario diagram of the face coding method based on region tracking and dynamic interpolation described in an embodiment of the present application.
[0019] Figure 2 Shown is a flowchart of the face coding method based on region tracking and dynamic interpolation described in an embodiment of the present application.
[0020] Figure 3 Shown is a hardware schematic diagram of the camera setup described in an embodiment of the present application.
[0021] Figure 4 Shown is a schematic diagram of the structure of the YOLO-Face model described in an embodiment of the present application.
[0022] Figure 5 Shown is a flowchart of region tracking and dynamic interpolation according to an embodiment of the present application.
[0023] Figure 6 Shown is a schematic diagram of image area division according to an embodiment of the present application.
[0024] Figure 7 Shown is a structural diagram of the face coding system based on region tracking and dynamic interpolation described in an embodiment of the present application.
[0025] Figure 8 Shown is a structural schematic diagram of an electronic device described in an embodiment of the present application.
[0026] Component number description
[0027] 200 300 electronic devices based on area tracking and dynamic interpolation
[0028] Face coding system
[0029] 201 Acquisition Module 301 Memory
[0030] 202 face detection module 302 processor
[0031] 203 Queue Maintenance Module 303 Display
[0032] 204 Processing modules S1 to S5 Step 205 Coding module DETAILED DESCRIPTION
[0033] The following describes the embodiments of the present application through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.
[0034] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application. Therefore, the illustrations only show components related to the present application and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0035] The following embodiments of the present application provide a method, system, device, and medium for face coding based on area tracking and dynamic interpolation, which solves the problem of facial privacy leakage caused by missing codes due to reasons such as workers lowering their heads, hand occlusion, fabric interference, and rapid movement, which cannot be solved in the prior art.
[0036] The face coding method based on region tracking and dynamic interpolation provided in the embodiment of the present application can be run on mobile terminals, computer terminals and other similar devices. Taking running on the mobile terminal as an example, Figure 1 is a hardware structure diagram of the mobile terminal, such as Figure 1 As shown, the mobile terminal may include: a processor and a memory, the processor may be a central processing unit, and the memory is used to store data. Figure 1 The mobile terminal in the figure is only used as an example and does not limit the specific structure of the mobile terminal.
[0037] Optionally, the mobile terminal may further include: a communication transmission device and an input / output device.
[0038] Optionally, the memory can be used to store computer programs, such as software programs and modules of application software. The memory may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the mobile terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0039] Optionally, the communication transmission device can be used to receive or send data via a network, which may include a wireless network provided by the communication provider of the mobile terminal. The communication transmission device may include a NIC (Network Interface Controller) that can be connected to other network devices through a base station so as to communicate with the Internet.
[0040] The technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings in the embodiments of the present application.
[0041] like Figure 2 As shown, this embodiment provides a face coding method based on region tracking and dynamic interpolation, including the following steps S1 to S5.
[0042] Step S1: Obtain a video stream captured by a camera installed on a sewing machine.
[0043] In some embodiments, a video stream captured by a camera provided on a sewing machine is obtained, wherein the camera is configured to Figure 3 As shown, it is installed on the column of the sewing machine and parallel to the table for video shooting.
[0044] Step S2: Perform face detection based on the video stream and the target detection model to obtain a face detection result; the face detection result includes a face detection frame of the frame image in the video stream. The target detection model is a YOLO-Face model.
[0045] In some embodiments, this application uses OpenCV to extract image frames from a video stream and input them into a face detection model. The first step in face deblurring requires considering the accuracy of the model's face detection. The face detection model used is the YOLO-Face model, which is based on an improved YOLOv8.
[0046] Figure 4 Shown is a schematic diagram of the structure of the YOLO-Face model described in the embodiment of this application. Figure 4As shown, the YOLO-Face model consists of a backbone network, a feature pyramid, and a multi-task detection head. The backbone network uses GhostNet-C3 to replace the original CSPDarknet53, so that the model can significantly reduce the amount of calculation while maintaining accuracy, making it more suitable for mobile devices; the EMA (Efficient Mutil-Attention, EMA) module is embedded between the feature pyramid (FPN) layer and the backbone network to enhance the feature response of small target faces by effectively fusing multi-scale features, thereby improving the model's detection rate for small target faces; on the detection head, a multi-task head design is added, such as adding 5-point facial key point regression to the detection head (for example, left / right eyes, nose tip, left / right corners of the mouth), and improving the robustness of detection under occlusion by constraining the detection frame through the geometric position of the key points. In terms of loss function, the newly added key point loss adopts the improved Smooth L1 loss L keypoints (as in formula (1)) and the geometric consistency constraint loss L geo (as shown in formula (2)), where the improved Smooth L1 loss is mainly used to enhance the robustness of the model to occluded key points and avoid training instability caused by excessive errors in the occluded area. The geometric consistency constraint loss emphasizes the geometric consistency between the key points and the detection box. Their calculation formulas are:
[0047]
[0048] in, represents the coordinates (x, y) of the i-th key point predicted by the model; k i Represents the coordinates of the i-th key point of the ground truth; (x box ,y box ) represents the coordinate of the upper left corner of the detection box, (w box , b box ) represents the width and height of the detection box, N represents the number of key points involved in the loss calculation, and γ represents the occlusion perception weight factor. γ = 0.5, by reducing the weight of visible key points, forces the model to pay more attention to the occluded area. During the training process, the joint loss function is formula (3):
[0049] L total =αL CIoU +βL keypoints +τL geo (α=0.7, β=0.2, τ=0.1) Formula (3)
[0050] Among them, α, β, and τ represent hyperparameters used to balance the various loss terms. CIoUThe CIOU loss function represents the detection box. The output of the detection head is 15-dimensional, including 5-dimensional detection box parameters (upper left corner coordinates (x, y), + width and height (w, h) + confidence) and 10-dimensional key point parameters (including 5 groups of key points, each key point has two coordinates of x and y).
[0051] In some embodiments, the YOLO-Face model training is first pre-trained using the public dataset WiderFace to obtain a face detection pre-training model. In the dataset, sewing cloth occlusion blocks and hand occlusion maps are randomly added to simulate the face occlusion in the workshop. In addition, random lighting changes and workshop noise (Gaussian noise + salt and pepper noise) are applied to enhance the environmental adaptability of the model. Based on the pre-trained model, Figure 3 The sewing machine shown in the figure is equipped with a camera to collect workers’ sewing videos, create a data set and use the pre-trained model for training to obtain the final face detection model.
[0052] Step S3: maintaining a detection result queue with a set number of frame images based on the face detection result; the detection result queue is used to store the image frames and face detection results of the latest set number of frame images;
[0053] In one embodiment of the present application, the detection result queue is implemented using a ring buffer, and the oldest frame face detection result is automatically eliminated each time a new frame is written.
[0054] In some embodiments, the present application first maintains a detection result queue with a set number of frame images (for example, K frames); the detection result queue is used to store the image frames and face detection results (detection frames) of the latest set number of frame images. track ={B1,B2,...,B k}, where B k The image and detection box of the kth frame are included. At this time, there are three cases: the first case is that the queue length is less than k, the second case is that the queue length is equal to k for the first time, and the rest are the third cases.
[0055] Among them, in the first case, no output is performed, in the second case, interpolation is performed and all pictures in the queue are output in sequence at once, but the queue only deletes the element with index 0, and in the third case, the picture with index k-1 in the queue is output and the element with index 0 is deleted, so that after each picture enters the queue, its previous k-1 pictures are in the queue. This application does not output when the queue length is less than k. When the queue reaches k frames for the first time, the images of these k frames are interpolated and all pictures in the queue are output in sequence at once, and the element with index 0 in the queue is deleted to ensure that when the next frame of image arrives, there are k-1 frames of image in the queue before it.
[0056] It should be noted that the length k of the detection result queue indicates that the first k frames of images enter the dynamic interpolation module for interpolation processing, and the images after k frames enter the region division and region tracking processing. The k value can be flexibly determined by the user, and this application is not limited to this.
[0057] Step S4: Processing is performed based on the detection result queue, the dynamic interpolation strategy, and the region tracking strategy to obtain the maximum interpolation frame of each face in the detection result queue.
[0058] In one embodiment of the present application, based on the detection result queue, dynamic interpolation strategy and region tracking strategy, obtaining the maximum interpolation frame of each frame image includes the following steps S41 to S42.
[0059] Step S41: Perform face matching based on the detection result queue and the dynamic interpolation strategy to obtain a matching result.
[0060] Step S42: Obtain the maximum interpolation frame of each face based on the matching results.
[0061] Specifically, for the second case in the above embodiment, the length of the detection result queue is equal to k frames for the first time, and the images and detection frames of these k frames are interpolated according to the dynamic interpolation strategy to obtain the maximum interpolation frame of each face, and the corresponding face of each frame image is coded according to the maximum interpolation frame of each face, and all images in the queue are output in sequence at one time, and the element with index 0 in the queue is deleted (see Figure 5 ).
[0062] In one embodiment of the present application, face matching is performed based on the detection result queue and the dynamic interpolation strategy, and obtaining the matching result includes the following steps S411 to S414.
[0063] Step S411: perform pairwise IOU calculations on detection boxes in different indexes based on the detection result queue.
[0064] Step S412: If the IOU value between the current detection frame and the comparison detection frame exceeds a preset threshold, the matching result is determined to be that the two correspond to the same face.
[0065] Step S413: If the IOU value between the current detection frame and the comparison detection frame does not exceed the preset threshold, the matching result is determined to be that the current detection frame is a newly added face.
[0066] IOU (Intersection over Union) is an indicator used in the field of computer vision to measure the degree of overlap between two bounding boxes (such as a predicted box and a true box). Its core idea is to evaluate the accuracy of the prediction by calculating the ratio of the intersection to the union.
[0067] In one embodiment of the present application, obtaining the maximum interpolation frame of each frame image based on the matching result includes the following steps S421 to S422.
[0068] Step S421: Obtain a detection frame for each face based on the matching result.
[0069] Step S422: Generate a maximum interpolation frame covering all detection frames of each face based on the detection frame of each face.
[0070] Step S421: If the matching result indicates that the two detection frames correspond to the same face, a maximum interpolation frame of the current face is generated based on the two detection frames;
[0071] Step S422: If the matching result is determined to be a new face, a maximum interpolation frame covering all detection frames of each independent face is generated.
[0072] Specifically, for a detection result queue including K frames, it is necessary to perform IOU matching on the detection frames of different indexes in the queue to determine whether the targets detected by the detection frames in different indexes are the same face.
[0073] Specifically, for the detection result queue including K frames, first perform pairwise IoU matching on the detection frames in different indexes. If the IoU exceeds the set threshold, it means that the targets of the two detection frames are the same face. Otherwise, it means that a new face has been added. The maximum number of faces is increased by 1 to obtain the detection frame of each face. Based on the detection frame of each face, the maximum enclosing rectangular frame B of each face is generated. max =(min x1,min y1,max x2,max y2), update the detection frame of all images in the index list to the corresponding maximum bounding rectangle (i.e., the maximum interpolation frame); the third case is the same as the second case, and the maximum bounding rectangle is dynamically adjusted after the frame image element output and input after the frame image is set in the video stream. Among them, the coordinates (x 1, y1) represents the coordinates of the upper left corner of the detection frame of each face, and the coordinates (x 2, y2) represents the coordinates of the lower right corner of the detection box of each face.
[0074] In one embodiment of the present application, the method further includes the following steps S441 to S443.
[0075] Step S441 : For a frame image after a set frame image in a video stream, the current frame image is divided into regions to obtain divided region information; the divided region information includes a key region and a non-key region.
[0076] Step S442: Perform face tracking based on the key area, the area tracking strategy, and the maximum interpolation frame of each face in the detection result queue, determine whether there is a face missing in the key area, and obtain the maximum interpolation frame of the key area based on the determination result.
[0077] Step S443: Perform dynamic interpolation based on the non-critical area and the dynamic interpolation strategy to obtain a maximum interpolation frame of the non-critical area.
[0078] In one embodiment of the present application, determining whether a face is missing in the key area and obtaining a maximum interpolation frame of the key area according to the determination result includes the following steps S4421 to S4424.
[0079] Step S4421: Based on the key area and the maximum interpolation frame of each face in the detection result queue, Kalman filtering is used to track the detection frame to obtain the predicted frame of the key area.
[0080] Step S4422: Perform detection frame matching based on the predicted frame of the key area using the Hungarian matching algorithm to obtain a face loss result of the key area.
[0081] Step S4423: If the face loss result of the key area is that the face is not lost, obtaining the maximum interpolation frame of the key area according to the detection frame of the key area and the prediction frame of the key area.
[0082] Step S4424: If the face loss result of the key area is face loss, dynamically interpolate the key area based on the dynamic interpolation strategy to obtain a maximum interpolation frame of the newly added face in the key area.
[0083] In some embodiments, for the frame image after the set frame image in the video stream, after obtaining the face detection result, the region division module is first entered, such as Figure 5 As shown. Since the camera is fixed, we can prioritize the areas first. The right half of the picture has no analytical value for IE, so we can directly mosaic it with the coordinates determined. The rest of the picture is divided into dynamic background area A and sewing machine operation area B, as shown. Figure 6 As shown in the figure, area A is the key area for face coding in this application, which requires face tracking and interpolation, while area B is a secondary area, which only requires interpolation. The center point of the detection frame is used to determine which area the detection frame is located in.
[0084] After region division, the detection frames after K frames of the key region enter the region tracking module, while the detection frames before K-1 frames of the key region and the detection frames of non-key regions enter the dynamic interpolation module directly. The region tracking module takes the detection frames and the current frame timestamp t0 as input. It uses Kalman filtering for detection frame tracking and the Hungarian matching algorithm for detection frame matching. If the face ID remains unchanged or lost, the censored region is directly output and the tracking queue is updated. Otherwise, the dynamic interpolation module is triggered for interpolation processing.
[0085] Step S5: Encode the frame images in the detection result queue based on the maximum interpolation frame.
[0086] In one embodiment, the present application continuously de-codes the first frame in the queue using the largest interpolated frame corresponding to the face and outputs it. For example, the pixels in the detection frame area to be de-coded are shuffled to complete the face de-coding. The face de-coding methods used in this application include, but are not limited to, Gaussian blur or pixelation. This application is not limited to the face de-coding methods described in the embodiment.
[0087] The face coding method based on area tracking and dynamic interpolation described in the above embodiments of the present application can immediately trigger interpolation when the face detection model fails to detect, thereby avoiding missing codes, and can dynamically adjust the coding area based on interpolation, which can effectively solve the problem of facial privacy leakage caused by occlusion and rapid movement in sewing factory workshops.
[0088] The scope of protection of the face coding method based on area tracking and dynamic interpolation described in the embodiment of the present application is not limited to the order of execution of the steps listed in this embodiment. All solutions implemented by adding, subtracting, or replacing steps in the prior art based on the principles of the present application are included in the scope of protection of the present application.
[0089] An embodiment of the present application also provides a face coding system based on region tracking and dynamic interpolation. The face coding system based on region tracking and dynamic interpolation can implement the face coding method based on region tracking and dynamic interpolation described in the present application. However, the implementation device of the face coding method based on region tracking and dynamic interpolation described in the present application includes but is not limited to the structure of the face coding system based on region tracking and dynamic interpolation listed in this embodiment. All structural variations and replacements of the prior art made according to the principles of the present application are included in the scope of protection of the present application.
[0090] like Figure 7 As shown, this embodiment provides a face coding system based on region tracking and dynamic interpolation. The system 200 includes: an acquisition module 201, a face detection module 202, a queue maintenance module 203, a processing module 204 and a coding module 205.
[0091] The acquisition module 201 is configured to acquire a video stream captured by a camera installed on the sewing machine.
[0092] The face detection module 202 is configured to perform face detection based on the video stream and the target detection model to obtain a face detection result; the face detection result includes a face detection frame of a frame image in the video stream.
[0093] The queue maintenance module 203 is configured to maintain a detection result queue with a set number of frame images based on the face detection result; the detection result queue is used to store the image frames and face detection results of the latest set number of frame images.
[0094] The processing module 204 is configured to perform processing based on the detection result queue, the dynamic interpolation strategy and the region tracking strategy to obtain the maximum interpolation frame of each face in the detection result queue.
[0095] The coding module 205 is configured to code each face in the detection result queue based on the maximum interpolation frame of each face.
[0096] It should be noted that the functions or operations of the acquisition module 201, face detection module 202, queue maintenance module 203, processing module 204 and coding module 205 described in the embodiment of the present disclosure correspond one-to-one to the steps in the sewing action recognition described above, so they will not be repeated here.
[0097] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices or methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules / units is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules or units, which can be electrical, mechanical or other forms.
[0098] The modules / units described as separate components may or may not be physically separate, and the components displayed as modules / units may or may not be physical modules, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules / units may be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, the functional modules / units in the various embodiments of the present application may be integrated into a processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into a single module / unit.
[0099] Those skilled in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0100] like Figure 8 As shown, this embodiment provides an electronic device, and the electronic device 300 includes a memory 301 and a processor 302.
[0101] The memory 301 is used to store computer programs; preferably, the memory 301 includes: ROM, RAM, disk, USB flash drive, memory card or optical disk, etc., various media that can store program codes.
[0102] Specifically, the memory 301 may include a computer system readable medium in the form of a volatile memory, such as a random access memory (RAM) and / or a cache memory. The electronic device 300 may further include other removable / non-removable, volatile / non-volatile computer system storage media. The memory 301 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the various embodiments of the present application.
[0103] The processor 302 is connected to the memory 301 and is used to execute the computer program stored in the memory 301 so that the electronic device 300 executes the face coding method based on area tracking and dynamic interpolation provided in any embodiment of the present application.
[0104] Optionally, the processor 302 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0105] Optionally, the electronic device 300 in this embodiment may further include a display 303. The display 303 is communicatively connected to the memory 301 and the processor 302, and is configured to display a GUI interface related to the face coding method based on region tracking and dynamic interpolation.
[0106] The embodiment of the present application also provides a computer-readable storage medium. Those skilled in the art will understand that all or part of the steps in the method for implementing the above embodiment can be completed by instructing the processor through a program, and the program can be stored in a computer-readable storage medium, and the storage medium is a non-transitory medium, such as a random access memory, a read-only memory, a flash memory, a hard disk, a solid-state drive, a magnetic tape, a floppy disk, an optical disc, and any combination thereof. The above storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a digital video disc (DVD)), or a semiconductor medium (for example, a solid-state drive (SSD)), etc.
[0107] The embodiment of the present application may also provide a computer program product, the computer program product including one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the process or function described in the embodiment of the present application is generated in whole or in part. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer or data center to another website, computer or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method.
[0108] When the computer program product is executed by a computer, the computer executes the method described in the above method embodiment. The computer program product can be a software installation package. When the above method is needed, the computer program product can be downloaded and executed on the computer.
[0109] In summary, the face coding method, system, device, and medium based on region tracking and dynamic interpolation described in this application have the following beneficial effects:
[0110] The present application obtains a video stream captured by a camera installed on a sewing machine; performs face detection based on the video stream and a target detection model to obtain a face detection result; the face detection result includes a face detection frame of a frame image in the video stream; maintains a detection result queue with a set number of frame images based on the face detection result; the detection result queue is used to store the image frames and face detection results of the latest set number of frame images; processes based on the detection result queue, dynamic interpolation strategy and regional tracking strategy to obtain the maximum interpolation frame of each face in the detection result queue; and codes each face in the detection result queue based on the maximum interpolation frame of each face. The present application utilizes a dynamic interpolation strategy and a regional tracking strategy to immediately trigger interpolation when the face detection model fails to detect, thereby avoiding missing codes, and can dynamically adjust the coding area based on interpolation, which can effectively solve the problem of facial privacy leakage in sewing factory workshops caused by workers lowering their heads, hand occlusion, fabric interference and rapid movement.
[0111] The descriptions of the processes or structures corresponding to the above figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.
[0112] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.
Claims
1. A face coding method based on region tracking and dynamic interpolation, characterized in that: include: Obtain a video stream captured by a camera installed on the sewing machine; Performing face detection based on the video stream and the target detection model to obtain a face detection result; The face detection result includes a face detection frame of a frame image in the video stream; Maintaining a detection result queue with a set number of frame images based on the face detection result; the detection result queue is used to store the image frames and face detection results of the latest set number of frame images; Processing based on the detection result queue, dynamic interpolation strategy and region tracking strategy to obtain the maximum interpolation frame of each face in the detection result queue; Each face in the detection result queue is coded based on the maximum interpolation frame of each face.
2. The face coding method based on region tracking and dynamic interpolation according to claim 1, characterized in that: The detection result queue is implemented using a ring buffer, and the oldest frame face detection result is automatically eliminated each time a new frame is written.
3. The face coding method based on region tracking and dynamic interpolation according to claim 1, characterized in that: Based on the detection result queue, dynamic interpolation strategy and region tracking strategy, obtaining the maximum interpolation frame of each frame image includes: Performing face matching based on the detection result queue and the dynamic interpolation strategy to obtain a matching result; Based on the matching results, the maximum interpolation frame of each face is obtained.
4. The face coding method based on region tracking and dynamic interpolation according to claim 3 is characterized in that: Performing face matching based on the detection result queue and the dynamic interpolation strategy, obtaining the matching result includes: Based on the detection result queue, IOU calculations are performed on the detection boxes in different indexes. If the IOU value between the current detection frame and the comparison detection frame exceeds the preset threshold, the matching result is determined to be that the two correspond to the same face; If the IOU value between the current detection frame and the comparison detection frame does not exceed the preset threshold, the matching result is determined to be the current detection frame is a new face.
5. The face coding method based on region tracking and dynamic interpolation according to claim 3 is characterized in that: The maximum interpolation frame of each frame image is obtained based on the matching results: Obtaining a detection frame for each face based on the matching results; Based on the detection frame of each face, a maximum interpolation frame covering all its detection frames is generated.
6. The face coding method based on region tracking and dynamic interpolation according to claim 1, characterized in that: The method further comprises: For a frame image after a set frame image in a video stream, dividing the current frame image into regions and obtaining divided region information; the divided region information includes a key region and a non-key region; Performing face tracking based on the key area, the area tracking strategy, and the maximum interpolation frame of each face in the detection result queue, determining whether there is a face missing in the key area, and obtaining the maximum interpolation frame of the key area based on the determination result; Dynamic interpolation is performed based on the non-critical area and a dynamic interpolation strategy to obtain a maximum interpolation frame of the non-critical area.
7. The face coding method based on region tracking and dynamic interpolation according to claim 6, characterized in that: Determining whether there is a face missing in the key area and obtaining a maximum interpolation frame of the key area according to the determination result includes: Based on the key area and the maximum interpolated frame of each face in the detection result queue, a detection frame is tracked using Kalman filtering to obtain a predicted frame of the key area; Performing detection frame matching based on the prediction frame of the key area using the Hungarian matching algorithm to obtain a face loss result of the key area; If the face loss result of the key area is that the face is not lost, obtaining a maximum interpolation frame of the key area according to the detection frame of the key area and the prediction frame of the key area; If the face loss result of the key area is face loss, the key area is dynamically interpolated based on the dynamic interpolation strategy to obtain a maximum interpolation frame of the newly added face in the key area.
8. A face coding system based on region tracking and dynamic interpolation, characterized in that: The system comprises: an acquisition module configured to acquire a video stream captured by a camera provided on the sewing machine; A face detection module is configured to perform face detection based on the video stream and the target detection model to obtain a face detection result; the face detection result includes a face detection frame of a frame image in the video stream; A queue maintenance module is configured to maintain a detection result queue having a set number of frame images based on the face detection result; the detection result queue is used to store the image frames and face detection results of the latest set number of frame images; a processing module configured to process based on the detection result queue, the dynamic interpolation strategy, and the region tracking strategy to obtain a maximum interpolation frame for each face in the detection result queue; The coding module is configured to code each face in the detection result queue based on the maximum interpolation frame of each face.
9. An electronic device, characterized in that: include: a memory configured to store a computer program; A processor is configured to run the computer program to implement the face coding method based on area tracking and dynamic interpolation according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the face coding method based on region tracking and dynamic interpolation according to any one of claims 1 to 7 is implemented.