Dynamic face tracking method and device, storage medium and program product
By creating face layers in the video stream and drawing recognition boxes, updating and superimposing them on the video stream in real time, the problems of slow video stream network transmission, low face recognition efficiency and late rendering feedback are solved, and efficient face dynamic tracking is achieved.
Patent Information
- Application Number
- CN202510256857.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, the problems of slow network transmission of video streams, low facial recognition efficiency, poor real-time performance of rendering feedback and high latency are particularly obvious when the network conditions are poor or the data transmission volume is large.
Collect the source video stream and transmit it to the server through the network, perform facial recognition application analysis, obtain face area coordinates and data, create face layers and draw recognition boxes, superimpose real-time updated layers on the video stream, reduce the amount of data transmitted on the network and improve the efficiency of dynamic face tracking.
Through efficient layer overlay and recognition box drawing technology, the amount of data transmitted on the network is reduced, the efficiency of face dynamic tracking and the real-time rendering feedback are improved, and the delay problem is solved.
Smart Images

Figure CN120236309A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly to a method, device, storage medium, and program product for dynamic face tracking. Background Art
[0002] With the development of artificial intelligence and deep learning technologies, face recognition technology has gradually shifted to deep learning-based methods. Since then, significant breakthroughs have been made in the field of face recognition using deep learning technologies, making face recognition technology a widely used recognition technology in various fields; thus enabling face recognition based on video streams captured by VR glasses cameras to achieve dynamic face tracking.
[0003] At the same time, there are also some problems at present. Video streams need to be transmitted over the network to the server for processing, and network transmission itself introduces a certain delay. This includes time delays in data transmission and increased delays caused by factors such as network congestion. When the network conditions are poor or the data transmission volume is large, this delay becomes more obvious; face detection algorithms are usually based on machine learning and deep learning technologies and need to process and analyze each frame in the video stream. Due to the complexity of face detection algorithms and the increase in computational requirements, the processing time for face detection increases, thereby introducing delays. If the algorithm is not optimized enough or there is insufficient hardware acceleration support, this delay will be even more significant. Therefore, there is a need for further improvement and development in the problems of slow network transmission of video streams, low face recognition efficiency, poor real-time rendering feedback, and high latency.
[0004] The above content is only used to assist in understanding the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main objective of this application is to provide a method for dynamic face tracking, aiming to solve the technical problems of slow network transmission of video streams, low face recognition efficiency, poor real-time rendering feedback, and high latency.
[0006] To achieve the above objective, this application proposes a method for dynamic face tracking, and the method includes:
[0007] Collect the source video stream and transmit the source video stream over the network to the server to obtain the first processed data;
[0008] Obtain the first processed data after processing from the server, and parse the first processed data through a face recognition application to obtain the second processed data;
[0009] Perform face detection on the second processed data to obtain the face region coordinates and face data of the facial image;
[0010] Create a face layer based on the coordinates of the face region and draw an identification box in the face layer;
[0011] Display the corresponding face data in the face layer, and superimpose and render the face layer with real-time updates on the portrait layer of the source video stream at the current moment.
[0012] In one embodiment, before the step of collecting the source video stream and transmitting the source video stream to the server through the network to obtain the first processed data, it includes:
[0013] Copy the source video stream to obtain a target video stream, and directly output the target video stream to the user terminal for display.
[0014] In one embodiment, the step of obtaining the first processed data from the server and parsing the first processed data through a face recognition application to obtain the second processed data includes:
[0015] Receive the first processed data processed by the server through a network transmission method;
[0016] Perform decoding processing on the first processed data;
[0017] Extract the decoded data of the first processed data frame by frame through a face recognition application to obtain video frames.
[0018] In one embodiment, the step of performing face detection on the second processed data to obtain the face region coordinates and face data of the facial image includes:
[0019] Receive the video frames generated by the face recognition application as the second processed data;
[0020] Load the face recognition module to recognize the facial image of the second processed data;
[0021] Match the facial image with the face database in the face recognition module;
[0022] Obtain the face region coordinates and face data of the facial image according to the data matching result.
[0023] In one embodiment, the step of creating a face layer based on the face region coordinates and drawing an identification box in the face layer includes:
[0024] Obtain the face region coordinates to determine the position and size of the face in the image;
[0025] Create a new transparent face layer according to the position and size;
[0026] Draw a rectangular box around the entire face region as the identification box according to the face region coordinates.
[0027] In one embodiment, the step of displaying the corresponding face data in the face layer includes:
[0028] Inside or near the drawn recognition frame, adjust the face state by adjusting the line thickness, color of the recognition frame, and the position and transparency of the data label, and display the corresponding face data.
[0029] In one embodiment, the step of superimposing and rendering the face layer with real-time updates on the portrait layer of the source video stream at the current moment includes:
[0030] Real-time update the face region coordinates and face data according to the result of the face detection;
[0031] Update the position and content of the face layer according to the real-time updated face region coordinates and face data;
[0032] Render and synthesize the updated face layer with the portrait layer of the source video stream at the current moment into a video stream;
[0033] Keep the update of the face layer synchronized with the frame rate of the source video stream at the current moment, and output the synthesized video stream to a display device.
[0034] In addition, to achieve the above object, the present application also proposes a face dynamic tracking device, and the face dynamic tracking device includes:
[0035] A data acquisition module, configured to acquire a source video stream and transmit the source video stream to a server through a network to obtain first processed data;
[0036] A data parsing module, obtains the first processed data after processing from the server, and parses the first processed data through a face recognition application to obtain second processed data;
[0037] A face detection module, performs face detection on the second processed data, and obtains the face region coordinates and face data of the facial image;
[0038] A data processing module, creates a face layer according to the face region coordinates and draws a recognition frame in the face layer;
[0039] A data display module, configured to display the corresponding face data in the face layer, and superimpose and render the face layer with real-time updates on the portrait layer of the source video stream at the current moment.
[0040] In addition, to achieve the above object, the present application further provides a face dynamic tracking device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the face dynamic tracking method as described above.
[0041] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the face dynamic tracking method as described above.
[0042] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the face dynamic tracking method as described above.
[0043] One or more technical solutions proposed by the present application have at least the following technical effects:
[0044] Collect the source video stream and transmit the source video stream to the server through the network to obtain the first processed data; obtain the processed first processed data from the server, and parse the first processed data through a face recognition application to obtain the second processed data; perform face detection on the second processed data to obtain the face area coordinates and face data of the facial image; create a face layer according to the face area coordinates and draw a recognition frame in the face layer; display the corresponding face data in the face layer, and superimpose and render the real-time updated face layer on the portrait layer of the source video stream at the current moment, and implement dynamic tracking of the face in the video stream through the high-efficiency solution of the superimposed layer, and only return the personnel and coordinate data of the recognition result, greatly reducing the amount of data transmitted over the network, and improving the efficiency of face dynamic tracking through face detection, adding a new layer and a recognition frame, and ensuring the real-time nature of the feedback display. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0047] Figure 1 It is a flowchart provided for Embodiment 1 of the face dynamic tracking method of the present application;
[0048] Figure 2 It is a schematic flowchart provided for the second embodiment of the applicant's face dynamic tracking method;
[0049] Figure 3 It is a schematic flowchart provided for the third embodiment of the applicant's face dynamic tracking method;
[0050] Figure 4 It is a schematic flowchart provided for the fourth embodiment of the applicant's face dynamic tracking method;
[0051] Figure 5 It is a schematic flowchart provided for the fifth embodiment of the applicant's face dynamic tracking method;
[0052] Figure 6 It is a schematic flowchart provided for the sixth embodiment of the applicant's face dynamic tracking method;
[0053] Figure 7 It is a tracking rendering schematic diagram provided for the seventh embodiment of the applicant's face dynamic tracking method;
[0054] Figure 8 It is a tracking rendering schematic diagram provided for the eighth embodiment of the applicant's face dynamic tracking method;
[0055] Figure 9 It is a schematic diagram of the module structure of the face dynamic tracking device in the embodiment of the present application;
[0056] Figure 10 It is a schematic diagram of the device structure of the hardware operating environment involved in the face dynamic tracking method in the embodiment of the present application.
[0057] The realization of the purpose, functional characteristics and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments
[0058] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0059] To better understand the technical solutions of the present application, the following will be described in detail in conjunction with the drawings of the specification and specific embodiments.
[0060] Due to the problems of slow network transmission of video streams, low face recognition efficiency, poor real-time performance of rendering feedback and high latency in the prior art.
[0061] This application provides a solution, which collects the source video stream and transmits the source video stream to the server through the network to obtain the first processed data; obtains the processed first processed data from the server, and analyzes the first processed data through a face recognition application to obtain the second processed data; performs face detection on the second processed data to obtain the face area coordinates and face data of the facial image; creates a face layer according to the face area coordinates and draws a recognition frame in the face layer; displays the corresponding face data in the face layer, and superimposes and renders the face layer with real-time updates on the portrait layer of the source video stream at the current moment.
[0062] Based on this, an embodiment of this application provides a face dynamic tracking method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the face dynamic tracking method of this application.
[0063] In this embodiment, the face dynamic tracking method includes steps S10 to S50:
[0064] Step S10, collect the source video stream and transmit the source video stream to the server through the network to obtain the first processed data;
[0065] It should be noted that a camera or VR glasses can be used to obtain the source video stream; the source video stream is subjected to necessary preprocessing, such as adjusting the resolution, frame rate, etc.
[0066] Step S20, obtain the processed first processed data from the server, and analyze the first processed data through a face recognition application to obtain the second processed data;
[0067] It should be noted that the face recognition application is used to analyze the first processed data, and the face recognition application uses an efficient face detection algorithm to analyze and obtain the second processed data video frame.
[0068] Step S30, perform face detection on the second processed data to obtain the face area coordinates and face data of the facial image;
[0069] It should be noted that through the face recognition module, according to the matching of the face database data, the face area is recognized from the captured facial image, and the face position can be accurately and quickly located;
[0070] Step S40, create a face layer according to the face area coordinates and draw a recognition frame in the face layer;
[0071] It should be noted that the drawing of the recognition frame needs to accurately correspond to the face area to provide intuitive visual feedback, and at the same time, layer management needs to ensure that the face layer is correctly superimposed on the source video stream.
[0072] Step S50, display the corresponding face data in the face layer, and superimpose and render the face layer with real-time updates on the portrait layer of the source video stream at the current moment.
[0073] It should be noted that information such as name, age, and expression status can be rendered on the face layer, and the face data of the current frame can be adjusted by updating the face layer in real time. The data rendering should be clearly visible and avoid conflicts or confusion with the background in the source video stream. The superimposed rendering must ensure that the face layer is updated synchronously with the video stream to achieve real-time tracking.
[0074] In this embodiment, the source video stream is collected and transmitted to the server through the network to obtain the first processed data; the first processed data after processing is obtained from the server, and the second processed data is obtained by parsing the first processed data through a face recognition application; face detection is performed on the second processed data to obtain the face area coordinates and face data of the facial image; a face layer is created according to the face area coordinates and a recognition frame is drawn in the face layer; the corresponding face data is displayed in the face layer, and the face layer with real-time updates is superimposed and rendered on the portrait layer of the source video stream at the current moment. By using the high-efficiency scheme of the superimposed layer, the face in the video stream is dynamically tracked. Only the personnel and coordinate data of the recognition result are returned, which greatly reduces the amount of data transmitted over the network. The efficiency of face dynamic tracking is improved through face detection, adding a new layer, and the recognition frame, ensuring the real-time nature of the feedback display.
[0075] In one embodiment, before the step of collecting the source video stream and transmitting the source video stream to the server through the network to obtain the first processed data, it includes:
[0076] Copy the source video stream to obtain a target video stream, and directly output the target video stream to the user terminal for display.
[0077] Specifically, as Figure 2 shown, by directly copying the source video stream, it is directly output as the target video stream to the user terminal, avoiding problems such as network latency and video stuttering caused by transmission to the server and then returned by the server. Through this method, the source video stream sent to the server alone can reduce the amount of data transmitted over the network by compressing pixels and other means, so that it can meet the pixel requirements for face recognition applications on the server.
[0078] Based on the above step method, the transfer from the video capture device to the server and the terminal is realized; for the video stream transmitted to the server over the network, it is parsed through a face recognition application to obtain video frames. Through the face recognition module, according to the data in the face database for matching, the face area is recognized from the captured image to obtain the coordinates and face data of the face area.
[0079] Further, referring to Figure 3 , the third embodiment of the intelligent diagnosis method of the present application provides a process schematic diagram. Based on the above Figure 3 illustrated embodiment, the step of "obtaining the processed first processed data from the server and parsing the first processed data through a face recognition application to obtain the second processed data" in step S20 is further refined, including steps A301 to A303:
[0080] Step A301: Receive the processed first processed data from the server through a network transmission method;
[0081] It should be noted that the client device establishes a network connection with the server, and the client receives the first processed data sent by the server; the network transmission protocol can support high-speed and reliable data transmission, such as TCP and UDP.
[0082] Step A302: Perform decoding processing on the first processed data;
[0083] It should be noted that the first processed data is scanned using an appropriate decoding algorithm to restore the data to its original format, such as a video stream or video frame, while ensuring that no key information is lost during the decoding process.
[0084] Step A303: Extract the decoded data of the first processed data frame by frame through a face recognition application to obtain video frames.
[0085] It should be noted that the decoded data is input into the face recognition application, and the application analyzes the data frame by frame, extracts the face images in each frame, and the face recognition algorithm in the application processes the extracted face images.
[0086] In this embodiment, the processed first processed data is obtained from the server, and the first processed data is parsed through a face recognition application to restore the data to its original format. Through a dedicated face recognition application, the system can quickly and accurately detect and identify the faces in the video frames, which helps to achieve face detection and tracking.
[0087] Further, referring to Figure 4 , the fourth embodiment of the intelligent diagnosis method of the present application provides a process schematic diagram. Based on the above Figure 4 illustrated embodiment, the step of "performing face detection on the second processed data and obtaining the face region coordinates and face data of the facial image" in step S30 is further refined, including steps A401 to A404:
[0088] Step A401: Receive the video frames generated by the face recognition application as the second processed data;
[0089] It should be noted that after obtaining the video frames processed frame by frame from the face recognition application, it is necessary to ensure that the quality of the video frames meets the requirements of face detection. For example, the video frames should contain clear face images, and the format and resolution of the input data should match the requirements of the face detection module. Then, the video frames are passed to the face detection module as input data.
[0090] Step A402: Load the face recognition module to recognize the facial image of the second processed data;
[0091] It should be noted that by loading a pre-trained face recognition model to analyze the facial images in the video frames, the face images in the video frames are recognized and located.
[0092] Step A403: Perform data matching between the facial image and the face database in the face recognition module;
[0093] It should be noted that the face database contains the feature data of known faces. By extracting the feature points or feature vectors of the facial image and comparing them with the data in the face database, the similarity or distance is calculated to determine whether there is a match.
[0094] Step A404: Obtain the face area coordinates and face data of the facial image according to the data matching result.
[0095] It should be noted that the face area coordinates are used to draw a recognition frame to determine the position and size of the face according to the matching result, record the coordinate information of the face area, and at the same time extract or generate face data related to the facial image, such as name, age, mood state, etc.
[0096] In this embodiment, by receiving high-quality video frames, the facial images in the video frames can be quickly and accurately recognized by loading the pre-trained face recognition model, improving the real-time performance of the system; by extracting the feature points or feature vectors of the facial image and comparing them with the face database, the system can quickly determine whether there is a match, improving the accuracy of the entire face recognition system; determining the position and size of the face according to the matching result can provide accurate coordinate information for subsequent drawing of the recognition frame and display of face data.
[0097] Further, referring to Figure 5 , the fifth embodiment of the intelligent diagnosis method of this application provides a flow schematic diagram. Based on the above Figure 5 shown embodiment, the step of "creating a face layer according to the face area coordinates and drawing a recognition frame in the face layer" in step S40 is further refined, including steps A501 to A503:
[0098] Step A501: Obtain the face area coordinates to determine the position and size of the face in the image;
[0099] It should be noted that the face region coordinates are usually in pixels. After determining the position and size of the face in the image, the accuracy of the coordinate data can be verified through image processing techniques.
[0100] Step A502: Create a new transparent face layer according to the position and size.
[0101] It should be noted that the creation of the layer can use a graphics processing library or framework. By creating a transparent person layer, the source video source is not blocked, and the transparent layer allows the user to see the person or the source video background at the same time, and the new layer is placed in the correct position to cover the face region.
[0102] Step A503: Draw a rectangular box around the entire face region according to the face region coordinates as the recognition box.
[0103] It should be noted that a rectangular box is created on the newly created transparent layer. The size and position of the rectangular box match the face region, and a color can also be added to the rectangular box.
[0104] In this embodiment, by obtaining and parsing the coordinate data of the face region, the system can accurately determine the position and size of the face in the image; creating a transparent face layer can ensure that the user does not block the background of the original video stream while viewing the face information, enhancing the user's visual experience; drawing a rectangular box around the entire face region as the recognition box can visually identify the face in the video, facilitating the user to quickly identify and providing a basis for face dynamic tracking.
[0105] In one embodiment, the step of displaying the corresponding face data in the face layer includes:
[0106] Inside or near the drawn recognition box, adjust the face state by adjusting the line thickness, color of the recognition box, and the position and transparency of the data label, and display the corresponding face data.
[0107] It should be noted that to emphasize or weaken the visual effect of the recognition box as needed and adjust the line thickness, the visibility of the recognition box in the face layer can be preset by the corresponding color to ensure that the color does not conflict with the background or the face data label; in addition, the best position of the data label is usually inside or near the recognition box, which helps the user quickly associate the face and the data, and the transparency of the data label can also be adjusted to ensure that the person image is not completely blocked; according to the face recognition result or user interaction, adjust the state of the recognition box, such as changing the color to represent different recognition states; update the face data in real time to reflect the information of the current video frame.
[0108] In this embodiment, by drawing the positional relationship between the recognition frame and the data label, the face data is ensured to be presented clearly and accurately, while not interfering with the user's viewing of the source video stream, effectively integrating the face data and the recognition frame into the video stream, and realizing efficient face dynamic tracking and recognition.
[0109] Further, referring to Figure 6 , the sixth embodiment of the intelligent diagnosis method of this application provides a process schematic diagram. Based on the above Figure 6 shown embodiment, the step of "overlaying and rendering the face layer updated in real time on the portrait layer of the source video stream at the current moment" in step S50 is further refined, including steps A601 to A604:
[0110] Step A601: Update the face area coordinates and face data in real time according to the result of the face detection.
[0111] It should be noted that continuously receive the real-time detection results from the face detection module, parse the detection results to obtain the latest face area coordinates and face data, and update the face information reflecting the current video frame; the real-time update needs to be fast and accurate to maintain the coherence of face tracking, and the update frequency should also match the frame rate of the video stream.
[0112] Step A602: Update the position and content of the face layer according to the face area coordinates and face data updated in real time.
[0113] It should be noted that adjust the position of the face layer to align it with the latest face area coordinates; update the layer content, including the recognition frame and the data label, to reflect the latest face data; optimize the visual flicker or jitter caused by the layer update to achieve smooth transition to provide a good user experience; the layer content update should be synchronized with the face data update.
[0114] Step A603: Render and synthesize the updated face layer with the portrait layer of the source video stream at the current moment into a video stream.
[0115] It should be noted that use the rendering engine to overlay the face layer on the portrait layer of the source video stream to ensure a natural overlay effect and the integration of the face layer and the portrait layer background; utilize image processing techniques, such as adjusting the transparency of the face layer, etc., to optimize the visual effect and ensure that the recognition frame and face data are clearly visible in the rendering effect.
[0116] Step A604: Keep the update of the face layer synchronized with the frame rate of the source video stream at the current moment, and output the synthesized video stream to the display device.
[0117] It should be noted that the update frequency of the synchronized face layer is synchronized with the frame rate of the source video stream to ensure that there is no delay or frame loss when the synthesized video stream is output to the display device, and the quality of the output video stream should meet the requirements of the display device.
[0118] Specifically, Figure 7 The following is a schematic diagram of traditional tracking rendering. By using the face region coordinates, drawing is performed on the original video frame to obtain a target frame with a tracking box and personnel information, and finally it is encapsulated into a target video stream and transmitted back to the terminal display device through the network; Figure 8 The following is a schematic diagram of the tracking rendering of the present application. Through the concept of adding a new layer, no drawing processing is performed on the video frame. It only needs to transmit the results of face recognition, that is, the coordinate values and personnel information, back to the front end through the network. The front-end device can draw a tracking box and personnel information at the specified position to achieve the same effect.
[0119] In this embodiment, by updating the face region coordinates and face data in real time, the system can quickly respond to changes in users or scenarios, ensure the real-time and accuracy of face tracking, and improve the tracking effect in dynamic scenarios; according to the real-time updated face region coordinates, the position of the face layer can be accurately adjusted to ensure that the layer is consistent with the actual position of the face; rendering and synthesizing the updated face layer with the portrait layer of the source video stream can create a video stream containing face recognition information without affecting the smoothness of the original video; ensuring that the update of the face layer is synchronized with the frame rate of the source video stream can avoid delays or frame losses during display, maintain the real-time nature of tracking, and output the synthesized video stream to the display device. Users can directly view the video with real-time face recognition information, improving the usability and user experience of the system.
[0120] In addition, the present application also provides a face dynamic tracking device. Please refer to Figure 9 , the face dynamic tracking device includes:
[0121] A data acquisition module 10, configured to acquire a source video stream and transmit the source video stream to a server through the network to obtain first processed data;
[0122] A data parsing module 20, which obtains the first processed data processed by the server and parses the first processed data through a face recognition application to obtain second processed data;
[0123] A face detection module 30, which performs face detection on the second processed data to obtain the face region coordinates and face data of the facial image;
[0124] A data processing module 40, which creates a face layer according to the face region coordinates and draws a recognition box in the face layer;
[0125] The data display module 50 is configured to display the corresponding face data in the face layer, and superimpose and render the face layer with real-time updates on the portrait layer of the source video stream at the current moment.
[0126] The face dynamic tracking device provided in this application adopts the face dynamic tracking method in the above embodiment, and can solve the technical problems of slow network transmission of the video stream, low face recognition efficiency, poor real-time performance of rendering feedback, and high latency. Compared with the prior art, the beneficial effects of the face dynamic tracking device provided in this application are the same as those of the face dynamic tracking method provided in the above embodiment, and other technical features in the face dynamic tracking device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.
[0127] In addition, this application provides a face dynamic tracking device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the face dynamic tracking method in the first embodiment above.
[0128] Next, refer to Figure 10 , which shows a schematic structural diagram of a face dynamic tracking device suitable for implementing the embodiments of this application. The face dynamic tracking device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 10 The face dynamic tracking device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of this application.
[0129] As Figure 10As shown, the face dynamic tracking device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the xxx device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the face dynamic tracking device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a face dynamic tracking device with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems can be alternatively implemented or had.
[0130] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0131] The face dynamic tracking device provided by the present application adopts the face dynamic tracking method in the above-mentioned embodiment, and can solve the technical problems of slow network transmission of video streams, low face recognition efficiency, poor real-time performance of rendering feedback, and high latency. Compared with the prior art, the beneficial effects of the face dynamic tracking device provided by the present application are the same as those of the face dynamic tracking method provided by the above-mentioned embodiment, and other technical features in the face dynamic tracking device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0132] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0133] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0134] In addition, this application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the face dynamic tracking method in the above embodiments.
[0135] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0136] The above computer-readable storage medium can be included in the face dynamic tracking device; or it can exist separately without being assembled into the face dynamic tracking device.
[0137] The above computer-readable storage medium carries one or more programs, which, when executed by the face dynamic tracking device, cause the face dynamic tracking device to: collect a source video stream and transmit the source video stream to a server via a network to obtain first processed data; obtain the first processed data after processing from the server, and parse the first processed data through a face recognition application to obtain second processed data; perform face detection on the second processed data to obtain the face area coordinates and face data of a facial image; create a face layer according to the face area coordinates and draw a recognition frame in the face layer; display the corresponding face data in the face layer, and superimpose and render the face layer with real-time updates on the portrait layer of the source video stream at the current moment.
[0138] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0140] The modules involved in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0141] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned face dynamic tracking method, which can solve the technical problems of slow network transmission of video streams, low face recognition efficiency, poor real-time performance of rendering feedback, and high latency. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the face dynamic tracking method provided by the above embodiments, and will not be elaborated here.
[0142] In addition, the present application also provides a computer program product, including a computer program, and the steps of the above-mentioned face dynamic tracking method are implemented when the computer program is executed by a processor.
[0143] The computer program product provided by the present application can solve the technical problems of slow network transmission of video streams, low face recognition efficiency, poor real-time performance of rendering feedback, and high latency. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the face dynamic tracking method provided by the above embodiments, and will not be elaborated here.
[0144] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A method for dynamic face tracking, characterized in that: The method comprises: Collecting a source video stream and transmitting the source video stream to a server via a network to obtain first processed data; Obtaining processed first processed data from the server, and parsing the first processed data through a face recognition application to obtain second processed data; Performing face detection on the second processed data to obtain face region coordinates and face data of the facial image; Create a face layer according to the face area coordinates and draw a recognition frame in the face layer; The corresponding face data is displayed in the face layer, and the face layer updated in real time is overlaid and rendered on the image layer of the source video stream at the current moment.
2. The method according to claim 1, characterized in that The step of collecting the source video stream and transmitting the source video stream to the server through the network to obtain the first processed data includes: The source video stream is copied to obtain a target video stream, and the target video stream is directly output to a user terminal for display.
3. The method according to claim 1, characterized in that The step of obtaining the processed first processed data from the server and parsing the first processed data by a face recognition application to obtain the second processed data comprises: receiving first processed data processed by the server through a network transmission method; Decoding the first processed data; The decoded data of the first processed data is extracted frame by frame by applying face recognition to obtain video frames.
4. The method according to claim 3, characterized in that The step of performing face detection on the second processed data to obtain face region coordinates and face data of the facial image comprises: receiving a video frame generated by a face recognition application as second processed data; Loading a face recognition module to recognize a facial image of the second processed data; Matching the facial image with the face database in the face recognition module; The face region coordinates and face data of the face image are obtained according to the data matching result.
5. The method according to claim 4, characterized in that The step of creating a face layer according to the face area coordinates and drawing an identification frame in the face layer comprises: Obtaining the coordinates of the face region to determine the position and size of the face in the image; Create a new transparent face layer at the position and size described; A rectangular frame surrounding the entire face area is drawn according to the face area coordinates as a recognition frame.
6. The method according to claim 5, characterized in that The step of displaying the corresponding face data in the face layer comprises: Inside or near the drawn identification frame, the face state is adjusted by adjusting the line thickness, color of the identification frame and the position and transparency of the data label to display the corresponding face data.
7. The method according to claim 6, characterized in that The step of overlaying and rendering the real-time updated face layer on the image layer of the source video stream at the current moment comprises: According to the result of the face detection, the face area coordinates and face data are updated in real time; Update the position and content of the face layer according to the real-time updated face area coordinates and face data; Rendering the updated face layer and the image layer of the source video stream at the current moment into a composite video stream; The updating of the face layer is kept synchronized with the frame rate of the source video stream at the current moment, and the synthesized video stream is output to a display device.
8. A face dynamic tracking device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for dynamic face tracking according to any one of claims 1 to 7.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the face dynamic tracking method according to any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the method for dynamic face tracking according to any one of claims 1 to 7 are implemented.