Main figure image segmentation method and device, equipment and medium

By detecting and tracking the regional anchor boxes in the live stream, combining the significance judgment conditions and lightweight model, the precise segmentation of the subject characters is achieved, solving the problems of unstable multi-person scene segmentation and high computational complexity in the prior art, and real-time segmentation effect on low-performance devices is achieved.

CN120163977APending Publication Date: 2025-06-17GUANGZHOU FANGGUI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510234751.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively distinguish between the main characters and the background-independent characters when dealing with multi-person scenarios, and the calculation complexity is high, making it difficult to achieve real-time segmentation, and there are also problems with stability.

Method used

By detecting the regional anchor box of each character in the image frame of the live stream, and filtering out the regional anchor box of the subject character based on the preset significance determination conditions, tracking and smoothing processing, combining the lightweight object detection model and image segmentation model, the precise segmentation of the subject character is achieved.

Benefits of technology

It solves the problem that traditional technology cannot distinguish between the main characters and the background and has no relationship with the characters, and realizes efficient and stable real-time segmentation on low-performance devices, and meets the scenario requirements of high real-time requirements such as video conferencing and live broadcasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163977A_ABST
    Figure CN120163977A_ABST
Patent Text Reader

Abstract

The invention relates to the field of network live broadcast, and discloses a main body figure image segmentation method and device, equipment and a medium, and the method comprises the steps: detecting an image region where each figure is located in an image frame of a live broadcast stream, and determining a region anchor frame corresponding to the image region where each figure is located; according to a preset significance judgment condition, determining a region anchor frame which correspondingly indicates the image region where the main body figure is located from all region anchor frames as a main body anchor frame; tracking and smoothing the main body anchor frame to stabilize the position information of the main body anchor frame; and according to the position information of the main body anchor frame, carrying out image segmentation on an image frame in the live broadcast stream to obtain a human body image of a main body person. According to the method, the limitation of a traditional portrait segmentation technology in a multi-person scene is overcome, precise recognition and stable segmentation of the main body character are realized through an innovative technical means, and a more efficient and accurate technical solution is provided for related application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to network live broadcast technology, and particularly to a method and apparatus, device, and medium for segmenting the image of a main person. Background Art

[0002] In the fields of computer vision and image processing, the technology of human figure segmentation has always been one of the research hotspots, and is widely applied to multiple scenarios such as image editing, virtual reality, augmented reality, video conferencing, etc. However, the existing traditional technologies have some obvious limitations when dealing with multi-person scenarios, and these problems are particularly prominent in practical applications.

[0003] First of all, traditional human figure segmentation technologies usually regard all the people in the picture as segmentation targets, and cannot effectively distinguish the main person in the foreground from the irrelevant people in the background. For example, in video conferencing or live broadcast scenarios, it is often necessary to exclude the irrelevant people in the background from the segmentation range to achieve better visual effects, such as background blurring or background replacement. However, the existing segmentation technologies cannot meet this requirement, resulting in the possibility that the segmentation result may include unnecessary background people, affecting the final visual effect.

[0004] Secondly, traditional technologies have significant deficiencies in terms of real-time performance. Most of the existing human figure segmentation algorithms have a high computational complexity and are difficult to meet the requirements of low latency in real-time applications, especially on devices with low computing power, such as mobile devices or PC hosts with low-end graphics cards. This problem of poor real-time performance limits the wide application of human figure segmentation technology in practical scenarios, especially in video live broadcast and video conferencing with high requirements for real-time performance.

[0005] In addition, the stability of traditional technologies in multi-person scenarios also has problems. When there are multiple moving targets in the picture, the existing tracking and segmentation technologies often have difficulty accurately distinguishing and continuously tracking the main person, and are easily affected by the interference of other people or the background, resulting in unstable segmentation results, and even misjudgment or loss of targets.

[0006] The existence of these problems has greatly limited the application of traditional human figure segmentation technologies in multi-person scenarios. Therefore, there is an urgent need for a new technical solution that can effectively solve the above problems to achieve accurate segmentation of the most prominent person in the picture, while meeting the requirements of real-time performance and stability, so as to better adapt to the needs of various practical application scenarios. Summary of the Invention

[0007] The purpose of the present application is to solve the above problems and provide a method and corresponding apparatus, device, and non-volatile readable storage medium for segmenting the image of a main person.

[0008] According to one aspect of the present application, a method for segmenting a subject person's image is provided, including the following steps: detecting the image region where each person is located in the image frame of the live stream, and determining the region anchor box corresponding to the image region where each person is located; according to the preset saliency determination condition, determining the region anchor box corresponding to the image region indicating the subject person from all the region anchor boxes as the subject anchor box; performing tracking and smoothing processing on the subject anchor box to stabilize the position information of the subject anchor box; and performing image segmentation on the image frame in the live stream according to the position information of the subject anchor box to obtain the human body image of the subject person.

[0009] According to another aspect of the present application, a device for segmenting a subject person's image is provided, including: an anchor box detection module configured to detect the image region where each person is located in the image frame of the live stream and determine the region anchor box corresponding to the image region where each person is located; a subject determination module configured to determine the region anchor box corresponding to the image region indicating the subject person from all the region anchor boxes as the subject anchor box according to the preset saliency determination condition; a tracking and smoothing module configured to perform tracking and smoothing processing on the subject anchor box to stabilize the position information of the subject anchor box; and a segmentation processing module configured to perform image segmentation on the image frame in the live stream according to the position information of the subject anchor box to obtain the human body image of the subject person.

[0010] According to another aspect of the present application, a device for segmenting a subject person's image is provided, including a central processing unit and a memory, and the central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the method for segmenting a subject person's image according to the present application.

[0011] According to another aspect of the present application, a non-volatile readable storage medium is provided, which stores a computer program implemented according to the method for segmenting a subject person's image in the form of computer-readable instructions, and when the computer program is called and run by a computer, it executes the steps included in the method.

[0012] Aiming at the limitations of traditional portrait segmentation technology in multi-person scenarios, the present application can accurately identify and stably track the subject person by detecting the region anchor boxes of each person in the live stream and screening out the region anchor boxes of the subject person according to the saliency determination condition. Combining tracking and smoothing processing ensures the stability and accuracy of the subject anchor box in consecutive frames, and accurately extracts the subject person's image through image segmentation.

[0013] Accordingly, the present application achieves remarkable beneficial effects, including but not limited to: First, it solves the problem that traditional technologies cannot distinguish the main character from irrelevant background characters, improving the accuracy and practicality of segmentation; Second, it realizes efficient and stable real-time segmentation on low-performance devices, meeting the requirements of scenarios with high real-time requirements such as video conferencing and live streaming; Third, through tracking and smoothing processing, it solves the problem of unstable segmentation results in multi-person scenarios, ensuring the stable tracking and accurate segmentation of the main character. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a network architecture exemplary suitable for applying the main character image segmentation method of the present application;

[0015] Figure 2 is a flowchart of an embodiment of the main character image segmentation method of the present application;

[0016] Figure 3 is a schematic block diagram of the principle of the main character image segmentation device of the present application;

[0017] Figure 4 is a schematic structural diagram of a main character image segmentation device adopted by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Before introducing the specific embodiments of the technical solution of the present application in detail, the network architecture and application scenarios suitable for supporting the implementation of the technical solution of the present application are first revealed.

[0019] The technical solution of the present application is applicable to a variety of video processing fields, especially suitable for scenarios that require accurate identification and segmentation of the main character in a video stream, such as network live streaming, video conferencing, virtual reality, and augmented reality. In this context, the technical solution of the present application can be applied in a typical network live streaming platform, such as Figure 1 shown, the platform consists of a live streaming server 83, a media server 85, the terminal device 90 of the live streamer user, and the terminal device 92 of the audience user.

[0020] The live streaming server 83 is responsible for managing the creation of the live streaming room and the reception and distribution of the live stream. As the central node of the live stream, it receives the live stream from the terminal device 90 of the live streamer user and distributes it to the media server 85 and the terminal device 92 of the audience user. The live streaming server 83 is also responsible for processing the management tasks of the live streaming room, such as user authentication, live streaming room settings, etc.

[0021] The media server 85 undertakes the tasks of encoding, transcoding, and storing the live stream. It performs necessary transcoding on the live stream according to the network conditions and device capabilities of the viewer users to ensure that the live stream can be transmitted to the viewers in a suitable format and quality. The media server 85 is also responsible for storing the live stream for replay or other subsequent processing.

[0022] The terminal device 90 of the host user is the source of the live content. The host captures video through the terminal device 90 and sends the original video stream to the live server 83. The terminal device 90 of the host user can be a smart phone, a tablet computer, a laptop computer, or a professional camera device, which are connected to the live server 83 through the Internet.

[0023] The terminal device 92 of the viewer user is the receiving end of the live content. The viewers watch the live stream through the terminal device 92, and these devices can be smart phones, tablet computers, personal computers, smart TVs, etc. The terminal device 92 of the viewer user is connected to the live server 83 through the Internet, receives the live stream, and plays it.

[0024] The transmission process of the live stream is as follows: The terminal device 90 of the host user captures video and notifies the live server 83, and the live server 83 distributes the live stream to the media server 85. The media server 85 performs encoding and transcoding on the live stream, and then sends the network address of the processed live stream back to the live server 83. Then, the live server 83 transmits the network address of the live stream to the terminal device 92 of the viewer user so that the terminal device 92 can load and play the live stream according to the network address.

[0025] Programming according to the main character image segmentation method of the present application can be implemented as a computer program product, which can be selectively deployed in nodes such as the terminal device 90 of the host user, the media server 85, or the terminal device 92 of the viewer user, so as to recognize the character image corresponding to the main character in the image frames of the live stream processed by the node. The program product contains the algorithms and logics required to implement the technical solution of the present application, and can accurately segment the main character in the video frames of the live stream, so as to carry out various subsequent applications, including but not limited to background blurring, background replacement, character beauty makeup, etc., and interactive functions such as virtual gift giving to specific main characters.

[0026] Please refer to Figure 2 , according to a main character image segmentation method provided by the present application, it can be implemented as a computer program product, installed and run in each node device of the network live broadcast platform, such as the terminal device of the host user, the media server, or the terminal device of the viewer user, to be responsible for implementing the main character image segmentation processing in the image frames of the live stream. In some embodiments of the method, the method includes the following steps:

[0027] Step S3100: Detect the image regions where each person is located in the image frames of the live stream, and determine the corresponding regional anchor boxes for the image regions where each person is located.

[0028] Taking the exemplary network live streaming scenario of this application as an example, when it is necessary to identify the main characters in the live stream and perform human image segmentation, first detect the image regions where each person is located in the image frames of the live stream, and determine the corresponding regional anchor boxes for the image regions where each person is located. The purpose is to accurately identify the positions and ranges of all people from each frame of the live stream image, providing basic data for subsequent main character identification and image segmentation. In actual application scenarios, there may be multiple people in the live stream at the same time. Therefore, it is necessary to be able to accurately distinguish and identify each person to ensure the accuracy and effectiveness of subsequent processing.

[0029] Specifically, the detection of regional anchor boxes can be achieved in various ways. For example, object detection models based on deep learning can be used, such as the YOLO (You Only Look Once) series, SSD (Single Shot MultiBoxDetector), or Faster R-CNN, etc. These models can quickly and accurately detect the people in the image and generate a bounding box for each person, that is, the regional anchor box. These models learn the characteristics of people through a large amount of training data, so as to accurately locate people in complex backgrounds. In actual applications, an appropriate detection model can be selected according to the resolution and frame rate of the live stream and the computing power of the device. For example, on mobile devices with limited computing resources, lightweight models such as YOLOv5 or NanoDet can be preferred. These models can run at a relatively high speed while ensuring detection accuracy, meeting the requirements of real-time performance.

[0030] In addition, in order to further improve the accuracy and robustness of detection, multi-scale detection technology can also be combined. Multi-scale detection technology samples and detects the image at different scales, which can effectively solve the problem of different person sizes. Especially in the live streaming scenario, people may appear in different sizes due to the distance from the camera. For example, the image can be downsampled by a pyramid to generate image copies with different resolutions, then people detection is performed separately on each copy, and finally the detection results are merged to obtain a more comprehensive and accurate detection result.

[0031] The regional anchor box of this application refers to the rectangular box output by the object detection model for identifying the position and range of a person. The coordinates of the regional anchor box usually take the upper left corner of the image as the origin and are represented by four parameters: the abscissa of the upper left corner, the ordinate of the upper left corner, the width of the box, and the height of the box. These parameters can uniquely determine the position and size of a person in the image. For example, assuming that in an image frame with a resolution of 1920×1080, the regional anchor box of a person is detected as (500, 300, 200, 400), which means that the person is located in the middle-right position of the image and their head is roughly in the upper half of the image.

[0032] In a multi-person scenario, multiple regional anchor boxes can be generated through the object detection model, and each anchor box corresponds to a person. These regional anchor boxes provide accurate position information for subsequent identification of the main person, ensuring that the main person can be accurately distinguished from the multi-person scenario.

[0033] In some embodiments, object detection can be performed on a certain number of image frames at intervals to determine the regional anchor boxes in these image frames, and then the main anchor box can be determined through step S3200. For the image frames that are not detected in the middle, the main anchor box obtained by tracking and smoothing the already determined main anchor box in step S3300 of this application can be used for replacement until a new main anchor box is determined at the next interval and then replaced, and so on, continuously iterating. This not only saves system overhead but also improves the efficiency of identifying the main person.

[0034] Step S3200: Determine the regional anchor box corresponding to the image area indicating the main person from all the regional anchor boxes as the main anchor box according to the preset significance determination condition;

[0035] After the positions and ranges of all the people in each frame of the live stream are accurately identified through the regional anchor boxes, the regional anchor box corresponding to the image area indicating the main person can be determined from all the regional anchor boxes as the main anchor box according to the preset significance determination condition, so as to accurately identify the most prominent person, that is, the main person, in the multi-person scenario, so that the subsequent image segmentation process can focus on the main person, thereby achieving an accurate image segmentation effect.

[0036] The saliency determination conditions can be predefined based on various factors, such as the position, size, and motion state of a person in the image. In one embodiment, the saliency determination conditions are determined based on the position and size of the person in the image. For example, a saliency determination condition can be preset to consider the person closest to the center of the image and with the largest area as the salient person. The rationale for this condition setting is that the anchor person usually locates at the center of the screen and occupies a relatively large proportion of the screen area. In practical applications, the salient person can be determined by calculating the distance between the center point of each regional anchor box and the center point of the image and comparing the area sizes of the regional anchor boxes.

[0037] In addition to the saliency determination conditions based on position and size, in another embodiment that is more inclined to support motion scenarios, the saliency determination conditions can also be set in combination with the motion state of the person. For example, a saliency determination condition can be preset to consider the person with a faster motion speed as the salient person. The rationale for this condition setting is that the anchor person may have more actions and movements during the live broadcast. In practical applications, the salient person can be determined by tracking the regional anchor boxes in consecutive frames and calculating the motion speed of each person.

[0038] In yet another embodiment, the saliency determination conditions can also be set by comprehensively considering multiple action information reflected in the person images in the live stream, such as setting these conditions according to any one or any combination of the facing angle of the person, the eye gaze duration, etc., so as to determine the main anchor box from multiple regional anchor boxes.

[0039] In some complex live broadcast scenarios, it may be necessary to comprehensively consider multiple factors to determine the saliency determination conditions. For example, factors such as the position, size, facing angle, eye gaze duration, and motion state of the person can be considered simultaneously. By setting weights, these factors can be combined for comprehensive judgment. For example, a comprehensive saliency determination condition can be preset, where the weight of the position factor is 0.4, the weight of the size factor is 0.3, and the weight of the motion state factor is 0.3. In this way, the salient person can be determined more accurately, thereby improving the accuracy and robustness of the main person image segmentation.

[0040] It can be seen that this application can accurately determine, from all the regional anchor boxes in each frame of the live stream image, the regional anchor box corresponding to the image area indicating the main person as the main anchor box according to the preset saliency determination conditions. This process not only considers the situation where there may be multiple people in the live stream at the same time, but also improves the accuracy and robustness of the salient person recognition through the flexible application of multiple saliency determination conditions, providing accurate main person position information for subsequent image segmentation processing.

[0041] Step S3300: Track and smooth the subject anchor box to stabilize the position information of the subject anchor box;

[0042] After determining the regional anchor box of the subject person in the live stream, in order to ensure the stability and accuracy of the position information of the subject person in the consecutive image frames of the live stream, it is necessary to track and smooth the subject anchor box. Thereby, the problem of jitter or discontinuity of the position of the subject anchor box caused by factors such as human movement, camera jitter, or detection noise in a multi-person scenario is solved.

[0043] Specifically, the technical means of tracking and smoothing can be implemented through a variety of technologies. In one embodiment, a Kalman Filter can be used to implement the tracking and smoothing of the position information of the subject anchor box. The Kalman Filter is an efficient self-recursive filter that can estimate the dynamic state of the system from a series of noisy measurements. In this application, the Kalman Filter first predicts the position of the subject anchor box in the next frame based on the historical position information of the subject anchor box, and then corrects the predicted position by combining the actual detection results of the current frame, thereby updating the position information of the subject anchor box. This process can not only effectively reduce the position jitter caused by detection noise, but also provide a stable predicted position in the image frames where the subject person is not detected.

[0044] In addition to the Kalman Filter, other smoothing techniques can also be used, such as moving average filtering or median filtering. Moving average filtering smooths the position information by taking the average of the positions of the subject anchor box in several consecutive frames, while median filtering takes the median of the positions in several consecutive frames. Both of these embodiments can effectively reduce the jitter of the position information.

[0045] Furthermore, in order to further improve the robustness of tracking, in some embodiments, multi-object tracking algorithms such as SORT (Simple Online and Realtime Tracking) or DeepSORT can also be combined. These algorithms can more accurately track and associate the subject person in consecutive frames by jointly considering the appearance features and motion information of the target, and can maintain a stable tracking effect even in the case of occlusion or loss of the target.

[0046] By tracking and smoothing the subject anchor box, it can be ensured that the position information of the subject anchor box of the subject person is stable and accurate in consecutive frames, which not only improves the robustness of the subject person image segmentation, but also provides high-quality input data for subsequent image segmentation processing, thereby achieving an accurate subject person image segmentation effect.

[0047] Step S3400: Perform image segmentation on the image frames in the live stream according to the position information of the subject anchor box to obtain the human body image of the subject person.

[0048] After the above steps, the subject anchor boxes corresponding to the main characters in each image frame of the live stream are determined. Among them, when the subject anchor boxes are detected from the image frames in an interval frame manner, the subject anchor boxes detected from the image frames can be used directly. For the image frames that have not been detected, the subject anchor boxes obtained by tracking and smoothing the previously detected subject anchor boxes can be used. After determining the subject anchor boxes corresponding to each image frame, the image frames in the live stream can be segmented according to the position information of the subject anchor boxes to obtain the human body images of the main characters.

[0049] Specifically, image segmentation can be achieved in various ways. One embodiment is to use a deep learning-based segmentation model, such as Mask R-CNN, DeepLab series, or U-Net, etc. These models learn the appearance features of people through a large amount of training data and can accurately segment the outlines of people in complex backgrounds. For example, the Mask R-CNN model can not only detect the positions of people but also generate a segmentation mask for each person. The mask is a binary image with the same resolution as the image, where the pixel values of the person area are 1 and the pixel values of the background area are 0. By multiplying the segmentation mask with the original image, the human body image of the main character and the corresponding background image can be extracted.

[0050] In another embodiment, in order to further improve the efficiency and accuracy of segmentation, an image segmentation model based on the attention mechanism can be used. The attention mechanism enables the model to pay more attention to the person areas in the image, thereby improving the segmentation accuracy. For example, the SeaTopFormer architecture can be used. By introducing the attention mechanism, it can better handle the person segmentation problem in multi-person scenarios and maintain lightweight. Through the optimized attention design, it improves the model's ability to recognize the outlines of people, especially the segmentation effect is more significant at the person edges and in complex backgrounds.

[0051] In practical applications, the process of image segmentation based on the position information of the main anchor box can be implemented as follows: First, according to the position information of the main anchor box (such as the upper left corner coordinates, width, and height), an image area containing the main person is cropped from the image frame of the live stream; then, the cropped image area is input into a pre-trained image segmentation model. The image segmentation model generates a segmentation mask through the learned human features. The segmentation mask is a binary image used to distinguish the person from the background. Subsequently, a per-pixel multiplication operation is performed between the segmentation mask and the original image frame to extract the human body image of the main person. At the same time, the background image can also be extracted according to needs for subsequent background processing. Further, the extracted human body image of the main person and the background image are further processed for application, such as applying background blurring, background replacement, or human special effects, etc., to enhance the visual effect of the live broadcast. Finally, taking the webcast scenario as an example, the processed image frame can replace the original image frame in the live stream, so that the live stream contains the corresponding processing effect and is output to the live broadcast room for playback presentation.

[0052] In addition, to improve the robustness and adaptability of segmentation, multi-scale segmentation technology can be combined. The multi-scale segmentation technology can better handle the situation of different human sizes by segmenting the image at different resolutions. For example, the image can be downsampled in a pyramid to generate image copies of different resolutions, then segmented separately on each copy, and finally the segmentation results are merged to obtain a more comprehensive and accurate segmentation effect.

[0053] Accordingly, this application can perform precise image segmentation on the image frames in the live stream based on the position information of the main anchor box and extract the human body image of the main person. The above process not only considers the situation where there may be multiple people in the live stream, but also improves the accuracy and real-time performance of segmentation through various technical means, providing high-quality human body images of the main person for subsequent image processing and applications.

[0054] Through the above embodiments, this application has achieved precise segmentation of the human body image of the main person in the live stream and obtained significant technical advantages, including but not limited to:

[0055] First, by detecting the image area where each person is located in the image frame of the live stream and determining the regional anchor box, this application can accurately identify the positions and ranges of all people, providing a solid foundation for subsequent main person recognition. Combining the preset saliency determination conditions, the main anchor box is screened out from all regional anchor boxes, further ensuring that the most prominent person, that is, the main person, can be accurately identified in a multi-person scenario. This process not only considers various factors such as the position, size, and motion state of the person in the image, but also improves the accuracy and robustness of prominent person recognition by flexibly applying various saliency determination conditions.

[0056] Secondly, in terms of real-time performance, the present application significantly reduces the consumption of computing resources and improves the processing speed by adopting a lightweight object detection model and an interval frame detection strategy, meeting the requirements of application scenarios with high real-time requirements such as webcasting. In addition, by tracking and smoothing the main anchor box through a Kalman filter with low computational overhead, the present application not only stabilizes the position information of the main anchor box, but also further enhances the tracking effect in consecutive frames. Even in scenarios where the person moves fast or the background is complex, accurate segmentation can be maintained, ensuring the real-time performance and stability of the main person image segmentation.

[0057] The technical advantages of the present application are also reflected in its rich scene application capabilities. By accurately positioning and stably tracking the main anchor box and efficiently segmenting the image frames, the present application can be widely applied to various video processing scenarios, such as webcasting, video conferencing, virtual reality, etc. In these scenarios, the present application can not only achieve special effects such as background blurring and background replacement, but also provide high-quality main person images for interactive functions such as person special effects and virtual gift giving, greatly improving the user experience and visual effects.

[0058] Based on any embodiment of the method of the present application, according to the preset saliency determination condition, determining a region anchor box corresponding to the image region indicating the main person from all region anchor boxes as the main anchor box, including:

[0059] Step S3211: Calculate the distance between each region anchor box and the image center of its corresponding image frame;

[0060] Calculating the distance between each region anchor box and the image center of its corresponding image frame can quantify the position information of the region anchor box corresponding to each detected person in the image. The image center usually refers to the geometric center point of the image, and the center point of the region anchor box can be calculated from its upper left corner coordinates, width, and height. By calculating the Euclidean distance between these two center points, the position deviation of each region anchor box relative to the image center can be obtained. For example, in an image frame with a resolution of 1920×1080, if the center point coordinates of a region anchor box are (800, 500) and the image center point coordinates are (960, 540), the distance between the region anchor box and the image center can be calculated through the formula Calculated.

[0061] Step S3212: Calculate the area size of each region anchor box;

[0062] Calculating the area of each regional anchor box can effectively measure the proportion of each person's image in the picture. The area of the regional anchor box can be calculated by multiplying its width and height, and this indicator reflects the space size occupied by the person in the image. For example, if the width of a regional anchor box is 200 pixels and the height is 400 pixels, then its area is 80,000 square pixels. The size of the area can be an important basis for judging the saliency of the person, because in most cases, the main person often occupies a larger proportion of the picture.

[0063] Step S3213: Sort the regional anchor boxes according to the distance and area.

[0064] When sorting all the regional anchor boxes according to the distance and area calculated above, it can be achieved in various ways. For example, in one embodiment, it can be sorted first from largest to smallest according to the area, and for the regional anchor boxes with the same area, then sorted from smallest to largest according to the distance. The rationale for this sorting method is that the regional anchor box with a larger area is more likely to correspond to the main person, and in the case of similar areas, the regional anchor box closer to the center of the image is more likely to be the main person.

[0065] In another embodiment, a weighted method can also be used to comprehensively consider the distance and area factors. For example, set the weight of the area to 0.6 and the weight of the distance to 0.4, and sort the regional anchor boxes by the weighted sum method.

[0066] Step S3214: Select the regional anchor box with the highest ranking as the main anchor box with the saliency determination condition being the highest ranking.

[0067] To determine the most significant regional anchor box as the main anchor box corresponding to the main person, the regional anchor box with the highest ranking according to the distance and area is used as the main anchor box. This saliency determination condition is based on the assumption that in a multi-person scenario, the most significant person is often the one located at the center of the picture and occupying a larger proportion of the picture. Through the above sorting process, the regional anchor box with the highest ranking is considered to be the anchor box of the area where the person most conforms to this saliency condition. For example, in a live broadcast scenario, the anchor usually locates at the center of the picture and occupies a larger proportion of the picture. Therefore, the main anchor box determined through the above steps can accurately indicate the location of the anchor.

[0068] In the above embodiments, the region anchor box of the main character is accurately determined by calculating the distance and area size between each region anchor box and the center of the image, and sorting according to these quantization metrics. This process not only has lightweight operations but also high-efficiency and rapid recognition. First, through the calculation of distance and area, the position and size of each person in the image can be quantitatively evaluated, providing an accurate basis for subsequent sorting. Second, various flexible sorting methods can be adopted, such as sorting by area first and then by distance, or using a weighted sum method to comprehensively consider distance and area. These methods can effectively improve the accuracy of sorting. Finally, taking the region anchor box with the highest ranking as the main anchor box can ensure the accurate recognition of the most prominent person, that is, the main character, in a multi-person scenario. This process has a small amount of computation and a fast recognition speed, and is suitable for scenarios with high real-time requirements such as network live broadcasts, and can provide efficient and accurate technical support for related applications.

[0069] Based on any embodiment of the method of the present application, according to the preset saliency determination condition, a region anchor box corresponding to the image region where the main character is located is determined from all the region anchor boxes as the main anchor box, including:

[0070] Step S3221: Continuously track the region anchor box of each person, and extract the motion feature information of each person, including the motion speed, the change frequency of the motion direction, and the residence time in the image of the live stream;

[0071] By continuously tracking the region anchor box of each person, the motion feature information of each person in the live stream can be extracted. Continuous tracking can be achieved through various technical means. For example, the Kalman filter or the SORT (Simple Online and Realtime Tracking) algorithm can be used. These algorithms can predict and correct the position information of a person according to the position changes of the person in consecutive frames, so as to achieve continuous tracking of the person. Through continuous tracking, the motion trajectory of each person can be obtained, and then its motion speed and the change frequency of the motion direction can be calculated. For example, the motion speed can be obtained by calculating the change rate of the person's position in consecutive frames, and the change frequency of the motion direction can be determined by analyzing the number of changes in the person's motion direction. In addition, the residence time of each person in the live stream image can also be counted, that is, the length of time the person remains stationary in a certain area. These motion feature information together constitute the dynamic behavior characteristics of the person.

[0072] Step S3222: According to the motion feature information of each person, apply the preset activity score model to determine the corresponding activity score of each person;

[0073] The motion feature information of each person can be obtained, and a preset activity score model can be applied to determine the corresponding activity score for each person. The activity score model can be designed according to different application scenarios and requirements. In one embodiment, features such as motion speed, motion direction change frequency, and stay time can be weighted and summed to obtain a comprehensive activity score. The weight assignment can be adjusted according to the degree of emphasis on different features in actual applications. For example, in a live broadcast scenario mainly featuring dynamic performances, the weights of motion speed and motion direction change frequency can be set relatively high, while in a live broadcast scenario mainly for explanations, the weight of stay time can be appropriately increased.

[0074] In another embodiment, a corresponding machine learning model or deep learning model can be pre-trained and used as the activity score model. Using the motion feature information as the input of the model, the activity score determined after the model's inference can be obtained.

[0075] It can be seen that with the help of the activity score model, it can flexibly adapt to different live broadcast scenarios and accurately reflect the activity level of each person.

[0076] Step S3223: Using the highest activity score as the significance determination condition, select the region anchor box with the highest score as the main anchor box.

[0077] After determining the activity score, the highest activity score can be used as the significance determination condition, and the region anchor box with the highest score can be selected as the main anchor box. This determination condition is based on the assumption that in a multi-person scenario, the most active person is often the main person in the live broadcast. The activity score calculated through the above steps can effectively reflect the activity state of the person. Therefore, the region anchor box with the highest score is considered to be the anchor box of the region where the person most conforms to the significance condition. For example, in a live broadcast scenario, the host may move frequently during the explanation, with a relatively high motion speed and motion direction change frequency, and will stay still for a while when explaining a key point. These behavioral characteristics make the host's activity score higher than that of other people, so that the host can be accurately identified as the main person.

[0078] Embodiments of the present application determine the main character by introducing dynamic behavior features, significantly improving the accuracy and adaptability of main character recognition. First, by continuously tracking the regional anchor boxes of each person and extracting motion feature information such as their movement speed, change frequency of movement direction, and residence time, the dynamic behavior of the person can be comprehensively reflected. Second, applying a preset activity score model to calculate the activity score of each person based on this motion feature information, this model can be flexibly designed, such as weighted summation or using a machine learning model, to meet the requirements of different live broadcast scenarios. Finally, taking the regional anchor box with the highest activity score as the main anchor box ensures that the most active main character can be accurately identified in a multi-person scenario. This process not only considers the dynamic behavior of the person but also improves the accuracy and real-time performance of main character recognition through a flexible scoring model and significance determination conditions, is applicable to various complex live broadcast scenarios, and provides high-quality position information of the main character for subsequent image segmentation processing.

[0079] Based on any embodiment of the method of the present application, tracking and smoothing processing is performed on the main anchor box, including:

[0080] Step S3310: Predict the position information of the main anchor box based on the Kalman filter;

[0081] For the anchor box set B composed of multiple regional anchor boxes detected from the image frame, assuming that the distance between each regional anchor box and the center of the screen and the area size of the regional anchor box are used as the significance determination conditions, select the regional anchor box with the most prominent significance, that is, the main anchor box B m As the target person, accordingly, the regional anchor box B corresponding to the main character with the most prominent significance can be obtained based on the Kalman filter m The corresponding prediction result and the result after smoothing processing.

[0082] The Kalman filter is a recursive filter based on the theory of linear dynamic systems and is widely used in the field of target tracking. In the present application, the Kalman filter models the historical position information of the main anchor box to predict its position in the next frame.

[0083] Before making a prediction, first perform the initialization settings required for the prediction. First, based on the main anchor box obtained at time t - 1, the state vector x corresponding to this moment can be calculated t-1 :

[0084]

[0085] where (x, y) are the coordinates corresponding to the center point of the main anchor box, w and h are the width and height corresponding to the main anchor box respectively, v x , v y , v w and vh is the velocity component corresponding to the above variables. These parameters together describe the state of the subject anchor frame at the current moment.

[0086] The following state transfer matrix F is defined to represent the temporal transformation of the state vector:

[0087]

[0088] At the same time, the measurement matrix H is defined to map the state vector to the measurement space:

[0089]

[0090] In addition, the noise covariance matrix Q = c is defined in advance Q I8. Measurement noise covariance matrix R = c R I4, the state covariance matrix P at the initial moment t-1 =c P *I8, where c Q , c R and c P are the coefficients corresponding to each matrix, which can be set to 0.01, 10, 0.1, and I respectively. n represents the identity matrix of size n × n. The process noise covariance matrix Q is used to simulate the uncertainty of model prediction, the measurement noise covariance matrix R is used to simulate the uncertainty of measurement, and the state covariance matrix P is used to simulate the uncertainty of predicted state.

[0091] In the prediction stage, the Kalman filter uses the state transfer matrix F to transform the state vector x at the previous moment t-1 Mapped to the state vector x at the next moment t|t-1 The state transfer matrix F describes the change law of the state vector in the time series and reflects the movement trend of the target. t|t-1 =Fx t-1 ,The Kalman filter can predict the position information of the subject anchor box in the next frame, providing a priori estimation for subsequent measurement updates.

[0092] At the same time, the Kalman filter also needs to predict the state covariance matrix P at the next moment t|t-1 , to quantify the uncertainty of the prediction. The state covariance matrix P t-1 represents the uncertainty of the state estimation at the previous moment, while the process noise covariance matrix Q reflects the uncertainty of the model prediction. t|t-1 =FP t-1 F T +Q, the Kalman filter can comprehensively consider the uncertainty in the state transfer process and the uncertainty of the model prediction, thereby obtaining an uncertainty description of the state estimation at the next moment.

[0093] It can be seen that in this step, the Kalman filter is used to predict the position information of the main anchor box. This not only takes into account the motion trend of the target, but also quantifies the uncertainty of the prediction through the state covariance matrix. This process provides prior information for the subsequent measurement update and is an important basis for realizing the stable tracking of the main anchor box.

[0094] Step S3320: Correct the predicted position information according to the position information of the main anchor box determined in the current image frame to obtain the predicted position information;

[0095] Under the framework of the Kalman filter, the key to correcting the predicted position information lies in calculating the measurement residual and the Kalman gain. The measurement residual reflects the difference between the actual measurement value and the predicted value. Specifically, assume that the actual position information of the main anchor box in the current image frame is z t , and the position information predicted by the Kalman filter is x t|t-1 , then the measurement residual y t can be expressed as:

[0096] y t = z t - Hx t|t-1

[0097] where H is the measurement matrix, which is used to map the predicted state vector to the measurement space. The measurement residual y t describes the deviation between the actual measurement value and the predicted value and is an important basis for correcting the predicted position information.

[0098] Next, calculate the Kalman gain K t , which is a key parameter for weighing the weights of the predicted value and the measurement value in the correction process. The calculation formula of the Kalman gain is:

[0099] K t = P t|t-1 H T (HP t|t-1 H T + R) -1

[0100] where P t|t-1 is the state covariance matrix obtained in the prediction stage, which represents the uncertainty of the predicted state; R is the measurement noise covariance matrix, which reflects the uncertainty of the measurement value. The Kalman gain K t determines the relative importance of the predicted value and the measurement value in the correction process by comprehensively considering the uncertainties of prediction and measurement.

[0101] Finally, use the Kalman gain and the measurement residual to correct the predicted position information to obtain the corrected predicted position information x t|t:

[0102] x t|t = x t|t-1 + K t y t

[0103] This correction process combines the actual measurement value with the predicted value, and dynamically adjusts the weights of the two in the correction result through the Kalman gain, so as to obtain a more accurate estimation of the target position. The corrected predicted position information x t not only considers the motion trend of the target, but also combines the actual measurement information of the current frame, so it can more accurately reflect the actual position of the target in the current frame.

[0104] Through the above correction process, the Kalman filter can effectively reduce the prediction error and improve the accuracy and stability of target tracking. This process has wide application value in the field of target tracking, especially when dealing with measurement data with noise and uncertainty, it can significantly improve the performance of the system.

[0105] Step S3330: According to the corrected predicted position information, update the state parameters of the Kalman filter to reflect the latest position of the main anchor box;

[0106] After obtaining the corrected predicted position information, the state parameters of the Kalman filter can be updated to reflect the latest position of the main anchor box, so as to incorporate the corrected predicted position information into the state of the filter, thereby providing more accurate prior information for subsequent prediction and correction.

[0107] In the update stage of the Kalman filter, first, the state vector x t|t , which has been completed in step S3320. Next, the state covariance matrix P t|t needs to be updated to reflect the uncertainty of the corrected state estimate. The update formula for the state covariance matrix is:

[0108]

[0109] where I is the identity matrix, K t is the Kalman gain, H is the measurement matrix, and P t|t-1 is the state covariance matrix obtained in the prediction stage. This formula updates the state covariance matrix by subtracting the influence of the product of the Kalman gain and the measurement matrix on the predicted covariance matrix. The updated state covariance matrix P t|t more accurately reflects the uncertainty of the corrected state estimate.

[0110] The purpose of updating the state covariance matrix is to more accurately estimate the uncertainty of the target position in the subsequent prediction stage. For example, when the Kalman gain is large, it indicates that the measurement value contributes more to the correction process. At this time, the updated state covariance matrix will decrease, indicating more confidence in the corrected state estimate; conversely, when the Kalman gain is small, it means that the predicted value is more reliable, and the change in the state covariance matrix is relatively small.

[0111] By updating the state covariance matrix, the Kalman filter can more accurately reflect the uncertainty of the target position in the subsequent prediction stage, thereby improving the accuracy and stability of target tracking. This process not only considers the motion trend of the target but also combines the actual measurement information of the current frame, enabling the Kalman filter to dynamically adjust the estimation accuracy of the target position.

[0112] In practical applications, the updated state parameters (including the state vector and the state covariance matrix) will be used in the prediction stage of the next frame to achieve continuous tracking of the target position. For example, in a live broadcast scenario, if the host moves quickly in the frame, the Kalman filter can accurately track the position change of the host by continuously updating the state parameters, and can maintain a stable tracking effect even in the presence of noise or occlusion.

[0113] Thus, by updating the state parameters of the Kalman filter, it is ensured that the filter can dynamically reflect the latest information and uncertainty of the target position, providing a more accurate basis for subsequent prediction and correction. This process is one of the key links for the Kalman filter to achieve target tracking and is of great significance for improving the accuracy and stability of target tracking.

[0114] Step S3340: In the image frame where detection has not been performed, use the predicted position information as the position of the main anchor box.

[0115] An important role of the Kalman filter in target tracking in this application is to ensure that when the main anchor box corresponding to the main person cannot be directly detected in some image frames, or in the case of an embodiment using interval frame detection to determine the main anchor box, for the situation where detection has not been performed or the main anchor box has not been detected, stable target position information can still be provided through the prediction of the Kalman filter to stabilize the position information of the main anchor box, thereby maintaining the continuity of target tracking.

[0116] In practical applications, especially when computing resources are limited or the detection algorithm fails to successfully detect the target in every frame, there may be cases where the main anchor box is not detected in some frames. At this time, the prediction function of the Kalman filter becomes particularly important. By making predictions based on the current state vector and the state transition matrix in step S3310, and correcting the prediction results and updating the state parameters in steps S3320 and S3330, the Kalman filter can provide a reliable predicted position in frames where the target is not detected.

[0117] Specifically, when a frame is not detected or the detection fails, the Kalman filter uses the previously updated state vector x t|t and the state covariance matrix P t|t to make a prediction and obtain the state vector x t+1|t at the next moment. This predicted position information can be directly used as the position of the main anchor box for this frame, ensuring the continuity of target tracking. For example, assume that the main anchor box is not detected in a certain frame. The Kalman filter makes a prediction through the state transition matrix F and the previous state vector x t|t to obtain the predicted position information:

[0118] x t+1|t = Fx t|t

[0119] This predicted position information will be used as the position of the main anchor box for this frame in subsequent image segmentation or other processing. In this way, even if the target main anchor box is not detected or cannot be directly detected in some frames, the Kalman filter can still provide a reasonable estimate of the target position, thus ensuring the stability and continuity of target tracking.

[0120] In addition, this prediction mechanism can also effectively reduce the problem of target loss caused by detection noise or instability of the detection algorithm. For example, in a live broadcast scenario, if the anchor moves quickly or is partially blocked, the detection algorithm may not be able to accurately detect the main anchor box in some frames. At this time, the prediction function of the Kalman filter can fill these detection gaps and ensure the stability of the position information of the main anchor box in consecutive frames.

[0121] In the above embodiments, the Kalman filter is used to track and smooth the main anchor box, significantly improving the stability and accuracy of object tracking. First, through the prediction and correction mechanism, the Kalman filter can provide reliable predicted positions in image frames where the main anchor box is not detected and where the main anchor box detection is not performed, effectively avoiding the problem of missed detection and ensuring the continuity of object tracking. Second, through the detection of interval frames and the prediction function of the Kalman filter, the need to detect each frame is reduced, thus significantly saving the overhead of system operation resources and improving the processing efficiency. In addition, the lightweight characteristics and efficient mathematical model of the Kalman filter enable it to be widely adapted to devices with different performances, including mobile devices and low-end PC hosts with limited computing resources, ensuring efficient operation on various devices. These technical advantages not only improve the robustness and real-time performance of the main body person image segmentation, but also provide an efficient and stable technical solution for related application scenarios.

[0122] Based on any embodiment of the method of the present application, image segmentation is performed on the image frames in the live stream according to the position information of the main anchor box to obtain a human body image of the main body person, including:

[0123] Step S3410: Crop the corresponding image area in the image frame according to the position information of the main anchor box to obtain a portrait area map;

[0124] The position information of the main anchor box includes its upper left corner coordinates, width, and height, and these parameters can uniquely determine the position and range of the main body person in the image. Through these parameters, a rectangular area can be cropped from the original image frame, and this area contains the complete image of the main body person. For example, assuming that the position information of the main anchor box is the upper left corner coordinates (x, y), width w, and height h, then a rectangular area of size w*h can be cropped from the original image frame, and the upper left corner of this area is located at the position (x, y).

[0125] The cropping process can be implemented through cropping functions in an image processing library or framework. These functions usually accept the coordinates and dimensions of the cropping area as input parameters and output the cropped image area. For example, in the Python image processing library Pillow, the crop() function can be used to implement the cropping operation, and its input parameters are the upper left and lower right coordinates of the cropping area.

[0126] The cropped portrait area image is a reduced image area, called the portrait area image, which only contains the main person and some background information around them. This cropped image area will be used as the input for the subsequent image segmentation model, thereby reducing the size of the image processed by the model and improving the processing speed. For example, if the resolution of the original image frame is 1920×1080 pixels and the size of the main anchor box is 300×400 pixels, then the resolution of the cropped portrait area image will be 300×400 pixels, significantly reducing the amount of image data.

[0127] In addition, the cropping process can also be combined with image preprocessing steps, such as scaling, normalization, etc., to further optimize the input of the image segmentation model. For example, if the image segmentation model requires the input image to be of a fixed size of 256×256 pixels, then the cropped portrait area image can be scaled after cropping to meet the input requirements of the model.

[0128] Step S3420: Input the cropped portrait area image into the attention mechanism-based image segmentation model to generate a segmentation mask for the human body image of the main person through the image segmentation model;

[0129] Input the cropped portrait area image into the attention mechanism-based image segmentation model to generate a segmentation mask for the human body image of the main person through this model, so as to accurately extract the contour and boundary of the main person from the cropped portrait area image according to the segmentation mask, providing high-quality segmentation results for subsequent image editing and applications.

[0130] The attention mechanism-based image segmentation model is a deep learning model that enhances the model's attention to the target area by introducing the attention mechanism, thereby improving the accuracy and efficiency of segmentation. The attention mechanism enables the model to automatically focus on the most important part of the image, that is, the main person, while ignoring the irrelevant information in the background. For example, the SeaTopFormer architecture can be used, which is a variant model that combines SeaFormer and TopFormer. Through optimized attention design, it improves the model's ability to recognize the human contour, especially the segmentation effect is more significant at the human edges and in complex backgrounds.

[0131] In specific implementation, the image segmentation model usually needs to be trained with a large amount of training data to learn the appearance features and shape information of the main person. The training data usually includes images with segmentation mask annotations, and these masks clearly identify the boundaries of the main person. Through training, the model can learn how to distinguish the main person from the background in the input image and generate the corresponding segmentation mask.

[0132] After inputting the cropped portrait area image into the segmentation model, the model will gradually extract image features and generate a segmentation mask through a series of convolutional layers, attention modules, and decoder structures. For example, the SeaTopFormer model first obtains the global features of the image through the feature extraction module of SeaFormer, then uses the attention mechanism of TopFormer to optimize the feature map, and finally restores the resolution of the segmentation mask through the decoder to output a segmentation mask with the same size as the portrait area image.

[0133] Step S3430: Extract the human body image and background image of the main person from the image frame according to the segmentation mask;

[0134] The segmentation mask is a binary image with the same resolution as the input image, where the pixel value of the main person area is 1 and the pixel value of the background area is 0. By multiplying the segmentation mask with the original image frame pixel by pixel, the human body image of the main person can be extracted. Specifically, for each pixel point in the original image frame, if the corresponding segmentation mask pixel value is 1, the value of that pixel point is retained; if the segmentation mask pixel value is 0, the value of that pixel point is set to 0. In this way, an image containing only the main person, that is, the human body image of the main person, can be obtained.

[0135] At the same time, by multiplying the inverted segmentation mask (i.e., 1 becomes 0 and 0 becomes 1) with the original image frame pixel by pixel, the background image can be extracted. This process is also based on the binary nature of the segmentation mask, but the operation is opposite. For each pixel point in the original image frame, if the corresponding segmentation mask pixel value is 0, the value of that pixel point is retained; if the segmentation mask pixel value is 1, the value of that pixel point is set to 0. In this way, an image containing only the background, that is, the background image, can be obtained.

[0136] For example, assume that the resolution of the original image frame is 1920×1080 pixels, and the segmentation mask is also a binary image of 1920×1080 pixels. Through the above operations, two images can be obtained: one is the human body image containing only the main person, and the other is the background image containing only the background. These two images can be used for subsequent image editing processes, such as background blurring, background replacement, or adding special effects, etc.

[0137] In practical applications, the process of extracting the main person and background images can be implemented through per-pixel operation functions in an image processing library or framework. For example, in the Python image processing library Pillow, the Image.point() function can be used to perform per-pixel operations on the image and extract the main person and background images in combination with the segmentation mask.

[0138] Step S3440: After performing image editing processing on the human body image or background image, transmit the live stream to the live broadcast room for playback.

[0139] Performing image editing processing on the extracted main character image or background image and transmitting the processed image frames to the live broadcast room for playback can achieve the final application link of this application. By using image editing technology to enhance the visual effect of the live broadcast, it can meet the requirements of different application scenarios.

[0140] Image editing processing can include various operations, such as background blurring, background replacement, adding special effects, or performing beauty treatments, etc. These operations are based on the extracted main character image and background image and are implemented through different image processing algorithms. For example, background blurring can be achieved by applying the Gaussian blur algorithm to the background image to make the background blurred, thereby highlighting the main character. Background replacement can replace the extracted background image with a preset virtual background, such as a virtual studio scene or a specific background picture, so as to provide a more creative visual effect for the live broadcast.

[0141] In addition, beauty treatments can also be performed on the main character image, such as skin smoothing, whitening, face slimming, etc., to enhance the visual effect of the anchor. These beauty treatments are usually implemented through image processing algorithms. For example, the bilateral filtering algorithm is used for skin smoothing, or the face slimming effect is achieved through affine transformation. These treatments can not only improve the visual quality of the live broadcast but also meet the personalized needs of different anchors and audiences.

[0142] After completing the image editing processing, the processed image frames need to be transmitted to the live broadcast room for playback. For example, in a network live broadcast scenario, the main character image of the anchor is accurately extracted and background blurred. The processed image frames are encoded into a video stream in H.264 format and transmitted to the viewer's terminal device through a streaming media server. When watching the live broadcast, the viewer can see the anchor being highlighted and the background being blurred, thus obtaining a better visual experience.

[0143] Through a series of optimized image processing steps in the above embodiments of the present application, accurate segmentation and efficient editing of the main characters in the live stream are achieved. First, by cropping the region corresponding to the main anchor box in the image frame, the amount of image data is significantly reduced, and the processing speed is improved. Second, an image segmentation model based on the attention mechanism is used to accurately extract the contours and boundaries of the main characters, enhancing the accuracy and robustness of the segmentation. In addition, the main characters and background images are extracted through the segmentation mask, providing high-quality materials for subsequent image editing. Finally, the main character image or background image is edited (such as background blurring, background replacement, beauty, etc.) and efficiently transmitted to the live broadcast room for playback, significantly improving the visual effect and user experience of the live broadcast. This process not only optimizes the efficiency of image segmentation and editing, but also meets the diverse application scenario requirements through flexible image processing functions, is applicable to various devices and network conditions, and has wide applicability and high efficiency.

[0144] Based on any embodiment of the method of the present application, the method further includes:

[0145] Step S31, call the target detection model that has been pruned and channel-compressed, and determine the region anchor box corresponding to each person from the image frames of the live stream at a preset interval. Among them, for the image frames that have not been detected, the main anchor box thereof adopts the main anchor box determined after corresponding tracking and smoothing processing;

[0146] Calling the target detection model that has been pruned and channel-compressed to determine the region anchor box corresponding to each person from the image frames of the live stream at a preset interval can optimize the efficiency of target detection, enable it to run efficiently on low-performance devices, and at the same time reduce the consumption of computing resources.

[0147] Specifically, the target detection model in this embodiment adopts a lightweight design model, such as the NanoDet model. NanoDet is a lightweight single-stage target detection model optimized for mobile devices and low-performance hardware. It significantly reduces the number of model parameters and computational complexity while maintaining high detection accuracy through model pruning and channel compression techniques. For example, the number of parameters of the NanoDet model is only 1.01M, and it can achieve fast detection on low-performance devices through the SimOTA label dynamic allocation strategy of the optimal transport theory for training.

[0148] During the implementation process, the object detection model runs according to a preset interval frame calling strategy. This means that not every frame in the live stream is detected, but rather detection is performed every several frames. For example, it can be set to detect every 3 frames or 5 frames, and the specific interval is adjusted according to device performance and real-time requirements. This interval frame detection strategy further reduces the computational load, and at the same time, the Kalman filtering algorithm in the subsequent steps is used to predict and track the main body anchor boxes of the undetected frames, ensuring the continuity of the main body person tracking.

[0149] For the image frames that are not detected, their main body anchor boxes are the main body anchor boxes determined after tracking and smoothing. This mechanism combines the efficiency of the object detection model and the stability of the Kalman filtering algorithm, ensuring stable tracking and accurate segmentation of the main body person even on low-performance devices.

[0150] Step S32: Apply the Kalman filtering algorithm to track and smooth the main body anchor boxes;

[0151] Applying the Kalman filtering algorithm to track and smooth the main body anchor boxes is also a key technical means to ensure the stable tracking of the main body person on low-performance devices. When the object detection model determines the main body anchor boxes from the image frames of the live stream according to the preset interval, for the image frames that are not detected, the position information of their main body anchor boxes needs to be predicted and corrected by the Kalman filtering algorithm to ensure the continuity and stability of the main body person tracking.

[0152] Referring to the specific algorithm demonstration revealed above, the Kalman filtering algorithm uses the motion information of the object in consecutive frames to predict the position of the main body anchor boxes. In the frames where the main body anchor boxes are detected, the Kalman filter updates its internal state according to the detection results, while in the undetected frames, the Kalman filter makes predictions based on the state of the previous frame and the motion model, thereby providing stable position information of the main body anchor boxes. This prediction and correction mechanism enables the position information of the main body person to remain continuous and accurate even when the detection interval is large.

[0153] For example, assume that the object detection model detects every 3 frames. In the undetected frames, the Kalman filter will predict the position of the main body anchor box in the current frame based on the position of the main body anchor box detected last time and its motion trend. When the main body anchor box is detected again, the Kalman filter will correct the predicted position according to the new detection results, thereby ensuring that the position information of the main body anchor box always accurately reflects the actual position of the main body person.

[0154] By applying the Kalman filtering algorithm, the stable tracking of the main character can be efficiently achieved on low-performance devices. This process not only reduces the need to detect each frame, thus reducing the consumption of computing resources, but also ensures the continuity and accuracy of the main character tracking through prediction and correction mechanisms. For example, on mobile devices or low-end PC hosts, this optimization strategy enables the main character image segmentation technology to operate stably in application scenarios with high real-time requirements, such as webcasting and video conferencing.

[0155] Step S33: Use an image segmentation model based on the attention mechanism to perform image segmentation on the image frames in the live stream.

[0156] As revealed above, the image segmentation model based on the attention mechanism is a deep learning model that enhances the model's attention to the main character by introducing the attention mechanism, thereby improving the accuracy and efficiency of segmentation. The attention mechanism enables the model to automatically focus on the most important part of the image, that is, the main character, while ignoring the irrelevant information in the background. For example, the SeaTopFormer architecture can be used. This is a variant model that combines SeaFormer and TopFormer. Through optimized attention design, it improves the model's ability to recognize the human contour, especially the segmentation effect is more significant at the human edge and in complex backgrounds.

[0157] In specific implementation, the image segmentation model usually needs to be trained with a large amount of training data to learn the appearance features and shape information of the main character. The training data usually includes images with segmentation mask annotations, and these masks clearly identify the boundaries of the main character. Through training, the model can learn how to distinguish the main character from the background in the input image and generate the corresponding segmentation mask.

[0158] After inputting the cropped portrait area map into the segmentation model, the model will gradually extract image features and generate a segmentation mask through a series of convolutional layers, attention modules, and decoder structures. For example, the SeaTopFormer model first obtains the global features of the image through the feature extraction module of SeaFormer, then uses the attention mechanism of TopFormer to optimize the feature map, and finally restores the resolution of the segmentation mask through the decoder to output a segmentation mask with the same size as the portrait area map.

[0159] To meet the requirements of low-performance devices, the image segmentation model can be optimized through pruning and channel compression techniques to reduce the number of model parameters and computational complexity. For example, the number of parameters of the SeaTopFormer model can be significantly reduced after optimization while maintaining a high segmentation accuracy. This optimization strategy enables the model to operate efficiently on low-performance devices and meet application scenarios with high real-time requirements, such as webcasting and video conferencing.

[0160] Through various optimization measures, the above embodiments ensure that the subject person image segmentation technology can operate efficiently on low-performance devices while maintaining accurate recognition and smooth performance. First, a lightweight object detection model is adopted and combined with an interval frame detection strategy, significantly reducing the computational amount and resource consumption. At the same time, the Kalman filter algorithm is used to predict and track the undetected frames, ensuring the continuity and stability of the subject person tracking. Second, the image segmentation model based on the attention mechanism further optimizes the segmentation accuracy, especially performing well in the processing of complex backgrounds and human edges. Finally, through model pruning and channel compression technologies, the number of parameters and computational complexity of the image segmentation model are reduced, enabling it to run quickly on low-performance devices. The superposition of these optimization measures not only ensures the accurate recognition of the subject person but also achieves efficient and smooth operation in application scenarios with high real-time requirements such as webcasting and video conferencing, demonstrating significant technical advantages.

[0161] Based on any embodiment of the method of the present application, before performing image segmentation on the image frames in the live stream using an image segmentation model based on the attention mechanism, it includes:

[0162] Step S2100: Obtain the foreground person real-scene image and the background person real-scene image, their corresponding segmentation masks, and the person depth information, and extract the corresponding foreground person image and background person image according to the segmentation mask.

[0163] The foreground person real-scene image and the background person real-scene image respectively refer to the images containing the subject person and the background person taken in the actual scene. These images can be obtained through professional photographic equipment or ordinary cameras to ensure that the image quality meets the training requirements. The segmentation mask refers to the binary image corresponding to these real-scene images. The segmentation mask can be generated manually or automatically by existing segmentation algorithms, and is used to clearly identify the boundaries of the foreground person and the background person.

[0164] The person depth information refers to the depth value of each pixel point relative to the camera, usually obtained through a depth camera or stereo vision technology. The depth information is crucial for determining the layer order of the foreground person and the background person in the image synthesis space because it can reflect the relative distance between the persons. Pixel points with smaller depth values usually represent the foreground person closer to the camera, while pixel points with larger depth values represent the background person.

[0165] According to the segmentation mask, the foreground person image and the background person image can be extracted from the real scene image. The extraction process can be achieved through pixel-by-pixel operations. For example, for the foreground person image, the areas with pixel values of 1 in the segmentation mask are retained, and the remaining areas are set to transparent or the background color; for the background person image, the areas with pixel values of 0 in the segmentation mask are retained. This extraction method ensures the independence of the foreground and background person images and provides a basis for subsequent image synthesis.

[0166] Step S2200: Determine the layer order of the foreground person image and the background person image in the image synthesis space based on the person depth information, and synthesize an initial image sample that includes both the foreground person and the background person;

[0167] When making the image samples required for training the image segmentation model, in the virtual image synthesis space jointly defined by the corresponding person depth information and the image plane, determine the layer order of the foreground person and the background person according to their person depth information. Based on the distance of their depth information from each other, place the foreground person image as the top layer above the corresponding bottom layer of the background person image to form an initial image sample. For example, assume that the depth value range of the foreground person is characterized as 1 - 5 meters, and the depth value range of the background person is characterized as 6 - 10 meters. Then, when synthesizing the image, the foreground person image will be placed above the background person image to simulate the situation where the foreground person blocks the background person in the real scene.

[0168] Step S2300: Based on the step-by-step adjustment of the relative position relationship between the foreground person and the background person in the image synthesis space, synthesize multiple derivative image samples to make the initial image sample and each derivative image sample exhibit a temporal gradient characteristic;

[0169] Based on the initial image sample, it is possible to perform step-by-step adjustments based on the relative position relationship between the foreground person and the background person in the image synthesis space to simulate the relative position changes of the foreground person and the background person at different time steps, so as to synthesize multiple derivative image samples. The resulting derivative image samples not only enrich the diversity of the training data but also provide training samples closer to the real scene for the image segmentation model, thereby accelerating the convergence of the model and improving its generalization ability.

[0170] Specifically, the objects of step-by-step adjustment can include various operations such as translation, rotation, and scaling. For example, by performing a translation operation on the foreground person image in the image synthesis space, the position changes of the person at different time steps can be simulated; by performing a rotation operation, the orientation changes of the person can be simulated; by performing a scaling operation, the distance changes between the person and the camera can be simulated. The combination of these operations can generate a series of image samples with visual continuity and gradual change, thus more realistically reflecting the dynamic behavior of the person in the actual scene.

[0171] Taking the translation operation as an example, assume that the depth value of the foreground human image at the initial position is 3 meters, and the depth value of the background human image is 7 meters. After synthesizing the initial image sample, the position of the foreground human image can be gradually adjusted so that it moves left or right in the synthesis space while keeping the position of the background human image unchanged. In this way, multiple derivative image samples can be generated, and each sample reflects the position change of the foreground human at different time steps.

[0172] Similarly, the rotation operation can be achieved by adjusting the orientation angle of the foreground human image. For example, starting from the initial front-facing orientation, gradually rotate the foreground human image by a certain angle (such as 5 degrees each time) to generate a series of derivative image samples with different orientations. These samples can simulate the turning actions of the human in the actual scene and provide richer training data for the image segmentation model.

[0173] The scaling operation can simulate the change in the distance between the human and the camera by changing the size of the foreground human image. For example, starting from the initial normal size, gradually enlarge or reduce the foreground human image to generate a series of derivative image samples with different sizes. These samples can simulate the dynamic process of the human approaching or moving away from the camera in the actual scene.

[0174] Through these stepwise adjustment operations, a series of derivative image samples with visual continuity and gradual change can be generated. These samples not only enrich the diversity of the training data but also provide training samples closer to the real scene for the image segmentation model. The derivative image samples generated by translation, rotation, and scaling operations can simulate various dynamic behaviors of the human in the actual scene, thereby improving the adaptability and generalization ability of the image segmentation model to human dynamic changes.

[0175] Step S2400: Use the synthesized image samples to train the image segmentation model, and adopt the corresponding foreground human image as the supervision sample to train the image segmentation model to the convergence state.

[0176] The initial image sample and its derivative image samples are used to train the image segmentation model. By adopting the corresponding foreground human image as the supervision sample, the model can learn how to accurately segment the main human from the complex background and maintain a high segmentation accuracy even when the position, orientation, and size of the human change. This training sample synthesis method based on the time-series gradual change characteristics not only improves the training efficiency of the image segmentation model but also significantly enhances the robustness and accuracy of the model in practical applications. After training the image segmentation model to the convergence state, it can be used to extract the human image of the main human in this application.

[0177] The above embodiments construct image samples with time-varying characteristics through an efficient synthesis method for model training, significantly improving the ability of the image segmentation model to accurately segment the main characters in the live stream. First, by obtaining the real-scene images of the foreground and background characters, their segmentation masks, and depth information, the foreground and background character images can be accurately extracted. Then, based on the character depth information, the layer order is determined, and an initial image sample containing the foreground and background characters is synthesized to simulate the occlusion relationship in the real scene. Further, through stepwise adjustment operations such as translation, rotation, and scaling, multiple derivative image samples are generated. These samples not only enrich the diversity of the training data but also simulate the dynamic behaviors of the characters, enabling the model to learn the changes of the main characters under different time and space conditions. Finally, using these synthesized samples for training, the model can converge quickly and have stronger generalization ability, so as to efficiently and accurately segment the human body image of the main character in the live stream and meet the requirements of application scenarios with high real-time requirements.

[0178] Please refer to Figure 3 , a main character image segmentation device provided according to an aspect of the present application includes an anchor box detection module 3100, a main body determination module 3200, a tracking and smoothing module 3300, and a segmentation processing module 3400. Among them, the anchor box detection module 3100 is configured to detect the image area where each character is located in the image frame of the live stream and determine the regional anchor box corresponding to the image area where each character is located; the main body determination module 3200 is configured to determine, according to a preset saliency determination condition, the regional anchor box corresponding to the image area indicating the main character among all the regional anchor boxes as the main body anchor box; the tracking and smoothing module 3300 is configured to perform tracking and smoothing processing on the main body anchor box to stabilize the position information of the main body anchor box; the segmentation processing module 3400 is configured to perform image segmentation on the image frame in the live stream according to the position information of the main body anchor box to obtain the human body image of the main character.

[0179] Based on any embodiment of the device of the present application, the main body determination module 3200 includes: a distance calculation module configured to calculate the distance between each regional anchor box and the image center of the image frame where it is located; an area calculation module configured to calculate the area size of each regional anchor box; a joint sorting module configured to sort the regional anchor boxes according to the distance and area; and an order screening module configured to select the regional anchor box with the top ranking as the main body anchor box with the top ranking as the saliency determination condition.

[0180] Based on any embodiment of the device in the present application, the main body determination module 3200 includes: an anchor box tracking module configured to continuously track the regional anchor boxes of each person, and extract the motion feature information of each person, including the motion speed, the change frequency of the motion direction, and the residence time in the image of the live stream; an activity scoring module configured to determine the corresponding activity score of each person by applying a preset activity scoring model according to the motion feature information of each person; a high score screening module configured to select the regional anchor box with the highest score as the main body anchor box with the highest activity score as the significance determination condition.

[0181] Based on any embodiment of the device in the present application, the tracking and processing module includes: a position prediction module configured to predict the position information of the main body anchor box based on a Kalman filter; a prediction correction module configured to correct the predicted position information according to the position information of the main body anchor box determined in the current image frame to obtain predicted position information; a parameter update module configured to update the state parameters of the Kalman filter according to the corrected predicted position information to reflect the latest position of the main body anchor box; a prediction application module configured to use the predicted position information as the position of the main body anchor box in the image frame where detection has not been performed.

[0182] Based on any embodiment of the device in the present application, the segmentation processing module 3400 includes: a portrait cropping module configured to crop the corresponding image area in the image frame according to the position information of the main body anchor box to obtain a portrait area map; a segmentation execution module configured to input the cropped portrait area map into an image segmentation model based on an attention mechanism, and generate a segmentation mask of the human body image of the main body person through the image segmentation model; a separation and extraction module configured to extract the human body image and the background image of the main body person from the image frame according to the segmentation mask; an application processing module configured to perform image editing processing on the human body image or the background image, and then transmit the live stream to the live broadcast room for playback.

[0183] Based on any embodiment of the device in the present application, the anchor box detection module 3100 is further configured to call a target detection model that has been pruned and channel compressed to determine the regional anchor box corresponding to each person from the image frames of the live stream at a preset interval. For the image frames where detection has not been performed, the main body anchor box thereof is the main body anchor box determined after corresponding tracking and smoothing processing; the tracking and smoothing module 3300 is further configured to perform tracking and smoothing processing on the main body anchor box by applying a Kalman filtering algorithm; the segmentation processing module 3400 is further configured to perform image segmentation on the image frames in the live stream by using an image segmentation model based on an attention mechanism.

[0184] Based on any embodiment of the device in the present application, prior to the segmentation processing module 3400, the device further includes: a material processing module configured to obtain a foreground person real scene image and a background person real scene image, their corresponding segmentation masks, and person depth information, and extract the corresponding foreground person image and background person image according to the segmentation masks; a base sample synthesis module configured to determine the layer order of the foreground person image and the background person image in the image synthesis space based on the person depth information, and synthesize an initial image sample that includes both the foreground person and the background person; a sample augmentation module configured to synthesize a plurality of derivative image samples based on the step-by-step adjustment of the relative position relationship between the foreground person and the background person in the image synthesis space, so that a temporal gradual change characteristic is reflected between the initial image sample and each derivative image sample; a model training module configured to train an image segmentation model using the synthesized image samples, and use the corresponding foreground person image as a supervision sample to train the image segmentation model to a converged state.

[0185] Another embodiment of the present application further provides a main body person image segmentation device. As Figure 4 shown, it is a schematic internal structure diagram of the main body person image segmentation device. The main body person image segmentation device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable non-volatile storage medium of the main body person image segmentation device stores an operating system, a database, and computer-readable instructions. Information sequences can be stored in the database. When the computer-readable instructions are executed by the processor, the processor can implement a main body person image segmentation method.

[0186] The processor of the main body person image segmentation device is used to provide computing and control capabilities to support the operation of the entire main body person image segmentation device. Computer-readable instructions can be stored in the memory of the main body person image segmentation device. When the computer-readable instructions are executed by the processor, the processor can execute the main body person image segmentation method of the present application. The network interface of the main body person image segmentation device is used to connect and communicate with a terminal.

[0187] Those skilled in the art can understand that Figure 4 the structure shown in

[0188] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the main body person image segmentation device to which the solution of the present application is applied. The specific main body person image segmentation device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. Figure 3For the specific functions of each module in [the device], the memory stores program codes and various types of data required to execute the above modules or sub-modules. The network interface is used to implement data transmission between user terminals or servers. In this embodiment, the non-volatile readable storage medium stores the program codes and data required to execute all modules in the main character image segmentation device of the present application, and the server can call the program codes and data of the server to execute the functions of all modules.

[0189] The present application also provides a non-volatile readable storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the main character image segmentation method according to any embodiment of the present application.

[0190] The present application also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by one or more processors, the steps of the method according to any embodiment of the present application are implemented.

[0191] Those of ordinary skill in the art can understand that to implement all or part of the processes in the methods of the above embodiments of the present application, it can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disc, a Read-Only Memory (ROM), or a Random Access Memory (RAM), etc.

[0192] In summary, through the comprehensive application of multiple innovative technologies, the present application significantly improves the efficiency, accuracy, and adaptability of main character image segmentation. First, through the optimization of the saliency determination conditions, the main character can be accurately identified in complex multi-person scenarios, avoiding misjudgments in traditional methods. Second, the introduction of the Kalman filter for tracking and smoothing the main anchor box not only improves the stability of the segmentation results but also enhances the tracking effect in consecutive frames, maintaining accurate segmentation even in scenarios where people move quickly or the background is complex. In addition, the present application adopts lightweight object detection models and image segmentation models. Through pruning and channel compression technologies, the consumption of computing resources is significantly reduced, enabling efficient operation when implemented on low-performance devices and meeting real-time requirements. At the same time, the training sample synthesis strategy based on depth information further improves the robustness and generalization ability of the model, solving the problem of insufficient training data. The comprehensive application of these technical advantages enables the present application to not only perform excellently in multi-person scenarios but also be widely applied to fields such as video conferencing, live streaming, and virtual reality, providing users with a higher-quality visual experience, while reducing hardware costs and improving the overall performance and practicality.

Claims

1. A method for segmenting a subject image, characterized in that: include: Detecting the image region where each person is located in the image frame of the live stream, and determining the regional anchor frame corresponding to the image region where each person is located; According to the preset saliency judgment condition, a regional anchor frame corresponding to the image region indicating the main character is located is determined from all regional anchor frames as the main anchor frame; Tracking and smoothing the main body anchor frame to stabilize the position information of the main body anchor frame; Image segments are performed on the image frames in the live stream according to the position information of the subject anchor frame to obtain a human body image of the subject person.

2. The subject person image segmentation method according to claim 1, characterized in that: According to the preset saliency judgment condition, a regional anchor frame corresponding to the image region where the subject person is located is determined from all regional anchor frames as the subject anchor frame, including: Calculate the distance between each region anchor box and the image center of the image frame in which it is located; Calculate the area size of each region anchor box; sorting the regional anchor boxes according to the distance and the area; The top ranked region is used as the saliency judgment condition, and the top ranked region anchor box is selected as the main anchor box.

3. The subject person image segmentation method according to claim 1, characterized in that: According to the preset saliency judgment condition, a regional anchor frame corresponding to the image region where the subject person is located is determined from all regional anchor frames as the subject anchor frame, including: Continuously track the regional anchor frame of each person and extract the motion feature information of each person, including the movement speed, movement direction change frequency and residence time of each person in the live stream image; According to the motion characteristic information of each character, a preset activity score model is applied to determine the activity score corresponding to each character; The highest activity score is used as a saliency determination condition, and the region anchor frame with the highest score is selected as the main anchor frame.

4. The subject person image segmentation method according to claim 1, characterized in that: Tracking and smoothing the subject anchor frame includes: Predict the position information of the subject anchor frame based on the Kalman filter; Correcting the predicted position information according to the position information of the subject anchor frame determined in the current image frame to obtain the predicted position information; According to the corrected predicted position information, the state parameters of the Kalman filter are updated to reflect the latest position of the subject anchor frame; In image frames where no detection is performed, the predicted position information is used as the position of the subject anchor box.

5. The subject person image segmentation method according to claim 1, characterized in that: Performing image segmentation on the image frames in the live stream according to the position information of the subject anchor frame to obtain a human body image of the subject person includes: According to the position information of the subject anchor frame, the corresponding image area in the image frame is cropped to obtain a portrait area map; The cropped portrait area image is input into an image segmentation model based on the attention mechanism, and a segmentation mask of the human body image of the subject person is generated by the image segmentation model; After performing image editing processing on the human body image or the background image, the live stream is transmitted to a live broadcast room for playback; A human body image of a subject person and a background image are extracted from the image frame according to the segmentation mask.

6. The subject person image segmentation method according to any one of claims 1 to 5, characterized in that: The method further comprises: The pruned and channel-compressed object detection model is called to determine the regional anchor frame corresponding to each person from the image frames of the live stream at preset intervals. For the image frames that have not been detected, the subject anchor frame adopts the subject anchor frame determined after corresponding tracking and smoothing processing; Using an attention-based image segmentation model to perform image segmentation on image frames in the live stream; A Kalman filter algorithm is applied to track and smooth the subject anchor frame.

7. The subject person image segmentation method according to claim 6, characterized in that: Before performing image segmentation on the image frames in the live stream using an attention mechanism-based image segmentation model, the following steps are included: Acquire a foreground person real scene image and a background person real scene image and their corresponding segmentation masks and person depth information, and extract the corresponding foreground person image and background person image according to the segmentation masks; Determining the layer order of the foreground character image and the background character image in the image synthesis space based on the character depth information, and synthesizing an initial image sample containing both the foreground character and the background character; Based on the stepwise adjustment of the relative position relationship between the foreground person and the background person in the image synthesis space, a plurality of derived image samples are synthesized so that the initial image sample and each derived image sample exhibit a time-sequential gradual change characteristic; The image segmentation model is trained using the synthesized image samples, and the corresponding foreground person images are used as supervision samples to train the image segmentation model to a convergence state.

8. A subject person image segmentation device, characterized in that: include: An anchor frame detection module, configured to detect the image region where each person is located in the image frame of the live stream, and determine the regional anchor frame corresponding to the image region where each person is located; A subject determination module is configured to determine, according to a preset saliency determination condition, a region anchor frame corresponding to the image region indicating the location of the subject person from all region anchor frames as the subject anchor frame; A tracking and smoothing module, configured to track and smooth the subject anchor frame to stabilize position information of the subject anchor frame; The segmentation processing module is configured to perform image segmentation on the image frames in the live stream according to the position information of the subject anchor frame to obtain a human body image of the subject character.

9. A subject person image segmentation device, comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.

10. A non-volatile readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.

Citation Information

Cited By

  • Subject person image segmentation method, apparatus and device, and medium

    WO2026179250A1