Image processing method and device, electronic equipment and storage medium

By acquiring and segmenting images on large-screen devices and using historical and current detection box information for background replacement, the problems of detection box jitter and insufficient accuracy are solved, generating high-quality display images.

CN114155268BActive Publication Date: 2026-07-24BEIJING SENSETIME TECH DEV CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING SENSETIME TECH DEV CO LTD
Filing Date
2021-11-24
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Large-screen devices suffer from issues such as detection frame jitter and insufficient accuracy in background replacement when replacing images.

Method used

By acquiring the image to be processed from the display device, limb detection is performed. The target detection box is determined using historical frame images and current detection box information. Background segmentation and replacement are then performed to generate a segmented image and apply a background replacement scheme.

Benefits of technology

It reduces detection box jitter, improves the accuracy and effect of background replacement, and generates target display images with better display effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155268B_ABST
    Figure CN114155268B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image processing method and device, electronic equipment and storage medium, the method comprising: acquiring a to-be-processed image collected by a display device; performing limb detection on the to-be-processed image to determine current bounding box information of a user included in the to-be-processed image; determining target bounding box information based on historical bounding box information corresponding to a historical frame image collected by the display device and the current bounding box information; performing background segmentation on the to-be-processed image based on the target bounding box information to generate a segmentation image corresponding to the to-be-processed image; the segmentation image is used to indicate at least one of a background region and a foreground region in the to-be-processed image; and performing background replacement on the to-be-processed image based on the segmentation image and a set background replacement scheme to generate a target display image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and more specifically, to an image processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] Image segmentation technology refers to the technique of dividing an image into non-overlapping regions with their own characteristics. Image segmentation technology can be applied to a variety of scenarios, such as image background replacement.

[0003] With the development of technology, more and more electronic devices with large display screens have the need for image background replacement. For example, during video conferencing, it is necessary to replace the background of the projected image displayed by the projector to meet the meeting requirements. However, image background replacement on large screens often results in poor display quality. Therefore, there is an urgent need for an image processing solution to meet the background replacement needs in large-screen scenarios. Summary of the Invention

[0004] In view of this, the present disclosure provides at least one image processing method, apparatus, electronic device, and storage medium to improve the accuracy of image processing and the display effect of the processed image.

[0005] In a first aspect, this disclosure provides an image processing method, including:

[0006] Acquire the image to be processed from the display device;

[0007] Perform limb detection on the image to be processed to determine the current detection box information of the user included in the image to be processed;

[0008] Based on the historical detection box information corresponding to the historical frame images collected by the display device, and the current detection box information, the target detection box information is determined;

[0009] Based on the target detection box information, the image to be processed is segmented to generate a segmented image corresponding to the image to be processed; the segmented image is used to indicate at least one of the background region and the foreground region in the image to be processed;

[0010] Based on the segmented image and the set background replacement scheme, the background of the image to be processed is replaced to generate the target display image.

[0011] In the above method, the target detection box information is determined by using the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information. This can alleviate the problem of repeated jumping of the detection box edges caused by detection box jitter. Based on the target detection box information, the background of the image to be processed can be segmented more accurately to generate the segmented image corresponding to the image to be processed. Then, based on the segmented image and the set background replacement scheme, the background of the image to be processed is replaced to generate the target display image. The background replacement effect is good, resulting in a better display effect of the generated target display image.

[0012] In one possible implementation, the process of replacing the background of the image to be processed based on the segmented image and the set background replacement scheme to generate a target display image includes at least one of the following:

[0013] The first pixel information of the pixel to be replaced in the background region of the image to be processed, as indicated by the segmented image, is adjusted to preset pixel information to generate the target display image;

[0014] The first pixel information of the pixel to be replaced in the background region of the image to be processed, as indicated by the segmented image, is adjusted to the second pixel information of the candidate pixel in the preset replacement image that matches the pixel position of the pixel to be replaced, thereby generating the target display image.

[0015] Here, you can easily replace the background of the image to be processed by setting preset pixel information, which is quite efficient. Alternatively, you can determine a preset replacement image, which allows for more flexible background replacement of the image to be processed, offering greater interest and variety.

[0016] In one possible implementation, the number of current detection boxes is multiple, and the step of determining the target detection box information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information includes:

[0017] For each current detection box included in the image to be processed, based on the detection box identifier indicated by the current detection box information corresponding to the current detection box, the historical detection box information in the historical frame image that matches the detection box identifier is determined; and based on the current detection box information and the historical detection box information that matches the detection box identifier, the target detection box information corresponding to the current detection box is determined.

[0018] The step of performing background segmentation on the image to be processed based on the target detection box information to generate a segmented image corresponding to the image to be processed includes:

[0019] Based on the obtained target detection box information, the background of the image to be processed is segmented to generate the segmented image corresponding to the image to be processed.

[0020] Here, the corresponding target detection box information can be determined for each current detection box. Subsequently, based on the obtained target detection box information, background segmentation is performed on the image to be processed, which improves the accuracy of background segmentation.

[0021] In one possible implementation, the number of current detection boxes is multiple, and after determining the current detection box information of the user included in the image to be processed, the method further includes:

[0022] Based on at least one of the confidence level, detection box size, and position information indicated by the current detection box information, multiple current detection boxes included in the image to be processed are filtered to obtain the filtered current detection boxes and the current detection box information corresponding to the filtered current detection boxes.

[0023] The determination of target detection box information based on historical detection box information corresponding to historical frame images acquired by the display device and the current detection box information includes:

[0024] Based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information corresponding to the current detection box after filtering, the target detection box information is determined.

[0025] In this embodiment, multiple current detection boxes included in the image to be processed are filtered based on at least one of the confidence level, detection box size, and position information indicated by the current detection box information. This results in the current detection box information corresponding to the filtered current detection box, making the filtered current detection box a detection box with higher importance. This allows for the generation of target detection box information containing more accurate limb information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information corresponding to the filtered current detection box.

[0026] In one possible implementation, the number of current detection boxes is multiple, and the step of determining the target detection box information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information includes:

[0027] Based on the position information indicated by each of the current detection boxes, generate intermediate detection box information that surrounds the multiple current detection boxes;

[0028] Based on the intermediate detection box information and the historical detection box information, the target detection box information is determined.

[0029] Here, by generating intermediate detection box information that surrounds multiple current detection boxes, and based on the intermediate detection box information and historical detection box information, the target detection box information is determined. The number of the target detection box is one, so that the background segmentation of the image to be processed can be performed more efficiently through the target detection box, thereby improving the efficiency of image processing.

[0030] In one possible implementation, determining the target detection box information based on the intermediate detection box information and the historical detection box information includes:

[0031] In response to the matching of the number of current detection boxes in the image to be processed with the number of historical detection boxes in the historical frame image, at least one set of average coordinate information corresponding to a point set is determined based on the intermediate detection box information and the historical detection box information; wherein, the point set includes multiple vertices that are in the same position as the detection boxes and belong to the intermediate detection boxes and the historical detection boxes respectively.

[0032] Based on the obtained average coordinate information, the target detection box information is determined.

[0033] Using the above method, when determining that the number of current detection boxes in the image to be processed matches the number of historical detection boxes in the historical frame image, at least one set of average coordinate information corresponding to the point set is determined based on the intermediate detection box information and the historical detection box information. This can alleviate the sliding frame phenomenon caused when the number of current detection boxes does not match the number of historical detection boxes in the historical frame image, and when the average coordinate information is determined using the intermediate detection box information and the historical detection box information.

[0034] In one possible implementation, when the at least one set of points includes two sets of points diagonally related to the detection box, determining the target detection box information based on the obtained average coordinate information includes:

[0035] For each set of points, a preset region corresponding to the target vertex indicated by the average coordinate information is determined, centered on the average coordinate information and with a preset length as the radius; and

[0036] If the historical coordinate information of a historical vertex that is in the same position as the target vertex on the historical detection frame is located in the preset area corresponding to the target vertex, the historical coordinate information of the historical vertex is determined as the target coordinate information of the target vertex.

[0037] If the historical coordinate information of a historical vertex that is in the same position as the target vertex on the historical detection box is not located in the preset area corresponding to the target vertex, the average coordinate information of the target vertex is determined as the target coordinate information of the target vertex.

[0038] Based on the determined target coordinate information, the target detection box information is determined.

[0039] To alleviate the frequent jitter of the detection box, a preset area corresponding to the target vertex indicated by the average coordinate information is determined. When the historical coordinate information of a historical vertex that is in the same position as the target vertex on the historical detection box is located in the preset area, the historical coordinate information of the historical vertex is determined as the target coordinate information of the target vertex. This ensures that the target detection box information of the image to be processed follows the historical detection box information corresponding to the historical frame image, maintaining the stability of the intermediate detection box and ensuring the stability of the background replacement.

[0040] The effects of the following devices, electronic equipment, etc., are described in the instructions above and will not be repeated here.

[0041] Secondly, this disclosure provides an image processing apparatus, comprising:

[0042] The acquisition module is used to acquire the image to be processed captured by the display device;

[0043] The first determining module is used to perform limb detection on the image to be processed and determine the current detection box information of the user included in the image to be processed;

[0044] The second determining module is used to determine the target detection box information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information;

[0045] The segmentation module is used to perform background segmentation on the image to be processed based on the target detection box information, and generate a segmented image corresponding to the image to be processed; the segmented image is used to indicate at least one of the background region and the foreground region in the image to be processed;

[0046] The generation module is used to perform background replacement on the image to be processed based on the segmented image and the set background replacement scheme, and generate the target display image.

[0047] In one possible implementation, the generation module, when performing background replacement on the image to be processed based on the segmented image and the set background replacement scheme to generate the target display image, is configured to:

[0048] The first pixel information of the pixels to be replaced belonging to the background region on the image to be processed, as indicated by the segmented image, is adjusted to preset pixel information to generate the target display image; or...

[0049] The first pixel information of the pixel to be replaced in the background region of the image to be processed, as indicated by the segmented image, is adjusted to the second pixel information of the candidate pixel in the preset replacement image that matches the pixel position of the pixel to be replaced, thereby generating the target display image.

[0050] In one possible implementation, the number of current detection boxes is multiple, and the second determining module, when determining the target detection box information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information, is used to:

[0051] For each current detection box included in the image to be processed, based on the detection box identifier indicated by the current detection box information corresponding to the current detection box, the historical detection box information in the historical frame image that matches the detection box identifier is determined; and based on the current detection box information and the historical detection box information that matches the detection box identifier, the target detection box information corresponding to the current detection box is determined.

[0052] The segmentation module, when performing background segmentation on the image to be processed based on the target detection box information to generate a segmented image corresponding to the image to be processed, is used for:

[0053] Based on the obtained target detection box information, the background of the image to be processed is segmented to generate the segmented image corresponding to the image to be processed.

[0054] In one possible implementation, the number of current detection boxes is multiple, and the device further includes a filtering module, which is used to filter the multiple current detection boxes included in the image to be processed based on at least one of the confidence level, detection box size, and position information indicated by the current detection box information after determining the current detection box information of the user included in the image to be processed, to obtain the filtered current detection boxes and the current detection box information corresponding to the filtered current detection boxes.

[0055] The second determining module, when determining the target detection box information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information, is used for:

[0056] Based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information corresponding to the current detection box after filtering, the target detection box information is determined.

[0057] In one possible implementation, the number of current detection boxes is multiple, and the second determining module, when determining the target detection box information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information, is used to:

[0058] Based on the position information indicated by each of the current detection boxes, generate intermediate detection box information that surrounds the multiple current detection boxes;

[0059] Based on the intermediate detection box information and the historical detection box information, the target detection box information is determined.

[0060] In one possible implementation, the second determining module, when determining the target detection box information based on the intermediate detection box information and the historical detection box information, is used to:

[0061] In response to the matching of the number of current detection boxes in the image to be processed with the number of historical detection boxes in the historical frame image, at least one set of average coordinate information corresponding to a point set is determined based on the intermediate detection box information and the historical detection box information; wherein, the point set includes multiple vertices that are in the same position as the detection boxes and belong to the intermediate detection boxes and the historical detection boxes respectively.

[0062] Based on the obtained average coordinate information, the target detection box information is determined.

[0063] In one possible implementation, when the at least one set of points includes two sets of points diagonally related to the detection box, the second determining module, when determining the target detection box information based on the obtained average coordinate information, is configured to:

[0064] For each set of points, a preset region corresponding to the target vertex indicated by the average coordinate information is determined, centered on the average coordinate information and with a preset length as the radius; and

[0065] If the historical coordinate information of a historical vertex that is in the same position as the target vertex on the historical detection frame is located in the preset area corresponding to the target vertex, the historical coordinate information of the historical vertex is determined as the target coordinate information of the target vertex.

[0066] If the historical coordinate information of a historical vertex that is in the same position as the target vertex on the historical detection box is not located in the preset area corresponding to the target vertex, the average coordinate information of the target vertex is determined as the target coordinate information of the target vertex.

[0067] Based on the determined target coordinate information, the target detection box information is determined.

[0068] Thirdly, this disclosure provides an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the image processing method as described in the first aspect or any of the embodiments above are performed.

[0069] Fourthly, this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the image processing method as described in the first aspect or any of the embodiments above.

[0070] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0071] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0072] Figure 1 A schematic flowchart of an image processing method provided by an embodiment of this disclosure is shown;

[0073] Figure 2 This illustration shows a schematic diagram of the middle detection box in an image processing method provided by an embodiment of the present disclosure;

[0074] Figure 3 A schematic flowchart of another image processing method provided by an embodiment of this disclosure is shown;

[0075] Figure 4 A schematic diagram of the architecture of an image processing apparatus provided in an embodiment of this disclosure is shown;

[0076] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0077] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0078] With the development of technology, more and more electronic devices with large display screens have the need for image background replacement; for example, when conducting video conferences, it is necessary to replace the background of the projected image displayed by the projector to meet the requirements of the meeting.

[0079] Generally, traditional image background replacement schemes are applicable to small-screen devices such as mobile phones. However, when applied to large-screen devices such as projectors, the larger size and richer information of the images displayed on these devices result in more interference, reducing the accuracy of background replacement. Therefore, this disclosure provides an image processing method, apparatus, electronic device, and storage medium.

[0080] The shortcomings of the above solutions are the result of the inventor's practical experience and careful research. Therefore, the discovery process of the above problems and the solutions proposed in this disclosure below should be considered as the inventor's contribution to this disclosure.

[0081] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0082] To facilitate understanding of the embodiments of this disclosure, a detailed description of an image processing method disclosed in this disclosure will be provided first. The image processing method provided in this disclosure is generally executed by a server or electronic device. The server may be a cloud server, a local server, etc.; the electronic device may be a user equipment (UE), a mobile device, a personal digital assistant (PDA), a handheld device, etc. In some possible implementations, the image processing method can be implemented by a processor calling computer-readable instructions stored in memory.

[0083] See Figure 1 The diagram shown is a flowchart illustrating an image processing method provided in this embodiment of the present disclosure. The method includes: S101-S105, wherein:

[0084] S101, acquire the image to be processed captured by the display device;

[0085] S102, perform limb detection on the image to be processed, and determine the current detection box information of the user included in the image to be processed;

[0086] S103, based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information, determine the target detection box information;

[0087] S104, based on the target detection box information, perform background segmentation on the image to be processed to generate a segmented image corresponding to the image to be processed; the segmented image is used to indicate at least one of the background region and the foreground region in the image to be processed;

[0088] S105, based on the segmented image and the set background replacement scheme, perform background replacement on the image to be processed to generate the target display image.

[0089] In the above method, the target detection box information is determined by using the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information. This can alleviate the problem of repeated jumping of the detection box edges caused by detection box jitter. Based on the target detection box information, the background of the image to be processed can be segmented more accurately to generate the segmented image corresponding to the image to be processed. Then, based on the segmented image and the set background replacement scheme, the background of the image to be processed is replaced to generate the target display image. The background replacement effect is good, resulting in a better display effect of the generated target display image.

[0090] The following provides a detailed explanation of S101-S105.

[0091] In S101, the display device can be an electronic device with a large screen, such as a television, projector, or computer. The image to be processed can be any image captured by the display device; for example, the display device can capture a real-world scene image in real time and use that image as the image to be processed; or, the display device can also capture any video frame from received video data and use that video frame as the image to be processed. For example, in a remote video conferencing scenario, the display device can receive video content transmitted by other devices over a network and capture the image to be processed from that video content. After capturing the image to be processed, the display device can display that image; that is, the image to be processed can be any frame of the image displayed by the display device.

[0092] In S102, during implementation, a trained first neural network for limb detection can be used to perform limb detection on the image to be processed, determining the current detection box information for each user included in the image. The current detection box information may include the position information of the current detection box, such as: the size information of the detection box and the coordinates of its center point; and / or the position information of the current detection box may include: the coordinates of the two diagonal vertices of the detection box, or the coordinates of all four vertices, etc.

[0093] The current detection frame information may also include a detection frame identifier. This identifier is determined by tracking users in multiple consecutive frames of images captured by the display device. Specifically, the detection frame identifier for the same user is consistent across multiple consecutive frames, while different users have different identifiers. If a user included in the current detection frame has a corresponding historical detection frame in a historical frame image, the detection frame identifier of the historical frame can be used as the detection frame identifier for the current frame. If the user included in the current detection frame does not exist in a historical frame image, a corresponding detection frame identifier can be generated for the current detection frame.

[0094] The current detection box information may also include the confidence score of the current detection box, which represents the probability that the object included in the current detection box belongs to a human. Generally, the more complete the user's limbs are in the current detection box, the higher the probability of detecting that the user is human, and the higher the corresponding confidence score; the larger the area of ​​the user's limbs that is occluded in the current detection box, the lower the probability of detecting that the user is human, and the lower the corresponding confidence score.

[0095] In one optional embodiment, when the number of current detection boxes included in the image to be processed is multiple, after determining the current detection box information of the user included in the image to be processed, the method further includes: filtering the multiple current detection boxes included in the image to be processed based on at least one of the confidence level indicated by the current detection box information, detection box size, and position information, to obtain the filtered current detection boxes and the current detection box information corresponding to the filtered current detection boxes.

[0096] When the image to be processed includes multiple current detection boxes, the multiple current detection boxes included in the image to be processed can be filtered based on at least one of the confidence level indicated by the current detection box information, the detection box size, and the position information, to obtain the filtered current detection boxes and the current detection box information corresponding to the filtered current detection boxes; then, in S103, the target detection box information can be determined based on the historical detection box information corresponding to the historical frame images collected by the display device and the filtered current detection box information.

[0097] During implementation, when filtering the current detection box using confidence level, a confidence level threshold can be set to delete the current detection box with a confidence level lower than the confidence level threshold, thus obtaining the filtered current detection box and the current detection box information corresponding to the filtered current detection box.

[0098] Considering that when the detection frame size is small, the user within the detection frame is farther from the image acquisition device (such as the display device), making the user more likely to be passing through the image acquisition device and thus less important, a size threshold can be set when filtering the current detection frames based on their size. Detection frames smaller than this threshold can be deleted, resulting in filtered current detection frame information. Alternatively, the ratio between the current detection frame size and the display device's screen size can be determined, and current detection frames with a ratio smaller than a set threshold can be deleted, resulting in filtered current detection frames and their corresponding information.

[0099] In real-world scenarios, users located at the edges of the image acquisition device's capture range are considered less important. For example, in human-computer interaction scenarios, users controlling electronic devices (such as televisions) are typically positioned in the center of the television's capture range; similarly, in video conferencing scenarios, users participating in the meeting are usually located in the center of the electronic device's (such as a computer's) capture range. Therefore, when using location information to filter the current detection bounding box, a central region range corresponding to the image to be processed can be set. Detection boxes whose center point coordinates are outside this central region range can be filtered out, resulting in the filtered current detection bounding box and its corresponding information.

[0100] When filtering the current detection box using confidence level and detection box size, the current detection box with confidence level greater than or equal to the confidence level threshold and detection box size greater than or equal to the size threshold can be retained, while the current detection box with confidence level less than the confidence level threshold or detection box size less than the size threshold can be filtered out, thus obtaining the filtered current detection box and the current detection box information corresponding to the filtered current detection box.

[0101] During implementation, the size information of the user's face frame in the current detection frame can also be determined. By setting a face frame threshold, the current detection frames whose face frame size information is smaller than the face frame threshold are deleted, thus obtaining the filtered current detection frame and the current detection frame information corresponding to the filtered current detection frame.

[0102] In this embodiment, multiple current detection boxes included in the image to be processed are filtered based on at least one of the confidence level, detection box size, and position information indicated by the current detection box information to obtain filtered current detection boxes. The filtered current detection boxes are detection boxes with higher importance, so that target detection box information containing more accurate limb information can be generated based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information of the filtered current detection boxes.

[0103] In step S103, when there is only one detection box, based on the detection box identifier of the current detection box, it is determined whether there is historical detection box information corresponding to that identifier in the historical frame image. If not, the current detection box information is used as the target detection box information. If it exists, the target detection box information can be determined based on the historical detection box information and the current detection box information in the historical frame image. The number of historical frame images can be set as needed, for example, the number of historical frame images can be 5 frames, 10 frames, etc.

[0104] For example, if there are 5 historical frames, and all 5 historical frames contain historical detection box information corresponding to the detection box identifier, and the current detection box information is the coordinate information of two vertices on the first diagonal, then for the first vertex located in the upper left position, the average coordinate information of the first vertex is determined based on the coordinate information of the first vertex indicated by the current detection box information and the coordinate information of the vertices located in the upper left position indicated by the 5 historical detection box information. This is obtained by averaging the coordinate information of the 6 vertices located in the upper left position. Similarly, for the second vertex located in the lower right position, the average coordinate information of the second vertex is determined based on the coordinate information of the second vertex indicated by the current detection box information and the coordinate information of the vertices located in the lower right position indicated by the 5 historical detection box information. Finally, the average coordinate information of the first vertex and the average coordinate information of the second vertex are used to determine the target detection box information.

[0105] The process of determining the average coordinates of the third vertex in the upper right position and the fourth vertex in the lower left position can be referred to the process of determining the average coordinates of the first and second vertices mentioned above, and will not be described in detail here.

[0106] When there are multiple current detection boxes, in the first method, for each current detection box, the target detection box information corresponding to the current detection box can be determined; the target detection box information corresponding to each current detection box can be obtained; then, based on the target detection box information, the background of the image to be processed can be segmented to generate the segmented image corresponding to the image to be processed; finally, based on the segmented image and the set background replacement scheme, the background of the image to be processed can be replaced to generate the target display image.

[0107] In one optional implementation, there are multiple current detection boxes. Determining target detection box information based on historical detection box information corresponding to historical frame images acquired by the display device and current detection box information may include: for each current detection box included in the image to be processed, determining historical detection box information in historical frame images that matches the detection box identifier based on the detection box identifier indicated by the current detection box information corresponding to the current detection box; and determining target detection box information corresponding to the current detection box based on the current detection box information and the historical detection box information that matches the detection box identifier.

[0108] Here, the corresponding target detection box information can be determined for each current detection box. Subsequently, based on the obtained target detection box information, background segmentation is performed on the image to be processed, which improves the accuracy of background segmentation.

[0109] First, for each current detection box, determine whether there is a historical detection box in the historical frame image that matches the detection box identifier of the current detection box. If not, the current detection box information is determined as the target detection box information; if it exists, the historical detection box information in the historical frame image that matches the detection box identifier is determined.

[0110] Secondly, when the current detection box information consists of the coordinates of two vertices on a diagonal, the average coordinates of these two vertices can be calculated based on the current detection box information and historical detection box information matching the detection box identifier. Based on the average coordinates of these two vertices, the target detection box information corresponding to the current detection box can be determined. This allows us to obtain the target detection box information corresponding to each current detection box.

[0111] Then, in S104, based on the target detection box information, background segmentation is performed on the image to be processed to generate a segmented image corresponding to the image to be processed. This may include: based on the obtained target detection box information, background segmentation is performed on the image to be processed respectively to generate a segmented image corresponding to the image to be processed.

[0112] For example, for each object detection box, a local image corresponding to the object detection box can be extracted from the image to be processed based on the object detection box information. Using a trained second neural network for background segmentation, the local image is segmented to obtain a local segmentation image. This yields the local segmentation images corresponding to each object detection box. Then, based on the local segmentation images corresponding to each object detection box, a segmentation image corresponding to the image to be processed is generated. For instance, the union of the local foreground regions in the local segmentation images corresponding to each object detection box can be taken to obtain the overall foreground region. Then, based on the determined overall foreground region and the image to be processed, the segmentation image corresponding to the image to be processed is obtained.

[0113] In this process, the segmented image and the image to be processed have the same size. The segmented image can be a binary image including a first pixel value and a second pixel value to distinguish the foreground and background regions of the image to be processed. The first pixel value and the second pixel value can be any different pixel values. That is, the first pixel point corresponding to the first pixel value in the segmented image is the background region, and the second pixel point corresponding to the second pixel value is the foreground region; the pixels in the image to be processed whose positions match the first pixel point in the segmented image belong to the background region, and the pixels whose positions match the second pixel point belong to the foreground region.

[0114] In the second approach, the union of all current detection boxes can be taken to obtain the information of the larger intermediate detection box; then the target detection box information corresponding to the intermediate detection box can be determined; based on the target detection box information, the local image corresponding to the target detection box can be extracted from the image to be processed; the background of the local image can be segmented to obtain the segmented image corresponding to the image to be processed; finally, the background of the image to be processed can be replaced based on the segmented image to generate the target display image.

[0115] In an optional implementation, when the number of current detection boxes included in the image to be processed is multiple, determining the target detection box information based on the historical detection box information corresponding to the historical frame images acquired by the display device and the current detection box information may include:

[0116] Step A1: Based on the position information indicated by each of the current detection box information, generate intermediate detection box information surrounding the multiple current detection boxes;

[0117] Step A2: Determine the target detection box information based on the intermediate detection box information and the historical detection box information.

[0118] Here, by generating intermediate detection box information that surrounds multiple current detection boxes, and based on the intermediate detection box information and historical detection box information, the target detection box information is determined. The number of the target detection box is one, so that the background segmentation of the image to be processed can be performed more efficiently through the target detection box, thereby improving the efficiency of image processing.

[0119] In step A1, see Figure 2 As shown, the coordinates of the vertex located in the upper left position in the current detection box 1 are (x1, y1), the coordinates of the vertex located in the lower right position in the current detection box 2 are (x2, y2), and the coordinates of the vertex located in the lower left position in the current detection box 3 are (x3, y3). Therefore, the information of the middle detection box is determined to include: the coordinates of the vertex located in the upper left position (x1, y1) and the coordinates of the vertex located in the lower right position (x2, y3).

[0120] In step A2, in one optional implementation, determining the target detection box information based on the intermediate detection box information and the historical detection box information may include:

[0121] Step A21: In response to the number of current detection boxes included in the image to be processed matching the number of historical detection boxes included in the historical frame image, at least one set of average coordinate information corresponding to a point set is determined based on the intermediate detection box information and the historical detection box information; wherein, the point set includes multiple vertices that are in the same position as the detection boxes and belong to the intermediate detection boxes and the historical detection boxes respectively.

[0122] Step A22: Based on the obtained average coordinate information, determine the target detection box information.

[0123] Using the above method, when determining that the number of current detection boxes in the image to be processed matches the number of historical detection boxes in the historical frame image, at least one set of average coordinate information corresponding to the point set is determined based on the intermediate detection box information and the historical detection box information. This can alleviate the sliding frame phenomenon caused when the number of current detection boxes does not match the number of historical detection boxes in the historical frame image, and when the average coordinate information is determined using the intermediate detection box information and the historical detection box information.

[0124] In implementation, it can be determined whether the number of current detection boxes in the image to be processed matches the number of historical detection boxes in the historical frame images. If they do not match, the historical frame images can be filtered out, and the information of the intermediate detection boxes can be determined as the target detection box information. If they match, the average coordinate information corresponding to at least one set of points can be determined based on the information of the intermediate detection boxes and the historical detection boxes. For example, a set of points may include the vertex located at the upper left position of the intermediate detection box and the vertex located at the upper left position of the historical detection box. Alternatively, a set of points may include the vertex located at the lower right position of the intermediate detection box and the vertex located at the lower right position of the historical detection box. In implementation, the coordinate information of each vertex in the set of points is averaged to obtain the average coordinate information corresponding to the set of points.

[0125] If the smoothing window is set to 5, then 5 historical frame images can be selected. The number of current detection boxes included in the image to be processed is determined to match the number of historical detection boxes included in each historical frame image. If they match, then based on the intermediate detection box information and the historical detection box information, at least one set of average coordinate information corresponding to the point set is determined to complete the smoothing operation of the intermediate detection boxes. If they do not match, then the 5 historical frame images are deleted, and the intermediate detection box smoothing operation is restarted. That is, for the current image to be processed, the intermediate detection box information can be determined as the target detection box information, and the current image to be processed is used as the historical frame image of the next image to be processed to smooth the next image to be processed.

[0126] In implementation, if the historical detection box information consists of information about multiple limb bounding boxes included in a historical frame image, historical intermediate bounding box information surrounding the multiple limb bounding boxes can be generated based on this information. Then, based on the intermediate detection box information and the historical intermediate bounding box information, the average coordinate information of at least one set of points can be determined. If the historical detection box information consists of historical intermediate bounding box information surrounding multiple limb bounding boxes corresponding to a historical frame image, the average coordinate information of at least one set of points can be determined directly based on the intermediate detection box information and the historical detection box information.

[0127] When determining the average coordinates of a set of points, the target detection box information can be determined based on the average coordinates of the set of points and the coordinates of another vertex diagonally opposite the detection box. For example, after determining the average coordinates of the set of points located at the top left position, the average coordinates of the top left position and the coordinates of the vertex located at the bottom right position indicated by the middle detection box information can be used as the target detection box information; or, the size information and center point coordinates of the target detection box can be determined based on the average coordinates of the set of points located at the top left position and the coordinates of the vertex located at the bottom right position indicated by the middle detection box information; the size information and center point coordinates of the target detection box can then be used as the target detection box information.

[0128] In one optional implementation, when the average coordinate information of a set of points (composed of vertices located in the upper left position) is determined, a preset region corresponding to the vertex located in the upper left position indicated by the set of points is determined with the average coordinate information corresponding to the set of points as the center and a preset length as the radius. Then, it is detected whether the historical coordinate information of a historical vertex (i.e., a historical vertex located in the upper left position on the historical detection box) that is in the same direction as the vertex in the historical detection box is located within the preset region. If so, the historical coordinate information of that historical vertex is determined as the target coordinate information of the vertex; otherwise, the average coordinate information corresponding to the set of points is determined as the target coordinate information of the vertex. Finally, based on the target coordinate information of the vertex and the coordinate information of the vertex located in the lower right position indicated by the middle detection box information, the target detection box information is determined.

[0129] In one optional implementation, when at least one set of points includes two sets of points diagonally related to the detection box, determining the target detection box information based on the obtained average coordinate information may include:

[0130] Step B1: For each set of points, take the average coordinate information corresponding to the set of points as the center and the preset length as the radius, and determine the preset region corresponding to the target vertex indicated by the average coordinate information.

[0131] Step B2: If the historical coordinate information of a historical vertex that is in the same position as the target vertex on the historical detection frame is located in the preset area corresponding to the target vertex, the historical coordinate information of the historical vertex is determined as the target coordinate information of the target vertex.

[0132] Step B3: If the historical coordinate information of a historical vertex that is in the same position as the target vertex on the historical detection box is not located in the preset area corresponding to the target vertex, the average coordinate information of the target vertex is determined as the target coordinate information of the target vertex.

[0133] Step B4: Based on the determined target coordinate information, determine the target detection box information.

[0134] Two sets of points diagonally opposite each other in the detection box can be defined as follows: the vertices in one set and the vertices in the other set are located on opposite sides of the diagonal of the detection box. The following explanation uses two sets of points as an example: a first set consisting of a first vertex located in the upper left position and a second set consisting of a second vertex located in the lower right position. For the first set of points, a preset region corresponding to the target vertex indicated by the average coordinate information is determined, centered on the average coordinate information and with a preset length as the radius. When there are multiple historical frame images, it can be determined whether the historical coordinate information of the historical vertices located in the upper left position (i.e., historical vertices in the same direction as the target vertex) in multiple historical detection boxes are all located within the preset region. If so, the historical coordinate information of the historical vertices (for example, the historical coordinate information of the historical vertices in the most recent historical frame image) is determined as the target coordinate information of the target vertex. If not, the average coordinate information of the target vertex is determined as the target coordinate information of the target vertex, thus obtaining the target coordinate information of the target vertex located in the upper left position. Based on the above process, for the second set of points, the target coordinate information of the target vertex located in the lower right position can be obtained. This gives us the target coordinates of the two vertices located on the same diagonal.

[0135] Finally, the target detection box information is determined based on the target coordinate information corresponding to the two vertices located on the same diagonal. For example, the target coordinate information corresponding to the two vertices can be used to determine the target detection box information; or, the size information and center point coordinate information of the target detection box can be determined based on the target coordinate information corresponding to the two vertices; and the size information and center point coordinate information of the target detection box can be used to determine the target detection box information.

[0136] To alleviate the frequent jitter of the detection box, a preset area corresponding to the target vertex indicated by the average coordinate information is determined. When the historical coordinate information of a historical vertex that is in the same position as the target vertex on the historical detection box is located in the preset area, the historical coordinate information of the historical vertex is determined as the target coordinate information of the target vertex. This ensures that the target detection box information of the image to be processed follows the historical detection box information corresponding to the historical frame image, maintaining the stability of the intermediate detection box and ensuring the stability of the background replacement.

[0137] In S104 and S105, a local image corresponding to the target detection box can be extracted from the image to be processed based on the target detection box information; the local image is then segmented using a trained second neural network for background segmentation to generate a segmented image with the same size as the image to be processed.

[0138] For example, when segmenting an image into a binary image, the first pixel with a first pixel value (e.g., 1) can be identified as a foreground pixel, and the second pixel with a second pixel value (e.g., 0) as a background pixel. Then, for each target pixel in the image to be processed, based on its target location, the pixel value of the pixel in the segmented image that matches that target location can be determined. If the pixel value is the first pixel value, the target pixel is determined to belong to the foreground region; if the pixel value is the second pixel value, the target pixel is determined to belong to the background region. Furthermore, using the segmented image and the set background replacement scheme, the background of the image to be processed can be replaced to generate the target display image.

[0139] In one optional implementation, the background of the image to be processed is replaced based on the segmented image and a set background replacement scheme to generate a target display image, which may include the following two methods:

[0140] Method 1: Adjust the first pixel information of the pixel to be replaced in the background region of the image to be processed, as indicated by the segmented image, to the preset pixel information, and generate the target display image.

[0141] Method 2: Adjust the first pixel information of the pixel to be replaced in the background region of the image to be processed, as indicated by the segmented image, to the second pixel information of the candidate pixel in the preset replacement image that matches the pixel position of the pixel to be replaced, and generate the target display image.

[0142] In Method 1, preset pixel information can be set in advance, such as setting the preset pixel value to 0; the first pixel information of the pixel to be replaced in the background area of ​​the image to be processed indicated by the segmented image is replaced with the preset pixel information, while keeping the pixel information of the foreground pixel in the foreground area other than the background area of ​​the image to be processed unchanged, and generating the target display image.

[0143] In Method 2, the first pixel information of the pixel to be replaced in the background region of the image to be processed, as indicated by the segmented image, is adjusted to the second pixel information of the candidate pixel in the preset replacement image that matches the pixel position of the pixel to be replaced. For example, the first pixel information of the pixel to be replaced in the first row and first column of the image to be processed can be adjusted to the second pixel information of the candidate pixel in the first row and first column of the preset replacement image. The pixel information of the foreground pixels in the foreground region other than the background region in the image to be processed remains unchanged, and the target display image is generated.

[0144] Here, you can easily replace the background of the image to be processed by setting preset pixel information, which is quite efficient. Alternatively, you can determine a preset replacement image, which allows for more flexible background replacement of the image to be processed, offering greater interest and variety.

[0145] Combination Figure 3 An image processing method is illustrated by way of example, which may include the following steps:

[0146] S301: Acquire the image to be processed captured by the display device.

[0147] S302: Perform limb detection on the image to be processed and determine the current detection box information of the user included in the image to be processed.

[0148] S303: Based on at least one of the confidence level, detection box size, and position information indicated by the current detection box information, filter multiple current detection boxes included in the image to be processed to obtain the filtered current detection boxes and the current detection box information corresponding to the filtered current detection boxes.

[0149] In one approach, S304 may be included: for each current detection box after filtering, based on the detection box identifier indicated by the current detection box information corresponding to the current detection box, determine the historical detection box information in the historical frame image that matches the detection box identifier; and based on the current detection box information and the historical detection box information that matches the detection box identifier, determine the target detection box information corresponding to the current detection box.

[0150] During implementation, for each current detection box, based on the current detection box information and historical detection box information matching the detection box identifier, the target detection box information corresponding to the current detection box is determined, which may include:

[0151] First, based on the current detection box information and the historical detection box information that matches the detection box identifier, determine the average coordinate information of two vertices located on the same diagonal.

[0152] Second, take the two vertices located on the same diagonal as the target vertices, and determine the preset area corresponding to the target vertices with the average coordinate information of the target vertices as the center and the preset length as the radius.

[0153] Third, check whether the historical coordinate information of historical vertices that are in the same position as the target vertex on the historical detection frame is located within the preset area corresponding to the target vertex.

[0154] Fourth, if yes, use the historical coordinate information of the historical vertex as the target coordinate information of the target vertex; if no, use the average coordinate information of the target vertex as the target coordinate information of the target vertex.

[0155] Fifth, based on the target coordinate information corresponding to the two vertices on the same diagonal, determine the target detection box information corresponding to the current detection box.

[0156] By repeating steps one through five, we can obtain the target detection box information corresponding to each current detection box, thus obtaining multiple target detection box information.

[0157] S305: Based on the obtained information of multiple target detection boxes, perform background segmentation on the image to be processed to generate the segmented image corresponding to the image to be processed.

[0158] In another approach, S306 may be included: generating intermediate detection box information that surrounds the multiple current detection boxes based on the position information indicated by the filtered current detection box information; and determining a target detection box information based on the intermediate detection box information and historical detection box information.

[0159] During implementation, based on intermediate detection box information and historical detection box information, a target detection box is determined, which may include:

[0160] First, determine whether the number of current detection boxes in the image to be processed matches the number of historical detection boxes in the historical frame images.

[0161] Second, if a match is found, the average coordinates of the two vertices on the same diagonal can be determined based on the intermediate detection box information and the historical detection box information; if a match is not found, the intermediate detection box information can be determined as the target detection box information.

[0162] Third, when the number of current detection boxes matches the number of historical detection boxes included in the historical frame image, two vertices on the same diagonal are taken as target vertices, and the preset region corresponding to the target vertex is determined with the average coordinate information of the target vertex as the center and the preset length as the radius.

[0163] Fourth, check whether the historical coordinate information of historical vertices that are in the same position as the target vertex on the historical detection frame are located within the preset area corresponding to the target vertex.

[0164] Fifth, if yes, use the historical coordinate information of the historical vertex as the target coordinate information of the target vertex; if no, use the average coordinate information of the target vertex as the target coordinate information of the target vertex.

[0165] Sixth, based on the target coordinate information corresponding to the two vertices on the same diagonal, a target detection box is determined.

[0166] S307: Based on the obtained target detection box information, perform background segmentation on the image to be processed to generate a segmented image corresponding to the image to be processed.

[0167] S308: Based on the segmented image and the set background replacement scheme, perform background replacement on the image to be processed to generate the target display image.

[0168] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0169] Based on the same concept, this disclosure also provides an image processing apparatus, see [link to relevant documentation]. Figure 4 The diagram shown is a schematic representation of the architecture of an image processing apparatus provided in this embodiment of the present disclosure, including an acquisition module 401, a first determination module 402, a second determination module 403, a segmentation module 404, and a generation module 405. Specifically:

[0170] The acquisition module 401 is used to acquire the image to be processed collected by the display device;

[0171] The first determining module 402 is used to perform limb detection on the image to be processed and determine the current detection box information of the user included in the image to be processed;

[0172] The second determining module 403 is used to determine the target detection box information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information;

[0173] The segmentation module 404 is used to perform background segmentation on the image to be processed based on the target detection box information, and generate a segmented image corresponding to the image to be processed; the segmented image is used to indicate at least one of the background region and the foreground region in the image to be processed;

[0174] The generation module 405 is used to perform background replacement on the image to be processed based on the segmented image and the set background replacement scheme to generate a target display image.

[0175] In one possible implementation, the generation module 404, when performing background replacement on the image to be processed based on the segmented image and the set background replacement scheme to generate the target display image, is used to:

[0176] The first pixel information of the pixels to be replaced belonging to the background region on the image to be processed, as indicated by the segmented image, is adjusted to preset pixel information to generate the target display image; or...

[0177] The first pixel information of the pixel to be replaced in the background region of the image to be processed, as indicated by the segmented image, is adjusted to the second pixel information of the candidate pixel in the preset replacement image that matches the pixel position of the pixel to be replaced, thereby generating the target display image.

[0178] In one possible implementation, the number of current detection boxes is multiple, and the second determining module 403, when determining the target detection box information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information, is used to:

[0179] For each current detection box included in the image to be processed, based on the detection box identifier indicated by the current detection box information corresponding to the current detection box, the historical detection box information in the historical frame image that matches the detection box identifier is determined; and based on the current detection box information and the historical detection box information that matches the detection box identifier, the target detection box information corresponding to the current detection box is determined.

[0180] The segmentation module 404, when performing background segmentation on the image to be processed based on the target detection box information to generate a segmented image corresponding to the image to be processed, is used for:

[0181] Based on the obtained target detection box information, the background of the image to be processed is segmented to generate the segmented image corresponding to the image to be processed.

[0182] In one possible implementation, the number of current detection boxes is multiple, and the device further includes a filtering module 406, which is used to filter the multiple current detection boxes included in the image to be processed based on at least one of the confidence level, detection box size, and position information indicated by the current detection box information after determining the current detection box information of the user included in the image to be processed, to obtain the filtered current detection boxes and the current detection box information corresponding to the filtered current detection boxes;

[0183] The second determining module 403, when determining the target detection box information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information, is used to:

[0184] Based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information corresponding to the current detection box after filtering, the target detection box information is determined.

[0185] In one possible implementation, the number of current detection boxes is multiple, and the second determining module 403, when determining the target detection box information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information, is used to:

[0186] Based on the position information indicated by each of the current detection boxes, generate intermediate detection box information that surrounds the multiple current detection boxes;

[0187] Based on the intermediate detection box information and the historical detection box information, the target detection box information is determined.

[0188] In one possible implementation, the second determining module 403, when determining the target detection box information based on the intermediate detection box information and the historical detection box information, is used to:

[0189] In response to the matching of the number of current detection boxes in the image to be processed with the number of historical detection boxes in the historical frame image, at least one set of average coordinate information corresponding to a point set is determined based on the intermediate detection box information and the historical detection box information; wherein, the point set includes multiple vertices that are in the same position as the detection boxes and belong to the intermediate detection boxes and the historical detection boxes respectively.

[0190] Based on the obtained average coordinate information, the target detection box information is determined.

[0191] In one possible implementation, when the at least one set of points includes two sets of points diagonally related to the detection box, the second determining module 403, when determining the target detection box information based on the obtained average coordinate information, is configured to:

[0192] For each set of points, a preset region corresponding to the target vertex indicated by the average coordinate information is determined, centered on the average coordinate information and with a preset length as the radius; and

[0193] If the historical coordinate information of a historical vertex that is in the same position as the target vertex on the historical detection frame is located in the preset area corresponding to the target vertex, the historical coordinate information of the historical vertex is determined as the target coordinate information of the target vertex.

[0194] If the historical coordinate information of a historical vertex that is in the same position as the target vertex on the historical detection box is not located in the preset area corresponding to the target vertex, the average coordinate information of the target vertex is determined as the target coordinate information of the target vertex.

[0195] Based on the determined target coordinate information, the target detection box information is determined.

[0196] In some embodiments, the functions or templates of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0197] Based on the same technical concept, this disclosure also provides an electronic device. (See also...) Figure 5 The diagram shows the structure of an electronic device provided in this embodiment, including a processor 501, a memory 502, and a bus 503. The memory 502 stores execution instructions and includes a main memory 5021 and an external memory 5022. The main memory 5021, also called internal memory, is used to temporarily store computational data in the processor 501 and data exchanged with external memory such as a hard disk. The processor 501 exchanges data with the external memory 5022 through the main memory 5021. When the electronic device 500 is running, the processor 501 and the memory 502 communicate through the bus 503, causing the processor 501 to execute the following instructions:

[0198] Acquire the image to be processed from the display device;

[0199] Perform limb detection on the image to be processed to determine the current detection box information of the user included in the image to be processed;

[0200] Based on the historical detection box information corresponding to the historical frame images collected by the display device, and the current detection box information, the target detection box information is determined;

[0201] Based on the target detection box information, the image to be processed is segmented to generate a segmented image corresponding to the image to be processed; the segmented image is used to indicate at least one of the background region and the foreground region in the image to be processed;

[0202] Based on the segmented image and the set background replacement scheme, the background of the image to be processed is replaced to generate the target display image.

[0203] The specific processing flow of the processor 501 can be referred to the description in the above method embodiment, and will not be repeated here.

[0204] Furthermore, this disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the image processing method described in the above-described method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.

[0205] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the image processing method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0206] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0207] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0208] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0209] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0210] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0211] The above are merely specific embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. An image processing method, characterized in that, include: Acquire the image to be processed from the display device; Perform limb detection on the image to be processed to determine the current detection box information of the user included in the image to be processed; Based on the historical detection box information corresponding to the historical frame images collected by the display device, and the current detection box information, the target detection box information is determined; Based on the target detection box information, the background of the image to be processed is segmented to generate a segmented image corresponding to the image to be processed; The segmented image is used to indicate at least one of the background and foreground regions in the image to be processed; Based on the segmented image and the set background replacement scheme, the background of the image to be processed is replaced to generate the target display image; The number of current detection boxes is multiple. The determination of target detection box information based on historical detection box information corresponding to historical frame images acquired by the display device and the current detection box information includes: Based on the position information indicated by each of the current detection box information, generate intermediate detection box information surrounding multiple current detection boxes; based on the intermediate detection box information and the historical detection box information, determine the target detection box information; The step of determining the target detection box information based on the intermediate detection box information and the historical detection box information includes: in response to the matching of the number of current detection boxes included in the image to be processed with the number of historical detection boxes included in the historical frame image, determining the average coordinate information corresponding to at least one set of points based on the intermediate detection box information and the historical detection box information; wherein, the set of points includes multiple vertices located in the same direction as the detection boxes and belonging to the intermediate detection boxes and the historical detection boxes respectively; and determining the target detection box information based on the obtained average coordinate information.

2. The method according to claim 1, characterized in that, The process of replacing the background of the image to be processed based on the segmented image and the set background replacement scheme to generate a target display image includes at least one of the following: The first pixel information of the pixel to be replaced in the background region of the image to be processed, as indicated by the segmented image, is adjusted to preset pixel information to generate the target display image; The first pixel information of the pixel to be replaced in the background region of the image to be processed, as indicated by the segmented image, is adjusted to the second pixel information of the candidate pixel in the preset replacement image that matches the pixel position of the pixel to be replaced, thereby generating the target display image.

3. The method according to claim 1 or 2, characterized in that, The number of current detection boxes is multiple. The determination of target detection box information based on historical detection box information corresponding to historical frame images acquired by the display device and the current detection box information includes: For each current detection box included in the image to be processed, based on the detection box identifier indicated by the current detection box information corresponding to the current detection box, the historical detection box information in the historical frame image that matches the detection box identifier is determined; and based on the current detection box information and the historical detection box information that matches the detection box identifier, the target detection box information corresponding to the current detection box is determined. The step of performing background segmentation on the image to be processed based on the target detection box information to generate a segmented image corresponding to the image to be processed includes: Based on the obtained target detection box information, the background of the image to be processed is segmented to generate the segmented image corresponding to the image to be processed.

4. The method according to claim 1, characterized in that, The number of current detection boxes is multiple. After determining the current detection box information of the user included in the image to be processed, the method further includes: Based on at least one of the confidence level, detection box size, and position information indicated by the current detection box information, multiple current detection boxes included in the image to be processed are filtered to obtain the filtered current detection boxes and the current detection box information corresponding to the filtered current detection boxes. The determination of target detection box information based on historical detection box information corresponding to historical frame images acquired by the display device and the current detection box information includes: Based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information corresponding to the current detection box after filtering, the target detection box information is determined.

5. The method according to claim 1, characterized in that, When at least one set of points includes two sets of points diagonally related to the detection box, determining the target detection box information based on the obtained average coordinate information includes: For each set of points, a preset region corresponding to the target vertex indicated by the average coordinate information is determined, centered on the average coordinate information and with a preset length as the radius; and If the historical coordinate information of a historical vertex that is in the same position as the target vertex on the historical detection frame is located in the preset area corresponding to the target vertex, the historical coordinate information of the historical vertex is determined as the target coordinate information of the target vertex. If the historical coordinate information of a historical vertex that is in the same position as the target vertex on the historical detection box is not located in the preset area corresponding to the target vertex, the average coordinate information of the target vertex is determined as the target coordinate information of the target vertex. Based on the determined target coordinate information, the target detection box information is determined.

6. An image processing apparatus, characterized in that, include: The acquisition module is used to acquire the image to be processed captured by the display device; The first determining module is used to perform limb detection on the image to be processed and determine the current detection box information of the user included in the image to be processed; The second determining module is used to determine the target detection box information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information; The segmentation module is used to perform background segmentation on the image to be processed based on the target detection box information, and generate a segmented image corresponding to the image to be processed; the segmented image is used to indicate at least one of the background region and the foreground region in the image to be processed; The generation module is used to perform background replacement on the image to be processed based on the segmented image and the set background replacement scheme, and generate a target display image; The number of current detection boxes is multiple. When determining the target detection box information based on the historical detection box information corresponding to the historical frame images collected by the display device and the current detection box information, the second determining module is used for: Based on the position information indicated by each of the current detection boxes, generate intermediate detection box information that surrounds the multiple current detection boxes; Based on the intermediate detection box information and the historical detection box information, the target detection box information is determined; The second determining module, when determining target detection box information based on the intermediate detection box information and the historical detection box information, is configured to: in response to a match between the number of current detection boxes included in the image to be processed and the number of historical detection boxes included in the historical frame image, determine at least one set of average coordinate information corresponding to a point set based on the intermediate detection box information and the historical detection box information; wherein, the point set includes multiple vertices located in the same direction as the detection boxes and belonging to the intermediate detection boxes and the historical detection boxes respectively; Based on the obtained average coordinate information, the target detection box information is determined.

7. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the image processing method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the image processing method as described in any one of claims 1 to 5.