Image processing method and device for high-precision automatic optical detection equipment

Through object detection and superposition and fusion technology, the image information of an object at multiple moments is combined into a target image, solving the problem of image data synthesis in the prior art, and achieving efficient synthesis and uniform layout of image information.

CN120374792APending Publication Date: 2025-07-25SHENZHEN TCL HIGH TECH DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410106205.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of combining the image data of an object at multiple times.

Method used

The object detection technology accurately extracts the image information of the object at multiple moments, and combines these image information into one target image information through superposition and fusion technology. The image layout is optimized using detection frame technology to avoid mutual overlay of images during the fusion process, and retains the details of each image to the greatest extent.

Benefits of technology

It is realized that the image information of the object at multiple moments is efficiently combined into a target image, and the layout of each subject in the image is uniform, avoiding the mutual overlay of the images during the fusion process and retaining the details of each image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374792A_ABST
    Figure CN120374792A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device for high-precision automatic optical detection equipment, and the method and device carry out the target detection and superposition fusion of image information in to-be-processed image information, and obtain target image information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine vision technology, and particularly to an image processing method and apparatus. Background Art

[0002] With the development of big data and AI technologies, the ability of the Internet to process and analyze data has been greatly improved at the present stage. The data volume owned by mobile devices is also diverse and complex, and various AI products emerge in an endless stream. How to use AI technology to extract image data of an object at multiple moments from a large amount of resource data, and how to synthesize the image data of the object at multiple moments into a synthesized image data containing the image data of the object at multiple moments has become an urgent problem to be solved. Summary of the Invention

[0003] Embodiments of this application provide an image processing method and apparatus.

[0004] In a first aspect, embodiments of this application provide an image processing method, including:

[0005] Obtaining image information to be processed;

[0006] Performing object detection on the image information in the image information to be processed to obtain first processed image information;

[0007] Performing superimposed fusion on the first processed image information to obtain target image information.

[0008] In a second aspect, embodiments of this application provide an image processing apparatus, including:

[0009] An obtaining module, configured to obtain image information to be processed;

[0010] An object detection module, configured to perform object detection on the image information in the image information to be processed to obtain first processed image information;

[0011] A fusion module, configured to perform superimposed fusion on the first processed image information to obtain target image information.

[0012] In a third aspect, embodiments of this application further provide an electronic device, including a memory storing multiple computer programs; a processor loads the computer programs from the memory to execute any one of the image processing methods provided by the embodiments of this application.

[0013] In a fourth aspect, embodiments of this application further provide a computer-readable storage medium storing multiple computer programs, and the computer programs are suitable for being loaded by a processor to execute any one of the image processing methods provided by the embodiments of this application.

[0014] Fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program, which when executed by a processor, implements any one of the image processing methods provided by the embodiments of the present application. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0016] Figure 1 is a schematic flowchart of the image processing method provided in the embodiment of the present application;

[0017] Figure 2 is an example diagram of the result output by the detection model provided in the embodiment of the present application;

[0018] Figure 3 is a schematic diagram of the detection frame selection frame provided in the embodiment of the present application;

[0019] Figure 4 is an example diagram of the cropping of a non-congested scene provided in the embodiment of the present application;

[0020] Figure 5 is an example diagram of the cropping of a congested scene provided in the embodiment of the present application;

[0021] Figure 6 is an example diagram of the stitching of multiple frames of images provided in the embodiment of the present application;

[0022] Figure 7 is a schematic structural diagram of the image processing device provided in the embodiment of the present application;

[0023] Figure 8 is a schematic structural diagram of the electronic device provided in the embodiment of the present application. Detailed Embodiments

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application. At the same time, in the description of the embodiments of the present application, terms such as "first" and "second" are only used for differential description and cannot be understood as indicating or implying relative importance. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, the meaning of "a plurality of" is two or more, unless otherwise specifically defined.

[0025] The embodiments of the present application provide an image processing method and apparatus. Specifically, the embodiments of the present application will be described from the perspective of an image processing apparatus, which can be specifically integrated in an electronic device, that is, the image processing method in the embodiments of the present application can be executed by the electronic device. Optionally, the electronic device includes a terminal device. The terminal device can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a game console, or a personal computer (PC) and other devices. Optionally, the electronic device includes a server, which can be an independent server or a server network or server cluster composed of servers, including but not limited to a computer, a network host, a single network server, a set of network servers, or a cloud server composed of servers. Among them, the cloud server is composed of a large number of computers or network servers based on cloud computing (Cloud Computing).

[0026] It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from that shown in the drawings.

[0027] In order to make the final target image information contain the image information of an object at multiple moments, the embodiments of the present application provide an image processing method. Optionally, the embodiments of the present application are described with an image processing apparatus as the execution subject, and a mobile phone terminal is used as an example of the image processing apparatus in the embodiments of the present application. The following will be described in detail with reference to the accompanying drawings respectively. Refer to Figure 1 , Figure 1 is a schematic flowchart of the image processing method provided in the embodiments of the present application. The specific flow steps of the image processing method provided in the embodiments of the present application can be as steps 10 to 30, including:

[0028] Step 10, obtain the image information to be processed.

[0029] Optionally, the image information to be processed includes a video data stream and a sequence of continuous-shot pictures. Therefore, it can be understood that the mobile terminal acquires the video data stream and the sequence of continuous-shot pictures through a shooting device such as a camera or a video camera. Further, the embodiment of the present application is described with a video data stream.

[0030] Step 20: Perform target detection on the image information in the image information to be processed to obtain first processed image information.

[0031] Optionally, the mobile terminal parses the video data stream to obtain each frame image in the video data stream, and performs target detection on each frame image in the video data stream, wherein the target detection may include subject detection and subject frame detection, and the subject detection may include human body detection, sphere detection, etc.

[0032] Therefore, the embodiment of the present application can be understood as that the mobile terminal performs human body detection, sphere detection, etc. on each frame of the video data stream, selects each frame of the video data stream containing subject information (human body, sphere, etc.), and obtains a set of images to be processed containing subject information in the video data stream. Further, the mobile terminal performs subject frame detection on the set of images to be processed to obtain a first processed image that meets the subject frame requirements, as described in steps 201 to 203.

[0033] Step 30: superimpose and fuse the first processed image information to obtain target image information.

[0034] Optionally, the mobile terminal superimposes and fuses each image in the first processed image to obtain a synthesized image, wherein it can be understood that each image in the first processed image is a portrait clone image of the subject. Therefore, the mobile terminal superimposes and fuses each image in the first processed image, which actually superimposes and fuses multiple portrait clone images of the subject to obtain a synthesized image superimposed by multiple portrait clones, as described in detail in steps 301 to 303.

[0035] The embodiment of the present application uses target detection technology to accurately extract image information of objects at multiple moments in the image information to be processed, and then uses superposition and fusion technology to superimpose and fuse the image information of the object at multiple moments into a target image information, so that the final target image information contains the image information of the object at multiple moments.

[0036] In an optional embodiment, in order to make the layout of various subjects in the composite image more uniform, the embodiment of the present application needs to extract the image data of the objects in the resource data at multiple times through the target detection technology and the detection frame technology. The description of steps 201 to 203 is as follows:

[0037] Step 201: Perform object detection on each frame of image information in the to-be-processed image information, and output the second processed image information in the to-be-processed image information;

[0038] Step 202: Based on the detection box information of each first to-be-processed image in the second processed image information, determine the detection box set information;

[0039] Step 203: Determine the first processed image information according to the detection box set information.

[0040] Optionally, the mobile terminal parses the video data stream to obtain each frame of image in the video data stream, and performs object detection on each frame of image in the video data stream. Among them, the object detection can be human detection, sphere detection, etc. In one embodiment, the object detection is human detection. Therefore, it can be understood that the mobile terminal performs human detection on each frame of image in the video data stream, selects each frame of image containing a human in the video data stream, and obtains the second processed image containing a human in the video data stream. Therefore, the second processed image includes several first to-be-processed images.

[0041] Further, the mobile terminal performs object box detection on each first to-be-processed image in the second processed image to obtain the detection box of each first to-be-processed image in the second processed image. In one embodiment, the object box detection can be performed through a detection model. The detection model in the embodiments of the present application can be a detection network based on YOLOX. The object categories can include humans, spheres, etc. Among them, the spheres include basketballs, footballs, etc. Basketballs, footballs, etc. can be regarded as a category and represented by ball. Therefore, the embodiments of the present application can be understood as that the mobile terminal inputs each first to-be-processed image in the second processed image into the detection model for object box detection, obtains the detection box of each first to-be-processed image in the second processed image output by the detection model, and obtains the detection box set. The upper left coordinate of the human detection box is represented as (x0_human, y0_human), the lower right coordinate is represented as (x1_human, y1_human), the upper left coordinate of the sphere detection box is represented as (x0_ball, y0_ball), and the lower right coordinate of the sphere detection box is represented as (x1_ball, y1_ball). Specifically, refer to Figure 2 , Figure 2 which is the schematic diagram of the detection box provided by the embodiments of the present application.

[0042] Further, the embodiments of the present application are described by taking only human body box detection as an example. Therefore, after performing human body box detection, the obtained set of detected box frames is a set of human body detection box frames. Therefore, the mobile terminal selects each first image to be processed in the second processed image according to the set of human body detection box frames to obtain the first processed image, as specifically described in steps 2031 to 2033.

[0043] The embodiments of the present application accurately extract the image data of an object in resource data at multiple moments through object detection technology and detection box technology, making the layout between each main body in the synthesized image more uniform.

[0044] In an alternative embodiment, the descriptions of steps 2031 to 2033 are as follows:

[0045] Step 2031: Obtain the vertex position coordinate information of each piece of detection box information in the set of detection box information;

[0046] Step 2032: Determine the central position coordinate information of each piece of detection box information according to the vertex position coordinate information of each piece of detection box information;

[0047] Step 2033: Determine the first processed image information according to the central position coordinate information of each piece of detection box information.

[0048] It should be noted that when a human body is moving, the moving direction may be from left to right, from right to left, or may walk back and forth from side to side. Therefore, adjacent frames may intersect or not intersect. For specific reference, see Figure 7 . Based on the above various situations and considering the main body layout as a priority, frames are uniformly selected for uniform layout based on the horizontal distance between the centers of the human body detection box frames, as follows:

[0049] Optionally, the mobile terminal obtains the vertex position coordinates of each human body detection box in the set of human body detection box frames, that is, the upper left corner coordinates and the lower right corner coordinates of the human body detection box.

[0050] Among them, the upper left corner coordinates of the human body detection box are represented as (x0_human, y0_human), and the lower right corner coordinates are represented as (x1_human, y1_human).

[0051] Further, the mobile terminal calculates the center position coordinates of each human detection box according to the upper left corner coordinates and the lower right corner coordinates of each human detection box. In the application embodiment, frame selection is evenly arranged by the horizontal distance. Therefore, it can be understood that the mobile terminal calculates the horizontal coordinate of the center position of each human detection box according to the upper left corner coordinates and the lower right corner coordinates of each human detection box. The horizontal coordinate of the center position can be expressed as Center, and the specific calculation formula is:

[0052] Center = (x0_human + x1_human) / 2

[0053] The mobile terminal selects the images that meet the requirements according to the horizontal coordinates of the center positions of each human detection box, and obtains the first processed image, as specifically described in steps 20331 to 20333.

[0054] In the application embodiment, the object image data at multiple moments in the resource data is accurately extracted through the object detection technology and the detection box technology, so that the layout between the main bodies in the synthesized image is more uniform.

[0055] In an optional embodiment, the descriptions of steps 20331 to 20333 are as follows:

[0056] Step 20331: Sort the center position coordinate information of each detection box information to obtain the maximum center position coordinate information and the minimum center position coordinate information;

[0057] Step 20332: Determine the position interval information according to the maximum center position coordinate information and the minimum center position coordinate information;

[0058] Step 20333: Determine the first processed image information according to the position interval information, the maximum center position coordinate information, the minimum center position coordinate information, and the center position coordinate information of each detection box information.

[0059] Optionally, the mobile terminal sorts the center position coordinates of each human detection box in ascending or descending order of the coordinate values to obtain a coordinate list center_points of the center position coordinates of the human detection box, and obtains the maximum center position coordinate and the minimum center position coordinate from the coordinate list center_points.

[0060] In the embodiment of the present application, the central position coordinate is the abscissa of the central position. Therefore, it can be understood that the abscissas of the central positions of each human detection box are sorted in ascending or descending order of the coordinate values to obtain a coordinate list center_points of the central position coordinates of the human detection boxes. The maximum abscissa of the central position and the minimum abscissa of the central position are obtained from the coordinate list center_points. Therefore, the maximum abscissa of the central position can be expressed as Center_max = max(center_points), and the minimum abscissa of the central position can be expressed as Center_min = min(center_points).

[0061] Further, the mobile terminal obtains the number of images H of the user on the terminal interface or the pre-set clone image, and calculates a position interval according to the maximum abscissa of the central position, the minimum abscissa of the central position, and the number of images H. Among them, the position interval can be understood as the average interval, and the position interval can be expressed as interval. The calculation formula for the position interval interval is:

[0062] interval = (max(center_points) - min(center_points)) / (H - 1)

[0063] Further, the mobile terminal determines the first processed image according to the position interval, the maximum abscissa of the central position, the minimum abscissa of the central position, and the abscissa of the central position of each human detection box. In one embodiment, when the number of images H is 4, the specific process is as follows:

[0064] Optionally, the mobile terminal determines the human detection box corresponding to the minimum abscissa of the central position as the first human detection box, that is, the leftmost human detection box, and determines the human detection box corresponding to the maximum abscissa of the central position as the second human detection box, that is, the rightmost human detection box. Therefore, the image corresponding to the leftmost human detection box is determined as the first clone image in the first processed image, and the image corresponding to the rightmost human detection box is determined as the fourth clone image in the first processed image. Therefore, it is also necessary to determine 2 human detection boxes from the remaining human detection box set except the first human detection box and the second human detection box, that is, the third human detection box and the fourth human detection box.

[0065] The specific process for determining the third human detection box is as follows:

[0066] First, determine the abscissa of the optimal center position of the third human detection box. The calculation formula for the abscissa of the optimal center position is: the minimum abscissa of the center position + (i - 1) * position interval, that is, min(center_points)+(i - 1)*interval, where i ∈ [2, H - 1], and i is determined by the number of images H. In the embodiment of the present application, H = 4. Therefore, i can be 2 or 3. At this time, i = 2. Therefore, the abscissa of the optimal center position of the third human detection box is min(center_points)+interval.

[0067] Further, calculate the abscissa distance from the abscissa of the center position of each human detection box in the remaining human detection box set to the abscissa of the optimal center position, and determine the human detection box corresponding to the smallest abscissa distance as the third human detection box. Therefore, the determination formula for the i-th human detection box is min{center_i - [min(center_points)+(i - 1)*interval]}, where center_i represents the abscissa of the center position of all human detection boxes in the remaining human detection box set except for the leftmost and rightmost human detection boxes. Further, determine the image corresponding to the third human detection box as the second clone image in the first processed image.

[0068] Similarly, the specific process for determining the fourth human detection box is as follows: Substitute i = 3 into the formula min{center_i - [min(center_points)+(i - 1)*interval]}, and the fourth human detection box in the remaining human detection box set can be determined. Further, determine the image corresponding to the fourth human detection box as the third clone image in the first processed image.

[0069] In the embodiment of the present application, the image data of the object in the resource data at multiple moments is accurately extracted through the object detection technology and the detection box technology, so that the layout between the main bodies in the synthesized image is more uniform.

[0070] In an alternative embodiment, in order to avoid mutual coverage of images during the fusion process and retain the image details of each clone image to the greatest extent, steps 301 to 303 are described as follows:

[0071] Step 301, obtain the first target image corresponding to the maximum center position coordinate information and the second target image corresponding to the minimum center position coordinate information;

[0072] Step 302: Determine the image data intersection area information and the image data union area information between the first target image and the second target image;

[0073] Step 303: Based on the image data intersection area information and the image data union area information, superimpose and fuse the first processed image information to obtain the target image information.

[0074] It should be noted that when splicing multiple pictures, it is necessary to consider that there should be overlapping areas and non-overlapping areas between two pictures. At the same time, in order to ensure that the portrait areas of the two pictures can be found in the corresponding frames, therefore, in the embodiments of the present application, each to-be-processed image in the first processed image needs to be cropped so that the background after splicing is more unified. The specific process is as follows:

[0075] Optionally, the mobile terminal obtains the first target image corresponding to the maximum center position coordinate in the first processed image, and the second target image corresponding to the minimum center position coordinate. The center position abscissa in the embodiments of the present application can thus be understood as that the mobile terminal obtains the first target image corresponding to the maximum center position abscissa in the first processed image, and the second target image corresponding to the minimum center position abscissa.

[0076] Further, the mobile terminal calculates the image intersection area and the image union area between the first target image and the second target image. The processes of calculating the image intersection area and the image union area are as follows:

[0077] Regarding the process of calculating the image intersection area:

[0078] The detection box corresponding to the maximum center position coordinate is box a, and the detection box corresponding to the minimum center position coordinate is box b. Among them, the coordinates of box a are (x1_a, y1_a, x2_a, y2_a), and the coordinates of box b are (x1_b, y1_b, x2_b, y2_b), where x1_a < x2_a and x1_b < x2_b. First, find the upper left corner coordinates of the overlapping part of the two detection boxes of box a and box b, that is, the upper left corner coordinates of the overlapping part are (max(x1_a, x1_b), max(y1_a, y1_b)), and find the lower right corner coordinates of the overlapping part of the two detection boxes of box a and box b, that is, the lower right corner coordinates of the overlapping part are (min(x2_a, x2_b), min(y2_a, y2_b)).

[0079] Further, calculate the width of the overlapping part of the two detection boxes of box a and box b in the horizontal direction, where the width can be expressed as x2_overlap - x1_overlap, and x2_overlap - x1_overlap = (min(x2_a, x2_b) - max(x1_a, x1_b)). Further, calculate the height of the overlapping part of the two detection boxes of box a and box b in the vertical direction, where the height can be expressed as y2_overlap - y1_overlap, and y2_overlap - y1_overlap = (min(y2_a, y2_b) - max(y1_a, y1_b)).

[0080] Further, calculate the intersection area intersection_area of the two detection boxes of box a and box b. The intersection area intersection_area = the width of the overlapping part * the height of the overlapping part. The finally obtained intersection area is the image intersection area.

[0081] Regarding the process of calculating the image union area:

[0082] Calculate the area area_a of box a and the area area_b of box b, where area_a = (x2_a - x1_a) * (y2_a - y1_a), and area_b = (x2_b - x1_b) * (y2_b - y1_b). Further, calculate the union area union_area of the two detection boxes of box a and box b, where union_area = area_a + area_b - intersection_area.

[0083] Further, the mobile terminal calculates the overlapping degree between the first target image and the second target image according to the image intersection area and the image union area between the first target image and the second target image, and superimposes and fuses all the split images in the first processed image according to the overlapping degree between the first target image and the second target image to obtain a synthesized image including multiple split images, as specifically described in steps 3031 to 3034.

[0084] In the embodiment of the present application, the image information of an object at multiple moments is superimposed and fused into one piece of target image information through the superimposing and fusing technology. Therefore, the finally obtained one piece of target image information contains the image information of the object at multiple moments. During the fusion process, cropping is performed through the image intersection area and the image union area between the first target image and the second target image, avoiding the mutual covering of images during the fusion process and retaining the image details of each split image as completely as possible to the greatest extent.

[0085] In an alternative embodiment, the descriptions of steps 3031 to 3034 are as follows:

[0086] Step 3031: Sort each second image to be processed in the first processed image information in the order from front to back according to the time information to obtain sorted image set information;

[0087] Step 3032: Based on the image data intersection area information and the image data union area information, crop each third image to be processed in the sorted image set information to obtain cropped image set information;

[0088] Step 3033: Obtain a background image based on the cropped image set information;

[0089] Step 3034: Using the background image as the background, superimpose and fuse each cropped image in the cropped image set information to obtain the target image information.

[0090] Optionally, the mobile terminal sorts each second image to be processed in the first processed image in the order from front to back according to the time to obtain a sorted image set.

[0091] Further, the mobile terminal calculates the intersection - union ratio between the first target image and the second target image according to the image intersection area and the image union area between the first target image and the second target image.

[0092] Further, compare the numerical size of the intersection - union ratio between the first target image and the second target image with the intersection - union ratio threshold to obtain a comparison result, where the intersection - union ratio threshold is, for example, 0.3, 0.35, etc.

[0093] If the comparison result is that the intersection - union ratio is greater than the intersection - union ratio threshold, it is determined that the second image to be processed in the first processed image is a crowded scene, and then each third image to be processed in the sorted image set information is cropped by the cropping strategy for the crowded scene to obtain a cropped image set. If the comparison result is that the intersection - union ratio is less than or equal to the intersection - union ratio threshold, it is determined that the second image to be processed in the first processed image is a non - crowded scene, and then each third image to be processed in the sorted image set information is cropped by the cropping strategy corresponding to the non - crowded scene to obtain a cropped image set, as specifically described in steps 30321 to 30323.

[0094] Further, the mobile terminal performs Alignment processing and Stitching processing on each cropped image in the cropped image set information to obtain the aligned image corresponding to each cropped image in the cropped image set information and the unified background image of all cropped images in the cropped image set respectively.

[0095] Further, the mobile terminal uses the background image as the background and superimposes and fuses each of the cropped image set information to obtain a synthesized image.

[0096] In the embodiment of the present application, the image information of an object at multiple moments is superimposed and fused into one piece of target image information through the superimposing and fusing technology. Therefore, the final one piece of target image information contains the image information of the object at multiple moments. During the fusion process, cropping is performed based on the image intersection area and the image union area between the first target image and the second target image, avoiding the mutual coverage of images during the fusion process and retaining the image details of each split image to the greatest extent.

[0097] In an alternative embodiment, the descriptions of steps 30321 to 30323 are as follows:

[0098] Step 30321: Determine the intersection-union ratio information of the image data based on the image data intersection area information and the image data union area information;

[0099] Step 30322: If the intersection-union ratio information of the image data is less than or equal to the intersection-union ratio threshold information, crop each of the third images to be processed in the sorted image set information with the first cropping information to obtain the cropped image set information;

[0100] Step 30323: If the intersection-union ratio information of the image data is greater than the intersection-union ratio threshold information, crop each of the third images to be processed in the sorted image set information with the second cropping information to obtain the cropped image set information.

[0101] Optionally, the mobile terminal calculates the intersection-union ratio between the first target image and the second target image according to the image intersection area and the image union area between the first target image and the second target image.

[0102] Further, compare the intersection-union ratio between the first target image and the second target image with the intersection-union ratio threshold to obtain a comparison result, where the intersection-union ratio threshold is, for example, 0.3, 0.35, etc.

[0103] Further, if the comparison result is that the intersection-union ratio is less than or equal to the intersection-union ratio threshold, it is determined that the images in the first processed image are in a non-congested scenario, and then each of the third images to be processed in the sorted image set information is cropped with the first cropping strategy corresponding to the non-congested scenario to obtain the cropped image set. The specific process is as follows:

[0104] Continuing with the above embodiment, the number H of the split images is 4. Therefore, the first split image, the second split image, the third split image, and the fourth split image in the first processed image obtained through target detection are specifically referred toFigure 4 The first row of images in . For the first split image, taking the right border of the first split image as the boundary, crop a preset proportion of the area to the left. The preset proportion = cropping width / total width of the first split image, such as 1 / 5, 1 / 4, etc. For the second, third, and fourth split images, taking the human body as the boundary, crop all the areas to the left of the human body. Figure 4 The second row of images in are the pre-cropped images of the first, second, third, and fourth split images. Figure 4 The third row of images in are the cropped images of the first, second, third, and fourth split images.

[0105] Further, if the comparison result shows that the intersection over union is greater than the intersection over union threshold, it is determined that the images in the first processed image are in a clustering scenario. Then, each third image to be processed in the sorted image set information is cropped according to the second cropping strategy corresponding to the clustering scenario to obtain a cropped image set. The specific process is as follows:

[0106] Continuing with the above embodiment, the number of images H of the split images is 4. Therefore, the first, second, third, and fourth split images in the first processed image obtained through object detection are specifically referred to Figure 5 The first row of images in . The widths of the first, second, third, and fourth split images are W, the height is h, the minimum center position abscissa is Center_min, the maximum center position abscissa is Center_max, and the left side is evenly divided into L = Center_min / (H - 1). In this scenario, the left side is evenly divided into L = Center_min / 3, and the right side is evenly divided into R = (W - Center_max) / (H - 1). In this scenario, the right side is evenly divided into R = (W - Center_max) / 3. Therefore, the cropping range of the first split image is [0, Center_max], the cropping range of the second split image is [Center_min / 3, W - (2R) / 3], the cropping range of the third split image is [(2 * Center_min) / 3, W - R / 3], and the cropping range of the fourth split image is [Center_min, W]. Figure 5 The second row of images in are the pre-cropped images of the first, second, third, and fourth split images. Figure 5 The third row of images in are the cropped images of the first, second, third, and fourth split images.

[0107] The embodiment of the present application performs cropping through the intersection-and-union ratio between the first target image and the second target image, thereby avoiding mutual overlap of images during the fusion process and retaining the image details of each clone image to the greatest extent possible.

[0108] In an optional embodiment, after obtaining a synthesized image including a plurality of clone images, the position of the human body at each moment can be tracked according to each clone image in the synthesized image. Therefore, after superimposing and fusing the first processed image information to obtain the target image information, the method further includes:

[0109] Obtaining each clone image in the target image information;

[0110] The position of the target at each moment is tracked based on each clone image.

[0111] Optionally, the mobile terminal responds to the location tracking request triggered by the user, and obtains each human clone image in the synthesized image according to the location tracking request, wherein each human clone image corresponds to its location information at different times. Therefore, the mobile terminal can track the position of the human body at each time according to the location information of each human clone image. In one embodiment, there are 4 human clone images in the synthesized image, the first human clone image corresponds to the location information position1 of 2023-12-25-10:00, the second human clone image corresponds to the location information position2 of 2023-12-25-10:01, the third human clone image corresponds to the location information position3 of 2023-12-25-10:03, and the fourth human clone image corresponds to the location information position4 of 2023-12-25-10:04. Therefore, if the location tracking request requires the location tracking of the third human clone image, the location tracking information obtained is {2023-12-25-10:03, position3}.

[0112] Therefore, it can be understood that the embodiments of the present application may include a subject detection part, a frame and image selection part, a portrait cropping and splicing part, and a portrait clone fusion part.

[0113] (1) Subject detection: Based on the continuous image sequence or video data stream, the target detection algorithm is used to detect the human body, sphere, etc.

[0114] (2) Frame selection: The interval frames are selected based on the stretchability of the human body and the uniform layout is prioritized.

[0115] (3) Cropping and splicing: According to the selected interval frames, each frame is first cropped and then aligned and spliced.

[0116] (4) Split body fusion part: According to the aligned image and the spliced image, combined with the segmentation result, perform the fusion of the human body split.

[0117] Continuing with the above embodiment, the number of images H of the body image is 4. Therefore, the first split image, the second split image, the third split image, and the fourth split image in the first processed image obtained through object detection are specifically referred to Figure 6 the first row of images in. Figure 6 The second row of images in are the pre-cropped images of the first split image, the second split image, the third split image, and the fourth split image. Figure 6 The third row of images in are the cropped images of the first split image, the second split image, the third split image, and the fourth split image. After obtaining the unified background image after splicing the four split images, as shown in Figure 6 the third row and fifth column. It can be seen from the unified background image that the heads or hands of some human bodies are missing. Then, combined with the aligned images corresponding to each frame in the splicing process and the results of human body segmentation corresponding to each frame (where the aligned images are shown in the images of the third row and columns 1 to 4 in Figure 6 , and the human body segmentation images are shown in the images of the fourth row and columns 1 to 4 in Figure 6 ), further recover the foreground areas of the missing heads or hands. According to the chronological order of the movement time, the human body foreground that is later in time is preferentially displayed on the top layer. Finally, the synthesized split image of the human body as shown in Figure 6 the fourth row and fifth column is obtained.

[0118] Next, the image processing device provided in the embodiments of the present application will be described. The image processing device described below can be correspondingly referred to the image processing method described above.

[0119] Refer to Figure 7 as shown in Figure 7 which is a schematic structural diagram of the image processing device provided in the embodiments of the present application. The image processing device may include:

[0120] An acquisition module 701, configured to acquire image information to be processed;

[0121] An object detection module 702, configured to perform object detection on the image information in the image information to be processed to obtain first processed image information;

[0122] A fusion module 703, configured to perform superimposed fusion on the first processed image information to obtain target image information.

[0123] In the embodiments of the present application, the target detection technology is used to accurately extract the image information of an object in the image information to be processed at multiple moments, and then the superposition fusion technology is used to superpose and fuse the image information of the object at multiple moments into one piece of target image information. Therefore, the final one piece of target image information contains the image information of the object at multiple moments.

[0124] In an optional example, the target detection module 702 is further configured to:

[0125] Perform body detection on each frame of image information in the image information to be processed, and output second processed image information in the image information to be processed; the second processed image information includes several first pieces of image to be processed;

[0126] Based on the detection box information of each first piece of image to be processed in the second processed image information, determine detection box set information;

[0127] Determine the first processed image information according to the detection box set information.

[0128] In an optional example, the target detection module 702 is further configured to:

[0129] Obtain the vertex position coordinate information of each piece of detection box information in the detection box set information;

[0130] According to the vertex position coordinate information of each piece of detection box information, determine the center position coordinate information of each piece of detection box information;

[0131] Determine the first processed image information according to the center position coordinate information of each piece of detection box information.

[0132] In an optional example, the target detection module 702 is further configured to:

[0133] Sort the center position coordinate information of each piece of detection box information to obtain the maximum center position coordinate information and the minimum center position coordinate information;

[0134] Determine position interval information according to the maximum center position coordinate information and the minimum center position coordinate information;

[0135] Determine the first processed image information according to the position interval information, the maximum center position coordinate information, the minimum center position coordinate information, and the center position coordinate information of each piece of detection box information.

[0136] In an optional example, the fusion module 703 is further configured to:

[0137] Obtain a first target image corresponding to the maximum center position coordinate information, and a second target image corresponding to the minimum center position coordinate information;

[0138] Determine the image data intersection area information and the image data union area information between the first target image and the second target image;

[0139] Based on the image data intersection area information and the image data union area information, superimpose and fuse the first processed image information to obtain the target image information.

[0140] In an optional example, the fusion module 703 is further configured to:

[0141] Sort each second image to be processed in the first processed image information in the order of time information from front to back to obtain sorted image set information;

[0142] Based on the image data intersection area information and the image data union area information, crop each third image to be processed in the sorted image set information to obtain cropped image set information;

[0143] Obtain a background image based on the cropped image set information;

[0144] Using the background image as the background, superimpose and fuse each cropped image in the cropped image set information to obtain the target image information.

[0145] In an optional example, the fusion module 703 is further configured to:

[0146] Based on the image data intersection area information and the image data union area information, determine the image data intersection-to-union ratio information;

[0147] If the image data intersection-to-union ratio information is less than or equal to the intersection-to-union ratio threshold information, crop each third image to be processed in the sorted image set information with the first cropping information to obtain cropped image set information; or,

[0148] If the image data intersection-to-union ratio information is greater than the intersection-to-union ratio threshold information, crop each third image to be processed in the sorted image set information with the second cropping information to obtain cropped image set information.

[0149] Furthermore, the image processing device is further configured to:

[0150] Obtain each clone image in the target image information;

[0151] Track the position of the target at each moment based on each clone image.

[0152] The specific embodiments of the image processing apparatus provided in this application are basically the same as those of the image processing method embodiments, and will not be elaborated here.

[0153] Optionally, as Figure 8 shown, Figure 8 is a schematic structural diagram of an electronic device provided by an embodiment of this application. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call a computer program in the memory 830 to execute the steps of the image processing method, for example, including:

[0154] Obtain image information to be processed;

[0155] Perform object detection on the image information in the image information to be processed to obtain first processed image information;

[0156] Superpose and fuse the first processed image information to obtain target image information.

[0157] In an optional embodiment, performing object detection on the image information in the image information to be processed to obtain first processed image information includes:

[0158] Perform body detection on each frame of image information in the image information to be processed, and output second processed image information in the image information to be processed; the second processed image information includes several first images to be processed;

[0159] Based on the detection box information of each first image to be processed in the second processed image information, determine detection box set information;

[0160] Determine the first processed image information according to the detection box set information.

[0161] In an optional embodiment, determining the first processed image information according to the detection box set information includes:

[0162] Obtain the vertex position coordinate information of each detection box information in the detection box set information;

[0163] According to the vertex position coordinate information of each detection box information, determine the center position coordinate information of each detection box information;

[0164] Determine the first processed image information according to the center position coordinate information of each detection box information.

[0165] In an alternative embodiment, determining the first processed image information according to the central position coordinate information of each detection frame information includes:

[0166] Sorting the central position coordinate information of each detection frame information to obtain the maximum central position coordinate information and the minimum central position coordinate information;

[0167] Determining the position interval information according to the maximum central position coordinate information and the minimum central position coordinate information;

[0168] Determining the first processed image information according to the position interval information, the maximum central position coordinate information, the minimum central position coordinate information, and the central position coordinate information of each detection frame information.

[0169] In an alternative embodiment, superimposing and fusing the first processed image information to obtain the target image information includes:

[0170] Obtaining a first target image corresponding to the maximum central position coordinate information and a second target image corresponding to the minimum central position coordinate information;

[0171] Determining the image data intersection area information and the image data union area information between the first target image and the second target image;

[0172] Based on the image data intersection area information and the image data union area information, superimposing and fusing the first processed image information to obtain the target image information.

[0173] In an alternative embodiment, based on the image data intersection area information and the image data union area information, superimposing and fusing the first processed image information to obtain the target image information includes:

[0174] Sorting each second image to be processed in the first processed image information in the order from front to back according to the time information to obtain the sorted image set information;

[0175] Based on the image data intersection area information and the image data union area information, cropping each third image to be processed in the sorted image set information to obtain the cropped image set information;

[0176] Obtaining a background image based on the cropped image set information;

[0177] Using the background image as the background, superimposing and fusing each cropped image in the cropped image set information to obtain the target image information.

[0178] In an alternative embodiment, based on the intersection area information of the image data and the union area information of the image data, each third image to be processed in the sorted image set information is cropped to obtain cropped image set information, including:

[0179] Based on the intersection area information of the image data and the union area information of the image data, determine the intersection-to-union ratio information of the image data;

[0180] If the intersection-to-union ratio information of the image data is less than or equal to the intersection-to-union ratio threshold information, then each third image to be processed in the sorted image set information is cropped with the first cropping information to obtain cropped image set information; or,

[0181] If the intersection-to-union ratio information of the image data is greater than the intersection-to-union ratio threshold information, then each third image to be processed in the sorted image set information is cropped with the second cropping information to obtain cropped image set information.

[0182] In an alternative embodiment, after the first processed image information is superimposed and fused to obtain target image information, it further includes:

[0183] Obtain each avatar image in the target image information;

[0184] Based on each avatar image, track the position of the target at each moment.

[0185] In addition, when the logical computer program in the above-mentioned memory 830 can be implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several computer programs for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0186] On the other hand, the embodiments of the present application also provide a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium includes a computer program. The computer program can be stored on the non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the steps of the image processing method provided in the above-mentioned various embodiments, for example, including:

[0187] Obtain the image information to be processed;

[0188] Perform object detection on the image information in the image information to be processed to obtain the first processed image information;

[0189] Superpose and fuse the first processed image information to obtain the target image information.

[0190] In an alternative embodiment, performing object detection on the image information in the image information to be processed to obtain the first processed image information includes:

[0191] Perform body detection on each frame of image information in the image information to be processed, and output the second processed image information in the image information to be processed; the second processed image information includes several first images to be processed;

[0192] Based on the detection box information of each first image to be processed in the second processed image information, determine the detection box set information;

[0193] Determine the first processed image information according to the detection box set information.

[0194] In an alternative embodiment, determining the first processed image information according to the detection box set information includes:

[0195] Obtain the vertex position coordinate information of each detection box information in the detection box set information;

[0196] According to the vertex position coordinate information of each detection box information, determine the center position coordinate information of each detection box information;

[0197] Determine the first processed image information according to the center position coordinate information of each detection box information.

[0198] In an alternative embodiment, determining the first processed image information according to the center position coordinate information of each detection box information includes:

[0199] Sort the center position coordinate information of each detection box information to obtain the maximum center position coordinate information and the minimum center position coordinate information;

[0200] Determine the position interval information according to the maximum center position coordinate information and the minimum center position coordinate information;

[0201] Determine the first processed image information according to the position interval information, the maximum center position coordinate information, the minimum center position coordinate information, and the center position coordinate information of each detection box information.

[0202] In an alternative embodiment, superimposing and fusing the first processed image information to obtain target image information includes:

[0203] Obtaining a first target image corresponding to the maximum center position coordinate information and a second target image corresponding to the minimum center position coordinate information;

[0204] Determining the image data intersection area information and the image data union area information between the first target image and the second target image;

[0205] Based on the image data intersection area information and the image data union area information, superimposing and fusing the first processed image information to obtain the target image information.

[0206] In an alternative embodiment, based on the image data intersection area information and the image data union area information, superimposing and fusing the first processed image information to obtain the target image information includes:

[0207] Sorting each second image to be processed in the first processed image information in the order of time information from front to back to obtain sorted image set information;

[0208] Based on the image data intersection area information and the image data union area information, cropping each third image to be processed in the sorted image set information to obtain cropped image set information;

[0209] Obtaining a background image based on the cropped image set information;

[0210] Using the background image as the background, superimposing and fusing each cropped image in the cropped image set information to obtain the target image information.

[0211] In an alternative embodiment, based on the image data intersection area information and the image data union area information, cropping each third image to be processed in the sorted image set information to obtain cropped image set information includes:

[0212] Based on the image data intersection area information and the image data union area information, determining image data intersection-union ratio information;

[0213] If the image data intersection-union ratio information is less than or equal to the intersection-union ratio threshold information, cropping each third image to be processed in the sorted image set information with first cropping information to obtain cropped image set information; or,

[0214] If the intersection over union (IoU) information of the image data is greater than the IoU threshold information, then each third image to be processed in the sorted image set information is cropped with the second cropping information to obtain the cropped image set information.

[0215] In an alternative embodiment, after superimposing and fusing the first processed image information to obtain the target image information, it further includes:

[0216] Obtain each clone image in the target image information;

[0217] Track the position of the target at each moment based on each clone image.

[0218] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0219] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, also by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several computer programs for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0220] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. An image processing method, characterized in that, Including: Obtain the image information to be processed; Perform object detection on the image information in the image information to be processed to obtain the first processed image information; Superpose and fuse the first processed image information to obtain the target image information.

2. The image processing method according to claim 1, wherein The performing object detection on the image information in the image information to be processed to obtain the first processed image information includes: Perform body detection on each frame of image information in the image information to be processed, and output the second processed image information in the image information to be processed; the second processed image information includes several first to-be-processed images; Determine the detection box set information based on the detection box information of each first to-be-processed image in the second processed image information; Determine the first processed image information according to the detection box set information.

3. The image processing method according to claim 2, wherein, The determining the first processed image information according to the detection box set information includes: Obtain the vertex position coordinate information of each detection box information in the detection box set information; Determine the center position coordinate information of each detection box information according to the vertex position coordinate information of each detection box information; Determine the first processed image information according to the center position coordinate information of each detection box information.

4. The image processing method according to claim 3, wherein The determining the first processed image information according to the center position coordinate information of each detection box information includes: Sort the center position coordinate information of each detection box information to obtain the maximum center position coordinate information and the minimum center position coordinate information; Determine the position interval information according to the maximum center position coordinate information and the minimum center position coordinate information; Determine the first processed image information according to the position interval information, the maximum center position coordinate information, the minimum center position coordinate information, and the center position coordinate information of each detection box information.

5. The image processing method according to claim 4, wherein The superposing and fusing the first processed image information to obtain the target image information includes: Obtain the first target image corresponding to the maximum center position coordinate information and the second target image corresponding to the minimum center position coordinate information; Determine the image data intersection area information and the image data union area information between the first target image and the second target image; Based on the image data intersection area information and the image data union area information, superpose and fuse the first processed image information to obtain the target image information.

6. The image processing method according to claim 5, wherein The superposing and fusing the first processed image information based on the image data intersection area information and the image data union area information to obtain the target image information includes: Sort each second to-be-processed image in the first processed image information in the order from front to back according to the time information to obtain the sorted image set information; Based on the image data intersection area information and the image data union area information, crop each third to-be-processed image in the sorted image set information to obtain the cropped image set information; Obtain the background image based on the cropped image set information; Using the background image as the background, superpose and fuse each cropped image in the cropped image set information to obtain the target image information.

7. The image processing method according to claim 6, wherein Based on the intersection area information and union area information of the image data, cropping each third image to be processed in the sorted image set information to obtain the cropped image set information, including: Determining the intersection-to-union ratio information of the image data based on the intersection area information and union area information of the image data; If the intersection-to-union ratio information of the image data is less than or equal to the intersection-to-union ratio threshold information, cropping each third image to be processed in the sorted image set information with the first cropping information to obtain the cropped image set information; or, If the intersection-to-union ratio information of the image data is greater than the intersection-to-union ratio threshold information, cropping each third image to be processed in the sorted image set information with the second cropping information to obtain the cropped image set information.

8. The image processing method according to any one of claims 1 to 7, characterized in that After superimposing and fusing the first processed image information to obtain the target image information, further including: Obtaining each clone image in the target image information; Tracking the position of the target at each moment based on each clone image.

9. An image processing apparatus, characterized in that, Including: An acquisition module for acquiring the image information to be processed; A target detection module for performing target detection on the image information in the image information to be processed to obtain the first processed image information; A fusion module for superimposing and fusing the first processed image information to obtain the target image information; Preferably, when the target detection module performs target detection on the image information in the image information to be processed to obtain the first processed image information, it includes: Performing main body detection on each frame of image information in the image information to be processed and outputting the second processed image information in the image information to be processed; the second processed image information includes several first images to be processed; Determining the detection box set information based on the detection box information of each first image to be processed in the second processed image information; Determining the first processed image information according to the detection box set information; Preferably, when the target detection module determines the first processed image information according to the detection box set information, it includes: Obtaining the vertex position coordinate information of each detection box information in the detection box set information; Determining the center position coordinate information of each detection box information according to the vertex position coordinate information of each detection box information; Determining the first processed image information according to the center position coordinate information of each detection box information; Preferably, when the target detection module determines the first processed image information according to the center position coordinate information of each detection box information, it includes: Sorting the center position coordinate information of each detection box information to obtain the maximum center position coordinate information and the minimum center position coordinate information; Determining the position interval information according to the maximum center position coordinate information and the minimum center position coordinate information; Determining the first processed image information according to the position interval information, the maximum center position coordinate information, the minimum center position coordinate information, and the center position coordinate information of each detection box information; Preferably, when the fusion module superimposes and fuses the first processed image information to obtain the target image information, it includes: Obtain a first target image corresponding to the maximum center position coordinate information and a second target image corresponding to the minimum center position coordinate information; Determine the image data intersection area information and the image data union area information between the first target image and the second target image; Based on the image data intersection area information and the image data union area information, superimpose and fuse the first processed image information to obtain the target image information; Preferably, the fusion module superimposes and fuses the first processed image information based on the image data intersection area information and the image data union area information to obtain the target image information, including: Sort each second image to be processed in the first processed image information in the order of time information from front to back to obtain sorted image set information; Based on the image data intersection area information and the image data union area information, crop each third image to be processed in the sorted image set information to obtain cropped image set information; Obtain a background image based on the cropped image set information; Using the background image as the background, superimpose and fuse each cropped image in the cropped image set information to obtain the target image information; Preferably, the fusion module crops each third image to be processed in the sorted image set information based on the image data intersection area information and the image data union area information to obtain cropped image set information, including: Based on the image data intersection area information and the image data union area information, determine the image data intersection-union ratio information; If the image data intersection-union ratio information is less than or equal to the intersection-union ratio threshold information, crop each third image to be processed in the sorted image set information with first cropping information to obtain cropped image set information; or, If the image data intersection-union ratio information is greater than the intersection-union ratio threshold information, crop each third image to be processed in the sorted image set information with second cropping information to obtain cropped image set information; Preferably, after the fusion module superimposes and fuses the first processed image information to obtain the target image information, it further includes: Obtain each clone image in the target image information; Track the position of the target at each moment based on each clone image.

10. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores multiple computer programs; the processor loads the computer programs from the memory to execute the image processing method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple computer programs, and the computer programs are suitable for being loaded by a processor to execute the image processing method according to any one of claims 1 to 8.