Image editing method and device, computer equipment and computer readable storage medium

By determining the image editing method of mobile signs, the sign information area is determined and replaced, and the problem of high-cost acquisition of image data in the prior art is solved, and efficient image expansion and sign information recognition accuracy are achieved.

CN120339458APending Publication Date: 2025-07-18SHENZHEN SMARTMORE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510329069.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, when identifying the sign information of a mobile object, a large amount of image data is required to be collected for training, resulting in high costs and prone to overfitting and misidentification problems.

Method used

By obtaining the original image containing the moving sign, determining the sub-image of the area where the sign information is located, performing mask processing, inputting it to the image editing model, replacing the sign information and fusing it into the original image, and generating a large number of simulated replacement samples.

Benefits of technology

It realizes the generation of a large number of replacement samples based on a small number of image samples, reduces the cost of data acquisition and labeling, and improves the accuracy of sign information recognition and the expansion efficiency of image data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339458A_ABST
    Figure CN120339458A_ABST
Patent Text Reader

Abstract

The invention relates to an image editing method and device, computer equipment and a computer readable storage medium. The method comprises the following steps: acquiring an original image containing a mobile label, wherein the mobile label is attached to a morphologically changing object; determining a sub-image of an area where the label information is located from the original image; performing mask processing on the label information in the sub-image to obtain a label mask image; inputting the sub-image, the label mask image and the acquired replacement information into a preset image editing model for processing, and outputting a target replacement image; the replacement information is used for replacing the label information in the sub-image; and fusing the target replacement image into the original image to obtain an image after label information replacement. By adopting the image expansion method and device, the image containing the mobile label can be replaced, so that a large number of simulation replacement samples are obtained based on a small number of image samples, and the purpose of image expansion is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technologies, and in particular, to an image editing method, apparatus, computer device, and computer-readable storage medium. Background Art

[0002] With the development of computer technologies, object detection technologies have emerged. Through object detection technologies, target objects in images can be identified.

[0003] In traditional technologies, a surveillance camera takes pictures of an indoor area, detects moving objects from the captured images, and identifies the moving objects through the sign information attached to the moving objects. Since the sign information may involve various situations, in order to make the object detection technology have a high accuracy in identifying target objects, it is necessary to collect training images to train the model corresponding to the object detection technology, and the training images need to cover as many sign information as possible, so that the model corresponding to the subsequent object detection technology can detect all sign information. However, taking pictures of all sign information requires a lot of time and manpower. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide an image editing method, apparatus, computer device, computer-readable storage medium, and computer program product, which can achieve editing and replacement of images containing moving signs, so as to obtain a large number of replacement samples based on a small number of image samples, and further achieve the purpose of image augmentation.

[0005] In a first aspect, the present application provides an image editing method, including:

[0006] Obtaining an original image containing a moving sign, where the moving sign is attached to an object with a changing form;

[0007] Determining a sub-image of the area where the sign information is located from the original image;

[0008] Performing masking processing on the sign information in the sub-image to obtain a sign mask image;

[0009] Inputting the sub-image, the sign mask image, and the obtained replacement information into a preset image editing model for processing, and outputting a target replacement image; the replacement information is used to replace the sign information in the sub-image;

[0010] Fusing the target replacement image into the original image to obtain an image after replacing the sign information.

[0011] In a second aspect, the present application provides an image editing apparatus, including:

[0012] An acquisition module, configured to acquire an original image including a mobile sign attached to an object with a changing form;

[0013] A determination module, configured to determine a sub-image of the area where the sign information is located from the original image;

[0014] A masking processing module, configured to perform masking processing on the sign information in the sub-image to obtain a sign mask image;

[0015] A replacement module, configured to input the sub-image, the sign mask image, and the acquired replacement information into a preset image editing model for processing, and output a target replacement image; the replacement information is used to replace the sign information in the sub-image;

[0016] A fusion module, configured to fuse the target replacement image into the original image to obtain an image after replacing the sign information.

[0017] In a third aspect, the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above method are implemented.

[0018] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are implemented.

[0019] In a fifth aspect, the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in the above method are implemented.

[0020] For the above image editing method, device, computer device, computer-readable storage medium, and computer program product, by acquiring an original image including a mobile sign, determining a sub-image of the area where the sign information is located from the original image, and performing masking processing on the sign information in the sub-image to obtain a sign mask image, the position of the sign information in the sub-image can be marked, and the sign information can be separated from other information in the image, which is convenient for subsequent editing of the sign information. And the form of the sign information in the sub-image can also be determined through the sign mask image. In this way, when the sub-image, the sign mask image, and the acquired replacement information are input into a preset image editing model for processing later, the form of the replacement information in the obtained target replacement image matches the form of the sign information in the sub-image. Then, when the target replacement image is fused into the original image, the obtained image after replacing the sign information is more realistic. By replacing the sign information in the original image through the above steps, the purpose of image expansion can be achieved. And the sign information in the replaced image fits well with the image, which is beneficial to the accuracy of the training of the downstream model. Description of the Drawings

[0021] Figure 1 It is a schematic flowchart of an image editing method provided by an embodiment of the present application;

[0022] Figure 2 It is a schematic diagram of the original image provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic flowchart of the step of obtaining the target replacement image provided by an embodiment of the present application;

[0024] Figure 4 It is a schematic flowchart of another image editing method provided by an embodiment of the present application;

[0025] Figure 5 It is a schematic diagram of the image after replacing the sign information provided by an embodiment of the present application;

[0026] Figure 6 It is a structural block diagram of an image editing device provided by an embodiment of the present application;

[0027] Figure 7 It is an internal structure diagram of a computer device provided by an embodiment of the present application;

[0028] Figure 8 It is an internal structure diagram of another computer device provided by an embodiment of the present application;

[0029] Figure 9 It is an internal structure diagram of a computer-readable storage medium provided by an embodiment of the present application. Detailed implementation manners

[0030] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0031] In some places, indoor pedestrians usually need to wear unified clothing such as protective clothing. To successfully identify each pedestrian, different digital signs are usually equipped for different pedestrians in this place. For example, characters such as "00", "16", "az" are provided on the sign. After the surveillance camera in this place captures the pedestrian, the pedestrian can be detected through the target detection algorithm and the information on the digital sign worn by the pedestrian can be recognized, so as to successfully identify the identity of the pedestrian.

[0032] However, in an actual scenario, if there are many types of sign information put into use, such as a large number of digital combinations, in order to improve the accuracy of the target detection algorithm, the collected digital combination image data needs to cover all digital combinations. If the target detection model is trained after annotating less image data, overfitting is likely to occur. For example, if the actual digital combinations in use are two-digit free combinations from 0 to 9, with a total of 100 combinations, due to data collection cost limitations, if the currently collected 10 digital combinations are {00, 11, 22, 33, 44, 55, 66, 77, 88, 99} as the training set, even though the 10 digits from 0 to 9 are covered, when the digital combination 38 appears during actual use, it is still very likely that it cannot be recognized or is misrecognized as 33 or 88. For three-digit combinations or 26-letter combinations, this situation is more likely to occur if the training data is insufficient. If the cost of real data collection and annotation is not considered, collecting a large amount of data of all combinations for training can basically solve the problem. However, in reality, the cost of such collection is extremely expensive. To solve or improve the above problems, the present application proposes the following embodiments.

[0033] As Figure 1 shown, an embodiment of the present application provides an image editing method, which is described by taking the method applied to a server as an example. It can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The method includes the following steps:

[0034] S102. Obtain an original image containing a mobile sign, where the mobile sign is attached to a mobile object with a changing form.

[0035] Among them, the mobile sign is used to identify the mobile object. It can be understood that the mobile signs corresponding to different mobile objects are different.

[0036] Exemplarily, a camera associated with the server takes a picture of a set scene and sends the captured original image to the server, so that the server obtains an original image containing a mobile sign.

[0037] It can be understood that there may be more than one mobile sign in an original image. The identification information in the mobile sign can be a pattern or a character.

[0038] S104. Determine a sub-image of the area where the sign information is located from the original image.

[0039] Among them, the sign information refers to the information corresponding to the mobile sign. The sign information can be, for example, characters or patterns in the mobile sign. Correspondingly, the area where the sign information is located can refer to the area occupied by the mobile sign. Specifically, the area where the sign information is located can also be the area occupied by the mobile object to which the mobile sign is attached.

[0040] Exemplarily, the server determines a sub-image of the area where the sign information is located from the original image, including: determining a sub-image corresponding to the area occupied by the mobile sign from the original image. After the server determines the area occupied by the mobile sign, it cuts the original image along the edge of the area to obtain the sub-image corresponding to the area.

[0041] Specifically, the server determines a sub-image of the area where the sign information is located from the original image, including: determining a sub-image corresponding to the area occupied by the mobile object to which the mobile sign is attached from the original image. In this way, it is beneficial to learn the characteristics of the mobile object subsequently to determine the characteristics of the mobile sign.

[0042] S106. Perform masking processing on the sign information in the sub-image to obtain a sign mask image.

[0043] Among them, the masking processing is a process for separating the sign information in the sub-image from other information in the sub-image. The pixel information corresponding to the sign information in the sign mask image is different from the pixel information corresponding to other information except the sign information. For example, referring to Figure 2 , the sign information in the sub-image is "19", and the value of the pixels occupied by "19" in the corresponding sign mask image obtained is 0, and the value of the pixels not occupied by "19" is 255. By performing masking processing on the sign information in the sub-image, the server can learn the form of the sign information in the original image, thereby improving the fitting degree of the sign information in the subsequently generated image with the image.

[0044] S108. Input the sub-image, the sign mask image, and the obtained replacement information into a preset image editing model for processing, and output a target replacement image.

[0045] Among them, the replacement information is used to replace the sign information in the sub-image. For example, the replacement information can be "21".

[0046] It can be understood that the replacement information can be a replacement character or a replacement pattern. The replacement information can be randomly generated.

[0047] Among them, the image editing model is a model for replacing the sign information to be edited in the input image. The sign mask image is used to determine the position of the replacement information in the sub-image and also to determine the morphological information of the replacement information in the sub-image.

[0048] Exemplarily, after inputting the sub-image, the sign mask image, and the obtained replacement information into a preset image editing model, the preset image editing model can process the replacement information based on the sign mask image to obtain the processed replacement information. Then, the area to be edited can be determined based on the sign mask image, and the processed replacement information can be fused into the area to be edited to obtain the target replacement image. Among them, the form of the processed replacement information matches the form of the sign information in the sub-image.

[0049] When the sub-image is an image of the area occupied by a moving sign, the steps of processing the replacement information based on the sign mask image may include: obtaining the form information corresponding to the replacement information based on the sign mask image; rendering the replacement information based on the form information. Among them, the form information includes the size information and angle information corresponding to the replacement information.

[0050] When the sub-image is an image of the area occupied by a moving object, the steps of processing the replacement information based on the sign mask image may include: obtaining the first form information corresponding to the sign information based on the sign mask image; obtaining the second form information corresponding to the sign information based on the sub-image; determining the target form information corresponding to the replacement information based on the first form information and the second form information; rendering the replacement information based on the target form information. Among them, the first form information includes direct information such as the size information and angle information of the sign information in the sub-image. The second form information includes the change information generated by the form change of the moving object for the sign information. The change information can be, for example, concavo-convex change information or bending change information. For example, when the moving object bends down, the sign information shows a protruding effect. Further, by analyzing the bending degree of the moving object, the protruding degree of the sign information can be determined, so that the protruding degree of the replacement information can be determined based on the protruding degree when rendering the replacement information.

[0051] S110. Fuse the target replacement image into the original image to obtain the image after replacing the sign information.

[0052] Exemplarily, fusing the target replacement image into the original image includes: replacing the sub-image in the original image with the target replacement image. Specifically, change the value of each pixel in the sub-image to the matching pixel value in the target replacement image. For example, the pixel at the 23rd row and 23rd column in the sub-image is the pixel at the 1000th row and 1000th column in the original image. The value of the pixel at the 23rd row and 23rd column in the sub-image is 255. After replacement, the value of the pixel at the 23rd row and 23rd column in the target replacement image is 0. After fusion, the value of the pixel at the 1000th row and 1000th column in the original image is changed to 0.

[0053] Specifically, the step of fusing the target replacement image into the original image includes: removing the sub-image in the original image along the edge of the sub-image; determining the region position corresponding to the sub-image in the original image; and splicing the target replacement image into the original image based on the region position. That is, splicing the target replacement image into the region position in the original image.

[0054] In some embodiments, when there is more than one moving sign in the original image, after obtaining each sub-image of the region where each sign information is located from the original image, it is necessary to obtain the association relationship between each sub-image and the corresponding region position. The input of the preset image editing model can be a combination of multiple images. Among them, the image combination includes a sub-image, a sign mask image corresponding to the sub-image, and replacement information. In this way, when there are multiple moving signs in the original image, the generated multiple target replacement images can be correctly mapped to the corresponding region positions. In this way, more different images can be obtained, thus realizing the expansion of the number of images.

[0055] It can be seen that in the embodiments of the present application, by obtaining the original image containing the moving sign, determining the sub-image of the region where the sign information is located from the original image, and performing mask processing on the sign information in the sub-image to obtain the sign mask image, the position of the sign information in the sub-image can be marked, separating the sign information from other information in the image, which is convenient for subsequent editing of the sign information. And through the sign mask image, the form of the sign information in the sub-image can also be determined. In this way, when the sub-image, the sign mask image, and the randomly generated replacement information are input into the preset image editing model later, the form of the replacement information in the obtained target replacement image matches the form of the sign information in the sub-image. Then, when the target replacement image is fused into the original image, the image after replacing the sign information is more realistic. By replacing the sign information in the original image through the above steps, the purpose of expanding the number of images can be achieved, and the efficiency of image data acquisition can be improved. And the sign information in the replaced image fits well with the image, which is beneficial to the accuracy of the training of the downstream model.

[0056] In some embodiments, the sign information includes the original characters. Inputting the sub-image, the sign mask image, and the obtained replacement information into a preset image editing model for processing to output a target replacement image includes:

[0057] Inputting the sub-image, the sign mask image, and the obtained replacement information into the preset image editing model;

[0058] Rendering the replacement characters based on the sign mask image to obtain target characters; the replacement characters are used to replace the original characters; the form of the target characters matches the form of the original characters in the original image;

[0059] Fuse the target character into the sub-image to obtain the target replacement image.

[0060] To make the replaced image more realistic and the replacement information in the replaced image fit well with the image, the preset image editing model can render the replacement character based on the sign mask image, and use the rendered replacement character as the target character. During the rendering process, the model can learn the form of the original character in the original image based on the sign mask image. The form of the original character in the original image includes the angle information, size information, etc. of the original character. The angle information refers to the camera view corresponding to the original character. The size information refers to the size of the original character. Then, render the replacement character based on the learned form so that the replacement character is aligned with the original character in form.

[0061] It can be understood that the form learned by the model can also include information such as the color, style, texture, and shadow of the original character.

[0062] Specifically, the preset image editing model can be a generative adversarial network model, so that the features of the original character can be fully learned to generate a target character that is highly consistent with the form of the original character.

[0063] After the above steps, the target character to be replaced can be obtained. Then, fuse the target character into the sub-image, and the target replacement image, that is, the updated sign information, can be obtained.

[0064] It can be seen that in this embodiment, an image different from the original image can be obtained based on the original image to achieve the expansion of the image corresponding samples. And rendering the replacement character based on the sign mask image can also make the replaced image highly match the original image, and the replacement information is completely and accurately integrated into the image, thereby improving the authenticity of the replaced image.

[0065] In some embodiments, as Figure 3 shown, when rendering the replacement character, in addition to considering the form of the original character itself, the form of the moving object to which the original character adheres can also be considered. This can improve the rendering accuracy of the replacement character and simulate a more realistic state. The specific steps can be referred to as follows:

[0066] S302. Determine the object image of the area where the object corresponding to the sign information is located from the original image.

[0067] That is to say, the object image includes both the image of the area where the sign information is located and the image of the area where the moving object is located (refer to Figure 2 ).

[0068] Exemplarily, the server can identify an object in the original image, determine the object rectangle annotation box corresponding to the object, and extract the image enclosed by the object rectangle annotation box in the original image to obtain an object image.

[0069] S304. Perform a masking process on the original characters in the object image to obtain a sign masking image.

[0070] Since the sign information is attached to the object, the sign information can be recognized in the object image. When the sign information is characters, the server can perform a masking process on the original characters in the object image to obtain a sign masking image.

[0071] S306. Based on the object image, determine the object form of the object.

[0072] Among them, the object form can be the body state of the object. For example, states such as bending over, stretching the arms backward, deflecting the body to the left, and deflecting the body to the right. When the object is in different object states, the sign information attached to the object will also be different. The sign information includes the form of the moving sign. When the form of the moving sign changes, the form of the characters or patterns correspondingly attached to the moving sign will also change accordingly. For example, when the object arches its back or stretches both hands backward, the moving sign may have a concave-convex form change. Correspondingly, the characters on the moving sign may also have corresponding form changes. Therefore, the object form of the object can be determined based on the object image, so as to further refine the rendering degree of the replacement characters based on the object form.

[0073] Exemplarily, the steps of determining the object form of the object based on the object image may include: determining the key points of the object based on the object image; determining the object form based on the key points. Among them, when the object is a person, the key points can be 17 key points of the human body. Since the 17 key points include the main joints and parts of the human body, the form of the human body can be determined by detecting the key points of the human body, that is, the object form can be determined by detecting the key points of the object. Among them, the joints and parts corresponding to the 17 key points are: nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left arm, right arm, left knee, right knee, left ankle, right ankle. After determining the key points of the object, the object form can be determined through the positional relationship between the key points.

[0074] It can be understood that if the area where the sign is attached to the object is fixed, the object image may also only contain a part of the body of the moving object, such as only the upper body of the human body.

[0075] S308. Render the replacement characters based on the object form and the sign masking image to obtain the target characters.

[0076] Among them, the form of the replacement character in the target character matches the form of the object and also matches the original character in the original image.

[0077] In this step, the server can first render the replacement character based on the sign mask image to obtain the rendered character; then render the rendered character based on the object form to obtain the target character. The object form can be used to refine the concave-convex effect of the target character, etc.

[0078] S310. Integrate the target character into the sub-image to obtain the target replacement image.

[0079] It can be seen that in this embodiment, the image range input to the preset image editing model is expanded, and the image information of the moving object attached to the sign is also input to the model, so that the model can learn the form of the moving object, and then render the replacement character based on the form of the moving object combined with the form of the original character, thereby further improving the adaptability of the replacement character to the scene in the original image, making the image after replacing the character more realistic and facilitating the training of the subsequent target detection model.

[0080] After the above steps, the target replacement image corresponding to the area where the sign information is located can be obtained. To reduce the interference of other information in the image on the model processing result, instead of inputting the entire original image into the preset image editing model, the sub-image corresponding to the area where the sign information is located is input into the preset image editing model. Therefore, after obtaining the target replacement image, the sub-image in the original image should be covered with the target replacement image to achieve the purpose of editing the sign information in the original image. Exemplarily, after determining the sub-image, the sub-image in the original image is extracted. Then, after obtaining the target replacement image, the sub-image in the original image is replaced with the target replacement image, so as to obtain the image after replacing the sign information.

[0081] In some embodiments, determining the sub-image of the area where the sign information is located from the original image includes:

[0082] Identifying multiple bounding boxes from the original image;

[0083] Performing clustering processing on the multiple bounding boxes to obtain at least one combined bounding box;

[0084] Extracting the image of the target area corresponding to the combined bounding box from the original image to obtain the sub-image of the area where the sign information is located.

[0085] Reference Figure 2, the annotation box can be a polygon box that frames a single character or pattern in the image. The combined annotation box is obtained based on at least one annotation box. The combined annotation box can frame all corresponding characters or patterns. When the sign information consists of multiple characters or patterns, it is crucial to accurately determine each piece of sign information. To accurately distinguish each piece of sign information in the original image, each character or pattern in the original image can be detected first. Then, the detected characters or patterns can be clustered using a clustering algorithm, and in this way, each piece of sign information in the original image can be determined based on the clustering result. It can be understood that, to obtain an accurate clustering result, before clustering, the server can first determine the compositional information of the sign information in the image. The compositional information can be, for example, the number of characters or patterns in the sign information.

[0086] For example, referring to Figure 2 , the annotation boxes of the digital characters corresponding to this pedestrian are "2" and "1". Taking the digital characters in all sign information as two-digit numbers as an example, the DBSCAN clustering algorithm is used to cluster the digital characters corresponding to all pedestrians in the current image, and all sign combinations containing two-digit numbers are obtained.

[0087] After determining the combined annotation box in the image through the above steps, the image of the target area corresponding to the combined annotation box can be segmented from the original image to obtain a sub-image. Among them, the size of the target area can be greater than or equal to the size of the combined annotation box.

[0088] It can be seen that in this embodiment, by first separately identifying the single annotation boxes in the original image and then clustering these annotation boxes to determine each piece of sign information, and then accurately determining the sub-images corresponding to each piece of sign information from the original image. And since only single characters or patterns need to be identified to achieve the detection of any combined sign information, the categories that the server needs to learn can be reduced. For example, for a two-digit number combination, in order to detect each number combination, if directly detecting each number combination, the features of 100 numbers from 00 to 99 need to be learned. If single characters are detected first and then clustered, it can be achieved by learning the features of 10 numbers from 0 to 9, and in this way, the load pressure on the server can be reduced.

[0089] In some embodiments, to improve the processing accuracy of the model for the image and since the area where the sign information is located is small, before inputting the sub-image corresponding to the area where the sign information is located into the model, the sub-image can be enlarged. This is convenient for the model to perform processing such as feature extraction on the image. After enlargement, the target replacement image output by the model also becomes larger than the original sub-image. Therefore, during the process of fusing the target replacement image into the original image, the size of the target replacement image needs to be reduced to the size of the original sub-image, and then the reduced-size target replacement image is used to cover the original sub-image. The specific description can be seen in the following content.

[0090] In some embodiments, extracting an image of the target region corresponding to the combined annotation box from the original image to obtain a sub-image of the region where the sign information is located includes:

[0091] Determining the center point of the combined annotation box in the original image;

[0092] Based on the size of the combined annotation box, determining the horizontal extension length and the vertical extension length;

[0093] Taking the center point as a reference, determining the image of the target region corresponding to the combined annotation box based on the horizontal extension length and the vertical extension length to obtain a sub-image of the region where the sign information is located.

[0094] That is to say, the server can perform cropping on the original image by expanding a certain length from the center point of the combined annotation box.

[0095] Among them, based on the size of the combined annotation box, determining the horizontal extension length and the vertical extension length includes:

[0096] Determining the first preset multiple of the length of the combined annotation box as the horizontal extension length;

[0097] Determining the second preset multiple of the length of the combined annotation box as the vertical extension length.

[0098] In this way, the finally determined target region can be determined according to the size of the sign in the image, so that the finally determined target region can not only pay attention to the background information around the sign, but also reduce the interference that may be caused by excessive background information.

[0099] Among them, the length of the target region can be equal to the horizontal extension length, and the width of the target region can be equal to the vertical extension length. The center point of the target region overlaps with the center point of the combined annotation box.

[0100] It can be seen that in this embodiment, by expanding the region based on the center of the combined annotation box, a target region containing more image information can be cropped. In this way, when the subsequent model processes the input sub-image, the model can also pay attention to some information (such as background information) around the sign information, thereby improving the matching degree between the target replacement image generated by the model and the original image.

[0101] It can be understood that if the regions occupied by each sign information in the original image are not very different, the horizontal extension length and the vertical extension length can also be determined by experience.

[0102] In some embodiments, after the server determines the sub-image of the region where the sign information is located from the original image, the method further includes:

[0103] Enlarge the sub-image to obtain the enlarged sub-image.

[0104] Correspondingly, perform masking processing on the sign information in the sub-image to obtain a sign mask image, including:

[0105] Perform masking processing on the original characters in the enlarged sub-image to obtain a sign mask image;

[0106] Correspondingly, input the sub-image, the sign mask image, and the obtained replacement information into a preset image editing model for processing, and output a target replacement image, including:

[0107] Input the enlarged sub-image, the sign mask image, and the obtained replacement information into a preset image editing model for processing, and output a target replacement image;

[0108] Correspondingly, fuse the target replacement image into the original image to obtain an image with the sign information replaced, including:

[0109] Perform shrinking processing on the target replacement image to obtain a shrunk target replacement image; wherein, the shrunk target replacement image has the same size as the sub-image;

[0110] Fuse the shrunk target replacement image into the original image to obtain an image with the original characters replaced.

[0111] In this embodiment, by enlarging the sub-image, the model can capture more features in the sub-image, thereby generating a target replacement image with a high degree of matching with the sub-image.

[0112] In some embodiments, perform super-resolution enlargement on the sub-image to obtain the enlarged sub-image, including:

[0113] Enlarge the sub-image by a preset multiple to obtain a temporarily enlarged image;

[0114] Perform normalization processing on the temporarily enlarged image to obtain a preprocessed image;

[0115] Input the preprocessed image into a preset super-resolution enlargement model for processing, and output the enlarged sub-image.

[0116] By performing super-resolution enlargement on the sub-image, compared with other enlargement processes, using the preset super-resolution enlargement model to enlarge the sub-image can make the image obtained after subsequent shrinking of the target replacement image have high clarity and retain the features and details of the shrunk target replacement image.

[0117] Since the original image captured or the image generated by editing the original image is subsequently used in the training process of the downstream model, in order to improve the accuracy of the downstream model and reduce the impact of data distribution changes caused by the images generated by the preset image editing model on the training of the downstream model, some target images can be added to the training data applied to the downstream model. Among them, the target image is an image obtained by enlarging the original image but not edited using the model. Specifically, after determining the sub-image, the sub-image is segmented from the original image and directly fused back into the original image to obtain the target image.

[0118] In some embodiments, the method further includes:

[0119] Determine multiple replacement elements from the replacement information;

[0120] Input the sub-image, the sign mask image, and the multiple replacement elements into a preset image editing model for processing, and output multiple target replacement images;

[0121] Fuse each target replacement image into the original image respectively to obtain multiple images after replacing the sign information;

[0122] Add the multiple images after replacing the sign information to the training image set to expand the training image set.

[0123] Among them, for the server to determine the sign information in the target replacement image, it can first determine the target number of elements included in the original characters corresponding to the single sign information in the original image, then perform permutations and combinations on the input multiple replacement elements to obtain multiple replacement characters composed of the target number of replacement elements, and then render each different replacement character respectively to obtain multiple target replacement images, so as to obtain multiple images after replacing the sign information, and add these images after replacing the sign information to the training image set to expand the training image set.

[0124] In an exemplary embodiment, referring to Figure 4 , the server can first obtain the original image containing the movable sign and identify multiple annotation boxes from the original image. For example Figure 2 in, the annotation boxes corresponding to the digital patterns of the pedestrians are "2" and "1". The annotation boxes in the original image can also be pre-annotated manually.

[0125] Then cluster the multiple annotation boxes to obtain at least one combined annotation box. Taking all the sign information in the image as two-digit numbers as an example, use the DBSCAN clustering algorithm to cluster all the sign combinations of the pedestrians in the current image to obtain all the two-digit sign combinations.

[0126] After that, determine the center point of the combined annotation box, and determine the horizontal expansion length and the vertical expansion length based on the size of the combined annotation box. Taking the center point as a reference, determine the target area corresponding to the combined annotation box based on the horizontal expansion length and the vertical expansion length. Extract the image of the target area corresponding to the combined annotation box from the original image to obtain a sub-image. That is to say, expand a certain length from the center point of each sign combined annotation box and crop it on the original image to obtain the sub-image in the original image.

[0127] To enable the model to capture more features and details in the image, the sub-image can be enlarged to obtain an enlarged image. Perform masking processing on the original characters in the enlarged image to obtain a sign mask image. For example, use a super-resolution model (such as the EDSR model) to enlarge the obtained sub-image to an image with a side length of 512, and perform masking processing on the original numbers on this image.

[0128] After obtaining the replacement information and determining the replacement characters based on the replacement information, render the replacement characters based on the sign mask image to obtain target characters; fuse the target characters into the sub-image to obtain a target replacement image. For example, the server can also randomly generate a two-digit random number from 00 to 99, and input it together with the masked image and the sub-image into a preset image editing model for processing, and output a target replacement image with the edited numbers.

[0129] To accurately fuse the target replacement image into the original image, the size of the target replacement image needs to be reduced to the size of the sub-image. Then fuse the reduced target replacement image into the original image to obtain an image that replaces the original characters (see Figure 5 ).

[0130] Through the above steps, it is possible to realize the true data of the sign information in a small number of images, that is, to obtain the effect of a large number of simulated images containing sign information, greatly reducing the manual collection and data annotation costs, and improving the efficiency of data collection. Moreover, by rendering the replacement characters based on the sign mask image, the angle, font, format, etc. corresponding to the original sign can be retained as much as possible. Compared with the method of directly copying and pasting to expand data, the risk of affecting the original data distribution is greatly reduced.

[0131] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless specifically stated herein, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0132] Based on the same inventive concept, an embodiment of the present application also provides an image editing device. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the image editing device provided below can refer to the limitations on the image editing method in the above text, and will not be repeated here.

[0133] As Figure 6 shown, an embodiment of the present application provides an image editing device 600, including:

[0134] An acquisition module 601, configured to acquire an original image including a mobile sign, where the mobile sign is attached to an object with a changing form;

[0135] A determination module 602, configured to determine a sub-image of the area where the sign information is located from the original image;

[0136] A masking processing module 603, configured to perform masking processing on the sign information in the sub-image to obtain a sign mask image;

[0137] A replacement module 604, configured to input the sub-image, the sign mask image, and the acquired replacement information into a preset image editing model for processing, and output a target replacement image; the replacement information is used to replace the sign information in the sub-image;

[0138] A fusion module 605, configured to fuse the target replacement image into the original image to obtain an image after replacing the sign information.

[0139] In some embodiments, the sign information includes original characters. When inputting the sub-image, the sign mask image, and the acquired replacement information into a preset image editing model for processing and outputting a target replacement image, the replacement module 604 is specifically configured to:

[0140] Input the sub-image, the sign mask image, and the acquired replacement information into a preset image editing model;

[0141] Determine replacement characters based on replacement information;

[0142] Render the replacement characters based on the sign mask image to obtain target characters; the replacement characters are used to replace the original characters; the form of the target characters matches the form of the original characters in the original image;

[0143] Fuse the target characters into the sub-image to obtain a target replacement image.

[0144] In some embodiments, in terms of determining the sub-image of the area where the sign information is located from the original image, the determining module 602 is specifically configured to:

[0145] Determine the object image of the area where the object corresponding to the sign information is located from the original image;

[0146] In terms of performing a masking process on the sign information in the sub-image to obtain a sign mask image, the masking processing module 603 is specifically configured to:

[0147] Perform a masking process on the original characters in the object image to obtain a sign mask image;

[0148] In terms of rendering the replacement characters based on the sign mask image to obtain target characters, the replacement module 604 is specifically configured to:

[0149] Based on the object image, determine the object form of the object;

[0150] Render the replacement characters based on the object form and the sign mask image to obtain target characters; the form of the replacement characters in the target characters matches the form of the object.

[0151] In some embodiments, in terms of determining the sub-image of the area where the sign information is located from the original image, the determining module 602 is specifically configured to:

[0152] Identify multiple annotation boxes from the original image;

[0153] Perform clustering processing on the multiple annotation boxes to obtain at least one combined annotation box;

[0154] Extract the image of the target area corresponding to the combined annotation box from the original image to obtain the sub-image of the area where the sign information is located.

[0155] In some embodiments, in terms of extracting the image of the target area corresponding to the combined annotation box from the original image to obtain the sub-image of the area where the sign information is located, the determining module 602 is specifically configured to:

[0156] Determine the center point of the combined annotation box from the original image;

[0157] Based on the size of the combined annotation box, determine the horizontal extension length and the vertical extension length;

[0158] Based on the center point, determine the image of the target area corresponding to the combined annotation box based on the horizontal expansion length and the vertical expansion length, so as to obtain the sub-image of the area where the sign information is located.

[0159] In some embodiments, the image editing device 600 further includes a scaling module 606, and the scaling module 606 is configured to: magnify the sub-image to obtain a magnified sub-image.

[0160] In terms of performing a masking process on the sign information in the sub-image to obtain a sign mask image, the masking process module 603 is specifically configured to:

[0161] Perform a masking process on the original characters in the magnified sub-image to obtain a sign mask image.

[0162] In terms of inputting the sub-image, the sign mask image, and the obtained replacement information into a preset image editing model for processing and outputting a target replacement image, the replacement module 604 is specifically configured to:

[0163] Input the magnified sub-image, the sign mask image, and the obtained replacement information into a preset image editing model for processing and output a target replacement image.

[0164] In terms of fusing the target replacement image into the original image to obtain an image after replacing the sign information, the fusion module 605 is specifically configured to:

[0165] Perform a reduction process on the size of the target replacement image to obtain a reduced target replacement image; the size of the reduced target replacement image is the same as that of the sub-image;

[0166] Fuse the reduced target replacement image into the original image to obtain an image after replacing the original characters.

[0167] In some embodiments, the replacement module 604 is further configured to: determine a plurality of replacement elements from the replacement information; input the sub-image, the sign mask image, and the plurality of replacement elements into a preset image editing model for processing and output a plurality of target replacement images;

[0168] The fusion module 605 is further configured to: respectively fuse each target replacement image into the original image to obtain a plurality of images after replacing the sign information; add the plurality of images after replacing the sign information to the training image set to expand the training image set.

[0169] Each module in the above image editing device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0170] In some embodiments, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to image editing. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements the steps in the above image editing method.

[0171] In some embodiments, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 8As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements the steps in the above-mentioned image editing method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen; the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0172] Those skilled in the art can understand that Figure 7 or Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0173] In some embodiments, a computer device is provided. The computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the steps in the above-mentioned method embodiments.

[0174] In some embodiments, as Figure 9 shown, an internal structure diagram of a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. When the computer program is executed by the processor, it implements the steps in the above-mentioned method embodiments.

[0175] In some embodiments, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by the processor, it implements the steps in the above-mentioned method embodiments.

[0176] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0177] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0178] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0179] The above embodiments merely illustrate several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. An image editing method, characterized in that, Including: Obtain an original image including a mobile sign attached to a morphing object; Determine a sub-image of the area where the sign information is located from the original image; Perform masking processing on the sign information in the sub-image to obtain a sign mask image; Input the sub-image, the sign mask image, and the obtained replacement information into a preset image editing model for processing, and output a target replacement image; the replacement information is used to replace the sign information in the sub-image; Fuse the target replacement image into the original image to obtain an image after replacing the sign information.

2. The method according to claim 1, wherein The sign information includes original characters. The step of inputting the sub-image, the sign mask image, and the obtained replacement information into a preset image editing model for processing and outputting a target replacement image includes: Input the sub-image, the sign mask image, and the obtained replacement information into a preset image editing model; Determine replacement characters based on the replacement information; Render the replacement characters based on the sign mask image to obtain target characters; the replacement characters are used to replace the original characters; the form of the target characters matches the form of the original characters in the original image; Fuse the target characters into the sub-image to obtain a target replacement image.

3. The method according to claim 2, wherein The step of determining a sub-image of the area where the sign information is located from the original image includes: Determine an object image of the area where the object corresponding to the sign information is located from the original image; The step of performing masking processing on the sign information in the sub-image to obtain a sign mask image includes: Perform masking processing on the original characters in the object image to obtain a sign mask image; The step of rendering the replacement characters based on the sign mask image to obtain target characters includes: Determine the object form of the object based on the object image; Render the replacement characters based on the object form and the sign mask image to obtain target characters; the form of the replacement characters in the target characters matches the form of the object.

4. The method according to claim 1, characterized in that, The step of determining a sub-image of the area where the sign information is located from the original image includes: Identify multiple annotation frames from the original image; Perform clustering processing on the multiple annotation frames to obtain at least one combined annotation frame; Extract an image of the target area corresponding to the combined annotation frame from the original image to obtain a sub-image of the area where the sign information is located.

5. The method according to claim 4, characterized in that, The step of extracting an image of the target area corresponding to the combined annotation frame from the original image to obtain a sub-image of the area where the sign information is located includes: Determine the center point of the combined annotation frame from the original image; Determine the horizontal expansion length and the vertical expansion length based on the size of the combined annotation frame; Based on the center point, determine an image of the target area corresponding to the combined annotation frame based on the horizontal expansion length and the vertical expansion length to obtain a sub-image of the area where the sign information is located.

6. The method according to claim 4, characterized in that The sign information includes original characters. After determining a sub-image of the area where the sign information is located from the original image, the method further includes: Enlarge the sub-image to obtain an enlarged sub-image; The masking process for the sign information in the sub-image to obtain a sign mask image includes: Perform a masking process on the original characters in the enlarged sub-image to obtain a sign mask image; The step of inputting the sub-image, the sign mask image, and the obtained replacement information into a preset image editing model for processing to output a target replacement image includes: Input the enlarged sub-image, the sign mask image, and the obtained replacement information into the preset image editing model for processing to output a target replacement image; The step of fusing the target replacement image into the original image to obtain an image after replacing the sign information includes: Perform a shrinking process on the target replacement image to obtain a shrunk target replacement image; the shrunk target replacement image has the same size as the sub-image; Fuse the shrunk target replacement image into the original image to obtain an image replacing the original characters.

7. The method according to claim 1, characterized in that, The method further includes: Determine a plurality of replacement elements from the replacement information; Input the sub-image, the sign mask image, and the plurality of replacement elements into a preset image editing model for processing to output a plurality of target replacement images; Fuse each of the target replacement images into the original image respectively to obtain a plurality of images after replacing the sign information; Add the plurality of images after replacing the sign information to the training image set to expand the training image set.

8. An image editing device, characterized in that, It includes: An acquisition module for acquiring an original image containing a moving sign, where the moving sign is attached to an object with a changing form; A determination module for determining a sub-image of the area where the sign information is located from the original image; A masking process module for performing a masking process on the sign information in the sub-image to obtain a sign mask image; A replacement module for inputting the sub-image, the sign mask image, and the obtained replacement information into a preset image editing model for processing to output a target replacement image; the replacement information is used to replace the sign information in the sub-image; A fusion module for fusing the target replacement image into the original image to obtain an image after replacing the sign information.

9. A computer device, the computer device includes a memory and a processor, the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.