Picture editing method and system, and related device
By identifying the user's editing area in the picture and generating recommended editing instructions, the problem of high threshold for existing image editing tools is solved, and a more efficient and simple image editing process is achieved.
Patent Information
- Application Number
- PCT/CN2024/140387
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-18
- Publication Date
- 2025-06-26
AI Technical Summary
The existing image editing tools have a high threshold for ordinary users, requiring professional knowledge and tools, and the editing process is complex, making it difficult to achieve high-quality editing results.
It provides an image editing method, which can detect user operations, identify editing areas, and generate recommended editing instructions based on image features of the area, thereby reducing the complexity of user operations and improving editing efficiency.
It significantly reduces the complexity of image editing, supports user-defined editing instructions, reduces the number of trial and errors of users, improves the efficiency of image production, and ensures the rationality of edited image content.
Smart Images

Figure CN2024140387_26062025_PF_FP_ABST
Abstract
Description
Image editing method, related equipment and system
[0001] This application claims priority to the Chinese patent application with application number 202311795993.1 filed with the State Intellectual Property Office of China on December 22, 2023, and priority to the Chinese patent application with the invention name “Image Editing Method, Related Equipment and System”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of electronic technology, and in particular to image editing methods, related equipment and systems. Background Art
[0003] Currently, many professional image editing tools or applications can assist users in editing and processing local areas of an image. When editing an image, users typically need to manually select the desired area within the original image using functions like lassoing and selecting a frame. They then proceed to perform operations such as cropping, redrawing, coloring, or adding new elements to achieve the final edited image. This entire process requires a high level of expertise and the use of specialized tools, as well as significant labor costs to complete high-quality editing tasks, making it unsuitable for the average user. Summary of the Invention
[0004] The embodiments of the present application provide a picture editing method, related equipment and system, which can lower the threshold for picture editing, support user-defined editing instructions, and understand the user's operating intentions, provide users with editing ideas, avoid users from experiencing a lot of trial and error, and significantly improve the efficiency of picture output.
[0005] In a first aspect, an embodiment of the present application provides a picture editing method, which can be applied to a picture editing system. The method may include: displaying a first picture, detecting a user operation on the first picture for selecting an editing area, determining an editing area from the first picture according to the operation position of the user operation on the first picture, and distinguishing and displaying the editing area in the first picture; then, generating a recommended editing instruction based on the image features of the editing area, and displaying the recommended editing instruction; detecting that the user inputs a first editing instruction for the editing area, the first editing instruction includes: a recommended editing instruction; then, performing image editing processing on the first picture according to the first editing instruction; and finally, displaying the first picture after image editing processing.
[0006] In the first aspect, the first image can be an image opened by a user through an application for browsing, managing, or processing images, such as a photo album (also known as a gallery), a photo editing application, a drawing design program, etc. The first image can be stored on a terminal device such as a mobile phone, or stored on a network.
[0007] The method provided in the first aspect is based on the understanding of the semantic content of the original image. It recommends editing instructions to users according to the image features of the editing area selected by the user, and provides users with editing ideas. This not only reduces the complexity of use, but also ensures the rationality of the content of the edited image, avoids users from experiencing a lot of trial and error, and significantly improves the efficiency of image output.
[0008] In combination with the first aspect, in some embodiments, generating recommended editing instructions based on the image features of the editing area may specifically include: using a fused feature vector of multiple image features of the editing area as one of the inputs of the first artificial intelligence algorithm; using one or more preset editing types as the second input of the first artificial intelligence algorithm; obtaining recommended editing instructions through the operation of the first artificial intelligence algorithm, the recommended editing instructions including editing parameters corresponding to the preset editing types; wherein the multiple image features may include multiple items of mask features, depth features, contour features, and color features.
[0009] In conjunction with the first aspect, in some embodiments, the preset editing type may include one or more of the following: delete, drag, replace, add, or color adjustment. The editing parameter corresponding to delete may include the deleted content, the editing parameter corresponding to replace may include the replaceable content, the editing parameter corresponding to drag may include the drag target location, the editing parameter corresponding to add may include the newly added content, the editing parameter corresponding to color adjustment may include the color adjustment value, etc.
[0010] In conjunction with the first aspect, in some embodiments, detecting that the user inputs a first editing instruction for the editing area may specifically include: detecting that the user selects to input a recommended editing instruction, and the recommended editing instruction selected by the user is determined as the first editing instruction.
[0011] In conjunction with the first aspect, in some embodiments, after identifying the editing area, the image editing system is not limited to displaying recommended editing instructions. It can also prompt the user to enter editing instructions in the following manner: display a first input box, which can be used to receive voice or text editing instructions; the first editing instructions also include voice or text editing instructions entered through the input box. When prompting the user to enter editing instructions in this manner, detecting that the user enters the first editing instruction for the editing area can specifically include: detecting the voice or text instruction entered by the user in the first input box, and the voice or text instruction being determined as the first editing instruction. That is, the first editing instruction can include the user entering a text instruction in the input box or pressing the voice key to enter a voice instruction.
[0012] In conjunction with the first aspect, in some embodiments, after identifying an editing area, the image editing system may not only display recommended editing instructions but may also prompt the user to enter an editing instruction by displaying one or more preset editing instructions, such as commonly used editing instructions or editing instructions previously saved by the user. When prompting the user to enter an editing instruction in this manner, detecting that the user enters a first editing instruction for the editing area may specifically include: detecting that the user selects to enter a preset editing instruction, and determining that the voice or text instruction is the first editing instruction.
[0013] In combination with the first aspect, in some embodiments, the user operation for selecting the editing area may include: a user operation of selecting a first object in a first picture, the operation position of the user operation on the first picture falls on the first object in the first picture, and the image area where the first object is located is the editing area; the image area where the first object is located is determined by performing image segmentation processing on the first picture.
[0014] In combination with the first aspect, in some embodiments, the user operation for selecting the editing area may include: a user operation for drawing the editing area in the first picture.
[0015] In combination with the first aspect, in some embodiments, distinctively displaying the editing area in the first picture may include one or more of the following methods: highlighting the outline of the editing area, highlighting the entire editing area, or displaying a dotted frame along the outline of the editing area.
[0016] In conjunction with the first aspect, in some embodiments, the method of the first aspect may further include: obtaining preprocessing information of the first image, and determining an editing region from the region indicated by the preprocessing information. In this way, the editing region can be determined based on the preprocessing information of the image, rather than image segmentation techniques, eliminating the need for repeated online calculations.
[0017] In conjunction with the first aspect, in some embodiments, the preprocessing information may include indicative information for multiple regions, where the preprocessing information includes coordinates of contour points of each of the multiple regions, a binary image, and a grayscale image, where the values of the multiple regions in the binary image are first values, and where the grayscale values of the multiple regions in the grayscale image are first grayscale values or first grayscale ranges. Thus, determining the editing region from the regions indicated by the preprocessing information may specifically include determining the region within the multiple regions where the user operates as the editing region, or determining the region within the multiple regions closest to the operating position as the editing region.
[0018] In conjunction with the first aspect, in some embodiments, the pre-processing information may include only information indicating a region, such as coordinates of contour points of the region, a binary image or grayscale image indicating the region, etc. Thus, determining the editing region from the region indicated by the pre-processing information may specifically include directly determining the region as the editing region.
[0019] In conjunction with the first aspect, in some embodiments, the pre-processing information of the first image may not directly indicate one or more regions, but may include other data, such as multi-layer information, contour information, or depth information. In this way, the other data may be used to first determine one or more regions, and then the editing area may be determined from these one or more regions.
[0020] Methods for determining one or more regions indirectly indicated by the pre-processing information may include, but are not limited to:
[0021] If the preprocessing information includes layer information of multiple layers, and the layer information of each layer includes the coordinates of opaque pixels in the layer, then: before determining the editing area from the area indicated by the preprocessing information, the opaque pixels connected in each layer can be determined as an area indicated by the preprocessing information.
[0022] If the pre-processing information includes contour information, then: before determining the editing area from the area indicated by the pre-processing information, the area surrounded by the contour indicated by the contour information may be determined as the area indicated by the pre-processing information.
[0023] If the preprocessing information includes depth information, then before determining the editing area from the area indicated by the preprocessing information, pixels with the same or similar depth values may be determined as an area indicated by the preprocessing information based on the depth information.
[0024] In conjunction with the first aspect, in some embodiments, the method of the first aspect may further include: before performing image editing processing on the first image according to the first editing instruction, if it is determined that the first editing instruction is not reasonably applied to the editing area, re-recommending the editing instruction or re-recommending the editing area. This solves the problem of the user arbitrarily inputting editing instructions, resulting in editing results that do not conform to common sense and logic, thereby reducing the number of trial and error times when the user edits the image.
[0025] In combination with the first aspect, in some embodiments, it can be determined whether the application of the first editing instruction to the editing area is reasonable by comparing the feature vector corresponding to the first editing instruction with the fused feature vector of each image feature of the editing area.
[0026] In combination with the first aspect, in some embodiments, re-recommending the editing area may specifically include: traversing the area outside the editing area in the first image, comparing the fused feature vector of the traversed area with the feature vector corresponding to the first editing instruction, finding an area that can be reasonably matched with the first editing instruction, and re-recommending the found area as the editing area.
[0027] In combination with the first aspect, in some embodiments, re-recommending editing instructions may specifically include: traversing the editing types and / or editing parameters in the recommendation pool, comparing the feature vectors corresponding to the traversed editing types and / or editing parameters with the fusion feature vector of the editing area, finding the editing types and / or editing parameters that can be reasonably matched with the editing area, and modifying the first editing instruction according to the found editing type and / or editing parameters, and re-recommending the modified first editing instruction.
[0028] In conjunction with the first aspect, in some embodiments, the editing instructions entered by the user into the editing area may involve adding new objects. These editing instructions may include: add, replace, drag, and other types of editing instructions. Replacement is equivalent to deleting an object from the original image and then adding another object; dragging is equivalent to deleting an object from a certain location in the original image and then adding it to another location. In other words, an editing instruction involving adding new objects may mean that the editing process corresponding to the editing instruction includes adding a new object to the editing area.
[0029] After undergoing such image editing processing, in the edited area of the first picture, the depth features of the first pixel area are replaced with the depth features of the new object, and the depth features of the second pixel area remain the depth features of the original object, wherein the perspective relationship of the first pixel area is that the new object is in front of the original object, and the perspective relationship of the second pixel area is that the original object is in front of the new object.
[0030] In combination with the first aspect, in some embodiments, performing image editing processing on the first picture according to the first editing instruction may include: correcting image features of the first picture, and regenerating the first picture using the corrected image features of the first picture.
[0031] In conjunction with the first aspect, in some embodiments, the image features may include depth features. Correcting the image features of the first image may specifically include: determining the perspective relationship between the new object and the original object based on the image features of the first image and the editing parameters in the first editing instruction; determining a baseline depth of the new object based on the perspective relationship between the new object and the original object and the depth value of the original object, and then correcting the depth features of the new object using the baseline depth; replacing the original depth features of the edited area with the corrected depth features of the new object; wherein, after the depth features of the new object are corrected, the average depth of the new object is close to or equal to the baseline depth, and the depth differences between various areas on the new object remain unchanged.
[0032] In combination with the first aspect, in some embodiments, the image feature may further include a first image feature, which is an image feature other than a depth feature and may include one or more of the following: a mask feature, a contour feature, and a color feature.
[0033] Correcting the image features of the first picture may also include: within the editing area, if the depth feature of a pixel area is replaced with the corrected depth feature of the new object, then using the first image feature of the pixel area on the new object to replace the original first image feature of the pixel area in the editing area.
[0034] In the second aspect, an embodiment of the present application provides a terminal device, which may include: a human-computer interaction module, a processor and a memory, wherein the human-computer interaction module is coupled to the processor, and the memory is coupled to the processor; the human-computer interaction module may include input and output components such as a touch screen; wherein the memory can be used to store computer program code, and the computer program code may include computer instructions. When the processor executes the computer instructions, the terminal device executes the method described in any one or more embodiments of the first aspect mentioned above.
[0035] When the terminal device provided in the second aspect possesses powerful computing capabilities, complex calculations can be performed directly on the computing module (including the processor) of the terminal device. When the terminal device possesses powerful computing capabilities, the human-computer interaction module and computing module mentioned in subsequent embodiments can be deployed on the terminal device. The steps performed by the two are the steps performed by the terminal device, and the communication or data exchange between the two constitutes intra-device communication.
[0036] In a third aspect, an embodiment of the present application provides a computer-readable storage medium comprising instructions, characterized in that when the instructions are executed on a terminal device, the terminal device executes the method described in any one or more embodiments of the first aspect.
[0037] In a fourth aspect, an embodiment of the present application provides a picture editing method, which can be applied to a human-computer interaction module. The human-computer interaction module is included in a picture editing system, and the picture editing system also includes: a computing module.
[0038] The method may include: a human-computer interaction module displays a first picture, detects a user operation on the first picture for selecting an editing area, and distinctively displays the editing area in the first picture; then, the human-computer interaction module receives a recommended editing instruction sent by a calculation module, and displays the recommended editing instruction, which is generated by the calculation module based on the image features of the editing area; the human-computer interaction module detects that the user inputs a first editing instruction for the editing area, and the first editing instruction includes: a recommended editing instruction; finally, the human-computer interaction module receives the first picture after image editing processing sent by the calculation module, and displays the first picture after image editing processing, and the image editing processing is performed by the calculation module based on the first editing instruction.
[0039] In the fourth aspect, the first image can be an image opened by a user through an application for browsing, managing, or processing images, such as a photo album (also known as a gallery), a photo editing application, or a drawing design program. The first image can be stored on a terminal device such as a mobile phone or on a network. For some technical details of the fourth aspect, reference can be made to any one or more embodiments of the first aspect.
[0040] In conjunction with the fourth aspect, in some embodiments, detecting that the user inputs a first editing instruction for the editing area may specifically include: detecting that the user selects to input a recommended editing instruction, and the recommended editing instruction selected by the user is determined as the first editing instruction.
[0041] In conjunction with the fourth aspect, in some embodiments, after identifying the editing area, the human-computer interaction module is not limited to displaying recommended editing instructions. It can also prompt the user to enter editing instructions in the following manner: display a first input box, which can be used to receive voice or text editing instructions; the first editing instruction also includes voice or text editing instructions entered through the input box. When prompting the user to enter an editing instruction in this manner, detecting that the user enters the first editing instruction for the editing area can specifically include: detecting the voice or text instruction entered by the user in the first input box, and determining the voice or text instruction as the first editing instruction. That is, the first editing instruction can include the user entering a text instruction in the input box or pressing the voice key to enter a voice instruction.
[0042] In conjunction with the fourth aspect, in some embodiments, after identifying an editing area, the human-computer interaction module may not only display recommended editing instructions but may also prompt the user to enter an editing instruction by displaying one or more preset editing instructions, such as commonly used editing instructions or editing instructions previously saved by the user. When prompting the user to enter an editing instruction in this manner, detecting that the user enters a first editing instruction for the editing area may specifically include: detecting that the user selects to enter a preset editing instruction, and determining that the voice or text instruction is the first editing instruction.
[0043] In combination with the fourth aspect, in some embodiments, the user operation for selecting the editing area may include: a user operation of selecting a first object in the first picture, the operation position of the user operation on the first picture falls on the first object in the first picture, and the image area where the first object is located is the editing area; the image area where the first object is located is determined by performing image segmentation processing on the first picture.
[0044] In conjunction with the fourth aspect, in some embodiments, the user operation for selecting the editing area may include: a user operation for drawing the editing area in the first picture.
[0045] In combination with the fourth aspect, in some embodiments, distinctively displaying the editing area in the first picture may include one or more of the following methods: highlighting the outline of the editing area, highlighting the entire editing area, or displaying a dotted box along the outline of the editing area.
[0046] In combination with the fourth aspect, in some embodiments, after the human-computer interaction module detects that the user inputs a first editing instruction for the editing area, the method may further include: if the first editing instruction is unreasonable to be applied to the editing area, the human-computer interaction module re-recommends the editing area; the re-recommended editing area is found by the calculation module from an area outside the editing area based on the feature vector corresponding to the first editing instruction.
[0047] In combination with the fourth aspect, in some embodiments, after the human-computer interaction module detects that the user enters a first editing instruction for the editing area, the method may further include: if the first editing instruction is unreasonable to be applied to the editing area, the human-computer interaction module recommends a modified first editing instruction; the editing type and / or editing parameters of the modified first editing instruction are found by the calculation module from the recommendation pool based on the fusion feature vector of the editing area.
[0048] In a fifth aspect, an embodiment of the present application provides a picture editing method, which is applied to a computing module, the computing module being included in a picture editing system, and the picture editing system further comprising: a human-computer interaction module;
[0049] The method may include: a computing module generates a recommended editing instruction based on image features of an editing area of a first image, and sends the recommended editing instruction to a human-computer interaction module, so that the human-computer interaction module displays the recommended editing instruction; then, the computing module receives the first editing instruction sent by the human-computer interaction module, and performs image editing processing on the first image according to the first editing instruction; finally, the computing module sends the first image after image editing processing to the human-computer interaction module, so that the human-computer interaction module displays the first image after image editing processing.
[0050] In conjunction with the fifth aspect, in some embodiments, the computing module generates recommended editing instructions based on the image features of the editing area. Specifically, the computing module may use a fused feature vector of multiple image features of the editing area as one input to a first artificial intelligence algorithm, and use one or more preset editing types as a second input to the first artificial intelligence algorithm. The computing module then calculates the recommended editing instructions through the operation of the first artificial intelligence algorithm, where the recommended editing instructions include editing parameters corresponding to the preset editing types. The multiple image features may include multiple features of mask features, depth features, contour features, and color features.
[0051] In combination with the fifth aspect, in some embodiments, the preset editing type includes one or more of the following: deletion, dragging, replacement, addition, or color adjustment.
[0052] In conjunction with the fifth aspect, in some embodiments, the method of the fifth aspect may further include: the computing module obtaining preprocessing information of the first image, and determining an editing region from the region indicated by the preprocessing information. In this way, the editing region can be determined based on the preprocessing information of the image, rather than image segmentation technology, eliminating the need for repeated online calculations.
[0053] In conjunction with the fifth aspect, in some embodiments, the preprocessing information may include indicative information for multiple regions, where the preprocessing information includes coordinates of contour points, a binary image, and a grayscale image for each of the multiple regions. In the binary image, the values of the multiple regions are first values, and in the grayscale image, the grayscale values of the multiple regions are first grayscale values or a first grayscale range. Thus, determining the editing region from the regions indicated by the preprocessing information may specifically include: determining the region within the multiple regions where the user's operation is located as the editing region, or determining the region within the multiple regions closest to the operation location as the editing region.
[0054] In conjunction with the fifth aspect, in some embodiments, the pre-processing information may include only information indicating a region, such as coordinates of contour points of the region, a binary image or grayscale image indicating the region, etc. Thus, determining the editing region from the region indicated by the pre-processing information may specifically include directly determining the region as the editing region.
[0055] In conjunction with the fifth aspect, in some embodiments, the pre-processing information of the first image may not directly indicate one or more regions, but may include other data, such as multi-layer information, contour information, or depth information. In this way, the other data may be used to first determine one or more regions, and then the editing area may be determined from these one or more regions.
[0056] Methods for determining one or more regions indirectly indicated by the pre-processing information may include, but are not limited to:
[0057] If the preprocessing information includes layer information of multiple layers, and the layer information of each layer includes the coordinates of opaque pixels in the layer, then: before determining the editing area from the area indicated by the preprocessing information, the calculation module can determine the connected opaque pixels in each layer as an area indicated by the preprocessing information.
[0058] If the preprocessing information includes contour information, then: before determining the editing area from the area indicated by the preprocessing information, the calculation module may determine the area surrounded by the contour indicated by the contour information as the area indicated by the preprocessing information.
[0059] If the preprocessing information includes depth information, the calculation module may determine pixels with the same or similar depth values as an area indicated by the preprocessing information based on the depth information before determining the editing area from the area indicated by the preprocessing information.
[0060] In conjunction with the fifth aspect, in some embodiments, the method of the first aspect may further include: before performing image editing processing on the first image according to the first editing instruction, if it is determined that the first editing instruction is not reasonably applied to the editing area, causing the computing module to re-recommend an editing instruction or re-recommend an editing area. This solves the problem of users randomly entering editing instructions, resulting in editing results that do not conform to common sense logic, thereby reducing the number of trial and error times when users edit images.
[0061] In combination with the fifth aspect, in some embodiments, the calculation module can determine whether it is reasonable to apply the first editing instruction to the editing area by comparing the feature vector corresponding to the first editing instruction with the fusion feature vector of each image feature of the editing area.
[0062] In combination with the fifth aspect, in some embodiments, re-recommending the editing area may specifically include: the computing module traverses the area outside the editing area in the first image, compares the fused feature vector of the traversed area with the feature vector corresponding to the first editing instruction, finds the area that can be reasonably matched with the first editing instruction, and re-recommends the found area as the editing area.
[0063] In combination with the fifth aspect, in some embodiments, re-recommending editing instructions may specifically include: a computing module traversing the editing types and / or editing parameters in the recommendation pool, comparing the feature vectors corresponding to the traversed editing types and / or editing parameters with the fusion feature vector of the editing area, finding the editing types and / or editing parameters that can be reasonably matched with the editing area, and modifying the first editing instruction according to the found editing type and / or editing parameters, and re-recommending the modified first editing instruction.
[0064] In conjunction with the fifth aspect, in some embodiments, the editing process corresponding to the first editing instruction may include: adding a new object to the editing area. After such image editing process, in the editing area of the first image, the depth features of the first pixel area are replaced with the depth features of the new object, and the depth features of the second pixel area remain the depth features of the original object, wherein the perspective relationship of the first pixel area is that the new object is in front of the original object, and the perspective relationship of the second pixel area is that the original object is in front of the new object.
[0065] In combination with the fifth aspect, in some embodiments, the computing module performs image editing processing on the first picture according to the first editing instruction, including: the computing module corrects the image features of the first picture, and regenerates the first picture using the corrected image features of the first picture.
[0066] In conjunction with the fifth aspect, in some embodiments, the image features may include depth features. The calculation module corrects the image features of the first image, specifically including: the calculation module determines the perspective relationship between the new object and the original object based on the image features of the first image and the editing parameters in the first editing instruction; determines the reference depth of the new object based on the perspective relationship between the new object and the original object, and the depth value of the original object, and then corrects the depth features of the new object using the reference depth; replaces the original depth features of the editing area with the corrected depth features of the new object; wherein, after the depth features of the new object are corrected, the average depth of the new object is close to or equal to the reference depth, and the depth difference between various areas on the new object remains unchanged.
[0067] In combination with the fifth aspect, in some embodiments, the image feature may further include a first image feature, which is an image feature other than a depth feature and includes one or more of the following: a mask feature, a contour feature, and a color feature.
[0068] The calculation module corrects the image features of the first image, and may also include: in the editing area, if the depth feature of a pixel area is replaced with the corrected depth feature of the new object, the calculation module uses the first image feature of the pixel area on the new object to replace the original first image feature of the pixel area in the editing area.
[0069] In the sixth aspect, an embodiment of the present application provides a terminal device, which may include: a human-computer interaction module, a processor and a memory, wherein the human-computer interaction module is coupled to the processor, and the memory is coupled to the processor; the human-computer interaction module may include input and output components such as a touch screen; wherein the memory can be used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the terminal device executes the method described in any one or more embodiments of the aforementioned fourth aspect.
[0070] Unlike the terminal device provided in the second aspect, which has powerful computing capabilities, the terminal device provided in the sixth aspect does not have powerful computing capabilities. It is necessary to submit complex computing tasks (such as editing instruction recommendation algorithms, image processing algorithms, image semantic understanding algorithms, etc.) to the cloud through the network for execution and wait for the task execution results returned by the cloud-side server. When the terminal device does not have powerful computing capabilities, the aforementioned human-computer interaction module can be deployed on the terminal device, and the aforementioned computing module can be deployed on the server. The steps executed by the human-computer interaction module are the steps executed by the terminal device, and the steps executed by the computing module are the steps executed by the server. The communication or data exchange between the two is considered inter-device communication.
[0071] In the seventh aspect, an embodiment of the present application provides a server, which may include: a processor and a memory, wherein the memory is coupled to the processor; wherein the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the terminal device executes the method described in any one or more embodiments of the aforementioned fifth aspect.
[0072] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium comprising instructions, characterized in that when the instructions are executed on a terminal device, the terminal device executes the method described in any one or more embodiments of the aforementioned fourth aspect.
[0073] In the ninth aspect, an embodiment of the present application provides a computer-readable storage medium, comprising instructions, characterized in that when the instructions are executed on a terminal device, the terminal device executes the method described in any one or more embodiments of the aforementioned fifth aspect.
[0074] In the tenth aspect, an embodiment of the present application provides a picture editing system, which may include: a terminal device and a server, wherein the terminal device is the terminal device described in the sixth aspect, and the server is the server described in the seventh aspect.
[0075] In an eleventh aspect, an embodiment of the present application provides a picture editing system, which may include: a human-computer interaction module and a computing module, wherein:
[0076] The human-computer interaction module may be configured to display the first image, detect a user operation on the first image for selecting an editing area, and notify the calculation module of an operation position of the user operation on the first image;
[0077] The calculation module may be configured to determine an editing area from the first image according to an operation position of a user operation on the first image, and inform the human-computer interaction module of the editing area;
[0078] The human-computer interaction module may also be used to distinguish and display the editing area in the first picture;
[0079] The calculation module may also be used to generate a recommended editing instruction based on the image features of the editing area, and inform the human-computer interaction module of the recommended editing instruction;
[0080] The human-computer interaction module may also be configured to display recommended editing instructions, then detect that a user inputs a first editing instruction for the editing area, and send the first editing instruction to the computing module; the first editing instruction includes: a recommended editing instruction;
[0081] The computing module may also be configured to perform image editing processing on the first image according to the first editing instruction, and send the first image after image editing processing to the human-computer interaction module;
[0082] Finally, the human-computer interaction module can also be used to display the first picture after image editing.
[0083] The image editing system provided in the eleventh aspect can be deployed in the same device (such as a terminal device such as a mobile phone with powerful computing capabilities) or in two devices (a terminal device and a server).
[0084] In the picture editing system provided in the eleventh aspect, the human-computer interaction module can execute the method described in any one or more embodiments of the aforementioned fourth aspect, and the computing module can execute the method described in any one or more embodiments of the aforementioned fifth aspect.
[0085] In the twelfth aspect, an embodiment of the present application provides a chip system, which is applied to a terminal device. The chip system includes one or more processors, which are used to call computer instructions so that the terminal device can execute the method described in the aforementioned first aspect, or any one or more embodiments of the fourth aspect.
[0086] In the thirteenth aspect, an embodiment of the present application provides a chip system, which is applied to a server. The chip system includes one or more processors, which are used to call computer instructions so that the server can execute the method described in any one or more embodiments of the aforementioned fifth aspect.
[0087] In the fourteenth aspect, the present application provides a computer program product comprising instructions, which, when run on an electronic device, enables the electronic device to execute the method described in any one or more embodiments of the first, fourth, or fifth aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] FIG1 shows a picture editing system provided by an embodiment of the present application;
[0089] FIG2 shows a terminal device provided in an embodiment of the present application;
[0090] FIG3 shows a server provided in an embodiment of the present application;
[0091] FIG4 shows the relationship between various method embodiments of the present application;
[0092] FIG5 shows a picture editing method provided by an embodiment of the present application;
[0093] FIG6 shows an example of a user opening a picture;
[0094] FIG7 shows an example in which a user selects the editing area “sky” in a picture;
[0095] FIG8 shows an example of a user drawing a heart-shaped editing area in a picture;
[0096] FIG9 shows an example of prompting the user to input an editing instruction by voice or text;
[0097] FIG10 shows an example of outputting a recommended editing instruction to a user;
[0098] FIG11 shows another example of outputting a recommended editing instruction to a user;
[0099] FIG12 exemplarily shows the digitized form of the mask feature of an image;
[0100] FIG13 exemplarily shows a visualization form of a mask feature;
[0101] FIG14 exemplarily shows a visualization form of deep features;
[0102] FIG15 exemplarily shows a visualization form of contour features;
[0103] FIG16 exemplarily shows a visualization form of color features;
[0104] FIG17 exemplarily shows the result of executing the editing command “add ‘flying geese’” on the editing area “sky”;
[0105] FIG18 exemplarily shows the result after executing the editing command “adjust hue: reduce brightness to 82%, reduce saturation to 75%, and change hue to red” for the editing area “sky”;
[0106] FIG19 exemplarily shows an image of “flying geese”;
[0107] FIG20 shows another picture editing method provided by an embodiment of the present application;
[0108] FIG21 exemplarily shows that the pre-processing information of an image only includes the coordinate data of the contour points of one region;
[0109] FIG22 exemplarily shows that the pre-processing information of an image only includes binary data of one region;
[0110] FIG23 exemplarily shows determining a plurality of regions indicated by pre-processing information based on pre-processing information including contour information;
[0111] FIG24 exemplarily shows determining a plurality of areas indicated by pre-processing information based on pre-processing information including depth information;
[0112] FIG25 shows another picture editing method provided by an embodiment of the present application;
[0113] FIG26 shows a method flow for determining whether a first editing instruction and an editing area are properly matched, provided by an embodiment of the present application;
[0114] FIG27 shows an example of prompting the user that the first editing instruction and the editing area are not properly matched;
[0115] FIG28 shows an example of prompting the user to change the editing area;
[0116] FIG29 shows an example of prompting the user to change the editing instruction;
[0117] FIG30 shows a method flow of how to execute an image involving a newly added object provided by an embodiment of the present application;
[0118] FIG31 exemplifies a situation where the perspective relationship between the new object and the original object in the image is unreasonable;
[0119] FIG32 exemplarily shows a new object image;
[0120] FIG33 exemplarily shows pixel areas in the editing area where the new object mask feature cannot be used;
[0121] FIG34 exemplarily shows a case where the perspective relationship between the new object and the original object in the picture is reasonable;
[0122] FIG35 exemplarily shows a size comparison between the new object image area and the area in the original image where feature replacement (such as depth feature replacement, mask feature replacement, contour feature replacement, etc.) occurs. DETAILED DESCRIPTION
[0123] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application.
[0124] To reduce the complexity of image editing, some simple image editing functions are widely used, such as erasing local areas of an image, moving the position of an object in an image, adjusting the color or brightness of an object in an image, etc. In these image editing functions, users do not need to perform complex editing processes such as cutting out, redrawing, and coloring local areas in the image. They only need to select the target area or target object by clicking and enter editing instructions such as deletion, color adjustment, and moving the object position. However, this type of image editing function only supports a few inherent editing operations and does not support user-defined editing operations. Moreover, it only supports users to process content that already appears in the image and does not support users to add new content to the image. In addition, it cannot prompt users whether the editing operation is reasonable. For example, "dragging a street bench to the sky" is unreasonable.
[0125] The embodiments of the present application provide a picture editing method, related equipment and system, which can lower the threshold for picture editing, support user-defined editing instructions, and understand the user's operating intentions, provide users with editing ideas, avoid users from experiencing a lot of trial and error, and significantly improve the efficiency of picture output.
[0126] The image editing method provided in the embodiment of the present application can be implemented based on the image editing system 10 shown in Figure 1. As shown in Figure 1, the image editing system 10 may include a human-computer interaction module 100 and a computing module 200. The human-computer interaction module 100 and the computing module 200 work together to implement the image editing method provided in the embodiment of the present application.
[0127] Among them, the human-computer interaction module 100 can serve as a human-computer interaction interface for users to use the picture editing function. Picture editing programs, such as photo albums (also known as gallery), photo editing applications, drawing design programs, etc., can be run on the human-computer interaction module 100. The human-computer interaction module 100 can provide input capabilities such as touch input, audio input, gesture input, and output capabilities such as display output and audio output, so that users can use picture editing functions by clicking, dragging, text input, gestures, etc., such as adding new objects to the picture, replacing an object in the picture, erasing a local area in the picture, moving the position of an object in the picture, adjusting the color or brightness of an object in the picture, etc. The human-computer interaction module 100 can also be used to communicate with the computing module 200 to transmit editing parameters (such as the position clicked by the user, the image features of the editing area, etc.) to the computing module 200, so that the computing module performs the complex calculations involved in the picture editing function according to the editing parameters. What are the editing parameters and what kind of calculations the computing module performs according to the editing parameters will be explained in detail in the following embodiments, and will not be expanded here.
[0128] The calculation module 200 is responsible for the complex calculations involved in image editing, such as editing instruction recommendation algorithms, image processing algorithms, image understanding algorithms, etc. The calculation module 200 can also be used to communicate with the human-computer interaction module 100 to receive editing parameters transmitted by the human-computer interaction module 100 and return processing results to the human-computer interaction module 100, so that the human-computer interaction module 100 can provide editing suggestions to the user or display the edited image based on the processing results.
[0129] The human-computer interaction module 100 and the computing module 200 can be respectively in different devices. For example, the human-computer interaction module 100 can be in a terminal device such as a mobile phone, tablet computer, smart screen, smart watch, etc., and the computing module 200 can be in a device such as a cloud-side server that can provide stronger computing power, or in other terminal devices with surplus computing power. In this case, the communication between the two is device-to-device communication, which can be mobile communication such as 2G / 3G / 4G / 5G, wireless fidelity (Wi-Fi) communication, satellite communication and other wireless communication, or Ethernet communication, Universal Serial Bus (USB) communication and other wired communication.
[0130] The human-computer interaction module 100 and the computing module 200 can also be integrated into the same device. For example, both are located in a terminal device such as a mobile phone, tablet computer, or personal computer that has both human-computer interaction capabilities and complex computing capabilities. In this case, the communication between the two is intra-device communication, which can be bus communication, shared memory communication, or other intra-device communication methods.
[0131] FIG2 exemplarily shows a terminal device 300 provided in an embodiment of the present application.
[0132] The terminal device 300 may have both human-computer interaction and computing capabilities. The terminal device 300 may be a mobile phone, tablet computer, handheld computer, desktop computer, laptop computer, ultra-mobile personal computer (UMPC), netbook, cellular phone, personal digital assistant (PDA), smart home devices such as smart screens, wearable devices such as smart watches and smart glasses, extended reality (XR) devices such as augmented reality (AR), virtual reality (VR), and mixed reality (MR), in-vehicle devices, or smart city devices, among others.
[0133] As shown in FIG2 , the terminal device 300 may include: a processor 110, a memory 120, a display 130, a display driver integrated circuit (DDIC) 140, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, and a subscriber identification module (SIM) card interface 195. The sensor module 180 may include a gyroscope sensor 180B, an acceleration sensor 180E, and a touch sensor 180K. The various components of the terminal device 300 may be connected via a bus.
[0134] Among them, the processor 110 can be responsible for providing computing power and can be used as a computing module of the terminal device; the display 130, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193 and other input and output components can be responsible for providing human-computer interaction capabilities and can be used as a human-computer interaction module of the terminal device. When the computing module in the terminal device has strong computing power, the terminal device 300 can be independently implemented as the picture editing system 10 shown in Figure 1; in this case, the part of the terminal device 300 responsible for providing computing power (such as the processor) can constitute the computing module 200 in the picture editing system 10, and the part of the terminal device 300 responsible for providing human-computer interaction capabilities (such as the display screen, touch sensor, etc.) can constitute the human-computer interaction module 100 in the picture editing system 10. When the computing module in the terminal device does not have strong computing capabilities, the terminal device 300 can also be implemented as only the human-computer interaction module 100 in the picture editing system 10, and the computing module 200 in the picture editing system 10 can be implemented by the cloud-side server.
[0135] The processor 110 may be one or more processors, which may be integrated into an integrated circuit of a system on chip (SOC). An SOC is a system-on-chip. The processor 110 may include a central processing unit (CPU), a graphics processing unit (GPU), a neural-network processing unit (NPU), etc. The CPU may include an application processor (AP), a baseband processor chip (BP), etc. The AP may be responsible for running the operating system, user interface, and application programs on the terminal device; the BP may be responsible for transmitting and receiving wireless signals and managing radio frequency services. The GPU may be responsible for graphics rendering, shading according to rendering instructions and data from the CPU, filling, rendering, and outputting materials, etc. The NPU draws on the structure of biological neural networks, such as the transmission mode between neurons in the human brain, to quickly process input information and can also continuously self-learn. The NPU can be used to run artificial intelligence algorithms, such as editing instruction recommendation algorithms, image processing algorithms, image understanding algorithms, etc. The CPU and GPU can be used to render and synthesize images to be sent to the display 130.
[0136] The processor 110 may include one or more interfaces, such as an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input and output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0137] The processor 110 may be provided with a cache memory for storing instructions or data that have just been used or are cyclically used by the processor 110. If the processor 110 needs to use the instruction or data again, it can be directly called from the cache memory, which can reduce the waiting time of the processor 110 and improve the efficiency of program operation.
[0138] The memory 120 may include a program storage area and a user data storage area. The program storage area may store an operating system and one or more application programs (such as game applications), and the data storage area may store data (such as photos and contacts) created by the user during the use of the terminal device 300. The memory 120 may be a high-speed random access memory or a non-volatile memory, such as a disk, flash memory, or universal flash storage (UFS). The memory 120 may also be an external memory card, such as a Micro SD card.
[0139] The memory 120 may also store code instructions of the image editing method provided in the embodiment of the present application. When the processor 110 reads the code instructions from the memory 120 and runs the code instructions, the terminal device 300 can execute the steps performed by the human-computer interaction module and / or computing module in the image editing method provided in the embodiment of the present application.
[0140] The memory 120 may also be integrated with the processor 110 into an integrated circuit of a SOC.
[0141] As shown in FIG. 2 , the terminal device 300 can implement a display function through the SOC, the DDIC 140 , and the display 130 .
[0142] Display 130 has multiple refresh rates. A refresh rate indicates how many times a display refreshes its image in one second. For example, a 60 Hz refresh rate indicates that the display refreshes its image 60 times in one second. Display 130 can utilize an LTPO display panel, allowing the refresh rate to be reduced to a low refresh rate, such as 10 Hz or 1 Hz, thereby reducing display power consumption.
[0143] The display driver integrated circuit (DDIC) 140 serves as the control core of the display 130, driving the display 130 and receiving data, such as image data and instructions, from the SOC (processor 110). DDIC 140 transmits drive signals and data to the display panel of the display 130 in the form of electrical signals, thereby controlling the screen brightness and color, allowing image information such as letters and pictures to appear on the screen and completing screen refresh.
[0144] The image data to be displayed sent by the SOC to the DDIC 140 may be stored in a frame buffer for display transmission (or image transmission). Then, the DDIC 140 retrieves the image data from the frame buffer and drives the display 130 for display.
[0145] The wireless communication function of the terminal device 300 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor.
[0146] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal device 300 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0147] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the terminal device 300. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0148] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display 130. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.
[0149] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the terminal device 300. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0150] In some embodiments, the antenna 1 of the terminal device 300 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the terminal device 300 can communicate with the network and other devices through wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).
[0151] The terminal device 300 can realize the shooting function through the ISP, camera 193, video codec, GPU, display 130 and application processor.
[0152] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.
[0153] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the terminal device 300 may include 1 or N cameras 193, where N is a positive integer greater than 1.
[0154] Video codecs are used to compress or decompress digital video. Terminal device 300 may support one or more video codecs. This allows terminal device 300 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0155] The terminal device 300 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0156] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0157] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The terminal device 300 can listen to music or listen to hands-free calls through the speaker 170A.
[0158] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the terminal device 300 receives a call or voice message, the user can hear the voice by placing the receiver 170B close to the ear.
[0159] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The terminal device 300 can be provided with at least one microphone 170C. In other embodiments, the terminal device 300 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the terminal device 300 can also be provided with three, four or more microphones 170C to realize sound signal collection, noise reduction, and identification of sound sources, and realize directional recording function, etc.
[0160] The headphone jack 170D is used to connect a wired headphone and can be a USB interface, or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0161] Keys 190 include a power button, volume button, and other buttons. Keys 190 can be mechanical or touch-sensitive. Terminal device 300 can receive key inputs and generate key signal inputs related to user settings and function control of terminal device 300. Motor 191 can generate vibration prompts. SIM card interface 195 is used to connect a SIM card. A SIM card can be connected to and disconnected from terminal device 300 by inserting or removing it from SIM card interface 195.
[0162] The structure shown in FIG2 does not constitute a specific limitation on terminal device 300. Terminal device 300 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The various components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0163] FIG3 exemplarily shows a server 400 provided in an embodiment of the present application.
[0164] The server 400 can provide complex computing capabilities for executing the steps performed by the computing module in the image editing method provided in the embodiment of the present application.
[0165] As shown in FIG3 , the server 400 may include: a processor 210 , a memory 220 , an input / output device 230 , a communication module 240 , etc. These components may be coupled via a bus.
[0166] The server 400 may have powerful computing resources, and the processor 210 thereon may include one or more processors with powerful computing power, such as a central processing unit (CPU), a neural network processing unit (NPU), a graphics processing unit (GPU), etc.
[0167] The processor 210 may include one or more interfaces, such as an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input and output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0168] The processor 210 may be provided with a cache memory for storing instructions or data that have just been used or are cyclically used by the processor 210. If the processor 210 needs to use the instruction or data again, it can be directly called from the cache memory, which can reduce the waiting time of the processor 210 and improve the efficiency of program operation.
[0169] The processor 210 may also be connected to an external memory. The memory may be a high-speed random access memory or a non-volatile memory, such as a disk, flash memory, universal flash memory (UFS), etc. The memory may also be an external memory card, such as a Micro SD card.
[0170] Processor 210 is the computing core of server 400 and possesses powerful computing capabilities. It is coupled to memory 220 and can be used to read and execute computer-readable instructions stored in memory 220, running an operating system and various programs. Specifically, CPU 210 can be used to call programs stored in memory 220, such as the program implementing the image editing method provided in the embodiments of the present application, and execute the instructions contained in the program.
[0171] The memory 220 may include a high-speed random access memory, a non-volatile memory, such as a disk, a flash memory or other non-volatile solid-state storage device. The memory 220 can be used to store various software programs and multiple sets of instructions. The memory 220 can store an operating system, such as an operating system such as Linux. The memory 220 can also store one or more programs, such as programs involved in patch production, such as a compiler and a linker. The memory 220 can also store code instructions of the image editing method provided in the embodiment of the present application. When the processor 210 reads the code instructions from the memory 220 and runs the code instructions, the server 400 can execute the steps performed by the computing module in the image editing method provided in the embodiment of the present application.
[0172] The input and output devices 230 may include a display screen, a keyboard, a mouse and other devices, which can be used to receive user input and output program execution results to the user.
[0173] The communication module 240 may include a wired communication module and a wireless communication module. The wired communication module may support wired communication protocols such as universal serial bus (USB), serial port, Ethernet, and other protocols, and communicate with other devices through physical communication cables. The wireless communication module may include 2G / 3G / 4G / 5G wireless communication modules, Wi-Fi communication modules, etc. The wireless communication module receives electromagnetic waves via an antenna, modulates and filters the electromagnetic wave signals, and sends the processed signals to the CPU 210. The wireless communication module may also receive signals to be transmitted from the CPU 210, modulate and amplify them, and convert them into electromagnetic waves for radiation via the antenna.
[0174] The structure shown in FIG3 does not constitute a limitation on the server 400. The server 400 may include more or fewer components than shown, or combine or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0175] Based on the various products introduced in the aforementioned embodiments, the picture editing method provided by the embodiments of the present application will be described in detail through four embodiments below. These four embodiments may be referred to as Example 1, Example 2, Example 3 and Example 4, respectively. As shown in Figure 4, Example 1 introduces the overall process of the picture editing method provided by the present application, and Example 2 is an alternative to Example 1. Examples 1 and 2 introduce two user operation methods, both of which can predict the user's editing intentions and give recommended editing operations. Example 3 introduces how to judge whether the editing operation input by the user is reasonable. Example 4 solves the problem of how to insert the element into the middle of an existing object in the original image when adding or replacing an element. Examples 1 and 2 are in a parallel relationship, and Examples 3 and 4 are supplements to Examples 1 and 2, and are an expanded introduction to certain steps in the overall process.
[0176] Example 1
[0177] Example 1 introduces the overall process of a user editing a picture through a terminal device. When the terminal device has powerful computing capabilities, complex image processing can be completed directly on the computing module of the terminal device. When the terminal device does not have powerful computing capabilities, the terminal device can submit complex computing tasks (such as complex image processing tasks) to the cloud through the network for execution and wait for the task execution results returned by the cloud-side server. In the subsequent method embodiments, the technical solution will be described with the human-computer interaction module and the computing module as the execution subjects. Then, when the terminal device has powerful computing capabilities, the human-computer interaction module and the computing module can both be located on the terminal device, and the steps performed by the two are the steps performed by the terminal device, and the communication or data interaction between the two belongs to intra-device communication; when the terminal device does not have powerful computing capabilities, the human-computer interaction module can be located on the terminal device, and the computing module can be located on the server. The steps performed by the human-computer interaction module are the steps performed by the terminal device, and the steps performed by the computing module are the steps performed by the server, and the communication or data interaction between the two belongs to inter-device communication.
[0178] As shown in FIG5 , the overall process of the image editing method provided in the first embodiment may include:
[0179] S10-S13: Open the first picture.
[0180] Specifically, as described in S10, the human-computer interaction module may detect a user operation of opening the first image. In response, as described in S11-S12, the human-computer interaction module may retrieve the first image from the storage module. As described in S13, after retrieving the first image, the human-computer interaction module may display the first image.
[0181] The terminal device may be installed with an application for browsing, managing or processing pictures, such as a photo album (also known as a gallery), a picture editing application, a drawing design program, etc. The user can open the first picture through the application.
[0182] In the embodiment of the present application, the user operation of opening the first picture can be referred to as the first user operation.
[0183] As shown in Figure 6, the first user operation may be, for example, a user clicking on a thumbnail 61 of a first image in an album. In response to this operation, the terminal device displays the original image 63 of the first image. In contrast to a thumbnail, the original image can also be referred to as a large image. Not limited to Figure 6, the user operation of opening the first image may also be a user operation of clicking on a thumbnail in a folder, a user operation of clicking on a thumbnail on a webpage, a user operation of clicking on a thumbnail in a chat interface, and so on. The first image user operation may also be a user operation of clicking on a link to the first image, a user operation of clicking on a file icon for the first image, and so on.
[0184] The first image can be stored in a local storage module of the terminal device or in a network storage module. When the first image is stored on the network, the human-computer interaction module can download the first image from the network storage module.
[0185] S14-S18: Identify the editing area.
[0186] Specifically, as described in S14, the human-computer interaction module can detect the user's operation of selecting an editing area in the first picture. In response to this, as described in S15, the human-computer interaction module can transmit the operation position of the user's operation on the first picture to the calculation module. Then, as described in S16, the calculation module can determine which area in the picture the editing area selected by the user is based on the operation position, and inform the human-computer interaction module of the contour information of the editing area selected by the user as described in S17. In this way, as described in S18, the human-computer interaction module can distinguish and display the editing area in the first picture based on the contour information of the editing area selected by the user. Not limited to the contour information of the editing area, the information returned by the calculation module to the human-computer interaction module in S17 can also be a binary image, grayscale image, etc. of the editing area. Like the contour information, they can uniquely indicate the editing area in the first picture. In the embodiment of the present application, this information is collectively referred to as indication information of the editing area.
[0187] Here, "distinguishing" means distinguishing the editing area from other areas in the first image, and its implementation methods may include but are not limited to: highlighting the outline of the editing area, highlighting the entire editing area, or displaying a dotted frame along the outline of the editing area, etc.
[0188] In the embodiment of the present application, the operation of the user selecting the editing area in the first picture can be called the second user operation, and the editing area selected by the user can be called the first editing area.
[0189] The second user operation may be an operation of selecting an object in the first image, such as clicking or long pressing an object, and the first editing area may be the image area where the user-selected object is located. For example, as shown in FIG7 , in first image 71 , the user clicks on the object "sky," and the first editing area is the image area of "sky." In response, the human-computer interaction module may display a dotted box around the outline of "sky" to distinguish "sky" from other areas as the first editing area.
[0190] The first editing area can specifically be identified by the computing module based on the location of the second user operation on the first image. Specifically, the human-computer interaction device can send the location of the second user operation on the screen to the computing module, which then identifies the first editing area based on the location. The identification of the first editing area can be based on existing image segmentation techniques. Image segmentation techniques can be used to identify and classify each object instance within an image and label each pixel in the image with a corresponding object category label. After using image segmentation techniques to identify and classify each object within the image, the computing module can determine which object within the image the user selected based on the location of the second operation on the image and determine the image area containing the object as the first editing area. This embodiment of the present application, based on an understanding of the semantic content of the image, predicts the scope of the editing area intended by the user based on the user operation. This allows users to simply and efficiently select an editing area within the image to be edited, eliminating the need for users to perform complex operations such as lassoing, cropping, and separating layers. This not only reduces user complexity but also improves editing efficiency.
[0191] An example of an existing image segmentation technique is the Segment Anything Model (SAM), a deep learning model for image segmentation tasks. SAM combines a convolutional neural network (CNN) and a Transformer-based architecture to process images in a hierarchical and multi-scale manner. The working principle of SAM can be summarized as follows: SAM uses a pre-trained Vision Transformer (VIT) as its backbone network. The backbone network is used to extract features from the input image. SAM uses a Feature Pyramid Network (FPN) to generate feature maps at multiple scales. The FPN is a series of convolutional layers that operate at different scales to extract features from the backbone network output. The FPN ensures that SAM can identify objects and boundaries at different levels of detail. SAM uses a decoder network to generate a segmentation mask for the input image. The decoder network takes the FPN output and upsamples it to the original image size. The upsampling process enables the model to generate a segmentation mask with the same resolution as the input image. SAM also uses a Transformer-based architecture to improve segmentation results. The Transformer is a neural network architecture that is very effective at processing sequential data, such as text or images. Using a Transformer-based architecture improves segmentation results by incorporating contextual information from the input image. SAM utilizes self-supervised learning to learn from unlabeled data. This involves training the model on a large dataset of unlabeled images to learn common patterns and features within the images. The learned features can then be used to improve the model's performance on specific image segmentation tasks. SAM can perform panoptic segmentation, which involves combining instance and semantic segmentation. Instance segmentation involves identifying and classifying each object instance within an image, while semantic segmentation involves labeling each pixel in the image with a corresponding class label. Panoptic segmentation combines these two approaches to provide a more comprehensive understanding of the image.
[0192] Another existing image segmentation technology is the "segment everything everywhere all at once" (SEEM) technology. In terms of model architecture, the SEEM model uses a common encoder-decoder architecture. It not only enables image segmentation but also supports multimodal input, enabling one-click segmentation of any object of any category in any image through different types of visual prompts and text prompts. A visual prompt can be a user-selected point on an image, a box, a randomly drawn doodle, a mask, or a referenced area from another image; while a text prompt can be the class the user wants to segment or a sentence that represents the task. In other words, for an image, users can simply point to a point, draw a box, or scribble a few times to segment the corresponding object. Furthermore, users can also tell the model the name of the object to be segmented or describe it in a sentence using the method described in this article, which also enables one-click segmentation.
[0193] The second user operation is not limited to directly selecting an object. It can also involve drawing an editable area within the first image. For example, as shown in Figure 8 , in the first image 81, the user draws a heart outline 82. The image area within the heart outline is the editable area selected by the user. Besides drawing an outline, the second user operation can also involve painting out the editable area, and so on.
[0194] S19: Receive the editing instruction input by the user.
[0195] Specifically, the human-computer interaction module may detect that the user inputs a first editing instruction for the first editing area.
[0196] In the embodiment of the present application, after the editing area is identified and before S19 , the human-computer interaction module may further prompt the user to input an editing instruction.
[0197] The methods for prompting the user to enter editing instructions may include but are not limited to the following:
[0198] Method 1. As shown in Figure 9 , the human-computer interaction module pops up an input box 121 for the user to enter an editing instruction, such as a text or voice command. For example, the user can enter the following text or voice command in the input box: "Add 'Soaring Geese'." In this example, "Add" is the editing type, and "Soaring Geese" is the editing parameter. Figure 9 is for example only. This embodiment of the present application does not limit the appearance, style, display location, etc. of the input box 121.
[0199] When the user is prompted to input an editing instruction in method 1, the first editing instruction may be a text instruction input by the user in the input box or a voice instruction input by pressing the voice key.
[0200] Mode 2: The human-computer interaction module may display one or more recommended editing instructions, which are generated by the calculation module according to the image features of the first editing area.
[0201] For example, as shown in FIG10 , the editing area selected by the user is a heart-shaped area 131 in the sky, and the recommended editing instruction generated for the editing area is “recommended adding new screen content ‘Soaring geese’”, where “add” is the editing type and “Soaring geese” is the editing parameter.
[0202] For another example, as shown in Figure 11, the editing area selected by the user is the sky area, and the recommended editing instruction generated for this editing area is "Recommended adjustment of hue: reduce brightness to 82%, saturation to 75%, and hue to red", where "adjust hue" is the editing type, and "reduce brightness to 82%, reduce saturation to 75%, and hue to red" are editing parameters.
[0203] Compared with the existing technology that requires users to independently determine the editing type and editing parameters, the embodiment of the present application is based on the understanding of the semantic content of the original image, and recommends editing types and editing parameters to users according to the image features of the editing area selected by the user, providing users with editing ideas. This not only reduces the complexity of use, but also ensures the rationality of the content of the edited image, avoids users from experiencing a lot of trial and error, and significantly improves the efficiency of image output.
[0204] The image features of the first editing area may include, but are not limited to, mask features, depth features, contour features, and color features. These features may be vector features extracted from the first editing area using various existing technologies.
[0205] Specifically, the calculation module can use an image segmentation algorithm (such as SAM, SEEM algorithm) to extract the mask features of the image. The mask features can carry the category information corresponding to each pixel in the image. Figure 12 exemplifies the matrix storage form of the mask features of a picture including the sky, castle, mountain, valley, forest and tree. The size of the matrix is the same as the size of the original image, and the coordinates of the matrix correspond one-to-one with the coordinates of the original image. The values in the matrix represent the category to which the pixels at the corresponding positions of the original image belong, such as "0" for "sky", "1" for "forest", etc. The storage form of the mask features can be varied, as long as it can reflect the category information of each pixel, and the embodiment of the present application does not limit this. Regardless of the storage form, the mask features can be converted to obtain the visualization results shown in Figure 13.
[0206] Taking Figure 10 as an example, when the user selects the heart-shaped area 131 as the editing area, the computing module determines that the pixel category within the heart-shaped area 131 is "sky" based on the mask features. Based on a specific recommendation rule, the recommended editing type is "add," and the corresponding editing parameter is "flying geese." In this example, the specific recommendation rule includes: when the first editing area is the sky, the editing type is "add," and the editing parameter is "flying geese." In actual applications, the recommendation rule can be set to other values, and this embodiment of the application does not limit this.
[0207] The calculation module can also use the depth estimation algorithm to extract depth features, use the edge detection and contour extraction algorithm to extract contour features, use the main color extraction algorithm (such as the Kmeans clustering algorithm) to extract color features, and so on. Taking the picture including the sky, castle, mountain, valley, forest and tree as an example, the visualization form of the depth feature can be shown in Figure 14, the visualization form of the contour feature can be shown in Figure 15, and the visualization form of the color feature can be shown in Figure 16. Figure 16 visualizes the color moment of the picture and intuitively describes the color distribution in the picture through the second-order rectangular form. The storage form of the three features of depth, contour and color can also be in matrix form, similar to Figure 12, except that the values in the matrix represent different meanings. The matrix value of the depth feature represents the depth value of each pixel, the matrix value of the contour feature represents whether each element is a contour, and the matrix value of the color feature represents the color of each element.
[0208] After extracting various features of the first edit area, such as mask features, depth features, contour features, and color features, the calculation module can fuse these features using techniques such as weighted summation to obtain a fused feature vector. This fused feature vector is then used as one input to a specific artificial intelligence algorithm. Furthermore, the calculation module can map some edit types to numbers (e.g., "add" is mapped to 1, and "delete" is mapped to 2) and use this as a second input to the specific artificial intelligence algorithm. Ultimately, the calculation module can calculate the editing parameters corresponding to each edit type through the specific artificial intelligence algorithm. For a certain edit type, if the artificial intelligence algorithm cannot output the corresponding edit parameters or the confidence level of the outputted edit parameters is low (e.g., less than 60%), the edit type is determined to be unsuitable for the edit area and is not recommended. Here, the specific artificial intelligence algorithm can be a trusted model trained with a large number of samples. Each training sample in its training sample set can include an input sample and an output sample. The input sample includes the fused feature vector of the image area and the mapping ID of the edit type, and the output sample includes the editing parameters used when applying the edit type to the image area. The more reasonable the training samples and the larger the sample set, the more reliable the trained model.
[0209] In some embodiments, the computing module can also generate editing questions based on the image features of the editing area and the preset editing type. The answer to the editing question is how to match the editing parameters of the editing type, and then input it into a large model such as chatGPT to obtain the answer. For example, the editing question generated based on the editing area "sky" and the editing type "add" is "What to add to the sky?" The chatGPT model can output an answer, such as: "Add 'clouds' to the sky", where 'clouds' is the answer to the question by the large model. This example is only used to explain the embodiments of the present application. In actual applications, the questions posed to chatGPT can be more complex and carry more details, such as adding additional picture descriptions.
[0210] In some embodiments, the computing module can also traverse the editing instructions in the recommendation pool and compare the feature vectors corresponding to the traversed editing instructions with the fusion feature vectors of the editing area. For details, please refer to S52 in Example 3 to find editing instructions that can be reasonably matched with the editing area, and finally recommend editing instructions.
[0211] When prompting the user to enter an editing instruction using method 2, the first editing instruction entered by the user can be selected from the recommended editing instructions generated by the calculation module. For example, the user clicks on a recommended editing instruction to confirm entering the recommended editing instruction as the first editing instruction, or the user drags a recommended editing instruction to the first editing area to confirm entering the recommended editing instruction as the first editing instruction. This example is only used to illustrate the embodiments of the present application and may be different in actual applications. The embodiments of the present application do not limit the method by which the user enters the first editing instruction.
[0212] Mode 3: The human-computer interaction module may display one or more preset editing instructions, such as commonly used editing instructions or editing instructions saved in advance by the user.
[0213] When prompting the user to enter an editing instruction using method 3, the first editing instruction entered by the user can be selected from preset editing instructions displayed by the human-computer interaction module. For example, the user clicks on a preset editing instruction to confirm the recommended editing instruction as the first editing instruction. Another example is that the user drags a preset editing instruction to the first editing area to confirm the recommended editing instruction as the first editing instruction. This example is only used to explain the embodiments of the present application. In actual applications, it may be different. The embodiments of the present application do not limit the method by which the user enters the first editing instruction.
[0214] In the embodiment of the present application, the several methods of prompting the user to enter editing instructions described above can also be implemented in combination. For example, the human-computer interaction module can pop up an input box to prompt the user to enter editing instructions as in method 1, display recommended editing instructions as in method 2, and display threshold editing as in method 3. The user can see prompts in multiple ways on the interface at the same time and choose to enter editing instructions according to a certain prompt.
[0215] S20-S23: Edit pictures.
[0216] In response to the user inputting a first editing instruction for the first editing area, as described in S20, the human-computer interaction module may send the first editing instruction to the computing module. Then, as described in S21, the computing module may perform corresponding image editing processing on the first image according to the first editing instruction and return the first image after the image editing processing to the human-computer interaction module as described in S22. Then, as described in S23, the human-computer interaction module may display the first image after the image editing processing.
[0217] Compared to the image before the image editing process, the image in the first editing area after the image editing process has been changed, and the change is determined by the first editing instruction. The editing instruction may include the following: the editing type and its corresponding editing parameters. The editing type may include deletion, dragging, replacement, addition, color adjustment, etc. The corresponding editing parameters may include replaceable content, drag target location, added content, color adjustment value, etc.
[0218] For example, if the user selected the "sky" area for editing and the first editing instruction was "Delete 'dark clouds'," then the dark clouds in the "sky" of the first image after editing will be deleted compared to the image before editing. For another example, if the user selected the "sky" area for editing and the first editing instruction was "Add 'flying geese'," then the "flying geese" will be added to the "sky" of the first image after editing compared to the image before editing.
[0219] Assuming that the first picture is picture 71 shown in Figure 7 and picture 81 shown in Figure 8, Figure 17 exemplarily shows the first picture after editing when the first editing instruction is "add 'soaring geese'", and Figure 18 exemplarily shows the first picture after editing when the first editing instruction is "adjust hue: reduce brightness to 82%, saturation to 75%, and hue to red".
[0220] When the editing type in the first editing instruction is "add", the computing module can also obtain the image content to be added to the first picture according to the editing parameters in the first editing instruction, such as "soaring geese", and add the image content to the first picture to complete the image editing process corresponding to the first editing instruction (such as "add 'soaring geese'"). The image content can be carried in a material picture, which can come from the Internet, a terminal device, or the material picture can be generated by the computing module using artificial intelligence. Figure 19 exemplifies the material picture obtained according to the editing parameters ("soaring geese") of the editing instruction "add 'soaring geese'".
[0221] Example 2
[0222] Example 2 is an alternative to Example 1 and also introduces the overall process of the image editing method. However, in Example 2, the editing area can be determined based on the pre-processing information of the image instead of image segmentation technology, and repeated online calculations are not required.
[0223] As shown in FIG20 , the overall process of the image editing method provided in the second embodiment may include:
[0224] S30-S33: Open the first picture.
[0225] For details, please refer to S10-S13 in the first embodiment, which will not be described in detail here.
[0226] S34-S39: Identify the editing area.
[0227] Specifically, as described in S34, the human-computer interaction module can detect the user's operation of selecting an editing area in the first picture, such as clicking on an object in the picture. In response to this, as described in S35, the human-computer interaction module can trigger the calculation module to obtain the preprocessing information of the first picture; in addition, as described in S36, the human-computer interaction module can also transmit the operation position of the user's operation on the first picture to the calculation module. In this way, as described in S37, the calculation module can determine which area in the picture the editing area selected by the user is based on the preprocessing information, or further combined with the operation position, and inform the human-computer interaction module of the indication information (such as contour information, binary image, grayscale image) of the editing area selected by the user as described in S38. Then, as described in S39, the human-computer interaction module can distinguish and display the editing area in the first picture based on the contour information of the editing area selected by the user.
[0228] For some implementation details of "identifying the editing area," please refer to S14-S18 in Example 1 and the related descriptions. For example, the user's operation of selecting the editing area in the first image may be to click or long-press an object in the first image. As shown in Figure 7, in the first image 71, the user clicks the object "sky" to select "sky" as the editing area. For another example, the method of distinguishing the editing area in the first image may include, but is not limited to: highlighting the outline of the editing area, highlighting the entire editing area, or displaying a dotted box along the outline of the editing area.
[0229] In the second embodiment, the preprocessing information can be used to indicate one or more regions. Here, the one or more regions refer to image regions in the first image. The image regions in the first image can be divided into units of objects, and all pixels of an object constitute an image region. The editing region can be determined from these one or more regions. These one or more regions can be regions that the user prefers to edit, or regions that are recommended for editing, or regions that are allowed to be edited, etc. The editing region can be quickly determined based on the preprocessing information without having to perform image segmentation processing on the first image.
[0230] After the human-computer interaction module detects the user's operation of selecting the editing area in the first picture, the calculation module can first read the preprocessing information:
[0231] 1. If the pre-processing information only includes indication information of one area, then the area is determined as the editing area.
[0232] Preprocessing information can simply contain the coordinates of a region's outline points. As shown in Figure 21, preprocessing information can be stored in a JSON array format. Each array element represents the outline of a region, and each element in the array is a tuple. The first value in the tuple represents the x-coordinate value of a contour point in the region, and the second value in the tuple represents the y-coordinate value of the contour point. Therefore, based on the contour recorded in the array, the area enclosed by the contour can be determined, and this area can be identified as the editing area.
[0233] The pre-processing information is not limited to the Jason format and may also adopt other data formats.
[0234] Not limited to the coordinates of the contour points of the region, the data content in the preprocessing information can also be a binary image, a grayscale image, etc.
[0235] The preprocessing information may also be a binary image containing only one area. In the binary image, only one area has a first value. The binary image may be exemplified in FIG22 , wherein one element corresponds to one pixel, the coordinates of the element may represent the position of the pixel, the value of the element may represent the pixel category of the pixel, and is binarized, such as “1” may represent that the pixel at this position belongs to “sky”, and “0” may represent that the pixel at this position does not belong to “sky”. Based on the binary image shown in FIG22 , an area with a value of “1” (“sky” area, the first value is “1”) or an area with a value of “0” (non-“sky” area, the first value is “0”) may be determined as an editing area.
[0236] The pre-processing information may be a grayscale image, in which there may be only one region whose grayscale value is a specific grayscale value or is within a specific grayscale range. Therefore, the only region can be determined as the editing region based on the grayscale image.
[0237] 2. If the pre-processing information includes information indicating multiple regions, the editing region can be further determined based on the location of the second user's action on the first image. Specifically, the region of the image where the action location is located, or which region it is closest to, can be determined, and that region can be determined as the editing region. In other words, among the multiple regions indicated by the pre-processing information, the region where the action location is located can be determined as the editing region, or the region closest to the action location among the multiple regions can be determined as the user-selected region. Furthermore, the user-selected region can be determined as the editing region.
[0238] Furthermore, after determining the area selected by the user, it is also possible to determine whether there are areas in the remaining areas that overlap more with the selected area. If so, the selected area and the area that overlaps more with it are merged into one area, and the merged area is finally determined to be the editing area. In some embodiments, a non-maximum suppression (NMS) algorithm can be used to merge multiple overlapping areas. In the embodiment of the present application, more overlap can mean that the area of the overlapping part exceeds a specific value, such as 10 pixels (px).
[0239] In order to directly indicate multiple areas, the preprocessing information may include the coordinates of the contour points of the multiple areas, wherein the coordinates of the contour of each area can be expressed as an array. The preprocessing information may also be a binary image including multiple areas, wherein the values of multiple areas in the binary image are the first value (such as "1"). For example, "1" indicates that the pixel at this position belongs to "flower", and "0" indicates that the pixel at this position does not belong to "flower". The preprocessing information may also be a grayscale image, wherein the grayscale values of multiple areas in the grayscale image are specific grayscale values or are in a specific grayscale range, for example, there are multiple areas whose grayscale values are in the grayscale range of 0-50.
[0240] The preprocessing information is not limited to the coordinates, binary images, and grayscale images of the contour points of the region and can be used to directly indicate multiple regions. The preprocessing information can also have other forms of data content, which is not limited in the embodiments of the present application.
[0241] 3. The pre-processing information of the first image may not directly indicate one or more regions, but may include other data, such as multi-layer information, contour information, or depth information. However, the calculation module can use this other data to determine one or more regions, and then determine the editing area from these one or more regions based on the method described in 1 or 2 above. In other words, the pre-processing information of the first image can indirectly indicate one or more regions.
[0242] Methods for determining one or more regions indirectly indicated by the pre-processing information may include, but are not limited to:
[0243] 3.1 If the pre-processing information includes layer information for multiple layers, and each layer's layer information includes the coordinates of opaque pixels within that layer, then for each layer in the layer information, a contiguous area of opaque pixels within that layer can be considered a region. Furthermore, overlapping layer regions can be merged into a single region, for example, using the NMS algorithm to merge multiple overlapping layer regions, and the merged region or regions can be used to determine the editing area.
[0244] 3.2 If the preprocessing information includes contour information (as shown in Figure 18), which records the pixels on the contour, techniques such as dilation and erosion can be used to filter out useless internal contours within the contour information, and based on graph theory and other techniques, one or more regions containing the contour information can be obtained. Dilation and erosion can be used to eliminate noise, segment independent image elements, and connect adjacent elements; graph theory and other techniques can be used for connected domain analysis, which involves finding and marking independent connected domains in an image. A connected domain in an image is a region of adjacent pixels with the same pixel value. Generally, a connected domain contains only one pixel value. Therefore, to prevent pixel value fluctuations from affecting the extraction of different connected domains, connected domain analysis often processes binarized images.
[0245] Still taking the picture including the sky, castle, mountain, valley, forest and tree as an example, FIG23 exemplarily shows a plurality of regions obtained based on the contour information of the picture.
[0246] 3.3 If the pre-processing information includes depth information (as shown in Figure 19), and the depth information can record the depth values of multiple pixels, the pixels can be divided into different regions based on the depth value of each pixel in the depth information. Specifically, pixels with the same or similar depth values can be divided into the same region.
[0247] Still taking the picture including the sky, castle, mountain, valley, forest and tree as an example, Figure 24 exemplifies multiple areas obtained based on the depth information of the picture, among which the "castle" area 241 can be composed of some pixels with the same or similar depth values, and the "tower" area 242 can be composed of some pixels with the same or similar depth values.
[0248] In the embodiments of the present application, similar depth values may mean that the difference in depth values does not exceed a specific value, such as 0.1. In an image, the depth value of each pixel can be expressed in a range of 0 to 1, where 0 represents the depth value of the pixel farthest from the camera that captured the first image, and 1 represents the depth value of the pixel closest to the camera that captured the first image. Depth values may also be expressed in other ways, and the embodiments of the present application are not limited thereto.
[0249] The preprocessing information can simultaneously include multiple contents described in 3.1 to 3.3 above. For example, the preprocessing information can include layer information, depth information, and contour information. After the area indicated by the preprocessing information is determined based on these three contents, they can be calibrated with each other to improve the accuracy of recognition. For example, if the layer information and depth information both indicate the same area, the recognition of that area is often accurate; conversely, if the areas indicated by the layer information and depth information conflict, it means that the area recognition based on the depth information or layer information is inaccurate, and the algorithm can be optimized and re-recognized.
[0250] It can be seen that embodiment 2 can determine the editing area from one or more areas directly or indirectly indicated by the preprocessing information, without first using image segmentation processing to identify and divide the image area where each object in the first picture is located and then identify the editing area according to the user's operation position. Therefore, the area that the user wants to edit can be predicted more quickly, and image segmentation calculations can be avoided.
[0251] S40: receiving an editing instruction input by the user.
[0252] For details, please refer to S19 in Example 1, which will not be described again here.
[0253] S41-S44: Edit pictures.
[0254] For details, please refer to S20-S23 in Example 1, which will not be repeated here.
[0255] Example 3
[0256] Example 3 is a supplementary refinement of Examples 1 and 2. Before executing the "Edit Picture" step, a step is added to determine whether the editing instructions entered by the user are reasonably compatible with the editing area selected by the user (see S41b in Figure 25). When the editing area selected by the user and the entered editing instructions do not match, not only will the user be prompted that the editing instructions cannot be executed, but the user can also be recommended to select a new editing area or helped to modify the editing instructions. Example 3 solves the problem that the editing effect caused by the user's arbitrary input of editing instructions does not conform to common sense logic, thereby reducing the number of trial and error times when users edit pictures.
[0257] Specifically, as shown in FIG26 , the specific implementation of determining whether the first editing instruction and the editing area are properly matched may include the following process:
[0258] S50. Calculate and obtain the feature vector of the editing parameters and editing type in the first editing instruction.
[0259] The edit types in the first edit instruction can be finite and enumerable. Different edit types can be mapped to different identification numbers (IDs) and can therefore be represented by corresponding IDs. The computing module can use a model such as a deep learning algorithm to convert the ID corresponding to the edit type to obtain a feature vector corresponding to the edit type. For example, the ID corresponding to "increase" is "001", which is converted into the following feature vector: [0.5250, 0.7937, 0.1356, 1.4893, -3.9651, 1.5068].
[0260] The editing parameters in the first editing instruction (such as "quiet lake") can also be mapped to IDs first and then converted into feature vectors. The difference is that the text content of the editing parameters is unpredictable and not easy to be directly mapped to IDs. In this embodiment, the calculation module can perform word segmentation on the phrases representing the editing parameters, and then find the IDs corresponding to each word segmentation result from the preset word list. For example, the editing parameter "quiet lake" can be word segmented as: ["quiet", "quiet", "of", "lake", "park"], and the ID array corresponding to the word segmentation result is: [40496,3152,2099,8024,3563,8024,40497,0,0,0]. Among them, 40496 and 40497 respectively represent the start and end of the description phrase of the editing parameter, and 0 represents supplement. The length of the ID array, that is, the number of elements it contains, can be preset to constrain the maximum length of the description phrase of the editing parameter. Then, the calculation module can generate a corresponding feature vector for each ID in the ID array through a model such as a deep learning algorithm. For example, the ID array in the previous example can be converted into the following array of feature vectors: [[0.8838,0.1570,0.5249,...,0.4278,0.1725,0.4225],[1.8143,-0.5514,0.0995,...,-4.7141,-1.3811,-1.1166],[0.0186,3.5949,1.1780,...,1.1433,2.7235,-0.5069],...,[1.1674,0.9497,1.8264,...,1.3671,0 .5551,-0.4302]], among which, [0.8838,0.1570,0.5249,...,0.4278,0.1725,0.4225] represents the eigenvector of 40496, [1.8143,-0.5514,0.0995,...,-4.7141,-1.3811,-1.1166] represents the eigenvector of 3152, [0.0186,3.5949,1.1780,...,1.1433,2.7235,-0.5069] represents the eigenvector of 2099, ..., [1.1674,0.9497,1.8264,...,1.3671,0.5551,-0.4302] represents the eigenvector of 0. The feature vector is a two-dimensional array, in which each element is a feature vector corresponding to an ID.
[0261] The deep learning algorithm model may be, for example, a Word2Vec model, a Transformers model, etc.
[0262] S51. Extract various features such as mask features, depth features, contour features, color features, etc. of the editing area selected by the user, and fuse these features using weighted summation and other techniques to obtain a fused feature vector.
[0263] The fused feature vector corresponding to the edited area can be expressed as a two-dimensional array. For example, the fused feature vector of the "sky" area is: [[0.2611,1.8726,...,-0.9721],...,[1.6888,2.6287,...,-5.5910]].
[0264] The feature vector corresponding to a feature (such as a depth feature) can be calculated by a deep learning algorithm model, and the deep learning algorithm model can be, for example, a convolutional neural network model (CNN). The fusion of multiple features can be achieved through multiple CNN models and weighted summation. For example, the mask feature shown in Figure 13 is obtained by CNN model 1, the depth feature shown in Figure 14 is obtained by CNN model 2, and the contour feature shown in Figure 15 is obtained by CNN model 3, etc. The dimensions of the feature vectors of these features are the same length, so these feature vectors can be obtained by weighted summation to obtain a fused feature vector.
[0265] S52. Use a model such as a machine learning or deep learning algorithm to compare the feature vector corresponding to the first editing instruction with the fused feature vector of the editing area selected by the user to determine whether it is reasonable to apply the first editing instruction to the editing area selected by the user.
[0266] Specifically, the calculation module can use the feature vector corresponding to the editing type in the first editing instruction as one of the inputs of a model such as a machine learning or deep learning algorithm, use the feature vector corresponding to the editing parameter in the first editing instruction as the second input of a model such as a machine learning or deep learning algorithm, and use the fusion feature vector of the editing area as the third input of a model such as a machine learning or deep learning algorithm. Finally, the calculation module can obtain a judgment result on whether it is reasonable to apply the first editing instruction to the editing area selected by the user through calculations by a model such as a machine learning or deep learning algorithm.
[0267] When it is determined that the first editing instruction is not reasonable to be applied to the editing area selected by the user, the human-computer interaction module may output an error prompt, which may be a visual prompt displayed on the screen, a tactile vibration prompt, a voice prompt, etc.
[0268] For example, as shown in FIG27 , when the editing area selected by the user is “sky” and the editing instruction input is “add ‘quiet lake’”, the human-computer interaction module may display an error prompt 311 to remind the user that adding “quiet lake” to “sky” is not in line with common sense and may not allow the execution of the editing instruction.
[0269] For another example, for picture 71 shown in Figure 7, when the editing area selected by the user is "sky" and the editing instruction entered is "delete 'sun'", the calculation module can determine that the editing instruction is applicable to the "sky", and if the editing instruction changes to "delete 'mountain'", the calculation module can determine that the new editing instruction is not applicable to the "sky".
[0270] For another example, for picture 71 shown in Figure 7, when the editing area selected by the user is "sun" and the editing instruction entered is "replace with 'cloud'", the calculation module can determine that the editing instruction is applicable to the editing area, and if the editing instruction changes to "replace with 'tree'", the calculation module can determine that the new editing instruction is not applicable to the editing area.
[0271] For another example, for picture 71 shown in Figure 7, when the editing area selected by the user is "tree" and the editing instruction entered is "drag to 'forest'", the calculation module can determine that the editing instruction is applicable to the editing area. If the editing instruction becomes "drag to 'lake'", the calculation module can determine that the new editing instruction is not applicable to the editing area.
[0272] The above examples are only used to explain the embodiments of the present application and should not be construed as limiting.
[0273] The error prompt 311 shown in Figure 27 is only an example. In actual applications, the error prompt can also be other styles, such as flashing erroneous editing instructions or flashing editing areas. The embodiments of the present application do not limit this.
[0274] S53. When it is determined that the first editing instruction is not reasonable to be applied to the editing area selected by the user, a new editing instruction and editing area are recommended.
[0275] Specifically, other areas in the first image can be recommended to the user as editing areas based on the edit type and feature vectors corresponding to the edit parameters in the first editing instruction. The calculation module can traverse other areas in the first image and compare the fused feature vectors of the other areas with the feature vectors corresponding to the first editing instruction using the method in S52 to find areas that can reasonably match the first editing instruction, and then recommend the found areas as editing areas. Here, other areas refer to areas in the first image outside the editing area selected by the user.
[0276] Specifically, the first editing instruction can be modified according to the editing area selected by the user, such as modifying the editing type and / or editing parameters, so that the modified editing instruction adapts to the editing area selected by the user. The other area refers to the area outside the editing area selected by the user in the first picture. The calculation module can traverse the editing types and / or editing parameters in the recommendation pool, and use the method in S52 to compare the feature vectors corresponding to the traversed editing types and editing parameters with the fused feature vector of the editing area selected by the user to find the editing type and / or editing parameters that can reasonably match the area selected by the user, and then modify the editing instruction, and finally recommend the modified editing instruction.
[0277] For example, as shown in Figure 28, when the editing area selected by the user is "sky" and the editing instruction entered is "add 'quiet lake'", the calculation module can determine that: the editing instruction is not applicable to the "sky" area 312, and modify the editing area, using the "forest" area 313 as the new editing area, and finally recommend the user to apply the editing instruction to the "forest" area 313.
[0278] For another example, as shown in FIG29 , when the editing area selected by the user is “sky” and the editing instruction input is “add ‘quiet lake’”, the calculation module can determine that the editing instruction is not applicable to the “sky” and modify the editing instruction to “add ‘soaring geese’” as a new editing instruction, and finally recommend the user to apply the new editing instruction to the area “sky”.
[0279] Example 4
[0280] Example 4 is a supplementary refinement of the "edit picture" step in Examples 1 and 2. It introduces how to change the perspective relationship of objects in the original image (i.e., the front-to-back relationship, the depth of field relationship) so that when new content is added to the original image, the front objects in the original image are not blocked, thereby ensuring that the perspective relationship between objects in the edited image is reasonable.
[0281] In an embodiment of the present application, the editing instructions entered by the user into the editing area involve adding a new object. The editing instructions involving adding a new object may include: adding, replacing, dragging, and other types of editing instructions. Replacing is equivalent to deleting an object from the original image and then adding another object; dragging is equivalent to deleting an object from a certain position in the original image and then adding it to another position. In other words, the editing instruction involving adding a new object may mean that the editing process corresponding to the editing instruction includes adding a new object to the editing area.
[0282] In this embodiment, the "editing an image" step may specifically include: in response to a user inputting a first editing instruction for an editing area, the human-computer interaction module may send the first editing instruction to the computing module. Here, the first editing instruction involves adding a new object to the editing area, and its editing type may be, for example, add, replace, drag, etc. The computing module may then perform corresponding image editing processing on the first image according to the first editing instruction and return the edited first image to the human-computer interaction module. In this way, the human-computer interaction module can display the edited first image.
[0283] As shown in FIG30 , the specific implementation of the image editing performed by the computing module may be as follows (S61-S62):
[0284] S61. Extract various image features of the new object image, such as depth features, mask features, contour features, color features, etc., extract various image features of the first image, and correct the various features of the first image in combination with the image features of the editing area and the editing parameters in the editing instructions.
[0285] Here, the term "new object" refers to the original object in the editing area before the new object is added. The new object replaces the original object in the editing area. For both add and replace edit types, the human-computer interaction module provides one or more images of the new object indicated by the edit parameters, allowing the user to select an image for the new object. For the drag edit type, the new object is essentially the object being dragged by the user, and its image is from the first image, not from outside the first image.
[0286] The following describes the modifications of various features.
[0287] Correction of deep features
[0288] Step 1: Determine the perspective relationship, i.e., the front-to-back relationship, between the new object and the original object in the first image based on the image features of the first image and the editing parameters in the first editing instruction. This is reflected in the data as different depth values.
[0289] Directly replacing the original object's image with the new object's image will result in an illogical perspective relationship between the objects, preventing the new object from properly integrating into the first image. This can cause the original object within the edited area to be completely obscured by the new object. For example, as shown in Figure 31 , if the new object "Quiet Lake" shown in Figure 32 were directly used to replace the original objects "Tree" and "Valley" within edited area 317 in image 315 , the "Quiet Lake" would appear in front of the original object "Tree," creating an illogical perspective relationship.
[0290] The embodiments of the present application may use artificial intelligence algorithms such as image semantic understanding to determine the perspective relationship between the new object and the original object, and based on this, correct the depth features of the image to present a reasonable perspective relationship between objects.
[0291] For example, if the first editing instruction for the editing area 317 of the picture 315 in Figure 31 is "add 'quiet lake'", and the image of the new object "quiet lake" is shown in Figure 32; then the calculation module can determine based on the masking features of the picture 315: the original objects in the editing area 317 are "tree" and "valley", and the new object "quiet lake" will block the original objects "tree" and "valley". Then, the calculation module can determine the front and back relationship between the new object "quiet lake" and the original objects "tree" and "valley" based on artificial intelligence algorithms such as image semantic understanding, such as: "tree" is in front of "quiet lake", and "quiet lake" is in front of "valley".
[0292] Step 2. Determine the reference depth of the new object based on the perspective relationship between the new object and the original object, as well as the depth value of the original object. Then use the reference depth to correct the depth feature of the new object.
[0293] First, interpolation or other methods can be used to determine the baseline depth of the new object. In the above example, assuming the average depths of the original objects "tree" and "valley" are 0.6 and 1.0, respectively, then, based on the foreground-and-background relationship between the new object "quiet lake" and the original objects "tree" and "valley," interpolation algorithms can be used to determine that the baseline depth of "quiet lake" is 0.8. 0.8 is greater than 0.6 but less than 1.0. Here, the further back an object is, the greater its depth. This ensures that the perspective relationship of the new object within the overall image of the first image is reasonable.
[0294] Secondly, the depth feature of the new object is corrected using the baseline depth of the new object and the depth difference between different regions on the new object.
[0295] After the new object's depth features are corrected, its average depth can approach or equal the baseline depth, while preserving the depth differences between its regions. "Approximate" can mean that the difference between the average depth and the baseline depth does not exceed a specific depth value, such as 0.05. This ensures that the new object as a whole forms a reasonable perspective relationship with the original object, while also preserving the perspective relationships between individual elements within the new object.
[0296] Step 3: Replace the original depth features of the edited area with the corrected depth features of the new object to correct the depth features of the first image.
[0297] Here, the essence of the replacement can be: traversing each pixel area within the editing area, if the new object is in front of the original object in the pixel area, the original depth data of the pixel area is replaced with the depth data of the new object in the pixel area to achieve the perspective relationship of the new object in front of the original object; if the original object in the pixel area is in front of the new object, the original depth data of the pixel area is retained to achieve the perspective relationship of the original object in front of the new object. The original depth data of a pixel area refers to the depth data of the pixel area on the original object.
[0298] The editing area can be divided into a first pixel area and a second pixel area, where the perspective relationship of the first pixel area is that the new object is in front of the original object, and the perspective relationship of the second pixel area is that the original object is in front of the new object. The depth characteristics of the first pixel area can be replaced with the depth characteristics of the new object, while the depth characteristics of the second pixel area can retain the depth characteristics of the original object. In actual applications, the new object may be entirely in front of the original object, excluding the second pixel area. In this case, the depth characteristics of the entire editing area can be directly replaced with the depth characteristics of the new object.
[0299] Therefore, in actual applications, only a portion of the pixel areas in the editing area may have their depth data replaced. The perspective relationship of these pixel areas is: the new object is in front of the original object; while other pixel areas may retain their original depth data. The perspective relationship of these other pixel areas is: the original object is in front of the new object. For example, as shown in Figure 33, in pixel area 319, assuming that the average depth value of the original object "tree" is 0.6 and the baseline depth of the new object "quiet lake" is 0.8, it can be seen that in pixel area 319, the original object "tree" is in front of the new object "quiet lake". Therefore, pixel area 319 retains the depth data, which ensures that: in area 319, the "tree" in front will not be blocked by the "quiet lake" behind, forming a reasonable perspective relationship.
[0300] FIG34 exemplarily shows a reasonable perspective relationship formed between the new object “quiet lake” and the original objects “tree” and “valley” through correction of depth features.
[0301] Modification of mask features
[0302] After depth feature correction, mask feature correction may not directly use the new object mask data to replace the original mask data in the edited area. Instead, the new object mask data is used to replace the original mask data only in the area where the depth data has been replaced. This ensures that the correction of the mask feature is consistent with the correction of the depth feature, so that the corrected mask and depth features both point to the same object, avoiding any conflicts between the two and ensuring the semantics of the regenerated image are reasonable.
[0303] For example, in FIG33 , the depth data of pixel area 319 in the editing area 317 is not replaced, that is, the depth data of the original object "tree" is still used, so that the original object "tree" is in front of the new object "quiet lake" in terms of perspective. In response to this, when correcting the mask feature, in area 319, the mask data of the new object "quiet lake" (such as "8") is not used to replace the original mask data of area 319 (such as "5", i.e., the mask data of "tree"), but the mask data of the original object "tree" is retained; otherwise, a contradiction will result, that is, from the perspective of the depth feature, area 319 is the original object "tree", but from the perspective of the mask feature, area 319 is indeed the new object "quiet lake".
[0304] In other words, within the editing area, if the depth data for a pixel region belongs to the new object and is no longer the original depth data for that pixel region, the original mask data for that pixel region can be replaced with the mask data for that pixel region on the new object. The original mask data for a pixel region refers to the mask data for that pixel region on the original object. Therefore, in actual applications, the area where the mask feature is replaced may only be a portion of the pixel region within the editing area, and the depth features of this portion of the pixel region are replaced with the depth features of the new object.
[0305] In the example shown in FIG34 , the area where the mask feature replacement occurs is smaller than the area of the new object “quiet lake”. FIG35 briefly shows a comparison between the two.
[0306] Correction of contour features
[0307] Similar to mask feature correction, after depth feature correction, contour feature correction may not directly use the new object's contour data to replace the original contour data in the edited area. Instead, the new object's contour data will only be used to replace the original contour data in the area where the depth data has been replaced. This ensures that contour feature correction is consistent with depth feature correction, ensuring that the corrected contour and depth features both point to the same object, avoiding conflicts and ensuring the semantics of the regenerated image are reasonable.
[0308] In other words, within the editing area, if the depth data for a pixel region belongs to the depth data of the new object and is no longer the original depth data for the pixel region, the outline data of the pixel region on the new object can be used to replace the original outline data of the pixel region. The original outline data of a pixel region refers to the outline data of the pixel region on the original object.
[0309] Similarly, in the example shown in FIG34 , the area where the contour feature replacement occurs is smaller than the area of the new object “quiet lake”. FIG35 can also be used to illustrate the comparison between the two.
[0310] Not limited to mask features and contour features, the correction of other features (such as color features) is also considered. That is: within the editing area, if the depth data of a pixel area belongs to the depth data on the new object, and is no longer the original depth data of the pixel area, then the color features of the pixel area on the new object and other features can be used to replace the original color features and other features of the pixel area.
[0311] S62. Regenerate the image using the corrected depth features, mask features, contour features, etc.
[0312] Specifically, images can be regenerated using artificial intelligence algorithm models, such as by combining the Stable Diffusion model with the ControlNet model. Stable Diffusion is a diffusion model used to generate images from text or images. During the image generation process, the ControlNet model can be used to introduce additional conditions (such as depth features, mask features, and contour features) to intervene in the image generation process.
[0313] The regenerated image can be shown in the right picture of Figure 34, which realizes the reasonable insertion of new objects (such as "quiet lake") between the original objects in the first image (such as "tree" and "valley"), presenting a reasonable perspective relationship between objects.
[0314] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the steps performed by the human-computer interaction module in the above-mentioned method embodiments, or the steps performed by the human-computer interaction module and the computing module.
[0315] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the steps performed by the computing module in the above-mentioned various method embodiments.
[0316] An embodiment of the present application also provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps performed by the human-computer interaction module in the above-mentioned various method embodiments, or the steps performed by the human-computer interaction module and the computing module.
[0317] An embodiment of the present application also provides a computer program product. When the computer program product is run on a server, the server can implement the steps performed by the computing module in the above-mentioned various method embodiments.
[0318] The present application also provides a chip system, comprising a processor coupled to a memory, the processor executing a computer program stored in the memory to implement the steps performed by the human-computer interaction module, or the steps performed by the human-computer interaction module and the computing module, in any method embodiment of the present application. The chip system can be a single chip or a chip module composed of multiple chips.
[0319] The present application also provides a chip system, comprising a processor coupled to a memory, the processor executing a computer program stored in the memory to implement the steps performed by the computing module in any of the method embodiments of the present application. The chip system can be a single chip or a chip module composed of multiple chips.
[0320] The term "user interface (UI), or interface for short" in the specification and drawings of this application refers to the media interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The user interface of an application is a source code written in a specific computer language such as Java and extensible markup language (XML). The interface source code is parsed and rendered on the terminal device, and finally presented as content that the user can recognize, such as pictures, text, buttons and other controls. Controls, also known as widgets, are the basic elements of the user interface. Typical controls include toolbars, menu bars, text boxes, buttons, scroll bars, pictures and text. The properties and contents of controls in the interface are defined by tags or nodes, such as XML through <textview> 、 <imgview> 、 <videoview>The controls contained in the interface are specified by nodes such as <head> and <body>. A node corresponds to a control or attribute in the interface, and the node is presented as user-visible content after parsing and rendering. In addition, many applications, such as hybrid applications, usually also contain web pages in their interfaces. A web page, also known as a page, can be understood as a special control embedded in the application interface. A web page is a source code written in a specific computer language, such as hypertext markup language (HTML), cascading style sheets (CSS), JavaScript (JS), etc. The web page source code can be loaded and displayed as user-recognizable content by a browser or a web page display component with similar functions to a browser. The specific content contained in a web page is also defined by tags or nodes in the web page source code, such as HTML through <body>. 、 、 <video> 、 <canvas>To define the elements and attributes of a web page.
[0321] A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operations that uses graphics. It can be an icon, window, control, or other interface element displayed on the display of an electronic device. Controls can include icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, widgets, and other visual interface elements.
[0322] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk).
[0323] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0324] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.< / canvas> < / video> < / videoview> < / imgview> < / textview>
Claims
1. A method for editing a picture, characterized in that: include: Display the first picture; detecting a user operation on the first picture for selecting an editing area; Determining the editing area from the first picture according to the operation position of the user operation on the first picture, and distinguishably displaying the editing area in the first picture; generating a recommended editing instruction according to the image features of the editing area; displaying the recommended editing instructions; detecting that a user inputs a first editing instruction for the editing area, the first editing instruction comprising: a recommended editing instruction; Performing image editing processing on the first picture according to the first editing instruction; The first picture after the image editing process is displayed.
2. The method according to claim 1, characterized in that The generating of the recommended editing instruction according to the image features of the editing area specifically includes: Using a fusion feature vector of multiple image features of the editing area as one of the inputs of the first artificial intelligence algorithm; the multiple image features include multiple features of mask features, depth features, contour features, and color features; Using one or more preset editing types as second input to the first artificial intelligence algorithm; The recommended editing instruction is obtained by the operation of the first artificial intelligence algorithm, and the recommended editing instruction includes editing parameters corresponding to the preset editing type.
3. The method according to claim 1 or 2, characterized in that The preset editing type includes one or more of the following: delete, drag, replace, add, or color adjustment.
4. The method according to any one of claims 1 to 3, characterized in that The detecting that the user inputs the first editing instruction for the editing area specifically includes: detecting that the user selects to input the recommended editing instruction; and the recommended editing instruction selected by the user is determined as the first editing instruction.
5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: displaying one or more preset editing instructions; the first editing instruction further comprises the preset editing instruction; The detecting that the user inputs the first editing instruction for the editing area specifically includes: detecting that the user selects to input the preset editing instruction; and the preset editing instruction selected by the user is determined as the first editing instruction.
6. The method according to any one of claims 1 to 5, characterized in that The method further includes: displaying a first input box, the first input box being used to receive a voice or text editing instruction; the first editing instruction also includes a voice or text editing instruction input through the input box; The detecting that the user inputs the first editing instruction for the editing area specifically includes: detecting a voice or text instruction input by the user in the first input box; and determining the voice or text instruction as the first editing instruction.
7. The method according to any one of claims 1 to 6, characterized in that The user operation for selecting the editing area includes: a user operation of selecting a first object in the first picture, the operation position of the user operation on the first picture falls on the first object in the first picture, and the image area where the first object is located is the editing area; the image area where the first object is located is determined by performing image segmentation processing on the first picture.
8. The method according to any one of claims 1 to 7, characterized in that The user operation for selecting the editing area includes: a user operation of drawing the editing area in the first picture.
9. The method according to any one of claims 1 to 8, characterized in that Also includes: Preprocessing information of the first picture is obtained, and the editing area is determined from the area indicated by the preprocessing information.
10. The method according to claim 9, characterized in that The preprocessing information includes indication information of multiple regions, wherein the preprocessing information includes coordinates of contour points of each of the multiple regions, a binary image, and a grayscale image, the values of the multiple regions in the binary image are first values, and the grayscale values of the multiple regions in the grayscale image are first grayscale values or first grayscale ranges; Determining the editing area from the area indicated by the preprocessing information specifically includes: determining the area where the operation position of the user operation is located among the multiple areas as the editing area, or determining the area among the multiple areas that is closest to the operation position as the editing area.
11. The method according to claim 9, characterized in that The preprocessing information includes layer information of multiple layers, and the layer information of each layer includes the coordinates of opaque pixels in the layer; the method also includes: before determining the editing area from the area indicated by the preprocessing information, each of the opaque pixels connected in the layer is determined as an area indicated by the preprocessing information.
12. The method according to claim 9 or 11, characterized in that The preprocessing information includes contour information; the method further comprises: before determining the editing area from the area indicated by the preprocessing information, determining the area surrounded by the contour indicated by the contour information as the area indicated by the preprocessing information.
13. The method according to claim 9, 11 or 12, characterized in that: The preprocessing information includes depth information; the method further includes: before determining the editing area from the area indicated by the preprocessing information, according to the depth information, determining pixels with the same or similar depth values as an area indicated by the preprocessing information.
14. The method according to any one of claims 1 to 13, characterized in that The distinguishing display includes one or more of the following methods: highlighting the outline of the editing area, highlighting the entire editing area, or displaying a dotted frame along the outline of the editing area.
15. The method according to any one of claims 1 to 14, characterized in that Also includes: Before performing image editing processing on the first picture according to the first editing instruction, if it is determined that the application of the first editing instruction to the editing area is unreasonable, then a new editing instruction or a new editing area is recommended.
16. The method according to claim 15, characterized in that Also includes: Whether the application of the first editing instruction to the editing area is reasonable is determined by comparing the feature vector corresponding to the first editing instruction with the fused feature vector of each image feature of the editing area.
17. The method according to claim 15 or 16, characterized in that The re-recommending the editing area specifically includes: traversing the area outside the editing area in the first picture, comparing the fused feature vector of the traversed area with the feature vector corresponding to the first editing instruction, finding an area that can reasonably match the first editing instruction, and re-recommending the found area as the editing area.
18. The method according to any one of claims 15 to 17, characterized in that The re-recommending editing instructions specifically includes: traversing the editing types and / or editing parameters in the recommendation pool, comparing the feature vectors corresponding to the traversed editing types and / or editing parameters with the fused feature vector of the editing area, finding the editing types and / or editing parameters that can reasonably match the editing area, and modifying the first editing instruction according to the found editing type and / or editing parameters, and re-recommending the modified first editing instruction.
19. The method according to any one of claims 1 to 18, characterized in that The editing process corresponding to the first editing instruction includes: adding a new object to the editing area; in the first picture after the image editing process, in the editing area, the depth features of the first pixel area are replaced with the depth features of the new object, and the depth features of the second pixel area remain the depth features of the original object, wherein the perspective relationship of the first pixel area is that the new object is in front of the original object, and the perspective relationship of the second pixel area is that the original object is in front of the new object.
20. The method of claim 19, wherein: The image editing processing of the first picture according to the first editing instruction includes: correcting the image features of the first picture, and regenerating the first picture using the corrected image features of the first picture.
21. The method of claim 20, wherein: The image feature includes a depth feature; and the correcting the image feature of the first picture specifically includes: determining a perspective relationship between the new object and the original object according to image features of the first picture and editing parameters in the first editing instruction; Determine a reference depth of the new object according to the perspective relationship between the new object and the original object, and the depth value of the original object, and then use the reference depth to correct the depth feature of the new object; Replacing the original depth feature of the edited area with the corrected depth feature of the new object; After the depth feature of the new object is corrected, the average depth of the new object is close to or equal to the reference depth, and the depth difference between various regions on the new object remains unchanged.
22. The method according to claim 20 or 21, characterized in that The image feature further includes a first image feature, which is an image feature other than a depth feature and includes one or more of the following: a mask feature, a contour feature, and a color feature; The correcting the image feature of the first picture further includes: In the editing area, if the depth feature of a pixel area is replaced with the corrected depth feature of the new object, the original first image feature of the pixel area in the editing area is replaced with the first image feature of the pixel area on the new object.
23. A terminal device, characterized in that: include: A human-computer interaction module, a processor and a memory, wherein the human-computer interaction module is coupled to the processor, and the memory is coupled to the processor; the human-computer interaction module includes a touch screen; The memory is used to store computer program codes, and the computer program codes include computer instructions. When the processor executes the computer instructions, the terminal device executes the method as described in any one of claims 1 to 22.
24. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on the terminal device, the terminal device executes the method according to any one of claims 1 to 22.
25. A picture editing method, the method being applied to a human-computer interaction module, characterized in that: The human-computer interaction module is included in the picture editing system, and the picture editing system further includes: a calculation module; The method comprises: The human-computer interaction module displays a first picture; The human-computer interaction module detects a user operation for selecting an editing area on the first picture; The human-computer interaction module distinguishably displays the editing area in the first picture; The human-computer interaction module receives the recommended editing instruction sent by the calculation module and displays the recommended editing instruction, where the recommended editing instruction is generated by the calculation module according to the image features of the editing area; The human-computer interaction module detects that a user inputs a first editing instruction for the editing area, the first editing instruction comprising: a recommended editing instruction; The human-computer interaction module receives the first image after image editing processing sent by the computing module, and displays the first image after image editing processing, wherein the image editing processing is performed by the computing module.
26. The method of claim 25, wherein: The method further comprises: the human-computer interaction module displays one or more preset editing instructions; The detecting that the user inputs the first editing instruction for the editing area specifically includes: detecting that the user selects to input the preset editing instruction; and the preset editing instruction selected by the user is determined as the first editing instruction.
27. The method according to any one of claims 25 to 26, characterized in that The method further comprises: the human-computer interaction module displays a first input box, the first input box being used to receive a voice or text editing instruction; The detecting that the user inputs the first editing instruction for the editing area specifically includes: detecting a voice or text instruction input by the user in the first input box; and determining the voice or text instruction as the first editing instruction.
28. The method according to any one of claims 25 to 27, characterized in that The user operation for selecting the editing area includes: a user operation of selecting a first object in the first picture, the operation position of the user operation on the first picture falls on the first object in the first picture, and the image area where the first object is located is the editing area; the image area where the first object is located is determined by performing image segmentation processing on the first picture.
29. The method according to any one of claims 25 to 28, characterized in that The user operation for selecting the editing area includes: a user operation of drawing the editing area in the first picture.
30. The method according to any one of claims 25 to 29, characterized in that The distinguishing display includes one or more of the following methods: highlighting the outline of the editing area, highlighting the entire editing area, or displaying a dotted frame along the outline of the editing area.
31. The method according to any one of claims 25 to 30, characterized in that After the human-computer interaction module detects that the user inputs a first editing instruction for the editing area, the method further includes: if it is unreasonable to apply the first editing instruction to the editing area, then re-recommend the editing area; the re-recommended editing area is found by the calculation module from an area outside the editing area based on a feature vector corresponding to the first editing instruction.
32. The method according to any one of claims 25 to 31, characterized in that After the human-computer interaction module detects that the user inputs a first editing instruction for the editing area, the method further includes: if it is unreasonable to apply the first editing instruction to the editing area, recommending a modified first editing instruction; the editing type and / or editing parameters of the modified first editing instruction are found by the calculation module from the recommendation pool based on the fused feature vector of the editing area.
33. A terminal device, characterized in that: include: A human-computer interaction module, a processor and a memory, wherein the human-computer interaction module is coupled to the processor, and the memory is coupled to the processor; the human-computer interaction module includes a touch screen; The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the terminal device executes the method as described in any one of claims 25-32.
34. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on the terminal device, the terminal device executes the method according to any one of claims 25 to 32.
Citation Information
Patent Citations
Depth image processing method and device, equipment and storage medium
CN112801907A
Image processing method, model training method and related device
CN114943789A
Image processing method and device, electronic equipment and readable storage medium
CN116883307A
Image processing model training method and device, electronic equipment and storage medium
CN116958325A
Image display device
JP2005234912A
Cited By
Video coloring model construction method, video coloring method, equipment and medium
CN122023560A
A video coloring model construction method, a video coloring method, a device, and a medium
CN122023560B