Method and device for modifying picture through computing power of intelligent computing center

Through the computing power of the intelligent computing center, combined with semantic understanding, reverse coding and multimodal large model inference, the problem of low accuracy and efficiency of picture modification in the existing technology is solved, and high accuracy and high efficiency of picture modification is achieved, improving the user experience.

CN119991881APending Publication Date: 2025-05-13DATACANVAS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510199231.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing image modification method based on large language models relies on single-modal prompt words, which leads to the user's real intention understanding easily missing, the accuracy and efficiency of modifying pictures, and the user experience is poor.

Method used

Through the computing power of the intelligent computing center, you can receive the user's pictures to be modified, refer to the attached drawings and the text of the modification intention to describe it, conduct semantic understanding, reverse coding and multimodal large model reasoning, and accurately modify the pictures.

Benefits of technology

It realizes the high accuracy of pictures based on user modification intentions, improves the efficiency of picture modification and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991881A_ABST
    Figure CN119991881A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for modifying a picture through computing power of an intelligent computing center. The modification method comprises the following steps: S1, receiving a to-be-modified picture, a reference attached drawing and a modification intention description text; s2, the computing power of the intelligent computing center is called to execute the step S21, the step S22 and the step S23 in parallel, and the step S21 comprises the steps that semantic understanding is conducted on the modification intention description text to obtain a modification target subject; determining a target area; the step S22 comprises the following steps: respectively carrying out reverse coding on the modification intention description text and the reference drawing, and fusing the coding to obtain a target prompt code; the step S23 comprises the following steps: determining a pre-trained multi-modal large model; and S3, modifying the to-be-modified picture according to the target region and the target prompt code by adopting a multi-modal large model. According to the method, high-accuracy modification can be performed according to the modification intention of the user, the execution efficiency is improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent computing centers, smart computing centers and computing power infrastructure, and in particular to a method and device for modifying images through the computing power of an intelligent computing center. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have come into being.

[0003] "Intelligent computing center" refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning) by using large-scale heterogeneous computing resources, including general computing power and intelligent computing power. The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.

[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to execute certain computing needs. It is the computing power to process information data and achieve target result output. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] Existing image modification methods based on large language models are often performed according to the prompt words of a single modality input by the user, for example, text that only describes the intention of modification. Under the prompt words of a single modality, the large language model is prone to miss the user's true intention, resulting in the modification of the image deviating from the user's modification intention and low accuracy of the modified image. In addition, the multimodal large model needs to perform high-intensity reasoning during the image modification process, which takes a lot of time, and the image modification accuracy and efficiency are low, resulting in a poor user experience. Summary of the invention

[0008] The present invention provides a method and device for modifying images through the computing power of an intelligent computing center, so as to solve the existing image modification method based on a large language model, which is often performed according to a single-modal prompt word input by the user, for example, a text that only describes the intention of modification. Under a single-modal prompt word, the large language model is prone to omissions in understanding the user's true intention, resulting in the deviation of the modified image from the user's modification intention and low accuracy of the modified image. In addition, the multimodal large model needs to perform high-intensity reasoning in the process of modifying the image, which takes a lot of time, and the image modification accuracy and efficiency are low, resulting in a poor user experience.

[0009] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:

[0010] In a first aspect, the present invention provides a method for modifying an image by using the computing power of an intelligent computing center, comprising:

[0011] Step S1: receiving a picture to be modified, a reference drawing, and a modification intention description text sent by a user through an interactive terminal;

[0012] Step S2: calling the computing power of the intelligent computing center to execute step S21, step S22 and step S23 in parallel, wherein step S21 includes: semantically understanding the modification intention description text to obtain the modification target subject that the user needs to modify; determining the area where the modification target subject is located on the image to be modified as the target area; step S22 includes: reverse encoding the modification intention description text and the reference figure respectively, and fusing the encodings of the two to obtain the target prompt encoding; step S23 includes: determining a pre-trained multimodal large model for modifying the image to be modified;

[0013] Step S3: using the multimodal large model, modifying the image to be modified according to the target area and the target prompt code.

[0014] Optionally, step S3 includes:

[0015] Step S31: configuring the multimodal large model with the first computing resource of the intelligent computing center required to modify the image to be modified;

[0016] Among them, the multimodal large model infers the target area and the target prompt code based on the first computing power resource to modify the image to be modified.

[0017] Optionally, the step S21 includes:

[0018] Step S211: using a computer vision (CV) method to identify the modification target subject in the modification area from the image to be modified, and determining the target area.

[0019] Optionally, the target area includes: a core area and a transition area;

[0020] The step S211 then includes:

[0021] Step S212: using the CV method to identify the edges of each modified target body from the target area, drawing a minimum convex set along the edge of the modified target body, and taking the area enclosed by the minimum convex set as the core area;

[0022] Step S213: determining other areas in the target area except the core area as the transition area;

[0023] The step S3 comprises:

[0024] Step S32: the multimodal large model infers the core area and the target prompt code to generate a modified replacement graphic of the modified target body; the modified replacement graphic is used to replace the modified target body in the core area; after replacing all the modified target bodies, the multimodal large model infers the transition area and the target prompt code to modify the light and shadow effects of the transition area.

[0025] Optionally, a virtual Kubernetes service deployed in the intelligent computing center is called to execute step S2, wherein required service containers are configured respectively for executing step S21, step S22, and step S23.

[0026] Optionally, step S23 includes:

[0027] Step S231: sending a model list to an interactive terminal associated with the user, the model list including: names and descriptions of all models in a preset picture modification multimodal large model set;

[0028] Step S232: receiving a model determination instruction sent by the interaction terminal, and in response to the model determination instruction, determining that the model indicated by the model determination instruction is the multimodal large model.

[0029] In a second aspect, the present invention provides a device for modifying an image by using the computing power of an intelligent computing center, comprising:

[0030] A receiving module, used to receive the picture to be modified, the reference drawing and the modification intention description text sent by the user through the interactive terminal;

[0031] An execution module is used to call the computing power of the intelligent computing center to perform semantic understanding of the modification intention description text in parallel to obtain the modification target subject that the user needs to modify; determine the area where the modification target subject is located on the image to be modified as the target area; perform reverse encoding of the modification intention description text and the reference figure in parallel, and fuse the encodings of the two to obtain the target prompt encoding; and perform the step of determining a pre-trained multimodal large model for modifying the image to be modified in parallel;

[0032] A modification module is used to use the multimodal large model to modify the image to be modified according to the target area and the target prompt code.

[0033] In a third aspect, the present invention provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps in the method for modifying an image by using the computing power of an intelligent computing center as described in any one of the first aspects.

[0034] In a fourth aspect, the present invention provides a readable storage medium storing a program or instruction, which, when executed by a processor, implements the steps of the method for modifying an image by using the computing power of an intelligent computing center as described in any one of the first aspects.

[0035] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps of the method for modifying an image by using the computing power of an intelligent computing center as described in any one of the first aspects.

[0036] In the present invention, step S1 is used to receive a picture to be modified, a reference drawing and a modification intention description text sent by a user through an interactive terminal; step S2 is used to call the computing power of an intelligent computing center to execute steps S21, S22 and S23 in parallel, wherein step S21 includes: semantically understanding the modification intention description text to obtain a modification target subject that the user needs to modify; determining that the area where the modification target subject is located on the picture to be modified is the target area; step S22 includes: reversely encoding the modification intention description text and the reference drawing respectively, and fusing the encodings of the two. The target prompt code is obtained by using the code; the step S23 includes: determining a pre-trained multimodal large model for modifying the image to be modified; step S3: using the multimodal large model to modify the image to be modified according to the target area and the target prompt code. The present invention can modify the image with high accuracy according to the user's modification intention, and, with the abundant computing power of the intelligent computing center and executing steps S21, S22 and S23 in parallel, the present invention improves the execution efficiency, realizes the accuracy of image modification, improves the efficiency of image modification, and improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0038] Figure 1 It is a flow chart of the method for modifying a picture by using the computing power of an intelligent computing center according to the present invention;

[0039] Figure 2 A schematic diagram of the picture to be modified input by the user;

[0040] Figure 3 A schematic diagram of the reference drawing input by the user;

[0041] Figure 4 A schematic diagram to determine the core area;

[0042] Figure 5 This is a schematic diagram of the image modification results;

[0043] Figure 6 It is a principle block diagram of the device for modifying pictures through the computing power of the intelligent computing center of the present invention;

[0044] Figure 7 It is a principle block diagram of the electronic device of the present invention. DETAILED DESCRIPTION

[0045] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] The terms "first", "second", etc. of the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of one type, and the number of objects is not limited, for example, the first object can be one or more. In addition, "or" in the present invention represents at least one of the connected objects. For example, "A or B" covers three schemes, namely, Scheme 1: including A but not including B; Scheme 2: including B but not including A; Scheme 3: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.

[0047] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0048] It should be noted that the collection, collection, updating, analysis, processing, use, transmission, storage and other aspects of personal information involved in the technical solution of the present invention are in compliance with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken for personal information to prevent illegal access to personal information data and maintain personal information security and network security.

[0049] First, the technical terms involved in the present invention are briefly explained below.

[0050] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data center to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and it mainly provides services to the society through computing power infrastructure.

[0051] The "computing power" (Computational Power, CP) mentioned in the present invention refers to: it is the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP 通用 +CP 智能 +CP 超级 .

[0052] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capacity of the computing power facilities, which includes the comprehensive capabilities of network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.

[0053] The "Storage Power" (SP) described in the present invention refers to: the comprehensive capabilities of a data center in terms of data storage capacity, performance, safety and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server built-in storage devices. The commonly used unit of measurement for storage capacity is exabyte (EB, 1EB = 2^60bytes), and the commonly used unit of measurement for performance is the number of reads and writes per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of safety and reliability.

[0054] The "computing power infrastructure" mentioned in the present invention refers to: a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, which can realize centralized computing, storage, transmission and application of information, and presents characteristics such as diversity and ubiquity, intelligence and agility, safety and reliability, green and low-carbon. It is of great significance to promote industrial transformation and upgrading, enable my country's scientific and technological innovation, meet people's better life and realize high-efficiency social governance.

[0055] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G network, fiber broadband network, backbone network, international communication network, satellite Internet, computing power infrastructure such as data center, general computing power center, intelligent computing center, supercomputing center, and new technology facilities such as artificial intelligence, blockchain, and quantum computing. With the emergence and promotion and application of new general technologies, the new information infrastructure will become more diverse.

[0056] The “computing power” mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.

[0057] The “general computing power” mentioned in the present invention refers to: the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0058] The "intelligent computing power" mentioned in the present invention refers to: a computing platform based on GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit) and other dedicated chips for various innovative artificial intelligence applications, such as natural language processing, machine vision, etc.

[0059] The "super computing power" mentioned in the present invention refers to: the computing power mainly provided by high-performance computing clusters such as supercomputers. It uses the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0060] The "intelligent computing center" mentioned in the present invention refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning) by using large-scale heterogeneous computing resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0061] The “intelligent computing center” mentioned in the present invention includes but is not limited to the “intelligent computing center”.

[0062] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0063] The "computing power center" mentioned in the present invention refers to a facility that is mainly composed of infrastructure such as wind, fire, water, and electricity, and IT hardware and software equipment, and has computing power, transportation power, and storage power, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0064] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide functions such as large-scale computing, storage and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.

[0065] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.

[0066] The “model” and “big model” mentioned in the present invention include but are not limited to “big language model” and “multimodal big model”.

[0067] The “large language model” mentioned in the present invention refers to a large language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0068] The “Multimodal Large Models” described in the present invention refers to models that combine multimodal information such as text, images, videos, and audio for training, including but not limited to multimodal large models.

[0069] The present invention provides a method for modifying images by using the computing power of an intelligent computing center, see Figure 1 As shown, Figure 1 The flowchart of the method for modifying an image by using the computing power of an intelligent computing center of the present invention includes:

[0070] Step S1: receiving a picture to be modified, a reference drawing, and a modification intention description text sent by a user through an interactive terminal;

[0071] Step S2: calling the computing power of the intelligent computing center to execute step S21, step S22 and step S23 in parallel, wherein step S21 includes: semantically understanding the modification intention description text to obtain the modification target subject that the user needs to modify; determining the area where the modification target subject is located on the image to be modified as the target area; step S22 includes: reverse encoding the modification intention description text and the reference figure respectively, and fusing the encodings of the two to obtain the target prompt encoding; step S23 includes: determining a pre-trained multimodal large model for modifying the image to be modified;

[0072] Step S3: Using a multimodal large model, modify the image to be modified according to the target area and the target hint encoding.

[0073] In the present invention, the modification intention description text is a text description input by the user of what kind of modification the user wants to make to the modified picture. For example, the modification intention description input by the user is "Please change the red flowers in the picture to yellow". Correspondingly, the reference figure is a picture of yellow flowers provided by the user. The reference figure can supplement the modification requirements that are not mentioned or not clearly stated in the modification intention description text, for example: the variety of yellow flowers is not mentioned in the modification intention description input by the user; or the user mentioned that it is a yellow azalea, but the azalea itself has multiple sub-varieties, and the user's real intention is actually to hope that the red flowers in the picture are modified to a specific sub-variety of yellow azalea. In the case that there are modification requirements that are not mentioned or clearly stated in the modification intention description text, the image features of the yellow azalea in the reference figure can supplement the user's requirements for the sub-variety of azalea, ensuring that the modification requirements that meet the user's real needs are obtained. It can be understood that the multimodal large model can generate yellow azaleas for replacing red flowers according to the image features of yellow azaleas, or, according to the image features of yellow azaleas, it can determine the sub-variety of yellow azaleas and replace the red flowers in the image to be modified with yellow azaleas of the sub-variety.

[0074] In step S2 of the present invention, based on the characteristic that the intelligent computing center can handle large-scale parallel tasks, the computing power of the intelligent computing center is called to execute steps S21, S22 and S23 in parallel, which can improve execution efficiency, shorten the time required to modify the image, and improve user experience.

[0075] Combined with the example explanation, for step S21, the modification intention description text is semantically understood to obtain the modification target subject that the user needs to modify; the area where the modification target subject is located on the image to be modified is determined as the target area. For example, the modification intention description input by the user is "Please change the red flowers in the picture to yellow", then the modification target subject determined after semantic understanding is "red flowers", and the area where the "red flowers" are located on the image to be modified is further determined as the target area.

[0076] It should be noted that a large number of tests have shown that only using natural language prompt words to describe cannot completely restore the desired modification result in the image. Therefore, step S22 of the present invention reversely encodes the modification intention description text and the reference figure respectively, and fuses the codes of the two to obtain the target prompt code. By reverse encoding the reference figure, it is ensured that the large model can obtain the basic concept of the objects in the reference figure; and the modification intention description text can also accurately express the modification intention of the large model after being reverse encoded; further, the target prompt code obtained by fusing the codes of the two is input into the large model, so that the large model can accurately recognize the modification intention, which can ensure that the large model completes the task of modifying the picture with high quality.

[0077] Back-encoding is the process of the backpropagation algorithm. Back-propagation is a technique used to train neural networks by updating the model's weights by calculating the gradient of the loss function with respect to the parameters of the network. Back-encoding can refer to the process of reconstructing the input data from the latent space in the model. Specifically, back-encoding may involve the process of converting the latent vector back to the original data space.

[0078] In step S3 of the present invention, a multimodal large model is used to modify the image to be modified according to the target area and the target prompt code, that is, the target area and the target prompt code are input into the multimodal large model, the multimodal large model infers the target area and the target prompt code, and the target area on the image to be modified is modified according to the modification requirements represented by the target prompt word.

[0079] In the present invention, step S1 is used to receive a picture to be modified, a reference drawing and a modification intention description text sent by a user through an interactive terminal; step S2 is used to call the computing power of an intelligent computing center to execute steps S21, S22 and S23 in parallel, wherein step S21 includes: semantically understanding the modification intention description text to obtain a modification target subject that the user needs to modify; determining that the area where the modification target subject is located on the picture to be modified is a target area; step S22 includes: reversely encoding the modification intention description text and the reference drawing respectively, and fusing the encodings of the two to obtain a target prompt code; step S23 includes: determining a pre-trained multimodal large model for modifying the picture to be modified; step S3: using the multimodal large model to modify the picture to be modified according to the target area and the target prompt code, the present invention can modify the picture with high accuracy according to the user's modification intention, and by utilizing the abundant computing power of the intelligent computing center and executing steps S21, S22 and S23 in a parallel manner, the present invention improves execution efficiency, realizes efficient modification of pictures, and improves user experience.

[0080] In some embodiments of the present invention, optionally, step S3 includes:

[0081] Step S31: configuring the first computing power resource of the intelligent computing center required to modify the image to be modified for the multimodal large model;

[0082] Among them, the multimodal large model infers the target area and target prompt encoding based on the first computing power resource to modify the image to be modified.

[0083] In some optional embodiments, the amount of computation required for the modification can be evaluated based on the image to be modified (image frame, image size), the target area (size of the target area) and the target prompt code (modification requirements), and then the required first computing power resources can be determined, thereby avoiding increased time spent on modifications due to insufficient computing power resources, improving user experience, avoiding waste of computing power resources due to excessive allocation of computing power resources, and improving the utilization efficiency of computing power resources.

[0084] In the present invention, the first computing power resources of the intelligent computing center required for configuring the multimodal large model to modify the image to be modified are configured through step S31, which helps to avoid redundant computing power resources and insufficient computing power resources and improve the computing power resource utilization efficiency of the intelligent computing center.

[0085] In some embodiments of the present invention, optionally, step S21 includes:

[0086] Step S211: using computer vision (CV) methods to identify the target subject to be modified from the image to be modified and determine the target area.

[0087] Computer Vision (CV) is an interdisciplinary field that aims to enable computers to understand and interpret visual information. The present invention can achieve high-efficiency and high-accuracy recognition of the modified target subject through step S211, and can determine the target area with high efficiency and high accuracy.

[0088] In some embodiments of the present invention, optionally, the target area includes: a core area, a transition area;

[0089] Step S211 and subsequent steps include:

[0090] Step S212: using the CV method to identify the edges of each modified target body from the target area, drawing a minimum convex set along the edge of the modified target body, and taking the area enclosed by the minimum convex set as the core area;

[0091] Step S213: determining other areas in the target area except the core area as transition areas;

[0092] Step S3 includes:

[0093] Step S32: The multimodal large model infers the core area and the target prompt code to generate a modified replacement graphic for the modified target body; the modified replacement graphic is used to replace the modified target body in the core area; after replacing all the modified target bodies, the multimodal large model infers the transition area and the target prompt code to modify the light and shadow effects of the transition area.

[0094] Minimal Convex Set is an important concept in convex analysis and optimization. Specifically in the present invention, the minimum convex set drawn along the edge of the modified target body is the range formed by all the outermost pixel points on the edge of the modified target body, that is, the minimum area of ​​the envelope modified target body.

[0095] The present invention uses the CV method to identify the edges of each modified target body from the target area through step S212, draws the minimum convex set along the edge of the modified target body, and takes the area enveloping the minimum convex set as the core area, so as to accurately determine the core area. Furthermore, the computing power resources required for the multimodal large model to perform reasoning on the core area (i.e., the multimodal large model in step S32 of the present invention infers the core area and the target prompt encoding to generate a modified replacement graphic of the modified target body; and uses the modified replacement graphic to replace the modified target body in the core area) are reduced as much as possible, thereby improving the utilization efficiency of the computing power resources of the intelligent computing center. In some optional embodiments, the first computing power resource is also reduced.

[0096] In the present invention, the target area is divided into a core area and a transition area, and for the core area, the multimodal large model infers the core area and the target prompt code to generate a modified replacement graphic for modifying the target body; the modified target body in the core area is replaced by the modified replacement graphic; for the transition area, after replacing all the modified target bodies, the multimodal large model infers the transition area and the target prompt code to modify the light and shadow effects of the transition area. While ensuring that the core area is modified efficiently and accurately, the present invention reduces the inference intensity of the multimodal large model on the transition area, realizes the precise inference application of the multimodal large model, and improves the utilization efficiency of computing resources; in addition, through the precise inference application of the multimodal large model, the present invention also improves the efficiency of modifying pictures, reduces the time required to modify pictures, and improves user experience.

[0097] In some embodiments of the present invention, optionally, a virtual Kubernetes service deployed in an intelligent computing center is called to execute step S2, wherein required service containers are configured for executing steps S21, S22, and S23, respectively.

[0098] Virtual Kubernetes Service, also known as VKS (Virtual Kubernetes Service), refers to a Kubernetes cluster deployed and managed in a virtual machine (VM) or containerized environment. This deployment method can provide an isolated and controllable Kubernetes environment suitable for a variety of scenarios, including development, testing, and production environments.

[0099] In the present invention, step S2 is executed by using a virtual Kubernetes service deployed in an intelligent computing center, wherein required service containers are configured for executing steps S21, S22, and S23, respectively, thereby improving the efficiency of executing steps S21, S22, and S23 in parallel, shortening the time required for modifying images, and improving user experience.

[0100] In some embodiments of the present invention, optionally, step S23 includes:

[0101] Step S231: sending a model list to an interactive terminal associated with the user, the model list including: names and descriptions of all models in a preset image modification multimodal large model set;

[0102] Step S232: receiving a model determination instruction sent by the interaction end, and in response to the model determination instruction, determining that the model indicated by the model determination instruction is a multimodal large model.

[0103] In some embodiments of the present invention, optionally, the image modification multimodal large model set includes at least one of the following multimodal large models: Stable Diffusion, flux.

[0104] Stable Diffusion is an image generation model based on deep learning, especially used to generate high-quality images and artworks. It belongs to the category of Generative Adversarial Network (GAN) and Diffusion Model, and has made significant progress in the field of image generation in recent years.

[0105] The Flux model is an AI model developed by Black Forest Labs that generates images from text. The Flux model can efficiently output high-resolution, detailed images without the need for additional plug-ins. This is especially important when users generate promotional posters and conduct conceptual design. Especially when rendering human anatomy, the images generated by the Flux model can better reflect real details, giving it a significant advantage in the professional field.

[0106] The present invention is described below in conjunction with specific embodiments.

[0107] See also Figure 2 and Figure 3 As shown, Figure 2 A schematic diagram of the picture to be modified input by the user. Figure 3 This is a schematic diagram of the reference figure input by the user, and the modification intention description text input by the user is "I want to change the white flowers drawn in the red circle into red roses."

[0108] Processing modules (1), (2) and (3) execute the image modification process in parallel, wherein:

[0109] Processing module (1):

[0110] Through the description of "the white flower circled in red" in the prompt word, the computer vision CV automatically identifies the object to be modified (i.e., the target object to be modified). Figure 4 The area enclosed by the blue circle in the middle red circle is the core area, and the other areas in the red circle are transition areas.

[0111] Processing module (2):

[0112] Use reverse encoding Figure 3 The reference picture shown in the figure is used to obtain the encoding result 1. The reverse encoding is performed to modify the intention description text ("I want to change the white flowers drawn in the red circle into red roses") to obtain the encoding result 2.

[0113] The fused encoding results 1 and 2 are used as the prompt words of the final input model (ie, the target prompt coding).

[0114] Processing module (3):

[0115] The selected model (i.e., the multimodal large model in the present invention) is used to automatically load the model into a suitable GPU (i.e., the first computing power resource of the intelligent computing center required for configuring the multimodal large model to modify the image to be modified).

[0116] Determine the model size -> select GPU resources -> pull the model and load it into the GPU -> listen for requests.

[0117] The target area and target prompt encoding are input into the selected model (i.e., the multimodal large model in the present invention), and the modified result is obtained. Figure 5 As shown, the white flowers originally located in the red circle are changed to red roses.

[0118] The present invention provides a device for modifying images through the computing power of an intelligent computing center, see Figure 6 As shown, Figure 6 The schematic diagram of the apparatus for modifying a picture by using the computing power of an intelligent computing center of the present invention is shown in FIG. 60 . The apparatus for modifying a picture by using the computing power of an intelligent computing center includes:

[0119] The receiving module 61 is used to receive the picture to be modified, the reference drawing and the modification intention description text sent by the user through the interactive terminal;

[0120] An execution module 62 is used to call the computing power of the intelligent computing center to perform semantic understanding of the modification intention description text in parallel to obtain the modification target subject that the user needs to modify; determine the area where the modification target subject is located on the image to be modified as the target area; perform reverse encoding of the modification intention description text and the reference figure in parallel, and merge the encodings of the two to obtain the target prompt encoding; and perform the step of determining a pre-trained multimodal large model for modifying the image to be modified in parallel;

[0121] The modification module 63 is used to modify the image to be modified according to the target area and the target prompt code by using the multimodal large model.

[0122] In some embodiments of the present invention, optionally, the modification module 63 is further used to configure the first computing power resource of the intelligent computing center required to modify the image to be modified for the multimodal large model;

[0123] Among them, the multimodal large model infers the target area and the target prompt code based on the first computing power resource to modify the image to be modified.

[0124] In some embodiments of the present invention, optionally, the execution module 62 is further configured to identify the modification target subject in the modification area from the image to be modified by using a computer vision CV method to determine the target area.

[0125] In some embodiments of the present invention, optionally, the target area includes: a core area, a transition area;

[0126] The execution module 62 is further configured to use the CV method to identify the edges of each modified target body from the target area, draw a minimum convex set along the edges of the modified target body, and use the area enclosed by the minimum convex set as the core area;

[0127] The execution module 62 is further configured to determine other areas in the target area except the core area as the transition area;

[0128] The modification module 63 is also used for the multimodal large model to infer the core area and the target prompt code to generate a modified replacement graphic of the modified target body; use the modified replacement graphic to replace the modified target body in the core area; after replacing all the modified target bodies, the multimodal large model infers the transition area and the target prompt code to modify the light and shadow effects of the transition area.

[0129] In some embodiments of the present invention, optionally, a virtual Kubernetes service deployed in the intelligent computing center is called to execute the computing power of the intelligent computing center to perform semantic understanding of the modification intention description text in parallel to obtain the modification target subject that the user needs to modify; determine the area where the modification target subject is located on the image to be modified as the target area; perform reverse encoding of the modification intention description text and the reference figure respectively in parallel, and fuse the encodings of the two to obtain the target prompt encoding; and perform the step of determining a pre-trained multimodal large model for modifying the image to be modified in parallel, wherein, in order to perform semantic understanding of the modification intention description text to obtain the modification target subject that the user needs to modify; determine the area where the modification target subject is located on the image to be modified as the target area; perform reverse encoding of the modification intention description text and the reference figure respectively, and fuse the encodings of the two to obtain the target prompt encoding, and configure the required service containers respectively for the step of determining the pre-trained multimodal large model for modifying the image to be modified.

[0130] In some embodiments of the present invention, optionally, the execution module 62 is further used to send a model list to an interactive terminal associated with the user, the model list including: the names and descriptions of all models in the preset picture modification multimodal large model set;

[0131] The execution module 62 is further configured to receive a model determination instruction sent by the interaction terminal, and in response to the model determination instruction, determine that the model indicated by the model determination instruction is the multimodal large model.

[0132] The device provided by the present invention for modifying pictures by using the computing power of an intelligent computing center can achieve Figures 1 to 5 The various processes implemented by the method embodiment and achieving the same technical effect are not described here to avoid repetition.

[0133] The present invention provides an electronic device 70, see Figure 7 As shown, Figure 7 It is a principle block diagram of an electronic device 70 of the present invention, including a processor 71, a memory 72, and a program or instruction stored in the memory 72 and executable on the processor 71. When the program or instruction is executed by the processor, any step in the method of modifying an image through the computing power of an intelligent computing center of the present invention is implemented.

[0134] The present invention provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of an embodiment of a method for modifying an image by the computing power of an intelligent computing center as described above is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0135] The readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc. In some examples, the readable storage medium may be a non-transitory readable storage medium.

[0136] The present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the various processes of any of the above-mentioned methods for modifying an image through the computing power of an intelligent computing center, and can achieve the same technical effect. To avoid repetition, they will not be described here.

[0137] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.

Claims

1. A method for modifying an image by using the computing power of an intelligent computing center, characterized in that: include: Step S1: receiving a picture to be modified, a reference drawing, and a modification intention description text sent by a user through an interactive terminal; Step S2: calling the computing power of the intelligent computing center to execute step S21, step S22 and step S23 in parallel, wherein step S21 includes: semantically understanding the modification intention description text to obtain the modification target subject that the user needs to modify; determining the area where the modification target subject is located on the image to be modified as the target area; step S22 includes: reverse encoding the modification intention description text and the reference figure respectively, and fusing the encodings of the two to obtain the target prompt encoding; step S23 includes: determining a pre-trained multimodal large model for modifying the image to be modified; Step S3: using the multimodal large model, modifying the image to be modified according to the target area and the target prompt code.

2. The method for modifying an image by using the computing power of an intelligent computing center according to claim 1, characterized in that: The step S3 comprises: Step S31: configuring the multimodal large model with the first computing resource of the intelligent computing center required to modify the image to be modified; Among them, the multimodal large model infers the target area and the target prompt code based on the first computing power resource to modify the image to be modified.

3. The method for modifying an image by using the computing power of an intelligent computing center according to claim 1, characterized in that: The step S21 comprises: Step S211: using a computer vision (CV) method to identify the modification target subject in the modification area from the image to be modified, and determining the target area.

4. The method for modifying an image by using the computing power of an intelligent computing center according to claim 3, characterized in that: The target area includes: a core area and a transition area; The step S211 then includes: Step S212: using the CV method to identify the edges of each modified target body from the target area, drawing a minimum convex set along the edge of the modified target body, and taking the area enclosed by the minimum convex set as the core area; Step S213: determining other areas in the target area except the core area as the transition area; The step S3 comprises: Step S32: the multimodal large model infers the core area and the target prompt code to generate a modified replacement graphic of the modified target body; the modified replacement graphic is used to replace the modified target body in the core area; after replacing all the modified target bodies, the multimodal large model infers the transition area and the target prompt code to modify the light and shadow effects of the transition area.

5. The method for modifying images by using the computing power of an intelligent computing center according to claim 1, characterized in that: The virtual Kubernetes service deployed in the intelligent computing center is called to execute step S2, wherein the required service containers are configured respectively for executing step S21, step S22 and step S23.

6. The method for modifying an image by using the computing power of an intelligent computing center according to claim 1, characterized in that: The step S23 comprises: Step S231: sending a model list to an interactive terminal associated with the user, the model list including: names and descriptions of all models in a preset picture modification multimodal large model set; Step S232: receiving a model determination instruction sent by the interaction terminal, and in response to the model determination instruction, determining that the model indicated by the model determination instruction is the multimodal large model.

7. A device for modifying an image by using the computing power of an intelligent computing center, characterized in that: include: A receiving module, used to receive the picture to be modified, the reference drawing and the modification intention description text sent by the user through the interactive terminal; An execution module is used to call the computing power of the intelligent computing center to perform semantic understanding of the modification intention description text in parallel to obtain the modification target subject that the user needs to modify; The step of determining the area where the modification target subject is located on the to-be-modified picture as the target area; and the step of performing reverse encoding on the modification intention description text and the reference figure respectively in parallel, and fusing the encodings of the two to obtain the target prompt encoding; And, executing in parallel the step of determining a pre-trained multimodal large model for modifying the image to be modified; A modification module is used to use the multimodal large model to modify the image to be modified according to the target area and the target prompt code.

8. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps in the method for modifying an image by using the computing power of an intelligent computing center as described in any one of claims 1 to 6.

9. A readable storage medium, characterized in that: The readable storage medium stores programs or instructions, and when the programs or instructions are executed by the processor, the steps in the method of modifying an image by using the computing power of an intelligent computing center as described in any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that It comprises computer instructions, which, when executed by a processor, implement the steps of the method for modifying a picture by using the computing power of an intelligent computing center as described in any one of claims 1 to 6.