Data processing method and apparatus therefor

By introducing the texture and semantic features of the image in the generative model as prior information, the problem that image super-resolution technology in the prior art cannot understand the input content, and the improvement of generated image quality and enhanced model generalization is achieved.

WO2025167429A1PCT designated stage Publication Date: 2025-08-14HUAWEI TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/070905
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-05
Filing Date
2025-01-07
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing image super-resolution technology cannot effectively understand the content of the input image, resulting in a low quality of generated images.

Method used

By obtaining the target information of the texture features and semantic features of the first image, it is input into the generative model as a priori, and the generative model is guided to optimize the image quality.

Benefits of technology

Improve the quality of generated images, enhance texture and semantic details in the images, and improve the generalization of the model in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025070905_14082025_PF_FP_ABST
    Figure CN2025070905_14082025_PF_FP_ABST
Patent Text Reader

Abstract

A data processing method, which is applied to the field of artificial intelligence and comprises: acquiring a first image; generating target information on the basis of the first image, wherein the target information is used for indicating texture features or semantic features of the first image; and on the basis of the first image and the target information, obtaining a second image by means of a first generative model, wherein the second image has the texture features or the semantic features. In the present application, by means of determining target information, which is related to texture features or semantic features, in a first image, and inputting the target information as a priori into a generative model, the generative model can be helped to understand the content and texture of an input image, such that the texture details and semantics details in a generated image can be better restored, thereby obtaining a generated image with higher quality. The generalization performance of a model in different scenarios can also be improved.
Need to check novelty before this filing date? Find Prior Art

Description

A data processing method and device thereof

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on February 5, 2024, with application number 202410171549.0 and application name “A data processing method and device thereof”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence, and in particular to a data processing method and device thereof. Background Art

[0003] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that seeks to understand the essence of intelligence and develop new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0004] Image enhancement technology aims to convert low-quality images into high-quality images. For example, image super-resolution technology is used to convert low-resolution images into high-resolution images, thereby improving the clarity and detail authenticity of the image.

[0005] Existing image super-resolution techniques typically fit a mapping between a low-resolution image and a high-resolution image. To better restore details, existing methods use pre-trained generative models.

[0006] However, existing methods usually do not understand the content of the input image, resulting in low quality of generated images. Summary of the Invention

[0007] In a first aspect, the present application provides a data processing method, comprising: acquiring a first image; generating target information based on the first image, the target information comprising a first text and / or a second image, the first text being used to indicate an optimization direction of the image quality of the first image, the image quality of the second image being higher than that of the first image, and the semantics of the second image corresponding to those of the first image; and obtaining a third image through a first generation model based on the first image and the target information, the third image having higher image quality than that of the first image, and the content of the third image being consistent with that of the first image.

[0008] Image quality refers to a person's subjective evaluation of the visual experience of an image. It is generally considered to refer to the degree of error in the human visual system between the image being measured (i.e., the target image) and the reference image (i.e., the original image). Another definition states that good image quality occurs when the human eye can clearly distinguish objects in the image, even without a reference image, and can effectively distinguish foreground and background, object outlines, and textures.

[0009] Among them, the "quality optimization direction" can be understood as an improvement in people's visual perception of images, with the foreground and background, object contours, textures, etc. in the image becoming clearer and better distinguishable.

[0010] Among them, the pixels of the first image (low-definition image) and the second image (reference image) do not match, and the pixels of the first image and the second image at corresponding positions are very different, while the pixels of the first image (low-definition image) and the third image are matched, and the pixels of the first image and the second image at corresponding positions are very small. Visually, it can be considered that the same picture indicates that there is a difference in the clarity of the image.

[0011] Among them, the quality can be evaluated by the degree of human visual satisfaction, or by some non-reference indicators, such as NIQE (Natural Image Quality Evaluator), MANIQA (Multi-dimension Attention Network for No-Reference Image Quality), and multi-scale image quality (MUSIQ).

[0012] The pixels of the first and third images (high-definition images) are basically matched, and the quality improvement can be manifested in improved clarity and richer details at the same location. The above-mentioned no-reference indicators are also improved accordingly.

[0013] The target information can indicate the direction of quality improvement of the generation results of the first generation model and be input into the generation model as prior information, thereby helping the generation model to be better optimized and improve the quality of the generated image.

[0014] In one possible implementation, generating target information based on the first image includes: generating target information based on the first image using a machine learning model. In other words, the target information is obtained based on the first information.

[0015] The first generation model is used to perform a target task, and the target information is also used to indicate the features that the first generation model needs to possess in order to obtain a high-quality result after performing the target task on the first image, and the required features are not possessed by the first image.

[0016] The quality improvement direction may be: based on the first image, what kind of generated image obtained by the first generative model is a higher quality generated image. The quality improvement direction may describe the image quality improvement direction at a finer granularity level. Optionally, this information may be input into the generative model in text form.

[0017] In one possible implementation, the target information is also used to indicate the texture features or semantic features of the first image. In the existing technology, when enhancing an image, such as during super-resolution processing, the noise of the image and the image itself are input into the generative network to obtain a super-resolved image. However, the generative network often does not understand the content of the input image and cannot use common sense to accurately restore the semantics and texture of the objects in the image, especially in real scenes, the generated image effect is even worse. In this application, by determining the target information related to the texture features or semantic features in the first image, and inputting the target information as a priori into the generative model, it can help the generative model understand the content and texture of the input image, so that the texture and semantic details in the generated image can be better repaired, thereby obtaining a higher quality generated image. It can also improve the generalization of the model in different scenarios.

[0018] Among them, texture is a visual feature that reflects homogeneous phenomena in an image. It reflects the surface structural organization and arrangement properties of an object with slow or periodic changes. Texture is expressed through the grayscale distribution of pixels and their surrounding spatial neighborhood, which is local texture information. In addition, local texture information is different from image features such as grayscale and color to varying degrees, and its repeatability is global texture information. While texture features reflect the properties of global features, they also describe the surface properties of the scene corresponding to the image or image region. However, since texture is only a characteristic of the surface of an object and cannot fully reflect the essential properties of the object, it is impossible to obtain high-level image content using only texture features. Unlike color features, texture features are not pixel-based features. They need to be statistically calculated in an area containing multiple pixels.

[0019] The semantic features of an image refer to the meaning of the image content. These features are not limited to descriptions in natural language terms but can encompass the way the human visual system understands images. For example, the semantics of an image of a puppy might be interpreted as containing the natural language word "puppy," along with attributes such as the puppy's breed and gender. Alternatively, the concept could be represented by a symbol that refers to a puppy with the same characteristics as the one in the image. Furthermore, people can derive an abstract impression from the image, which helps them connect it to known images when they see other similar puppy images.

[0020] In a possible implementation, the first generative model is used to perform a target task, which is one of the following: image enhancement, image noise addition, image style transformation, and image editing.

[0021] In one possible implementation, generating target information based on the first image includes: identifying target information in the first image; or receiving target information in the first image input by a user. In other words, based on user-provided prompt information, the object's structure and texture can be better restored, thereby improving image fidelity and quality.

[0022] In a possible implementation, the target information may include an encoding of texture features of the first image.

[0023] For example, the first image may be processed by an encoder (eg, an image encoder) to obtain an encoding of texture features of the first image.

[0024] In a possible implementation, the target information may include an encoding of semantic features of the first image.

[0025] For example, the first image may be processed by an encoder (eg, a text encoder) to obtain an encoding of texture features of the first image.

[0026] In a possible implementation, the target information may include an encoding of the first image that is a mixture of the texture feature and the semantic feature.

[0027] For example, the first image can be processed by an encoder (for example, a text encoder and an image encoder) to obtain the encoding of the texture features and the encoding of the semantic features of the first image respectively, and then the encoding of the texture features and the encoding of the semantic features of the first image are fused using a modality adaptation module to obtain an encoding that is a mixture of the texture features and the semantic features.

[0028] In a possible implementation, the target information may include a natural language description of texture features of the first image.

[0029] For example, the first image can be processed by a machine learning model, and the machine learning model can be trained to have the ability to obtain a natural language description of texture features in the image based on the image.

[0030] In a possible implementation, the target information may include a natural language description of semantic features of the first image.

[0031] For example, the first image can be processed by a machine learning model, and the machine learning model can be trained to have the ability to obtain a natural language description of the semantic features in the image based on the image.

[0032] In a possible implementation, the target information may include a natural language description of the first image that is a mixture of the texture features and the semantic features.

[0033] For example, the first image can be processed by a machine learning model, and the machine learning model can be trained to have the ability to obtain a natural language description of texture features and semantic features in the image based on the image.

[0034] In addition, in addition to the encoding or natural language description introduced above, the semantic features can also be semantic segmentation maps, semantic labels of images, main areas of images, classification labels of images, etc., which are not limited in the embodiments of this application.

[0035] In a possible implementation, the target information includes an image that has texture features or semantic features of the first image and is different from the first image.

[0036] Through the above approach, another image (also referred to as a reference image in this application) with the texture features or semantic features of the first image is used as prior information for the generative model and input into the generative model in the form of an image. This explicitly guides the generation of texture details and helps the generative model understand the content and texture of the input image, thereby better restoring the texture and semantic details in the generated image, resulting in a higher-quality generated image. This can also improve the generalization of the model in different scenarios.

[0037] In a possible implementation, generating target information based on the first image includes: obtaining, through a second generative model, the image having the texture features or semantic features of the first image according to an encoding or natural language description of texture features of the first image, or an encoding or natural language description of semantic features of the first image; or

[0038] According to the encoding or natural language description of the texture feature of the first image, or the encoding or natural language description of the semantic feature of the first image, an image having the texture feature or semantic feature of the first image is selected from multiple images.

[0039] In addition, the target information can also indicate the direction of quality improvement of the generation results of the first generation model and be input into the generation model as prior information, thereby helping the generation model to be better optimized and improve the quality of the generated image.

[0040] In one possible implementation, the first generation model is used to perform a target task, and the target information is also used to indicate the features that the first generation model needs to possess in order to obtain a high-quality result after performing the target task on the first image, and the required features are not possessed by the first image.

[0041] The quality improvement direction may be: based on the first image, what kind of generated image obtained by the first generative model is a higher quality generated image. The quality improvement direction may describe the image quality improvement direction at a finer granularity level. Optionally, this information may be input into the generative model in text form.

[0042] In one possible implementation, the first generation model is a diffusion model, and obtaining the second image based on the first image and the target information through the first generation model includes: using the target information and the first image as conditional inputs of a noise addition module in the diffusion model, and obtaining the second image through the diffusion model.

[0043] In a second aspect, the present application provides a data processing device, comprising:

[0044] An acquisition module, configured to acquire a first image;

[0045] A processing module is used to generate target information based on the first image, where the target information includes a first text and / or a second image, where the first text is used to indicate an optimization direction of the image quality of the first image, the image quality of the second image is higher than that of the first image, and the semantics of the second image corresponds to that of the first image; and based on the first image and the target information, a third image is obtained through a first generation model, where the image quality of the third image is higher than that of the first image, and the content of the third image is consistent with that of the first image.

[0046] In one possible implementation, the first generative model is used to perform an image enhancement task.

[0047] In a possible implementation, the processing module is specifically configured to:

[0048] The target information is obtained according to the first image through a machine learning model.

[0049] In a possible implementation, the image quality is at least one of the following definitions: NIQE, MANIQA, and MUSIQ.

[0050] In a possible implementation, the target information is further used to indicate a texture feature or a semantic feature of the first image.

[0051] In one possible implementation, the target information includes at least one of the following:

[0052] An encoding of the texture features of the first image, or an encoding of the semantic features, or an encoding mixed with the texture features and the semantic features, or a natural language description of the texture features of the first image, or a natural language description of the semantic features, or a natural language description mixed with the texture features and the semantic features.

[0053] In a possible implementation, the processing module is specifically configured to:

[0054] Obtaining, according to the first image, an encoding of texture features of the first image through an image encoder; or,

[0055] Obtaining an encoding of the semantic features of the first image through a text encoder according to the text description of the first image; or

[0056] According to the coding of the texture feature and the coding of the semantic feature, a coding mixed with the texture feature and the semantic feature is obtained through a modality adaptation module.

[0057] In a possible implementation, the target information includes an image that has texture features or semantic features of the first image and is different from the first image.

[0058] In a possible implementation, the processing module is specifically configured to:

[0059] Obtaining the image having the texture features or semantic features of the first image through a second generative model according to the encoding or natural language description of the texture features of the first image, or the encoding or natural language description of the semantic features of the first image; or

[0060] According to the encoding or natural language description of the texture feature of the first image, or the encoding or natural language description of the semantic feature of the first image, an image having the texture feature or semantic feature of the first image is selected from multiple images.

[0061] In one possible implementation, the first generation model is used to perform a target task, and the target information is also used to indicate the features that the first generation model needs to possess in order to obtain a high-quality result after performing the target task on the first image, and the required features are not possessed by the first image.

[0062] In a possible implementation, the first generation model is a diffusion model, and the processing module is specifically configured to:

[0063] The target information and the first image are used as conditional inputs of a noise adding module in the diffusion model, and a second image is obtained through the diffusion model.

[0064] In a third aspect, an embodiment of the present application provides a data processing device, which may include a memory, a processor, and a bus system, wherein the memory is used to store programs, and the processor is used to execute the programs in the memory to perform the first aspect and any optional method thereof.

[0065] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned first aspect and any optional method thereof.

[0066] In a fifth aspect, an embodiment of the present application provides a computer program, which, when executed on a computer, enables the computer to execute the above-mentioned first aspect and any optional method thereof.

[0067] In a sixth aspect, the present application provides a chip system comprising a processor configured to support the execution of a data processing device to implement the functions described in the aforementioned aspects, such as transmitting or processing the data or information described in the aforementioned methods. In one possible design, the chip system further comprises a memory configured to store program instructions and data necessary for executing the device or training the device. The chip system may consist of a single chip or may include a chip and other discrete components. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] FIG1A is a schematic diagram of a structure of an artificial intelligence main framework;

[0069] 1B and 1C are schematic diagrams of the application system framework of the present invention;

[0070] FIG1D is a schematic diagram of an optional hardware structure of a terminal;

[0071] FIG2 is a schematic diagram of the structure of a server;

[0072] FIG3 is a schematic diagram of a system architecture of the present application;

[0073] Figure 4 shows a process of cloud services;

[0074] FIG5 is a flowchart of a data processing method provided in an embodiment of the present application;

[0075] FIG6 is a schematic diagram of a data processing method provided in an embodiment of the present application;

[0076] FIG7 is a schematic diagram of an effect provided by an embodiment of the present application;

[0077] FIG8 is a schematic diagram of a data processing method provided in an embodiment of the present application;

[0078] FIG9A is a schematic diagram of an effect provided by an embodiment of the present application;

[0079] FIG9B is a schematic diagram of a data processing method provided in an embodiment of the present application;

[0080] FIG9C is a schematic diagram of an effect provided by an embodiment of the present application;

[0081] FIG9D is a schematic diagram of a data processing method provided in an embodiment of the present application;

[0082] FIG10 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;

[0083] FIG11 is a schematic diagram of the structure of an execution device provided in an embodiment of the present application;

[0084] FIG12 is a schematic diagram of a structure of a training device provided in an embodiment of the present application;

[0085] FIG13 is a schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION

[0086] The following describes the embodiments of the present invention in conjunction with the accompanying drawings. The terms used in the embodiments of the present invention are only used to explain the specific embodiments of the present invention, and are not intended to limit the present invention.

[0087] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0088] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0089] As used herein, the terms "substantially," "about," and similar terms are used as terms of approximation, not as terms of degree, and are intended to take into account the inherent variations in measurements or calculations that one of ordinary skill in the art would recognize. Furthermore, the use of "may" when describing embodiments of the present invention refers to "one or more possible embodiments." As used herein, the terms "use," "using," and "used" may be considered synonymous with the terms "utilize," "utilizing," and "utilized," respectively. Additionally, the term "exemplary" is intended to refer to an example or illustration.

[0090] First, let's describe the overall workflow of an AI system. See Figure 1A, which shows a schematic diagram of the main AI framework. This AI framework will be explained from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed journey from "data-information-knowledge-wisdom." The "IT value chain," spanning the underlying infrastructure of human intelligence, information (provided and processed by technology), and the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.

[0091] (1) Infrastructure

[0092] Infrastructure provides computing power for AI systems, enabling communication with the outside world and supporting this through a foundational platform. External communication occurs through sensors; computing power is provided by intelligent chips (CPUs, NPUs, GPUs, ASICs, FPGAs, and other hardware accelerators). The foundational platform includes a distributed computing framework and network-related platform guarantees and support, including cloud storage and computing, and interconnected networks. For example, sensors communicate with the outside world to acquire data, which is then fed into the intelligent chips within the distributed computing system provided by the foundational platform for computation.

[0093] (2) Data

[0094] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.

[0095] (3) Data processing

[0096] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0097] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.

[0098] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.

[0099] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.

[0100] (4) General ability

[0101] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0102] (5) Smart products and industry applications

[0103] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart transportation, smart medical care, autonomous driving, smart cities, etc.

[0104] This application can be applied to the field of natural language processing in the field of artificial intelligence. Taking natural language processing as an example, the following will introduce multiple application scenarios that have been implemented in products.

[0105] First, we will introduce the application scenarios of this application. This application can be applied to, but is not limited to, applications with image generation functions (hereinafter referred to as generation applications) or cloud services provided by cloud-side servers. The following are introduced respectively:

[0106] 1. Generate application

[0107] The product form of the embodiment of the present application can be a generation type application. The generation type application can be run on a terminal device or a cloud-side server.

[0108] Among them, the image generation in the embodiment of the present application can be to regenerate an image based on the input image, the image generation task can be image enhancement (for example, denoising), image noise addition, etc., and the image editing can be to replace part of the subject in the image (such as foreground or background), etc.

[0109] In a possible implementation, a generation application can implement an image generation task and obtain a processing result.

[0110] For example, the generation application may implement at least an image generation task based on a diffusion method, but is not limited thereto.

[0111] In one possible implementation, the user can open a generation application installed on the terminal device and input image data. The generation application can use a model trained by the method provided in the embodiment of the present application, or process the image data by the method provided in the embodiment of the present application, and present the processing result (generated image) to the user (the presentation method can be but is not limited to display, playback, saving, uploading to the cloud side, etc.).

[0112] In one possible implementation, a user can open a generation application installed on a terminal device and input image data. The generation application can send the image data to a server on the cloud side. The server on the cloud side processes the image data using a model trained using the method provided in an embodiment of the present application, and transmits the processing results back to the terminal device. The terminal device can present the processing results to the user (the presentation method can be, but is not limited to, display, playback, saving, uploading to the cloud side, etc.).

[0113] Next, the generation type application in the embodiment of this application is introduced from the functional architecture and the product architecture that realizes the function.

[0114] Referring to FIG. 1B , FIG. 1B is a schematic diagram of the functional architecture of a generation-type application in an embodiment of the present application:

[0115] In one possible implementation, as shown in FIG1B , a generation application 102 may receive input parameters 101 (e.g., including image data) and generate a processing result 103. The generation application 102 may be executed on, for example, at least one computer system and include computer code that, when executed by one or more computers, causes the computers to execute a model trained using the method provided in the embodiments of the present application.

[0116] Referring to FIG. 1C , FIG. 1C is a schematic diagram of the entity architecture for running a generation-type application in an embodiment of the present application:

[0117] Referring to FIG1C , FIG1C shows a schematic diagram of a system architecture. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers (FIG1C illustrates one server as an example), and the server 200 may provide image generation or natural language generation functions for one or more terminals.

[0118] Among them, the terminal 100 can be installed with a generation application, or a web page related to the image generation or natural language generation function can be opened. The above application and web page can provide an interface. The terminal 100 can receive the relevant parameters entered by the user on the image generation or natural language generation function interface, and send the above parameters to the server 200. The server 200 can obtain the processing results based on the received parameters and return the processing results to the terminal 100.

[0119] It should be understood that in some optional implementations, the terminal 100 can also complete the action of obtaining the processing result based on the received parameters by itself without the need for the cooperation of the server, and the embodiments of the present application are not limited to this.

[0120] Next, the product form of the terminal 100 in FIG1C is described;

[0121] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a wearable device, an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any restrictions on this.

[0122] FIG1D shows a schematic diagram of an optional hardware structure of the terminal 100 .

[0123] 1D , the terminal 100 may include components such as a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), a processor 170, an external interface 180, and a power supply 190. Those skilled in the art will appreciate that FIG1D is merely an example of a terminal or multi-function device and does not limit the terminal or multi-function device. The terminal or multi-function device may include more or fewer components than shown, or may combine certain components or have different components.

[0124] The input unit 130 can be used to receive input digital or character information and generate key signal input related to user settings and function control of the portable multifunction device. Specifically, the input unit 130 may include a touch screen 131 (optional) and / or other input devices 132. The touch screen 131 can detect user touch operations on or near it (for example, operations performed on or near the touch screen using a finger, joint, stylus, or any other suitable object) and drive corresponding connected devices according to pre-set programs. The touch screen can detect user touch actions on the touch screen, convert the touch actions into touch signals and transmit them to the processor 170. It can also receive and execute commands sent by the processor 170; the touch signals include at least touch point coordinate information. The touch screen 131 provides an input interface and an output interface between the terminal 100 and the user. Touch screens can be implemented using various types, including resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch screen 131, the input unit 130 may also include other input devices. Specifically, the other input devices 132 may include, but are not limited to, one or more of a physical keyboard, function keys (such as a volume control key, a switch key, etc.), a trackball, a mouse, a joystick, and the like.

[0125] Among them, other input devices 132 can receive input image data or text data.

[0126] The display unit 140 may be used to display information input by the user or provided to the user, various menus of the terminal 100, interactive interfaces, file display, and / or playback of any multimedia file. In an embodiment of the present application, the display unit 140 may be used to display the interface of a generated application, processing results, etc.

[0127] Memory 120 can be used to store instructions and data. It primarily includes an instruction storage area and a data storage area. The data storage area can store various data, such as multimedia files and text. The instruction storage area can store software units such as the operating system, applications, and instructions required for at least one function, or subsets or extensions thereof. It may also include non-volatile random access memory (RAM). It provides processor 170 with management functions for the hardware, software, and data resources within the computing and processing device, supporting control software and applications. It is also used to store multimedia files and running programs and applications.

[0128] The processor 170 is the control center of the terminal 100. It connects all components of the terminal 100 using various interfaces and circuits. By executing instructions stored in the memory 120 and accessing data stored therein, it executes various functions of the terminal 100 and processes data, thereby providing overall control of the terminal device. Optionally, the processor 170 may include one or more processing units. Preferably, the processor 170 may integrate an application processor and a modem processor, with the application processor primarily processing the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 170. In some embodiments, the processor and memory may be implemented on a single chip; in other embodiments, they may be implemented on separate chips. The processor 170 may also generate corresponding operational control signals and send them to the corresponding components of the computing and processing device. It may also read and process data in the software, particularly the data and programs in the memory 120, to enable the various functional modules therein to perform their corresponding functions, thereby controlling the corresponding components to operate as instructed.

[0129] Among them, the memory 120 can be used to store software codes related to the data processing method, the processor 170 can execute the steps of the chip's data processing method, and can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to achieve corresponding functions.

[0130] The RF unit 110 (optional) can be used to send and receive information or receive and send signals during a call. For example, after receiving downlink information from the base station, it is passed to the processor 170 for processing; in addition, it sends the designed uplink data to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF unit 110 can also communicate with network devices and other devices via wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0131] In this embodiment of the present application, the RF unit 110 may send image data or text data to the server 200 and receive a processing result sent by the server 200.

[0132] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network port.

[0133] The terminal 100 also includes a power supply 190 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 170 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.

[0134] The terminal 100 further includes an external interface 180 , which may be a standard Micro USB interface or a multi-pin connector, and may be used to connect the terminal 100 to other devices for communication, or to connect a charger to charge the terminal 100 .

[0135] Although not shown, the terminal 100 may also include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which are not described in detail here. Some or all of the methods described below may be applied to the terminal 100 shown in FIG1D .

[0136] Next, the product form of the server 200 in FIG1C is described;

[0137] FIG2 provides a schematic diagram of the structure of a server 200. As shown in FIG2, the server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate with each other via the bus 201.

[0138] Bus 201 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, FIG2 shows only one thick line, but this does not imply that there is only one bus or only one type of bus.

[0139] The processor 202 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0140] The memory 204 may include volatile memory, such as random access memory (RAM). The memory 204 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard drive (HDD), or solid state drive (SSD).

[0141] The memory 204 may be used to store software codes related to the data processing method, and the processor 202 may execute the steps of the data processing method of the chip, and may also schedule other units to implement corresponding functions.

[0142] It should be understood that the above-mentioned terminal 100 and server 200 can be centralized or distributed devices, and the processors in the above-mentioned terminal 100 and server 200 (such as processor 170 and processor 202) can be hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processor (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.

[0143] It should be understood that the steps related to the model reasoning process in the embodiments of this application involve AI-related operations. When performing AI operations, the instruction execution architecture of the terminal device and server is not limited to the processor-memory architecture described above. The system architecture provided in the embodiments of this application is described in detail below with reference to Figure 3.

[0144] FIG3 is a schematic diagram of the system architecture provided by an embodiment of the present application. As shown in FIG3 , the system architecture 500 includes an execution device 510 , a training device 520 , a database 530 , a client device 540 , a data storage system 550 , and a data acquisition system 560 .

[0145] The execution device 510 includes a calculation module 511, an I / O interface 512, a pre-processing module 513, and a post-processing module 514. The calculation module 511 may include the target model / rule 501, and the pre-processing module 513 and the post-processing module 514 are optional.

[0146] The execution device 510 may be a terminal device or a server that runs the aforementioned generated application program.

[0147] The data acquisition device 560 is used to collect training samples. The training samples can be image data, etc. After collecting the training samples, the data acquisition device 560 stores these training samples in the database 530.

[0148] The training device 520 can train the neural network to be trained (such as the neural network model in the embodiment of the present application (such as including a modality adaptation module, a first generation model, a second generation model, etc.)) based on the training samples maintained in the database 530 to obtain the target model / rule 501.

[0149] It should be understood that the training device 520 can perform a pre-training process on the neural network to be trained based on the training samples maintained in the database 530, or fine-tune the model based on the pre-training.

[0150] It should be noted that, in actual applications, the training samples maintained in the database 530 may not all be collected by the data acquisition device 560, but may also be received from other devices. It should also be noted that the training device 520 may not train the target model / rule 501 entirely based on the training samples maintained in the database 530, but may also obtain training samples from the cloud or other places for model training. The above description should not be used as a limitation on the embodiments of the present application.

[0151] The target model / rule 501 obtained through training with the training device 520 can be applied to different systems or devices, such as the execution device 510 shown in FIG3 . The execution device 510 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, an in-vehicle terminal, etc., or a server, etc.

[0152] Specifically, the training device 520 may transfer the trained model to the execution device 510 .

[0153] In Figure 3, the execution device 510 is configured with an input / output (I / O) interface 512 for data interaction with an external device. The user can input data (such as image data or text data in the embodiment of the present application) into the I / O interface 512 through the client device 540.

[0154] Preprocessing module 513 and preprocessing module 514 are used to preprocess the input data received by I / O interface 512. It should be understood that preprocessing module 513 and preprocessing module 514 may be absent or only one preprocessing module may be present. If preprocessing module 513 and preprocessing module 514 are absent, computing module 511 may be used directly to process the input data.

[0155] When the execution device 510 preprocesses the input data, or when the computing module 511 of the execution device 510 performs calculations and other related processing, the execution device 510 can call the data, code, etc. in the data storage system 550 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 550.

[0156] Finally, the I / O interface 512 provides the processed results to the client device 540 and thus to the user.

[0157] In the scenario shown in FIG3 , the user can manually input data, and this "manual input data" can be operated through the interface provided by I / O interface 512. In another scenario, client device 540 can automatically send input data to I / O interface 512. If user authorization is required for client device 540 to automatically send input data, the user can set the corresponding permissions in client device 540. The user can view the results output by execution device 510 on client device 540, and the specific presentation form can be a display, sound, action, or other specific method. Client device 540 can also serve as a data acquisition terminal, collecting input data input into I / O interface 512 and output results output from I / O interface 512 as new sample data, and storing them in database 530. Of course, collection can also be performed without client device 540, and instead the I / O interface 512 directly stores the input data input into I / O interface 512 and output results output from I / O interface 512 as new sample data in database 530.

[0158] It is worth noting that FIG3 is merely a schematic diagram of a system architecture provided by an embodiment of the present application, and the positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in FIG3 , the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 can also be placed in the execution device 510. It should be understood that the execution device 510 can be deployed in the client device 540.

[0159] From the inference side of the model:

[0160] In the embodiment of the present application, the computing module 511 of the above-mentioned execution device 510 can obtain the code stored in the data storage system 550 to implement the steps related to the model reasoning process in the embodiment of the present application.

[0161] In an embodiment of the present application, the computing module 511 of the execution device 510 may include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.

[0162] Specifically, the computing module 511 of the execution device 510 can be a hardware system with an execution instruction function, and the steps related to the model reasoning process provided in the embodiment of the present application can be software codes stored in the memory. The computing module 511 of the execution device 510 can obtain the software code from the memory and execute the obtained software code to implement the steps related to the model reasoning process provided in the embodiment of the present application.

[0163] It should be understood that the computing module 511 of the execution device 510 can be a combination of a hardware system that does not have the function of executing instructions and a hardware system that has the function of executing instructions. Some of the steps related to the model reasoning process provided in the embodiment of the present application can also be implemented by the hardware system that does not have the function of executing instructions in the computing module 511 of the execution device 510, which is not limited here.

[0164] From the training side of the model:

[0165] In an embodiment of the present application, the above-mentioned training device 520 can obtain the code stored in the memory (not shown in Figure 3, which can be integrated into the training device 520 or deployed separately from the training device 520) to implement the steps related to model training in the embodiment of the present application.

[0166] In an embodiment of the present application, the training device 520 may include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.

[0167] It should be understood that the training device 520 can be a combination of a hardware system that does not have the function of executing instructions and a hardware system that has the function of executing instructions. Some of the steps related to model training provided in the embodiments of the present application can also be implemented by the hardware system in the training device 520 that does not have the function of executing instructions, which is not limited here.

[0168] 3. Image generation cloud services provided by the server:

[0169] In a possible implementation, the server may provide the image generation function service to the terminal side through an application programming interface (API).

[0170] Among them, the terminal device can send relevant parameters (such as images and other data) to the server through the API provided by the cloud. The server can obtain processing results based on the received parameters, etc., and return the processing results to the terminal.

[0171] The description of the terminal and the server can be the same as that of the above embodiments, and will not be repeated here.

[0172] FIG4 shows a process of using an image generation function cloud service provided by a cloud platform.

[0173] 1. Activate and purchase the image generation service.

[0174] 2. Users can download the software development kit (SDK) corresponding to the image generation service. Cloud platforms usually provide multiple development versions of the SDK for users to choose based on their development environment requirements, such as Java version SDK, Python version SDK, PHP version SDK, Android version SDK, etc.

[0175] 3. After the user downloads the corresponding version of the SDK to the local computer as needed, import the SDK project into the local development environment, configure and debug it in the local development environment. The local development environment can also be used to develop other functions, forming an application that integrates image generation functional capabilities.

[0176] 4. When an image generation application is used and needs to perform image generation, it can trigger an API call for this function. When the application triggers the image generation function, it initiates an API request to the running instance of the image generation service in the cloud environment. The API request carries the image data, and the running instance in the cloud environment processes the image to obtain the processing result.

[0177] 5. The cloud environment returns the processing results to the application, thus completing a call to the image generation function.

[0178] Since the embodiments of the present application involve the application of a large number of neural networks, in order to facilitate understanding, the relevant terms and related concepts such as neural networks involved in the embodiments of the present application are first introduced below.

[0179] (1) Neural Network

[0180] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input. The output of the operation unit can be:

[0181] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal of the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.

[0182] (2) Convolutional neural network (CNN) is a deep neural network with a convolutional structure. Convolutional neural network contains a feature extractor consisting of a convolution layer and a subsampling layer, which can be regarded as a filter. The convolution layer refers to the neuron layer in the convolutional neural network that performs convolution processing on the input signal. In the convolution layer of the convolutional neural network, a neuron can only be connected to some neurons in the adjacent layer. A convolution layer usually contains several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weights here are convolution kernels. Shared weights can be understood as the way of extracting features is independent of position. The convolution kernel can be formalized as a matrix of random size, and the convolution kernel can obtain reasonable weights through learning during the training process of the convolutional neural network. In addition, the direct benefit of shared weights is to reduce the connections between the layers of the convolutional neural network, while reducing the risk of overfitting.

[0183] CNN is a very common neural network. The following section will focus on a detailed description of its structure, using Figure 4. As mentioned in the previous section about basic concepts, a convolutional neural network is a deep neural network with a convolutional structure. It is a deep learning architecture, which uses machine learning algorithms to perform multiple levels of learning at different levels of abstraction. As a deep learning architecture, a CNN is a feed-forward artificial neural network, in which individual neurons respond to input images.

[0184] (3) Deep Neural Networks

[0185] Deep Neural Network (DNN), also known as multi-layer neural network, can be understood as a neural network with many hidden layers. There is no special metric for "many" here. Based on the position of different layers in DNN, the neural network inside DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since there are many DNN layers, the coefficient W and the offset vector The definition of these parameters in DNN is as follows: Take the coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscript corresponds to the output of the third layer index 2 and the input of the second layer index 4. In summary, the coefficient from the kth neuron in the L-1th layer to the jth neuron in the Lth layer is defined as It's important to note that the input layer has no W parameter. In deep neural networks, more hidden layers allow the network to better capture complex real-world situations. Theoretically, a model with more parameters has higher complexity and greater "capacity," meaning it can handle more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrices for all layers of a trained deep neural network (a weight matrix formed by the vectors W across many layers).

[0186] (4) Loss function

[0187] During the training of a deep neural network, because we want the output of the deep neural network to be as close as possible to the desired predicted value, we can compare the current network's predicted value with the desired target value and then update the weight vector of each layer of the neural network based on the difference between the two. (Of course, before the first update, there is usually an initialization process, which pre-configures the parameters for each layer in the deep neural network.) For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value. This adjustment is continued until the deep neural network can predict the desired target value or a value very close to the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value." This is the loss function (or objective function), which is an important equation used to measure the difference between the predicted value and the target value. For example, the loss function output value (loss) indicates a greater difference, so training a deep neural network becomes a process of minimizing this loss.

[0188] (5) Backpropagation algorithm

[0189] The back propagation (BP) algorithm can be used to correct the size of the initial model parameters during training, reducing the model's error loss. Specifically, forward propagation of the input signal to the output generates error loss. This error loss information is then backpropagated to update the parameters in the initial model, thereby converging the error loss. The BP algorithm is a backward propagation movement driven by error loss, aiming to obtain optimal model parameters, such as the weight matrix.

[0190] (6) Diffusion model

[0191] A generative model used to generate data, such as images and text. The core idea of ​​the diffusion model is to diffuse noise into the data and then gradually remove the noise to restore the original data. The diffusion model consists of two stages: the forward process (noise diffusion) and the reverse process (noise removal and restoration).

[0192] Image enhancement technology aims to convert low-quality images into high-quality images. For example, image super-resolution technology is used to convert low-resolution images into high-resolution images, thereby improving the clarity and detail authenticity of the image.

[0193] Existing image super-resolution techniques typically fit a mapping between a low-resolution image and a high-resolution image. To better restore details, existing methods use pre-trained generative models.

[0194] However, existing methods usually do not understand the content of the input image, resulting in low quality of generated images.

[0195] In order to solve the above problems, the present invention provides a data processing method. The data processing method of the present invention is described in detail below with reference to the accompanying drawings.

[0196] Refer to Figure 5, which is a flow chart of a data processing method provided in an embodiment of the present application. As shown in Figure 5, a data processing method provided in an embodiment of the present application may include steps 501 to 503, and these steps are described in detail below.

[0197] 501. Acquire a first image.

[0198] In an embodiment of the present application, a target task may be performed on the first image, where the target task is a task of generating a new image based on the original image, and a new image needs to be generated in the task.

[0199] For example, the target task is one of the following: image enhancement, image noise addition, image style conversion, and image editing. Taking image enhancement as an example, image enhancement can include, but is not limited to, denoising, deblurring, rain removal, defogging, low-light enhancement, contrast enhancement, and high dynamic range.

[0200] 502. Generate target information based on the first image, where the target information is used to indicate texture features or semantic features of the first image.

[0201] In existing technologies, when performing image enhancement, such as super-resolution processing, image noise and the image itself are fed into a generative network to produce a super-resolved image. However, generative networks often lack understanding of the input image content and are unable to use common sense to accurately restore the semantics and textures of objects in the image. This results in poor quality images, especially in real-world scenarios.

[0202] Among them, texture is a visual feature that reflects homogeneous phenomena in an image. It reflects the surface structural organization and arrangement properties of an object with slow or periodic changes. Texture is expressed through the grayscale distribution of pixels and their surrounding spatial neighborhood, which is local texture information. In addition, local texture information is different from image features such as grayscale and color to varying degrees, and its repeatability is global texture information. While texture features reflect the properties of global features, they also describe the surface properties of the scene corresponding to the image or image region. However, since texture is only a characteristic of the surface of an object and cannot fully reflect the essential properties of the object, it is impossible to obtain high-level image content using only texture features. Unlike color features, texture features are not pixel-based features. They need to be statistically calculated in an area containing multiple pixels.

[0203] Image semantics refers to the meaning of the image's content. It's not limited to descriptions in natural language terms; it encompasses how the human visual system understands images. For example, the semantics of an image of a puppy might be interpreted as containing the natural language word "puppy," along with attributes like the puppy's breed and gender. Alternatively, the concept could be represented by a symbol that refers to a puppy with the same characteristics as the one in the image. Furthermore, people can derive an abstract impression from the image, which helps them connect it to known images when they see other similar puppy images.

[0204] In this application, by determining target information related to texture features or semantic features in the first image and inputting this target information into the generative model as a priori, the generative model can understand the content and texture of the input image, thereby better restoring the texture and semantic details in the generated image, thereby obtaining a higher-quality generated image. This can also improve the generalization of the model in different scenarios.

[0205] In one possible implementation, when generating target information based on the first image, target information in the first image input by a user can be received. In other words, the structure and texture of the object can be better restored based on the prompt information provided by the user, thereby improving image fidelity and quality.

[0206] In a possible implementation, the target information may include an encoding of texture features of the first image.

[0207] For example, the first image may be processed by an encoder (eg, an image encoder) to obtain an encoding of texture features of the first image.

[0208] In a possible implementation, the target information may include an encoding of semantic features of the first image.

[0209] For example, the first image may be processed by an encoder (eg, a text encoder) to obtain an encoding of texture features of the first image.

[0210] In a possible implementation, the target information may include an encoding of the first image that is a mixture of the texture feature and the semantic feature.

[0211] For example, the first image can be processed by an encoder (for example, a text encoder and an image encoder) to obtain the encoding of the texture features and the encoding of the semantic features of the first image respectively, and then the encoding of the texture features and the encoding of the semantic features of the first image are fused using a modality adaptation module to obtain an encoding that is a mixture of the texture features and the semantic features.

[0212] In a possible implementation, the target information may include a natural language description of texture features of the first image.

[0213] For example, the first image can be processed by a machine learning model, and the machine learning model can be trained to have the ability to obtain a natural language description of texture features in the image based on the image.

[0214] In a possible implementation, the target information may include a natural language description of semantic features of the first image.

[0215] For example, the first image can be processed by a machine learning model, and the machine learning model can be trained to have the ability to obtain a natural language description of the semantic features in the image based on the image.

[0216] In a possible implementation, the target information may include a natural language description of the first image that is a mixture of the texture features and the semantic features.

[0217] For example, the first image can be processed by a machine learning model, and the machine learning model can be trained to have the ability to obtain a natural language description of texture features and semantic features in the image based on the image.

[0218] In addition, in addition to the encoding or natural language description introduced above, the semantic features can also be semantic segmentation maps, semantic labels of images, main areas of images, classification labels of images, etc., which are not limited in the embodiments of this application.

[0219] In one possible implementation, when generating target information based on the first image, the target information in the first image can be identified based on the first image. For example, the target information in the first image can be obtained by processing the first image using a pre-trained machine learning model.

[0220] For example, in one possible implementation, an encoding of the texture features of the first image can be obtained through an image encoder based on the first image; or, an encoding of the semantic features of the first image can be obtained through a text encoder based on the text description of the first image; or, an encoding mixed with the texture features and the semantic features can be obtained through a modality adaptation module based on the encoding of the texture features and the encoding of the semantic features.

[0221] Next, we introduce the network used to obtain the encoding of semantic features and texture features in the embodiment of the present application, as well as its training process:

[0222] 6 , which is a schematic diagram of a model of an embodiment of the present application, wherein the image content understanding and generation module may be the network for obtaining the encoding of semantic features and the encoding of texture features in the above embodiment.

[0223] Taking image enhancement as an example, the structure of the image content understanding module can be depicted in the dashed box in Figure 6. The input is a low-resolution image, and the network architecture consists of two branches. The first branch uses a visual encoder to encode the texture features of the low-resolution image. The second branch uses a text description generator to generate a text description of the image, and then uses a text encoder to encode the semantic features of the low-resolution image. The encoding results of the two branches are input to the modality adapter, which outputs a fused texture and semantic encoding. In practical applications, the two branches can be used separately or simultaneously. This condition helps the generative model understand low-resolution images and better recover the structure and details of objects in the images.

[0224] During training, high-definition images are sequentially passed through a text description generator and a text encoder to generate corresponding text encodings. Because valid text encodings vary in length, only the valid portion is used as the supervision target, a process known as dynamic supervision.

[0225] In addition, in practical applications, the text description generator of the second branch can be replaced with a user-specified text description, thus supporting user interaction.

[0226] As shown in FIG7 , after introducing image content understanding, the shape, structure, texture, and details of the generated chicken are better restored in the embodiment of the present application.

[0227] In one possible implementation, the target information includes an image having texture features or semantic features of the first image and being different from the first image. Different from the first image may be understood as meaning that the content of the image is not completely identical, for example, pixel values ​​of parts or corresponding locations are different.

[0228] Through the above approach, another image (also referred to as a reference image in this application) with the texture features or semantic features of the first image is used as prior information for the generative model and input into the generative model in the form of an image. This explicitly guides the generation of texture details and helps the generative model understand the content and texture of the input image, thereby better restoring the texture and semantic details in the generated image, resulting in a higher-quality generated image. This can also improve the generalization of the model in different scenarios.

[0229] In one possible implementation, an image can be generated based on the encoding of the texture features and / or semantic features of the first image through a generative model (e.g., a Vincent graph model). The generated image can have the texture features or semantic features of the first image and be different from the first image. Although the input is the encoding of the semantic features or texture features of the first image, the generative network can often generate images with richer textures and more accurate semantics. Therefore, the image obtained by the generative network based on the encoding of the texture features or semantic features of the first image can have richer and more accurate texture details and semantics than the first image, and thus the quality of the image generated based on the high-resolution image is higher.

[0230] In one possible implementation, the image having the texture features or semantic features of the first image can be obtained through a second generative model based on the encoding or natural language description of the texture features of the first image, or the encoding or natural language description of the semantic features of the first image.

[0231] In one possible implementation, an image having the texture features or semantic features of the first image can be selected from multiple images based on an encoding or natural language description of the texture features of the first image, or an encoding or natural language description of the semantic features of the first image. There are various ways to obtain a reference image, such as extracting features of the first image (e.g., an encoding or natural language description of the texture features, or an encoding or natural language description of the semantic features of the first image) and matching the most similar image from an image library.

[0232] A reference image generation module based on content understanding is introduced to automatically generate reference images that are semantically and texture-aligned with low-resolution images, explicitly guiding detail generation. This module generates high-quality reference images that are semantically aligned and texture-consistent with low-resolution images, guiding the model to generate better texture details and improving model generalization.

[0233] Refer to Figure 8, which illustrates a process for generating a reference image based on image content understanding. Specifically, the texture semantic encoding results of a trained image content understanding module are input into a text-based graph model to generate a reference image that is semantically aligned and texture-consistent with the low-resolution image. This is then used to explicitly guide texture detail generation. In practice, this module can be replaced with a user-specified reference image, thus supporting user interaction.

[0234] As shown in Figure 9A, the reference image generated by this embodiment of the present application based on image content understanding is more consistent with the semantics and texture of the low-definition image than a reference image generated purely from text, which helps guide the restoration of the low-definition image. As shown in Figure 9A, after introducing the reference image guidance, the present invention can produce results with rich texture details and greater authenticity.

[0235] In addition, the target information can also indicate the direction of quality improvement of the generation results of the first generation model and be input into the generation model as prior information, thereby helping the generation model to be better optimized and improve the quality of the generated image.

[0236] In one possible implementation, the first generation model is used to perform a target task, and the target information is also used to indicate the features that the first generation model needs to possess in order to obtain a high-quality result after performing the target task on the first image, and the required features are not possessed by the first image.

[0237] The quality improvement direction can be: based on the first image, what kind of generated image, obtained by the first generative model, is a higher-quality generated image? The quality improvement direction can describe the image quality improvement direction at a fine-grained level. Optionally, this information can be input into the generative model as text. For example, in an image enhancement task, for a low-resolution image of a cat, the target information may include: "a cat with clear fur."

[0238] In one possible implementation, the first image can be processed through a machine learning model to obtain information on the direction of quality improvement (that is, the characteristics required in the above embodiment to indicate the high-quality results obtained by the first generation model after performing the target task on the first image).

[0239] The machine learning model used to obtain information on quality improvement directions can be referred to as a multimodal reasoning module. Referring to Figure 9B , the structure of the modal reasoning module is shown within the dashed box. The input is a low-resolution image. After passing through the multimodal model, the module outputs a text description of the image quality improvement directions, guiding the generative model to improve image quality. During training, high-resolution images pass through the text description generator and the text enricher, generating a text description of the image quality improvement directions as a supervision target.

[0240] In addition, in practical applications, the multimodal model can be replaced with a text description of the image quality improvement direction specified by the user, thus supporting user interaction.

[0241] As shown in FIG9C , after the present invention introduces multimodal reasoning, it clarifies the optimization direction of “clear and smooth hair”, and obtains results with richer texture details and better authenticity.

[0242] 503. Obtain a second image through a first generation model according to the first image and the target information, where the second image has the texture feature or semantic feature.

[0243] In an embodiment of the present application, the target information and the first image can be input into the first generation model as generation conditions, and the second image is obtained through the first generation model, and the second image has the texture feature or semantic feature.

[0244] Taking the image enhancement task as an example, an application architecture of an embodiment of the present application can be shown in FIG9D , where the input of the generation model can include noise (z T , condition 1) and low-definition image (condition 2) input, and can support one or more conditions of semantic texture coding or description (condition 3), quality improvement direction description (condition 4), and reference image (condition 5).

[0245] In a possible implementation, the first generation model is a diffusion model, and the target information and the first image can be used as conditional inputs of a noise adding module in the diffusion model to obtain a second image through the diffusion model.

[0246] In the diffusion model architecture, the denoising module of the diffusion model performs multiple noise additions (i.e., noise additions at multiple steps) on the original image (or the feature representation of the original image). The denoising module of the diffusion model can predict the noise added at each step and perform denoising based on the predicted noise to generate a new image.

[0247] In the embodiment of the present application, the image and target information are used as conditions, and the above conditions can be input into the noise addition module of the diffusion model. Optionally, the denoising model can be a neural network with a UNet architecture. Unet is an end-to-end convolutional neural network mainly used for image segmentation. It is characterized by a U-shaped structure, which includes an encoder and a decoder. The encoder uses convolutional layers and maximum pooling layers to achieve feature extraction and downsampling; the decoder performs feature fusion and recovery through upsampling and jump connections. This architecture is mainly used to predict the noise added to the image and text feature forward in the diffusion framework. Specifically, it can be based on the β t The variance given is gradually added to the sample Gaussian noise to produce a series of latent variables x1,x2,…,x T , where t is the number of steps to add noise (or it can be called step length), and N is the normal distribution:

[0248] Then learn a denoising module Φ θ (x t ,t) (that is, the denoising module in the diffusion model) to approximate the true posterior distribution, where u θ and Σ θ are the mean and variance predicted by the model:

[0249] p θ (x t-1 |x t )=N(u θ (x t ,t),Σ θ (x t ,t));

[0250] In order to perform the inverse denoising process, we sample from random noise x T Starting from N(0,I), the noise is gradually reduced and finally a sample x0 is obtained. The existing method uses a denoising network Φ θ (x t ,t), which predicts the noise component of the noise sample and the following training objectives:

[0251] Among them, t can be uniformly sampled from 1, 2, …, T, and ε is the noise added to each step of the model.

[0252] It should be understood that the generation model in the embodiment of the present application can also be other model architectures besides the diffusion model, and the present application is not limited thereto.

[0253] Next, we will introduce the beneficial effects of the embodiments of this application with experimental results. As shown in Table 1, on the RealSR and DRealSR public datasets, the understanding and reasoning-based image super-resolution method proposed in this application outperforms existing optimal algorithms, including Real-ESRGAN+ based on a generative adversarial network and StableSR based on a diffusion model. At the same time, this invention significantly improves the model's understanding, reasoning, and generalization capabilities.

[0254] Table 1

[0255] 10 , which is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. As shown in FIG10 , a data processing device provided in an embodiment of the present application, the device 1000 includes:

[0256] An acquisition module 1001 is configured to acquire a first image;

[0257] For a detailed description of the acquisition module 1001 , reference may be made to the description of step 501 in the above embodiment, which will not be repeated here.

[0258] Processing module 1002 is used to generate target information based on the first image, and the target information is used to indicate the texture features or semantic features of the first image; according to the first image and the target information, a second image is obtained through a first generation model, and the second image has the texture features or semantic features.

[0259] For a detailed description of the processing module 1002 , reference may be made to the description of steps 502 to 503 in the above embodiment, which will not be repeated here.

[0260] In one possible implementation, the first generation model is used to perform a target task, and the target task is one of the following:

[0261] Image enhancement, image noise addition, image style transfer, image editing.

[0262] In a possible implementation, the processing module 1002 is specifically configured to:

[0263] According to the first image, identifying and obtaining target information in the first image; or,

[0264] Receive target information in the first image input by a user.

[0265] In one possible implementation, the target information includes at least one of the following:

[0266] An encoding of the texture features of the first image, or an encoding of the semantic features, or an encoding mixed with the texture features and the semantic features, or a natural language description of the texture features of the first image, or a natural language description of the semantic features, or a natural language description mixed with the texture features and the semantic features.

[0267] In a possible implementation, the processing module 1002 is specifically configured to:

[0268] Obtaining, according to the first image, an encoding of texture features of the first image through an image encoder; or,

[0269] Obtaining an encoding of the semantic features of the first image through a text encoder according to the text description of the first image; or

[0270] According to the coding of the texture feature and the coding of the semantic feature, a coding mixed with the texture feature and the semantic feature is obtained through a modality adaptation module.

[0271] In a possible implementation, the target information includes an image that has texture features or semantic features of the first image and is different from the first image.

[0272] In a possible implementation, the processing module 1002 is specifically configured to:

[0273] Obtaining the image having the texture features or semantic features of the first image through a second generative model according to the encoding or natural language description of the texture features of the first image, or the encoding or natural language description of the semantic features of the first image; or

[0274] According to the encoding or natural language description of the texture feature of the first image, or the encoding or natural language description of the semantic feature of the first image, an image having the texture feature or semantic feature of the first image is selected from multiple images.

[0275] In one possible implementation, the first generation model is used to perform a target task, and the target information is also used to indicate the features that the first generation model needs to possess in order to obtain a high-quality result after performing the target task on the first image, and the required features are not possessed by the first image.

[0276] In a possible implementation, the first generation model is a diffusion model, and the processing module 1002 is specifically configured to:

[0277] The target information and the first image are used as conditional inputs of a noise adding module in the diffusion model, and a second image is obtained through the diffusion model.

[0278] Next, a terminal device provided in an embodiment of the present application is introduced. Please refer to Figure 11. Figure 11 is a structural diagram of a terminal device provided in an embodiment of the present application. The terminal device 1100 can be specifically manifested as a virtual reality VR device, a mobile phone, a tablet, a laptop computer, a smart wearable device, etc., which is not limited here. Specifically, the terminal device 1100 includes: a receiver 1101, a transmitter 1102, a processor 1103 and a memory 1104 (wherein the number of processors 1103 in the terminal device 1100 can be one or more, and Figure 11 takes one processor as an example), wherein the processor 1103 may include an application processor 11031 and a communication processor 11032. In some embodiments of the present application, the receiver 1101, the transmitter 1102, the processor 1103 and the memory 1104 may be connected via a bus or other means.

[0279] The memory 1104 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1103. A portion of the memory 1104 may also include non-volatile random access memory (NVRAM). The memory 1104 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.

[0280] Processor 1103 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all of these buses are referred to as a bus system in the figure.

[0281] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1103. Processor 1103 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1103. The above processor 1103 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1103 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly implemented as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 1104. Processor 1103 reads information from memory 1104 and, in conjunction with its hardware, completes the steps involved in the model training or model inference process in the above method.

[0282] Receiver 1101 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 1102 can be used to output digital or character information through the first interface. Transmitter 1102 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1102 can also include a display device such as a display screen.

[0283] The embodiment of the present application also provides a server. Please refer to Figure 12. Figure 12 is a schematic diagram of the structure of the server provided in the embodiment of the present application. The server 1200 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 1212 (for example, one or more processors) and a memory 1232, and one or more storage media 1230 (for example, one or more mass storage devices) for storing application programs 1242 or data 1244. Among them, the memory 1232 and the storage medium 1230 can be temporary storage or permanent storage. The program stored in the storage medium 1230 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1212 can be configured to communicate with the storage medium 1230 to execute a series of instruction operations in the storage medium 1230 on the server 1200.

[0284] The server 1200 may also include one or more power supplies 1226, one or more wired or wireless network interfaces 1250, one or more input and output interfaces 1258; or one or more operating systems 1241, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0285] In an embodiment of the present application, the central processing unit 1212 is used to execute actions related to model training or model reasoning in the above embodiments.

[0286] An embodiment of the present application also provides a computer program product, which, when running on a computer, enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.

[0287] A computer-readable storage medium is also provided in an embodiment of the present application, which stores a program for signal processing. When the computer-readable storage medium is run on a computer, it enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.

[0288] The execution device, training device or terminal device provided in the embodiments of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the data processing method described in the above embodiment, or so that the chip in the training device executes the data processing method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0289] Specifically, see Figure 13 , which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip may be a neural network processor (NPU) 1300. NPU 1300 is mounted on a host CPU (host CPU) as a coprocessor, with tasks assigned by the host CPU. The core of the NPU is arithmetic circuit 1303, which is controlled by controller 1304 to extract matrix data from memory and perform multiplication operations.

[0290] In some implementations, the arithmetic circuit 1303 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1303 is a two-dimensional systolic array. The arithmetic circuit 1303 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1303 is a general-purpose matrix processor.

[0291] For example, assume there are input matrix A, weight matrix B, and output matrix C. The computation circuit retrieves the corresponding data of matrix B from weight memory 1302 and caches it on each PE in the computation circuit. The computation circuit then retrieves the data of matrix A from input memory 1301 and performs a matrix operation on it with matrix B. The partial or final matrix result is stored in accumulator 1308.

[0292] Unified memory 1306 is used to store input and output data. Weight data is directly transferred to weight memory 1302 through the Direct Memory Access Controller (DMAC) 1305. Input data is also transferred to unified memory 1306 through the DMAC.

[0293] BIU stands for Bus Interface Unit 1310 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1309 .

[0294] The bus interface unit 1310 (BIU) is used for the instruction fetch memory 1309 to obtain instructions from the external memory, and is also used for the storage unit access controller 1305 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0295] DMAC is mainly used to move input data in the external memory DDR to the unified memory 1306 or to move weight data to the weight memory 1302 or to move input data to the input memory 1301.

[0296] The vector calculation unit 1307 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1303, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0297] In some implementations, the vector calculation unit 1307 can store the processed output vector in the unified memory 1306. For example, the vector calculation unit 1307 can apply a linear function or a nonlinear function to the output of the operation circuit 1303, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, the vector calculation unit 1307 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1303, for example, for use in subsequent layers in a neural network.

[0298] An instruction fetch buffer 1309 connected to the controller 1304 is used to store instructions used by the controller 1304;

[0299] Unified memory 1306, input memory 1301, weight memory 1302, and instruction fetch memory 1309 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0300] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0301] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0302] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.

[0303] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0304] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

Claims

1. An image processing method, characterized in that: The method comprises: Get the first image, generating target information based on the first image, the target information including first text and / or a second image, the first text being used to indicate an optimization direction for image quality of the first image, the image quality of the second image being higher than that of the first image, and the semantics of the second image corresponding to those of the first image; Based on the first image and the target information, a third image is obtained through a first generation model. The image quality of the third image is higher than that of the first image, and the content of the third image is consistent with that of the first image.

2. The method according to claim 1, characterized in that The first generative model is used to perform an image enhancement task.

3. The method according to claim 1 or 2, characterized in that The generating target information based on the first image includes: Based on the first image, target information is generated through a machine learning model.

4. The method according to any one of claims 1 to 3, characterized in that The image quality is at least one of the following clarity: PSNR, SSIM, LPIPS, DISTIS, NIQE, MANIQA, MUSIQ, CLIPIQA.

5. The method according to any one of claims 1 to 4, characterized in that: The target information is also used to indicate a texture feature or a semantic feature of the first image.

6. The method according to claim 5, characterized in that The target information also includes at least one of the following: An encoding of the texture features of the first image, or an encoding of the semantic features, or an encoding mixed with the texture features and the semantic features, or a natural language description of the texture features of the first image, or a natural language description of the semantic features, or a natural language description mixed with the texture features and the semantic features.

7. The method according to claim 5 or 6, characterized in that The generating target information based on the first image includes: Obtaining, according to the first image, an encoding of texture features of the first image through an image encoder; or, Obtaining an encoding of the semantic features of the first image through a text encoder according to the text description of the first image; or According to the coding of the texture feature and the coding of the semantic feature, a coding mixed with the texture feature and the semantic feature is obtained through a modality adaptation module.

8. The method according to any one of claims 5 to 7, characterized in that: The generating target information based on the first image includes: Obtaining the image having the texture features or semantic features of the first image through a second generative model according to the encoding or natural language description of the texture features of the first image, or the encoding or natural language description of the semantic features of the first image; or According to the encoding or natural language description of the texture features of the first image, or the encoding or natural language description of the semantic features of the first image, an image having the texture features or semantic features of the first image is selected from multiple images.

9. The method according to any one of claims 1 to 8, characterized in that: The first generation model is a diffusion model, and obtaining the second image through the first generation model according to the first image and the target information includes: The target information and the first image are used as conditional inputs of a noise adding module in the diffusion model, and a second image is obtained through the diffusion model.

10. A data processing device, characterized in that: The device comprises: An acquisition module, configured to acquire a first image; A processing module is used to generate target information based on the first image, where the target information includes a first text and / or a second image, where the first text is used to indicate an optimization direction for the image quality of the first image, the image quality of the second image is higher than that of the first image, and the semantics of the second image correspond to those of the first image; and based on the first image and the target information, a third image is obtained through a first generation model, where the image quality of the third image is higher than that of the first image, and the content of the third image is consistent with that of the first image.

11. The device according to claim 10, characterized in that The first generative model is used to perform an image enhancement task.

12. The device according to claim 10 or 11, characterized in that The processing module is specifically used to: The target information is obtained according to the first image through a machine learning model.

13. The device according to any one of claims 10 to 12, characterized in that The image quality is at least one of the following clarity: PSNR, SSIM, LPIPS, DISTIS, NIQE, MANIQA, MUSIQ, CLIPIQA.

14. The device according to any one of claims 10 to 13, characterized in that The target information is also used to indicate a texture feature or a semantic feature of the first image.

15. The device according to claim 14, characterized in that The target information also includes at least one of the following: An encoding of the texture features of the first image, or an encoding of the semantic features, or an encoding mixed with the texture features and the semantic features, or a natural language description of the texture features of the first image, or a natural language description of the semantic features, or a natural language description mixed with the texture features and the semantic features.

16. The device according to claim 14 or 15, characterized in that The processing module is specifically used to: Obtaining, according to the first image, an encoding of texture features of the first image through an image encoder; or, Obtaining, by a text encoder, an encoding of semantic features of the first image according to the text description of the first image; or, According to the coding of the texture feature and the coding of the semantic feature, a coding mixed with the texture feature and the semantic feature is obtained through a modality adaptation module.

17. The device according to any one of claims 14 to 16, characterized in that The processing module is specifically used to: Obtaining the image having the texture features or semantic features of the first image through a second generative model according to the encoding or natural language description of the texture features of the first image, or the encoding or natural language description of the semantic features of the first image; or According to the encoding or natural language description of the texture features of the first image, or the encoding or natural language description of the semantic features of the first image, an image having the texture features or semantic features of the first image is selected from multiple images.

18. The device according to any one of claims 10 to 17, characterized in that The first generation model is a diffusion model, and the processing module is specifically configured to: The target information and the first image are used as conditional inputs of a noise adding module in the diffusion model, and a second image is obtained through the diffusion model.

19. A computer storage medium, characterized in that The computer storage medium stores one or more instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the method of any one of claims 1 to 9.

20. A computer program product, characterized in that The method comprises computer-readable instructions, which, when executed on a computer device, cause the computer device to execute the method according to any one of claims 1 to 9.

21. A system comprising at least one processor and at least one memory; the processor and the memory are connected via a communication bus and communicate with each other; The at least one memory is used to store code; The at least one processor is configured to execute the code to perform the method according to any one of claims 1 to 9.

22. A chip comprising a processor, characterized in that: The processor is used to support a data processing device to implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Data processing method and device

    CN120430987A

  • Image generation method and terminal device

    CN110136216A

  • Text-guided natural image super-resolution reconstruction method

    CN114494007A

  • Underwater image enhancement method and device for reinforcement learning parameter optimization, and medium

    CN115423724A

  • Image generation processing method and device, electronic equipment and storage medium

    CN116012481A