Image Processing Method, Apparatus, Electronic Device, and Computer-Readable Storage Medium

Through data augmentation, a similarity annotation result of the first training data set and a second training data set manually marked are generated, and the neural network model is trained, which solves the problem of limited performance of the image similarity model, and achieves the improvement of model performance and the enhancement of generalization capabilities.

CN113704531BActive Publication Date: 2025-07-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110261801.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-10
Publication Date
2025-07-29
Estimated Expiration
2041-03-10

AI Technical Summary

Technical Problem

In the prior art, the performance of the image similarity model is limited by the number of artificially labeled sample image pairs, resulting in limited model performance and waste of manpower.

Method used

The similarity labeling result of the first training data set is generated by the data augmentation method, and combined with the second training data set manually annotated, the neural network model is trained to improve the model performance.

Benefits of technology

A large number of training samples with similarity annotation results were generated, which improved the performance and generalization capabilities of the image similarity model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113704531B_ABST
    Figure CN113704531B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method, apparatus, electronic device, and computer-readable storage medium, relating to the fields of artificial intelligence, cloud technology, and image processing technology. The method includes: obtaining a first training data set and a second training data set, pre-training an initial neural network model based on the first training data set to obtain a pre-trained neural network model; and training the pre-trained neural network model based on the second training data set to obtain an image similarity model. According to the method of the present application, since the first training data set is automatically determined by performing data augmentation on each initial image, a large number of training samples with similarity annotation results can be generated, providing data support for the training of the model. Further, since the manually annotated second sample image set and its similarity annotation results are more accurate, the performance of the image similarity model trained based on the second training data set is better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of artificial intelligence, big data processing, and image processing. Specifically, this application relates to an image processing method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] Image similarity is widely used in common image processing application scenarios such as image retrieval and image recognition. In practical applications, the similarity between two images can be determined based on a trained image similarity model. In order to obtain an image similarity model that can determine the similarity between two images, in the prior art, the above image similarity model is usually trained based on a fully supervised learning method, that is, the above image similarity model is trained based on a large number of manually labeled image pairs. In order to obtain a model with higher accuracy, a large number of sample image pairs usually need to be manually labeled. Therefore, the number of sample image pairs using the manual labeling method is relatively limited, resulting in limited performance of the image similarity model and wasting a lot of manpower. Summary of the Invention

[0003] The purpose of this application aims to solve at least one of the above technical defects, and specifically provides the following technical solutions to solve the problem of improving the performance of the image similarity model.

[0004] According to one aspect of this application, an image processing method is provided, and the method includes:

[0005] Obtain a first training data set and a second training data set. The first training data set includes multiple first sample image sets, and the second training data set includes multiple second sample image sets. Among them, the first sample image set and its similarity annotation result are determined by performing data augmentation on each initial image, and the similarity annotation result of the second sample image set is a manually annotated result;

[0006] Pre-train an initial neural network model based on the first training data set to obtain a pre-trained neural network model;

[0007] Train the pre-trained neural network model based on the second training data set to obtain an image similarity model, so as to determine the similarity of an image pair through the image similarity model.

[0008] According to another aspect of this application, an image processing method is provided, and the method includes:

[0009] Obtain at least two images to be processed;

[0010] Process at least two pairs of images to be processed by calling the image similarity model to obtain the similarity of each image pair in the at least two images to be processed, so as to process the at least two images to be processed based on the similarity;

[0011] Among them, the image similarity model is obtained by the method shown in the first aspect of this application.

[0012] According to another aspect of this application, an image processing device is provided, and the device includes:

[0013] A training data acquisition module, configured to acquire a first training data set and a second training data set. The first training data set includes multiple first sample image sets, and the second training data set includes multiple second sample image sets. Among them, the first sample image set and its similarity annotation result are determined by performing data augmentation on each initial image, and the similarity annotation result of the second sample image set is an artificially annotated result;

[0014] A model training module, configured to pre-train an initial neural network model based on the first training data set to obtain a pre-trained neural network model, and train the pre-trained neural network model based on the second training data set to obtain an image similarity model, so as to determine the similarity of an image pair through the image similarity model.

[0015] In a possible implementation manner, when the training data acquisition module acquires the first training data set, it is specifically configured to:

[0016] Acquire multiple initial images;

[0017] For each initial image, perform data augmentation processing on the initial image to obtain at least two sub-images corresponding to the initial image;

[0018] Based on two sub-images belonging to the same initial image among the sub-images of each initial image, obtain multiple first positive sample image sets and the similarity annotation result of the first positive sample image sets;

[0019] Based on two sub-images belonging to different initial images among the sub-images of each initial image, obtain multiple first negative sample image sets and the similarity annotation result of the first negative sample image sets;

[0020] Among them, the multiple first sample image sets include multiple first positive sample image sets and multiple first negative sample image sets.

[0021] In a possible implementation manner, the second sample image set is a positive sample image set, and the second training data set further includes multiple first negative sample image sets.

[0022] In a possible implementation manner, the multiple second sample image sets include at least one of multiple second positive sample image sets or multiple second negative sample image sets. The similarity of the second positive sample image set is less than or equal to a first threshold; the similarity of the second negative sample image set is greater than or equal to a second threshold;

[0023] Among them, the similarity of the second sample image set is determined by the pre-trained neural network model.

[0024] In a possible implementation, the data augmentation process includes at least one of the following:

[0025] Image cropping; Smearing processing; Blurring processing; Color transformation; Grayscale transformation; Image rotation; Image flipping.

[0026] In a possible implementation, when the model training module pre-trains the initial neural network model based on the first training data set to obtain the pre-trained neural network model, it is specifically used for:

[0027] Repeatedly execute the following training steps until the pre-training loss value meets the pre-training end condition to obtain the pre-trained neural network model:

[0028] Input each first sample image set into the initial neural network model, extract the image features of two images in each first sample image set through the initial neural network model respectively, and predict the predicted similarity of the first sample image set based on the image features of the two images;

[0029] Determine the pre-training loss value according to the predicted similarity of each first sample image set and the similarity annotation result;

[0030] If the pre-training loss value meets the pre-training end condition, end the pre-training; if not, adjust the model parameters of the initial neural network model and repeat the training steps.

[0031] According to another aspect of the present application, an image processing device is provided, and the device includes:

[0032] An image acquisition module, configured to acquire at least two images to be processed;

[0033] An image processing module, configured to process at least two pairs of images to be processed by invoking an image similarity model to obtain the similarity of each pair of images in the at least two images to be processed, so as to process the at least two images to be processed based on the similarity;

[0034] Among them, the image similarity model is obtained by the method shown in the first aspect of the present application.

[0035] In a possible implementation, when the image acquisition module acquires at least two images to be processed, it is specifically used for:

[0036] Obtain an image retrieval request, where the image retrieval request includes a retrieved image;

[0037] Obtain the image database corresponding to the image retrieval request. At least two images to be processed include the retrieval image and the retrieved images in the image database. An image pair includes the retrieval image and one retrieved image;

[0038] The device further includes:

[0039] An image retrieval module, configured to determine the target image corresponding to the image retrieval request from the image database according to the similarity of each image pair, and provide the target image to the retriever.

[0040] In a possible implementation manner, when the image acquisition module acquires at least two images to be processed, it is specifically configured to:

[0041] Acquire a set of images to be processed. At least two images to be processed are each image in the set of images to be processed, and an image pair is any two images in the set of images to be processed;

[0042] The device further includes:

[0043] An image classification module, configured to classify the images in the set of images to be processed according to the similarity of each image pair.

[0044] According to another aspect of the present application, there is provided an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the image processing method of the present application is implemented.

[0045] According to yet another aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the image processing method of the present application is implemented.

[0046] The embodiments of the present invention further provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations of the above image processing method.

[0047] The beneficial effects brought by the technical solution provided by the present application are:

[0048] The image processing method, apparatus, electronic device, and computer-readable storage medium provided by the present application. When obtaining an image similarity model for determining the similarity of image pairs, the first sample image set in the first training data set of the model and its similarity annotation results are automatically determined by performing data augmentation on each initial image. Therefore, based on the data augmentation method, a large number of training samples with similarity annotation results can be generated, providing data support for the training of the model. Further, the solution of the embodiment of the present application also provides a second training data set for the model. Since the second sample image set in the second training data set and its similarity annotation results are manually annotated, the manually annotated similarity annotation results are more accurate. Therefore, the performance of the image similarity model trained based on the second training data set is better.

[0049] Additional aspects and advantages of the present application will be given in part in the following description, which will become apparent from the following description, or can be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.

[0051] Figure 1 Schematic flowchart of an image processing method provided by an embodiment of the present application;

[0052] Figure 2 Schematic diagram of an image after smear processing provided by an embodiment of the present application;

[0053] Figure 3 Schematic diagram of a network structure provided by an embodiment of the present application;

[0054] Figure 4 Schematic flowchart of the training process of an image similarity model in an image processing method provided by an embodiment of the present application;

[0055] Figure 5 Schematic flowchart of another image processing method provided by an embodiment of the present application;

[0056] Figure 6 Schematic flowchart of an image processing method provided by an embodiment of the present application;

[0057] Figure 7 Schematic diagram of the implementation environment of an image processing method provided by an embodiment of the present application;

[0058] Figure 8 Schematic diagram of the implementation environment of another image processing method provided by an embodiment of the present application;

[0059] Figure 9 The structural schematic diagram of an image processing device provided by an embodiment of the present application;

[0060] Figure 10 The structural schematic diagram of another image processing device provided by an embodiment of the present application;

[0061] Figure 11 The structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0062] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application.

[0063] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0064] The embodiment of the present application provides an image processing method for improving the efficiency of obtaining training samples and improving the model performance. This method can be applied to any scenario that needs to determine the similarity of image pairs. This method involves artificial intelligence, big data processing and cloud technology, and specifically involves fields such as machine learning and computer vision technology in artificial intelligence technology.

[0065] In one embodiment of the present application, the solution provided in this embodiment can be implemented based on cloud technology. The data processing involved in each optional embodiment (including but not limited to data calculation) can be implemented using cloud computing. Cloud technology refers to a hosting technology that unifies hardware, software, network, and other resources within a wide area network or local area network to enable data calculation, storage, processing, and sharing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology applied in the cloud computing business model. It can form a resource pool that can be used on demand with flexibility and convenience. Cloud computing technology will become a key support. Backend services of technical network systems, such as video websites, image websites, and more portals, require a large amount of computing and storage resources. With the rapid development and application of the internet industry, every item in the future may have its own unique identification and need to be transmitted to backend systems for logical processing. Data of different levels will be processed separately. Data from various industries requires strong system support, which can only be achieved through cloud computing.

[0066] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network that provides these resources is called the "cloud." To users, these resources appear infinitely scalable and can be accessed at any time, used on demand, expanded at any time, and paid for on a per-use basis.

[0067] As a provider of cloud computing infrastructure, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) is established. Various types of virtual resources are deployed within the resource pool for external customers to choose from. The cloud computing resource pool primarily includes computing devices (virtualized machines, including operating systems), storage devices, and network devices. Based on logical functional divisions, the PaaS (Platform as a Service) layer can be deployed on top of the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. SaaS can also be deployed directly on top of IaaS. PaaS is a platform for software running, such as databases and web containers. SaaS is a variety of business software, such as web portals and text message senders. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.

[0068] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0069] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0070] Among them, Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and adversarial learning.

[0071] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0072] Big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time range. It is a vast, high-growth, and diverse information asset that requires new processing models to have stronger decision-making power, insight discovery ability, and process optimization ability. With the advent of the cloud era, big data has also attracted increasing attention. Big data requires special technologies to effectively process large amounts of data that can tolerate elapsed time. Technologies applicable to big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.

[0073] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition and measurement in machine vision, and further performing graphics processing to make the images processed by the computer more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0074] The solution provided by the embodiments of the present application can be executed by any electronic device, which can be a user terminal device or a server. Among them, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can include at least one of the following: smart phones, tablets, laptop computers, desktop computers, smart speakers, smart watches, smart TVs, and intelligent in-vehicle devices.

[0075] The source of the training data required for model training involved in the embodiments of the present application is not limited in this application. It can include existing training data sets, big data obtained from the Internet, and training data obtained through data augmentation methods in the image processing method of the present application.

[0076] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0077] An embodiment of the present application provides a possible implementation manner. As Figure 1 shown, a flowchart of an image processing method is provided. This solution can be executed by any electronic device. For example, the solution of the embodiment of the present application can be executed on a terminal device or a server, or jointly executed by the terminal device and the server. For the convenience of description, hereinafter, the server will be taken as the execution subject as an example to illustrate the method provided by the embodiment of the present application. As Figure 1 shown in the flowchart, the method may include the following steps:

[0078] Step S110, obtain a first training data set and a second training data set. The first training data set includes multiple first sample image sets, and the second training data set includes multiple second sample image sets.

[0079] Among them, the first sample image set and its similarity annotation result are determined by performing data augmentation on each initial image, and the similarity annotation result of the second sample image set is an artificial annotation result.

[0080] Among them, the similarity annotation result characterizes the similarity degree between two images. The specific manner of similarity annotation is not limited in the present application. The first sample image set includes at least two images, and any two images in the at least two images are similar, or are not similar. The second sample image set also includes at least two images, and any two images in the at least two images are similar, or are not similar. Then, the similarity annotation result of the first sample image set characterizes that any two images in the image set are similar, or are not similar, and the similarity annotation result of the second sample image set characterizes that any two images in the image set are similar, or are not similar.

[0081] As an example, the similarity annotation result is represented by the numbers 1 and 0. When the number is 1, it means that the two images are similar. When the number is 0, it means that the two images are not similar.

[0082] Among them, two images (i.e., an image pair) being similar means that the similarity degree of the two images is greater than the similarity threshold. For example, in an example of the present application, the similarity threshold is 0.6. Then, when the similarity degree of the two images is greater than 0.6, the two images are similar. Otherwise, the two images are not similar. It should be noted that the similarity threshold can be configured based on actual requirements. For example, in a scenario with higher similarity requirements, the similarity threshold can be appropriately increased. For example, the similarity threshold is 0.8.

[0083] Data augmentation processing refers to increasing the data volume without changing the image category, that is, on the basis of the initial image, a data volume-rich and diverse image set can be obtained as training data through data augmentation processing.

[0084] Since the first sample image set and its similarity annotation results are determined by data augmentation of each initial image, data augmentation cannot cover image sets of various scenes, and the similarity annotation results of the image set determined by data augmentation may not be completely correct. Therefore, in the present application, manual annotation is used to determine the second training data set, that is, multiple second sample image sets and their similarity annotation results are determined by manual annotation. In this way, by manually annotating a portion of the sample data, the data used for training can cover as many scenes as possible, and the trained image similarity model is more generalized and has better performance.

[0085] Optionally, the second sample image set can be a set of images that are difficult to determine whether they are similar, such as a set of images obtained by re-photographing, a set of images with different layouts, or a set of images obtained by screenshots. Training the image similarity model based on such image sets can further improve the model's generalization ability.

[0086] It is understandable that the initial image can be an image collected manually or an image selected from an image library. In the embodiments of the present application, the source of the initial image is not limited.

[0087] Step S120: pre-training the initial neural network model based on the first training data set to obtain a pre-trained neural network model.

[0088] Step S130 : training the pre-trained neural network model based on the second training data set to obtain an image similarity model, so as to determine the similarity of the image pair through the image similarity model.

[0089] The image similarity model obtained through the above training is used to determine the similarity of an image pair, that is, the similarity between two images in an image set can be determined through the image similarity model.

[0090] As an optional solution, the model architecture of the initial neural network model is not limited in the embodiment of the present application, and can be any initial neural network model that can be used to determine the similarity of image pairs, such as SimCLR (A Simple Framework for Contrastive Learning of Visual Representations) or SimSiam (Simple Siamese, twin) network.

[0091] In the solution of this application, when obtaining an image similarity model for determining the similarity of an image pair, the first sample image set in the first training data set of the model and its similarity annotation results are automatically determined by performing data augmentation on each initial image. Therefore, based on the data augmentation method, a large number of training samples with similarity annotation results can be generated, providing data support for the training of the model. Further, the solution of the embodiment of this application also provides a second training data set for the model. Since the second sample image set in the second training data set and its similarity annotation results are manually annotated, and the manually annotated similarity annotation results are more accurate, the performance of the image similarity model trained based on the second training data set is better.

[0092] In one embodiment of this application, obtaining the first training data set may include:

[0093] Obtain multiple initial images;

[0094] For each initial image, perform data augmentation processing on the initial image to obtain at least two sub-images corresponding to the initial image;

[0095] Based on two sub-images belonging to the same initial image among the sub-images of each initial image, obtain multiple first positive sample image sets and the similarity annotation results of the first positive sample image sets;

[0096] Based on two sub-images belonging to different initial images among the sub-images of each initial image, obtain multiple first negative sample image sets and the similarity annotation results of the first negative sample image sets;

[0097] Among them, the multiple first sample image sets include multiple first positive sample image sets and multiple first negative sample image sets.

[0098] Optionally, the multiple initial images may be multiple images covering as many application scenarios as possible. Different images among the multiple initial images may not be similar, that is, the similarity between any two initial images is relatively low.

[0099] Among them, a sub-image refers to an image derived from an initial image. For example, an image obtained after performing image transformation on the initial image. Taking one initial image as an example, after performing data augmentation processing on the initial image, at least two sub-images are obtained. Each image in the at least two sub-images is similar to the initial image. Then, based on the at least two sub-images corresponding to the initial image set, the first positive sample image set corresponding to the initial image set and the similarity annotation results of the first positive sample image set can be determined, and at least one pair of first positive sample image sets can be determined for each initial image.

[0100] Among the sub - graphs of each initial image, two sub - graphs belonging to different initial images are not similar. As an example, the two initial images are initial image x1 and initial image x2 respectively, and the sub - graph corresponding to initial image x1 and the sub - graph corresponding to initial image x2 are not similar. Then, based on two sub - graphs belonging to different initial images among the sub - graphs of each initial image, multiple first negative sample image sets and the similarity annotation results of the first negative sample image sets can be obtained.

[0101] Among them, since data augmentation processing refers to increasing the amount of data without changing the image category, that is, on the basis of the initial image, only changing the display form of the initial image. For example, changing the size of the initial image, changing the color of the initial image, etc., without changing the image content in the initial image. Therefore, two sub - graphs corresponding to the same initial image are similar images. Therefore, the similarity annotation result of the first positive sample image set is similar, that is, any two images in the first positive sample image set are similar. And the image content between different initial images is different. Therefore, two sub - graphs corresponding to different initial images are not similar. Therefore, the similarity annotation result of the first negative sample image set is not similar, that is, any two images in the first negative sample image set are not similar.

[0102] If multiple first sample image sets include multiple first positive sample image sets, and any two images in each first positive sample image set are similar, then the similarity annotation result of each first positive sample image set is similar. If multiple first sample image sets include multiple first positive sample image sets and multiple negative sample image sets, and any two images in each negative sample image set are not similar, then the similarity annotation result of each negative sample image set is not similar.

[0103] In an alternative solution of the present application, the data augmentation processing includes at least one of the following:

[0104] Image cropping; Smearing processing; Blurring processing; Color transformation; Grayscale transformation; Image rotation; Image flipping.

[0105] Next, based on the above data augmentation processing, how to obtain the first training data set will be further described:

[0106] Obtain N initial images. Optionally, N is greater than or equal to 20000. These N initial images cover various different scenes as much as possible, and these N initial images do not require manual annotation.

[0107] For each of the N initial images, perform data augmentation processing on each initial image to obtain at least two sub - graphs corresponding to each initial image.

[0108] Taking an initial image as an example, a solution for performing different data augmentations on the initial image to obtain at least two corresponding sub - graphs will be described as follows:

[0109] (1) Image cropping

[0110] Perform image cropping on the initial image to obtain the cropped image, and the cropped image can be directly used as a sub - graph of the initial image.

[0111] Among them, when cropping the initial image, the proportion of the cropped sub - graph in the original image (initial image) can be controlled by parameters. Optionally, this proportion can be between 0.2 and  1.

[0112] Since the cropped images may have different sizes, for subsequent convenient processing of the sub - graphs, the cropped images can be resized to a fixed - size image, and then the cropped image with the fixed size is used as a sub - graph of the initial image.

[0113] As an example, the fixed size can be set to 256, so the size of the sub - graph obtained through image cropping processing is 256.

[0114] (2) Smearing processing

[0115] Perform smearing processing on the initial image, and the smeared image is used as the sub - graph corresponding to the initial image. Performing smearing processing on the initial image can enhance the robustness against smearing interference.

[0116] Among them, smearing processing means adding different types of elements to the initial image, such as at least one of an image, text, characters, symbols, special effects, or lines; a sub - graph can contain one element or at least two elements at the same time.

[0117] See Figure 2 the smeared image shown in the figure. Among them, Figure a is the sub - graph obtained by adding line a to the initial image, and Figure b is the sub - graph obtained by adding line b to the initial image. Line a and line b are lines of different colors.

[0118] It should be noted that the text content in Figure a and Figure b is the same, and Figure a is only a schematic diagram of the sub - graph obtained by adding line a, and Figure b is only a schematic diagram of the sub - graph obtained by adding line b to the initial image. The content in the image does not limit this solution.

[0119] (3) Blurring processing

[0120] Perform blurring processing on the initial image, and the blurred image is used as the sub - graph corresponding to the initial image. Performing blurring processing on the initial image can enhance the robustness of similarity learning for blurred images.

[0121] As an alternative, the blurring process can be Gaussian blurring.

[0122] (4) Color transformation

[0123] Perform color transformation on the initial image, and the processed image is used as the sub-image corresponding to the initial image. Performing color transformation on the initial image can enhance the robustness of similarity learning for image illuminance, color, and other transformations.

[0124] Among them, the color transformation process includes but is not limited to brightness transformation, contrast transformation, saturation transformation, and RGB (red green blue) color transformation.

[0125] (5) Grayscale transformation

[0126] Perform grayscale transformation on the initial image, and the processed image is used as the sub-image corresponding to the initial image. When the initial image is a color RGB image, grayscale transformation can be performed on the initial image, and performing grayscale transformation on the initial image can enhance the robustness of similarity learning for grayscale images.

[0127] (6) Image rotation

[0128] Perform image rotation on the initial image, for example, clockwise rotation or counterclockwise rotation, and the processed image is used as the sub-image corresponding to the initial image. Performing image rotation on the initial image can enhance the robustness of similarity learning for the rotated image.

[0129] (7) Image flipping

[0130] Perform image flipping on the initial image, for example, horizontal flipping or vertical flipping, and the processed image is used as the sub-image corresponding to the initial image. Performing image flipping on the initial image can enhance the robustness of similarity learning for the flipped image.

[0131] In an embodiment of the present application, the second sample image set is a positive sample image set, and the second training data set further includes a plurality of first negative sample image sets.

[0132] Among them, the second positive sample image set in the second training data set is determined by manual annotation. The second training data set may further include a negative sample image set, and the negative sample image set may be manually annotated or the first negative sample image set determined by the data augmentation method in the first training data set.

[0133] Using the first negative sample image set in the first training data set as the negative sample image set in the second training data set can increase the data volume of the training data, thereby improving the performance of the model.

[0134] In one embodiment of the present application, the multiple second sample image sets include at least one of multiple second positive sample image sets or multiple second negative sample image sets, the similarity of the second positive sample image set is less than or equal to a first threshold; the similarity of the second negative sample image set is greater than or equal to a second threshold;

[0135] Wherein, the similarity of the second sample image set is determined by a pre-trained neural network model.

[0136] Among them, two images with a similarity less than or equal to the first threshold are actually similar, however, based on the pre-trained neural network model, it is determined that the two images are not similar. Two images with a similarity greater than or equal to the second threshold are actually not similar, but through the pre-trained neural network model, it is determined that the two images are similar. It can be seen from this that for the pre-trained neural network model, the image pairs with a similarity less than or equal to the first threshold and the image pairs with a similarity greater than or equal to the second threshold are image pairs that are not easy to determine whether they are similar. Furthermore, the image pairs with a similarity less than or equal to the first threshold can be used as the second positive sample image set, and the image pairs with a similarity greater than or equal to the second threshold can be used as the second negative sample image set. Training the pre-trained neural network model based on the second sample image set can further improve the performance of the model, so that the model trained based on the second sample image set can accurately judge the similarity of the image pairs with a similarity less than or equal to the first threshold and the image pairs with a similarity greater than or equal to the second threshold.

[0137] In one embodiment of the present application, pre-training the initial neural network model based on the first training data set to obtain the pre-trained neural network model may include:

[0138] Repeatedly execute the following training steps until the pre-training loss value meets the pre-training end condition to obtain the pre-trained neural network model:

[0139] Input each first sample image set into the initial neural network model, extract the image features of two images in each first sample image set through the initial neural network model respectively, and predict the predicted similarity of the first sample image set based on the image features of the two images;

[0140] Determine the pre-training loss value according to the predicted similarity of each first sample image set and the similarity annotation result;

[0141] If the pre-training loss value meets the pre-training end condition, end the pre-training; if not, adjust the model parameters of the initial neural network model and repeat the training steps.

[0142] During the model training process, the initial neural network model can be a model based on a siamese neural network. To further deepen the understanding of the solution of the present application, seeFigure 3 Schematic diagram of the network structure shown. In this example, taking an image pair in a first sample image set as an example, the two images in the image pair are image x1 and image x2 respectively. The initial neural network model includes a first feature extraction layer (as an example, the feature extraction layer can be a convolutional neural network ConvNets), a first fully connected layer fc (fully connected layer), a second feature extraction layer, a second fully connected layer fc, and a classification layer (as an example, the classification layer can also be a fully connected layer).

[0143] Input the image pair into the initial neural network model. Extract the image feature f1 of image x1 through the first feature extraction layer, and extract the image feature f2 of image x2 through the second feature extraction layer. The image feature f1 is input into the first fully connected layer, and the image feature f2 is input into the second fully connected layer. Further process the image features f1 and f2 through the fully connected layer. The output of the first fully connected layer and the output of the second fully connected layer are input into the classification layer, and the predicted similarity of image x1 and image x2 is output through this classification layer.

[0144] Through the same processing process as above, the predicted similarity of each first sample image set can be obtained, that is, whether any two images in the first sample image set are similar or not. Based on the predicted similarity of each first sample image set and the similarity annotation result, determine the pre-training loss value. The pre-training loss value characterizes the difference between the predicted similarity and the similarity annotation result. For the first positive sample image set, the smaller the difference, the closer the predicted similarity between the two images in the first positive sample image set and the similarity annotation result. Conversely, the larger the difference, the greater the difference between the predicted similarity and the similarity annotation result.

[0145] Among them, the similarity annotation result can be identified by a class label. For example, the class label is y, y = 1 means similar, and y = 0 means not similar.

[0146] In an alternative solution of the present application, the predicted similarity of two images can be determined based on the feature distance (such as, Euclidean distance) between the two images. For example, for the images x1 and x2 in the above example, calculate the Euclidean distance d between the image features f1 and f2, as shown in formula (1), and characterize the similarity degree between image x1 and image x2 through the Euclidean distance.

[0147] (1)

[0148] In an alternative solution of this application, the pre-training end condition can be configured according to actual requirements. For example, the pre-training loss value is less than a first set threshold. When the pre-training loss value is less than the first set threshold, it means that the pre-training loss value meets the pre-training end condition, and the pre-training is ended. When the pre-training loss value is not less than the first set threshold, it means that the pre-training loss value does not meet the pre-training end condition, and the model parameters of the initial neural network model need to be adjusted, and the adjusted model is continuously trained based on the training data until the obtained pre-loss value meets the pre-training end condition, and the pre-training is ended.

[0149] In an alternative solution of this application, the pre-training end condition can also be the convergence of the loss function. For example, for the contrastive loss function Contrastive Loss as shown in formula (2), when this loss function converges, the pre-training is ended.

[0150] (2)

[0151] Among them, L is the value of the loss function corresponding to an image pair in a first sample image set (i.e., the training loss value), y represents the similarity annotation result, that is, the class label indicating whether two images in the first sample image set are similar (i.e., the similarity annotation result), y = 1 indicates similarity, and y = 0 indicates dissimilarity; m is a set threshold (the margin value that restricts the feature distance range of negative sample pairs). In this example, m can be set to 1.

[0152] It should be noted that the above loss function is for an image pair in the first sample image set. For multiple image pairs, the loss function of the initial neural network model can be N×S×L, where S is the number of image pairs in a first sample image set, and N is the number of first sample image sets.

[0153] It can be understood that in the solution of this application, training the pre-trained neural network model based on the second training data set to obtain the image similarity model can also be the same as the training process of training the pre-trained neural network model described above.

[0154] Specifically, the following training steps are repeatedly executed until the training loss value meets the training end condition to obtain the image similarity model:

[0155] Input each second sample image set into the pre-trained neural network model, extract the image features of two images in each second sample image set through the pre-trained neural network model, and predict the predicted similarity of the second sample image set based on the image features of the two images;

[0156] Determine the training loss value according to the predicted similarity and the similarity annotation result of each second sample image set;

[0157] If the training loss value meets the training end condition, the training is ended; if not, the model parameters of the neural network model after pre-training are adjusted, and the training steps are repeated.

[0158] Among them, the training end condition can be the same as or different from the pre-training end condition. For example, the training end condition is that the training loss value is less than the second set threshold. The first set threshold and the second set threshold can be the same or different.

[0159] The following combines Figure 4 The image processing method shown below further describes the solution of the present application in detail. The method includes the following steps:

[0160] Step S210, obtain multiple initial images.

[0161] Among them, the initial images are preferably selected to cover various scenes.

[0162] Step S220, for each initial image, perform data augmentation processing on the initial image to obtain at least two sub-images corresponding to the initial image.

[0163] Among them, the data augmentation processing includes at least one of the following: image cropping; smearing processing; blurring processing; color transformation; grayscale transformation; image rotation; image flipping. How to specifically process the initial image based on the above data augmentation processing to obtain at least two sub-images corresponding to the initial image has been described above and will not be elaborated here.

[0164] It should be noted that for multiple initial images, the data augmentation processing methods for each initial image are the same. After each initial image undergoes data augmentation processing, at least two sub-images corresponding to it are obtained.

[0165] Step S230, based on two sub-images belonging to the same initial image among the sub-images of each initial image, obtain multiple first positive sample image sets and similarity annotation results of the first positive sample image sets.

[0166] Among them, after obtaining at least two sub-images of each initial image, multiple first positive sample image sets and similarity annotation results of the first positive sample image sets can be determined based on two sub-images belonging to the same initial image among the sub-images of each initial image. The two sub-images belonging to the same initial image are similar. Therefore, after the initial image is augmented, the similarity annotation results of the first positive sample image sets can be directly obtained based on two sub-images belonging to the same initial image among the sub-images of each initial image, without manual annotation.

[0167] Step S240 : obtaining a plurality of first negative sample image sets and similarity annotation results of the first negative sample image sets based on two sub-images belonging to different initial images in the sub-images of each initial image.

[0168] Among them, the two subimages belonging to different initial images are not similar. Therefore, after the initial image is augmented, the similarity annotation results of the first negative sample image set can be directly obtained based on the two subimages belonging to the same initial image in the subimages of each initial image, and no manual annotation is required.

[0169] In step S250 , each first sample image set is input into the initial neural network model, and the image features of the two images in each first sample image set are extracted by the initial neural network model. Based on the image features of the two images, the predicted similarity of the first sample image set is predicted.

[0170] The predicted similarity of the first sample image set based on the image features of the two images has been described above and will not be repeated here.

[0171] Step S260 : determining a pre-training loss value based on the predicted similarities and similarity labeling results of each first sample image set.

[0172] Step S270: Determine whether the pre-training loss value meets the pre-training end condition.

[0173] If satisfied, execute steps SA71 to SA72:

[0174] Step SA71, end pre-training and obtain the pre-trained neural network model.

[0175] Step SA72: Based on multiple second positive sample image sets and multiple second negative sample image sets, the pre-trained neural network model is trained to obtain an image similarity model. The similarity annotation results of the second sample image sets are manual annotation results.

[0176] In practical applications, considering that it is difficult to judge whether some image pairs are similar, for example, image pairs obtained by reshooting the same image, image pairs obtained by screenshots of the same image, image pairs with different layouts, etc., these image pairs can be used as the second positive sample image set pairs to train the pre-trained neural network model, so that the model has better robustness for these images that are difficult to judge whether they are similar.

[0177] Among them, the similarity annotation results of the second positive sample image set are manually annotated. The second negative sample image set used for training the pre-trained neural network model described above can use the first negative sample image set described above, or the second negative sample image set and the corresponding similarity annotation results of the second negative sample image set can be determined by manual annotation.

[0178] If not satisfied, execute step B71: Adjust the model parameters of the initial neural network model, and repeat to execute steps S250 to S270 until the pre-training loss value satisfies the pre-training end condition.

[0179] Based on the same principle as the method shown in Figure 1 In the present application embodiment, an image processing method is further provided. Taking the server as the execution subject, this method will be described below. As shown in Figure 5 shown in, this method may include the following steps:

[0180] Step S310, obtain at least two images to be processed.

[0181] Step S320, process at least two images to be processed by calling the image similarity model to obtain the similarity of each image pair in the at least two images to be processed, so as to process the at least two images to be processed based on the similarity;

[0182] Among them, the image similarity model is obtained by the method described above.

[0183] Among them, the image pair in the at least two images to be processed refers to the image pair corresponding to any two images in the at least two images to be processed. If the at least two images to be processed are two images, the image pair in the at least two images to be processed refers to these two images.

[0184] Since the image similarity model trained above can be used to determine the similarity of image pairs, in practical applications, the trained image similarity model can be stored. When obtaining at least two images to be processed, the similarity of each image pair in the at least two images to be processed can be determined by the image similarity model, and then the at least two images to be processed can be subsequently processed based on the similarity. For example, classify similar image pairs and delete dissimilar image pairs, etc.

[0185] In the solution of the present application, since the image similarity model trained based on the method described above has good robustness and can accurately judge the similarity of image pairs, therefore, when determining the similarity based on this image similarity model, the accuracy of the similarity can be improved.

[0186] In one embodiment of the present application, by invoking an image similarity model to process at least two images to be processed, the similarity of each image pair in the at least two images to be processed is obtained, including:

[0187] Through the image similarity model, extract the image features of each image in the at least two images to be processed;

[0188] Based on the image features of each image, determine the similarity of each image pair in the at least two images to be processed.

[0189] Among them, the similarity of the image pair can be characterized by a feature distance (for example, Euclidean distance). The smaller the distance, the more similar the two images are. On the contrary, the larger the distance, the less similar the two images are. Then, based on the image features of each image, determine the feature distance of each image pair in the at least two images to be processed, and determine the similarity of each image pair in the at least two images to be processed through the feature distance.

[0190] To better understand the solution of the present application, the following is a further description of the solution in combination with Figure 6 The flow schematic diagram of an image processing method shown below:

[0191] Input the first training data set into the initial neural network model, and determine the predicted similarity of each first sample image set in the first training data set through this model.

[0192] Based on the similarity annotation results of each first sample image set in the first training data set and the predicted similarity of each first sample image set, determine the pre-training loss value.

[0193] If the pre-training loss value meets the pre-training end condition, obtain the pre-trained neural network model; if not, adjust the model parameters and retrain the initial neural network model until the pre-training loss value meets the pre-training end condition.

[0194] Input the second training data set into the pre-trained neural network model, and determine the predicted similarity of each second sample image set in the second training data set through this model.

[0195] Based on the similarity annotation results of each second sample image set in the second training data set and the predicted similarity of each second sample image set, determine the training loss value.

[0196] If the training loss value meets the training end condition, obtain the image similarity model; if not, adjust the model parameters and retrain the pre-trained neural network model until the training loss value meets the training end condition.

[0197] In practical applications, at least two images to be processed can be input into an image similarity model, and the similarity of each image pair in the at least two images to be processed can be determined through the image similarity model.

[0198] In one embodiment of the present application, obtaining at least two images to be processed includes:

[0199] Obtaining an image retrieval request, where the image retrieval request includes a retrieval image;

[0200] Obtaining an image database corresponding to the image retrieval request, the at least two images to be processed include the retrieval image and the retrieved images in the image database, and the image pair includes the retrieval image and one retrieved image;

[0201] The method further includes:

[0202] Determining a target image corresponding to the image retrieval request from the image database according to the similarity of each image pair, and providing the target image to the retriever.

[0203] Among them, the retrieval image is the image to be retrieved, which can be an image containing a certain target object. For example, an image containing a certain piece of clothing (target object), an image containing a certain pair of shoes (target object).

[0204] The image retrieval request can be initiated by a user based on the user's terminal device, and the terminal device can include at least one of the following: smart phone, tablet computer, notebook computer, desktop computer, smart speaker, smart watch, smart TV, smart vehicle-mounted device.

[0205] The image database includes the retrieval image and the retrieved images. Based on the image retrieval request, a target image matching the retrieval image can be retrieved from the image database based on the retrieval image. Specifically, through the image similarity model, the similarity between the retrieval image and each image in the image database can be determined, and a target image corresponding to the image retrieval request can be determined from the image database according to the similarity of each image pair.

[0206] The target image can be displayed through the retriever's terminal device. Among them, the terminal device can run a client that provides a picture display function. The client provides a picture display function, and the specific form of the client is not limited. For example: media player, browser, etc. The client can be in the form of an application or a web page, which is not limited here.

[0207] In one embodiment of the present application, to determine a target image corresponding to the image retrieval request from the image database according to the similarity of each image pair, specifically, the similarities of each image pair can be sorted from high to low, and the image corresponding to the image pair with the highest similarity can be selected as the target image.

[0208] Figure 7 A schematic diagram of the implementation environment of the image processing method provided by an embodiment of this application. The implementation environment in this example may include, but is not limited to, a retrieval server 101, a network 102, and a terminal device 103. The retrieval server 101 can communicate with the terminal device 103 through the network 102, send the received image retrieval request to the retrieval server 101, and the retrieval server 101 can send the retrieved target image to the terminal device 103 through the network.

[0209] The above terminal device 103 includes a human-computer interaction screen 1031, a processor 1032, and a memory 1033. The human-computer interaction screen 1031 is used to display the target image. The memory 1033 is used to store relevant data such as retrieved images and target images. The retrieval server 101 includes a database 1011 and a processing engine 1012. The processing engine 1012 can be used to train an image similarity model. The database 1011 is used to store the trained image similarity model and the image database. The terminal device 103 can upload the image retrieval request to the retrieval server 101 through the network. The processing engine 1012 in the retrieval server 101 can obtain the image database corresponding to the image retrieval request, determine the target image corresponding to the image retrieval request from the image database according to the similarity of each image pair, obtain the target image corresponding to the image retrieval request, and provide the target image to the terminal device 103 of the retriever for display.

[0210] The processing engine in the above retrieval server 101 has two main functions. The first function is to train an image similarity model, and the second function is to process the image retrieval request based on the image similarity model and the image database to obtain the target image corresponding to the image retrieval request (retrieval function). It can be understood that the above two functions can be implemented by two servers respectively. See Figure 8 , the two servers are a training server 201 and a retrieval server 202 respectively. The training server 201 is used to train an image similarity model, and the retrieval server 202 is used to implement the retrieval function. The image database is stored in the retrieval server 202.

[0211] In practical applications, the two servers can communicate with each other. After the training server 201 trains the image similarity model, it can store the image similarity model in the training server 201, or send the image similarity model to the retrieval server 202. Alternatively, when the retrieval server 202 needs to call the image similarity model, it sends a model call request to the training server 201, and the training server 201 sends the image similarity model to the retrieval server 202 based on this request.

[0212] As an example, the terminal device 204 sends an image retrieval request to the retrieval server 202 through the network 203. The retrieval server 202 invokes the image similarity model in the training server 201. Based on the image similarity model, after the retrieval server 202 completes the retrieval function, it sends the retrieved target image to the terminal device 204 through the network 203, so that the terminal device 204 can display the target image.

[0213] In one embodiment of the present application, obtaining at least two images to be processed includes:

[0214] Obtaining a set of images to be processed, where at least two images to be processed are each image in the set of images to be processed, and an image pair is any two images in the set of images to be processed;

[0215] The method further includes:

[0216] Classifying the images in the set of images to be processed according to the similarity of each image pair.

[0217] Among them, for the images in the set of images to be processed, they can be classified based on the similarity of each image pair. One implementable way is: those with similarities satisfying a preset condition are classified into one category. For example, those with similarities greater than a first set value and less than a second set value are classified into one category, those with similarities not less than the second set value and less than a third set value are classified into one category, and the remaining images in the set of images to be processed with similarities not less than the third set value are classified into one category. Among them, the first set value is less than the second set value is less than the third set value, and the first set value, the second set value, and the third set value can all be configured based on actual requirements.

[0218] In an alternative solution of the present application, content recommendation can also be performed according to the similarity of each image pair. For example, images with similarities greater than a set value are recommended to users as images to be recommended. The reference of the similarity between images is very extensive and will not be elaborated one by one here. Any solution related to determining the similarity between two images can be determined through the image similarity model in the present application.

[0219] Based on the same principle as the method shown in Figure 1 In the present application embodiment, an image processing apparatus 40 is further provided, as shown in Figure 9 shown in, the image processing apparatus 40 may include a training data acquisition module 410 and a model training module 420, where:

[0220] A training data acquisition module 410 is configured to acquire a first training data set and a second training data set. The first training data set includes multiple first sample image sets, and the second training data set includes multiple second sample image sets. Among them, the first sample image sets and their similarity annotation results are determined by performing data augmentation on each initial image, and the similarity annotation results of the second sample image sets are manually annotated results.

[0221] A model training module 420 is configured to pre-train an initial neural network model based on the first training data set to obtain a pre-trained neural network model; and train the pre-trained neural network model based on the second training data set to obtain an image similarity model, so as to determine the similarity of an image pair through the image similarity model.

[0222] In one embodiment of the present application, when the training data acquisition module 410 acquires the first training data set, it is specifically configured to:

[0223] Acquire multiple initial images;

[0224] For each initial image, perform data augmentation processing on the initial image to obtain at least two sub-images corresponding to the initial image;

[0225] Based on two sub-images belonging to the same initial image among the sub-images of each initial image, obtain multiple first positive sample image sets and the similarity annotation results of the first positive sample image sets;

[0226] Based on two sub-images belonging to different initial images among the sub-images of each initial image, obtain multiple first negative sample image sets and the similarity annotation results of the first negative sample image sets;

[0227] Among them, the multiple first sample image sets include multiple first positive sample image sets and multiple first negative sample image sets.

[0228] In one embodiment of the present application, the second sample image set is a positive sample image set, and the second training data set further includes multiple first negative sample image sets.

[0229] In one embodiment of the present application, the multiple second sample image sets include at least one of multiple second positive sample image sets or multiple second negative sample image sets. The similarity of the second positive sample image set is less than or equal to a first threshold; the similarity of the second negative sample image set is greater than or equal to a second threshold;

[0230] Among them, the first threshold is not less than the second threshold, and the similarity of the second sample image set is determined by the pre-trained neural network model.

[0231] In one embodiment of the present application, the data augmentation processing includes at least one of the following:

[0232] Image cropping; Smearing processing; Blurring processing; Color transformation; Grayscale transformation; Image rotation; Image flipping.

[0233] In one embodiment of the present application, when the model training module 420 pre-trains the initial neural network model based on the first training data set to obtain the pre-trained neural network model, it is specifically used for:

[0234] Repeatedly execute the following training steps until the pre-training loss value meets the pre-training end condition to obtain the pre-trained neural network model:

[0235] Input each first sample image set into the initial neural network model, extract the image features of two images in each first sample image set through the initial neural network model respectively, and predict the predicted similarity of the first sample image set based on the image features of the two images;

[0236] Determine the pre-training loss value according to the predicted similarity of each first sample image set and the similarity annotation result;

[0237] If the pre-training loss value meets the pre-training end condition, end the pre-training; if not, adjust the model parameters of the initial neural network model and repeat the training steps.

[0238] Based on the same principle as the method shown in Figure 5 In the embodiments of the present application, an image processing device 50 is further provided. As shown in Figure 10 As shown in, the image processing device 50 may include an image acquisition module 510 and an image processing module 520, where:

[0239] The image acquisition module 510 is used to acquire at least two images to be processed;

[0240] The image processing module 520 is used to process at least two pairs of images to be processed by calling the image similarity model to obtain the similarity of each image pair in the at least two images to be processed, so as to process the at least two images to be processed based on the similarity;

[0241] Among them, the image similarity model is obtained by the method in the foregoing text.

[0242] In one embodiment of the present application, when the image acquisition module acquires at least two images to be processed, it is specifically used for:

[0243] Acquire an image retrieval request, where the image retrieval request includes a retrieval image;

[0244] Acquire the image database corresponding to the image retrieval request. The at least two images to be processed include the retrieval image and the retrieved images in the image database, and the image pair includes the retrieval image and one retrieved image;

[0245] The device further includes:

[0246] An image retrieval module, configured to determine a target image corresponding to an image retrieval request from an image database according to the similarity of each pair of images, and provide the target image to a retriever.

[0247] In one embodiment of the present application, when the image acquisition module acquires at least two images to be processed, it is specifically configured to:

[0248] Acquire a set of images to be processed, where at least two images to be processed are each image in the set of images to be processed, and a pair of images is any two images in the set of images to be processed;

[0249] The device further includes:

[0250] An image classification module, configured to classify the images in the set of images to be processed according to the similarity of each pair of images.

[0251] The image processing device according to an embodiment of the present application can execute the image processing method provided by the embodiment of the present application, and the implementation principle is similar. The actions performed by each module and unit in the image processing device in each embodiment of the present application correspond to the steps in the image processing method in each embodiment of the present application. For the detailed function descriptions of each module of the image processing device, reference can specifically be made to the descriptions in the corresponding image processing method shown above, and details are not described herein again.

[0252] Wherein, the image processing device can be a computer program (including program code) running in a computer device. For example, the image processing device is an application software; the device can be used to execute the corresponding steps in the method provided by the embodiment of the present application.

[0253] In some embodiments, the image processing device provided by the embodiment of the present invention can be implemented in a combination of software and hardware. As an example, the image processing device provided by the embodiment of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the image processing method provided by the embodiment of the present invention. For example, a processor in the form of a hardware decoding processor can adopt one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components.

[0254] In some other embodiments, the image processing apparatus provided by the embodiments of the present invention may be implemented in software. Figure 9 The image processing apparatus stored in the memory is shown, which may be software in the form of a program, a plug-in, etc., and includes a series of modules, including a training data acquisition module 410 and a model training module 420, for implementing the image processing method provided by the embodiments of the present invention.

[0255] Based on the same principle as the method shown in the embodiments of the present application, an electronic device is further provided in the embodiments of the present application. The electronic device may include, but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the image processing method shown in any embodiment of the present application by calling the computer program.

[0256] When obtaining the image similarity model for determining the similarity of image pairs by the image processing method provided by the present application, the first sample image set and its similarity annotation result in the first training data set of the model are automatically determined by performing data augmentation on each initial image. Therefore, based on the data augmentation method, a large number of training samples with similarity annotation results can be generated to provide data support for the training of the model. Further, the solution of the embodiments of the present application also provides a second training data set of the model. Since the second sample image set and its similarity annotation result in the second training data set are manually annotated, the manually annotated similarity annotation result is more accurate. Therefore, the performance of the image similarity model trained based on the second training data set is better.

[0257] In an alternative embodiment, an electronic device is provided, such as Figure 11 shown Figure 11 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 may be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.

[0258] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0259] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0260] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0261] The memory 4003 is used to store the application program code (computer program) for executing the solution of this application, and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the application program code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.

[0262] Among them, the electronic device can also be a terminal device. Figure 11 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0263] Among them, the image processing method provided in this application can also be implemented in the way of cloud computing. Cloud computing refers to the delivery and usage model of IT infrastructure, which means obtaining the required resources in a demand-driven and easily expandable manner through the network; in a broad sense, cloud computing refers to the delivery and usage model of services, which means obtaining the required services in a demand-driven and easily expandable manner through the network. Such services can be related to IT and software, the Internet, or other services. Cloud computing is the product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balance.

[0264] With the development of the Internet, real-time data streams, and diverse connected devices, as well as the promotion of demands such as search services, social networks, mobile commerce, and open collaboration, cloud computing has developed rapidly. Different from the previous parallel distributed computing, the emergence of cloud computing will drive a revolutionary change in the entire Internet model and enterprise management model conceptually.

[0265] The image processing method provided by this application can also be implemented through an artificial intelligence cloud service. An artificial intelligence cloud service, generally also known as AIaaS (AI as a Service, which means "AI is a service" in Chinese). This is a current mainstream service method for artificial intelligence platforms. Specifically, the AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service model is similar to opening an AI-themed mall: all developers can access and use one or more artificial intelligence services provided by the platform through the API interface. Some senior developers can also use the AI framework and AI infrastructure provided by the platform to deploy and operate their own exclusive cloud artificial intelligence services. In this application, the AI framework and AI infrastructure provided by the platform can be used to implement the image processing method provided by this application.

[0266] The embodiment of this application provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiment.

[0267] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps do not necessarily execute in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction and can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily execute at the same time, but can be executed at different times. Their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0268] The computer-readable storage medium provided by the embodiment of this application can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or component.

[0269] The above computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to execute the methods shown in the above embodiments.

[0270] According to another aspect of the present application, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the image processing method provided in the above various embodiments.

[0271] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through an Internet service provider using the Internet).

[0272] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the methods and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0273] The modules described in the embodiments of the present application may be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the module itself.

[0274] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.

Claims

1. An image processing method, characterized in that, Including: Obtain a first training dataset, where the first training dataset includes multiple first sample image sets. Among them, the first sample image sets and their similarity annotation results are determined by performing data augmentation processing on each initial image; Pre-train an initial neural network model based on the first training dataset to obtain a pre-trained neural network model; Obtain multiple second sample image sets, and the similarity annotation results of the multiple second sample image sets are manually annotated results; For each of the second sample image sets, input the images in the second sample image set into the pre-trained neural network model to obtain the predicted similarity between the images in the second sample image set; According to the manual annotation results of each second sample image set and the predicted similarity between the images in each second sample image set, obtain a second training dataset, where the second training dataset includes at least one of the following: A second sample image set with a manual annotation result of similar and a predicted similarity less than or equal to a first threshold; a second sample image set with a manual annotation result of dissimilar and a predicted similarity greater than or equal to a second threshold; Train the pre-trained neural network model based on the second training dataset to obtain an image similarity model, so as to determine the similarity of an image pair through the image similarity model.

2. The method according to claim 1, characterized in that, The obtaining of the first training dataset includes: Obtain multiple initial images; For each of the initial images, perform data augmentation processing on the initial image to obtain at least two sub-images corresponding to the initial image; Based on two sub-images belonging to the same initial image among the sub-images of each initial image, obtain multiple first positive sample image sets and the similarity annotation results of the first positive sample image sets; Based on two sub-images belonging to different initial images among the sub-images of each initial image, obtain multiple first negative sample image sets and the similarity annotation results of the first negative sample image sets; Among them, the multiple first sample image sets include the multiple first positive sample image sets and the multiple first negative sample image sets.

3. The method according to claim 2, wherein The second training dataset also includes multiple of the first negative sample image sets.

4. The method according to any one of claims 1 to 3, characterized in that The data augmentation processing includes at least one of the following: Image cropping; Smearing processing; Blurring processing; Color transformation; Grayscale transformation; Image rotation; Image flipping.

5. The method according to any one of claims 1 to 3, characterized in that, The pre-training of the initial neural network model based on the first training dataset to obtain a pre-trained neural network model includes: Repeatedly execute the following training steps until the pre-training loss value meets the pre-training end condition to obtain the pre-trained neural network model: Input each of the first sample image sets into the initial neural network model, extract the image features of two images in each of the first sample image sets through the initial neural network model, and predict the predicted similarity of the first sample image set based on the image features of the two images; Determine the pre-training loss value according to the predicted similarity and the similarity annotation result of each first sample image set; If the pre-training loss value meets the pre-training end condition, end the pre-training; if not, adjust the model parameters of the initial neural network model and repeat the training step.

6. An image processing method, characterized in that, It includes: Obtain at least two images to be processed; Process the at least two images to be processed by calling an image similarity model to obtain the similarity of each image pair in the at least two images to be processed, so as to process the at least two images to be processed based on the similarity; Wherein, the image similarity model is obtained by the method described in any one of claims 1-5.

7. The method according to claim 6, characterized in that The obtaining of at least two images to be processed includes: Obtain an image retrieval request, and the image retrieval request contains a retrieval image; Obtain the image database corresponding to the image retrieval request, the at least two images to be processed include the retrieval image and the retrieved images in the image database, and the image pair includes the retrieval image and one retrieved image; The method further includes: Determine the target image corresponding to the image retrieval request from the image database according to the similarity of each image pair, and provide the target image to the retriever.

8. The method according to claim 6, wherein The obtaining of at least two images to be processed includes: Obtain a set of images to be processed, the at least two images to be processed are each image in the set of images to be processed, and the image pair is any two images in the set of images to be processed; The method further includes: Classify the images in the set of images to be processed according to the similarity of each image pair.

9. An image processing apparatus, characterized in that, It includes: A training data acquisition module, configured to acquire a first training data set, the first training data set includes a plurality of first sample image sets, wherein the first sample image set and its similarity annotation result are determined by performing data augmentation processing on each initial image; A second training data acquisition module, configured to acquire a plurality of second sample image sets, and the similarity annotation results of the plurality of second sample image sets are manually annotated results; for each second sample image set, input the images in the second sample image set into the pre-trained neural network model to obtain the predicted similarity between the images in the second sample image set; according to the manual annotation results of each second sample image set and the predicted similarity between the images in each second sample image set, obtain a second training data set, wherein the second training data set includes at least one of the following: a second sample image set with a manual annotation result of similar and a predicted similarity less than or equal to a first threshold; a second sample image set with a manual annotation result of dissimilar and a predicted similarity greater than or equal to a second threshold; A model training module, configured to pre-train an initial neural network model based on the first training data set to obtain a pre-trained neural network model; and train the pre-trained neural network model based on the second training data set to obtain an image similarity model, so as to determine the similarity of an image pair through the image similarity model.

10. The device according to claim 9, characterized in that When the training data acquisition module acquires the first training data set, it is specifically configured to: Obtain multiple initial images; For each of the initial images, perform data augmentation processing on the initial image to obtain at least two sub-images corresponding to the initial image; Based on two sub-images belonging to the same initial image among the sub-images of each initial image, obtain a plurality of first positive sample image sets and similarity annotation results of the first positive sample image sets; Based on two sub-images belonging to different initial images among the sub-images of each initial image, obtain a plurality of first negative sample image sets and similarity annotation results of the first negative sample image sets; Among them, the plurality of first sample image sets include the plurality of first positive sample image sets and the plurality of first negative sample image sets.

11. The device according to claim 10, wherein The second training data set further includes a plurality of the first negative sample image sets.

12. The device according to any one of claims 9-11, characterized in that, The data augmentation processing includes at least one of the following: Image cropping; Smearing processing; Blurring processing; Color transformation; Grayscale transformation; Image rotation; Image flipping.

13. The device according to any one of claims 9-11, characterized in that, When the model training module pre-trains the initial neural network model based on the first training data set to obtain the pre-trained neural network model, it is specifically used for: Repeatedly execute the following training steps until the pre-training loss value meets the pre-training end condition to obtain the pre-trained neural network model: Input each of the first sample image sets into the initial neural network model, extract the image features of two images in each of the first sample image sets through the initial neural network model, and predict the predicted similarity of the first sample image set based on the image features of the two images; Determine the pre-training loss value according to the predicted similarity and similarity annotation results of each of the first sample image sets; If the pre-training loss value meets the pre-training end condition, end the pre-training; if not, adjust the model parameters of the initial neural network model and repeat the training steps.

14. An image processing apparatus, characterized in that, Including: An image acquisition module, configured to acquire at least two images to be processed; An image processing module, configured to process each image pair of the at least two images to be processed by calling an image similarity model to obtain the similarity of each image pair of the at least two images to be processed, so as to process the at least two images to be processed based on the similarity; Among them, the image similarity model is obtained by the method described in any one of claims 1-7.

15. The device according to claim 14, characterized in that When the image acquisition module acquires at least two images to be processed, it is specifically used for: Acquire an image retrieval request, where the image retrieval request includes a retrieval image; Acquire the image database corresponding to the image retrieval request, the at least two images to be processed include the retrieval image and the retrieved images in the image database, and the image pair includes the retrieval image and one retrieved image; The apparatus further includes: An image retrieval module, configured to determine the target image corresponding to the image retrieval request from the image database according to the similarity of each image pair, and provide the target image to the retriever.

16. The device according to claim 14, characterized in that, When the image acquisition module acquires at least two images to be processed, it is specifically used for: Obtain a set of images to be processed, where the at least two images to be processed are each image in the set of images to be processed, and the image pair is any two images in the set of images to be processed; The apparatus further includes: An image classification module, configured to classify the images in the set of images to be processed according to the similarity of each image pair.

17. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method according to any one of claims 1-8.

18. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • MRI image segmentation method and device based on coarse and fine training, and storage medium

    CN110853048A

  • Image pre-labeling method and device and electronic equipment

    CN111753114A