Image difference detection method, device and equipment, computer readable storage medium and computer program product
By using a large visual-language model for difference detection, this method solves the problems of long image difference detection time and insufficient detection of subtle differences in existing technologies, and achieves automated and efficient image difference text description.
Patent Information
- Application Number
- CN202410547530.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-28
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies are labor-intensive and time-consuming when detecting differences between multiple images. They cannot directly output semantic differences, are not sensitive enough to subtle differences, require high computing resources, and cannot automatically process and provide detailed explanations.
A large vision-language model is used to detect image differences, and difference sub-images are obtained through differential processing. Feature embedding and text mapping are performed to generate intuitive text descriptions.
It realizes automated comparative analysis of large numbers of images, reduces manual review work, quickly and accurately detects image differences, and provides detailed text descriptions.
Smart Images

Figure CN120852812A_ABST
Abstract
Description
Technical Field
[0001] This application relates to image processing technology, and more particularly to a method, apparatus, device, computer-readable storage medium, and computer program product for detecting differences in images. Background Technology
[0002] With the development of image processing technology, it is necessary to detect the differences between different images in many scenarios, such as image resource comparison in game scenarios and medical image comparison analysis in medical image analysis scenarios.
[0003] In related technologies, manual comparison methods are used to detect differences between multiple images. However, this method involves a large workload and is time-consuming when processing hundreds of thousands of images, and it cannot directly output the semantic differences between images. Summary of the Invention
[0004] This application provides an image difference detection method, apparatus, device, computer-readable storage medium, and computer program product, which are at least applicable to the fields of artificial intelligence or image processing. They can convert differences between images into intuitive text descriptions and improve the detection rate of differences between images.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] This application provides a method for detecting image differences. The method includes: acquiring a first image and a second image to be detected; performing difference processing on the first image and the second image to obtain a difference image; the difference image includes multiple difference sub-images for characterizing the differences between the first image and the second image; performing feature embedding processing on the difference image based on the multiple difference sub-images to obtain an image feature sequence; performing text mapping on the image feature sequence based on a preset image-text dictionary to obtain a text feature sequence; the text feature sequence is a feature sequence corresponding to the multiple difference sub-images; and determining target text for characterizing the differences between the first image and the second image based on the text feature sequence.
[0007] This application provides an image difference detection device, comprising: an image acquisition module for acquiring a first image and a second image to be detected; an image difference module for performing difference processing on the first image and the second image to obtain a difference image; the difference image includes multiple difference sub-images for characterizing the differences between the first image and the second image; a feature embedding module for performing feature embedding processing on the difference image based on the multiple difference sub-images to obtain an image feature sequence; a text mapping module for performing text mapping on the image feature sequence based on a preset image-text dictionary to obtain a text feature sequence; the text feature sequence is a feature sequence corresponding to the multiple difference sub-images; and a text determination module for determining target text characterizing the differences between the first image and the second image based on the text feature sequence.
[0008] In the above scheme, the image difference module is further configured to subtract the pixel value of the corresponding position in the second image from each pixel value of the first image to obtain a third image; for each pixel in the third image, when the pixel value of the pixel is greater than a preset threshold, the pixel is marked as a difference region; contour detection is performed on the third image marked with difference regions to obtain the contour of each difference sub-image; the contour of each difference sub-image is marked in the third image with a preset style to obtain the difference image.
[0009] In the above scheme, the feature embedding module is further configured to perform convolution processing on the sub-image features of the multiple difference sub-images to obtain a difference enhancement matrix; extract features from the difference image to obtain a first image feature sequence; and determine the image feature sequence based on the first image feature sequence and the difference enhancement matrix.
[0010] In the above scheme, the feature embedding module is further configured to perform self-attention processing on the first image feature sequence through a preset set of multiple first linear transformation matrices respectively, thereby obtaining multiple second image feature sequences; concatenate the multiple second image feature sequences to obtain an intermediate image feature sequence; and perform weighted processing on the intermediate image feature sequence through the difference enhancement matrix to obtain the image feature sequence.
[0011] In the above scheme, the first linear transformation matrix set includes three first linear transformation matrices; the feature embedding module is further configured to perform the following operations on the first image feature sequence through each of the first linear transformation matrix sets to obtain the corresponding second image feature sequence: performing linear transformations on the first image feature sequence through the three first linear transformation matrices in the first linear transformation matrix set to obtain an image feature query sequence, an image feature key sequence, and an image feature value sequence; determining a first weight matrix based on the image feature query sequence and the image feature key sequence; and determining the second image feature sequence based on the first weight matrix and the image feature value sequence.
[0012] In the above scheme, the feature embedding module is further configured to determine a first similarity matrix between the image feature query sequence and the image feature key sequence; obtain the dimension of the image feature query sequence; determine an intermediate matrix based on the first similarity matrix and the dimension; and normalize the intermediate matrix to obtain the first weight matrix.
[0013] In the above scheme, the text mapping module is further used to perform text mapping on the image feature sequence based on a preset image-text dictionary to obtain an initial visual word sequence; the initial visual word sequence includes vectors of multiple visual words, and each of the difference sub-images corresponds to one visual word; the initial visual word sequence is subjected to position embedding processing to obtain a target visual word sequence; the target visual word sequence is subjected to encoding processing to obtain the text feature sequence.
[0014] In the above scheme, the text mapping module is further configured to perform self-attention processing on the target visual word sequence through a preset second set of linear transformation matrices to obtain a first text feature sequence; perform layer normalization processing on the sum of the first text feature sequence and the target visual word sequence to obtain a second text feature sequence; perform activation processing on the second text feature sequence to obtain a third text feature sequence; and perform layer normalization processing on the sum of the third text feature sequence and the second text feature sequence to obtain the text feature sequence.
[0015] In the above scheme, the second linear transformation matrix set includes three second linear transformation matrices; the text mapping module is further configured to perform linear transformations on the target visual word sequence using the three second linear transformation matrices in the second linear transformation matrix set, respectively, to obtain a text feature query sequence, a text feature key sequence, and a text feature value sequence; determine a second weight matrix based on the text feature query sequence and the text feature key sequence; and determine the first text feature sequence based on the second weight matrix and the text feature value sequence.
[0016] In the above scheme, the image difference detection method is applied to a pre-trained difference detection model. The image difference detection device further includes a model training module, used to train the difference detection model through the following steps: acquiring training data, the training data including multiple sample difference images and label text corresponding to each sample difference image, the sample difference images including multiple sample difference sub-images for characterizing the difference between two sample images; using the feature embedding layer of the difference detection model, performing feature embedding processing on the sample difference images through multiple sample difference sub-images to obtain a sample image feature sequence; using the feature mapping layer of the difference detection model, processing the sample difference images... The sample image feature sequence is mapped to text to obtain a sample text feature sequence; the sample text feature sequence is a feature sequence corresponding to the plurality of sample difference sub-images; using the text decoding layer of the difference detection model, based on the sample text feature sequence, sample text used to characterize the difference between the sample first image and the sample second image is determined; a first loss result is determined based on the sample image feature sequence and the sample text feature sequence; a second loss result is determined based on the label text and the sample text; the model parameters in the difference detection model are updated using the first loss result and the second loss result to obtain the trained difference detection model.
[0017] This application provides an electronic device, which includes: a memory for storing computer-executable instructions or computer programs; and a processor for executing the computer-executable instructions or computer programs stored in the memory to implement the image difference detection method provided in this application.
[0018] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the image difference detection method provided in this application when executed by a processor.
[0019] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the image difference detection method provided in this application.
[0020] The embodiments of this application have the following beneficial effects:
[0021] In a scenario involving difference detection between a first image and a second image, firstly, the first and second images are differentially processed to obtain a difference image. Then, based on multiple difference sub-images representing the differences between the first and second images in the difference image, feature embedding processing is performed on the difference image to obtain an image feature sequence. This allows the image feature sequence to focus on the differences between the first and second images. Consequently, the text feature sequence obtained after text mapping of the image feature sequence is a feature sequence corresponding to the multiple difference sub-images, that is, a feature sequence of text used to describe the differences between the first and second images. Therefore, based on the text feature sequence, the target text representing the differences between the first and second images can be directly determined. Thus, the embodiments of this application can achieve automated comparative analysis of a large number of images, thereby greatly reducing the workload of manual comparison, enabling rapid and accurate analysis of a large number of images, improving the detection rate of differences between images, and also converting the differences between images into intuitive text descriptions. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the architecture of the image difference detection system provided in the embodiments of this application;
[0023] Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;
[0024] Figure 3 This is a first flowchart illustrating the image difference detection method provided in this application embodiment;
[0025] Figure 4 This is a schematic diagram of the second process of the image difference detection method provided in the embodiments of this application;
[0026] Figure 5 This is a schematic diagram of the third process of the image difference detection method provided in the embodiments of this application;
[0027] Figure 6 This is a schematic diagram of the fourth process of the image difference detection method provided in the embodiments of this application;
[0028] Figure 7 This is a schematic diagram of the fifth step of the image difference detection method provided in the embodiments of this application;
[0029] Figure 8 This is another optional flowchart illustrating the image difference detection method provided in the embodiments of this application;
[0030] Figure 9 This is a schematic diagram of a game scene provided in an embodiment of this application;
[0031] Figure 10This is another game scene illustration provided in the embodiments of this application;
[0032] Figure 11 This is a schematic diagram of the semantic matching mechanism of the visual-language big model in the embodiments of this application;
[0033] Figure 12 This is a schematic diagram of the difference detection process for images using the visual-language large model provided in the embodiments of this application.
[0034] It should be noted that the terms "first", "second", "third", "fourth" and "fifth" mentioned above are only used to distinguish different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0036] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0037] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0038] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0039] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0040] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0041] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0042] 1) Visual-Language Big Model: A multimodal model applicable to many fields, such as computer vision, natural language processing, and artificial intelligence. The visual-language big model is used to simultaneously understand visual images and related textual descriptions to provide a deeper level of image understanding. For example, it can be used in automatic image captioning scenarios, generating text describing image content; it can also be used in visual question answering scenarios, extracting information from an image and answering a question about it.
[0043] 2) Transformer Model: A deep learning model based on attention mechanisms, used for processing sequential data, and widely used, especially in natural language processing tasks. Compared to traditional Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) networks, the Transformer model achieves more efficient sequence modeling through self-attention, and has the advantages of shorter training time and parallel computation capability. The core components of the Transformer model include self-attention mechanism and positional encoding.
[0044] 3) Self-attention mechanism: The self-attention mechanism allows interaction between elements in the input sequence, representing the importance of each element to other elements by calculating the attention weight of each element. The "Query-Key-Value (QKV)" attention mechanism is the core of the self-attention mechanism, where the query Q is used to find a relevant vector, the key K is used to calculate the similarity between the query vector and other elements, and the value V is used to represent the specific content information of each element. The concepts of Query, Key, and Value originate from information retrieval systems. For example, when searching for a product, the information entered in the search bar is called the Query. The system then matches the Query with the Key, and obtains the matching content based on the similarity between the Query and the Key. Taking "searching for gray men's sweaters" as an example: the Query (to match others) serves as guiding information, containing the idea of what information is needed. For example, entering "gray men's sweaters" in the search bar, the system hopes to return some similar products; the Key (to be matched) represents other products to be matched. When the system receives the information "gray" and "men's sweaters," it matches all products in the database; Attention (Q, K) represents the degree of matching between the Query and the Key. There are many products (Keys) in the system, and the matching degree of products that match the description (Query) will be higher; Value (information to be extracted) is the information itself, which simply expresses the information of the input features.
[0045] In the field of scene image resource comparison, due to technological limitations and complexity, manually comparing differences between images becomes extremely tedious when processing hundreds of thousands of images, and it is difficult to directly output semantic differences between images. Besides this, several other commonly used comparison methods exist: image feature-based comparison methods, hash algorithm comparison methods, deep learning models, and Generative Adversarial Networks (GANs). Image feature-based comparison methods utilize computer vision techniques to extract image features such as color, texture, and shape, and then compare these features to determine the similarity between images. However, this method mainly focuses on low-level image features rather than high-level semantic information, and its effectiveness is limited when dealing with scenes with significant semantic differences. Hash algorithm comparison methods hash images, mapping them to fixed-length strings, and then compare these strings to determine image similarity. This method can quickly determine image similarity in some scenarios, but it is not sensitive enough to small differences and it is difficult to output detailed semantic difference information. Image comparison methods using deep learning models: The development of deep learning technology has made image comparison based on neural networks possible. Models such as Convolutional Neural Networks (CNNs) can learn high-level semantic features of images, thus achieving more accurate comparisons. This method can better capture semantic differences between images, but model training and deployment require significant computational resources and time, resulting in low efficiency. Image comparison methods using Generative Adversarial Networks (GANs): GANs are powerful deep learning structures that can be used to generate new images. In comparison scenarios, GANs can be used to generate images similar to the original image but with subtle differences, and then compared with the original image. However, this method has low accuracy in complex image scenes.
[0046] In summary, the relevant technologies still have many shortcomings, especially when processing large-scale image data, the following problems exist:
[0047] 1. Lack of automated processing and semantic understanding, and inability to provide detailed explanations. Related technologies primarily rely on extracting low-level image features, such as color and texture, while ignoring semantic information. This results in the inability to directly output textual differences when processing image scenarios with highly semantically diverse images, making manual review unavoidable. The inability to directly output detailed textual differences or explanations adds an extra burden to manual review, while in some applications, understanding detailed information about image differences is crucial for timely problem-solving. Therefore, these technologies require manual intervention in detecting image differences, making the entire process relatively cumbersome and unable to completely eliminate reliance on manual review. Processing hundreds of thousands of images involves a significant workload, is time-consuming, and cannot directly output the semantic differences between images.
[0048] 2. Insufficient sensitivity to changes. Among related technologies, especially comparison methods based on hash algorithms, there is insufficient sensitivity to minute changes or subtle differences in images. When the scene of an image undergoes some minor changes, these methods cannot accurately detect these image differences.
[0049] 3. High computational resource requirements. In related technologies, deep learning models typically require a large amount of computational resources and time for training, making them unsuitable for scenarios requiring real-time or rapid comparisons.
[0050] 4. Sensitivity to noise and deformation. Since images may be affected by noise, changes in lighting, or deformation, the performance of image comparison methods and deep learning models in related technologies will degrade. Therefore, the above method has low robustness to interference factors.
[0051] Based on the problems existing in related technologies, this application provides an image difference detection method, device, electronic device, computer-readable storage medium, and computer program product. The method is a scheme for automated image difference detection by combining a visual-language large model. It combines visual and semantic text for automated detection, which can transform the differences between images into intuitive text descriptions, reduce the tedious work of manual review, and improve the detection rate of differences between images.
[0052] In the image difference detection method provided in this application embodiment, firstly, a first image and a second image to be detected are acquired; then, the first image and the second image are differentially processed to obtain a difference image; the difference image includes multiple difference sub-images used to characterize the differences between the first image and the second image; then, feature embedding processing is performed on the difference image based on the multiple difference sub-images to obtain an image feature sequence; then, based on a preset image-text dictionary, text mapping is performed on the image feature sequence to obtain a text feature sequence; the text feature sequence is a feature sequence corresponding to the multiple difference sub-images; finally, based on the text feature sequence, target text used to characterize the differences between the first image and the second image is determined.
[0053] The following describes an exemplary application of the image difference detection device provided in this application embodiment. This image difference detection is an electronic device used to implement an image difference detection method. The electronic device provided in this application embodiment can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or it can be implemented as a server. The server can be an independent physical server, a server cluster composed of multiple physical servers, or a distributed system. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application embodiment. The following will describe exemplary applications when the image difference detection device is implemented as a server or terminal.
[0054] See Figure 1 , Figure 1This is a schematic diagram of the architecture of the image difference detection system provided in this application embodiment. To achieve accurate difference detection between the first image and the second image, an image difference detection application can be provided. This application can be a dedicated application for image difference detection, or it can be a functional module in other applications (such as a scene resource comparison module in a real-time game development application). The image difference detection system 100 in this application embodiment includes at least a terminal 400, a network 300, and a server 200, where the server 200 is the server for the image difference detection application. The server 200 can constitute the image difference detection device of this application embodiment, that is, the image difference detection method of this application embodiment is implemented through the server 200. The terminal 400 is connected to the server 200 through the network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.
[0055] See Figure 1 When performing difference detection on the first image and the second image, the user can perform interactive operations on the client of the image difference detection application through the terminal 400. These interactive operations can include, for example, image input operations or click detection operations. After receiving the user's interactive operation, the client can encapsulate the first image and the second image into an image difference detection request and send the image difference detection request to the server 200 through the network 300. Upon receiving the image difference detection request, the server 200, in response to the request, acquires the first image and the second image to be detected. The server 200 performs difference processing on the first image and the second image to obtain a difference image. The difference image includes multiple difference sub-images used to characterize the differences between the first image and the second image. The server 200 performs feature embedding processing on the difference image based on the multiple difference sub-images to obtain an image feature sequence. The server 200 performs text mapping on the image feature sequence based on a preset image-text dictionary to obtain a text feature sequence. The text feature sequence is a feature sequence corresponding to the multiple difference sub-images. Finally, the server 200 determines the target text used to characterize the differences between the first image and the second image based on the text feature sequence. After the target text is determined, the server 200 can also send the target text to the terminal 400 to display the target text to the user, so that the user can understand the differences between the images based on the target text.
[0056] In some embodiments, the image difference detection method of this application embodiment can also be executed by the terminal 400 itself. That is, after the terminal 400 receives the interactive operation input by the user through the client, the terminal 400 acquires the first image and the second image to be detected; the terminal 400 performs difference processing on the first image and the second image to obtain a difference image; the difference image includes multiple difference sub-images used to characterize the differences between the first image and the second image; the terminal 400 performs feature embedding processing on the difference image based on the multiple difference sub-images to obtain an image feature sequence; the terminal 400 performs text mapping on the image feature sequence based on a preset image-text dictionary to obtain a text feature sequence; the text feature sequence is a feature sequence corresponding to the multiple difference sub-images; the terminal 400 determines the target text used to characterize the differences between the first image and the second image based on the text feature sequence. After the target text is determined, the target text is displayed on the client interface of the terminal 400, so that the user can understand the differences between the images based on the target text.
[0057] The image difference detection method provided in this application embodiment can also be implemented based on a cloud platform and through cloud technology. For example, the server 200 mentioned above can be a cloud server. The cloud server acquires the first image and the second image to be detected; or, the server performs differential processing on the first image and the second image to obtain a difference image; the difference image includes multiple difference sub-images used to characterize the differences between the first image and the second image; or, the server performs feature embedding processing on the difference image based on the multiple difference sub-images to obtain an image feature sequence; or, the server performs text mapping on the image feature sequence based on a preset image-text dictionary to obtain a text feature sequence; the text feature sequence is a feature sequence corresponding to the multiple difference sub-images; or, the server determines the target text used to characterize the differences between the first image and the second image based on the text feature sequence.
[0058] In some embodiments, a cloud storage device may also be included, where the first and second images to be detected can be stored, as well as the difference image and a preset image-text dictionary. Thus, upon receiving a request for image difference detection, the difference image and the preset image-text dictionary can be directly retrieved from the cloud storage, enabling feature embedding processing of the difference image and text mapping of the image feature sequence, thereby improving the efficiency of image difference detection.
[0059] It's important to clarify that cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Cloud technology is a collective term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies applied in the cloud computing business model. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will require robust system support, which can be achieved through cloud computing.
[0060] See Figure 2 , Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Figure 2 The illustrated electronic device may be an image difference detection device, which includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in terminal 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2 The general labeled all buses as Bus System 440.
[0061] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0062] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0063] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.
[0064] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.
[0065] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0066] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;
[0067] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.
[0068] Presentation module 453 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with user interface 430;
[0069] The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.
[0070] In some embodiments, the image difference detection device provided in this application can be implemented in software. Figure 2A difference detection device 455 for images stored in memory 450 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: an image acquisition module 4551, an image difference module 4552, a feature embedding module 4553, a text mapping module 4554, and a text determination module 4555. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.
[0071] In other embodiments, the image difference detection device provided in this application can be implemented in hardware. As an example, the device provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the image difference detection method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0072] It should be noted that the example of image difference detection method in the following text is based on the game field, which is to automatically detect differences in images in the game's resource scene. By automatically determining the difference image, the semantic differences of images in the resource scene can be quickly output. When the game is updated, it can quickly detect whether there are any errors in the updated resource scene. Based on the understanding of the following, those skilled in the art can also apply the image difference detection method provided in the embodiments of this application to any other scenario that requires the detection of image differences. For example, in the field of medical imaging, the image difference detection method can be used to compare images of the same object taken at different times to detect changes in lesions; in video detection systems, the image difference detection method can be used to detect abnormal events in the detection area, such as people entering and leaving, vehicles moving, etc. By comparing the differences between consecutive image frames, the system can identify abnormal situations in real time; in the manufacturing field, the image difference detection method can be used for product quality inspection, by comparing the differences between product images and standard images to detect whether the product meets quality standards, thereby improving production efficiency; in geographic information systems, the image difference detection method can be used for map updates and change detection, by comparing satellite images at different times to automatically identify urban development, land use changes, etc., to help update map data; the image difference detection method can also play a role in many other fields, such as agriculture and environmental monitoring.
[0073] Figure 3 This is a schematic diagram of the first process of image difference detection provided in the embodiments of this application. The following will be combined with... Figure 3 The steps shown are explained as follows: Figure 3 As shown, taking the server as the execution subject for image difference detection as an example, the method includes the following steps S101 to S105:
[0074] Step S101: Obtain the first image and the second image to be detected.
[0075] Here, the first image and the second image are images input by the user through the terminal for difference detection. The first image and the second image can also be two identical images. For example, in a scenario where difference detection is performed on images within a game's resource scene, the first image can be an image of resource scene a before the game version update, and the second image can be an image of resource scene a after the game version update. Difference detection is performed on the first image and the second image, and text representing the difference is output. This allows game operators to determine whether the update of resource scene a meets the preset game update expectations, i.e., whether the game has been updated normally.
[0076] Step S102: Perform differential processing on the first image and the second image to obtain a differential image.
[0077] The difference image includes multiple difference sub-images used to characterize the differences between the first image and the second image.
[0078] Here, the first image and the second image can be segmented separately to obtain multiple sub-images. These sub-images of the first image are compared with those of the second image. At least one sub-image in the first image differs from the corresponding sub-image in the second image; the region containing the different sub-image is called the difference region. After performing difference processing on the first and second images to obtain a difference image, the sub-images corresponding to the difference regions in the difference image are the difference sub-images. Multiple difference sub-images can be marked in the difference image using a preset style, such as using red boxes. The embodiments of this application do not limit the method of difference processing. For example, it can be a pixel difference method: subtracting corresponding pixels of the first image and the second image to obtain a difference image; or, it can be a histogram difference method: subtracting the histograms of the first image and the second image to obtain a difference image, where the histogram difference can reflect the distribution of pixel values between the two images; or, it can be a feature difference method: extracting feature vectors of the first image and the second image and comparing the differences between the feature vectors of the first image and the second image; or, it can be a structured difference method: performing structured processing (such as edge detection, filtering, etc.) on the two images before difference processing; or, it can also be a deep learning method: using a deep learning model to compare the first image and the second image and calculate the differences between them.
[0079] For example, both the first and second images represent a scene of a road in the game. The difference between the first and second images is that a sidewalk is added to the second image. The first image can be segmented into 100 sub-images, and the second image can also be segmented into 100 sub-images. The difference image obtained by performing difference processing on the first and second images also includes 100 sub-images, where the multiple sub-images at the sidewalk location are the difference sub-images.
[0080] In some embodiments, see Figure 4 , Figure 4 The difference processing of the first image and the second image in step S102 to obtain the difference image can be achieved through the following steps S1021 to S1024:
[0081] Step S1021: Subtract the corresponding pixel value in the second image from each pixel value in the first image to obtain the third image.
[0082] Here, multiple pixel values from the first image and multiple pixel values from the second image can be obtained. The first and second images have the same number and position of pixels. For any given pixel, the pixel value in the first image can be subtracted from the pixel value in the second image to obtain the pixel value in the third image. After performing the above operation on each pixel, the third image is obtained. Noise removal, contrast adjustment, and other operations can also be performed on the obtained third image to enhance its clarity.
[0083] Step S1022: For each pixel in the third image, when the pixel value of the pixel is greater than a preset threshold, the pixel is marked as a difference region.
[0084] Here, this embodiment does not specifically limit the value of the preset threshold; it can be set according to actual needs. For each pixel in the third image, when the pixel value is greater than the preset threshold, it indicates that the image of the pixel's location in the first image differs from the image of the pixel's location in the second image. When the pixel value is less than or equal to the preset threshold, it indicates that the image of the pixel's location in the first image is the same as the image of the pixel's location in the second image. Pixels with pixel values greater than the preset threshold are marked as difference regions.
[0085] Step S1023: Perform contour detection on the third image with marked difference regions to obtain the contour of each difference sub-image.
[0086] Here, contour detection refers to using image processing libraries or algorithms to detect the contours of differing regions in a third image. The contour detection process is as follows: First, edge detection can be performed on the third image to find the boundaries between differing and non-dissimilar regions. This application embodiment does not limit the edge detection algorithm; for example, it can use Sobel, Canny, etc. After edge detection, the contours of the boundaries can be detected using methods such as connectivity analysis. The findContours function can be used to detect the contours of each differing sub-image.
[0087] Step S1024: Mark the outline of each difference sub-image in the third image using a preset style to obtain the difference image.
[0088] This application embodiment does not specifically limit the preset style; different colors, lines, etc., can be used. For example, after obtaining the outline of each difference sub-image, a drawing library or function can be used to draw the boundary points by connecting them, forming a complete outline, by traversing the coordinates of each outline point. Different colors or line thicknesses can be selected to draw the outline to enhance the visualization effect. In this application embodiment, a red box is used to mark the difference sub-images; that is, if the outline of a difference sub-image is located within a sub-image of the third image, red lines are drawn on the four sides of that sub-image to obtain a red box mark.
[0089] This application embodiment marks the difference sub-images in the difference image by using contour detection, thereby enhancing the attention to the differences between images, facilitating accurate text output of image differences in the subsequent process, and improving the efficiency and accuracy of image difference detection.
[0090] Step S103: Perform feature embedding processing on the difference image based on multiple difference sub-images to obtain the image feature sequence.
[0091] Here, after obtaining the difference image, it can be input into a pre-trained difference detection model, which outputs text representing the difference. The pre-trained difference detection model can be a large vision-language model. This model can consist of a Transformer model and a Large Language Model (LLM). The Transformer model processes the difference image and predicts the feature vector sequence corresponding to the text representing the difference. The LLM decodes the feature vector sequence corresponding to the text, obtaining the text representing the image difference. Feature embedding based on multiple difference sub-images involves applying a self-attention mechanism to the difference image and then weighting the results of the self-attention mechanism based on multiple difference sub-images using a multi-head attention mechanism to enhance attention to the difference sub-images. The image feature sequence includes the image vector corresponding to each sub-image.
[0092] In some embodiments, see Figure 5 , Figure 5 The step S103, which involves feature embedding of the difference image based on multiple difference sub-images to obtain an image feature sequence, can be achieved through the following steps S1031 to S1033:
[0093] Step S1031: Convolution processing is performed on the sub-image features of multiple difference sub-images to obtain the difference enhancement matrix.
[0094] Here, features can first be extracted from multiple difference sub-images to obtain the sub-image features of each sub-image. A convolutional neural network (CNN) is then used to convolve these sub-image features to obtain the difference enhancement matrix Wo. The difference enhancement matrix is used to apply attention weights to the difference sub-images. For example, in the difference enhancement matrix, the values of elements corresponding to the difference sub-images are larger, while the values of elements corresponding to the non-difference sub-images are smaller. Specifically, if the difference sub-images are marked with red boxes, and if the difference image contains red boxes, the difference enhancement matrix can enhance the coefficients of the image feature vectors of the difference sub-images, thereby significantly strengthening attention. Wo = CNN(Features) cnn (Frame diff ), kernel), where Frame diff For difference images, Feature cnn (Frame diff The image features are obtained by extracting features from the difference image using a convolutional neural network (CNN), where kernel is the convolution kernel.
[0095] Step S1032: Extract features from the differential image to obtain the first image feature sequence.
[0096] In this embodiment, feature extraction can be performed on the difference image to obtain the image feature vector of each sub-image in the difference image. The image feature vectors of multiple sub-images form a first image feature sequence. In practical applications, the difference image is input into a pre-trained difference detection model, and the visual encoder in the difference detection model will extract features from the difference image to obtain the first image feature sequence.
[0097] Step S1033: Determine the image feature sequence based on the first image feature sequence and the difference enhancement matrix.
[0098] In this embodiment, the first image feature sequence can be processed by self-attention first, and then the result of the self-attention processing can be weighted by the difference enhancement matrix to obtain the image feature sequence.
[0099] This application embodiment enhances the image feature vectors of the difference sub-images by using a difference enhancement matrix to significantly enhance attention, which facilitates accurate text output of image differences in the subsequent process, improves the efficiency and accuracy of image difference detection, and can also successfully detect differences even when the differences are small, thus solving the problem of insufficient sensitivity to changes in technical problem 2.
[0100] In some embodiments, see Figure 6 , Figure 6 The step S1033, which determines the image feature sequence based on the first image feature sequence and the difference enhancement matrix, can be achieved through the following steps S10331 to S10333:
[0101] Step S10331: Self-attention processing is performed on the first image feature sequence using a set of preset first linear transformation matrices to obtain a set of second image feature sequences.
[0102] Here, the difference detection model includes multiple self-attention mechanism modules. In each self-attention mechanism module, a first image feature sequence is processed using a pre-defined set of first linear transformation matrices to obtain a second image feature sequence. It should be noted that the multiple sets of first linear transformation matrices are learnable parameters during the training process of the difference detection model; the pre-defined sets of first linear transformation matrices become the model parameters in the trained difference detection model. Each set of first linear transformation matrices is different.
[0103] In some embodiments, the first linear transformation matrix set includes three first linear transformation matrices. Step S10331, which involves performing self-attention processing on the first image feature sequence using multiple preset sets of first linear transformation matrices to obtain multiple second image feature sequences, can be implemented as follows: First, perform the following operations on the first image feature sequence using each set of first linear transformation matrices to obtain the corresponding second image feature sequence; then, perform linear transformations on the first image feature sequence using the three first linear transformation matrices in the first linear transformation matrix set to obtain an image feature query sequence, an image feature key sequence, and an image feature value sequence; then, determine the first weight matrix based on the image feature query sequence and the image feature key sequence; finally, determine the second image feature sequence based on the first weight matrix and the image feature value sequence.
[0104] For example, the three first linear transformation matrices in the set of first linear transformation matrices can be matrix W. Q W K and W V In a self-attention mechanism module, matrix W can be used. Q A linear transformation is performed on the first image feature sequence to obtain the image feature query sequence Q1; using matrix W... K A linear transformation is performed on the first image feature sequence to obtain the image feature key sequence K1; using matrix W VA linear transformation is performed on the first image feature sequence to obtain the image feature value sequence V1. A first weight matrix W1 is calculated using the image feature query sequence Q1 and the image feature key sequence K1. The first weight matrix W1 is multiplied by the image feature value sequence V1 to obtain the second image feature sequence. It should be noted that in this embodiment, the feature sequences are composed of multiple feature vectors; therefore, a feature sequence can also be considered a matrix and can be directly multiplied by other matrices. For example, if the image feature query sequence Q1 includes 100 vectors, each with a dimension of 100, then the image feature query sequence Q1 is a 100×100 matrix.
[0105] This application embodiment obtains multiple second image feature sequences by performing self-attention processing on the first image feature sequence, calculates the correlation between each element in the first image feature sequence, realizes the interaction and correlation of information between different positions, effectively utilizes the global information of the difference image, helps the difference detection model better understand the structure and correlation between the image and the text, and improves the representation ability and generalization ability of the difference detection model.
[0106] In this embodiment, the first weight matrix is determined based on the image feature query sequence and the image feature key sequence. This can be achieved by: first, determining the first similarity matrix between the image feature query sequence and the image feature key sequence; and obtaining the dimension of the image feature query sequence; then, determining the intermediate matrix based on the first similarity matrix and the dimension; and finally, normalizing the intermediate matrix to obtain the first weight matrix.
[0107] For example, the first similarity matrix Q1K1 can be determined by multiplying the image feature query sequence Q1 by the transpose of the image feature key sequence K1. T The dimension of the image feature query sequence is the dimension d of the vectors in the image feature query sequence. k The first similarity matrix Q1K1 T With respect to dimension d k The result after taking the square root The ratio of these values is used to determine the intermediate matrix. The softmax function can be used to normalize the intermediate matrix to obtain the first weight matrix W1.
[0108] This application embodiment determines the first weight matrix by using the first similarity matrix between the image feature query sequence and the image feature key sequence, calculates the correlation between each element in the first image feature sequence, realizes the interaction and correlation of information between different positions, effectively utilizes the global information of the difference image, helps the difference detection model better understand the structure and correlation between the image and the text, and improves the representation ability and generalization ability of the difference detection model.
[0109] Step S10332: Concatenate multiple second image feature sequences to obtain an intermediate image feature sequence.
[0110] Here, after obtaining multiple second image feature sequences, the multiple second image feature sequences can be directly concatenated end to end to obtain an intermediate image feature sequence. For example, the second image feature sequence A is (a1, a2, a3), the second image feature sequence B is (b1, b2, b3), and the intermediate image feature sequence obtained by concatenating the second image feature sequences A and B is (a1, a2, a3, b1, b2, b3).
[0111] Step S10333: The intermediate image feature sequence is weighted by the difference enhancement matrix to obtain the image feature sequence.
[0112] Here, the intermediate image feature sequence can be multiplied by the difference enhancement matrix to obtain the image feature sequence. The number of vectors in the intermediate image feature sequence is the same as the number of rows in the difference enhancement matrix, and the dimension of the vectors in the intermediate image feature sequence is the same as the number of columns in the difference enhancement matrix.
[0113] This application embodiment implements a multi-head attention mechanism by weighting a difference enhancement matrix after concatenating multiple second image feature sequences. This significantly enhances the attention to the difference sub-images, facilitating accurate text output of image differences and improving the efficiency and accuracy of image difference detection.
[0114] Step S104: Based on a preset image-text dictionary, perform text mapping on the image feature sequence to obtain a text feature sequence.
[0115] The text feature sequence is the feature sequence corresponding to multiple differential sub-images.
[0116] Here, the image-text dictionary includes multiple images and one or more visual word vectors that have a mapping relationship with each image. A visual word is the word corresponding to an image. For example, if image 1 is a cat, then the visual word that has a mapping relationship with image 1 is "cat". The visual word vector that has a mapping relationship with image 1 is the feature vector of the visual word "cat". The image-text dictionary can be a sample set pre-constructed manually based on a large number of images and texts, or it can be a standard set constructed through a model. This application embodiment does not limit the method of obtaining the image-text dictionary. In this application embodiment, when performing text mapping on the image feature sequence, only the labeled difference sub-images are text mapped to obtain the text feature sequence. The text feature sequence includes the text feature vector corresponding to each difference sub-image.
[0117] In some embodiments, see Figure 7 , Figure 7The step S104, which involves mapping the image feature sequence to text based on a preset image-text dictionary, to obtain the text feature sequence, can be achieved through the following steps S1041 to S1043:
[0118] Step S1041: Based on a preset image-text dictionary, perform text mapping on the image feature sequence to obtain an initial visual word sequence.
[0119] The initial visual word sequence consists of vectors of multiple visual words, with each differential sub-image corresponding to one visual word.
[0120] Here, the initial visual word sequence includes the vector of the visual word corresponding to each differential sub-image, obtained directly from the image-text dictionary. Text mapping of the image feature sequence refers to retrieving the visual word vector that has a mapping relationship with each labeled differential sub-image from a pre-defined image-text dictionary.
[0121] Step S1042: Perform position embedding processing on the initial visual word sequence to obtain the target visual word sequence.
[0122] Here, for each visual word vector in the initial visual word sequence, we obtain the number N of that visual word vector (i.e., which vector in the initial visual word sequence it is, counting from 0) and the vector dimension C of that visual word vector. Based on the number N and the vector dimension C of the visual word vectors, we determine the positional encoding vector PE(N,C) of that visual word vector. We then directly add the visual word vector to each element of the positional encoding vector PE(N,C) to obtain the positionally embedded visual word vector. After performing the above positional embedding operation on each visual word vector in the initial visual word sequence, multiple positionally embedded visual word vectors form the target visual word sequence. It should be noted that the positional encoding vector PE is a mathematical function composed of a sine function and a cosine function, and the positional encoding vector PE satisfies the following formulas (I) and (II).
[0123]
[0124]
[0125] Here, pos represents the position of each visual word vector within the entire initial visual word sequence (i.e., N), d represents the dimension of the visual word vector (i.e., C), for example, 512 dimensions, and i is the index of the visual word vector divided by 2 and rounded down, with a value range of [0, ..., d / 2]. That is, 2i refers to the even-numbered dimensions of the vector, i.e., the 0th, 2nd, 4th, ..., 510th dimensions, calculated using the sine function; 2i+1 refers to the odd-numbered dimensions of the vector, i.e., the 1st, 3rd, 5th, ..., 511th dimensions, calculated using the cosine function.
[0126] For example, if the text is "It is a cat", then the visual word vector of the visual word "cat" has a pos (or N) of 4, i = 2, and the dimension d (or C) of the visual word vector is 512. The position encoding vector of the visual word vector of "cat" can be calculated using the formula.
[0127] Step S1043: Encode the target visual word sequence to obtain the text feature sequence.
[0128] Here, the encoder layer in the differential detection model is used to encode the target visual word sequence to obtain the text feature sequence. The differential detection model can have one or more encoder layers.
[0129] This application embodiment obtains the target visual word sequence by performing position embedding processing on the initial visual word sequence, thereby realizing the relative position information between the visual words in the steps and enhancing the expressive power of the visual words.
[0130] In this embodiment, encoding the target visual word sequence to obtain a text feature sequence can be achieved in the following way: First, the target visual word sequence is subjected to self-attention processing through a preset set of second linear transformation matrices to obtain a first text feature sequence; then, the sum of the first text feature sequence and the target visual word sequence is subjected to layer normalization processing to obtain a second text feature sequence; then, the second text feature sequence is subjected to activation processing to obtain a third text feature sequence; finally, the sum of the third text feature sequence and the second text feature sequence is subjected to layer normalization processing to obtain a text feature sequence.
[0131] Here, the encoder layer in the differential detection model has a self-attention sublayer and a feedforward sublayer. The self-attention sublayer performs three different linear transformations on the input target visual word sequence to obtain a text feature query sequence, a text feature key sequence, and a text feature value sequence. Then, it calculates the dot product of the text feature query sequence and the text feature key sequence, and scales, masks, and normalizes this dot product to obtain a second weight matrix. Finally, it calculates the product of the second weight matrix and the text feature value sequence to obtain the first text feature sequence, which is the output matrix of the self-attention sublayer. The feedforward sublayer performs two linear transformations and one non-linear activation on the output matrix of the self-attention sublayer to obtain its output, which represents the non-linear transformation of the visual words. However, the differential detection model also has a residual connection between the output and input of each sublayer, which adds the sublayer's output and input, and then performs a layer normalization operation, normalizing the feature vector of each visual word to have a mean of 0 and a variance of 1. The target visual word sequence is processed by multiple encoder layers in the differential detection model to obtain the text feature sequence.
[0132] In this embodiment, the specific implementation of obtaining the first text feature sequence by performing self-attention processing on the target visual word sequence using a preset set of second linear transformation matrices can be referred to the embodiment of obtaining the second image feature sequence in step S10331 above, and will not be repeated here. Since both the first text feature sequence and the target visual word sequence can be represented in matrix form, and the number of rows and columns are the same, the elements at the same position in the first text feature sequence and the target visual word sequence can be directly added to obtain the sum of the first text feature sequence and the target visual word sequence. The sum of the first text feature sequence and the target visual word sequence is normalized using the LayerNorm() function to obtain the second text feature sequence. The second text feature sequence is activated using the ReLU activation function to obtain the third text feature sequence. The elements at the same position in the third text feature sequence and the second text feature sequence are added to obtain the sum of the third text feature sequence and the second text feature sequence, and the sum of the third text feature sequence and the second text feature sequence is normalized using the LayerNorm() function to obtain the text feature sequence.
[0133] It should be noted that the process of performing self-attention processing on the first image feature sequence using the preset first linear transformation matrix set and the process of performing self-attention processing on the target visual word sequence using the preset second linear transformation matrix set can be performed in the same self-attention mechanism module in the difference detection model. In this case, the first linear transformation matrix set and the second linear transformation matrix set are the same.
[0134] This application embodiment improves the convergence speed and generalization ability of the difference detection model by performing residual connection and layer normalization on the first text feature sequence, thereby improving the image difference detection speed and solving the problem of low robustness to interference factors in technical problem 4.
[0135] In some embodiments, the second linear transformation matrix set includes three second linear transformation matrices. The first text feature sequence is obtained by performing self-attention processing on the target visual word sequence using a preset set of second linear transformation matrices, which can be achieved as follows: First, the target visual word sequence is linearly transformed using the three second linear transformation matrices in the second linear transformation matrix set, respectively, to obtain a text feature query sequence, a text feature key sequence, and a text feature value sequence; then, a second weight matrix is determined based on the text feature query sequence and the text feature key sequence; finally, the first text feature sequence is determined based on the second weight matrix and the text feature value sequence.
[0136] Here, the specific process of using three second linear transformation matrices to perform linear transformation on the target visual word sequence, and the specific process of determining the second weight matrix based on the text feature query sequence and the text feature key sequence, can all refer to the embodiment of obtaining the second image feature sequence in step S10331 above, and will not be repeated here.
[0137] This application embodiment obtains a first text feature sequence by performing self-attention processing on the target visual word sequence, calculates the correlation between each element in the target visual word sequence, realizes the interaction and correlation of information between different positions, helps the difference detection model to better understand the structure and correlation between images and text, and improves the representation ability and generalization ability of the difference detection model.
[0138] Step S105: Based on the text feature sequence, determine the target text used to characterize the difference between the first image and the second image.
[0139] Here, the text feature sequence can be input into the Large Language Model (LLM) in the pre-trained difference detection model. The LLM can directly compile the text feature sequence into target text that represents the difference between the first and second images.
[0140] In some embodiments, the text feature sequence can be concatenated with the image feature sequence and then input into the Large Language Model (LLM) in a pre-trained difference detection model to obtain the target text. This method adds information from the image feature sequence, which can improve the semantic accuracy of the target text in expressing the differences between the first and second images.
[0141] This application embodiment performs differential processing on a first image and a second image to obtain a differential image. Based on multiple difference sub-images representing the differences between the first and second images in the differential image, feature embedding processing is performed on the differential image to obtain an image feature sequence. This image feature sequence can focus on the difference points between the first and second images, and thus the text feature sequence obtained after text mapping of the image feature sequence is a feature sequence corresponding to the multiple difference sub-images, that is, a feature sequence of text used to describe the difference points between the first and second images. Therefore, based on the text feature sequence, the target text representing the difference between the first and second images can be directly determined. This application embodiment greatly reduces the workload of manual comparison, can quickly and accurately analyze a large number of images, improves the detection rate of differences between images, and can also transform the differences between images into intuitive text descriptions.
[0142] In some embodiments, the image difference detection method provided in this application is applied to a pre-trained difference detection model. The difference detection model can be trained in the following way: First, training data is acquired, which includes multiple sample difference images and label text corresponding to each sample difference image. The sample difference images include multiple sample difference sub-images used to characterize the difference between two sample images. Then, using the feature embedding layer of the difference detection model, the sample difference images are processed by feature embedding through the multiple sample difference sub-images to obtain a sample image feature sequence. Next, using the feature mapping layer of the difference detection model, the sample image feature sequence is mapped to text to obtain a sample text feature sequence. The sample text feature sequence is a feature sequence corresponding to the multiple sample difference sub-images. Then, using the text decoding layer of the difference detection model, the sample text used to characterize the difference between the first sample image and the second sample image is determined based on the sample text feature sequence. A first loss result is determined based on the sample image feature sequence and the sample text feature sequence. A second loss result is determined based on the label text and the sample text. Finally, the model parameters in the difference detection model are updated using the first loss result and the second loss result to obtain the trained difference detection model.
[0143] Here, the label text is manually annotated semantic text used to identify the differences between two sample images. The specific implementation method for obtaining the sample difference image based on the two sample images can refer to step S102 above, and will not be repeated here. The feature embedding layer of the difference detection model can include multiple self-attention mechanism modules and multi-head attention modules. The specific method for obtaining the sample image feature sequence by performing feature embedding processing on the sample difference image through multiple sample difference sub-images can refer to step S103 above, and will not be repeated here. The process of using the feature mapping layer of the difference detection model to perform text mapping on the sample image feature sequence to obtain the sample text feature sequence can refer to step S104 above, and will not be repeated here. The text decoding layer of the difference detection model can be a Large Language Model (LLM). The dot product of the transpose of the sample image feature sequence and the sample text feature sequence can be calculated. The ratio of this dot product to the square root of the dimension of the vector in the sample image feature sequence is normalized to obtain the similarity between the sample image feature sequence and the sample text feature sequence, and this similarity is used as the first loss result. The cross-entropy loss between the label text and the sample text can be calculated as the second loss result. By using the average of the first and second loss results, the learnable model parameters in the difference detection model are iterated backward using stochastic gradient descent until the average of the first and second loss results reaches its minimum value, thus obtaining the trained difference detection model.
[0144] The embodiments of this application train the difference detection model by using two different loss calculations, namely the first loss result and the second loss result, which can improve the inference performance of the difference detection model and thus improve the accuracy of image difference detection.
[0145] Figure 8 This is another optional flowchart illustrating the image difference detection method provided in the embodiments of this application, such as... Figure 8 As shown, the method includes the following steps S201 to S210:
[0146] Step S201: The terminal receives the user's interactive operation.
[0147] Here, interactive operations can include image input operations, clicking to start detection operations, etc. Users can perform any kind of interactive operation on the client of the image difference detection application through the terminal, such as inputting a large number of images to be detected.
[0148] In step S202, the terminal responds to the interactive operation by generating an image difference detection request.
[0149] When the terminal receives a user interaction, it determines that image difference detection is needed. Therefore, in response to the interaction, it generates an image difference detection request. This request asks the server to perform image difference detection on the first and second images input by the user.
[0150] In step S203, the terminal sends the image difference detection request to the server.
[0151] In step S204, the server responds to the image difference detection request and obtains the first image and the second image to be detected.
[0152] Here, in the scenario of batch image difference detection, the server can randomly select two images from multiple images to be detected input by the user as the first image and the second image.
[0153] In step S205, the server performs differential processing on the first image and the second image to obtain a differential image.
[0154] Here, the difference image includes multiple difference sub-images used to characterize the differences between the first image and the second image. For a specific implementation of the difference processing of the first image and the second image, please refer to the description of step S102 above, which will not be repeated here.
[0155] In step S206, the server performs feature embedding processing on the differential image based on multiple differential sub-images to obtain an image feature sequence.
[0156] For a detailed implementation of feature embedding processing of the difference image based on multiple difference sub-images, please refer to the description of step S103 above, which will not be repeated here.
[0157] In step S207, the server performs text mapping on the image feature sequence based on a preset image-text dictionary to obtain a text feature sequence.
[0158] Here, the text feature sequence is the feature sequence corresponding to multiple differential sub-images. For the specific implementation of text mapping of image feature sequences based on a preset image-text dictionary, please refer to the description of step S104 above, which will not be repeated here.
[0159] In step S208, the server determines the target text used to characterize the difference between the first image and the second image based on the text feature sequence.
[0160] For a specific implementation of determining the target text used to characterize the difference between the first image and the second image, please refer to the description of step S105 above, which will not be repeated here.
[0161] In step S209, the server sends the target text to the terminal.
[0162] Step S210: The terminal displays the target text on the current interface.
[0163] This application embodiment performs differential processing on a first image and a second image to obtain a differential image. Based on multiple difference sub-images representing the differences between the first and second images in the differential image, feature embedding processing is performed on the differential image to obtain an image feature sequence. This image feature sequence can focus on the difference points between the first and second images, and thus the text feature sequence obtained after text mapping of the image feature sequence is a feature sequence corresponding to the multiple difference sub-images, that is, a feature sequence of text used to describe the difference points between the first and second images. Therefore, based on the text feature sequence, the target text representing the difference between the first and second images can be directly determined. This application embodiment greatly reduces the workload of manual comparison, can quickly and accurately analyze a large number of images, improves the detection rate of differences between images, and can also transform the differences between images into intuitive text descriptions.
[0164] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0165] This application provides an image difference detection method, which is an automated detection method for game scene differences combining a large visual-language model. Applied to a large-scale automated detection system based on a vision-to-text model, it focuses on efficiently detecting and analyzing large numbers of images in game scenes. It utilizes a powerful vision-to-text model to perform efficient and accurate scene analysis of game resources and scenes. The large visual-language model employed in this method can transform information in images into semantically understood text descriptions, not only improving detection speed but also providing a deeper understanding and description of each image. By converting visual information into text, it achieves a more detailed and accurate capture of image semantics, making this image difference detection method perform exceptionally well when processing large-scale scene images.
[0166] The image difference detection method provided in this application first performs visual-lexical modeling of the game scene content, and then performs image differencing on subtle differences. Next, an attention mechanism is established to focus on the differences during visual detection. Furthermore, the image difference detection method provided in this application has automated detection capabilities, greatly reducing the workload of manual comparison. It can quickly and accurately analyze large numbers of images and output easily understandable text descriptions, enabling game operators to quickly identify and resolve problems. This automation not only improves efficiency but also reduces the risk of human error. Overall, this application provides a large-scale automated detection system for vision-to-text models, offering a reliable solution for processing images in large-scale game scenes, and significantly improving processing efficiency and detection accuracy.
[0167] This application embodiment can be used for the automation of resource scene differences in games. By automatically providing difference images (corresponding to the first and second images in the above embodiment), semantic differences (corresponding to the target text in the above embodiment) can be quickly output. Figure 9 This is a schematic diagram of a game scene provided in an embodiment of this application. Figure 10 This is another game scene illustration provided in an embodiment of this application. Embodiments of this application can... Figure 9 and Figure 10 The game scene diagram is used for difference detection to obtain semantic output results: Figure 10 ( Figure 10 The area marked by the middle grid is a red pedestrian walkway (not shown in color in the image). There is a red brick pedestrian walkway, while... Figure 9 No; Figure 9 ( Figure 9 The dashed line represents the green outline of the sidewalk (not shown in color in the diagram). The sidewalk is surrounded by a green outline. Figure 10 No.
[0168] Figure 11 This is a schematic diagram of the semantic matching mechanism of the visual-language big model in the embodiments of this application. Figure 11In this framework, Q represents the image and T represents the text. First, features are extracted from the image. Then, an attention mechanism is used to build a rich QKV combination. The Query (to match others) serves as the input information, guiding the search and indicating what information is needed. For example, entering "gray men's sweater" in the search bar aims to return similar products. The Key (to be matched) represents other products to be matched; after receiving the information "gray" and "men's sweater," the system matches all products in the database. Attention(Q, K) represents the degree of matching between the Query and Key (since there are many products (Keys) in the system, products matching the description (Query) have a higher degree of matching). The Value (information to be extracted) is the information itself, simply expressing the input features. Image-text matching is performed using the established QKV combination to obtain a standard image-text set. Then, image-text difference learning is performed to finally obtain a pre-defined dictionary set. The pre-defined dictionary set contains a mapping relationship between images and text.
[0169] Figure 12 This is a schematic diagram illustrating the image difference detection process using the visual-language large model provided in this application embodiment. See also... Figure 12 First, differential labeling is performed on the two differencing resources (input image 1 and input image 2), marking each difference point to obtain the processed difference image. Frame diff =Frame1-Frame2, Frame diff The images are divided into two parts: Frame 1 represents input image 1, and Frame 2 represents input image 2. Multiple bounding boxes in the divided images mark the positions of the difference sub-images between input image 1 and input image 2. The divided images are then input into a pre-trained visual-language large model to achieve direct semantic output of the differences. The visual-language large model includes a visual-language learner and a Large Language Model (LLM). The visual-language learner includes a visual encoding module and a Transformer model. The visual-language large model can be trained using learnable query vectors (queries) and a pre-trained corpus of text-image models. The pre-trained corpus of text-image models includes text T and image sequences Q. The Transformer algorithm is used to fuse text T and image sequences Q, matching feature regions and text portions. Pre-training is performed using a large amount of scene text and images to obtain the trained Transformer model and the Large Language Model LLM, which is the pre-trained visual-language large model. The following section details the processing steps of the large vision-language model.
[0170] CNNs are a crucial component in the Transformer model for extracting massive image features. The visual encoding module can be a pre-trained Convolutional Neural Network (CNN). The CNN extracts features from the differenced images, resulting in an image feature sequence Q. Here, Lq is the length of the image feature sequence Q, and D is the dimension of each feature.
[0171] The Transformer model is fused. First, Queries, Keys, and Values are constructed. To introduce image features into the Transformer model, an attention mechanism for Queries, Keys, and Values needs to be constructed. For example, the same image feature sequence Q can be used as Queries, Keys, and Values, where Queries = Q, Keys = Q, and Values = Q. Then, the self-attention mechanism is calculated. Different weights are assigned to each position based on the similarity between Queries and Keys, and these weights are applied to Values to obtain the final attention output. The self-attention mechanism calculation can satisfy the following formula (1).
[0172]
[0173] Where Attention(Q,K,V) is the output of the self-attention mechanism (corresponding to the second image feature sequence in the above embodiment), softmax is the normalization exponential function, and K... T d is the transpose of K. k It is the dimension of Queries and Keys.
[0174] Next, the multi-head attention mechanism is calculated. To enhance the model's expressive power, a multi-head attention mechanism can be used, applying the self-attention mechanism to multiple subspaces and concatenating their outputs to obtain the final multi-head attention output. The calculation of multiple attention mechanisms can satisfy the following formula (2).
[0175] MultiHead(Q,K,V)=Concat(head1,…,head h W O Formula (2);
[0176] Where MultiHead(Q,K,V) is the multi-head attention output (corresponding to the image feature sequence in the above embodiment), Concat is concatenation, head1 is the output of the first self-attention module, and head... h For the output of the h-th self-attention module, head1 = Attention(QW Qi KW KiVW Vi ), W Qi W Ki W Vi This is the weight matrix corresponding to Queries, Keys, and Values. Wo is a matrix that weights the attention weights for the difference sub-images (corresponding to the difference enhancement matrix in the above embodiment). Wo = CNN(Feature) cnn (Frame diff ), kernel), where Feature cnn (Frame diff The image features are obtained by feature extraction from the difference image using a CNN, and the kernel is the convolution kernel. An operation similar to high-frequency enhancement is used to enhance the red boxes (marked difference sub-images) in the difference image. Specifically, if a red box is present in the convolution region, the coefficients are enhanced, thus significantly increasing attention.
[0177] Transformer processing of semantic path embedding is performed. For text T, it needs to be embedded into a sequence, where Lt is the length of the text sequence and d is the dimension of each embedded feature. First, position embedding is performed. In order for the Transformer model to capture the relative position information between visual words, some position embeddings need to be added to the visual words. These position embeddings can be fixed or learnable, with the aim of enhancing the expressive power of the visual words. The position embeddings satisfy the following formula (3).
[0178] Z (0) =V+PE(N,C) Formula (3);
[0179] Where V is the matrix of visual words, PE is the position embedding operation, N is the number of visual words, C is the dimension of the visual words, and Z is the position embedding operation. (0) It is a visual word matrix with location information added.
[0180] Then, the encoder layer is processed. The encoder layer is the core component of the Transformer model. The Transformer model consists of multiple encoder layers. Each encoder layer contains a self-attention sublayer and a feedforward sublayer. Both sublayers have a residual connection and a layer normalization operation, the purpose of which is to extract the global semantic relationship between visual words. The processing of the encoder layer satisfies the following formula (4).
[0181] Z (l) =Encoder(Z) (l-1) ) formula (4);
[0182] Encoder() is an encoder layer, Z (l-1)It is the input of the (l-1)th encoder layer, Z (l) It is the output of the l-th encoder layer, where l = 1, 2, ..., L, and L is the number of encoder layers.
[0183] The self-attention sublayer is the first sublayer in the encoder layer. It performs three different linear transformations on the input visual word matrix to obtain three matrices, called the query matrix, key matrix, and value matrix, respectively. Then, by calculating the dot product of the query matrix and the key matrix, an attention score matrix is obtained. After scaling, masking, and normalizing the attention score matrix, an attention weight matrix (corresponding to the second weight matrix in the above embodiment) is obtained. Finally, by calculating the product of the attention weight matrix and the value matrix, an output matrix is obtained. This output matrix is the result of the self-attention sublayer, which can represent the dependencies between visual words. The processing of the self-attention sublayer satisfies the following formulas (5), (6), and (7).
[0184] Q = Z (l-1) W Q K = Z (l-1) W Q V = Z (l-1) W V Formula (5);
[0185]
[0186]
[0187] Among them, W Q W Q W V It is a learnable linear transformation matrix, where Q, K, and V are the query matrix, key matrix, and value matrix, respectively. A is the attention weight matrix, and M is a masking matrix used to mask some irrelevant visual words. It is the output matrix of the self-attention sublayer.
[0188] The feedforward sublayer is the second sublayer in the encoder layer. It obtains an output matrix by performing two linear transformations and one nonlinear activation on the output matrix of the self-attention sublayer. This output matrix is the result of the feedforward sublayer, which can represent the nonlinear transformation of the visual word. The processing of the feedforward sublayer satisfies the following formula (8).
[0189]
[0190] in, is the output matrix of the feedforward sublayer, ReLU is the activation function, and W1, b1, W2 and b2 are all learnable model parameters.
[0191] Residual connections and layer normalization: To prevent gradient vanishing or exploding caused by the depth of the Transformer model, and to improve the convergence speed and generalization ability of the model, a residual connection needs to be added between the output and input of each sub-layer. That is, the output and input of the sub-layer are added together, and then a layer normalization operation is performed, that is, the feature vector of each visual word is normalized so that its mean is 0 and its variance is 1. The residual connection between the output and input of the self-attention sub-layer satisfies the following formula (9). The residual connection between the output and input of the feedforward sub-layer satisfies the following formula (10).
[0192]
[0193]
[0194] LayerNorm() is a layer normalization operation, Z (1) This is the final output of the first encoder layer.
[0195] During model training, image-text semantic comparison can also be performed, comparing the multi-head attention output of image features with the text embedding. A comparison score can be obtained by calculating their similarity. The comparison score can be used to guide the model in establishing correspondences between images and text. The similarity calculation satisfies the following formula (11).
[0196]
[0197] Where d is the dimension of the multi-head attention output of the image features and the text embedding, and Similarity(Q,T) is the similarity.
[0198] Finally, the semantics and corresponding image regions are output. Similar to object detection methods, regions in the image corresponding to the text can be determined by comparing scores. By introducing positional encoding, the model can learn the relative positional information between different regions in the image.
[0199] The embodiments of this application can replace the manual review method in the field of resource scene comparison in games, improve the efficiency of large-scale image difference detection, and can directly output semantic differences.
[0200] It is understood that in the embodiments of this application, data such as user information are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0201] The following description continues to illustrate the exemplary structure of the image difference detection device 455 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules in the image difference detection device 455 stored in the memory 450 may include: an image acquisition module 4551, used to acquire a first image and a second image to be detected; an image difference module 4552, used to perform difference processing on the first image and the second image to obtain a difference image; the difference image includes multiple difference sub-images for characterizing the differences between the first image and the second image; a feature embedding module 4553, used to perform feature embedding processing on the difference image based on the multiple difference sub-images to obtain an image feature sequence; a text mapping module 4554, used to perform text mapping on the image feature sequence based on a preset image-text dictionary to obtain a text feature sequence; the text feature sequence is a feature sequence corresponding to the multiple difference sub-images; and a text determination module 4555, used to determine target text for characterizing the differences between the first image and the second image based on the text feature sequence.
[0202] In some embodiments, the image difference module 4552 is further configured to subtract the pixel value of the corresponding position in the second image from each pixel value of the first image to obtain a third image; for each pixel in the third image, when the pixel value of the pixel is greater than a preset threshold, the pixel is marked as a difference region; contour detection is performed on the third image marked with difference regions to obtain the contour of each difference sub-image; and the contour of each difference sub-image is marked in the third image with a preset style to obtain the difference image.
[0203] In some embodiments, the feature embedding module 4553 is further configured to perform convolution processing on the sub-image features of the plurality of difference sub-images to obtain a difference enhancement matrix; perform feature extraction on the difference image to obtain a first image feature sequence; and determine the image feature sequence based on the first image feature sequence and the difference enhancement matrix.
[0204] In some embodiments, the feature embedding module 4553 is further configured to perform self-attention processing on the first image feature sequence through a preset set of multiple first linear transformation matrices respectively, to obtain multiple second image feature sequences; to concatenate the multiple second image feature sequences to obtain an intermediate image feature sequence; and to perform weighted processing on the intermediate image feature sequence through the difference enhancement matrix to obtain the image feature sequence.
[0205] In some embodiments, the first linear transformation matrix set includes three first linear transformation matrices; the feature embedding module 4553 is further configured to perform the following operations on the first image feature sequence through each of the first linear transformation matrix sets to obtain a corresponding second image feature sequence: performing linear transformations on the first image feature sequence through the three first linear transformation matrices in the first linear transformation matrix set to obtain an image feature query sequence, an image feature key sequence, and an image feature value sequence; determining a first weight matrix based on the image feature query sequence and the image feature key sequence; and determining the second image feature sequence based on the first weight matrix and the image feature value sequence.
[0206] In some embodiments, the feature embedding module 4553 is further configured to: determine a first similarity matrix between the image feature query sequence and the image feature key sequence; obtain the dimension of the image feature query sequence; determine an intermediate matrix based on the first similarity matrix and the dimension; and normalize the intermediate matrix to obtain the first weight matrix.
[0207] In some embodiments, the text mapping module 4554 is further configured to perform text mapping on the image feature sequence based on a preset image-text dictionary to obtain an initial visual word sequence; the initial visual word sequence includes vectors of multiple visual words, and each of the differential sub-images corresponds to one visual word; the initial visual word sequence is subjected to position embedding processing to obtain a target visual word sequence; the target visual word sequence is subjected to encoding processing to obtain the text feature sequence.
[0208] In some embodiments, the text mapping module 4554 is further configured to perform self-attention processing on the target visual word sequence through a preset second set of linear transformation matrices to obtain a first text feature sequence; perform layer normalization processing on the sum of the first text feature sequence and the target visual word sequence to obtain a second text feature sequence; perform activation processing on the second text feature sequence to obtain a third text feature sequence; and perform layer normalization processing on the sum of the third text feature sequence and the second text feature sequence to obtain the text feature sequence.
[0209] In some embodiments, the second linear transformation matrix set includes three second linear transformation matrices; the text mapping module 4554 is further configured to perform linear transformations on the target visual word sequence using the three second linear transformation matrices in the second linear transformation matrix set, respectively, to obtain a text feature query sequence, a text feature key sequence, and a text feature value sequence; determine a second weight matrix based on the text feature query sequence and the text feature key sequence; and determine the first text feature sequence based on the second weight matrix and the text feature value sequence.
[0210] In some embodiments, the image difference detection method is applied to a pre-trained difference detection model. The image difference detection device 455 further includes a model training module for training the difference detection model through the following steps: acquiring training data, the training data including multiple sample difference images and label text corresponding to each sample difference image, the sample difference images including multiple sample difference sub-images for characterizing the difference between two sample images; using the feature embedding layer of the difference detection model, performing feature embedding processing on the sample difference images through the multiple sample difference sub-images to obtain a sample image feature sequence; using the feature mapping layer of the difference detection model, ... The sample image feature sequence is mapped to text to obtain a sample text feature sequence; the sample text feature sequence is a feature sequence corresponding to the plurality of sample difference sub-images; using the text decoding layer of the difference detection model, based on the sample text feature sequence, sample text is determined to characterize the difference between the first sample image and the second sample image; a first loss result is determined based on the sample image feature sequence and the sample text feature sequence; a second loss result is determined based on the label text and the sample text; the model parameters in the difference detection model are updated using the first loss result and the second loss result to obtain the trained difference detection model.
[0211] This application provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the image difference detection method described in this application.
[0212] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the image difference detection method provided in this application. For example, ... Figure 3 The image difference detection method is shown.
[0213] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0214] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0215] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0216] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0217] In summary, the embodiments of this application can quickly and accurately analyze a large number of images, improve the detection rate of differences between images, and transform the differences between images into intuitive text descriptions. It also has good accuracy in detecting images with minor differences and is highly robust to interference factors. When applied to resource scene comparison in games, it can replace manual review and improve the speed of image comparison.
[0218] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A method for detecting differences in images, characterized in that, The method includes: Acquire the first and second images to be detected; The first image and the second image are subjected to difference processing to obtain a difference image; the difference image includes a plurality of difference sub-images used to characterize the differences between the first image and the second image; The difference image is subjected to feature embedding processing based on multiple difference sub-images to obtain an image feature sequence; Based on a preset image-text dictionary, the image feature sequence is mapped to text to obtain a text feature sequence; the text feature sequence is a feature sequence corresponding to the plurality of differential sub-images; Based on the text feature sequence, target text is determined to characterize the differences between the first image and the second image.
2. The method according to claim 1, characterized in that, The step of performing differential processing on the first image and the second image to obtain a differential image includes: Subtract the corresponding pixel value in the second image from each pixel value in the first image to obtain the third image; For each pixel in the third image, when the pixel value of the pixel is greater than a preset threshold, the pixel is marked as a difference region; Contour detection is performed on the third image with marked difference regions to obtain the contour of each difference sub-image; The contours of each difference sub-image are marked in the third image using a preset style to obtain the difference image.
3. The method according to claim 1, characterized in that, The step of performing feature embedding processing on the difference image based on multiple difference sub-images to obtain an image feature sequence includes: The sub-image features of the multiple differential sub-images are convolved to obtain the differential enhancement matrix; Feature extraction is performed on the difference image to obtain a first image feature sequence; The image feature sequence is determined based on the first image feature sequence and the difference enhancement matrix.
4. The method according to claim 3, characterized in that, Determining the image feature sequence based on the first image feature sequence and the difference enhancement matrix includes: The first image feature sequence is subjected to self-attention processing by a set of preset first linear transformation matrices, thereby obtaining a set of second image feature sequences. The multiple second image feature sequences are concatenated to obtain an intermediate image feature sequence; The intermediate image feature sequence is obtained by weighting the difference enhancement matrix.
5. The method according to claim 4, characterized in that, The first set of linear transformation matrices includes three first linear transformation matrices; The process involves performing self-attention processing on the first image feature sequence using a preset set of multiple first linear transformation matrices, resulting in multiple second image feature sequences, including: Perform the following operations on the first image feature sequence for each of the first set of linear transformation matrices to obtain the corresponding second image feature sequence: The first image feature sequence is linearly transformed by the three first linear transformation matrices in the first linear transformation matrix set, respectively, to obtain the image feature query sequence, the image feature key sequence, and the image feature value sequence. Based on the image feature query sequence and the image feature key sequence, a first weight matrix is determined; The second image feature sequence is determined based on the first weight matrix and the image feature value sequence.
6. The method according to claim 5, characterized in that, The step of determining the first weight matrix based on the image feature query sequence and the image feature key sequence includes: Determine the first similarity matrix between the image feature query sequence and the image feature key sequence; Obtain the dimension of the image feature query sequence; The intermediate matrix is determined based on the first similarity matrix and the dimension; The intermediate matrix is normalized to obtain the first weight matrix.
7. The method according to claim 1, characterized in that, The text feature sequence is obtained by mapping the image feature sequence to text based on a preset image-text dictionary, including: Based on a pre-defined image-text dictionary, the image feature sequence is mapped to text to obtain an initial visual word sequence; the initial visual word sequence includes vectors of multiple visual words, and each of the differential sub-images corresponds to one visual word; The initial visual word sequence is subjected to position embedding processing to obtain the target visual word sequence; The target visual word sequence is encoded to obtain the text feature sequence.
8. The method according to claim 7, characterized in that, The process of encoding the target visual word sequence to obtain the text feature sequence includes: The target visual word sequence is subjected to self-attention processing by a preset set of second linear transformation matrices to obtain a first text feature sequence. The sum of the first text feature sequence and the target visual word sequence is subjected to layer normalization to obtain the second text feature sequence; The second text feature sequence is activated to obtain the third text feature sequence; The sum of the third text feature sequence and the second text feature sequence is subjected to layer normalization to obtain the text feature sequence.
9. The method according to claim 8, characterized in that, The second set of linear transformation matrices includes three second linear transformation matrices; The step of performing self-attention processing on the target visual word sequence through a preset set of second linear transformation matrices to obtain a first text feature sequence includes: The target visual word sequence is linearly transformed by the three second linear transformation matrices in the second linear transformation matrix set, respectively, to obtain the text feature query sequence, text feature key sequence and text feature value sequence. Based on the text feature query sequence and the text feature key sequence, determine the second weight matrix; The first text feature sequence is determined based on the second weight matrix and the text feature value sequence.
10. The method according to any one of claims 1 to 9, characterized in that, The image difference detection method is applied to a pre-trained difference detection model, and the method further includes: training the difference detection model through the following steps: Acquire training data, which includes multiple sample difference images and label text corresponding to each sample difference image. The sample difference images include multiple sample difference sub-images used to characterize the differences between two sample images. Using the feature embedding layer of the difference detection model, the sample difference image is processed by feature embedding through multiple sample difference sub-images to obtain a sample image feature sequence; Using the feature mapping layer of the difference detection model, text mapping is performed on the feature sequence of the sample image to obtain the sample text feature sequence; the sample text feature sequence is the feature sequence corresponding to the multiple sample difference sub-images; Using the text decoding layer of the difference detection model, based on the sample text feature sequence, determine the sample text used to characterize the difference between the first sample image and the second sample image; The first loss result is determined based on the sample image feature sequence and the sample text feature sequence; A second loss result is determined based on the labeled text and the sample text; The model parameters in the difference detection model are updated using the first loss result and the second loss result to obtain the trained difference detection model.
11. An image difference detection device, characterized in that, The device includes: The image acquisition module is used to acquire the first and second images to be detected. An image difference module is used to perform difference processing on the first image and the second image to obtain a difference image; the difference image includes multiple difference sub-images used to characterize the differences between the first image and the second image; The feature embedding module is used to perform feature embedding processing on the difference image based on multiple difference sub-images to obtain an image feature sequence; The text mapping module is used to perform text mapping on the image feature sequence based on a preset image-text dictionary to obtain a text feature sequence; the text feature sequence is a feature sequence corresponding to the plurality of differential sub-images; The text determination module is used to determine target text that characterizes the differences between the first image and the second image based on the text feature sequence.
12. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the image difference detection method according to any one of claims 1 to 10.
13. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the image difference detection method according to any one of claims 1 to 10 is implemented.
14. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the image difference detection method according to any one of claims 1 to 10 is implemented.