Screen shooting image processing method and device, storage medium and electronic equipment

By performing image matching and perspective change operations on the screen-capturing image, accurately positioning the target area and extracting watermark information, the problem of noise interference in the screen-capturing image is solved, and efficient watermark information acquisition and copyright protection are achieved.

CN120031702APending Publication Date: 2025-05-23CHINA CONSTRUCTION BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510119698.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Screen capture images are often accompanied by noise such as angle shift, light changes and molar patterns, making it difficult for traditional watermark extraction algorithms to accurately parse hidden watermark information, especially when there are a large number of unrelated areas in the image.

Method used

By acquiring the original image and screen capture image, performing image matching operations and perspective change operations, determining the target area image, and performing a watermark extraction operation to obtain screen capture watermark information.

Benefits of technology

Accurately position and adjust the areas related to the original image watermark information in the screen-capturing image to ensure the accurate extraction and comparison of watermark information, and improve the application efficiency of watermark technology in complex image acquisition environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031702A_ABST
    Figure CN120031702A_ABST
Patent Text Reader

Abstract

The invention discloses a screen shooting image processing method and device, a storage medium and electronic equipment. The method comprises the steps that an original image and a screen shooting image are acquired, the original image comprises original watermark information, and the screen shooting image represents an image obtained by acquiring the original image through image acquisition equipment; performing image matching operation and perspective change operation on the original image and the screen shooting image to obtain a target area image; and comparing the screen shooting watermark information with the original watermark information to obtain a target comparison result. The technical problem that the watermark information of the screen shooting image cannot be accurately obtained is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and more specifically, to a method and device for processing screen images, a storage medium, and an electronic device. Background Art

[0002] With the widespread dissemination of digital images, copyright holders' watermark information is often embedded in the generated images to protect copyright and verify image authenticity. However, captured images are often accompanied by noise such as angle deviation, lighting changes, and moiré patterns. This makes it difficult for traditional watermark extraction algorithms to accurately parse the hidden watermark information. This is especially true when there are large amounts of irrelevant areas in the image, where the watermark signal becomes blurred, seriously affecting the accurate acquisition and interpretation of the watermark information. In summary, the relevant technologies have the technical problem of being unable to accurately obtain the watermark information of captured images. Summary of the Invention

[0003] The embodiments of the present application provide a method and device for processing a captured screen image, a storage medium, and an electronic device, so as to at least solve the technical problem of being unable to accurately obtain the watermark information of the captured screen image.

[0004] According to one aspect of an embodiment of the present application, a method for processing a screen shot image is provided, comprising: acquiring an original image and a screen shot image, wherein the original image includes original watermark information, and the screen shot image represents an image obtained by acquiring the original image through an image acquisition device; performing an image matching operation and a perspective change operation on the original image and the screen shot image to obtain a target area image, wherein the image matching operation is used to determine a first coordinate point set and a second coordinate point set, the first coordinate point set corresponding to the original image, the second coordinate point set corresponding to the screen shot image, the first coordinate point set and the second coordinate point set being used to generate a first homography matrix, the first homography matrix being used to determine an initial area image in the screen shot image corresponding to the original image, and the perspective change operation being used to adjust the image range and tilt angle of the initial area image to obtain the target area image; performing a watermark extraction operation on the target area image to obtain screen shot watermark information; and comparing the screen shot watermark information with the original watermark information to obtain a target comparison result.

[0005] According to another aspect of an embodiment of the present application, a device for processing a screen shot image is further provided, including: an acquisition module, configured to acquire an original image and a screen shot image, wherein the original image includes original watermark information, and the screen shot image represents an image obtained by acquiring the original image through an image acquisition device; an execution module, configured to perform an image matching operation and a perspective change operation on the original image and the screen shot image to obtain a target area image, wherein the image matching operation is configured to determine a first coordinate point set and a second coordinate point set, the first coordinate point set corresponding to the original image, the second coordinate point set corresponding to the screen shot image, the first coordinate point set and the second coordinate point set used to generate a first homography matrix, the first homography matrix used to determine an initial area image in the screen shot image corresponding to the original image, and the perspective change operation used to adjust the image range and tilt angle of the initial area image to obtain the target area image; an extraction module, configured to perform a watermark extraction operation on the target area image to obtain screen shot watermark information; and a comparison module, configured to compare the screen shot watermark information with the original watermark information to obtain a target comparison result.

[0006] Optionally, the device is used to perform a watermark extraction operation on the target area image to obtain screen capture watermark information in the following manner, including: converting the target area image into a grayscale image, and performing a fast Fourier transform operation on the grayscale image to obtain a spectrum diagram of the grayscale image; when there is a component higher than a target threshold in the spectrum diagram, performing a moiré removal operation on the target area image to obtain the denoised area image; performing the watermark extraction operation on the denoised area image to obtain the screen capture watermark information.

[0007] Optionally, the device is used to perform a watermark extraction operation on the target area image to obtain screen capture watermark information in the following manner: perform a scale transformation and normalization operation on the target area image to obtain an update area image; extract high-dimensional features of the update area image, wherein the high-dimensional features are used to indicate the screen capture watermark information in the update area image; perform a decoding operation on the high-dimensional features to obtain the screen capture watermark information, the decoding operation corresponds to the encoding operation used in the generation process of the original image, and the encoding operation is used to embed the original watermark information in the original image.

[0008] Optionally, the device is used to compare the screen capture watermark information and the original watermark information to obtain a target comparison result in the following manner: performing a codeword error correction operation on the screen capture watermark information to obtain a decoded string, and obtaining a watermark string corresponding to the original watermark information; comparing the decoded string and the watermark string to obtain the target comparison result.

[0009] Optionally, the device is used to perform image matching operations and perspective change operations on the original image and the captured image in sequence to obtain a target area image in the following manner: performing the image matching operation on the original image and the captured image to obtain the first coordinate point set and the second coordinate point set, and determining the first homography matrix based on the first coordinate point set and the second coordinate point set; using the first homography matrix to perform a perspective forward transformation operation on the original image to obtain a third coordinate point set, and determining a second homography matrix based on the third coordinate point set and the first coordinate point set, wherein the third coordinate point set is used to indicate the image range of the initial area image, the second homography matrix is ​​used to adjust the tilt angle of the initial area image, and the perspective transformation operation includes the perspective forward transformation operation; using the second homography matrix to perform a perspective inverse transformation operation on the initial area image to obtain the target area image, wherein the perspective transformation operation includes the perspective inverse forward transformation operation.

[0010] Optionally, the device is also used to: obtain a watermark string and set the encoding length and version data of the watermark string; perform an encoding operation on the watermark string to obtain a payload; pad the length of the payload to obtain padding data, and generate a check code for the padding data; splice the payload, the check code and the version data to obtain the original watermark information.

[0011] Optionally, the device is also used to: obtain an initial watermark-free image, and perform scale transformation and normalization operations on the initial watermark-free image to obtain a target watermark-free image; obtain a watermark-free feature map of the target watermark-free image, and input the target watermark-free image and the original watermark information into a target encoder to obtain a watermark-containing feature map; perform residual calculation on the watermark-containing feature map and the watermark-free feature map to obtain a residual image; and add the scale-transformed residual image and the target watermark-free image to obtain the original image.

[0012] Optionally, the device is further used to: obtain the original image and the screen image; perform the image matching operation on the original image and the screen image to obtain the first coordinate point set and the second coordinate point set, and determine the first homography matrix based on the first coordinate point set and the second coordinate point set; use the first homography matrix to perform a perspective forward transformation operation on the original image to obtain a third coordinate point set, and determine a second homography matrix based on the third coordinate point set and the first coordinate point set, wherein the third coordinate point set is used to indicate the image range of the initial area image, the second homography matrix is ​​used to adjust the tilt angle of the initial area image, and the perspective transformation operation includes the perspective forward transformation operation; use the second homography The matrix performs an inverse perspective transformation operation on the initial area image to obtain a target area image, wherein the perspective transformation operation includes the inverse perspective transformation operation; converts the target area image into a grayscale image, and performs a fast Fourier transform operation on the target grayscale image to obtain a spectrum diagram of the grayscale image; when there is a component higher than a target threshold in the spectrum diagram, performs a moiré removal operation on the target area image to obtain a denoised area image; extracts high-dimensional features of the denoised area image, and determines the screen watermark information based on the high-dimensional features; performs a codeword error correction operation on the screen watermark information to obtain a decoded string, and obtains the watermark string corresponding to the original watermark information; compares the decoded string and the watermark string to obtain the target comparison result.

[0013] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned method for processing the captured screen image when running.

[0014] According to another aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described method for processing a captured screen image.

[0015] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-mentioned method for processing the captured screen image through the computer program.

[0016] In this embodiment, an original image and a captured screen image are first acquired. The original image is pre-embedded with the original watermark information, while the captured screen image is acquired by an image acquisition device such as a mobile phone or camera. Image matching and perspective change operations are then performed on the original and captured screen images. This process accurately locates and adjusts the area in the captured screen image that is relevant to the watermark information in the original image, resulting in an image of the target area, ensuring accurate extraction of the watermark information.

[0017] Specifically, the image matching operation determines the corresponding sets of key point coordinates in the original image and the captured image, namely the first set of coordinate points and the second set of coordinate points, by analyzing the similarities between the images. These coordinate points are used to generate a first homography matrix. The first homography matrix reflects the geometric transformation relationship between the original image and the captured image and is the basis for determining the area in the captured image that corresponds to the watermark information. The first homography matrix can accurately determine the position of the initial area image in the captured image, even if there are a large number of irrelevant areas or the viewing angle is skewed. The perspective change operation further adjusts the determined initial area image. By changing the image range and correcting the tilt angle, the target area image is brought closer to the state of the original image, reducing interference from irrelevant areas and ensuring the integrity and clarity of the watermark information. Subsequently, a watermark extraction operation is performed on the target area image to obtain the captured watermark information. By comparing it with the original watermark information, the target comparison result can be obtained, further verifying the consistency of the watermark information.

[0018] In summary, by precisely locating the target area, eliminating irrelevant noise, and accurately extracting watermark information, the embodiments of the present application solve the technical problem of being unable to accurately obtain watermark information in screen capture scenarios, improve the application efficiency of watermark technology in complex image acquisition environments, and provide reliable technical support for image copyright protection and content authentication. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0020] Figure 1 is a schematic diagram of an application environment of an optional method for processing a screen shot image according to an embodiment of the present application;

[0021] Figure 2 is a flowchart of an optional method for processing a screen shot image according to an embodiment of the present application;

[0022] Figure 3 is a schematic diagram of an optional method for processing a screen shot image according to an embodiment of the present application;

[0023] Figure 4 is a schematic diagram of another optional method for processing a screen shot image according to an embodiment of the present application;

[0024] Figure 5 is a schematic structural diagram of an optional device for processing screen images according to an embodiment of the present application;

[0025] Figure 6 is a schematic structural diagram of an optional screen image processing product according to an embodiment of the present application;

[0026] Figure 7 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0029] The present application will be described below with reference to the following embodiments:

[0030] According to one aspect of an embodiment of the present application, a method for processing a screen shot image is provided. Optionally, in this embodiment, the method for processing a screen shot image can be applied to Figure 1 In the hardware environment composed of the server 101 and the terminal device 103 shown in FIG. Figure 1As shown, the server 101 is connected to the terminal device 103 via a network and can be used to provide services for the terminal device or the application 107 installed on the terminal device. The application can be a video application, instant messaging application, browser application, educational application, game application, etc. A database 105 may be set up on the server or independently of the server to provide data storage services for the server 101, for example, a game data storage server. The above-mentioned network may include but is not limited to: a wired network, a wireless network, wherein the wired network includes: a local area network, a metropolitan area network and a wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that realize wireless communication. The terminal device 103 may be a terminal configured with an application, and may include but is not limited to at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop computer, a tablet computer, a PDA, a MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a mixed reality (MR) terminal and other computer devices. The above-mentioned server may be a single server, a server cluster consisting of multiple servers, or a cloud server.

[0031] Combine Figure 1 As shown, the above-mentioned method for processing the captured screen image can be executed by an electronic device, which can be a terminal device or a server. The above-mentioned method for processing the captured screen image can be implemented separately by the terminal device or the server, or jointly by the terminal device and the server.

[0032] For example, the terminal device 103 is a camera. The server 101 obtains the screen shot image uploaded by the terminal, and then performs image matching and perspective transformation operations with the original image stored in itself to obtain the target area image, and performs subsequent watermark extraction and information comparison on the target area image.

[0033] For another example, the terminal device 103 is a smart phone, notebook, computer, etc., and the original image can be generated by the terminal device 103. Then, the original image is published to the target platform. After a period of time, a user uses a mobile phone camera to shoot the original image without the permission of the image owner, obtains a screenshot image, and shares the screenshot image on other platforms. At this time, if the owner of the original image finds the screenshot image, the screenshot image is downloaded through the terminal device 103, and then image matching and perspective transformation operations are performed with the original image to obtain a target area image. Subsequent watermark extraction and information comparison are performed on the target area image to generate a target comparison result. If the target comparison result indicates that the extracted screenshot watermark information is the same as the original watermark information, it means that the screenshot image is indeed from the original image protected by the watermark, and is shared without the permission of the image owner, constituting copyright infringement. Then, the original image owner can perform subsequent processing based on the target comparison result, including but not limited to sending infringement warnings, requiring relevant platforms to remove the screenshot image, etc.

[0034] The above is only an example and is not specifically limited in this embodiment.

[0035] Alternatively, as an optional implementation, Figure 2 As shown, the above-mentioned method for processing the screen image includes:

[0036] S202, acquiring an original image and a screen image, wherein the original image includes original watermark information, and the screen image represents an image obtained by capturing the original image through an image capture device;

[0037] Optionally, in the embodiments of the present application, the original watermark information refers to a specific identifier hidden in the image for copyright protection and content verification, and can be any form of digital watermark, such as text, pattern, or signal sequence, used to verify the true source and copyright ownership of the image. A screen capture image refers to a screen display image captured by a user or a third party using an image acquisition device in actual usage scenarios. Screen capture images are easily affected by various factors such as lighting, angle, and device performance, resulting in changes in image quality and blurred watermark information.

[0038] Optionally, in an embodiment of the present application, the above-mentioned image acquisition device can obtain the screen image in the following ways, including but not limited to using the image acquisition device to acquire the image in real time and uploading it to the image storage device, or, using the image acquisition device to acquire a video, obtaining the initial image by intercepting the video frame, and uploading the initial image to the image storage device, or, using the image acquisition device to acquire a video, uploading the video to the image storage device, and having the image storage device intercept the video frame to obtain the initial image, wherein the image acquisition device may include but not limited to a smart phone, a computer, a notebook, a smart wearable device, a camera, etc., and the image storage device may include but not limited to a server, a computer, a laptop, a smart phone, a tablet computer, a digital camera, a scanner, a printer, etc.

[0039] Specifically, before publishing the original image, Company A embedded the original watermark information in the original image, and then published the original image containing the original watermark information on the target platform. After a period of time, user B used a computer to log in to the target platform. While browsing the image resources on the target platform, he used a mobile phone (that is, the above-mentioned image acquisition device) to take a picture of the above-mentioned original image displayed on the computer screen, generated the screenshot image on the mobile phone, and published the screenshot image on other social media platforms. At this time, an employee of Company A found the above-mentioned screenshot image on the other social media platform. Therefore, Company A can directly download the above-mentioned screenshot image from the other social media platform, and directly query and obtain the above-mentioned original image from the company's image resource database.

[0040] It should be noted that the types of target platforms and upload mechanisms are also diverse. They can be cloud storage services, content management systems, copyright verification platforms, or any digital platform that requires verification of the image source. Upload methods can include but are not limited to automatic uploads, user manual confirmation uploads, etc. The specific upload process and permission control mechanism are determined by the platform's business rules and are not limited in this application.

[0041] It should also be noted that the specific content, format and size of the original image and the captured image can vary in many ways, and the embedding method and type of the original watermark information can also be adjusted according to actual needs. For example, it can be a perceptible or imperceptible watermark, and it can be a watermark technology based on the pixel domain, transform domain or deep learning model. This application does not limit this.

[0042] In addition, there may also be differences in the type of image acquisition equipment and the conditions during the acquisition process (such as light intensity, shooting angle, etc.). These factors may affect the final quality of the captured image, but the image matching and perspective change operations proposed in this application are intended to overcome these influences and ensure the accurate acquisition of watermark information.

[0043] S204: Perform an image matching operation and a perspective change operation on the original image and the captured screen image to obtain a target area image, wherein the image matching operation is used to determine a first set of coordinate points and a second set of coordinate points, the first set of coordinate points corresponding to the original image, and the second set of coordinate points corresponding to the captured screen image, the first set of coordinate points and the second set of coordinate points being used to generate a first homography matrix, the first homography matrix being used to determine an initial area image in the captured screen image that corresponds to the original image, and the perspective change operation being used to adjust the image range and tilt angle of the initial area image to obtain the target area image;

[0044] Optionally, in an embodiment of the present application, the above-mentioned image matching operation refers to finding corresponding points in the two images by comparing the similarity of feature points between the two images, including but not limited to feature point detection, descriptor matching and robustness estimation algorithms, such as feature point matching technologies such as SIFT, SURF or ORB, and the RANSAC algorithm for estimating the reliability of the homography matrix. The above-mentioned first homography matrix refers to a matrix that describes the geometric transformation between corresponding points in the original image and the captured image, which can reflect the translation, rotation, scaling and perspective effects between the two images, and is used to determine the watermark information area in the captured image.

[0045] It should be noted that the specific implementation of image matching and perspective change operations can adopt a variety of technical solutions, including but not limited to feature matching based on deep learning, traditional image processing algorithms or hybrid methods, and this application does not limit this.

[0046] In addition, when performing image matching operations, the key feature points selected and their matching accuracy, as well as the method of generating the homography matrix, may vary depending on the specific application scenario (such as the resolution of the image, the complexity of the content, or the characteristics of the acquisition device).

[0047] Similarly, the strategy for adjusting the image range and tilt angle in the perspective change operation can also be flexibly adjusted according to the embedding position of the watermark information and the degree of image deformation.

[0048] S206, performing a watermark extraction operation on the target area image to obtain screen watermark information;

[0049] Optionally, in an embodiment of the present application, the above-mentioned watermark extraction operation refers to extracting preset watermark information from an image using a watermark decoding algorithm or a neural network model, including but not limited to a deep learning-based watermark detector, traditional frequency domain analysis or spatial domain analysis technology, and a specific digital watermark decoder.

[0050] It should be noted that the specific implementation of the watermark extraction operation can be very diverse. For example, a decoder based on a convolutional neural network can be used to identify and extract watermark information by training the network; frequency domain analysis techniques, such as discrete Fourier transform or discrete wavelet transform, can be used to identify watermark signals in specific frequency bands in the image; spatial domain analysis can also be used to directly search for specific patterns or textures in the image to locate watermark information.

[0051] In addition, the architecture design of the decoder or detector, the selection of training data, and the parameter setting of the extraction algorithm may be adjusted according to different watermark types and image characteristics, which is not limited in this application.

[0052] S208, comparing the captured watermark information with the original watermark information to obtain a target comparison result.

[0053] Optionally, in an embodiment of the present application, the above-mentioned comparison of the captured screen watermark information and the original watermark information refers to checking the similarity or consistency of the two watermark information through a comparison algorithm, including but not limited to Hamming distance comparison, bit string matching or watermark pattern recognition, to determine whether the captured screen image originates from the original image embedded with a specific watermark identifier.

[0054] It should be noted that the specific method for generating the target comparison result can be very flexible. For example, a threshold can be set according to the degree of accurate matching of the watermark information. When the similarity between the two exceeds the threshold, it is confirmed that the match is successful; or a more complex probability model, such as a Bayesian classifier, can be used to evaluate the probability of watermark information matching.

[0055] In one exemplary embodiment, using the copyright protection system as an example, assume a user shares a photographic work containing a copyright watermark on a social media platform. This image has been embedded with specific watermark information by the image generation module before being uploaded. However, other users may capture and re-share this work using their mobile phone screens, resulting in the original copyright information being blurred or tampered with.

[0056] To verify and protect the copyright of this image, after the original image is released, the content creator discovers the screenshot image on the Internet. At this time, an independent verification system can be designed to verify the copyright information using the screenshot image processing method proposed in this application. Specifically:

[0057] The system first obtains the original image and the captured image. Next, it performs the image matching operation on the original image and the captured image. By analyzing the image feature points, the system determines a first coordinate point set and a second coordinate point set. These two sets correspond to the key points in the original image and the captured image, respectively, and are used to calculate the first homography matrix that describes the geometric transformation between the two images.

[0058] The system then uses the first homography to perform a forward perspective transformation on the original image, generating a third set of coordinate points indicating the possible locations of the watermarked area in the original image. Based on the third set of coordinate points and the first set of coordinate points, a second homography is further determined. This matrix is ​​used to adjust the tilt angle of the initial area image obtained from the first homography. By performing the forward perspective transformation and the inverse perspective transformation, the system obtains a geometrically corrected image of the target area, ensuring accurate identification of the watermarked area.

[0059] To remove any moiré noise that may exist in the target region image, the system converts the target region image into a grayscale image and performs a Fast Fourier Transform (FFT) on the grayscale image to obtain a spectrogram of the image. If the spectrogram contains components above the target threshold, indicating the presence of moiré, the system performs a moiré removal operation on the target region image, utilizing a specially designed denoising network such as ESDNet to obtain a denoised region image, thereby ensuring the clarity and readability of the watermark information.

[0060] Next, the system performs high-dimensional feature extraction on the de-noised area image and, by analyzing these features, determines the screen capture watermark embedded in the image. To further improve the reliability of the watermark, the system performs codeword error correction on the screen capture watermark to obtain a decoded string. The system then retrieves the watermark string corresponding to the original watermark from storage.

[0061] Finally, when the decoded string and the watermark string are identical, that is, when the watermark information extracted from the screenshot image is consistent with the watermark information embedded in the original image, the system will automatically or after confirmation by the user determine that the producer of the screenshot image privately photographed the original image without obtaining authorization from the content creator, causing damage to its own rights and interests. Therefore, the relevant platform can be requested to directly remove the screenshot image.

[0062] In the embodiments of the present application, an original image and a captured screen image are first obtained. The original image is pre-embedded with the original watermark information, while the captured screen image is obtained by capturing the original image using an image acquisition device such as a mobile phone or camera. Image matching and perspective change operations are performed on the original image and the captured screen image. This process accurately locates and adjusts the area in the captured screen image that is related to the watermark information of the original image, obtaining an image of the target area and ensuring accurate extraction of the watermark information.

[0063] Specifically, the image matching operation analyzes the similarities between the images to determine the corresponding coordinate sets of key points in the original image and the captured image, namely the first coordinate point set and the second coordinate point set. These coordinate points are used to generate a first homography matrix. The first homography matrix reflects the geometric transformation relationship between the original image and the captured image and is the basis for determining the area in the captured image corresponding to the watermark information. The first homography matrix accurately determines the position of the initial region image in the captured image, even if the image contains a large amount of irrelevant areas or the viewing angle is skewed. The perspective change operation further adjusts the determined initial region image by changing the image range and correcting the tilt angle, bringing the target region image closer to the original image, reducing interference from irrelevant areas, and ensuring the integrity and clarity of the watermark information. Subsequently, a watermark extraction operation is performed on the target region image to obtain the captured watermark information. By comparing it with the original watermark information, the target comparison result is obtained, further verifying the correctness and consistency of the watermark information. This not only verifies the source and authenticity of the image but also ensures the rights of the content creator.

[0064] In summary, by precisely locating the target area, eliminating irrelevant noise, and accurately extracting watermark information, the embodiments of the present application solve the technical problem of being unable to accurately obtain watermark information in screen capture scenarios, improve the application efficiency of watermark technology in complex image acquisition environments, and provide reliable technical support for image copyright protection and content authentication.

[0065] As an optional solution, the above-mentioned watermark extraction operation is performed on the above-mentioned target area image to obtain the screen watermark information, including: converting the above-mentioned target area image into a grayscale image, and performing a fast Fourier transform operation on the above-mentioned grayscale image to obtain a spectrum diagram of the above-mentioned grayscale image; when there is a component higher than the target threshold in the above-mentioned spectrum diagram, performing a moiré elimination operation on the above-mentioned target area image to obtain the above-mentioned denoised area image; performing the above-mentioned watermark extraction operation on the above-mentioned denoised area image to obtain the screen watermark information.

[0066] Optionally, in an embodiment of the present application, the above-mentioned target area image refers to the image area containing watermark information that is accurately located and separated from the captured image through image matching and perspective transformation operations, including but not limited to the image part after cropping, rotation or orthogonalization.

[0067] Optionally, in an embodiment of the present application, the grayscale image refers to converting the color information of the target area image into grayscale to form a single-channel image format to simplify image processing and improve algorithm efficiency.

[0068] Optionally, in this embodiment of the present application, the aforementioned Fast Fourier Transform operation refers to analyzing the frequency components of a grayscale image using a Fast Fourier Transform algorithm, including but not limited to frequency domain conversion of the image signal and generation of a spectrogram. A spectrogram is a visual representation of the frequency components of a grayscale image after the Fast Fourier Transform, and is used to detect specific noise patterns in the image.

[0069] Optionally, in an embodiment of the present application, the above-mentioned target threshold refers to a preset value used to determine whether there is noise such as moiré in the spectrum diagram, and can be any value set according to the image resolution, device characteristics and characteristics of the watermark information.

[0070] Optionally, in an embodiment of the present application, the above-mentioned moiré removal operation refers to the use of a specific image processing algorithm or neural network model to remove or weaken the effect of moiré noise in the target area image, including but not limited to frequency domain filtering, spatial domain smoothing or deep learning-based denoising technology.

[0071] Optionally, in an embodiment of the present application, the above-mentioned denoised area image refers to a target area image with no or significantly reduced moiré noise obtained after the moiré removal operation, providing higher quality image data for subsequent watermark information extraction.

[0072] Optionally, in an embodiment of the present application, the above-mentioned watermark extraction operation refers to the process of recovering preset watermark information from the processed image, including but not limited to feature matching, bit string detection or watermark signal decoding.

[0073] It should be noted that the specific algorithm implementation of the fast Fourier transform operation on the grayscale image, the setting standard of the target threshold, the detailed technical process of the moiré removal operation, the accuracy requirements of the watermark information extraction, etc. may vary depending on the specific application scenario, the complexity of the image and the type of watermark information, and this application does not limit this.

[0074] For example, the target area image is converted to a grayscale image and a fast Fourier transform is performed to obtain a spectrogram. If the high-frequency components in the spectrogram exceed a preset target threshold, it indicates the presence of moiré noise in the image. Subsequently, a moiré removal operation is performed on the target area image to obtain a denoised area image. A watermark extraction operation is then performed on the denoised area image to obtain the captured watermark information.

[0075] In one exemplary embodiment, taking the workflow of a copyright verification system as an example, after receiving a captured screen image uploaded by a user, the system first performs image matching and perspective transformation to obtain a target area image. Subsequently, the target area image is converted to a grayscale image, and the spectrum is analyzed using a fast Fourier transform to identify the presence of moiré patterns. If moiré patterns are present and their components exceed a target threshold, the system performs a moiré removal operation on the target area image to generate a denoised area image, providing clearer image data for subsequent watermark extraction. Finally, the system accurately extracts the watermark information from the denoised area image and compares it with the original watermark information, ensuring accurate verification of the image copyright.

[0076] Through the embodiments of the present application, a fast Fourier transform operation of a grayscale image is used to determine whether moiré patterns exist in the target area image. If so, a conditional moiré removal operation is performed on the target area image to obtain the above-mentioned denoised area image, thereby achieving the technical effect of accurately extracting watermark information from the captured image, and achieving the purpose of effectively verifying image copyright and preventing infringement even in complex capture scenes and noise interference.

[0077] As an optional solution, the above-mentioned watermark extraction operation is performed on the above-mentioned target area image to obtain the screen watermark information, including: performing scale transformation and normalization operations on the above-mentioned target area image to obtain the update area image; extracting high-dimensional features of the above-mentioned update area image, wherein the above-mentioned high-dimensional features are used to indicate the above-mentioned screen watermark information in the above-mentioned update area image; performing decoding operations on the above-mentioned high-dimensional features to obtain the above-mentioned screen watermark information.

[0078] Optionally, in an embodiment of the present application, the above-mentioned target area image refers to an image area containing watermark information that has been processed in advance, that is, accurately extracted from the screen capture image through matching algorithm and homography transformation, including but not limited to part or all of the original watermark image.

[0079] Optionally, in an embodiment of the present application, the above-mentioned scale transformation and normalization operations refer to adjusting the target area image to a size suitable for processing by the watermark extraction algorithm, and normalizing its pixel values ​​to a fixed range, such as 0 to 1, to ensure the consistency and accuracy of watermark information extraction, including but not limited to using bilinear interpolation for size adjustment and minimum and maximum scaling for normalization.

[0080] Optionally, in an embodiment of the present application, the above-mentioned updated region image refers to a target region image having a uniform size and pixel value range after scale transformation and normalization operations, so as to prepare for subsequent high-dimensional feature extraction.

[0081] Optionally, in an embodiment of the present application, the above-mentioned high-dimensional features refer to complex features for identifying watermark information obtained from the update area image through a deep learning model or other advanced feature extraction methods, including but not limited to feature maps of deep neural networks, frequency domain features or spatial domain features.

[0082] Optionally, in an embodiment of the present application, the above-mentioned decoding operation refers to the process of reverse parsing the watermark information based on high-dimensional features, which can be based on a specific decoding algorithm or model, such as a deep learning decoder, to recover the watermark string or bit stream hidden in the image.

[0083] It should be noted that the specific methods of scale transformation and normalization operations, such as the setting of target size and normalization range, and the specific algorithm or model used in the decoding operation can be adjusted according to actual needs and application scenarios, and this application does not limit this.

[0084] Exemplarily, rescaling and normalization operations are performed on the target area image to obtain an updated area image, and then the high-dimensional features of the image are extracted through a deep neural network model. The high-dimensional features point to the watermark information in the target area image. Finally, the high-dimensional features are parsed using a decoding algorithm to obtain the screen watermark information. This process ensures the accuracy and robustness of the watermark information.

[0085] In one exemplary embodiment, using the image copyright verification process of a digital rights management platform as an example, after acquiring a screen capture image, an image matching algorithm is first used to determine the target area. Then, scaling and normalization operations are performed to obtain an updated region image suitable for watermark extraction. A deep learning model is then used to extract high-dimensional features from the updated region image, which indicate the presence of watermark information. Finally, a decoding operation is used to recover the watermark information from the high-dimensional features for copyright verification and tracking, ensuring that copyright information can be accurately identified and protected even in the face of challenging screen captures.

[0086] Through the embodiments of the present application, by adopting scaling and normalization operations, as well as high-dimensional feature extraction and decoding technology, the technical effect of accurately identifying and extracting watermark information from screen-shot images is achieved, thereby achieving the purpose of effectively ensuring image copyright security and rights protection even in complex screen-shooting scenarios.

[0087] As an optional solution, the above-mentioned comparison of the above-mentioned screen watermark information and the above-mentioned original watermark information to obtain the target comparison result includes: performing a codeword error correction operation on the above-mentioned screen watermark information to obtain a decoded string, and obtaining the watermark string corresponding to the above-mentioned original watermark information; comparing the above-mentioned decoded string with the above-mentioned watermark string to obtain the above-mentioned target comparison result.

[0088] Optionally, in an embodiment of the present application, the above-mentioned screen capture watermark information refers to watermark information extracted from the screen capture image, which may contain noise or distortion, including but not limited to a bit sequence or character string containing copyright, brand or author identification.

[0089] Optionally, in an embodiment of the present application, the above-mentioned original watermark information refers to the watermark information initially embedded in the original image to identify and protect the copyright of the image, including but not limited to the original bit sequence or character string.

[0090] Optionally, in an embodiment of the present application, the above-mentioned codeword error correction operation refers to the process of using the error correction mechanism in coding theory, such as BCH code, Hamming code or Turbo code, to detect and correct errors in the screen watermark information to ensure the accuracy of the watermark information.

[0091] Optionally, in an embodiment of the present application, the above-mentioned decoded character string refers to a character string that is correct or has its errors corrected, which is recovered from the screen capture watermark information after a codeword error correction operation.

[0092] Optionally, in an embodiment of the present application, the above-mentioned watermark character string refers to a character string obtained from the original watermark information and used for comparison, which is used to verify the authenticity of the screen capture watermark information.

[0093] Optionally, in the embodiment of the present application, the target comparison result refers to a result of determining whether the decoded character string and the watermark character string are consistent by comparing them.

[0094] It should be noted that the specific mechanism of codeword error correction and the comparison method of the decoded string and the watermark string may vary according to the specific application scenario and requirements, and this application does not limit this.

[0095] Exemplarily, the screen capture watermark information extracted from the screen capture image performs a codeword error correction operation to obtain an error-free or low-error decoding string, and then compares this string with the watermark string corresponding to the original watermark information. If the two are the same, it is determined that the screen capture image originates from the original image, and the original image is allowed to be uploaded to the target platform.

[0096] Through the embodiments of the present application, codeword error correction technology and precise comparison strategy are adopted to achieve accurate verification of the watermark information extracted from the screen-shot image and the watermark information of the original image, thereby achieving the purpose of preventing image infringement, protecting copyright and effectively managing digital assets.

[0097] As an optional scheme, the above-mentioned image matching operation and perspective change operation are performed on the above-mentioned original image and the above-mentioned screen image in sequence to obtain the target area image, including: performing the above-mentioned image matching operation on the above-mentioned original image and the above-mentioned screen image to obtain the above-mentioned first coordinate point set and the above-mentioned second coordinate point set, and determining the above-mentioned first homography matrix based on the above-mentioned first coordinate point set and the above-mentioned second coordinate point set; using the above-mentioned first homography matrix to perform a perspective forward transformation operation on the above-mentioned original image to obtain a third coordinate point set, and determining the second homography matrix based on the above-mentioned third coordinate point set and the above-mentioned first coordinate point set, wherein the above-mentioned third coordinate point set is used to indicate the image range of the above-mentioned initial area image, the above-mentioned second homography matrix is ​​used to adjust the inclination angle of the above-mentioned initial area image, and the above-mentioned perspective transformation operation includes the above-mentioned perspective forward transformation operation; using the above-mentioned second homography matrix to perform a perspective inverse transformation operation on the above-mentioned initial area image to obtain the above-mentioned target area image, wherein the above-mentioned perspective transformation operation includes the above-mentioned perspective inverse forward transformation operation.

[0098] Optionally, in an embodiment of the present application, the above-mentioned perspective transformation operation refers to applying a perspective transformation on the image to correct the viewing angle and shape of the image, including but not limited to forward perspective transformation and inverse perspective transformation, in order to accurately extract the target area image.

[0099] It should be noted that the accuracy of the image matching operation, the calculation method of the homography matrix, the determination strategy of the third coordinate point set, and the adjustment effect of the second homography matrix may vary due to differences in specific application scenarios, and this application does not limit this.

[0100] For example, the system uses image matching to determine the correspondence between key points in the original image and the captured image. It then calculates a first homography matrix, which is used to correct the viewing angle of the captured image. This results in a third set of coordinate points, which in turn determines the initial range of the target area. The second homography matrix is ​​then used to adjust the tilt angle of this initial area to ensure accurate image extraction of the target area.

[0101] In one exemplary embodiment, using a digital copyright verification system as an example, when a user submits a captured screen image, the system first performs image matching on the original image and the captured screen image to obtain key point matching information. Based on this information, the first homography matrix is ​​calculated using the following formula, and a forward perspective transformation is applied to determine the range and position of the initial region image in the captured screen image, thereby obtaining a third set of coordinate points:

[0102] The coordinate set of the matching points in the screen image is the second coordinate point set setA={(x i ,y i )|i∈{1,…,N}}, the coordinate set of the matching points in the original image is the first coordinate point set setB={(xi ′,y i ′)|i∈{1,…,N}}, where N is the number of matching points;

[0103] Furthermore, the homography matrix H of the original image matching point setB and the screen image matching point setA is calculated to reflect the transformation matrix from the screen image to the original image. i ,y i , 1) T =H(x i ′,y i ′,1) T Perform the above-mentioned perspective transformation.

[0104] Next, the first homography matrix H is used to determine the coordinate range of the four boundary points of the original image, which are (0,0), (w-1,0), (w-1,d-1), and (0,d-1), respectively. w and d represent the length and width, respectively. The coordinate points in the third coordinate point set are multiplied by the first homography matrix H to obtain the corresponding coordinate range in the captured image after the inverse perspective transformation, which is the target area image mentioned above.

[0105] Specifically, the inverse transmission transform of the second homography matrix can be understood as the inverse operation of the forward transmission transform:

[0106] Use the first homography matrix H calculated in the previous step to perform perspective transformation on the four boundary points of the original image, and obtain ((x, y, 1)) through the coordinates (x, y) of these points. T ), and use ((x, y, 1) T ) is multiplied by the homography matrix H to obtain the coordinates after perspective transformation ((x', y')). Based on these new coordinates, the image range of the target area in C1 is determined to obtain the corresponding target area image.

[0107] Through the embodiments of the present application, image matching, homography matrix determination and multi-step perspective transformation technology are adopted to achieve the technical effect of accurately locating and extracting the target area image from the captured image, thereby improving the accuracy and robustness of watermark information extraction and effectively supporting image copyright protection and digital asset tracking.

[0108] As an optional solution, the above method also includes: obtaining a watermark string and setting the encoding length and version data of the above watermark string; performing an encoding operation on the above watermark string to obtain a payload; padding the length of the above payload to obtain padding data, and generating a check code for the above padding data; splicing the above payload, the above check code and the above version data to obtain the above original watermark information.

[0109] Optionally, in an embodiment of the present application, the watermark string refers to copyright information, a verification code, or any text used to identify the source of an image, including but not limited to "Copyright", "Generator ID", or "Name of Work".

[0110] It should be noted that the encoding length can be adjusted according to the accuracy and robustness requirements of watermark embedding and extraction, the version data can be used to distinguish different versions of watermark information or algorithms, the filling rule of the payload can be fixed zero filling, random bit filling or other predefined patterns, the check code generation method can be BCH, Hamming code or CRC, etc., and the splicing method can be direct concatenation or combination according to a specific format, which is not limited in this application.

[0111] Exemplarily, the system first obtains a watermark string, such as copyright information, then sets the encoding length and version data of the string, performs an encoding operation to convert the string into a binary payload, then pads the payload with length and generates a check code, and finally concatenates the payload, check code, and version data together to form the complete original watermark information.

[0112] In one exemplary embodiment, using a digital image copyright verification system as an example, when an artist uploads a digital image of a work, the system automatically generates a unique watermark string, such as the artist's name and work ID. The system then sets this string to a 40-bit encoding length and appends version data to distinguish different versions of the watermark. The payload is then generated through encoding. To ensure information integrity, the system appends padding data after the payload and generates a checksum based on BCH encoding. Finally, the system concatenates the payload, checksum, and version data into the original watermark information, preparing it for image embedding to ensure accurate copyright verification during subsequent image transmission and use.

[0113] Through the embodiments of the present application, a solution of encoding, length padding, checksum generation and data splicing of the watermark character string is adopted to achieve effective embedding and extraction of watermark information in the image, thereby achieving the precise control purpose of protecting the copyright of digital content and preventing illegal copying and use.

[0114] As an optional solution, the above method also includes: obtaining an initial watermark-free image, and performing scale transformation and normalization operations on the above initial watermark-free image to obtain a target watermark-free image; obtaining a watermark-free feature map of the above target watermark-free image, and inputting the above target watermark-free image and the above original watermark information into a target encoder to obtain a watermark-containing feature map; performing residual calculation on the above watermark-containing feature map and the above watermark-free feature map to obtain a residual image; adding the above residual image after scale transformation and the above target watermark-free image to obtain the above original image.

[0115] Optionally, in the embodiment of the present application, the initial watermark-free image refers to the original image uploaded by the artist or any image that has not been watermarked, including but not limited to high-definition pictures, photographic works or digital artworks.

[0116] In one exemplary embodiment, using a digital image copyright verification system as an example, when an artist uploads a digital image of a work, the system first obtains the initial watermark-free image, for example, a 1920x1080 pixel high-definition image. The system then resizes this image to a fixed size, such as 256x256 pixels, and normalizes the pixel values ​​to ensure that it meets the model input requirements, resulting in a target watermark-free image. The system then extracts the feature map of the target watermark-free image and simultaneously inputs the target watermark-free image, along with the previously generated original watermark information, into a target encoder based on a U-Net architecture. The encoder embeds the watermark information into the image features, generating a watermarked feature map. Next, the system compares the watermarked feature map with the watermark-free feature map, calculating the difference between the two and obtaining a residual image. Finally, the system adds the residual image to the rescaled target watermark-free image to restore it to its original size, resulting in the original watermarked image—the image in which the watermark information has been successfully embedded.

[0117] Through the embodiments of the present application, the initial watermark-free image is preprocessed and the feature map is compared, combined with the scheme of embedding watermark information by the encoder, to achieve deep embedding of watermark information and visual integrity of the image, thereby achieving the dual technical goals of protecting copyright without affecting image quality and ensuring the robustness and accuracy of watermark information.

[0118] As an optional solution, the above method further includes:

[0119] Obtain the original image and the captured screen image;

[0120] Performing the image matching operation on the original image and the captured screen image to obtain the first coordinate point set and the second coordinate point set, and determining the first homography matrix based on the first coordinate point set and the second coordinate point set;

[0121] Performing a forward perspective transformation operation on the original image using the first homography matrix to obtain a third set of coordinate points, and determining a second homography matrix based on the third set of coordinate points and the first set of coordinate points, wherein the third set of coordinate points is used to indicate an image range of the initial region image, the second homography matrix is ​​used to adjust a tilt angle of the initial region image, and the perspective transformation operation includes the forward perspective transformation operation;

[0122] Performing an inverse perspective transformation operation on the initial region image using the second homography matrix to obtain a target region image, wherein the perspective transformation operation includes the inverse perspective transformation operation;

[0123] Converting the target area image into a grayscale image, and performing a fast Fourier transform operation on the target grayscale image to obtain a frequency spectrum of the grayscale image;

[0124] When there is a component higher than the target threshold in the above-mentioned spectrum graph, a moiré removal operation is performed on the above-mentioned target area image to obtain a denoised area image;

[0125] Extracting high-dimensional features of the denoised area image, and determining the screen watermark information based on the high-dimensional features;

[0126] Perform codeword error correction on the screen capture watermark information to obtain a decoded string, and obtain the watermark string corresponding to the original watermark information;

[0127] Compare the decoded string and the watermark string to obtain the target comparison result.

[0128] In an exemplary embodiment, before the image owner publishes the image, in order to ensure that the watermark in the generated image can clearly mark the copyright of the image owner, before publishing the image, it is detected whether the image containing its own watermark can be photographed and the captured image can be obtained to accurately extract its own watermark information, so as to indicate that it is the owner of the image and prevent the image from being maliciously abused after publication. The processing method of the captured image proposed in this application can be used in the detection process. In view of the problem that the captured image still has a large number of irrelevant areas and moiré noise, which make the watermark extraction algorithm unable to accurately extract the watermark from the captured image, an image watermark processing system for the captured image scenario is realized.

[0129] It should be noted that image watermarking technology refers to the addition of specific text or graphics to images to indicate the image's copyright, brand, or author's identity. Currently, AI-generated content is experiencing rapid growth and widespread dissemination, but AI-generated content carries the risk of being easily confused, misidentified, and misused. Watermarking technology can be used to tag, identify, and trace images, thereby protecting the copyright of content creators.

[0130] Image watermarking technology generally consists of two stages: watermark embedding and watermark extraction. The watermark embedding process refers to embedding the watermark information into the carrier image to obtain the carrier image containing the watermark and make it visually imperceptible. The watermark extraction process refers to extracting the watermark information from the carrier image containing the watermark. By comparing the embedded watermark information with the extracted watermark information, the image can be traced back.

[0131] Traditional watermarking methods are typically based on signal processing techniques, which hide information by changing certain image properties in the spatial or transform domain. However, this method can be quite robust. Currently, deep learning-based watermarking algorithms are common. These utilize deep learning models such as neural networks to process watermark information. These algorithms primarily consist of an encoder, a noise layer, and a decoder. The encoder embeds the watermark into the image, while the decoder extracts it from the original image. However, because noise is introduced during image use and transmission, a noise layer is designed during model training to improve the robustness of the watermarking algorithm.

[0132] In addition to leaking original images, screen captures can also, to a certain extent, lead to the leakage of the copyright of the generated images. Attackers can use electronic devices such as mobile phones and cameras to capture the target images and distribute them online, thereby infringing the copyright of the image generator. In this case, the recaptured images contain a large amount of irrelevant areas and moiré noise, making it difficult to accurately extract watermark information from the screen capture images.

[0133] To address the above issues, the embodiments of the present application implement an image watermark processing solution for screen capture scenarios. First, an image watermark embedding and extraction method is designed based on the encoder-decoder structure. This solution can cope with common noises during transmission, such as image compression and white noise. However, since the screen capture image contains a large number of irrelevant areas and moiré noise, it is impossible to accurately extract the watermark from the screen capture image. Therefore, to address the problem of irrelevant areas, a matching-based target area extraction method is designed to extract the original image area from the screen capture image containing a large number of irrelevant areas; to address the problem of moiré noise, a moiré noise processing method is designed to remove the influence of moiré noise in the capture image.

[0134] For example, Figure 3 is a schematic diagram of an optional method for processing a screen shot image according to an embodiment of the present application, such as Figure 3 As shown, the detection system may include an image generation module, a watermark information generation module, an image watermark embedding module, a target area extraction module, a moiré noise processing module, a watermark extraction module, and a watermark information correction module, and mainly includes the following steps:

[0135] S1, image generation module: The user inputs a prompt to the image generation model to generate a photo P1.

[0136] S2, the watermark information generation module, encodes the 5-bit character string s1 input by the user into 100 random bits b1, specifically including the following steps:

[0137] S2-1, the user inputs a random string s1 to be embedded in the image.

[0138] S2-2, set the total length of the data packet generated by BCH encoding, the total length of the version bit, and the version information.

[0139] S2-2, converts the ASCII code corresponding to each bit string in s1 into the corresponding binary code as the payload.

[0140] S2-3, pad the length of the payload to a multiple of 8, and use the BCH encoder to perform ECC encoding on the padded data to obtain ECC check bits.

[0141] S2-4, concatenate the payload, the ECC check code and the version information as the binary information b1 to be embedded in the image watermark.

[0142] S3, the watermark embedding module, uses the U-Net model to build a watermark embedding model to embed the watermark. The specific steps include:

[0143] S3-1, scale transform and normalize the input non-watermarked image P1.

[0144] S3-2 inputs the normalized image into the encoder built based on U-Net to obtain the watermarked feature map, and uses the watermarked feature map and the non-watermarked feature map to perform residual calculation to obtain the residual image R1.

[0145] S3-3, scale-transform the residual image R1, and add the scale-transformed residual image to the normalized original image to obtain an original image P2 with the same scale as the original image.

[0146] S4, target area extraction module: accurately extract the area that matches the original image area from the screen shooting area, Figure 4 is a schematic diagram of an optional method for processing a screen shot image according to an embodiment of the present application, such as Figure 4 As shown, the specific steps include:

[0147] S4-1, input the screen image C1 and the original watermark image P2 into the image matching network, and obtain the coordinate sets of matching points in the screen image and the original watermark image, respectively, as the screen image matching point coordinate set (the above-mentioned second coordinate point set) setA = {(x i ,y i )|i∈{1,…,N}}, and the coordinate set of the original watermark image matching point (the first coordinate point set mentioned above) setB={(x i ′,y i ′)|i∈{1,…,N}}, where N is the number of matching points.

[0148] S4-2, calculate the homography matrix H of the screen image matching point setA and the original watermark image matching point setB, which is used to reflect the conversion matrix from the screen image to the original watermark image. It can be calculated according to the formula (x i ,y i , 1) T =H(x i ′,y i ′,1) T Perform a perspective transformation.

[0149] S4-3, use the first homography matrix H to determine the coordinate range of the four boundary points of the original watermark image, which are (0,0), (W-1,0), (W-1, H-1), and (0, H-1). The coordinates within the range are multiplied by the homography matrix H to obtain the corresponding coordinate range in the captured image after perspective transformation, and then obtain the target area image area that needs to be cropped.

[0150] S4-4, due to the shooting angle during the shooting process, if the target area is directly cropped, the cropped image will have a certain tilt angle. In order to map the image to the vertical and horizontal directions, a perspective transformation needs to be performed on the cropped area.

[0151] Furthermore, the second homography H2 from the cropping area to the vertical and horizontal directions is calculated, where the coordinate range corresponding to the vertical and horizontal positions is determined by the four coordinate points (0,0), (W-1,0), (W-1, H-1), and (0, H-1). H2 can be obtained according to the formula in 4.2. The coordinates in the cropping area can be multiplied by H2 to transform the cropping area into the vertical and horizontal perspective inverse transformation to obtain the target area image C2.

[0152] S5, Moiré Noise Processing Module: This module removes moiré noise from captured images due to varying device frequencies. It first converts image C2 into a grayscale image and then uses a Fast Fourier Transform to calculate a spectrogram. If the spectrogram contains components above a threshold, moiré removal is considered necessary. If not, moiré removal is not necessary. The specific threshold is determined based on the usage scenario. Moiré removal is implemented using an ESDNet network based on an encoder-decoder architecture, yielding the original image C2'.

[0153] S6, watermark extraction module: Use the residual network to build a decoder module to extract watermark information from the original image. The specific process is as follows:

[0154] S6-1, scale transform and normalize the input original image C2.

[0155] S6-2, the normalized image is input into the spatial transformation network to extract the high-dimensional features of the original image.

[0156] S6-3, input the high-dimensional features into a high-dimensional feature extraction network composed of multiple convolutional layers and fully connected layers to extract the watermark information b2.

[0157] S7, watermark information correction module: This is implemented using an error correction algorithm that matches the watermark information generation module. The specific steps are as follows:

[0158] S7-1, separating the information bits, check bits and version bits in the decoding process from the extracted watermark information according to the payload bits, ECC check code bits and version information bits in the watermark information generation module.

[0159] S7-2, use the BCH decoding algorithm to correct errors in the payload bits and obtain the corrected decoded string s2.

[0160] S7-3, by comparing the decoded string s2 with the user-embedded string s1, the source of the screen capture image can be traced.

[0161] It should be noted that for the watermark embedding and extraction modules, different network models can be selected and different network structures can be built for implementation based on factors such as time efficiency, hardware support conditions, and accuracy.

[0162] It is understandable that in the specific implementation of this application, related data such as user information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0163] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0164] According to another aspect of the embodiment of the present application, a screen image processing device for implementing the above-mentioned screen image processing method is also provided. Figure 5 As shown, the device includes:

[0165] An acquisition module 502 is configured to acquire an original image and a screen capture image, wherein the original image includes original watermark information, and the screen capture image represents an image obtained by acquiring the original image through an image acquisition device;

[0166] An execution module 504 is configured to perform an image matching operation and a perspective change operation on the original image and the captured screen image to obtain a target area image, wherein the image matching operation is used to determine a first set of coordinate points and a second set of coordinate points, the first set of coordinate points corresponding to the original image, and the second set of coordinate points corresponding to the captured screen image, the first set of coordinate points and the second set of coordinate points being used to generate a first homography matrix, the first homography matrix being used to determine an initial area image in the captured screen image corresponding to the original image, and the perspective change operation being used to adjust the image range and tilt angle of the initial area image to obtain the target area image;

[0167] Extraction module 506, used to perform watermark extraction operation on the target area image to obtain screen watermark information;

[0168] The comparison module 508 is used to compare the captured watermark information with the original watermark information to obtain a target comparison result.

[0169] As an optional solution, the above-mentioned device is used to perform a watermark extraction operation on the target area image to obtain the screen capture watermark information in the following manner, including: converting the target area image into a grayscale image, and performing a fast Fourier transform operation on the grayscale image to obtain a spectrum diagram of the grayscale image; when there is a component higher than the target threshold in the spectrum diagram, performing a moiré removal operation on the target area image to obtain a denoised area image; performing a watermark extraction operation on the denoised area image to obtain the screen capture watermark information.

[0170] As an optional solution, the above-mentioned device is used to perform a watermark extraction operation on the target area image in the following manner to obtain the screen watermark information: perform a scale transformation and normalization operation on the target area image to obtain the update area image; extract high-dimensional features of the update area image, wherein the high-dimensional features are used to indicate the screen watermark information in the update area image; perform a decoding operation on the high-dimensional features to obtain the screen watermark information.

[0171] As an optional solution, the above-mentioned device is used to compare the screen capture watermark information and the original watermark information in the following manner to obtain the target comparison result: perform codeword error correction operation on the screen capture watermark information to obtain a decoded string, and obtain the watermark string corresponding to the original watermark information; compare the decoded string and the watermark string to obtain the target comparison result.

[0172] As an optional solution, the above-mentioned device is used to perform image matching operations and perspective change operations on the original image and the captured image in sequence in the following manner to obtain a target area image: performing an image matching operation on the original image and the captured image to obtain a first coordinate point set and a second coordinate point set, and determining a first homography matrix based on the first coordinate point set and the second coordinate point set; using the first homography matrix to perform a perspective forward transformation operation on the original image to obtain a third coordinate point set, and determining a second homography matrix based on the third coordinate point set and the first coordinate point set, wherein the third coordinate point set is used to indicate the image range of the initial area image, the second homography matrix is ​​used to adjust the inclination angle of the initial area image, and the perspective transformation operation includes a perspective forward transformation operation; using the second homography matrix to perform a perspective inverse transformation operation on the initial area image to obtain a target area image, wherein the perspective transformation operation includes a perspective inverse forward transformation operation.

[0173] As an optional solution, the above-mentioned device is also used to: obtain a watermark string and set the encoding length and version data of the watermark string; perform an encoding operation on the watermark string to obtain a payload; pad the length of the payload to obtain padding data, and generate a check code for the padding data; splice the payload, check code and version data to obtain the original watermark information.

[0174] As an optional solution, the above-mentioned device is also used to: obtain an initial watermark-free image, and perform scale transformation and normalization operations on the initial watermark-free image to obtain a target watermark-free image; obtain a watermark-free feature map of the target watermark-free image, and input the target watermark-free image and the original watermark information into a target encoder to obtain a watermark-containing feature map; perform residual calculation on the watermark-containing feature map and the watermark-free feature map to obtain a residual image; and add the scale-transformed residual image and the target watermark-free image to obtain the original image.

[0175] As an optional solution, the above-mentioned device is also used to: obtain an original image and a screen image; perform an image matching operation on the original image and the screen image to obtain a first coordinate point set and a second coordinate point set, and determine a first homography matrix based on the first coordinate point set and the second coordinate point set; use the first homography matrix to perform a perspective forward transformation operation on the original image to obtain a third coordinate point set, and determine a second homography matrix based on the third coordinate point set and the first coordinate point set, wherein the third coordinate point set is used to indicate the image range of the initial area image, the second homography matrix is ​​used to adjust the tilt angle of the initial area image, and the perspective transformation operation includes a perspective forward transformation operation; use the second homography matrix Perform an inverse perspective transform operation on the initial area image to obtain a target area image, wherein the perspective transform operation includes an inverse perspective transform operation; convert the target area image into a grayscale image, and perform a fast Fourier transform operation on the target grayscale image to obtain a spectrum diagram of the grayscale image; when there are components higher than the target threshold in the spectrum diagram, perform a moiré removal operation on the target area image to obtain a denoised area image; extract high-dimensional features of the denoised area image, and determine the screen watermark information based on the high-dimensional features; perform a codeword error correction operation on the screen watermark information to obtain a decoded string, and obtain the watermark string corresponding to the original watermark information; compare the decoded string and the watermark string to obtain a target comparison result.

[0176] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0177] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0178] According to one aspect of the present application, a computer program product is provided, which includes a computer program.

[0179] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0180] Figure 6 The block diagram schematically shows a computer system structure of an electronic device used to implement an embodiment of the present application.

[0181] It should be noted that Figure 6The computer system 600 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0182] like Figure 6 As shown, the computer system 600 includes a central processing unit 601 (CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 602 (ROM) or the program loaded from the storage part 608 into the random access memory 603 (RAM). Various programs and data required for system operation are also stored in the random access memory 603. The central processing unit 601, the read-only memory 602 and the random access memory 603 are connected to each other via a bus 604. An input / output interface 605 (i.e., an I / O interface) is also connected to the bus 604.

[0183] The following components are connected to the input / output interface 605: an input section 606 including a keyboard, a mouse, and the like; an output section 607 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a local area network card or a modem. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output interface 605 as needed. Removable media 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like, are installed in the drive 610 as needed, so that computer programs read therefrom can be installed into the storage section 608 as needed.

[0184] In particular, according to an embodiment of the present application, the processes described in the various method flow charts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the methods shown in the flow charts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609 and / or installed from a removable medium 611. When the computer program is executed by the central processing unit 601, the various functions defined in the system of the present application are performed.

[0185] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit 601, various functions provided by the embodiments of the present application are performed.

[0186] According to another aspect of the embodiment of the present application, an electronic device for implementing the above-mentioned screen image processing method is also provided. The electronic device may be Figure 1 The terminal device or server shown in FIG. This embodiment is described by taking the electronic device as a terminal device as an example. Figure 7 As shown, the electronic device includes a memory 702 and a processor 704. The memory 702 stores a computer program, and the processor 704 is configured to execute the steps in any of the above method embodiments through the computer program.

[0187] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0188] Optionally, in this embodiment, the above-mentioned processor can be configured to execute the methods in each embodiment of the present application through a computer program.

[0189] Alternatively, those skilled in the art will appreciate that Figure 7 The structure shown is for illustration only. Figure 7 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 7 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 7 Different configurations shown.

[0190] Among them, the memory 702 can be used to store software programs and modules, such as the program instructions / modules corresponding to the method and device for processing screen images in the embodiments of the present application. The processor 704 executes various functional applications and data processing by running the software programs and modules stored in the memory 702, that is, realizing the above-mentioned method for processing screen images. The memory 702 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 702 may further include a memory remotely located relative to the processor 704, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 702 can be used specifically, but not limited to, to store information such as screen images and original images. As an example, if Figure 7As shown, the memory 702 may include, but is not limited to, the acquisition module 502, execution module 504, extraction module 506, and comparison module 508 in the device for processing the captured screen image. Furthermore, the memory 702 may also include, but is not limited to, other module units in the device for processing the captured screen image, which will not be described in detail in this example.

[0191] Optionally, the transmission device 706 is configured to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 706 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 706 is a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0192] In addition, the electronic device further includes: a display 708 for displaying the captured image and the original image; and a connection bus 710 for connecting various module components in the electronic device.

[0193] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes via network communication. The nodes may form a peer-to-peer network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.

[0194] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the screen capture image processing method provided in various optional implementation methods of the above-mentioned screen capture image processing.

[0195] Optionally, in this embodiment, the above-mentioned computer-readable storage medium can be configured to store data for executing the methods in various embodiments of the present application.

[0196] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0197] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0198] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing one or more electronic devices to execute all or part of the steps of the method described in each embodiment of the present application.

[0199] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0200] In the several embodiments provided in this application, it should be understood that the disclosed applications can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0201] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0202] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0203] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for processing a screen shot image, characterized in that: include: Acquire an original image and a screen shot image, wherein the original image includes original watermark information, and the screen shot image represents an image obtained by acquiring the original image through an image acquisition device; Performing an image matching operation and a perspective change operation on the original image and the screen image to obtain a target area image, wherein the image matching operation is used to determine a first coordinate point set and a second coordinate point set, the first coordinate point set corresponds to the original image, the second coordinate point set corresponds to the screen image, the first coordinate point set and the second coordinate point set are used to generate a first homography matrix, the first homography matrix is ​​used to determine an initial area image in the screen image that corresponds to the original image, and the perspective change operation is used to adjust the image range and tilt angle of the initial area image to obtain the target area image; Performing a watermark extraction operation on the target area image to obtain screen watermark information; Compare the captured screen watermark information with the original watermark information to obtain a target comparison result.

2. The method according to claim 1, characterized in that The performing of a watermark extraction operation on the target area image to obtain screen capture watermark information includes: Converting the target area image into a grayscale image, and performing a fast Fourier transform operation on the grayscale image to obtain a frequency spectrum of the grayscale image; When there is a component higher than a target threshold in the spectrum graph, performing a moiré removal operation on the target region image to obtain the denoised region image; The watermark extraction operation is performed on the denoised area image to obtain the screen capture watermark information.

3. The method according to claim 1, characterized in that The performing of a watermark extraction operation on the target area image to obtain screen capture watermark information includes: Performing scale transformation and normalization operations on the target region image to obtain an updated region image; Extracting high-dimensional features of the update region image, wherein the high-dimensional features are used to indicate the screen-shot watermark information in the update region image; A decoding operation is performed on the high-dimensional features to obtain the screen capture watermark information, wherein the decoding operation corresponds to an encoding operation used in the generation process of the original image, and the encoding operation is used to embed the original watermark information in the original image.

4. The method according to claim 1, characterized in that: The comparing the screen shot watermark information with the original watermark information to obtain a target comparison result includes: Performing a codeword error correction operation on the screen-captured watermark information to obtain a decoded string, and obtaining a watermark string corresponding to the original watermark information; The decoded character string is compared with the watermark character string to obtain the target comparison result.

5. The method according to claim 1, characterized in that The step of sequentially performing an image matching operation and a perspective change operation on the original image and the captured screen image to obtain a target area image includes: Performing the image matching operation on the original image and the captured screen image to obtain the first coordinate point set and the second coordinate point set, and determining the first homography matrix according to the first coordinate point set and the second coordinate point set; Using the first homography matrix to perform a perspective forward transformation operation on the original image to obtain a third coordinate point set, and determining a second homography matrix according to the third coordinate point set and the first coordinate point set, wherein the third coordinate point set is used to indicate an image range of the initial region image, the second homography matrix is ​​used to adjust the tilt angle of the initial region image, and the perspective transformation operation includes the perspective forward transformation operation; The second homography matrix is ​​used to perform an inverse perspective transformation operation on the initial region image to obtain the target region image, wherein the perspective transformation operation includes the inverse perspective transformation operation.

6. The method according to claim 1, characterized in that The method further comprises: Obtain a watermark string, and set the encoding length and version data of the watermark string; Performing encoding operation on the watermark character string to obtain a payload; Filling the length of the effective load to obtain filling data, and generating a check code of the filling data; The payload, the check code and the version data are concatenated to obtain the original watermark information.

7. The method according to claim 6, characterized in that The method further comprises: Acquire an initial watermark-free image, and perform a scale transformation and normalization operation on the initial watermark-free image to obtain a target watermark-free image; Acquire a watermark-free feature map of the target watermark-free image, and input the target watermark-free image and the original watermark information into a target encoder to obtain a watermark-containing feature map; Performing residual calculation on the watermarked feature map and the non-watermarked feature map to obtain a residual image; The scale-transformed residual image and the target watermark-free image are added to obtain the original image.

8. The method according to claim 1, characterized in that The method further comprises: Acquire the original image and the captured screen image; Performing the image matching operation on the original image and the captured screen image to obtain the first coordinate point set and the second coordinate point set, and determining the first homography matrix according to the first coordinate point set and the second coordinate point set; Using the first homography matrix to perform a perspective forward transformation operation on the original image to obtain a third coordinate point set, and determining a second homography matrix according to the third coordinate point set and the first coordinate point set, wherein the third coordinate point set is used to indicate an image range of the initial region image, the second homography matrix is ​​used to adjust the tilt angle of the initial region image, and the perspective transformation operation includes the perspective forward transformation operation; Using the second homography matrix, performing an inverse perspective transformation operation on the initial region image to obtain a target region image, wherein the perspective transformation operation includes the inverse perspective transformation operation; Converting the target area image into a grayscale image, and performing a fast Fourier transform operation on the target grayscale image to obtain a frequency spectrum of the grayscale image; When there is a component higher than a target threshold in the spectrum graph, performing a moiré removal operation on the target region image to obtain a denoised region image; Extracting high-dimensional features of the denoised area image, and determining the screen watermark information according to the high-dimensional features; Performing a codeword error correction operation on the screen-captured watermark information to obtain a decoded string, and obtaining a watermark string corresponding to the original watermark information; The decoded character string is compared with the watermark character string to obtain the target comparison result.

9. A screen image processing device, characterized in that: include: An acquisition module, used to acquire an original image and a screen shot image, wherein the original image includes original watermark information, and the screen shot image represents an image obtained by acquiring the original image through an image acquisition device; an execution module, configured to perform an image matching operation and a perspective change operation on the original image and the screen image to obtain a target area image, wherein the image matching operation is used to determine a first coordinate point set and a second coordinate point set, the first coordinate point set corresponds to the original image, the second coordinate point set corresponds to the screen image, the first coordinate point set and the second coordinate point set are used to generate a first homography matrix, the first homography matrix is ​​used to determine an initial area image in the screen image corresponding to the original image, and the perspective change operation is used to adjust the image range and tilt angle of the initial area image to obtain the target area image; An extraction module, used to perform a watermark extraction operation on the target area image to obtain screen watermark information; The comparison module is used to compare the screen shot watermark information with the original watermark information to obtain a target comparison result.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein the computer program can be executed by an electronic device to perform the method described in any one of claims 1 to 8.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 8 are implemented.

12. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 8 through the computer program.