Image processing method and device, equipment, storage medium and program product
By automatically selecting structurally stable and semantically important regions in an image as watermark embedding locations using a watermark calibration model, the problem of low robustness of watermarks in existing technologies is solved, achieving efficient anti-attack capabilities and reliable watermark detection.
Patent Information
- Application Number
- CN202511769275.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-02-27
AI Technical Summary
In existing technologies, improper selection of digital watermark embedding positions leads to low watermark robustness, making it susceptible to attacks such as rotation, translation, and scaling, resulting in detection failure or extraction errors.
A watermark calibration model is used to generate an image content confidence mask image. The system automatically selects structurally stable and semantically important regions in the image as watermark embedding locations. By dynamically selecting highly robust embedding locations, the binding strength between the watermark and the original image is enhanced, thus resisting attacks.
It improves the robustness of watermarks, reduces the extraction error rate, enhances the integrity and detectability of watermarks, and avoids the limitations of manual point selection and the defects of preset templates.
Smart Images

Figure CN121582048A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of image processing, and particularly relates to an image processing method, device, equipment, storage medium and program product. BACKGROUND
[0002] The popularity of computers makes digital products such as digital images have the characteristics of easy and rapid transmission in the Internet, but this characteristic can also be used by pirates to threaten the development of the digital product industry. In order to prevent others from stealing pictures and protect the works of the original author, digital watermarking emerges as the times require. Among them, the selection of the digital watermark embedding position is the key to ensuring that the watermark generation method is effectively applied to copyright protection.
[0003] In the related art, the digital watermark embedding position can be selected manually or determined according to a preset template, so as to embed the watermark into the original image based on the digital watermark embedding position to obtain a watermark image. If the position of the watermark embedding is selected improperly by the foregoing manner, the watermark image is more likely to be attacked by rotation, translation, scaling, clipping and the like. These attacks will change the content and structure of the watermark image, destroy the spatial relationship between the watermark and the original image, cause the watermark detection in the watermark image to fail or the extraction to be wrong, and result in low watermark robustness. SUMMARY
[0004] Embodiments of the present application provide an image processing method, device, equipment, storage medium and program product, which can solve the problem of low watermark robustness caused by inaccurate determination of the watermark embedding position.
[0005] In a first aspect, embodiments of the present application provide an image processing method, which can include: obtaining a first image and digital watermark content; inputting the first image into a watermark calibration model, and obtaining an image content confidence mask image of the first image output by the watermark calibration model through the watermark calibration model, the image content confidence mask image being an image used to represent the attention degree of the watermark calibration model to different regions in the first image; determining a first position according to the image content confidence mask image, the first position being a position of embedding the digital watermark content into the first image, and a value of the attention degree corresponding to the first position in the image content confidence mask image being greater than or equal to a value of a preset attention degree; fusing the digital watermark content and the first image based on the first position to obtain a second image, the second image being an image after the first image embeds the digital watermark content.
[0006] In a second aspect, embodiments of the present application provide an image processing device, which includes: an obtaining module configured to obtain a first image and digital watermark content; The model module is configured to input the first image into a watermark calibration model, and obtain an image content confidence mask image of the first image output by the watermark calibration model through the watermark calibration model, the image content confidence mask image being an image used to represent a degree of attention of the watermark calibration model to different regions in the first image. The determination module is configured to determine a first position according to the image content confidence mask image, the first position being a position where the digital watermark content is embedded into the first image, and a value of the degree of attention corresponding to the first position in the image content confidence mask image being greater than or equal to a preset value of the degree of attention. The fusion module is configured to fuse the digital watermark content and the first image based on the first position, and obtain a second image, the second image being an image in which the digital watermark content is embedded into the first image.
[0007] In a third aspect, an embodiment of the present application provides a computer device, which comprises a processor and a memory storing computer program instructions; and the processor implements the image processing method according to any one of the first aspect when executing the computer program instructions.
[0008] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores computer program instructions; and the computer program instructions are executed by a processor to implement the image processing method according to any one of the first aspect.
[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, which comprises computer programs or instructions; and the computer programs or instructions are executed by a processor to implement the image processing method according to any one of the first aspect.
[0010] In a sixth aspect, an embodiment of the present application provides a chip, which comprises a processor and a display interface, the display interface being coupled to the processor, and the processor being configured to run programs or instructions to implement the image processing method according to any one of the first aspect.
[0011] The image processing method, device, equipment, storage medium and program product provided by the embodiments of the present application obtain a first image and digital watermark content; input the first image into a watermark calibration model, obtain an image content confidence mask image of the first image output by the watermark calibration model through the watermark calibration model, and the image content confidence mask image is an image used to represent the attention degree of the watermark calibration model to different regions in the first image; determine a first position according to the image content confidence mask image, the first position is a position where the digital watermark content is embedded in the first image, and the value of the attention degree corresponding to the first position in the image content confidence mask image is greater than or equal to a preset value of the attention degree; and fuse the digital watermark content and the first image based on the first position to obtain a second image, the second image being an image after the first image is embedded with the digital watermark content. In this way, the image content confidence mask image is generated through the watermark calibration model, the region with stable structure and important semantics in the image, such as the part with rich texture and clear edge, is automatically selected as the watermark embedding position, the sensitivity of these regions to geometric transformations such as rotation, translation and scaling is relatively low, the destruction of the attack to the watermark-image space relationship can be effectively resisted, the robustness defects caused by the traditional manual selection or fixed position of the template are avoided, and the anti-attack capability is enhanced. The first position is dynamically determined according to the value of the attention degree in the mask image, so that the watermark can be embedded in the region with high image content confidence, the adaptive mechanism avoids the limitation of the preset template, the watermark position is optimized according to the image content, the detection failure or extraction error caused by improper position is reduced, and the embedding position is dynamically optimized. When the second image is generated by fusion, the watermark is embedded in the region with high confidence, the combination strength of the watermark and the original image is enhanced, the attack is difficult to destroy these key regions, so that the integrity and detectability of the watermark are maintained, the extraction error rate is significantly reduced, and the watermark detection reliability is improved. In this way, the watermark calibration model is used to replace the manual selection of the position, the automatic calibration of the embedding position is realized, the efficiency is improved and the human bias is reduced, the image content confidence mask image is generated by analyzing the image content features such as texture and edge through the watermark calibration model, the embedding position is naturally integrated with the image structure, the anti-attack performance is further improved, the high-robustness embedding position is dynamically selected, and the problem that the watermark detection failure or extraction error occurs in the watermark image caused by the attack in the traditional method is effectively solved, and the watermark robustness is improved. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced as follows, and other drawings can also be obtained by those of ordinary skill in the art without creating any creative labor on the premise that the drawings are not attached.
[0013] Figure 1 A flowchart of an image processing method provided by some embodiments of the present application is shown; FIG. 2(a) shows a schematic diagram of a first image of an image processing method according to some embodiments of the present application; FIG. 2(b) shows a schematic diagram of a digital watermark content of an image processing method according to some embodiments of the present application; Figure 3 FIG. 3 shows a schematic diagram of a flow of processing an image by a watermark calibration model of an image processing method according to some embodiments of the present application; Figure 4 FIG. 4 shows a schematic diagram of a flow of processing an image by an internal watermark calibration model of an image processing method according to some embodiments of the present application; FIG. 5(a) shows a schematic diagram of an image content confidence mask image of an image processing method according to some embodiments of the present application; FIG. 5(b) shows a schematic diagram of an image content confidence mask image of an image processing method according to some embodiments of the present application; Figure 6 FIG. 6 shows a schematic diagram of a second image of an image processing method according to some embodiments of the present application; Figure 7 FIG. 7 shows a schematic diagram of a digital watermark content after HSL color data adjustment of an image processing method according to some embodiments of the present application; Figure 8 FIG. 8 shows a schematic diagram of a second image after HSL color data adjustment of an image processing method according to some embodiments of the present application; Figure 9 FIG. 9 shows a schematic diagram of a hue circle of HSL color data of an image processing method according to some embodiments of the present application; Figure 10 FIG. 10 shows a schematic diagram of a local digital watermark content after HSL color data adjustment of an image processing method according to some embodiments of the present application; Figure 11 FIG. 11 shows a schematic diagram of a local digital watermark enlargement after HSL color data adjustment of an image processing method according to some embodiments of the present application; Figure 12 FIG. 12 shows a schematic diagram of an adjustment control of an image processing method according to some embodiments of the present application; Figure 13 FIG. 13 shows a schematic diagram of an effect of a watermark image generated by a conventional watermark generation method; Figure 14 FIG. 14 shows a schematic diagram of a structure of an image processing apparatus according to some embodiments of the present application; Figure 15 FIG. 15 shows a schematic diagram of a structure of a computer device according to some embodiments of the present application. DETAILED DESCRIPTION
[0014] The features and exemplary embodiments of the various aspects of the present application will be described in detail below with reference to the drawings. The following detailed description is merely exemplary in nature and is not intended to limit the present application or the application and uses of the present application. Furthermore, there is no intention to be bound by any expressed or implied theory presented in the preceding technical field, background, brief summary or the following detailed description.
[0015] It should be noted that the relational terms herein, such as first and second, and the like, are used solely to distinguish one from another entity or action without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0016] It should be noted that the acquisition, storage, use and processing of data in the embodiments of the present application comply with the relevant provisions of national laws and regulations.
[0017] It should be noted that in the embodiments of the present application, some software, components, models and other industry existing solutions may be mentioned, which should be considered as exemplary, and the purpose is only to illustrate the feasibility of the implementation of the technical solutions of the present application, but does not mean that the applicant has or will necessarily use the solution.
[0018] Before describing the technical solutions provided by the embodiments of the present application, in order to facilitate the understanding of the embodiments of the present application, the present application first specifically describes the related technologies involved.
[0019] At present, the popularity of computers makes digital products such as digital images have the characteristics of easy and rapid spread on the Internet, but this characteristic can also be used by pirates to threaten the development of the digital product industry. In addition, due to the increasing demand for digital products, digital product processing technology is widely used in people's life and work, and image processing software such as Photoshop with high intelligence and simple operation enables many ordinary people without any professional knowledge to easily modify and edit images, thereby causing a series of potential information security problems. In order to prevent others from stealing pictures, digital watermarking emerges as the times require. Among them, the selection of the embedding position of the digital watermark is the key to ensuring that the watermark generation method is effectively applied to copyright protection.
[0020] In the embodiments of the present application, digital watermarking, simply referred to as watermarking, refers to the addition of specific information markers such as logos, icons, etc. on pictures, videos, application programs, web pages or other media, and is an information protection and content authentication means for preventing others from stealing pictures, copyright protection, anti-counterfeiting, etc. At present, digital watermarking is an invisible marker embedded in digital content, which hides the watermark information in the host data through a specific algorithm without affecting the normal use of the host data, but can be extracted through a specific detection means. Digital watermarking is widely used in the fields of copyright protection, content authentication, data tracking, etc. Digital watermarking technology focuses on designing various watermark embedding algorithms for copyright protection, which is of great significance to protect digital resources from infringement. The selection of watermark information expression and embedding position is the key to ensuring that the watermark algorithm is effectively applied to copyright protection. The determination of the watermark embedding position is crucial to ensure the robustness and invisibility of the watermark. Watermark embedding technology embeds watermarks in the specific frequency domain or spatial domain of digital works, so that the watermark is closely combined with the original work, thereby playing a key role in copyright protection and content authentication. The selection of the embedding position needs to ensure the security of the watermark information to prevent it from being easily extracted or tampered with by unauthorized third parties. At the same time, selecting a suitable embedding position can enhance the robustness of the watermark, so that it can maintain its integrity after the image undergoes various processing such as compression, cropping, noise addition, etc., thereby ensuring the effectiveness of the watermark. Therefore, it is crucial to select a suitable watermark embedding position according to the image content.
[0021] In the related art, the digital watermark embedding position can be determined by manual selection or according to a preset template, so as to embed the watermark into the original image based on the digital watermark embedding position to obtain a watermark image. If the position of the watermark embedding is selected improperly in the foregoing manner, the watermark image will be more susceptible to attacks such as rotation, translation, scaling, cropping, etc. These attacks will change the content and structure of the watermark image, destroy the spatial relationship between the watermark and the original image, and cause the watermark detection in the watermark image to fail or the watermark to be extracted incorrectly, resulting in low watermark robustness.
[0022] However, the determination method of the watermark embedding position has the following disadvantages. Disadvantage 1: Many watermark generation methods only determine the specific embedding position of the watermark in the medium-high frequency domain with the naked eye or simple analysis, so that the selected position parameters are often random or non-optimal, and the image content characteristics are not combined for watermark embedding, which will directly affect the quality and effectiveness of watermark embedding. For multi-party shared data, traditional single watermark technology is difficult to guarantee the copyright protection demand of multiple key contents of the image, so it is necessary to balance the interests of each party. Disadvantage 2: The anti-attack ability of image watermark is low, because geometric attacks destroy the spatial relationship between the watermark and the picture, greatly affecting the robustness of the watermark. Geometric attacks include rotation, translation, scaling, clipping and the like, which change the content and structure of the image, so that the information originally embedded in the watermark is easily weakened, resulting in that the watermark information cannot be accurately extracted or recognized when detecting the watermark. In addition, image watermark research has its own particularity and complexity, including the large amount of image data, the diversification of image compression standards and the particularity of data structure, thereby affecting the ability to resist geometric attacks. Disadvantage 3: Visible watermark can be clearly seen, which greatly destroys the picture effect and easily affects information acquisition. The conventional watermark usually has fixed color and is superimposed on the subject layer, and the visibility is controlled through opacity, but such watermark generation method: when the watermark color is too different from the subject content, the overall effect of the picture of the subject layer will be affected, causing visual interference; on the other hand, when the watermark color is similar or the same as the subject content, the watermark content cannot be clearly seen.
[0023] To solve the problems in the related art, the embodiments of the present application provide an image processing method, device, equipment, storage medium and program product. The image processing method provided by some embodiments of the present application is described in detail below in combination with the accompanying Figure 1 to the accompanying Figure 15 , the image processing method provided by the embodiments of the present application is described in detail through specific embodiments and application scenarios.
[0024] Figure 1 The flowchart of the image processing method provided by some embodiments of the present application is shown. As shown in the figure, Figure 1 The image processing method can include steps 110 to 130.
[0025] Step 110, obtaining a first image and digital watermark content; Step 120, inputting the first image into a watermark calibration model, obtaining an image content confidence mask image of the first image output by the watermark calibration model through the watermark calibration model, the image content confidence mask image being an image for representing the attention degree of the watermark calibration model to different regions in the first image; Step 130, determining a first position according to the image content confidence mask image, the first position being a position of embedding the digital watermark content into the first image, and a value of the attention degree corresponding to the first position in the image content confidence mask image being greater than or equal to a preset value of the attention degree; Step 140, fusing the digital watermark content and the first image based on the first position to obtain a second image, the second image being an image after the first image is embedded with the digital watermark content.
[0026] Thus, the image content confidence mask image is generated by the watermark calibration model, the region with stable structure and important semantics in the image such as the part with rich texture and clear edge is automatically selected as the watermark embedding position, the sensitivity of these regions to geometric transformations such as rotation, translation and scaling is low, which can effectively resist the destruction of the attack on the watermark-image space relationship, avoid the robustness defects caused by the traditional manual selection or fixed position of the template, and enhance the anti-attack ability. The first position is dynamically determined according to the attention degree value in the mask image, which can ensure that the watermark is embedded in the region with high image content confidence. This adaptive mechanism avoids the limitations of the preset template, optimizes the watermark position according to the image content, reduces the detection failure or extraction error caused by improper position, and dynamically optimizes the embedding position. When the second image is generated by fusion, the watermark is embedded in the high-confidence region, which enhances the combination strength of the watermark and the original image, makes it difficult for the attack to destroy these key regions, thereby maintaining the integrity and detectability of the watermark, significantly reducing the extraction error rate, and improving the reliability of watermark detection. In this way, the watermark calibration model is used to replace the manual selection of the position, the embedding position is automatically calibrated, the efficiency is improved and the human bias is reduced, and the image content confidence mask image is generated by analyzing the image content features such as texture and edge through the watermark calibration model, which ensures that the embedding position is naturally integrated with the image structure, further consolidates the anti-attack performance, dynamically selects the high-robustness embedding position, and effectively solves the problem of watermark detection failure or extraction error in the watermark image caused by the attack in the traditional method, thereby improving the watermark robustness.
[0027] The above steps will be described in detail as follows.
[0028] Firstly, step 110 is related to, the first image and the digital watermark content can be obtained. Wherein, the first image can refer to FIG. 2(a), and the digital watermark content can refer to FIG. 2(b), such as “watermark watermark”.
[0029] Then, step 120 is involved, in the embodiments of the present application, the spatial coordinate information in the change feature can be efficiently integrated into the attention feature by decomposing the channel feature into two parallel 1D features, and then the coordinate feature is encoded by using the Transformer structure to mine deeper global spatial dependency between the change pixels, so as to implicitly construct the position expression of the elements in the image space, that is, the decision-level position information, so as to visualize the important content of the image, as shown below.
[0030] In some embodiments of the present application, the watermark calibration model comprises a spatial attention mechanism layer and a spatial attention mechanism reinforcement layer, based on which, step 120 can specifically comprise steps 1201 to 1203, as shown below.
[0031] Step 1201 inputs the first image into the watermark calibration model, and captures the decision-level position information of the first image by the spatial attention mechanism layer.
[0032] Exemplarily, as shown in Figure 3 , it is assumed that the first image 310 is a landscape photo on the grassland, there are two elephants in the picture, an adult elephant and a small elephant standing side by side, and the background is trees and sky, the first image 310 is input into the watermark calibration model 320, and the spatial attention mechanism layer in the watermark calibration model 320 is activated, which is like an experienced image analyst quickly scanning the entire picture and preliminarily judging which regions are the most important and most recognizable in vision, the mechanism layer will preliminarily lock the key subjects in the image, it will notice the outlines of the two elephants, especially their heads, tusks and huge bodies, because these regions contain rich edges, textures and unique structure information. These preliminarily identified key regions are the decision-level position information, that is, it can be understood that the watermark calibration model 320 initially considers that these places are more worthy of attention than the uniform grassland or sky background.
[0033] Step 1202 mines the global spatial dependency of the pixel points in the first image according to the decision-level position information of the first image by the spatial attention mechanism reinforcement layer, to obtain an image content confidence mask image.
[0034] Exemplarily, as shown in Figure 3As shown, the spatial attention mechanism reinforcement layer starts working, and it is not satisfied with merely identifying isolated key points, but like a strategic analyst, it deeply analyzes the global spatial dependency between these key points and all other pixels in the image, for example, it analyzes the spatial correlation between the back contour line of the adult elephant and the contour of the small elephant, and the trees behind them. It will be found that the back area of the elephant is not only rich in texture itself, but also forms a stable and coherent visual structure with the surrounding environment such as the small elephant and the trees. In contrast, a single shadow area under the abdomen of the small elephant and the grassland, although it is also part of the small elephant's body, has a weak spatial correlation with the whole, and the watermark may be lost if the picture is slightly cropped. After this complex global relationship analysis, the reinforcement layer finally generates an image content confidence mask image 330, which is a heat map with the same size as the original image, in which the contours of the two elephants, the elephant tusks, and the wrinkle texture of the elephant ears are highlighted in high brightness, for example, red, yellow, and cyan, indicating that the watermarking model 320 pays great attention to these areas and considers them to be excellent locations for embedding digital watermarks. The sky and part of the flat grassland area are in low brightness, for example, blue-violet, indicating that these areas have low attention and are not suitable for embedding.
[0035] In step 1203, an image content confidence mask image of the first image output by the spatial attention mechanism reinforcement layer is obtained.
[0036] Exemplarily, as Figure 3 As shown, the spatial attention mechanism reinforcement layer starts working, and it is not satisfied with merely identifying isolated key points, but like a strategic analyst, it deeply analyzes the global spatial dependency between these key points and all other pixels in the image, for example, it analyzes the spatial correlation between the back contour line of the adult elephant and the contour of the small elephant, and the trees behind them. It will be found that the back area of the elephant is not only rich in texture itself, but also forms a stable and coherent visual structure with the surrounding environment such as the small elephant and the trees. In contrast, a single shadow area under the abdomen of the small elephant and the grassland, although it is also part of the small elephant's body, has a weak spatial correlation with the whole, and the watermark may be lost if the picture is slightly cropped. After this complex global relationship analysis, the reinforcement layer finally generates an image content confidence mask image 330, which is a heat map with the same size as the original image, in which the contours of the two elephants, the elephant tusks, and the wrinkle texture of the elephant ears are highlighted in high brightness, for example, red, yellow, and cyan, indicating that the watermarking model 320 pays great attention to these areas and considers them to be excellent locations for embedding digital watermarks. The sky and part of the flat grassland area are in low brightness, for example, blue-violet, indicating that these areas have low attention and are not suitable for embedding.
[0037] In this way, a clear and data-based image content confidence mask image 330 with a heat relationship is obtained, which no longer relies on simple intuition but on complex spatial relationship calculations, and clearly marks which areas in the "two elephants" image are the most robust digital watermark embedding locations.
[0038] Thus, by extracting key information features in image vision through a watermarking model, performing watermark embedding, visualizing important image content according to a heat map, automatically generating and displaying a watermark using a hue, saturation, and lightness (HSL) color mode, and providing an adjustable interface for user adjustment.
[0039] In the embodiments of the present application, the above step 1201 can specifically include steps 12011 to 12014, as shown below.
[0040] Step 12011, input the first image into the watermark calibration model, encode the first image along the horizontal coordinate direction and the vertical coordinate direction respectively through the pooling kernel of the first size and the pooling kernel of the second size in the spatial attention mechanism layer, to obtain the first perceptual attention feature in the horizontal coordinate direction and the second perceptual attention feature in the vertical coordinate direction.
[0041] Exemplarily, the precise position information of the input feature is encoded along the horizontal coordinate and the vertical coordinate direction respectively using the one-dimensional pooling kernel with the size of H and W, to generate the attention feature with direction perception from two different dimensions. Specifically, as shown in Figure 4 , the first image x is encoded using the first size W, that is, W-dimensional average pooling, to obtain the first perceptual attention feature , and the first image x is encoded using the second size H, that is, H-dimensional average pooling, to obtain the second perceptual attention feature .
[0042] Step 12012, channel concatenation is performed on the first perceptual attention feature and the second perceptual attention feature through the convolution in the spatial attention mechanism layer to fuse the cross-channel relationship, to obtain the comprehensive perceptual attention feature map.
[0043] Exemplarily, still referring to Figure 4 , the first perceptual attention feature and the second perceptual attention feature are concatenated in the channel dimension, the cross-channel relationship is embedded into the feature information through 1 Figure 4 1 convolution, and then the comprehensive perceptual attention feature map f is decomposed into two separate first tensor and second tensor
[0044] .
[0045] Exemplarily, the two kinds of transformations are concatenated in the channel dimension, 1 1 convolution is used for fusion, the cross-channel relationship is embedded into the feature information, and then the comprehensive perceptual attention feature map f is decomposed into two separate first tensor and second tensor , that is, in Figure 4 , the first tensor can be represented by C H
[0046] Step 12014, the first tensor and the second tensor are respectively subjected to convolution transformation and activation function processing to obtain the decision-level position information of the first image.
[0047] In the embodiment of the present application, the first tensor and the second tensor are respectively subjected to convolution transformation and activation function processing to obtain two weights, the two weights are multiplied to realize weighting operation, and the decision-level position information of the first image is obtained to extract the element coordinate information. Based on this, the decision-level position information of the first image can be obtained through the following formula (1) and formula (2): Wherein, δ is a nonlinear activation function, Cov is 1 1 convolution function, is a sigmoid activation function, , represents 1 1 convolution operation in two directions.
[0048] Exemplarily, as Figure 4 shown, the first tensor and the second tensor can be processed through 1 1 convolution and activation function (sigmoid) respectively, two weights are obtained, the two weights are multiplied to realize weighting operation, and the element coordinate information, i.e. the decision-level position information, is extracted. The decision-level position information can be represented by C (CxHxW) in Figure 4 .
[0049] In the embodiment of the present application, the above step 1202 can specifically include steps 12021 to 12026, which are specifically shown as follows.
[0050] Step 12021, through the spatial attention mechanism reinforcement layer, the vector set is generated according to the decision-level position information of the first image, and the vector set includes key vectors, value vectors and query vectors.
[0051] Exemplarily, still referring to Figure 4 , the decision-level position information C of the first image is input into the spatial attention mechanism reinforcement layer, the interdependence relationship of the decision-level position information can be determined through the spatial attention mechanism reinforcement layer, i.e. through 3 1 convolution to process the decision-level position information C, to obtain a vector set, which includes key vectors Q, value vectors K and query vectors .
[0052] Step 12022, each vector in the vector set is subjected to deformation processing to obtain a deformed vector set, and the deformed vector set includes deformed key vectors of key vectors, deformed value vectors of value vectors and deformed query vectors of query vectors.
[0053] Exemplarily, still referring to Figure 6 , a deformation key vector of the key vector is obtained by deforming the key vector , a deformation value vector of the value vector , a deformation query vector of the query vector , wherein .
[0054] Step 12023, a similarity matrix is constructed based on the deformation key vector and the deformation value vector, and the similarity matrix is used to represent the association degree of at least two pixel points in the first image.
[0055] Exemplarily, the similarity matrix A is calculated by using the deformation key vector and the deformation value vector, wherein A .
[0056] Step 12024, the deformation query vector is transformed by the similarity matrix to obtain a first position feature, and the first position feature is used to represent the relative position relationship between at least two pixel points in the first image.
[0057] Exemplarily, the first position feature is obtained by multiplying the deformation query vector and the similarity matrix A, and the specific process can be referred to the following formula (3): Step 12025, the space dimension of the first position feature is reshaped to obtain a second position feature, and the second position feature is used to represent the spatial position relationship between at least two pixel points in the first image.
[0058] Step 12026, the first image is enhanced by the second position feature to obtain an image content confidence mask image.
[0059] In the embodiments of the present application, the above steps 12025 and 12026 can be implemented by the following formula (4) and formula (5), and the specific process is as follows: , wherein is a softmax activation function, and the first position feature is scaled according to , as shown in formula (4), the first position feature Y is reshaped to the second position feature , and the second position feature is multiplied by the first image X to obtain an image content confidence mask image after being enhanced by a channel-spatial attention module (CSAM), as shown in formula (5).
[0060] Thus, this process can mine the global spatial dependence between pixels, and construct the position expression of elements in the image space.
[0061] It should be noted that the watermark calibration model in the embodiments of the present application can be obtained by training the imageNet dataset. Thus, the first image to be embedded with a watermark is input into the watermark calibration model, and Grad-CAM technology is used to visualize the bottom-level decision-level features of the model to generate an image content confidence mask image. The imageNet dataset includes a sample image embedded with a watermark and a sample image content confidence mask image corresponding to the sample image. Thus, the proposed watermark calibration model can be used to learn the embedded watermark image, and the watermark is embedded by visualizing the decision-level features instead of the traditional visual calibration or rule calibration method, which better meets the objective needs of important content calibration.
[0062] Secondly, step 130 is involved. In some embodiments of the present application, the image content confidence mask image includes N regions, and the value of the attention degree of each region is different. Based on this, step 130 can specifically include step 1301 and step 1302, which are specifically as follows.
[0063] Step 1301: determining the value of the attention degree of each region according to the value of the attention degree of the pixel points in each region.
[0064] For example, referring to FIG. 5(a), assuming that there is an image content confidence mask image with a resolution of 6x6, which can be divided into N=4 regions, i.e., region A, region B, region C and region D. The goal is to find out which region is most suitable for embedding a watermark. Based on this, the average value of the attention degrees of all pixel points in the region is used as the attention degree value of the region. For example, the average values of the pixel point attention degrees of the four regions (A, B, C, D) are as follows: the attention degree value of region A: 0.15; the attention degree value of region B: 0.35; the attention degree value of region C: 0.75; and the attention degree value of region D: 0.40.
[0065] Alternatively, as shown in FIG. 5(b), assuming that there is an image content confidence mask image with a resolution of 6x6, which can be divided into N=3 regions, i.e., region A, region B and region C. The goal is to find out which region is most suitable for embedding a watermark. Based on this, the average value of the attention degrees of all pixel points in the region is used as the attention degree value of the region. For example, the average values of the pixel point attention degrees of the three regions (A, B, C) are as follows: the attention degree value of region A: 0.15; the attention degree value of region B: 0.95; and the attention degree value of region C: 0.75.
[0066] Step 1302, the position corresponding to the target region with the highest value of attention degree in the N regions is determined as the first position.
[0067] For example, referring to FIG. 5(a), from the result of step 1301, the region with the highest value is found to determine the first position. In the above data, the attention value of region C is 0.75, which is the highest. Therefore, the first position is the pixel range covered by region C.
[0068] Alternatively, as shown in FIG. 5(b), from the result of step 1301, the region with the highest value is found to determine the first position. In the above data, the attention value of region B is 0.95, which is the highest. Therefore, the first position is the pixel range covered by region B.
[0069] Then, step 140 is involved, which in some embodiments of the present application can specifically include steps 1401-1403, as shown below.
[0070] Step 1401, based on the first position, superimpose the digital watermark content on the subject layer of the first image to obtain the intersection region of the subject layer and the digital watermark content.
[0071] For example, the subject layer of the first image can be divided into an n*n pixel grid. Taking 1920*1080 as an example, superimpose the digital watermark content on the subject layer of the first image to obtain the intersection region of the digital watermark content and the subject layer. That is, superimpose the digital watermark content "watermark watermark" on the subject layer of the first image, and the intersection region of the digital watermark content in the subject layer, such as Figure 7 .
[0072] Step 1402, according to the HSL color data of the pixel points corresponding to the intersection region, adjust the HSL color data of the pixel points in the watermark content to obtain the adjusted digital watermark content.
[0073] For example, the HSL color data of the pixel points corresponding to the coordinates of the intersection region can be obtained, and the HSL color data of the pixel points in the watermark content is adjusted to obtain the adjusted digital watermark content, which can be shown as Figure 10 . As shown in Figure 11 , taking luminance as an example, the local data of the adjusted digital watermark content is illustrated, such as the HSL color data of the pixel points in the red circle, which can be shown as Figure 8 (x=6, y=1, L=75) (x=5, y=2, L=60) (x=6, y=2, L=55) … (x=12, y=8, L=46). By analogy, the luminance data of the pixel points of the subject layer corresponding to all intersection regions and their corresponding coordinate data and HSL color data can be obtained.
[0074] Step 1403, the adjusted digital watermark content and the main layer are fused to obtain a second image. Exemplarily, as shown in the following. Figure 9
[0075] In the embodiments of the present application, before the above step 1402 is described in detail, the HSL color data is described in combination with Figure 9 .
[0076] As shown in the following, the medium corresponding to the HSL color is the human eye, which uses a way closer to human sensory intuition to describe the color, divides the color into hue, saturation, and brightness three factors, i.e., extends the "depth" concept of the human brain to saturation and brightness. H hue (chroma): the appearance and tone of the color, on the standard color wheel, the hue is measured by position, and the value is 0-360 degrees (black and white have no hue) S saturation: the purity of the color, the value is 0-100 degrees, the higher the saturation, the more brilliant the color, and the lower the saturation, the closer the color to gray, L brightness (lightness): the light and dark degree of the color. The value is 0-100 degrees. The higher the brightness, the brighter the color, and the lower the brightness, the dimmer the color. Figure 12 Based on this, the HSL color data includes data of at least one of the following channels: hue, saturation, and lightness. Based on this, the step 1402 can specifically include steps 14021 and 14022, and the specific embodiments are as shown in the following.
[0077] Step 14021, according to the first channel in the HSL color data of the pixel point corresponding to the intersection region, the segmented gain rule information corresponding to the first channel is obtained, and the first channel includes at least one of the following: hue, saturation, and lightness.
[0078] Step 14022, according to the segmented gain rule information corresponding to the first channel, the data of the second channel of the pixel point in the watermark content is adjusted to obtain the adjusted digital watermark content, and the first channel and the second channel are of the same channel type.
[0079] In the embodiments of the present application, the data of the second channel, i.e., the lightness of the pixel point in the watermark content, is taken as an example for description. It should be noted that the data of the pixel point in the watermark content includes but is not limited to: adjusting the saturation; adjusting the lightness; adjusting the hue and the saturation; adjusting the hue and the lightness; adjusting the saturation and the lightness, etc.
[0080]
[0081] The HSL color data adjustment is performed in the intersection region, and the lightness of each unit pixel in the intersection region is adjusted without changing the hue and saturation. Affected by the visual characteristics of the human eye to lightness, when adjusting, if the lightness value of the unit pixel is selected to be increased, the higher the lightness value of the unit pixel, the greater the increase amplitude; if the lightness value of the unit pixel is selected to be decreased, the higher the lightness value of the unit pixel, the smaller the decrease amplitude. That is, the first channel is also lightness.
[0082] Based on this, the segmentation gain rule information of lightness can include increasing the lightness value, and the lightness value is divided into five equal parts as an example for illustration.
[0083] When the lightness of the first channel is 0%-20%, the data of the second channel of the pixel point in the watermark content is increased by 2% per unit pixel each time; When the lightness of the first channel is 20%-40%, the data of the second channel of the pixel point in the watermark content is increased by 3% per unit pixel each time; When the lightness of the first channel is 40%-60%, the data of the second channel of the pixel point in the watermark content is increased by 4% per unit pixel each time; When the lightness of the first channel is 60%-80%, the data of the second channel of the pixel point in the watermark content is increased by 5% per unit pixel each time; When the lightness of the first channel is 80%-100%, the data of the second channel of the pixel point in the watermark content is increased by 6% per unit pixel each time.
[0084] Therefore, by adjusting the lightness of the main layer of the digital watermark region, the user is less likely to recognize the situation, the main layer is less damaged and the damage is more difficult, and the readability and security are guaranteed. The display processing of the digital watermark content is not limited to lightness, that is, the first channel can also be saturation and hue, both of which are also applicable. Therefore, the first channel can be one of the three, or any combination of the three.
[0085] In addition, in some embodiments of the present application, after step 140, the image processing method in the embodiments of the present application can further include steps 1501 to 1504, as shown below.
[0086] Step 1501, display the second image and the adjustment control, and the adjustment control is used to adjust the HSL color data of the digital watermark content in the second image.
[0087] Step 1502, receive the input of the user to the adjustment control.
[0088] Step 1503, in response to the input, adjust the HSL color data of the digital watermark content in the second image through the adjustment parameter corresponding to the input.
[0089] At step 1504, the adjusted second image is displayed.
[0090] As shown in the examples, the default adjustment scale is generated, and an adjustment control is provided to adjust the interface for the user to select to increase or decrease the brightness, adjust the brightness variation degree, and support local adjustment according to the selected area. Figure 8 For example, the default adjustment scale is 5 times, and the unit pixel increase scale is 1. In addition to the increase / decrease interaction form of the above adjustment control, the user interaction modes include but are not limited to dragging, scrolling, selecting, inputting, etc. The user interaction scenarios include but are not limited to Web, mobile, pad, etc.
[0091] Therefore, the final generation effect of the watermark can be artificially controlled to ensure the generation effect in different scenarios and improve the success rate of the watermark.
[0092] In summary, based on the image processing method provided in the embodiment, the watermark position can be calibrated, and the display of the watermark is processed by HSL color data, the digital watermark content and the main layer are merged, and the image after embedding the digital watermark content is generated. The effect diagram can be referred to as shown in Figure 13 The conventional watermark visual effect diagram is shown in Figure 14 .
[0093] Specifically, the first image can be learned by using a deep learning algorithm, and a watermark calibration model is used to capture decision-level information in the image space as a position parameter for watermark embedding. The Grad-CAM technology is used to visualize the bottom layer features to generate a heat map mask and visualize the importance of the content to calibrate the watermark embedding position. Then, the intersection area of the digital watermark content and the main layer is obtained, the brightness data of the main layer unit pixel corresponding to the intersection area and the corresponding coordinate data are obtained, the brightness value of the watermark and the main layer area is adjusted by performing HSL color data in the intersection area, the default adjustment scale is generated, and an adjustable interface is provided for the user to select to increase or decrease the brightness, adjust the brightness variation degree, and support local adjustment according to the selected area. The digital watermark content and the main layer are merged, and the watermark generation is completed.
[0094] Therefore, combined with the watermark calibration model, the image watermark embedding load is adaptively adjusted, the watermark perception can be reduced by combining the HSL color mode, and the influence of the watermark information expression on the original image is effectively balanced. The decision features of the image to be embedded are extracted by using the computer vision algorithm, and the Grad-CAM technology is used for visual expression to guide the watermark to be embedded, which can reduce the watermark production cost, make the watermark quality more objective and effective, and ensure the watermark identification under the condition of good concealment. Based on the HSL color data, the watermark is generated, and according to different color characteristics, the watermark does not affect the overall effect of the picture on the basis of visual visibility, so that the problem of the influence of the watermark on the image visual effect can be effectively solved, including the problem of reduced image visual experience and information acquisition obstacle. In addition, the user can adjust the default effect of the watermark through the interface, and the interaction is simple and convenient, and the experience is good. And the number of adjustments, the adjustment ratio, the adjustment form and the like are agreed by the user himself, and the agreed form can improve the adaptability to a certain extent, so that the usability of the watermark is higher. Moreover, the high integration of the image and the watermark can effectively resist attacks and reduce the risk of being cropped, wiped out, ignored, modified and edited. The embodiments of the present application do not need complex technical verification of the watermark, do not need additional cost, ensure safety while saving expenses.
[0095] It should be noted that the image key content can be calibrated by using the algorithm in the embodiments of the present application, and an image content confidence mask image can be generated to guide the watermark to be embedded, which is also applicable to other image classification, image segmentation, target recognition and the like. The way of generating the watermark by adjusting the brightness in the HSL color mode in the embodiments of the present application can be diversified into HSL three options, such as adjusting the saturation or brightness; or can be upgraded into a form of two-by-two combination of the three, adjusting the brightness and saturation of the watermark in two dimensions or other two-by-two combination forms to adapt to different scenes. The way of generating the watermark by adjusting the HSL color mode in the embodiments of the present application is also applicable to HSB, RGB, CMYK, LAb, HSL and the like. The adjustment control provided in the embodiments of the present application can allow the user to adjust in forms not limited to: full automatic generation (without user adjustment), user overall adjustment, user local selection adjustment and the like.
[0096] The image processing method in the embodiments of the present application can be widely applied to application scenarios requiring watermark to achieve copyright protection, prevent others from stealing pictures, anti-counterfeiting and the like, by reducing the influence of the watermark on the main layer, increasing the use scenarios of the watermark, promoting the development of information protection, content authentication and the like, without additional cost, without the need for complex technical identification of watermark content, and being able to guarantee the obtainability of watermark information while not affecting the main effect, providing favorable technical support for copyright protection, information tracking, anti-counterfeiting and anti-theft. In addition, the application scenarios in the embodiments of the present application include but are not limited to the watermark adding method for related images such as pictures, videos, APPs and webpages.
[0097] Based on the image processing method provided in the above embodiments, the application further provides a specific implementation of an image processing device. Please refer to the following embodiments.
[0098] Based on the same inventive concept, the application further provides an image processing device. The specific implementation of the image processing device will be described in detail in combination with Figure 14
[0099] Figure 14 A structural schematic diagram of a data processing device provided in the embodiments of the application.
[0100] In some embodiments of the application, Figure 14 The data processing device shown can be arranged in a computer device. As Figures 1 to 13 The data processing device 1400 can specifically include: An acquisition module 1401 configured to acquire a first image and digital watermark content; A model module 1402 configured to input the first image into a watermark calibration model, and acquire an image content confidence mask image of the first image output by the watermark calibration model through the watermark calibration model, the image content confidence mask image being an image used to represent the attention degree of the watermark calibration model to different regions in the first image; A determination module 1403 configured to determine a first position according to the image content confidence mask image, the first position being a position where the digital watermark content is embedded into the first image, and a value of the attention degree corresponding to the first position in the image content confidence mask image being greater than or equal to a preset value of the attention degree; A fusion module 1404 configured to fuse the digital watermark content and the first image based on the first position, and obtain a second image, the second image being an image after the first image is embedded with the digital watermark content.
[0101] Thus, the image processing apparatus in the embodiments of the present application acquires a first image and digital watermark content; inputs the first image into a watermark calibration model, acquires an image content confidence mask image of the first image output by the watermark calibration model through the watermark calibration model, and the image content confidence mask image is an image for representing the attention degree of the watermark calibration model to different regions in the first image; determines a first position according to the image content confidence mask image, the first position is a position where the digital watermark content is embedded in the first image, and a value of the attention degree corresponding to the first position in the image content confidence mask image is greater than or equal to a preset value of the attention degree; and fuses the digital watermark content and the first image based on the first position to obtain a second image, the second image being an image after the first image embeds the digital watermark content. In this way, the image content confidence mask image is generated through the watermark calibration model, and a region with stable structure and important semantics in the image, such as a part with rich texture and clear edge, is automatically selected as a watermark embedding position. The sensitivity of these regions to geometric transformations such as rotation, translation and scaling is relatively low, which can effectively resist the destruction of the attack on the watermark-image space relationship, avoid the robustness defects caused by traditional manual point selection or template fixed position, and enhance the anti-attack ability. The first position is dynamically determined according to the attention degree value in the mask image, which can ensure that the watermark is embedded in a region with high image content confidence. This adaptive mechanism avoids the limitations of the preset template, optimizes the watermark position according to the image content, reduces the detection failure or extraction error caused by improper position, and dynamically optimizes the embedding position. When the second image is generated by fusion, the watermark is embedded in the high-confidence region, which enhances the combination strength of the watermark and the original image, making it difficult for attacks to destroy these key regions, thereby maintaining the integrity and detectability of the watermark, significantly reducing the extraction error rate, and improving the watermark detection reliability. In this way, the watermark calibration model is used to replace manual point selection to realize automatic calibration of the embedding position, improve the efficiency and reduce human bias, and the image content confidence mask image is generated by analyzing the image content features such as texture and edge through the watermark calibration model, which ensures that the embedding position is naturally integrated with the image structure, further consolidates the anti-attack performance, dynamically selects the high-robustness embedding position, and effectively solves the problem of watermark detection failure or extraction error in the watermark image caused by the vulnerability of traditional methods, thereby improving the watermark robustness.
[0102] The image processing apparatus 140 in the embodiments of the present application will be described in detail below.
[0103] In one or more optional embodiments, the model module 1402 can be specifically configured to, in the case where the watermark calibration model includes a spatial attention mechanism layer and a spatial attention mechanism strengthening layer, input the first image into the watermark calibration model, capture decision-level position information of the first image through the spatial attention mechanism layer; The image content confidence mask image of the first image output by the spatial attention mechanism reinforcement layer is obtained. The image content confidence mask image of the first image output by the spatial attention mechanism reinforcement layer is obtained.
[0104] In one or more optional embodiments, the model module 1402 can be specifically configured to input the first image into a watermark calibration model, encode the first image along a horizontal coordinate direction and a vertical coordinate direction through a first-size pooling kernel and a second-size pooling kernel in the spatial attention mechanism layer respectively, and obtain a first perceptual attention feature in the horizontal coordinate direction and a second perceptual attention feature in the vertical coordinate direction. The first perceptual attention feature and the second perceptual attention feature are channel-cascaded through convolution in the spatial attention mechanism layer to obtain a comprehensive perceptual attention feature map. The comprehensive perceptual attention feature map is decomposed along a spatial dimension to obtain a first tensor and a second tensor, the first tensor being used to represent a tensor subjected to global pooling or feature encoding in a height direction, and the second tensor being used to represent a tensor subjected to global pooling or feature encoding in a width direction. The first tensor and the second tensor are respectively subjected to convolution transformation and activation function processing to obtain decision-level position information of the first image.
[0105] In one or more optional embodiments, the model module 1402 can be specifically configured to generate a vector set including a key vector, a value vector and a query vector through the spatial attention mechanism reinforcement layer according to the decision-level position information of the first image. Each vector in the vector set is subjected to deformation processing to obtain a deformed vector set including a deformed key vector of the key vector, a deformed value vector of the value vector and a deformed query vector of the query vector. A similarity matrix is constructed based on the deformed key vector and the deformed value vector, the similarity matrix being used to represent an association degree of at least two pixel points in the first image. The deformed query vector is transformed through the similarity matrix to obtain a first position feature, the first position feature being used to represent a relative position relationship between the at least two pixel points in the first image. The first position feature is reshaped in a spatial dimension to obtain a second position feature, the second position feature being used to represent a spatial position relationship between the at least two pixel points in the first image. The first image is subjected to feature enhancement through the second position feature to obtain an image content confidence mask image.
[0106] In one or more optional embodiments, the determining module 1403 can be specifically configured to, in a case where the image content confidence mask image includes N regions and a value of the attention degree of each of the N regions is different, determine the value of the attention degree of each of the N regions according to the value of the attention degree of the pixel point in each of the N regions. determine a position corresponding to a target region with the highest value of the attention degree in the N regions as the first position.
[0107] In one or more optional embodiments, the fusing module 1404 can be specifically configured to, based on the first position, superimpose the digital watermark content on the subject layer of the first image to obtain an intersection region of the subject layer and the digital watermark content. adjust the HSL color data of the pixel point in the watermark content according to the HSL color data of the pixel point corresponding to the intersection region to obtain adjusted digital watermark content. fuse the adjusted digital watermark content and the subject layer to obtain a second image.
[0108] In one or more optional embodiments, the determining module 1403 can be specifically configured to, in a case where the HSL color data includes data of at least one of the following channels: hue, saturation, and lightness, obtain, according to a first channel in the HSL color data of the pixel point corresponding to the intersection region, segment gain rule information corresponding to the first channel, the first channel including at least one of the following: hue, saturation, and lightness. adjust data of a second channel of the pixel point in the watermark content according to the segment gain rule information corresponding to the first channel to obtain adjusted digital watermark content, the first channel and the second channel being of the same type.
[0109] In one or more optional embodiments, the data processing apparatus 1400 in the embodiments of the present application can further include a display module configured to display the second image and an adjustment control, the adjustment control being configured to adjust the HSL color data of the digital watermark content in the second image. The data processing apparatus 1400 in the embodiments of the present application can further include a receiving module configured to receive an input of the adjustment control by a user. The data processing apparatus 1400 in the embodiments of the present application can further include an adjusting module configured to, in response to the input, adjust the HSL color data of the digital watermark content in the second image by an adjusting parameter corresponding to the input. The display module can be further configured to display the adjusted second image.
[0110] The various modules of the image processing apparatus 1400 provided in the embodiments of the present application can realize Figure 15 the functions of the various steps of the image processing method provided in the embodiments of the present application and achieve the corresponding technical effects. For brevity, the functions and technical effects will not be described here.
[0111] Figure 15 A hardware structure diagram of a computer device provided by some embodiments of the present application is shown.
[0112] As shown in Figure 15 , the computer device can include a processor 1501 and a memory 1502 storing computer program instructions.
[0113] In particular, the processor 1501 can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement one or more embodiments of the present application.
[0114] The memory 1502 can include a mass storage for data or instructions. By way of example and not limitation, the memory 1502 can include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. The memory 1502 can include removable or non-removable (or fixed) media, where appropriate. The memory 1502 can be internal or external to the integrated gateway disaster recovery device, as appropriate. In particular embodiments, the memory 1502 is non-volatile, solid-state memory.
[0115] In particular embodiments, the memory 1502 can include read-only memory (ROM), random access memory (RAM), a disk storage medium device, an optical storage medium device, a flash memory device, electrical, optical, or other physical / tangible memory storage devices. Thus, in general, the memory 1502 includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software that, when executed (by one or more processors), is operable to perform operations described with reference to the image processing methods according to the above-described embodiments of the present application.
[0116] The processor 1501 implements any of the image processing methods in the above-described embodiments by reading and executing the computer program instructions stored in the memory 1502.
[0117] In one example, the computer device can further include an image processing interface 1503 and a bus 1510. As shown in Figures 1 to 14 , the processor 1501, the memory 1502, and the image processing interface 1503 are connected by the bus 1510 and complete image processing among each other.
[0118] The image processing interface 1503 is mainly used to implement image processing between modules, devices, units and / or equipment in the embodiments of the present application.
[0119] The bus 1510 includes hardware, software, or both, that couples components of the computer device to each other. By way of example, and not limitation, a bus can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand (IB) interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus or a combination of two or more of these. Where appropriate, bus 1510 can include one or more buses. Although the present application is described and illustrated with a particular bus, the present application contemplates any suitable bus or interconnect.
[0120] The computer device can execute the image processing method in the embodiments of the present application, thereby realizing the image processing method and device described in combination with the above embodiments.
[0121] In addition, in combination with the image processing method in the above embodiments, the embodiments of the present application can provide a computer readable storage medium to implement. The computer readable storage medium has computer program instructions stored thereon; the computer program instructions are executed by a processor to implement any one of the image processing methods in the above embodiments. Examples of the computer readable storage medium include non-transitory computer readable storage media, such as a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, etc.
[0122] In addition, in combination with the image processing method in the above embodiments, the embodiments of the present application can provide a computer program product to implement. The program product is stored in a storage medium, and specifically can include a computer program or instructions, which are executed by a processor to implement any one of the image processing methods in the above embodiments. The program product is executed by at least one processor to implement various processes of the above image processing method embodiments, and can achieve the same technical effects. To avoid repetition, details are not described here.
[0123] It is to be understood that the present application is not limited to particular configurations and processes described herein and as illustrated in the drawings. For the sake of brevity and clarity, detailed descriptions of well known methods and apparatuses will not be repeated herein. In the above embodiments, several specific steps are described and illustrated as examples. However, the methods process of the present application is not limited to the specific steps described and illustrated, and one of ordinary skill in the art can make various changes, modifications, and additions, or can change the order of steps, after having the benefit of this description.
[0124] The functional blocks shown in the structural block diagrams above can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, and the like. When implemented in software, the elements of the present application are program or code segments that are used to perform the required tasks. The program or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or a picture processing link. The "machine-readable medium" can include any medium that can store or transmit information. Examples of the machine-readable medium include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, and the like. The code segments can be downloaded via a computer network such as the Internet, an intranet, and the like.
[0125] It is also to be understood that the example embodiments described in this application are based on a series of steps or apparatuses to describe some methods or systems. However, the present application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiments, or in an order different from the embodiments, or several steps can be performed simultaneously.
[0126] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer program instructions can also be stored in a computer readable medium that can direct a computer, other programmable data processing apparatus, or other processing devices to operate in a particular manner, such that the instructions stored in the computer readable medium produce an article of manufacture including instructions which implement the function / act specified in the flowchart and / or block diagram block or blocks.
[0127] The above is merely specific implementation of the present application, and those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein again. It should be understood that the protection scope of the present application is not limited to this, and any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed in the present application, and these modifications or replacements should be covered within the protection scope of the present application.
Claims
1. An image processing method, characterized in that, include: Obtain the first image and digital watermark content; The first image is input into the watermark calibration model. The image content confidence mask image of the first image is obtained through the watermark calibration model. The image content confidence mask image is used to characterize the degree of attention of the watermark calibration model to different regions in the first image. Based on the image content confidence mask image, a first position is determined. The first position is the position where the digital watermark content is embedded in the first image. The attention level value corresponding to the first position in the image content confidence mask image is greater than or equal to the preset attention level value. Based on the first position, the embedded digital watermark content and the first image are fused to obtain a second image, which is the image after the digital watermark content is embedded in the first image.
2. The image processing method according to claim 1, characterized in that, The watermarking model includes a spatial attention mechanism layer and a spatial attention mechanism enhancement layer; the step of inputting the first image into the watermarking model and obtaining the image content confidence mask image of the first image output by the watermarking model includes: The first image is input into the watermarking model, and the decision-level location information of the first image is captured through the spatial attention mechanism layer; The spatial attention mechanism enhancement layer is used to mine the global spatial dependency relationship of pixels in the first image based on the decision-level location information of the first image, and the image content confidence mask image is obtained. Obtain the image content confidence mask image of the first image output by the spatial attention mechanism enhancement layer.
3. The image processing method according to claim 2, characterized in that, The step of inputting the first image into the watermarking model and capturing the decision-level location information of the first image through the spatial attention mechanism layer includes: The first image is input into the watermarking model. The first image is encoded by the pooling kernel of the first size and the pooling kernel of the second size in the spatial attention mechanism layer along the horizontal coordinate direction and the vertical coordinate direction, respectively, to obtain the first perceptual attention feature in the horizontal coordinate direction and the second perceptual attention feature in the vertical coordinate direction. By fusing cross-channel relationships through convolution in the spatial attention mechanism layer, the first perceptual attention feature and the second perceptual attention feature are channel-concatenated to obtain a comprehensive perceptual attention feature map; Along the spatial dimension, the comprehensive perceptual attention feature map is decomposed to obtain a first tensor and a second tensor. The first tensor is used to represent the tensor that performs global pooling or feature encoding in the height direction, and the second tensor is used to represent the tensor that performs global pooling or feature encoding in the width direction. The first tensor and the second tensor are subjected to convolution transformation and activation function processing respectively to obtain the decision-level position information of the first image.
4. The image processing method according to claim 2 or 3, characterized in that, The step of using the spatial attention mechanism enhancement layer to mine the global spatial dependencies of pixels in the first image based on the decision-level location information of the first image to obtain the image content confidence mask image includes: Through the spatial attention mechanism enhancement layer, a vector set is generated based on the decision-level location information of the first image. The vector set includes a key vector, a value vector, and a query vector. Each vector in the vector set is transformed to obtain a transformed vector set, which includes the transformed key vector of the key vector, the transformed value vector of the value vector, and the transformed query vector of the query vector. Based on the deformed key vector and the deformed value vector, a similarity matrix is constructed, which is used to characterize the degree of association between at least two pixels in the first image; The deformation query vector is transformed by the similarity matrix to obtain a first position feature, which is used to characterize the relative positional relationship between at least two pixels in the first image. The first positional feature is reshaped in the spatial dimension to obtain a second positional feature, which is used to characterize the spatial positional relationship between at least two pixels in the first image. The first image is enhanced using the second location feature to obtain the image content confidence mask image.
5. The image processing method according to claim 1, characterized in that, The image content confidence mask image includes N regions, each of which has a different attention level value; determining the first position based on the image content confidence mask image includes: The attention level value of each region is determined based on the attention level value of the pixels in each region; The location corresponding to the target region with the highest attention value among the N regions is determined as the first location.
6. The image processing method according to claim 1, characterized in that, The step of fusing the digital watermark content embedding and the first image based on the first position to obtain the second image includes: Based on the first position, the digital watermark content is superimposed on the main body layer of the first image to obtain the intersecting area of the main body layer with the digital watermark content; Based on the HSL color data of the pixels corresponding to the intersecting region, the HSL color data of the pixels in the watermark content is adjusted to obtain the adjusted digital watermark content. The adjusted digital watermark content is then fused with the main layer to obtain the second image.
7. The image processing method according to claim 6, characterized in that, The HSL color data includes data for at least one of the following channels: hue, saturation, and lightness; The step of adjusting the HSL color data of pixels in the watermark content based on the HSL color data of the pixels corresponding to the intersecting region to obtain the adjusted digital watermark content includes: Based on the first channel in the HSL color data of the pixels corresponding to the intersecting region, obtain the segmented gain rule information corresponding to the first channel, wherein the first channel includes at least one of the following: hue, saturation, and brightness. According to the segmented gain rule information corresponding to the first channel, the data of the second channel of the pixels in the watermark content is adjusted to obtain the adjusted digital watermark content. The first channel and the second channel have the same channel type.
8. The image processing method according to claim 6 or 7, characterized in that, The method further includes: Display the second image and adjustment controls, wherein the adjustment controls are HSL color data for adjusting the digital watermark content in the second image; Receive user input for the adjustment controls; In response to the input, the HSL color data of the digital watermark content in the second image is adjusted by the adjustment parameters corresponding to the input; The adjusted second image is displayed.
9. An image processing apparatus, characterized in that, include: The acquisition module is used to acquire the first image and the digital watermark content; The model module is used to input the first image into the watermark calibration model, and obtain the image content confidence mask image of the first image output by the watermark calibration model through the watermark calibration model. The image content confidence mask image is used to characterize the degree of attention of the watermark calibration model to different regions in the first image. The determining module is used to determine a first position based on the image content confidence mask image, wherein the first position is the position where the digital watermark content is embedded in the first image, and the attention level value corresponding to the first position in the image content confidence mask image is greater than or equal to a preset attention level value. The fusion module is used to fuse the embedded digital watermark content and the first image based on the first position to obtain a second image, wherein the second image is the image after the digital watermark content is embedded in the first image.
10. A computer device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the image processing method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the image processing method as described in any one of claims 1-8.
12. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device is able to perform the image processing method as described in any one of claims 1-8.