Character image ugly detection method based on image-text multi-mode

Through the multi-modal detection method of graphic and text, combined with image and text features, the problem of modal singleness in the existing technology is solved, and high-precision detection and specific content recognition of character images are achieved.

CN120279397AInactive Publication Date: 2025-07-08HANGZHOU ZHONGKE RUIJIAN TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510758545.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology has a single modality in character ugliness detection, making it difficult to identify specific ugliness content, and the model is not robust and accurate, especially when facing new ugliness styles and types, it cannot be effectively detected.

Method used

The ugliness detection method based on multimodal graphics and text is adopted, and the image-side model and text-side model are combined, and the multimodal graphics and text understanding basic model and OCR feature extractor are used to construct an ugliness discriminator to conduct multi-source fusion feature judgment.

Benefits of technology

It realizes high-precision detection of character ugliness, can identify specific ugly content, and transfer learning to improve detection accuracy under small sample data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279397A_ABST
    Figure CN120279397A_ABST
Patent Text Reader

Abstract

The invention relates to a character ugly detection method based on image-text multi-modality, and is suitable for the technical field of image detection and recognition. The method comprises the following steps: inputting a to-be-detected image-text into a trained ugly discriminator to obtain a character image ugly detection result of the to-be-detected image-text; the ugly discriminator comprises an image side model used for extracting image semantic features of an image in an image-text; the text OCR feature extractor is used for identifying a text in the image and extracting text semantic features of the text; the text side model is used for extracting a text vector feature of a text in the image-text; and the ugly discrimination network is used for outputting a character ugly detection result based on multi-source fusion features, and the multi-source fusion features are formed based on the image semantic features extracted by the image side model and the text semantic features extracted by the text OCR feature extractor. According to the invention, the high-precision judgment of the ugly result of the figure image is realized, and the specific content description of the ugly can be identified at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting the uglification of human images based on multi-modal graphics and text, which is applicable to the technical field of image detection and recognition. Background Art

[0002] With the rapid development of the Internet, the proportion of pictures in the traffic data at home and abroad is increasing, and image and video editing tools are becoming more and more advanced and easy to use, resulting in the current uglification of human images not only being limited to text, but also the spread of uglified images being more and more extensive. Uglified images are often related to human attributes, and the uglified content can be direct or hidden, and it not only includes traditional harmful elements, but also artificial tampering and AIGC generation means that are not easy to detect. The pictures, videos, texts, and voices of human images obtained through artificial production, synthesis, deep forgery and other technologies have the characteristics of uglification, distortion, falsity, etc., with a certain bad purpose, and are likely to cause serious adverse effects to society and individuals.

[0003] Currently, there is little research on the detection of human image uglification. Most of the mainstream methods are based on deep learning, which are often limited to image features and have a single modality. After using a single or multiple neural networks to extract image-level features and then performing classification processing, such methods are simple in design and can only simply judge whether it is an uglified picture, unable to obtain specific uglified content, categories and other information, lacking an understanding of the image content, resulting in insufficient robustness and accuracy of the model. Whenever a new uglification style and type appear, they often cannot be detected. And such methods rely on a large amount of data training, and it is difficult to collect a large amount of uglified data in reality. Most traditional methods use artificial feature design or establish a template library for image comparison, retrieval, etc., which have great limitations and extremely poor generalization ability. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: in view of the above problems, to provide a method for detecting the uglification of human images based on multi-modal graphics and text.

[0005] The technical solution adopted by the present invention is: a method for detecting the uglification of human images based on multi-modal graphics and text, including: Inputting the text and picture to be detected into a trained uglification discriminator to obtain the detection result of the uglification of the human image in the text and picture to be detected; The uglification discriminator includes: An image-side model, taken from a trained multi-modal graphics and text understanding base model, for extracting the image semantic features of the image in the text and picture; A text OCR feature extractor, for recognizing the text in the image and extracting the text semantic features of the text; The text-side model, taken from a trained multi-modal image-text understanding base model, is used to extract the text vector features of the text in the image and text. The uglification discrimination network is used to output the detection result of the uglification of the character image based on the multi-source fusion features, where the multi-source fusion features are formed based on the image semantic features extracted by the image-side model and the text semantic features extracted by the text OCR feature extractor. The multi-modal image-text understanding base model is constructed based on contrastive learning, and includes an image-side model for extracting image semantic features and a text-side model for extracting text vector features. Samples with consistent image semantics and text representations are used as positive samples, and samples with inconsistent image semantics and text representations are used as negative samples.

[0006] The image-side model includes: The first encoder, constructed based on Transformer, is used to extract the first features of the image. The second encoder, constructed based on CNN, is used to extract the second features of the image. The image semantic features of the image are formed based on the first features and the second features of the image.

[0007] The training of the multi-modal image-text understanding base model includes: In the first training stage, the image-side model and the text-side model are initialized using the image pre-training model and the text pre-training model respectively. Subsequently, the parameters of the image-side model are frozen, and the text-side model is associated with the existing image pre-training representation space. In the second training stage, the parameters of the image-side model are unfrozen, and while the parameters of the image-side model and the text-side model are associated, the data distribution is modeled.

[0008] The training of the multi-modal image-text understanding base model adopts contrastive loss: ; ; Among them, image_features are the features encoded by the image-side model, text_features are the features encoded by the text-side model, and norm represents the L2 norm. ; ; Among them, scale is the scaling factor during training. 、 is the cosine similarity matrix. ; ; Among them, For the label, loss is the cross - entropy loss operator, which is specifically expressed as: ; The loss of the defamed image - text pair is the combined loss of the text side and the image side, as follows: 。

[0009] The training of the defamed discriminator includes: Based on a small - sample defamed dataset, multi - branch joint training is performed on the multi - modal image - text understanding base model and the defamed discriminator.

[0010] The loss of the multi - branch joint training consists of the discriminator Loss and the loss of the defamed image - text pair, where the discriminator Loss uses cross - entropy loss.

[0011] The defamed discriminator further includes: A human feature analysis module for determining whether there is a human face in the image based on a face detection method.

[0012] A storage medium stores a computer program executable by a processor. When the computer program is executed, the steps of the method for detecting the defaming of a human image based on image - text multi - modality are implemented.

[0013] A device for detecting the defaming of a human image has a memory and a processor. The memory stores a computer program executable by the processor. When the computer program is executed, the steps of the method for detecting the defaming of a human image based on image - text multi - modality are implemented.

[0014] The beneficial effects of the present invention are as follows: The present invention adopts a multi - modal image - text understanding base model based on contrastive learning to establish a connection between vision and text, so that the image - side model can extract semantic features associated with the text content from the image; the text in the image is recognized by a text OCR feature extractor, and the text semantic features of the text are extracted; the present invention forms multi - source fusion features based on the semantic features and the text semantic features, and then determines the detection result of the defaming of the human image based on the multi - source fusion features.

[0015] In the process of identifying the defamed human image of the present invention, by using the basic model of image - text multi - modality trained with a large amount of data, selecting the image - side model of the multi - modal image - text understanding base model, and introducing the text OCR feature extraction ability, a defamed discriminator is constructed, and transfer learning is performed on the defamed samples through multi - task joint training, realizing high - precision judgment of the defaming result of the human image while being able to identify the specific content description of the defaming. Description of the Drawings

[0016] Figure 1It is the flow chart of the training and inference of the detection of the uglification of the human image based on the multi-modal text and image in the embodiment.

[0017] Figure 2 It is the model framework of the multi-modal text and image understanding basic model in the embodiment.

[0018] Figure 3 It is the schematic diagram of the multi-branch joint training in the embodiment. Detailed implementation manners

[0019] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of description and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0020] In the description of the present invention, the meaning of "a plurality of" is two or more. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the sequence of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention.

[0021] Embodiment 1: As Figure 1 shown, this embodiment is a method for detecting the uglification of the human image based on the multi-modal text and image, and specifically includes the following steps: S100. Obtain the image to be detected; S200. Input the image to be detected into the trained uglification discriminator to obtain the detection result of the uglification of the human image in the image.

[0022] In this embodiment, the uglification discriminator includes a human feature analysis module, an image-side model, a text-side model, a text OCR feature extractor, an uglification discrimination network, etc.

[0023] In this example, the human feature analysis module is used to determine whether there is a human face based on the face detection method. The uglification of the human image is usually related to the attributes of the person himself. Combining the face attributes gives an overall feature analysis of the person, and the result can be used to guide the subsequent modules.

[0024] In this embodiment, the image-side model is taken from the trained multi-modal text and image understanding basic model and is used to extract the image semantic features of the image .

[0025] In this example, the multi-modal graphic and text understanding base model is constructed based on contrastive learning, including an image-side model for extracting image semantic features and a text-side model for extracting text vector features. Samples with consistent image semantics and text representations are used as positive samples, and samples with inconsistent image semantics and text representations are used as negative samples to establish connections between vision and text.

[0026] In this example, the image-side model includes: a first encoder and a second encoder, where the first encoder is constructed based on Transformer and is used to extract the first features of the image; the second encoder is constructed based on CNN and is used to extract the second features of the image.

[0027] Current multi-modal models mainly use the Transformer structure as the backbone network. Common ones include the vit series. The self-attention in Transformer can capture long-term dependencies (i.e., global context), but it is very likely to ignore the structural information and local relationships in each image patch. Relatively speaking, CNN can pay more attention to local information and utilize the local connectivity in translational invariance, making CNN more dependent on texture when classifying visual objects. The convolutional and self-attention mechanisms have a complementary effect.

[0028] In this embodiment, the text-side model of the multi-modal graphic and text understanding base model selects the classic Bert structure to extract text vector features.

[0029] In this embodiment, the multi-modal graphic and text understanding base model uses a two-stage training process, including: In the first training stage, the model uses existing image pre-training models and text pre-training models to initialize the image-side model and the text-side model respectively. Subsequently, the parameters of the image-side model are frozen, allowing the text-side model to be associated with the existing image pre-training representation space while reducing the training cost; In the second training stage, the parameters of the image-side model are unfrozen, allowing the image-side model and the text-side model to be associated while modeling the data distribution. In this process, a large amount of graphic and text data is used to obtain a powerful basic representation model. The framework of the multi-modal graphic and text understanding base model is as Figure 2 shown.

[0030] In the training of the multi-modal graphic and text understanding base model in this embodiment, a contrastive loss is adopted: ; ; where, image_features is the feature map encoded by the image-side model, text_features is the feature map encoded by the text-side model, and norm represents the L2 norm; ; ; where scale is the scaling factor during the training process and is the cosine similarity matrix; ; ; where is the label, loss is the cross-entropy loss operator, which is different from the conventional cross-entropy calculation method, and is specifically expressed as: ; The finally trained and optimized loss is the combined loss of the text side and the image side (ugly text-image pair loss), as follows: .

[0031] The text OCR feature extractor in this embodiment is used to recognize text in an image and extract the text semantic features of the text . The text OCR feature extractor is divided into two parts. The ugly text detection part is based on the DBnet structure and extracts text position information from the image; the recognition model is based on the classic CRNN structure and is used to recognize text content from the text position and extract specific text semantic features .

[0032] The text OCR feature extractor in this example has been trained based on the OCR dataset, trained using the supervised CTC loss on the open-source dataset, and will be frozen in the subsequent steps and not participate in the training

[0033] In this embodiment, the ugly discrimination network is used to judge and output the ugly detection result of the character image based on the multi-source fusion feature , where the multi-source fusion feature is based on the overall semantic feature extracted by the image-side model and the text semantic feature extracted by the text OCR feature extractor to form

[0034] ;

[0035] In this embodiment, for the training of the uglification discriminator, transfer learning of the uglification sample dataset with multi-branch joint training is adopted. Based on the multi-modal graphic-text understanding base model that has been trained, and based on the small-sample uglification dataset, multi-branch joint training is performed on the multi-modal graphic-text understanding base model and the uglification discriminator. The purpose is to improve the uglification discrimination accuracy after joint multi-task training, and after transfer learning with small-sample uglification data, the text-side model of the graphic-text understanding base model can synchronously output the specific uglification content. Specifically, as Figure 3 , the Loss of multi-branch joint training consists of the discriminator Loss and the uglification graphic-text pair loss. The latter is the uglification graphic-text loss, which is calculated in the same way as the multi-modal graphic-text understanding base model. The former uses the cross-entropy loss and is specifically expressed as follows: ; Among them, is the true training label, is the probability that the model predicts the image as an uglified image.

[0036] The overall loss function of multi-branch joint training: .

[0037] In this embodiment, it is determined whether there is uglification through the uglification discriminator. If there is uglification, the uglification content is generated based on the text vector features extracted by the text-side model.

[0038] Embodiment 2: This embodiment is a storage medium, on which a computer program executable by a processor is stored. When the computer program is executed, the steps of the method for detecting the uglification of a character image based on graphic-text multi-modal in Embodiment 1 are implemented.

[0039] Embodiment 3: This embodiment is a device for detecting the uglification of a character image, which has a memory and a processor. A computer program executable by the processor is stored on the memory. When the computer program is executed, the steps of the method for detecting the uglification of a character image based on graphic-text multi-modal in Embodiment 1 are implemented.

[0040] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0041] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0042] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0043] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the above programs can be printed, because the above programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing them in a computer memory.

[0044] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0045] In the above description of this specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0046] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0047] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A method for detecting the uglification of a character image based on text and image multi-modalities, characterized in that, Including: Input the text and image to be detected into the trained uglification discriminator to obtain the detection result of the uglification of the human image in the text and image to be detected; The uglification discriminator includes: An image-side model, taken from a trained multi-modal text and image understanding base model, for extracting the image semantic features of the image in the text and image; A text OCR feature extractor, for recognizing the text in the image and extracting the text semantic features of the text; A text-side model, taken from a trained multi-modal text and image understanding base model, for extracting the text vector features of the text in the text and image; An uglification discrimination network, for outputting the detection result of the uglification of the human image based on the multi-source fusion features, where the multi-source fusion features are formed based on the image semantic features extracted by the image-side model and the text semantic features extracted by the text OCR feature extractor; The multi-modal text and image understanding base model is constructed based on contrastive learning, and includes an image-side model for extracting image semantic features and a text-side model for extracting text vector features. Samples with consistent image semantics and text representations are used as positive samples, and samples with inconsistent image semantics and text representations are used as negative samples.

2. The method for detecting the uglification of a character image based on graphic and text multi-modalities according to claim 1, characterized in that, The image-side model includes: A first encoder, constructed based on Transformer, for extracting the first features of the image; A second encoder, constructed based on CNN, for extracting the second features of the image; The image semantic features of the image are formed based on the first features and the second features of the image.

3. The method for detecting uglification of human images based on multi-modality of images and texts according to claim 1, characterized in that: The training of the multi-modal text and image understanding base model includes: In the first training stage, use the image pre-training model and the text pre-training model to initialize the image-side model and the text-side model respectively. Subsequently, freeze the parameters of the image-side model and let the text-side model be associated with the existing image pre-training representation space; In the second training stage, unfreeze the parameters of the image-side model, and while associating the parameters of the image-side model and the text-side model, model the data distribution.

4. The method for detecting the uglification of a character image based on graphic and text multi-modalities according to claim 1, wherein, The training of the multi-modal text and image understanding base model adopts a contrastive loss: ; ; where image_features are the features encoded by the image-side model, text_features are the features encoded by the text-side model, and norm represents the L2 norm; ; ; where scale is the scaling factor during the training process, 、 is the cosine similarity matrix; ; ; Among them, is a label, and loss is a cross-entropy loss operator, which is specifically expressed as: ; The finally trained and optimized loss is the combined loss of the text side and the image side, as follows: 。 5. The method for detecting the uglification of a character image based on graphic and text multi-modalities according to claim 1, characterized in that, The training of the uglification discriminator includes: Based on a small-sample uglification data set, perform multi-branch joint training on the multi-modal text and image understanding base model and the uglification discriminator.

6. The method for detecting the uglification of a character image based on graphic and text multi-modalities according to claim 5, wherein The loss of the multi-branch joint training consists of the discriminator Loss and the uglification text and image pair loss, where the discriminator Loss uses the cross-entropy loss.

7. The method for detecting the uglification of a character image based on text and image multi-modalities according to claim 1, wherein, The uglification discriminator further includes: A human feature analysis module, for determining whether there is a human face in the image based on a face detection method.

8. A storage medium having stored thereon a computer program executable by a processor, characterized in that, When the computer program is executed, it implements the steps of the method for detecting the uglification of a human image based on text and image multi-modalities according to any one of claims 1 to 7.

9. A character image uglification detection device, having a memory and a processor, wherein a computer program capable of being executed by the processor is stored on the memory, characterized in that When the computer program is executed, it implements the steps of the method for detecting the uglification of a human image based on text and image multi-modalities according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Convolution and self-attention-based shielded pedestrian re-identification method

    CN115909404A

  • False information identification model training method and device, equipment and storage medium

    CN118152812A

  • Image description generation method and device, computer equipment and storage medium

    CN119206723A

  • Multi-modal image-text tampering detection and positioning method based on feature enhancement

    CN119513743A