Apparatus, method, and recording medium for detecting irrelevance between news image and text
The news image-text non-relevance detection device uses unsupervised learning to construct datasets with positive and negative samples, including counterfactual text, effectively addressing the challenge of detecting irrelevant images in news content and improving news quality.
Patent Information
- Application Number
- PCT/KR2025/099346
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-23
- Filing Date
- 2025-02-07
- Publication Date
- 2025-09-04
AI Technical Summary
Existing techniques struggle to detect image-text inconsistencies in news content, particularly due to the lack of effective methods for learning image-text representations and the high cost and limited generalizability of supervised learning, leading to the spread of misleading information through unrelated images.
A news image-text non-relevance detection device and method utilizing unsupervised learning-based contrastive learning techniques, constructing datasets with positive and negative samples, including counterfactual text, to effectively represent visual-linguistic meanings and detect images with low relevance to text.
Reduces labeling costs and improves the quality of news articles by accurately identifying images that do not represent the text, addressing issues like clickbait and fake news.
Smart Images

Figure KR2025099346_04092025_PF_FP_ABST
Abstract
Description
News image-text non-relevance detection device, method, and recording medium
[0001] The present invention relates to a news image-text non-relevance detection device, method, and recording medium for detecting images with low relevance to text in news content.
[0002] In the online environment, much information is shared based on fragmented information like headlines and images. If this fragmented information doesn't represent the main content, those who haven't read the original text may misinterpret the content and spread false information, creating a social problem. Clickbait articles that use sensational content different from the main text are a prime example.
[0003] Furthermore, in the digital environment, news links are often shared based on headlines, summaries, and images, making it difficult to determine the authenticity of a news story without clicking on the link. This vulnerability is often exploited by using images unrelated to the content to attract attention or support one's opinion. Fake news, which spreads false information, is a prime example.
[0004] In particular, because visual information, among fragments of information, leaves a stronger and more lasting impression than text, sharing images that lack relevance to the main text can have a greater negative impact on users than other types of information. Therefore, images that effectively represent the content of the text should be used.
[0005] However, techniques for detecting text-type inconsistencies based on deep learning models have been mainly studied, and techniques for detecting inconsistencies that utilize other modalities, such as image-text, are lacking. In addition, when using supervised learning models for detection, they require hand-labeled data, which is time-consuming and costly.
[0006] Recently, techniques for learning image-text semantic representations, such as CLIP (Contrastive Language-Information Pretraining), have been proposed. However, conventional techniques only learn general image-caption pairs, making it difficult to understand non-explicit representations between images and text.
[0007] Therefore, to understand the subtle semantic differences contained in news text and the relevance of images and text, research is needed on learning methods and detection techniques that enable effective image-text representations based on counterfactual information that contradicts the facts covered in the news. Furthermore, because learning methods based on hand-crafted labels are expensive and have limited generalizability, research based on unsupervised learning, which enables learning from unlabeled data, is essential.
[0008] [Prior Art Literature]
[0009] [Patent Document]
[0010] Korean Patent Publication No. 10-2024-0006314
[0011] The present invention has been devised to solve the above problems, and an object of the present invention is to provide a news image-text non-relevance detection device, method, and recording medium.
[0012] According to one embodiment of the present invention for achieving the above object, a news image-text non-relevance detection device includes a pre-learning unit that constructs a dataset by extracting image-text pairs constituting news content as positive samples and extracting counterfactual text generated based on the text of the image-text pairs as negative samples; an embedding unit that, when news content is input, extracts a vector corresponding to an image of the news content and a vector corresponding to a text through a visual-language embedding model that is pre-trained based on the constructed dataset; and a detection unit that detects an image having low relevance to the text based on a similarity between a vector corresponding to the image and a vector corresponding to the text.
[0013] In addition, a news image-text non-relevance detection method according to one embodiment of the present invention for achieving the above object includes the steps of: constructing a dataset by extracting image-text pairs constituting news content as positive samples and extracting counterfactual text generated based on the text of the image-text pairs as negative samples; when news content is input, extracting a vector corresponding to an image of the news content and a vector corresponding to a text through a visual-language embedding model pre-trained based on the constructed dataset; and detecting an image having low relevance to the text based on the similarity between the vector corresponding to the image and the vector corresponding to the text.
[0014] In addition, a recording medium according to one embodiment of the present invention for achieving the above purpose is a computer-readable recording medium having recorded thereon a computer program for performing a news image-text non-relevance detection method according to one embodiment of the present invention.
[0015] According to one aspect of the present invention, by providing a news image-text non-relevance detection device, method, and recording medium, a learning method capable of effectively representing visual-linguistic meaning for content detection using images that do not represent text can be proposed, utilizing unsupervised learning-based contrastive learning techniques. Since data without correct labels is used during embedding model training, labeling costs can be reduced.
[0016] Furthermore, AI can detect news images with low relevance to the text, enabling the selection of images that represent the news article text from among possible candidates, thereby improving the quality of news articles. Furthermore, this technology can be used to detect news articles using images with low relevance to the text, thereby addressing the problem of clickbait images in online news and social media environments.
[0017] Figure 1 is a conceptual diagram of a news image-text non-relevance detection device according to an embodiment of the present invention;
[0018] FIG. 2 is a device diagram showing the internal blocks of a news image-text non-relevance detection device according to an embodiment of the present invention.
[0019] Figure 3 is a diagram illustrating how the pre-learning unit of Figure 2 constructs difficult negative samples using a masked language model.
[0020] FIG. 4 is a diagram for explaining the overall operation of a news image-text non-relevance detection device according to an embodiment of the present invention;
[0021] Figure 5 is a drawing showing the detailed configuration of the pre-learning unit and embedding unit of Figure 2.
[0022] And, Fig. 6 is a flowchart showing a news image-text non-relevance detection process according to an embodiment of the present invention.
[0023] The following detailed description of the present invention refers to the accompanying drawings, which illustrate specific embodiments in which the present invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the present invention. It should be understood that the various embodiments of the present invention, while different from each other, are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the present invention. Furthermore, it should be understood that the positions or arrangements of individual components within each disclosed embodiment may be modified without departing from the spirit and scope of the present invention. Accordingly, the following detailed description is not intended to be limiting, and the scope of the present invention is defined only by the appended claims, along with the full scope of equivalents to which such claims are entitled, if properly described. Like reference numerals in the drawings designate the same or similar functionality throughout the several aspects.
[0024] The components according to the present invention are defined by functional distinctions rather than physical distinctions, and can be defined by the functions each component performs. Each component may be implemented as hardware or program code and processing units that perform each function, and the functions of two or more components may be implemented by including them in a single component. Therefore, the names given to the components in the following embodiments are not intended to physically distinguish each component, but rather to suggest the representative functions performed by each component, and it should be noted that the technical spirit of the present invention is not limited by the names of the components.
[0025] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the drawings.
[0026] Fig. 1 is a conceptual diagram of a news image-text non-relevance detection device according to an embodiment of the present invention.
[0027] The urban news image-text non-relevance detection device (100) detects and outputs images with low relevance to the text of the input news content when news content consisting of images and text is input. Here, the image with low relevance to the text refers to an image that does not express or incorrectly represents the subject or the subject's actions of the event covered in the text. In other words, the relevance can be understood in relation to the "who" and "what" of the Six Ws.
[0028] FIG. 2 is a device diagram showing the internal blocks of a news image-text irrelevance detection device according to an embodiment of the present invention, FIG. 3 is a diagram for explaining a method by which the pre-learning unit of FIG. 2 constructs a difficult negative sample (i.e., counterfactual text) using a masked language model, FIG. 4 is a diagram for explaining the overall operation of the news image-text irrelevance detection device according to an embodiment of the present invention, and FIG. 5 is a diagram showing the detailed configuration of the pre-learning unit and the embedding unit of FIG. 2.
[0029] The news image-text non-relevance detection device (100) illustrated in FIG. 2 includes a pre-learning unit (110), an embedding unit (120), and a detection unit (130), and the embedding unit (120) is learned by the pre-learning unit (110).
[0030] The pre-learning unit (110) constructs a dataset without correct labels for learning the embedding unit (120), and the constructed dataset is applied to contrastive learning for learning the embedding unit (120) based on positive samples and negative samples. That is, the pre-learning unit (110) constructs a dataset by extracting image-text pairs constituting news content as positive samples and extracting counterfactual text generated based on the text of the image-text pair as negative samples, and then pre-trains a visual-language embedding model using the constructed dataset.
[0031] The above contrastive learning is a method of learning data expressed as a vector by moving closer to positive samples that are expected to have similar meanings and farther away from negative samples that are expected to have different meanings. This contrastive learning has been used as a technique to train visual-language embedding models with data without correct answers in existing studies such as CLIP (Contrastive Language Image Pretraining). However, it is difficult to understand the representativeness between images and texts by learning only general image-text pairs. Therefore, the present invention is significant in that it enables effective meaning expression from image-text data that does not have correct answers by utilizing counterfactual text in contrastive learning. Since news deals with factual information, it enables effective expression by learning to increase the correlation between the same image and factual information text by comparing the generated counterfactual text with image pairs.
[0032] The above pre-learning unit (110) may be composed of a counterfactual text generator (110-1) and an objective function (contrastive object) (110-2) as shown in FIG. 5. In this case, the counterfactual text generator (110-1) generates a counterfactual text (T) based on the text (T) of the news content. ) is generated, and contrastive learning is performed by applying the objective function (110-2). The objective function (110-2) will be described in more detail in mathematical expression 1 described below.
[0033] When news content consisting of images and text is input, the embedding unit (120) extracts vectors corresponding to images and vectors corresponding to text of the news content, respectively, through a visual-language embedding model that has been pre-trained based on the data set constructed in the pre-training unit (110).
[0034] An example of the input news content can be represented as in Fig. 4. That is, when news content consisting of text such as "Biden calls Mexican president an equal partner amid surge in border crossings" and an image representing such text is input to the embedding unit (120), the embedding unit (120) generates a vector v corresponding to the image through a pre-trained visual-language embedding model. I and the vector v corresponding to the above text T Prints out.
[0035] The above visual-language embedding model may be composed of an image encoder (120-3) and text encoders (120-1, 120-2) as shown in Fig. 5, in which case the image encoder (120-3) embeds an image (I) of news content to generate an image corresponding vector v. I , and the text encoder (120-2) embeds the text (T) of the news content to generate a vector v corresponding to the text. T, and the text encoder (120-1) outputs a counterfactual text ( ) as a vector corresponding to the counterfactual text . In addition, the visual-language embedding model may be a pre-prepared CLIP (Contrastive Language Image Pretraining) model.
[0036] Here, we describe in more detail the method of generating counterfactual texts to more effectively learn image representation using image-text pairs through Fig. 3. First, the visual-language embedding model extracts image-text pairs of a given data set as positive samples, extracts text embeddings of other samples in a mini-batch of the data set as easy negative samples, and extracts counterfactual text embeddings as difficult negative samples.
[0037] At this time, the counterfactual text embedding is generated by composing a candidate token set for a subject that can represent the text of a given image-text pair, selecting and masking a main token corresponding to the subject from the candidate token set, and predicting the masked token through a pre-prepared Masked Language Model (MLM).
[0038] Here, the above candidate token set is formed through a natural language processing preprocessing method such as a named entity recognizer based on the above text, and masking is performed by selecting tokens to be changed below a predefined threshold (e.g., 30% of the original text length) in order to not completely lose the meaning of the original text from the above candidate token set.
[0039] Additionally, to prevent the original token from being predicted when predicting masked tokens and to ensure that a variety of tokens can be predicted, the model's output is divided by a temperature constant to sample tokens from a softmax normalized distribution. The counterfactual text generated in this way is a difficult negative sample that resembles the original text but has subtly different content.
[0040] In Figure 3, the selected primary token is assumed to be {Biden}, and therefore the MLM predicts the masked token as {Trump}. The MLM may be a BERT (Bidirectional Encoder Representations from Transformers), which is based on a Transformer architecture used to predict words by considering context in both directions.
[0041] The above embedding unit (120) is pre-trained through contrastive learning according to the objective function using the dataset constructed through FIG. 3, and the objective function can be expressed as in the following mathematical expression 1.
[0042]
[0043] Here , represents a positive sample, person represents a difficult negative sample, person represents an easy negative sample, represents a temperature parameter that controls the similarity value. The pre-learning unit (110) trains the embedding unit (120) to lower the objective function, thereby enabling effective output of image and text expressions.
[0044] Afterwards, the detection unit (130) outputs a vector corresponding to the image output from the embedding unit (120). and the vector corresponding to the text The cosine similarity is calculated, and the calculated cosine similarity is set to a preset threshold. By comparing with the above text, an image with low relevance is detected and output. That is, the detection unit (130) detects and outputs an image with low relevance to the text if the cosine similarity is lower than the threshold. Here, the threshold is set by referring to the cosine similarity distribution for the verification data set using the embedding unit (120) learned by the above-mentioned pre-learning unit (110).
[0045] As another embodiment, the detection unit (130) detects a vector corresponding to the image output from the embedding unit (120). and the vector corresponding to the text By inputting the image-text pair into a pre-arranged classification model, the probability that the given image-text pair has a low correlation is calculated based on the correlation between the two vectors, and based on the calculated probability, an image with a low correlation with the text is detected and output. Here, the classification model is a deep neural network model and is prepared based on supervised learning.
[0046] Figure 6 is a flowchart illustrating a news image-text non-relevance detection process according to an embodiment of the present invention.
[0047] The news image-text non-relevance detection device pretrains a visual-language embedding model using a dataset built based on image-text pairs that constitute news content. (S501)
[0048] Then, when news content is input, the news image-text non-relevance detection device extracts a vector corresponding to the image of the news content and a vector corresponding to the text of the news content through a pre-trained visual-language embedding model through S501 (S503).
[0049] Thereafter, the news image-text non-relevance detection device detects an image with low relevance to the text based on the similarity between the vector corresponding to the image extracted in S503 and the vector corresponding to the text. (S505) Here, the image with low relevance to the text means an image whose relevance to the text is lower than a preset threshold, or an image for which a classification model based on supervised learning predicts a probability of low relevance higher than the threshold.
[0050] The news image-text non-relevance detection method of the present invention, as described above, may be implemented in the form of program commands that can be executed by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, and the like, either singly or in combination.
[0051] The program commands recorded on the above computer-readable recording medium may be specially designed and configured for the present invention or may be known and available to those skilled in the art of computer software.
[0052] Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions such as ROM, RAM, and flash memory.
[0053] Examples of program instructions include not only machine language codes, such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter or the like. The hardware device may be configured to operate as one or more software modules to perform processing according to the present invention, and vice versa.
[0054] Although various embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.
[0055] [Explanation of symbols]
[0056] 100: News Image-Text Inconsistency Detection Device
[0057] 110: Pre-learning Department
[0058] 120: Embedding section
[0059] 130: Detection Department
Claims
1. Pre-training unit that pre-trains a visual-language embedding model using a dataset built based on image-text pairs that constitute news content; When news content is input, an embedding unit extracts a vector corresponding to an image of the news content and a vector corresponding to text through the pre-trained visual-language embedding model; and A news image-text non-relevance detection device, comprising: a detection unit that detects an image having low relevance to the text based on the similarity between a vector corresponding to the image and a vector corresponding to the text.
2. In paragraph 1, A news image-text non-relevance detection device characterized in that the above-mentioned pre-learning unit constructs the dataset by extracting image-text pairs constituting the news content as positive samples and extracting counterfactual text generated based on the text of the image-text pairs as negative samples.
3. In paragraph 2, The above visual-language embedding model is, A news image-text non-relevance detection device characterized by being pre-trained through contrastive learning that learns to be close to the positive sample and far from the negative sample.
4. In paragraph 2, The above pre-learning section is, A news image-text non-relevance detection device characterized in that it creates a set of candidate tokens for a subject that can represent the above text, selects and masks a main token corresponding to the subject from the set of candidate tokens, and predicts the masked token through a pre-prepared masked language model (MLM) to generate the counterfactual text.
5. In paragraph 1, The above detection unit, A news image-text non-relevance detection device that calculates the cosine similarity between a vector corresponding to the image and a vector corresponding to the text, and compares the calculated cosine similarity with a preset threshold to detect an image having low relevance to the text.
6. In paragraph 1, The above detection unit, A news image-text non-relevance detection device that inputs a vector corresponding to the image and a vector corresponding to the text into a pre-arranged classification model, calculates a probability that the image-text pair has a low relevance based on the degree of relevance between the two vectors, and detects an image having a low relevance to the text based on the calculated probability.
7. In paragraph 1, A news image-text non-relevance detection device, characterized in that an image with low relevance to the above text means an image that does not express or incorrectly represents the subject or the subject's actions of the event covered in the above text.
8. A step of pre-training a visual-language embedding model using a dataset built based on image-text pairs that constitute news content; When news content is input, a step of extracting a vector corresponding to an image of the news content and a vector corresponding to text through the pre-trained visual-language embedding model; and A news image-text non-relevance detection method, comprising: a step of detecting an image having low relevance to the text based on the similarity between a vector corresponding to the image and a vector corresponding to the text.
9. In paragraph 8, A news image-text non-relevance detection method, characterized in that the step of pre-training the visual-language embedding model includes a step of constructing the dataset by extracting image-text pairs constituting the news content as positive samples and extracting counterfactual text generated based on the text of the image-text pairs as negative samples.
10. In paragraph 9, A news image-text non-relevance detection method, characterized in that the above visual-language embedding model is pre-trained through contrastive learning that learns to be close to the positive samples and far from the negative samples.
11. In paragraph 9, The steps to build the above dataset are: A news image-text non-relevance detection method characterized by comprising the steps of: composing a candidate token set for a subject that can represent the above text, selecting and masking a main token corresponding to the subject from the candidate token set, and generating the counterfactual text by predicting the masked token through a pre-prepared masked language model (MLM).
12. In paragraph 8, The step of detecting images with low relevance to the above text is: A news image-text non-relevance detection method comprising: calculating the cosine similarity between a vector corresponding to the image and a vector corresponding to the text, and comparing the calculated cosine similarity with a preset threshold to detect an image having low relevance to the text.
13. In paragraph 8, The step of detecting images with low relevance to the above text is as follows: A news image-text non-relevance detection method, which inputs a vector corresponding to the image and a vector corresponding to the text into a pre-arranged classification model, calculates a probability that the image-text pair has a low relevance based on the degree of relevance between the two vectors, and detects an image having a low relevance to the text based on the calculated probability.
14. In paragraph 8, A news image-text non-relevance detection method, characterized in that an image with low relevance to the above text means an image that does not express or incorrectly represents the subject or the subject's actions of the event covered in the above text.
15. A recording medium having recorded thereon a computer program for performing the news image-text non-relevance detection method of claim 8.
Citation Information
Patent Citations
Method for store database and smart phone interlock based refrigerator food information storage
KR1020210045724A
Method and apparatus for train neural networks for image training
KR1020230071719A
Manufacturing system of display device and manufacturing method of display device using the same
KR1020250031985A
Refill type lipstick container
KR1020250037876A
Tape gripping device
KR1020250057360A