A method for detecting fake text and image information in smart elderly care terminals
By employing multimodal feature encoding and hypergraph propagation reasoning, this study addresses the challenge of detecting cross-modal spoofing behavior in smart elderly care scenarios, achieving efficient identification and risk analysis of false information in images and text, and improving detection accuracy and interpretability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH OF CHINA
- Filing Date
- 2026-06-30
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to effectively identify cross-modal spoofing behavior in smart elderly care scenarios, and their large-scale image and text data retrieval efficiency is low. They also lack the ability to jointly model complex risk relationships, leading to frequent false positives and false negatives.
We employ a multimodal feature encoding, two-stage retrieval, and multimodal hypergraph construction approach. By using multimodal feature encoding of a text-image sample database and hypergraph propagation inference, we construct a unified multimodal representation to achieve efficient candidate sample retrieval and complex cross-modal relationship modeling. We then combine this with a visual large language model for risk analysis.
It improves the retrieval efficiency and cross-modal semantic modeling capabilities for detecting false information in images and text, effectively identifies hidden risks in complex elderly care fraud scenarios, and enhances detection accuracy and interpretability of results.
Smart Images

Figure CN122493472A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal information security detection technology, specifically to a method for detecting false text and image information in smart elderly care terminals. Background Technology
[0002] Because the elderly have relatively weaker ability to discern complex online information, in recent years, there has been a rapid increase in the amount of false information, deepfake information, and fraudulent content targeting the elderly.
[0003] Existing methods for detecting image and text forgery mostly rely on single-modal features or simple image-text matching, typically using image authenticity detection or text semantic classification for identification. These methods struggle to effectively identify cross-modal spoofing behaviors in smart elderly care scenarios. For example, attackers can use real images of elderly care facilities paired with fake text, or use real news text combined with forged images to generate misleading promotional content, leading to false positives or false negatives with traditional methods.
[0004] Furthermore, existing methods typically perform similar sample retrieval directly across the entire sample space. As the scale of smart elderly care image and text data grows, retrieval efficiency and reasoning complexity increase significantly. Simultaneously, they often rely on single similarity matching, lacking the ability to jointly model risk relationships, cross-modal conflict relationships, and multi-node associations, making it difficult to effectively reason about the implicit risks in complex elderly care fraud scenarios. Although multimodal large-scale models and graph structure reasoning techniques provide new technical paths for detecting image and text fraud, existing graph neural networks or ordinary graph structure methods typically only model binary relationships between nodes, making it difficult to characterize higher-order associations between entities, attributes, risk types, and image and text semantics. Moreover, traditional graph propagation mechanisms have limited collaborative reasoning capabilities for multi-source risk evidence.
[0005] Therefore, how to construct a method for detecting forged text and image content suitable for smart elderly care scenarios, so as to achieve efficient candidate sample retrieval, complex cross-modal relationship modeling, and joint reasoning of multi-source risk evidence, thereby improving the detection capabilities of smart elderly care terminals against risk scenarios such as false text and image information, identity impersonation, health product fraud, false medical advertising, and inducement to transfer money, has become an urgent technical problem to be solved. Summary of the Invention
[0006] To address the shortcomings of existing technologies in detecting misinformation in smart elderly care scenarios, such as insufficient cross-modal semantic association modeling capabilities, limited ability to reason about complex risk relationships, and low efficiency in large-scale image and text sample retrieval, this invention proposes a method for detecting misinformation in smart elderly care terminals. This method fully integrates multimodal feature encoding, two-stage retrieval, multimodal hypergraph construction, and hypergraph propagation reasoning mechanisms to achieve efficient detection and risk analysis of misinformation in smart elderly care scenarios.
[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A method for detecting misleading text and image information in smart elderly care terminals includes: A text-image sample library is constructed, and multimodal feature encoding is performed on the images and text in the text-image samples to obtain image features, text features, and text-image interaction features that can reflect the differences and interactions between text and images. These features are then fused to obtain a unified multimodal representation. Based on the unified multimodal representation, the image and text sample library is clustered to establish a first-level cluster index; the unified multimodal representation is extracted from the image and text samples to be detected, and the similarity retrieval is performed through the first-level cluster index to determine the query-related first-level clusters. Secondary clustering is performed within the query-related first-level clusters to obtain relevant local sub-clusters, and candidate reference sample sets are obtained within the relevant local sub-clusters based on multimodal weighted similarity. A multimodal hypergraph is constructed based on the image and text samples to be detected and the candidate reference sample set; Construct a node set and a node-hyperedge association matrix, perform propagation calculations on the hypergraph, and obtain the risk association strength between each node and the text and image sample to be detected; The strength of the hyperedge risk evidence is calculated and sorted based on the risk association strength. The set of key risk evidence hyperedges is selected and the enhanced reference sample set is obtained by reverse indexing. The text and image samples to be detected, the set of key risk evidence hyperedges and the enhanced reference sample set are input into the visual big language model, and the true or false judgment result and risk type are output.
[0008] In one embodiment, the construction of the image-text sample library involves multimodal feature encoding of the images and text in the image-text samples to obtain image features, text features, and image-text interaction features that reflect the differences and interactions between images and text, which are then fused to obtain a unified multimodal representation. Specifically, this includes: Image and text sample library ; express The first in One image and text sample, express The image in express The text in This indicates the number of image and text samples in the image and text sample library; For each image and text sample in the image and text sample library Image and text feature encoding; Image features for: Text features for: ; Indicates an image encoder. Indicates a text encoder; Normalize the image features and text features separately: ; ; For normalized image features, Normalized text features; Represents the L2 norm; Features of text and image interaction for: ; This represents a vector concatenation operation. Indicates element-wise multiplication. Used to represent the difference information between image features and text features. Used to represent the interaction information between image features and text features; By fusing image features, text features, and graphic-text interaction features, a unified multimodal representation is generated. ; in, These represent the weights of image features, text features, and image-text interaction features, respectively.
[0009] In one embodiment, the step of clustering the image and text sample library based on unified multimodal representation and establishing a first-level cluster index specifically includes: Unified multimodal representation based on all image and text samples Perform first-level clustering on the image and text sample database to obtain a global cluster set. : ; in, Indicates the first One first-level cluster, Indicates the number of first-level clusters; Create a first-level cluster index : ; Each index unit includes a first-level cluster. , Corresponding central sample and the unified multimodal representation of the central samples .
[0010] In one embodiment, the image and text samples to be detected are extracted with a unified multimodal representation, and a similarity retrieval is performed through a first-level cluster index to determine the query-related first-level cluster. Secondary clustering is then performed within this query-related first-level cluster to obtain relevant local sub-clusters. A set of candidate reference samples is then retrieved within these local sub-clusters. Specifically, this includes: Sample images to be tested Constructing a unified multimodal representation Based on the first-level cluster index of the sample database Calculate the similarity between the image / text sample to be detected and the center sample of each first-level cluster: ; Represents cosine similarity. express With the Similarity between samples of the first-level cluster centers; Select the P first-level clusters with the highest similarity to the image and text sample to be detected as the relevant first-level clusters for the query. ; In querying the relevant first-level cluster Internal secondary clustering: ; in, This indicates a query for the first-level cluster. A local sub-cluster Indicates the number of local subclusters; Subsequently, the similarity between the image / text sample to be detected and the sample at the center of the local sub-cluster is calculated. : ; in, For the first Local sub-cluster center samples Unified multimodal representation; Select B local sub-clusters with the highest similarity to the image / text sample to be detected; for each image / text sample in the selected local sub-clusters Calculate multimodal weighted similarity Select multimodal weighted similarity The K highest-ranking image and text samples are used as the candidate reference sample set.
[0011] In one embodiment, the multimodal weighted similarity is calculated as follows: ; in, , , , Representing image similarity respectively Text similarity Image and text interaction similarity and unified multimodal similarity The weight, The image features, text features, image-text interaction features, and unified multimodal representation of the image-text samples to be detected are used. This represents the image features, text features, image-text interaction features, and unified multimodal representation of the i-th image-text sample in a local subcluster. Let be the cosine similarity.
[0012] In one embodiment, the construction of a multimodal hypergraph based on the image / text sample to be detected and the candidate reference sample set specifically includes: Construct each candidate reference sample Seed vector relating to the image / text sample to be detected : ; in, A unified multimodal representation of the image and text samples to be detected. Indicates candidate reference sample Unified multimodal representation, This represents the image / text sample to be detected and the candidate reference sample. Multimodal weighted similarity; This represents a vector concatenation operation. Indicates element-wise multiplication; Constructing image-text difference seed vectors : ; in, Used to characterize the differences between the image and text samples to be detected. Used to characterize the textual and graphical differences of the candidate reference sample itself. Used to characterize the differences in image-text interaction features between the image-text sample to be detected and the candidate reference sample; The image features, text features, and image-text interaction features of the image-text sample to be detected are used to identify the image features, text features, and image-text interaction features. for Image features, text features, and graphic-text interaction features; calculate The corresponding hyperedge representation vector : ; It is a linear mapping matrix. This indicates a normalization operation; For candidate reference sample set All candidate reference samples By performing vector expansion, we obtain the hyperedge set. : ; remember Candidate reference sample The corresponding hyperedge, where They are represented as sample nodes. Related entities, attributes, and risk types are represented as related semantic nodes, and hyperedges Used to connect sample nodes and related semantic nodes; both sample nodes and related semantic nodes are hypergraph nodes.
[0013] In one embodiment, the method further includes: introducing redundancy removal constraints between hyperedges, for any two hyperedges and Define hyperedge similarity : ; like Then it is considered that the super-edge and There are redundant super-edges in the data. The redundant super-edges are removed according to the preset clipping strategy. Finally, we obtain the adaptive hyperedge set. ;in, This represents the hyperedge similarity threshold. This indicates an over-edge clipping operation.
[0014] In one embodiment, the construction of the node set and node-hyperedge association matrix, and the propagation calculation of the hypergraph to obtain the risk association strength between each node and the text / graph sample to be detected, specifically includes: Build a node set and node-hyperedge incidence matrix The node set includes entity nodes, attribute nodes, and risk type nodes; The weight of each hyperedge is jointly determined by multimodal weighted similarity, hyperedge internal consistency, graph-text conflict intensity, and risk node activation intensity: ; Candidate reference sample Corresponding hyperedge The weight, For multimodal weighted similarity, This represents the activation function. As weight; Hyperedge internal consistency for: ; in, Represents a node Feature representation, Indicates the superedge The correlation cardinality is used to normalize the internal consistency of the hyperedge; Intensity of conflict between images and text for: ; Risk node activation intensity for: ; in, This represents a smoothing term used to avoid a denominator of zero. This represents the set of nodes corresponding to the risk types in the hypergraph; Hyperedge weight matrix for: ; in Describes the constructor for a diagonal matrix; Define the diagonal elements of the node degree matrix : ; Represents the node-hyperedge incidence matrix Middle node With super-edge The relationship between them; Define the diagonal elements of the hyperedge degree matrix : ; Construct the hypergraph propagation matrix based on the node-hyperedge association matrix, hyperedge weight matrix, node degree matrix, and hyperedge degree matrix. : ; For transpose; Using the image / text sample node to be detected as the initial activated node, construct the initial propagation state vector for graph propagation: ; in, The initial state vector for graph propagation indicates that the node of the graph sample to be detected is set as the propagation source point; the other nodes initially have no propagation signal. Then iterative propagation is carried out: ; in, Indicates the current iteration number. For the propagation coefficient, Indicates the first The node state of the step, Indicates by The updated next node propagation state vector will be passed through The propagation state vector obtained after the next iteration of propagation As a representation node The strength of the risk association between the image and text sample to be detected.
[0015] In one embodiment, the step of calculating and sorting the strength of superedge risk evidence based on the risk association strength, filtering the key risk evidence superedge set, and obtaining the enhanced reference sample set through reverse indexing specifically includes: For each superedge Taking into account the weight of the hyperedge, the risk association strength of the nodes inside the hyperedge, and the consistency within the hyperedge, the risk evidence strength of the hyperedge is calculated. : ; in, Indicates the superedge The weight, Indicates consistency within the hyperedge; For the set of superedges All hyperedges are sorted in descending order of risk evidence strength, and the M hyperedges with the highest risk evidence strength are selected as the set of key risk evidence hyperedges. According to the customs The candidate reference samples corresponding to the inverted index are used to obtain the enhanced reference sample set. .
[0016] In one embodiment, the step of inputting the image / text sample to be detected, the set of key risk evidence hyperedges, and the set of enhanced reference samples into the visual large language model, and outputting the authenticity judgment result and risk type, specifically includes: The image and text sample to be detected Enhance the reference sample set and the super-edge set of key risk evidence Both are input into the visual large language model, which performs comprehensive reasoning based on the image semantics, text semantics, image-text consistency relationship, and knowledge conflict relationship between the image-text sample to be detected and the enhanced reference sample. Finally, it outputs the authenticity judgment result, risk type, and judgment basis of the image-text sample to be detected: ; in, Represents a visual large language model. This indicates the result of determining whether the image or text sample to be tested is real or fake. This indicates the risk type corresponding to the image / text sample to be tested. Indicates the basis for judgment.
[0017] Compared with the prior art, the beneficial technical effects of the present invention are: Improved retrieval efficiency: By constructing a first-level cluster index for the sample library and combining it with a two-stage retrieval strategy, this invention avoids high-complexity retrieval in the full sample space, significantly improving the retrieval efficiency of candidate samples in a large-scale smart elderly care image and text sample library, and meeting the real-time detection needs of smart elderly care terminals.
[0018] Enhanced cross-modal semantic modeling capabilities: In the feature encoding stage, this invention not only extracts image and text features, but also constructs image-text interaction features based on the difference and interaction information between normalized image features and text features, and then fuses them to generate a unified multimodal representation, effectively capturing the semantic consistency or conflict relationship between images and text, and improving the detection accuracy of image-text forgery content.
[0019] Achieving high-order risk relationship reasoning: This invention introduces a multimodal hypergraph structure, which connects entity nodes, attribute nodes, and risk type nodes simultaneously through hyperedges, overcoming the limitation of ordinary graphs that can only model binary relationships. Combined with the hypergraph propagation mechanism, risk information can be jointly propagated among multiple nodes and hyperedges, effectively uncovering implicit high-order risk relationships in complex elderly care fraud scenarios.
[0020] Enhancing the interpretability of detection results: This invention scores and ranks hyperedge risk evidence to output a set of key risk evidence hyperedges and corresponding candidate reference samples. Finally, combined with a visual large language model, it not only outputs the authenticity judgment results of images and text and the risk type, but also provides the judgment criteria (such as semantic conflicts, knowledge conflicts, risk propagation paths, etc.), thereby improving the interpretability of detection results and the ability to assist decision-making. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation
[0022] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.
[0023] like Figure 1 As shown, a method for detecting false information in images and text for smart elderly care terminals according to the present invention includes the following steps: S1. Construct a text and image sample library, encode the images and text in the text and image samples using multimodal features to obtain image features, text features, and text and image interaction features that can reflect the differences and interactions between text and images, and fuse them to obtain a unified multimodal representation; S2, cluster the image and text sample library based on unified multimodal representation and establish a first-level cluster index; extract unified multimodal representation of the image and text samples to be detected, perform similarity retrieval through the first-level cluster index to determine the relevant first-level cluster, perform secondary clustering within the relevant first-level cluster to obtain relevant local sub-clusters, and obtain the candidate reference sample set within the relevant local sub-clusters based on multimodal weighted similarity; S3, Construct a multimodal hypergraph based on the image and text samples to be detected and the candidate reference sample set; S4. Construct a set of nodes and a node-hyperedge association matrix, perform propagation calculations on the hypergraph, and obtain the risk association strength between each node and the text / graph sample to be detected. S5: Calculate and sort the super-edge risk evidence strength based on the risk association strength, filter the key risk evidence super-edge set, and obtain the enhanced reference sample set through reverse indexing; input the image and text sample to be detected, the key risk evidence super-edge set, and the enhanced reference sample set into the visual big language model, and output the true / false judgment result and risk type.
[0024] This invention constructs a smart elderly care image and text sample library, encodes the image and text samples using multimodal features, and establishes a first-level cluster index for the sample library based on a unified multimodal representation. Subsequently, a two-stage retrieval is performed on the image and text samples to be detected to obtain a set of candidate reference samples related to the image and text samples to be detected. Further, a multimodal hypergraph is constructed based on the image and text samples to be detected and the candidate reference samples. The risk association strength of nodes is calculated through hypergraph propagation, and the risk evidence strength of hyperedges is ranked. The image and text samples to be detected, the set of key risk evidence hyperedges, and the enhanced reference sample set are input into a visual big language model to output the authenticity judgment result and risk type.
[0025] The present invention will be described in detail below in several parts.
[0026] 1. Construct a smart elderly care image and text sample library and perform multimodal feature encoding on the sample library.
[0027] First, collect text and image data from smart elderly care scenarios to construct a smart elderly care text and image sample library. This text and image data may include medical insurance reimbursement instructions, community service announcements, elderly care institution promotions, health promotion content, remote consultation text and image materials, emergency notification text and image materials, health product promotion text and image materials, and text and image content historically received by smart elderly care terminals.
[0028] In practical applications, misinformation in smart elderly care scenarios often manifests not only as falsified text content but also as semantic inconsistencies between images and text. For example, attackers might use genuine community announcement images paired with fake transfer notification text, or genuine medical promotional slogans paired with counterfeit health product images. Therefore, this approach does not simply save text or images but rather uniformly represents each sample as a graphic-text sample composed of both images and text, enabling subsequent modeling of the relationship between the two modalities.
[0029] The image and text sample library is represented as follows: ; in, Indicates the first in the sample library One image and text sample, This indicates the image corresponding to the text and image sample. This indicates the text corresponding to the image and text sample. This indicates the number of image and text samples in the sample library.
[0030] Furthermore, for each image and text sample in the sample library Image and text encoding are performed. The image encoder extracts visual semantic information from the input image, such as people, organization logos, scenes, objects, poster structures, and potential risk visual cues within the image; the text encoder extracts semantic information from the input text, such as monetary information, medical promotional language, suggestive behaviors, contact information, and risk keywords. In a preferred embodiment, the image encoder may employ the ViT model, and the text encoder may employ the bidirectional encoder representation model BERT, thereby extracting features from the input image and input text respectively.
[0031] Image features are represented as follows: Text features are represented as: ;in, Indicates an image encoder. Indicates a text encoder. Representing image features, Represents text features.
[0032] Since image features and text features may originate from different encoders, their feature distributions, numerical scales, and vector dimensions may differ. To reduce the impact of scale differences between different modal features on subsequent similarity calculations and feature fusion, image features and text features are normalized separately: ; ; in, This represents the L2 norm.
[0033] After obtaining the normalized image and text features, we further construct image-text interaction features. These features include not only the image and text features themselves, but also the difference and interaction information between them. The difference information is used to characterize whether there is a semantic deviation between the image and the text, while the interaction information is used to characterize whether there is a consistent or complementary relationship between the image and the text.
[0034] The characteristics of text-image interaction are represented as follows: ; in, This represents a vector concatenation operation. Indicates element-wise multiplication. Used to represent the difference information between images and text. Used to represent interactive information between images and text.
[0035] Furthermore, by fusing image features, text features, and graphic-text interaction features, a unified multimodal representation is generated: ; in, These represent the weights of image features, text features, and image-text interaction features, respectively.
[0036] Through the above processing, each image-text sample in the sample library is mapped to a unified multimodal representation. This representation simultaneously preserves image semantics, text semantics, and the interaction relationships between images and text, providing fundamental feature support for subsequent clustering indexing, similar sample retrieval, hypergraph construction, and risk propagation.
[0037] 2. Establish a first-level cluster index for the sample library based on a unified multimodal representation.
[0038] After completing the multimodal feature encoding of the sample library, a unified multimodal representation based on all image and text samples is then developed. Perform first-level clustering on the sample database to obtain a global cluster set: ; in, Indicates the first One first-level cluster, This indicates the number of first-level clusters.
[0039] The purpose of setting up a first-level cluster index is to improve retrieval efficiency in a large-scale smart elderly care image and text sample database. If the similarity calculation is performed directly between the image and text sample to be detected and all image and text samples in the database, the computational cost will increase significantly as the size of the database expands. By pre-dividing the database into multiple semantically similar first-level clusters, the semantic region to which the sample to be detected may belong can be determined first, and then a refined search can be performed within the corresponding region, thereby reducing interference from irrelevant samples and improving retrieval efficiency.
[0040] In one specific implementation, the K-Medoids clustering algorithm is used to perform first-level clustering on the unified multimodal representation. Compared with ordinary mean clustering, K-Medoids clustering selects true samples as cluster centers, thus making it more suitable for subsequent sample indexing and reference sample interpretation. Its optimization objective is: ; in, Indicates the first The central sample of a first-order cluster This represents the unified multimodal representation of the central sample. Represents the distance function.
[0041] Furthermore, a first-level cluster index for the sample database is established: ; Each index unit includes a first-level cluster. The central sample corresponding to this cluster and the unified multimodal representation of the central sample .
[0042] This index structure allows for the prior calculation of similarity between received image / text samples and primary cluster centers, without requiring a direct traversal of all samples. This enables the rapid identification of candidate retrieval regions within the global sample space, providing a foundation for subsequent refined retrieval.
[0043] 3. Construct a candidate reference sample set.
[0044] The first stage of the image and text sample search is performed. The image and text sample to be detected is represented as follows: ; in, This represents the image to be detected. This indicates the text to be tested.
[0045] First, image and text features are extracted from the image and text samples to be detected using the same encoding method as the sample library. This ensures that the samples to be detected and the samples in the sample library are in the same feature space, which facilitates subsequent similarity calculations.
[0046] The features of the image to be detected are represented as follows: ; The features of the text to be detected are represented as follows: ; Subsequently, the features of the image and text to be detected are normalized: ; ; Furthermore, the text-image interaction features of the text-image samples to be detected are constructed: .
[0047] Furthermore, a unified multimodal representation of the image and text samples to be detected is constructed: .
[0048] After completing the encoding of the image and text samples to be detected, the data is then processed based on the first-level cluster index of the sample database. Calculate the similarity between the image / text sample to be detected and the center of each first-level cluster: ; Expanded to: ; in, This indicates that the image / text sample to be detected is related to the first... The similarity between the centers of each primary cluster. The higher the similarity, the closer the sample in that primary cluster is to the image / text example to be detected in the multimodal semantic space.
[0049] Select the P first-level clusters that are most relevant to the image and text samples to be detected: ; This indicates that the highest similarity score will be returned. A first-level cluster.
[0050] In one specific implementation, when only the most relevant first-level cluster is selected, we have: ; in: ; This process first identifies the candidate retrieval regions most relevant to the images and text samples to be detected within the global cluster structure of the sample library, thus avoiding unconstrained retrieval throughout the entire sample library.
[0051] Furthermore, in order to obtain more refined candidate reference samples within the relevant first-level clusters, when querying the relevant first-level clusters... Internal secondary clustering or local sub-clustering: ; in, This indicates a query for the first-level cluster. A local sub-cluster This indicates the number of local subclusters.
[0052] The objective of quadratic clustering is: ; in, Indicates the first A local sub-cluster center sample.
[0053] Subsequently, the similarity between the image / text sample to be detected and the center of the local sub-cluster is calculated: ; And select the Top-B local sub-clusters that are most relevant to the image and text sample to be detected: ; in This indicates that the highest similarity score will be returned. Indexes.
[0054] Through the above secondary search, the search can be further narrowed down from the primary cluster to a more granular local similarity region, thereby improving the relevance of the candidate reference sample.
[0055] Furthermore, within the selected local sub-clusters, for each candidate image / text sample... Calculate multimodal weighted similarity: ; in, , , , These represent the weights of image similarity, text similarity, image-text interaction similarity, and unified multimodal similarity, respectively.
[0056] This multimodal weighted similarity method considers image similarity, text similarity, image-text interaction relationships, and overall multimodal semantic similarity. Compared to using only a single similarity metric for retrieval, this method can more comprehensively measure the relevance between candidate reference samples and the image-text examples to be detected.
[0057] Based on multimodal weighted similarity, the Top-K set of most relevant candidate reference samples is obtained: ; in: .
[0058] Each candidate reference sample is an image and text sample: ; Phase 1 Output for: .
[0059] 4. Construct a multimodal hypergraph.
[0060] After the first stage of retrieval, which involves constructing a multimodal hypergraph based on candidate reference samples, the multimodal vectors of the text and image samples to be detected are obtained. , and the multimodal vectors of candidate reference samples .
[0061] In this invention, ordinary graph structures can typically only describe binary relationships between two nodes. However, false information in the context of smart elderly care often involves complex high-order relationships between multiple entities, attributes, and risk types. For example, a false health product advertisement may simultaneously involve multiple elements such as "elderly people," "disease treatment," "low-price discounts," "expert endorsements," and "inducement to purchase." To better model these complex relationships, this invention constructs a multimodal hypergraph, enabling a single hyperedge to connect multiple related nodes simultaneously, thereby expressing the joint risk relationships between multiple nodes.
[0062] For each candidate reference sample obtained in the first stage of retrieval Construct a seed vector relating the image and text samples to be detected: ; in, Represents the unified multimodal vector of the image and text sample to be detected. Represents the unified multimodal vector of the candidate reference sample. Indicate the differences between the two. This represents the interaction characteristics between the two. This represents the score of the candidate sample retrieval in the first stage.
[0063] To further highlight cross-modal anomalous relationships in text-image misinformation, a text-image difference seed vector is constructed: ; in, Used to characterize the differences between the image and text samples to be detected. Used to characterize the textual and graphical differences of the candidate reference sample itself. Used to characterize the differences in text-image interaction features between the text-image sample to be detected and the candidate reference sample.
[0064] Finally, the hyperedge representation vector corresponding to the second stage is obtained. : ; in, It is a linear mapping matrix. This indicates a normalization operation.
[0065] Furthermore, vector expansion is performed on all candidate reference samples output from the first stage to obtain the hyperedge set: ; Each super edge It consists of corresponding candidate reference samples and their related entity nodes, attribute nodes, and risk type nodes, and is used to describe the higher-order risk relationship between candidate reference samples and text / image samples to be detected.
[0066] To avoid excessive redundancy in hyperedges, redundancy removal constraints are introduced between hyperedges. For any two hyperedges... and Define hyperedge similarity: ; like If the two superedges are considered to have a high degree of similarity, then redundant superedges are removed according to a preset pruning strategy.
[0067] The final adaptive hyperedge set is obtained as follows: ; in, This represents the hyperedge similarity threshold. This indicates an over-edge clipping operation.
[0068] This process enables the construction of a more compact and efficient multimodal hypergraph structure based on candidate reference samples, reducing the interference of redundant risk relationships on subsequent propagation inference.
[0069] 5. Hypergraph propagation.
[0070] Construct a node-hyperedge incidence matrix and perform hypergraph propagation. After constructing an adaptive hyperedge set, further construct a node set. and node-hyperedge incidence matrix The node set can include entity nodes, attribute nodes, and risk type nodes automatically extracted or inferred from images, text, and candidate reference samples.
[0071] Entity nodes can represent objects such as people, institutions, medicines, products, communities, and hospitals involved in the text and image content; attribute nodes can represent information such as price, time, location, contact information, preferential conditions, disease description, and service items; risk type nodes can represent risk categories such as identity impersonation, false medical advertising, inconsistency between text and images, and inducement to transfer money.
[0072] Each super edge The weights are jointly determined by multimodal weighted similarity, hyperedge internal consistency, graph-text conflict intensity, and risk node activation intensity: ; in, For multimodal weighted similarity, This represents the activation function.
[0073] The internal consistency of a hyperedge is defined as: ; in, Represents a node Feature representation, Indicates the superedge The cardinality of the set of associated nodes is used to determine the hyperedge. The internal consistency is normalized.
[0074] The intensity of text-image conflict is defined as: .
[0075] Image-text conflict intensity measures whether there are semantic inconsistencies between the image and text within the detected image-text sample and candidate reference samples. A larger difference between the image content and the text content increases the score for this item, indicating that the image-text sample may pose a higher risk.
[0076] The activation intensity of a risk node is defined as: ; in, This represents a smoothing term, used to avoid a denominator of zero.
[0077] Therefore, the hyperedge weight matrix is expressed as: ; in Describes the constructor for a diagonal matrix; Furthermore, define the node degree matrix: .
[0078] Define the hyperboundary degree matrix: ; Construct the hypergraph propagation matrix based on the node-hyperedge association matrix, hyperedge weight matrix, node degree matrix, and hyperedge degree matrix: .
[0079] Using the nodes of the image and text samples to be detected as the initial activated nodes, construct the initial propagation state vector for graph propagation. : ; in The initial state vector for graph propagation indicates that the node of the graph sample to be detected is set as the propagation source point; the other nodes initially have no propagation signal.
[0080] Then iterative propagation is carried out: ; in, Indicates the current iteration number. For the propagation coefficient, Indicates the first The node state of the step, Indicates by Update the propagation state vector of the next node obtained from the update.
[0081] Through the hypergraph propagation process described above, risk information can be jointly propagated among entity nodes, attribute nodes, risk type nodes, and hyperedges. When certain risk nodes, suspicious attribute nodes, or related candidate sample nodes have a strong correlation with the text / graph example to be detected, their risk information can be further diffused to related nodes through hyperedges, thereby helping to discover hidden higher-order risk relationships.
[0082] 6. Construct an enhanced reference sample set.
[0083] This invention employs vector-extended hyperedge risk evidence retrieval, further scoring the risk evidence on the hyperedges after hypergraph propagation. Unlike methods that only output classification results, this invention not only determines whether the content of the detected image or text is fraudulent, but also extracts key risk evidence leading to that determination, thereby improving the interpretability of the detection results.
[0084] For each superedge Taking into account the weight of the hyperedge, the risk propagation results of the nodes inside the hyperedge, and the consistency within the hyperedge, the strength of the hyperedge risk evidence is calculated. Specifically, the strength of the risk evidence is defined as: ; in, Indicates the superedge The evidence weight is used to measure the reference value of the superedge in the risk assessment of the current image and text sample to be detected. Used to represent the average risk association strength between all nodes inside the hyperedge and the text / image sample to be detected. Used to indicate the degree of consistency between the nodes inside the hyperedge and the hyperedge representation vector.
[0085] Subsequently, the adaptive hyperedge set All hyperedges are sorted in descending order of risk evidence strength, and the M hyperedges with the highest scores are selected as the set of key risk evidence hyperedges. This represents the set of superedges containing key risk evidence.
[0086] Furthermore, based on the hyper-edge set of key risk evidence The candidate reference samples corresponding to the inverted index are used to obtain the second-stage enhancement reference sample set: ; in, This represents the second-stage enhanced reference sample set.
[0087] 7. Generate the final detection results.
[0088] After the first-stage candidate reference sample retrieval and the second-stage hypergraph evidence retrieval, the query image and text samples to be detected, the second-stage enhanced reference sample set, and the key risk evidence hyperedge set are obtained.
[0089] The example of the text and image to be detected is represented as follows: ; The final set of reference images and text examples is represented as follows: ; Each reference image and text example consists of an image and text, namely: ; in, Indicates the first The images in the final reference illustration, Indicates the first The text in the final reference image and text sample.
[0090] Furthermore, the image and text samples to be tested will be... Second-stage enhanced reference sample set and the super-edge set of key risk evidence Both are input into the visual large language model. The visual large language model performs comprehensive reasoning based on the image semantics, text semantics, image-text consistency relationship, knowledge conflict relationship, and historical risk sample support between the image-text sample to be detected and the reference image-text sample. Finally, it outputs the true or false judgment result, risk type, and judgment basis of the image-text sample to be detected.
[0091] This process can be represented as: ; in, Represents a visual large language model. This indicates the result of determining whether the image or text sample to be tested is real or fake. This indicates the risk type corresponding to the image / text sample to be tested. Indicates the basis for judgment.
[0092] Furthermore, risk types include, but are not limited to, inconsistencies between images and text, identity impersonation, health product fraud, false medical advertising, inducement to transfer funds, false community notifications, and false remote consultation information. The criteria for judgment include, but are not limited to, the similarity between the sample to be detected and historical fake samples, the semantic conflict between images and text, the inconsistency between text and trusted knowledge, the propagation path corresponding to key risk hyperedges, and the degree of support for the judgment results from candidate reference samples.
[0093] Therefore, this invention can output the authenticity judgment result, risk type and judgment basis of the image and text sample to be detected, realize the detection of false information in images and text in smart elderly care scenarios, risk identification and result interpretation, and improve the ability of smart elderly care terminals to identify relevant risk content.
[0094] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0095] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple steps or stages, which are not necessarily completed at the same time, but may be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but may be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0096] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0097] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.
[0098] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for detecting false information of images and texts for a smart elderly care terminal, characterized in that, include: A text-image sample library is constructed, and multimodal feature encoding is performed on the images and text in the text-image samples to obtain image features, text features, and text-image interaction features that can reflect the differences and interactions between text and images. These features are then fused to obtain a unified multimodal representation. Based on the unified multimodal representation, the image and text sample library is clustered to establish a first-level cluster index; the unified multimodal representation is extracted from the image and text samples to be detected, and the similarity retrieval is performed through the first-level cluster index to determine the query-related first-level clusters. Secondary clustering is performed within the query-related first-level clusters to obtain relevant local sub-clusters, and candidate reference sample sets are obtained within the relevant local sub-clusters based on multimodal weighted similarity. A multimodal hypergraph is constructed based on the image and text samples to be detected and the candidate reference sample set; Construct a node set and a node-hyperedge association matrix, perform propagation calculations on the hypergraph, and obtain the risk association strength between each node and the text and image sample to be detected; The strength of the hyperedge risk evidence is calculated and sorted based on the risk association strength. The set of key risk evidence hyperedges is selected and the enhanced reference sample set is obtained by reverse indexing. The text and image samples to be detected, the set of key risk evidence hyperedges and the enhanced reference sample set are input into the visual big language model, and the true or false judgment result and risk type are output. 2.The method of claim 1, wherein, The construction of the image-text sample library involves multimodal feature encoding of the images and text in the samples to obtain image features, text features, and image-text interaction features that reflect the differences and interactions between images and text. These features are then fused to obtain a unified multimodal representation, specifically including: Image and text sample library ; express The first in One image and text sample, express The image in express The text in This indicates the number of image and text samples in the image and text sample library; For each image and text sample in the image and text sample library Image and text feature encoding; Image features for: Text features for: ; Indicates an image encoder. Indicates a text encoder; Normalization is performed on both image features and text features: ; ; For normalized image features, Normalized text features; Represents the L2 norm; Features of text and image interaction for: ; This represents a vector concatenation operation. Indicates element-wise multiplication. Used to represent the difference information between image features and text features. Used to represent the interaction information between image features and text features; By fusing image features, text features, and graphic-text interaction features, a unified multimodal representation is generated: ; in, These represent the weights of image features, text features, and image-text interaction features, respectively.
3. The method for detecting false text and image information for smart elderly care terminals according to claim 1, characterized in that, The method of clustering the image and text sample database based on unified multimodal representation and establishing a first-level cluster index specifically includes: Unified multimodal representation based on all image and text samples Perform first-level clustering on the image and text sample database to obtain a global cluster set. : ; in, Indicates the first A first-level cluster, Indicates the number of first-level clusters; Create a first-level cluster index : ; Each index unit includes a first-level cluster. , Corresponding central sample and the unified multimodal representation of the central samples .
4. The method for detecting false text and image information in smart elderly care terminals according to claim 3, characterized in that, The process involves extracting a unified multimodal representation from the image and text samples to be detected, determining the relevant first-level clusters through similarity retrieval using a first-level cluster index, performing secondary clustering within these clusters to obtain relevant local sub-clusters, and retrieving a set of candidate reference samples within these local sub-clusters. Specifically, this includes: Sample images to be tested Constructing a unified multimodal representation Based on the first-level cluster index of the sample database Calculate the similarity between the image / text sample to be detected and the center sample of each first-level cluster: ; Represents cosine similarity. express With the Similarity between samples of the center of each primary cluster; Select the P first-level clusters with the highest similarity to the image and text sample to be detected as the relevant first-level clusters for the query. ; In querying the relevant first-level cluster Internal secondary clustering: ; in, This indicates a query for the first-level cluster. A local sub-cluster Indicates the number of local subclusters; Subsequently, the similarity between the image / text sample to be detected and the sample at the center of the local sub-cluster is calculated. : ; in, For the first Local sub-cluster center samples Unified multimodal representation; Select B local sub-clusters with the highest similarity to the image / text sample to be detected; for each image / text sample in the selected local sub-clusters Calculate multimodal weighted similarity Select multimodal weighted similarity The K highest-ranking image and text samples are used as the candidate reference sample set.
5. A method for detecting false text and image information in smart elderly care terminals according to claim 4, characterized in that, The multimodal weighted similarity is calculated as follows: ; in, , , , Representing image similarity respectively Text similarity Image and text interaction similarity and unified multimodal similarity The weight, The image features, text features, image-text interaction features, and unified multimodal representation of the image-text samples to be detected are used. This represents the image features, text features, image-text interaction features, and unified multimodal representation of the i-th image-text sample in a local subcluster. Let be the cosine similarity.
6. A method for detecting false text and image information in smart elderly care terminals according to claim 1, characterized in that, The construction of a multimodal hypergraph based on the image and text samples to be detected and the candidate reference sample set specifically includes: Construct each candidate reference sample Seed vector relating to the image / text sample to be detected : ; in, A unified multimodal representation of the image and text samples to be detected. Indicates candidate reference sample A unified multimodal representation, This represents the image / text sample to be detected and the candidate reference sample. Multimodal weighted similarity; This represents a vector concatenation operation. Indicates element-wise multiplication; Constructing image-text difference seed vectors : ; in, Used to characterize the differences between the image and text samples to be detected. Used to characterize the textual and graphical differences of the candidate reference sample itself. Used to characterize the differences in text-image interaction features between the text-image sample to be detected and the candidate reference sample; The image features, text features, and image-text interaction features of the image-text sample to be detected are used to identify the image features, text features, and image-text interaction features. for Image features, text features, and graphic-text interaction features; calculate The corresponding hyperedge representation vector : ; It is a linear mapping matrix. This indicates a normalization operation; For candidate reference sample set All candidate reference samples By performing vector expansion, we obtain the hyperedge set. : ; remember Candidate reference sample The corresponding hyperedge, where They are represented as sample nodes. Related entities, attributes, and risk types are represented as related semantic nodes, and hyperedges Used to connect sample nodes and related semantic nodes; both sample nodes and related semantic nodes are hypergraph nodes.
7. A method for detecting false text and image information in smart elderly care terminals according to claim 6, characterized in that, Also includes: introduction Redundancy removal constraints between hyperedges, for any two hyperedges and Define hyperedge similarity : ; like Then it is considered that the super-edge and There are redundant super-edges in the data. The redundant super-edges are removed according to the preset clipping strategy. Finally, we obtain the adaptive hyperedge set. ;in, This represents the hyperedge similarity threshold. This indicates an over-edge clipping operation.
8. A method for detecting false text and image information in smart elderly care terminals according to claim 6, characterized in that, The construction of the node set and node-hyperedge association matrix, and the propagation calculation of the hypergraph, yields the risk association strength between each node and the text / graph sample to be detected, specifically including: Build a node set Node-hyperedge incidence matrix The node set includes entity nodes, attribute nodes, and risk type nodes; The weight of each hyperedge is jointly determined by multimodal weighted similarity, hyperedge internal consistency, graph-text conflict intensity, and risk node activation intensity: ; Candidate reference sample Corresponding hyperedge The weight, For multimodal weighted similarity, This represents the activation function. As weight; Hyperedge internal consistency for: ; in, Represents a node Feature representation, Indicates the superedge The correlation cardinality is used to normalize the internal consistency of the hyperedge; Intensity of conflict between images and text for: ; Risk node activation intensity for: ; in, This represents a smoothing term used to avoid a denominator of zero. This represents the set of nodes corresponding to the risk types in the hypergraph; Hyperedge weight matrix for: ; in Describes the constructor for a diagonal matrix; Define the diagonal elements of the node degree matrix : ; Represents the node-hyperedge incidence matrix Middle node With super-edge The relationship between them; Define the diagonal elements of the hyperedge degree matrix : ; Construct the hypergraph propagation matrix based on the node-hyperedge association matrix, hyperedge weight matrix, node degree matrix, and hyperedge degree matrix. : ; For transpose; Using the image / text sample node to be detected as the initial activated node, construct the initial propagation state vector for graph propagation: ; in, The initial state vector for graph propagation indicates that the node of the graph sample to be detected is set as the propagation source point; the other nodes initially have no propagation signal. Then iterative propagation is carried out: ; in, Indicates the current iteration number. For the propagation coefficient, Indicates the first The node state of the step, Indicates by The updated next node propagation state vector will be passed through The propagation state vector obtained after the next iteration of propagation As a representation node The strength of the risk association between the image and text sample to be detected.
9. A method for detecting false text and image information in smart elderly care terminals according to claim 8, characterized in that, The process of calculating and sorting the strength of superedge risk evidence based on the risk correlation strength, filtering the key risk evidence superedge set, and obtaining the enhanced reference sample set through reverse indexing specifically includes: For each superedge Taking into account the weight of the hyperedge, the risk association strength of the nodes inside the hyperedge, and the consistency within the hyperedge, the risk evidence strength of the hyperedge is calculated. : ; in, Indicates the superedge The weight, Indicates consistency within the hyperedge; For the set of superedges All hyperedges are sorted in descending order of risk evidence strength, and the M hyperedges with the highest risk evidence strength are selected as the set of key risk evidence hyperedges. According to the customs The candidate reference samples corresponding to the inverted index are used to obtain the enhanced reference sample set. .
10. A method for detecting false text and image information for smart elderly care terminals according to claim 1, characterized in that, The process involves inputting the image / text sample to be detected, the set of key risk evidence hyperedges, and the set of enhanced reference samples into the visual large language model, and outputting the authenticity judgment result and risk type, specifically including: The image and text sample to be detected Enhance the reference sample set and the super-edge set of key risk evidence Both are input into the visual large language model, which performs comprehensive reasoning based on the image semantics, text semantics, image-text consistency relationship, and knowledge conflict relationship between the image-text sample to be detected and the enhanced reference sample. Finally, it outputs the authenticity judgment result, risk type, and judgment basis of the image-text sample to be detected: ; in, Represents a visual large language model. This indicates the result of determining whether the image or text sample to be tested is real or fake. This indicates the risk type corresponding to the image / text sample to be tested. Indicates the basis for judgment.