Social function identification method and system based on field sensing and semantic conflict detection

By combining field awareness and semantic conflict detection methods with multi-view coding and anchor point constraints, the problems of semantic conflict misjudgment and uncontrollable output in social media content are solved. This achieves highly deterministic and interpretable social function recognition, which is applicable to multimodal content understanding and content risk control.

CN121998630APending Publication Date: 2026-05-08QINGDAO TECHCAL UNIV QINDAO COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO TECHCAL UNIV QINDAO COLLEGE
Filing Date
2026-01-27
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as failure in multimodal semantic conflict recognition, lack of social context, and uncontrollable output of generative models when processing social media content. In particular, they are difficult to achieve high deterministic and interpretable recognition when dealing with satirical content.

Method used

By combining field perception and semantic conflict detection methods with scene depth perception, expert collaborative reasoning, and anchor point stabilization decision-making, and employing multi-view encoding, adaptive gating fusion, neural routing networks, and Logits bias intervention algorithms, a structured semantic decision vector is generated and anchor point matching and lexical probability correction are performed to ensure the determinism and interpretability of the output.

Benefits of technology

It achieves highly deterministic and interpretable identification of social content, significantly reduces the false positive rate, improves the accuracy and stability of multimodal content understanding, and meets the requirements of industrial-grade content risk control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998630A_ABST
    Figure CN121998630A_ABST
Patent Text Reader

Abstract

The invention provides a social function identification method and system based on field sensing and semantic conflict detection, and relates to the technical field of multi-modal content understanding, computational social science and content function identification. The method aims at solving the problems that in the prior art, multi-modal semantic conflict recognition fails, social field domain contexts are lost, and generative model input is uncontrollable. According to the method, to-be-recognized social media content is acquired, multi-view coding and self-adaptive gating fusion are performed on the to-be-recognized social media content, different expert agents are called for reasoning through a neural routing network based on fusion feature vectors and gating weights, semantic judgment vectors are generated, the semantic judgment vectors are matched with a pre-constructed social function anchor point library, a target anchor point is determined, and the social media content is identified. And performing offset correction on the vocabulary probability related to the target anchor point by adopting a Logits offset dry pre-algorithm, and outputting a function label and an interpretation text. According to the method, the problems in the prior art are solved, and automatic and high-certainty identification required by industrial-grade content risk control is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of multimodal content understanding, computational social science, and content function recognition technology, and particularly relates to a method and system for social function recognition based on field perception and semantic conflict detection. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] With the deep penetration of social media platforms, user-generated content exhibits highly multimodal, context-dependent, and rapidly evolving characteristics. Typical social media content often simultaneously contains images, text, and even information such as emojis, hashtags, and repost / comment relationships. Its true meaning depends not only on the content itself but also on the complex social interaction context, including the community section where it is posted, user identity attributes, platform culture, and popular discourse. Therefore, achieving automated and highly deterministic identification of the social functions carried by social content (such as information dissemination, emotional catharsis, group identity, irony / dark humor, etc.) has become a key technology in the fields of social computing and content moderation.

[0004] However, existing technologies still have certain shortcomings when dealing with social content such as satirical images and text: First, existing models mainly rely on shallow feature splicing mechanisms. When faced with scenarios where the meaning of images and text is inconsistent, such as memes, satire, or metaphors, they focus on semantic alignment and lack a structured perception of the social context, making it difficult to capture cross-modal emotional opposition, leading to frequent misjudgments of deep semantic conflicts. Second, existing methods often ignore key contextual features such as community culture, user personas, and public opinion atmosphere, causing semantic understanding to be divorced from the actual communication scenario, resulting in the model making one-sided inferences without contextual anchors. Third, visual language large models (VLMs) based on probability sampling have problems such as output uncertainty, semantic illusion, and uninterpretability, resulting in the inability to reproduce the results of the same input, making it difficult to meet the strict requirements of industrial-grade content risk control and speech analysis for result certainty and compliance.

[0005] In summary, existing technologies suffer from problems such as failure in multimodal semantic conflict recognition, lack of social context, and uncontrollable input to generative models. Summary of the Invention

[0006] To overcome the shortcomings of the existing technologies, this invention provides a social function recognition method and system based on field perception and semantic conflict detection. By combining scene depth perception, expert collaborative reasoning and anchor point stable decision-making, it solves the problems of semantic conflict misjudgment, lack of context and uncontrollable output faced by the existing technologies when processing social content, and realizes the automated and highly deterministic recognition required for industrial-grade content risk control.

[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of this invention provides a method for social function recognition based on field awareness and semantic conflict detection, comprising: Acquire the social media content to be identified, which includes original images, text, and social interaction context information; The social media content to be identified is subjected to multi-view encoding and adaptive gating fusion to obtain fused feature vectors and gating weights; Based on the fusion feature vector and gating weights, the neural routing network calls different expert agents in the hybrid expert module to perform inference, and generates semantic decision vectors based on the inference results. The semantic decision vector is matched with a pre-built social function anchor library to determine the target anchor and its corresponding function category index. The target anchor point is input as a constraint into the decoding stage of the visual language model, and the Logits bias intervention algorithm is used to bias and correct the probability of words related to the target anchor point. Under the constraints of the target anchor point, output function labels and explanatory text.

[0008] As one implementation method, multi-view encoding and adaptive gating fusion are performed on the social media content to be identified. The specific process is as follows: The original image is input into a visual encoder to extract visual features; The text is input into a text encoder to extract text features; The social interaction context information is input into the context encoder to extract social interaction context features; Visual features, text features, and social interaction field features are input into an adaptive gating system for fusion, resulting in a fused feature vector and gating weights.

[0009] As one implementation method, social interaction field features are extracted, including explicit circle vectors, interaction profile vectors, and state atmosphere vectors. The specific process is as follows: Embed and map topic tags or super topic IDs to generate explicit circle vectors; if missing, select the general default vector. Based on the publisher's historical content, an interaction profile vector is obtained; We use retrieval enhancement generation techniques to obtain popular comment summaries of similar content, and encode them to obtain a state atmosphere vector.

[0010] As one implementation method, in the batch training of adaptive gating, a supervised contrastive loss function is constructed to separate and optimize hard negative samples that meet the conditions of visual or text similarity being higher than a preset threshold and having opposite social function labels.

[0011] As one implementation method, the supervised comparison loss function formula is: ; in, This is a temperature coefficient used to control the model's attention to difficult samples; Indicates the sample index in the current batch; For the sample fused feature vectors; Let be the fused feature vector of sample j; Let p be the fused feature vector of sample p; Let i be the set of positive samples. The number of positive samples; The set of candidate samples to be compared; This represents the candidate sample index.

[0012] As one implementation, the different expert agents in the hybrid expert module include: The role of a direct expert agent is defined as describing visual elements of an image and transcribing text content. Conflict expert agents are defined as agents who detect emotional or semantic conflicts between images and text. Field expert agents are defined as those who interpret the underlying meaning of content based on social circles and the publisher's persona.

[0013] As one implementation method, inference is performed by calling different expert agents in the hybrid expert module through a neural routing network. The specific process is as follows: Based on the fused feature vector and gating weights, the credibility weight vector of each expert agent is calculated; Calculate the activation probability of each expert agent based on the credibility weight vector; The corresponding expert agent is invoked for inference only when the activation probability of the expert agent is greater than the preset activation threshold.

[0014] As one implementation method, the Logits bias intervention algorithm is used to bias and correct the probabilities of words related to the target anchor point. The specific process is as follows: Calculate the semantic drift between the current decoded hidden state and the target anchor point; Based on semantic drift and preset intervention intensity, the probability of words related to the target anchor is biased and corrected to generate a biased probability distribution.

[0015] As one implementation method, the probability distribution formula after bias correction is: ; in, For the intervention intensity hyperparameter, =5.0; For only with locked anchor points Relevant security vocabulary, i.e., a vocabulary related to the target anchor point; For the original Logits; This corresponds to the semantic drift degree; These are the candidate output words at time t during the decoding phase; Candidate output words after bias correction The probability of.

[0016] A second aspect of the present invention provides a social function recognition system based on field awareness and semantic conflict detection, comprising: The perception and feature fusion module is used to acquire the social media content to be identified, which includes the original image, text and social interaction field information; multi-view encoding and adaptive gating fusion are performed on the social media content to be identified to obtain the fused feature vector and gating weight; The cognitive reasoning module is used to perform reasoning by calling different expert agents in the hybrid expert module through a neural routing network based on fused feature vectors and gating weights, and to generate semantic decision vectors based on the reasoning results. The expression output module is used to perform similarity matching between the semantic decision vector and the pre-built social function anchor library to determine the target anchor and its corresponding function category index; the target anchor is input as a constraint condition into the decoding stage of the visual language model, and the Logits bias intervention algorithm is used to bias the probability of words related to the target anchor; under the constraint of the target anchor, the function label and explanatory text are output.

[0017] The above one or more technical solutions have the following beneficial effects: In this embodiment, to address the output instability and illusion issues of the General Visual Language Model (VLM), a three-layer control mechanism for stable output generation with anchor constraints is designed. First, a neural routing network schedules role-based expert agents to perform interpretable collaborative reasoning, generating a structured semantic decision vector. Second, this vector is matched with a pre-built social function anchor library to lock in the target category. Finally, a Logit bias intervention algorithm is introduced during the decoding stage to bias the word probability distribution when preset trigger conditions are met, thereby constraining the generation trajectory to converge towards the conclusion corresponding to the target anchor. This mechanism transforms the open-ended reasoning of generative models into stable outputs with deterministic boundaries, reproducible results, and traceable decision-making processes, fully meeting the stringent compliance and reliability standards of content risk control and speech analysis.

[0018] In this embodiment, the present invention defines social interaction field features, and encodes explicit circles, user profiles, and simulated atmospheres into feature vectors in a multi-dimensional and structured manner, and performs adaptive gating fusion with text and image content. This enables the model's reasoning process to have real contextual anchors, and can understand the real social function of the content in combination with specific dissemination scenarios, avoiding one-sided interpretations that are out of context.

[0019] In this embodiment, multimodal representation extraction with enhanced field is achieved through multi-view encoding and adaptive gating in the perception layer. Furthermore, a feature space capable of differentiating samples with similar images and text but opposite intentions is constructed through supervised training and optimization of the contrastive loss function. In particular, a hard negative sample mining strategy is employed to optimize these difficult sample pairs during training, enabling the model to fundamentally master the ability to identify cross-modal semantic conflicts, thereby significantly reducing the misjudgment rate of memes, irony, and other content.

[0020] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0021] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0022] Figure 1 This is a schematic diagram of the overall framework of the social function recognition method based on field perception and semantic conflict detection according to Embodiment 1 of the present invention. Detailed Implementation

[0023] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0024] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0025] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0026] Example 1 This embodiment discloses a social function recognition method based on field awareness and semantic conflict detection.

[0027] To more clearly illustrate this embodiment, the implementation process of social function recognition based on field awareness and semantic conflict detection can be specifically described as follows: Social function recognition methods based on field awareness and semantic conflict detection include: S1. Obtain the social media content to be identified, including the original image, text, and social interaction field information; The social media content to be identified is subjected to multi-view encoding and adaptive gating fusion to obtain fused feature vectors and gating weights; S2. Based on the fused feature vector and gating weights, different expert agents in the hybrid expert module are invoked through the neural routing network to perform inference, and semantic decision vectors are generated based on the inference results. S3. Perform similarity matching between the semantic decision vector and the pre-built social function anchor point library to determine the target anchor point and its corresponding function category index; The target anchor point is input as a constraint into the decoding stage of the visual language model, and the Logits bias intervention algorithm is used to bias and correct the probability of words related to the target anchor point. Under the constraints of the target anchor point, output function labels and explanatory text.

[0028] like Figure 1 As shown, in step S1, the social media content to be identified is obtained, wherein the social media content includes the original image, text and social interaction field information; The social media content to be identified is subjected to multi-view encoding and adaptive gating fusion to obtain fused feature vectors and gating weights.

[0029] S101. Obtain the social media content to be identified, wherein the social media content includes original images, text, and social interaction context information.

[0030] (1) Construct a social function recognition model based on field perception and semantic conflict detection.

[0031] A three-layer recognition framework consisting of a perception layer, a cognition layer, and an expression layer is constructed.

[0032] The perception layer includes a multimodal input module, multi-view encoding, and adaptive gating, which in turn obtains a complete input representation that includes explicit and implicit social context. It also forces a distance between literally similar but functionally similar samples in the feature space to solve the problem of reverse meaning representation.

[0033] The cognitive layer includes a neural routing network, a hybrid expert module, and a text encoder. It utilizes the feature signals and weights extracted from the perception layer to intelligently schedule frozen large visual language model (VLM) expert agents, thus solving the problem of confused thinking in a single model.

[0034] The expression layer includes a social function anchor library and a decoder for the visual language model (VLM), which transforms the open-ended reasoning of generative models into deterministic classification results, thus solving the problems of illusion and unstable output.

[0035] Through the above steps, a three-layer progressive architecture of perception-cognition-expression is obtained, which systematically solves the problems of failure in multimodal semantic conflict recognition, lack of social context, and uncontrollable input of generative models in existing technologies for social content recognition. It not only significantly improves the accuracy of social function recognition of complex multimodal content (especially ironic and metaphorical content), but also realizes the contextualization of the analysis process, the stabilization of decision results, and the interpretability of output logic, realizing the automated and highly deterministic recognition required for industrial-grade content risk control.

[0036] (2) Obtain the social media content to be identified.

[0037] In the perception layer, the social media content triples to be identified are obtained through the multimodal input module, using the following formula: (1) in, Original image; The text content to be published, i.e., the text itself; This refers to the social interaction context information contained within the surrounding text.

[0038] Contextual information about the social interaction field It consists of three parts: a summary of popular comments on the topic tag or super topic ID, the poster's historical content, and similar content obtained using search enhancement generation technology.

[0039] S102. Perform multi-view encoding and adaptive gating fusion on the social media content to be identified to obtain the fused feature vector and gating weights.

[0040] (1) Multi-view encoding of the social media content to be identified.

[0041] 1) Input the original image into the visual encoder to extract visual features.

[0042] Visual features extracted using ViT-Base (patch size = 16) .

[0043] 2) Input the text into the text encoder and extract the text features.

[0044] Extracting text features using RoBERTa (Chinese) .

[0045] 3) Input the social interaction context information into the context encoder to extract social interaction context features. .

[0046] It is composed of three parts: explicit circle vector, interaction profile vector, and state atmosphere vector, to obtain the characteristics of the social interaction field. The specific process is as follows: First, embed and map the topic tags or super topic IDs to generate explicit circle vectors. If they are missing, select the general default vector.

[0047] Embedding mapping is performed on hashtags or super topic IDs to obtain explicit circle vectors. If missing, a learnable general default vector is used. .

[0048] Secondly, based on the publisher's historical content, an interaction profile vector is obtained.

[0049] Extract the mean of the text features from N historical posts by the publisher, and use this as the interaction profile vector. The formula is: (2) Finally, the popular comment summaries of similar content are encoded to obtain the state atmosphere vector.

[0050] Using Retrieval Enhancement Generation (RAG) technology, retrieve typical comment summaries of Top-K similar popular posts. Encoding yields the pseudo-atmosphere vector The formula is: (3) (2) Input visual features, text features and social interaction field features into an adaptive gating system for fusion to obtain a fused feature vector and gating weights.

[0051] The adaptive gating formula is: ; The gating weights and fused feature vectors are given by the following formulas: ; (3) in, For learnable gating parameters, For the sample The fused feature vector For the sample Gating weights, For the sample Visual features For the sample Textual features, For the sample Social interaction and interaction characteristics , , Samples Gating weights for visual features, text features, and social interaction context features.

[0052] (3) In the batch training of adaptive gating, a supervised contrastive loss function is constructed to separate and optimize hard negative samples that meet the conditions of visual or text similarity being higher than a preset threshold and opposite social function labels.

[0053] First, in batch training, priority is given to sampling sample pairs that simultaneously meet the following conditions. : 1) Visual or textual similarity is higher than a preset threshold.

[0054] The visual / text similarity is extremely high, that is It's like a meme.

[0055] 2) The social function tags are opposite.

[0056] One is self-deprecation, the other is attack.

[0057] Then, the supervised contrastive loss function is constructed, with the following formula: (4) in, This is a temperature coefficient used to control the model's attention to difficult samples; Indicates the sample index in the current batch; For the sample fused feature vectors; Let be the fused feature vector of sample j; Let p be the fused feature vector of sample p; Let i be the set of positive samples. The number of positive samples; The set of candidate samples to be compared; This represents the candidate sample index.

[0058] =0.07, the smaller the value, the more attention is paid to difficult samples, which can effectively widen the distance between hard negative samples that are similar in image and text but have opposite intentions.

[0059] After the above steps, hard negative samples are mined, and the loss function forces the model to push away hard negative samples that are synonymous with each other in the feature space.

[0060] like Figure 1As shown, in step S2, based on the fused feature vector and gating weights, different expert agents in the hybrid expert module are invoked through the neural routing network to perform inference, and a semantic decision vector is generated based on the inference results.

[0061] S201. Based on the fused feature vector and gating weights, inference is performed by calling different expert agents in the hybrid expert module through a neural routing network.

[0062] (1) Load three sets of predefined role prompts onto the same visual language model (VLM) to obtain three expert agent outputs with different reasoning behaviors, namely, literal, conflict and field, which together constitute a hybrid expert module.

[0063] Specifically, in the hybrid expert module, multiple sets of role-based prompts are pre-constructed for different analysis dimensions, and the prompts are loaded onto the same visual language model (VLM) to generate multiple inference outputs for the same input sample. Each inference output is defined as a different expert agent to analyze social media content from three dimensions: literal understanding, semantic conflict detection, and field context interpretation, thus forming the hybrid expert module.

[0064] Among them, the three expert agents are: 1) Straightforward expert agents, whose role is defined as describing visual elements of images and transcribing text content.

[0065] Specifically, for the direct expert agent E1, the system prompt word is: "You are an objective visual recorder. Tasks: 1. Describe the visual elements of an image in detail; 2. Transcribe the text verbatim. Speculation of intent or the use of subjective terms such as 'sarcasm' is strictly prohibited. Output only objective facts." 2) Conflict expert agent, whose role is defined as detecting emotional or semantic conflicts between images and text.

[0066] Specifically, the conflict expert agent E2 received the following system prompt: "You are a multimodal conflict detective. Task: 1. Determine visual emotion (e.g., sadness); 2. Determine textual emotion (e.g., joy); 3. Detect conflict: Are the emotions of the image and text contradictory? Is there a sense of absurdity caused by 'image and text not matching'? Output a conflict intensity assessment." 3) Field expert agents, whose role is defined as interpreting the potential meaning of content based on social circles and the publisher's persona.

[0067] Specifically, the system prompt for the field expert agent E3 is: "You are a social network cultural anthropologist. Input context: social circles." Character design Task: Explain what kind of meme this image and text usually represents within this circle. Based on the persona, determine whether it's a genuine statement or a sarcastic remark. (2) Neural routing network scheduling expert agent, the specific process is as follows: 1) Calculate the credibility weight vector of each expert agent based on the fused feature vector and gating weight.

[0068] The neural routing network receives the fused feature vector output from the perceptual layer. and gating weights .

[0069] Input-based fusion feature vector and gating weights Calculate the credibility weight vector of each expert in the current sample. The formula is: (5) in, For the sample The fused feature vector For the sample Gating weights, , , These are the credibility weights for the semantic expert agent E1, the semantic conflict expert agent E2, and the field expert agent E3, respectively.

[0070] If the gating weight show >0.7, or fused feature vectors In the feature space, where the distance to the ironic anchor point is very close, the neural routing network will automatically and significantly increase the credibility weight of the field expert. ) and the credibility weight of conflict experts ( The value of ) is reduced, while the credibility weight of the intuitive expert is lowered ( The value of ).

[0071] 2) Calculate the activation probability of each expert agent based on the credibility weight vector.

[0072] The formula for calculating the expert activation probability is: (6).

[0073] 3) The corresponding expert agent is invoked for inference only when the activation probability of the expert agent is greater than the preset activation threshold.

[0074] Set activation threshold Only when the activation probability of a certain expert... Only when the condition is met will the corresponding expert agent in the hybrid expert module (a role-based expert agent based on VLM) be invoked for inference; otherwise, its output will be set to zero.

[0075] S202. Based on the reasoning results, generate semantic decision vectors.

[0076] (1) Based on the credibility weight vector and the comprehensive reasoning results of various experts, a contextual adjudication mechanism is adopted to generate a natural language form of judgment summary.

[0077] The referee agent uses the confidence weight vector output by the neural routing network. Based on the combined reasoning of various experts, a context-based adjudication mechanism is employed to generate a natural language summary of the judgment. .

[0078] (1) The decision summary in natural language form is input into the text encoder for encoding to obtain the semantic decision vector.

[0079] The semantic decision vector is obtained through the text encoder. The formula is: (7).

[0080] After the above steps, a semantic decision vector is obtained, which encapsulates the complete logic of multimodal reasoning and provides a data foundation for subsequent anchor point matching.

[0081] like Figure 1 As shown, in step S3, the semantic decision vector is matched with the pre-built social function anchor library to determine the target anchor and its corresponding function category index. The target anchor point is input as a constraint into the decoding stage of the visual language model, and the Logits bias intervention algorithm is used to bias and correct the probability of words related to the target anchor point. Under the constraints of the target anchor point, output function labels and explanatory text.

[0082] (1) Input the semantic decision vector into the visual language model and match it with the pre-built social function anchor library to determine the target anchor.

[0083] 1) Pre-construct a social function anchor point library.

[0084] Feature centers of K standard social function categories (e.g., political satire, group identity, emotional catharsis) are extracted beforehand from the labeled dataset, and an anchor point matrix is ​​constructed. Specifically, The anchor point library is built offline, using the same feature extraction / encoding process as steps S1 and S2. High-dimensional representations (preferably semantic decision vectors) are generated for matching from each sample in the manually labeled dataset. Then, statistical centers (e.g., mean vectors) are calculated for each of the K social function categories to form an anchor point matrix. This yields the social function anchor point library, as shown in Table 1.

[0085] Table 1 Examples of Social Function Anchor Point Database anchor_id function_label description_rationale anchor_vector_centroid_preview A_01 Sarcastic Venting Using mismatched images and text, exaggeration, or irony to express dissatisfaction, pressure, or criticism usually carries a negative emotional core. "[0.821 -0.314 0.452... 0.127]" A_02 Political Satire To comment on and criticize political figures, events, or policies using humor, metaphor, or satire. "[-0.253 0.791 -0.118... 0.054]" A_03 DailyShare (Life Log) It involves sharing everyday events, objects, or experiences objectively and frankly, without any hidden intentions. "[-0.882 -0.021 -0.055 ... 0.103]" A_04 In-group Bonding (Slang) The use of memes, emojis, or language that are only understood within a specific circle aims to strengthen internal connections and promote exclusivity. "[0.512 0.487 0.229... -0.301]" A_05 Emotional Resonance Expressing negative emotions such as sadness, anxiety, or loneliness directly aims to seek understanding, comfort, and empathy from others. "[0.156 -0.258 0.853... 0.022]" A_06 Humor and comedy (Pure Humor) Content posted purely for entertainment and to amuse others, without obvious aggression or deep satire. "[0.334 0.671 0.112... -0.415]" A_07 Objective Statement / News When publishing news reports, official announcements, or scientific facts, maintain a neutral stance and avoid personal bias. "[-0.501 0.109 -0.203... 0.758]" A_08 Hostile Provocation Intentionally publishing highly provocative, insulting, or discriminatory content with the aim of inciting conflict and confrontation. "[0.605 0.702 -0.308... 0.159]" A_09 Healing & Encouragement We publish warm, positive, and uplifting content, aiming to bring hope and psychological comfort to people. "[-0.307 -0.405 0.601... 0.206]" A_10 Marketing / Promotion Content with a clear commercial purpose, intended to promote products, services, or brands. "[-0.122 -0.554 -0.109 ... -0.673]" 2) Match the semantic decision vector with the pre-built social function anchor library.

[0086] Calculate the cosine similarity between the semantic decision vector and the feature center vectors in the social function anchor library, and determine the matching target anchor and its category index from the social function anchor library based on the similarity.

[0087] Specifically, the semantic decision vector is calculated. With all K feature centers in the anchor point library The cosine similarity directly identifies the target anchor point with the highest similarity. and its corresponding index The formula is: (8).

[0088] (2) Based on the target anchor point, the Logits bias intervention algorithm is used to bias the probability of words related to the target anchor point.

[0089] A semantic snapping term is introduced when the Visual Language Model (VLM) decoder decodes and generates the final label and explanatory text.

[0090] The specific process of the Logit bias intervention algorithm is as follows: 1) Calculate the semantic drift between the current decoded hidden state and the target anchor point.

[0091] The cosine similarity between the current generated state and the nearest social functional anchor is calculated using the following formula: (9) in, This represents the current hidden decoding state.

[0092] 2) Based on semantic drift and preset intervention intensity, bias correction is applied to the probability of words related to the target anchor point to generate a bias-corrected probability distribution.

[0093] The formula for the probability distribution after bias correction is: (10) in, For the intervention intensity hyperparameter, =5.0; For only with locked anchor points Relevant security vocabulary, i.e., a vocabulary related to the target anchor point; For the original Logits; This corresponds to the semantic drift degree; These are the candidate output words at time t during the decoding phase; Candidate output words after bias correction The probability of.

[0094] Intervention intensity hyperparameter Only if the semantic drift between the current decoded hidden state and the target anchor point It will only start at that time. Intervene to avoid forcibly classifying things when the semantics are extremely ambiguous.

[0095] Vocabulary related to target anchors Only load and lock anchor points A specific list of related safety words is used. For example, if "irony" is locked, only words such as "sarcasm" and "irony" will be rewarded, while interfering words such as "happy" and "praise" will be blocked.

[0096] Through the above steps, the probability of words related to the target social function is increased, irrelevant words are suppressed, and the generated trajectory is forced to "attach" to the predetermined track.

[0097] (3) Generate functional labels and explanatory text under the constraints of the target anchor point after bias correction.

[0098] The formula for the function label is: (11) Natural language explanations are generated under the bias-corrected target anchor constraints to restate the judge's reasoning logic.

[0099] The multimodal sarcasm detection dataset (data-of-multimodal-sarcasm-detection) was used as the test set. The social function recognition model based on field awareness and semantic conflict detection was tested according to steps S1-S3. The model was evaluated by measuring the accuracy, precision, recall, and F1 score of the inference five times. The results were compared with the baseline model, as shown in Table 2.

[0100] Table 2 Comparison results between this model and the baseline model Group Model Name Accuracy Precision Recall F1-Score Output consistency Plain text detection methods BERT(bert-base-multilingual-cased) 0.89 0.87 0.85 0.86 100% Multimodal fusion detection method A Thousand Questions on General Principles (qwen3-vl-plus) 0.75 0.74 0.78 0.76 94% This article's method Ours 0.92 0.92 0.90 0.91 99.8% As shown in Table 2, the method in this embodiment comprehensively surpasses mainstream baseline methods in key performance indicators, demonstrating the significant superiority of its architecture. Compared with pure text detection methods (such as BERT), the method in this embodiment, after fusing multimodal information, not only maintains extremely high output consistency (99.8% vs 100%), but also achieves significant improvement in various recognition accuracy indicators (such as F1 increasing from 0.86 to 0.91), proving that it can effectively utilize complementary information between text and images to overcome the inherent limitations of pure text models in understanding visual irony, etc. Compared with general multimodal fusion methods (such as Tongyi Qianwen VL), the method in this embodiment achieves a significant lead in both accuracy and stability (F1 increasing from 0.76 to 0.91, consistency increasing from 94% to 99.8%). This verifies that the field perception, expert collaboration, and anchor point constraint mechanisms designed in this embodiment can systematically overcome the shortcomings of traditional multimodal models in semantic conflict recognition and generation stability. In summary, this model achieves a breakthrough in both accuracy and reliability in the multimodal social content understanding task.

[0101] As can be seen, tests on authoritative multimodal irony detection datasets demonstrate that the model in this embodiment shows significant and substantial improvements in recognition accuracy, output stability, and decision interpretability compared to existing baseline methods. This fully proves the effectiveness of the technical approach of this embodiment, which addresses contextual gaps through field awareness, resolves conflict misjudgments through expert collaboration, and addresses output illusions through anchor point constraints. This provides a solid technical foundation for building a next-generation highly reliable social content analysis system.

[0102] An ablation experiment was conducted on the social function recognition model based on field awareness and semantic conflict detection in this example, as shown in Table 3.

[0103] Table 3 Ablation Experiment Results Model type Accuracy Precision Recall F1-Score Output consistency Full (S1+S2+S3) 0.92 0.92 0.90 0.91 99.8% w / o S3 (S1+S2) 0.91 0.91 0.90 0.90 95% w / o S2 (S1+S3, single expert) 0.88 0.87 0.86 0.86 99.5% w / o S1-key (no field + S2 + S3) 0.86 0.86 0.84 0.85 99.5% As shown in Table 3, through ablation experiments, the perception-cognition-expression three-layer architecture in this embodiment is a complementary and synergistically beneficial organic whole. The structured field features introduced by the perception layer are the cornerstone for the model to achieve accurate semantic understanding; their absence leads to the largest performance drop. The multi-expert collaborative reasoning mechanism of the cognition layer significantly improves the ability to resolve complex conflicts and is key to improving accuracy. While the anchor constraint and intervention algorithm of the expression layer have a relatively small impact on absolute accuracy, they are the core technical guarantee for transforming the random output of the generative model into deterministic results and meeting the stability requirements of industrial-grade risk control. The complete model achieves optimal accuracy, F1 score, and output consistency, proving that this integrated scheme effectively solves the three major problems of missing context, misjudgment of conflicts, and uncontrollable output in existing technologies.

[0104] Example 2 The purpose of this embodiment is to provide a social function recognition system based on field awareness and semantic conflict detection, including: The perception and feature fusion module is used to acquire the social media content to be identified, which includes the original image, text and social interaction field information; multi-view encoding and adaptive gating fusion are performed on the social media content to be identified to obtain the fused feature vector and gating weight; The cognitive reasoning module is used to perform reasoning by calling different expert agents in the hybrid expert module through a neural routing network based on fused feature vectors and gating weights, and to generate semantic decision vectors based on the reasoning results. The expression output module is used to perform similarity matching between the semantic decision vector and the pre-built social function anchor library to determine the target anchor and its corresponding function category index; the target anchor is input as a constraint condition into the decoding stage of the visual language model, and the Logits bias intervention algorithm is used to bias the probability of words related to the target anchor; under the constraint of the target anchor, the function label and explanatory text are output.

[0105] The perception and feature fusion module is the perception layer, which includes a multimodal input module, multi-view encoding and adaptive gating, thereby obtaining a complete input representation that includes explicit and implicit social context, and forcibly widening the distance between literally similar but functionally similar samples in the feature space to solve the problem of reverse meaning representation.

[0106] The cognitive reasoning module is the cognitive layer, which includes a neural routing network, a hybrid expert module, and a text encoder. It uses the feature signals and weights extracted by the perception layer to intelligently schedule the frozen large visual language model (VLM) expert agent to solve the problem of confused thinking in a single model.

[0107] The expression output module is the expression layer, which includes a social function anchor library and a decoder for the visual language model (VLM). It transforms the open reasoning of generative models into deterministic classification results, solving the problems of hallucination and unstable output.

[0108] Based on the provision of a social function recognition system based on field awareness and semantic conflict detection, the method steps in Embodiment 1 are implemented.

[0109] Example 3 The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the above-described method.

[0110] Example 4 The purpose of this embodiment is to provide a computer-readable storage medium.

[0111] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the above method.

[0112] Example 5 The purpose of this embodiment is to provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods and functions involved in any of the above embodiments.

[0113] The steps and methods involved in the apparatus of the above embodiments correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0114] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0115] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A social function recognition method based on field awareness and semantic conflict detection, characterized in that, include: Acquire the social media content to be identified, which includes original images, text, and social interaction context information; The social media content to be identified is subjected to multi-view encoding and adaptive gating fusion to obtain fused feature vectors and gating weights; Based on the fusion feature vector and gating weights, the neural routing network calls different expert agents in the hybrid expert module to perform inference, and generates semantic decision vectors based on the inference results. The semantic decision vector is matched with a pre-built social function anchor library to determine the target anchor and its corresponding function category index. The target anchor point is input as a constraint into the decoding stage of the visual language model, and the Logits bias intervention algorithm is used to bias and correct the probability of words related to the target anchor point. Under the constraints of the target anchor point, output function labels and explanatory text.

2. The social function recognition method based on field awareness and semantic conflict detection as described in claim 1, characterized in that, The process of performing multi-view encoding and adaptive gating fusion on the social media content to be identified is as follows: The original image is input into a visual encoder to extract visual features; The text is input into a text encoder to extract text features; The social interaction context information is input into the context encoder to extract social interaction context features; Visual features, text features, and social interaction field features are input into an adaptive gating system for fusion, resulting in a fused feature vector and gating weights.

3. The social function recognition method based on field awareness and semantic conflict detection as described in claim 2, characterized in that, Extracting social interaction field features, which include explicit circle vectors, interaction profile vectors, and state atmosphere vectors, is a specific process as follows: Embed and map topic tags or super topic IDs to generate explicit circle vectors; if missing, select the general default vector. Based on the publisher's historical content, an interaction profile vector is obtained; We use retrieval enhancement generation techniques to obtain popular comment summaries of similar content, and encode them to obtain a state atmosphere vector.

4. The social function recognition method based on field awareness and semantic conflict detection as described in claim 1, characterized in that, In the batch training of adaptive gating, a supervised contrastive loss function is constructed to separate and optimize hard negative samples that meet the conditions of visual or text similarity being higher than a preset threshold and opposite social function labels.

5. The social function recognition method based on field awareness and semantic conflict detection as described in claim 4, characterized in that, The formula for the supervised contrastive loss function is: ; in, This is a temperature coefficient used to control the model's attention to difficult samples; Indicates the sample index in the current batch; For the sample fused feature vectors; Let be the fused feature vector of sample j; Let p be the fused feature vector of sample p; Let i be the set of positive samples. The number of positive samples; The set of candidate samples to be compared; This represents the candidate sample index.

6. The social function recognition method based on field awareness and semantic conflict detection as described in claim 1, characterized in that, The different expert agents in the hybrid expert module include: The role of a direct expert agent is defined as describing visual elements of an image and transcribing text content. Conflict expert agents are defined as agents who detect emotional or semantic conflicts between images and text. Field expert agents are defined as those who interpret the underlying meaning of content based on social circles and the publisher's persona.

7. The social function recognition method based on field awareness and semantic conflict detection as described in claim 1, characterized in that, Inference is performed by calling different expert agents in the hybrid expert module through a neural routing network. The specific process is as follows: Based on the fused feature vector and gating weights, the credibility weight vector of each expert agent is calculated; Calculate the activation probability of each expert agent based on the credibility weight vector; The corresponding expert agent is invoked for inference only when the activation probability of the expert agent is greater than the preset activation threshold.

8. The social function recognition method based on field awareness and semantic conflict detection as described in claim 1, characterized in that, The Logits bias intervention algorithm is used to bias and correct the probabilities of words related to the target anchor point. The specific process is as follows: Calculate the semantic drift between the current decoded hidden state and the target anchor point; Based on semantic drift and preset intervention intensity, the probability of words related to the target anchor is biased and corrected to generate a biased probability distribution.

9. The social function recognition method based on field awareness and semantic conflict detection as described in claim 8, characterized in that, The formula for the probability distribution after bias correction is: ; in, For the intervention intensity hyperparameter, =5.0; For only with locked anchor points Relevant security vocabulary, i.e., a vocabulary related to the target anchor point; For the original Logits; This corresponds to the semantic drift degree; These are the candidate output words at time t during the decoding phase; Candidate output words after bias correction The probability of.

10. A social function recognition system based on field perception and semantic conflict detection, characterized in that, include: The perception and feature fusion module is used to acquire the social media content to be identified, which includes the original image, text and social interaction field information; multi-view encoding and adaptive gating fusion are performed on the social media content to be identified to obtain the fused feature vector and gating weight; The cognitive reasoning module is used to perform reasoning by calling different expert agents in the hybrid expert module through a neural routing network based on fused feature vectors and gating weights, and to generate semantic decision vectors based on the reasoning results. The expression output module is used to perform similarity matching between the semantic decision vector and the pre-built social function anchor library to determine the target anchor and its corresponding function category index; the target anchor is input as a constraint condition into the decoding stage of the visual language model, and the Logits bias intervention algorithm is used to bias the probability of words related to the target anchor; under the constraint of the target anchor, the function label and explanatory text are output.