News image value evaluation method and system based on predicate logic and multi-modal fusion

Through the news image value assessment method based on predicate logic and multimodal fusion, the problems of modal fragmentation and rigid evaluation are solved, multimodal semantic fusion, dynamic adaptive evaluation and interpretability are realized, and the accuracy and efficiency of news value assessment are improved. It is suitable for UGC content management and public opinion monitoring.

CN120654189AInactive Publication Date: 2025-09-16WENKE INFORMATION (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510769301.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing news value assessment technologies have problems such as modal fragmentation, rigid assessment and black-box decision-making, which lead to insufficient cross-modal semantic fusion, poor dynamic adaptability and lack of explainability, and cannot meet the needs of real-time news value judgment.

Method used

A method based on predicate logic and multimodal fusion is adopted. By acquiring multimodal data of news images, key frames are extracted using an adaptive trigger mechanism, and subject-action-scene triples are generated in combination with predicate logic analysis. The news value is calculated through a three-dimensional dynamic evaluation model, and the public opinion factor is introduced to dynamically adjust the weight, providing a visual tuning mechanism.

Benefits of technology

It achieves improved multimodal semantic fusion accuracy, dynamic adaptive evaluation, provides an explainable decision-making process, reduces the workload of manual review, improves evaluation accuracy and efficiency, and is suitable for the screening and grading of massive UGC content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654189A_ABST
    Figure CN120654189A_ABST
Patent Text Reader

Abstract

The invention discloses a news image value evaluation method and system based on predicate logic and multi-modal fusion, and the method comprises the steps: obtaining multi-modal data of a news image, and extracting an audio and visual data key frame through a self-adaptive triggering mechanism; semantic analysis is performed on the key frame through a predicate logic analysis engine, a subject-action-scene triple is generated in combination with text data, and dual verification is performed through spatial position and semantic association; calculating a news value score of the verified triple data based on a three-dimensional dynamic evaluation model; the three-dimensional dynamic evaluation model fuses propagation intensity, information entropy compensation and ontology value; and performing visual adjustment and optimization on evaluation parameters and rules to realize optimization of the dynamic evaluation model. According to the method, massive multi-modal data can be efficiently processed, and the response time to emergencies is greatly shortened; the overall event checking efficiency is improved; the evaluation result is accurate, efficient, transparent and reliable, and the user can understand and trust the model decision conveniently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the interdisciplinary field of artificial intelligence and media technology, and more specifically to a news image value assessment method and system based on predicate logic and multimodal fusion, which is suitable for automated news value assessment, screening, grading and dynamic optimization management of massive news images and videos in the context of user-generated content (UGC). Background Art

[0002] Currently, with the popularity of social media and mobile terminals, UGC news image data has exploded, posing a huge challenge to media organizations' content review and value assessment. Existing news value assessment technologies have the following shortcomings:

[0003] Modal fragmentation: Traditional methods typically process information from different modalities, such as text, images, and audio, separately, lacking unified semantic fusion. This results in cross-modal semantic association errors exceeding 28% (according to data from an ICCV 2022 report). For example, a news article may contain text descriptions and video footage, but existing systems struggle to accurately align the meanings of the two, reducing evaluation accuracy.

[0004] Rigid evaluation: Most existing systems are based on pre-set, static models, making them slow to respond to emergencies and shifts in public opinion. When hot news stories emerge, because the models can't dynamically adjust weights, evaluation results often lag behind actual public opinion changes. Manual evaluation can cause even longer delays, measured in hours. This rigid evaluation mechanism fails to meet the demand for real-time news value assessment.

[0005] Black-box decision-making: Many news value assessments rely on deep learning black-box models, lacking a transparent decision-making process, making the results difficult to interpret. Industry feedback indicates that this lack of explanation has increased manual review workload by approximately 35%, reducing overall efficiency.

[0006] Therefore, with the development of the media environment, there is an urgent need for a news image value assessment method or system that can integrate multimodal data, respond to changes in public opinion in a timely manner, and provide explainable decisions, so as to overcome the many subjective errors caused by workers' experience and judgment. Summary of the Invention

[0007] In view of this, the present invention provides a news image value assessment method and system based on predicate logic and multimodal fusion, aiming to solve technical problems in existing news image value assessment technologies such as insufficient cross-modal semantic fusion, poor dynamic adaptability, and lack of interpretability.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] In a first aspect, an embodiment of the present invention provides a news image value assessment method based on predicate logic and multimodal fusion, comprising the following steps:

[0010] S10, acquiring multimodal data of news images, including associated audio, visual, and text data, and extracting key frames of audio and visual data using an adaptive triggering mechanism;

[0011] S20, semantically analyze the keyframes using a predicate logic parsing engine, and generate subject-action-scene triples based on text data, which are then verified using both spatial location and semantic association.

[0012] S30. Calculate the news value score of the verified triplet data based on a three-dimensional dynamic evaluation model; the three-dimensional dynamic evaluation model integrates communication intensity, information entropy compensation, and ontological value;

[0013] S40: Optimizing the dynamic evaluation model by visually tuning the evaluation parameters and rules.

[0014] Furthermore, in step S10, the adaptive trigger mechanism extracts the audio-visual data key frames, including:

[0015] Detect changes in visual data features by analyzing image mutations through the color histogram differences between adjacent frames, triggering key frame extraction;

[0016] For audio data feature change detection, the energy changes between adjacent audio frames are detected through short-time energy analysis to trigger key frame extraction.

[0017] Furthermore, the step S20 includes:

[0018] S201. Use the pre-trained target detection model and the non-maximum suppression strategy to filter redundant detection boxes and perform subject detection.

[0019] S202. Extract subject motion features through the image sequence optical flow method, feed the visual features into the trained action classification model, and integrate the spatiotemporal attention mechanism to capture the subject's actions;

[0020] S203, using the pre-trained CNN model to perform semantic label judgment on the entire frame; at the same time, performing instance segmentation on the background area in the frame, extracting scene elements, and generating a brief semantic description of the scene;

[0021] S204: Verify the spatial position by calculating whether the spatial overlap between the subject and the scene exceeds a first threshold; and verify the semantic association by using cross-modal semantic analysis to calculate whether the CLIP similarity between the triple text description and the key frame image is greater than a second threshold.

[0022] Furthermore, in step S30, the three-dimensional dynamic evaluation model is represented as follows:

[0023]

[0024] In the formula, I(t), H(t), and S(t) represent the transmission intensity, information entropy compensation, and ontological value, respectively; It reflects the changing trend of event importance over time; coefficients α, β, and γ are used to balance the contribution of each part to the total value score.

[0025] Furthermore, the expression formula of the propagation intensity I(t) is as follows:

[0026] I(t)=D(t)*I channel (t)*H p (t)

[0027] Where D(t) represents the time decay function of the influence of information; when exponentially decaying, D(t) = e -δt or power law decay δ represents the exponential decay coefficient of the influence of information; ε represents the power law decay coefficient of the influence of information; t represents time;

[0028] I channel (t) represents the weighted sum of channel communication volume;

[0029] N represents the total number of information dissemination channels; w i represents the weight of the i-th channel; f i (t) represents the propagation volume of the i-th channel at time t; H p (t) represents the propagation entropy, which is used to measure the uniformity of information distribution among channels;

[0030]

[0031] Among them, p i (t) represents the propagation ratio of the i-th channel at time t, f j (t) represents the propagation volume of the jth channel at time t.

[0032] Furthermore, the information entropy compensation H(t) is expressed as follows:

[0033] Countdown form: Or in logarithmic form:

[0034] Indicates the current event E t Average similarity to historical events;

[0035]

[0036] Sim(·) represents the event similarity calculation, and ∈ represents the zero-proof constant.

[0037] Furthermore, the ontological value S(t) is expressed as follows:

[0038]

[0039] Where w k represents the weight of the k-th event evaluation dimension; ∑w k =1;x k represents the standardized value of the kth event evaluation dimension; M represents the total number of event evaluation dimensions; the evaluation dimensions include: person status, economic impact, social attention, geographical influence range and time type.

[0040] Furthermore, the step S30 further includes:

[0041] Obtain the real-time public opinion factor P(t) associated with the triplet data, and collect event attention indicators through the social media API interface;

[0042] Dynamically adjust the model weight coefficient according to the public opinion factor P(t):

[0043] When P(t)>θ hot When , increase the propagation intensity weight α;

[0044] When P(t)<θ cold When , increase the information entropy compensation weight β;

[0045] Among them, θ hot is the high thermal threshold, θ cold is the low thermal threshold.

[0046] Furthermore, the step S40 includes:

[0047] Build a parameter sandbox to visualize the α, β, and γ weight trajectories, display the historical adjustment path of the model weight coefficients and the real-time changes of the scores;

[0048] Adjust any weight parameter based on natural language rules, compile into machine constraints, and support reverse parsing;

[0049] The logical basis and weight contribution ratio of each score adjustment are recorded to optimize the dynamic evaluation model.

[0050] In a second aspect, an embodiment of the present invention further provides a news image value assessment system based on predicate logic and multimodal fusion, using the news image value assessment method based on predicate logic and multimodal fusion as described in any one of the first aspects, the system comprising:

[0051] The multimodal acquisition module is used to obtain multimodal data of news images, including associated audio, visual and text data, and uses an adaptive trigger mechanism to extract key frames of audio and visual data;

[0052] The predicate logic parsing engine module is used to perform semantic parsing on key frames using the predicate logic parsing engine, generate subject-action-scene triples based on text data, and perform dual verification through spatial location and semantic association;

[0053] A news value scoring module is used to calculate the news value score of the verified triplet data based on a three-dimensional dynamic evaluation model; the three-dimensional dynamic evaluation model integrates communication intensity, information entropy compensation and ontological value;

[0054] The human-machine collaborative optimization module is used to optimize the dynamic evaluation model by visually tuning the evaluation parameters and rules.

[0055] It can be seen from the above technical solutions that compared with the prior art, the present invention has the following technical advantages:

[0056] 1. Multimodal semantic fusion improves accuracy: By jointly analyzing multimodal data such as vision, text, and audio and using predicate logic modeling to achieve deep semantic alignment, the system effectively reduces errors caused by inconsistent information from different modalities.

[0057] 2. Dynamic adaptive evaluation: By introducing a dynamic weight network that uses public opinion factors and information entropy compensation, the system can automatically adjust evaluation parameters as real-time events develop, enabling rapid responses to hot events. The response delay to public opinion changes is significantly shortened, significantly outperforming traditional fixed models.

[0058] 3. Explainable Decision Process: Utilizing a hybrid reasoning architecture combining symbolic logic and deep learning, the system provides a clear decision path (for example, by visualizing triples and weighted contributions), supports rule-level intervention, and makes evaluation results traceable. Compared to pure black-box models, this system reduces the workload of manual result review and improves decision transparency.

[0059] 4. Improved assessment accuracy and efficiency: The introduction of human-machine collaborative optimization has improved the efficiency of manual review and greatly reduced labor costs.

[0060] 5. Wide range of applications: This invention can efficiently screen and classify massive amounts of UGC content, helping media organizations to promptly identify high-value news leads. It can also be applied to scenarios such as content review on social media platforms and public opinion monitoring and early warning. It has broad application prospects and social value in the fields of news dissemination and content management. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0062] Figure 1 This is a flow chart of the news image value assessment method based on predicate logic and multimodal fusion provided by the present invention.

[0063] Figure 2 This is a working principle diagram of the predicate extraction logic provided by the present invention.

[0064] Figure 3 This is a framework diagram of the news image value assessment system based on predicate logic and multimodal fusion provided by the present invention. DETAILED DESCRIPTION

[0065] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0066] Example 1:

[0067] This invention provides a news image value assessment method based on predicate logic and multimodal fusion, which realizes intelligent assessment of news pictures and videos through an innovative predicate logic parsing engine and a dynamic weight assessment model. Figure 1 As shown, it specifically includes steps S10 to S40:

[0068] S10. Acquire multimodal data of news images, including associated audio, visual, and text data, and use an adaptive trigger mechanism to extract key frames of audio and visual data.

[0069] This step is responsible for the acquisition and preprocessing of news image data, and can provide efficient hierarchical storage to support subsequent analysis. Data sources include video streams, images, and related text descriptions, audio, etc. uploaded by users. Data can be collected through multi-source access interfaces: for example, an API interface that follows the RESTful specification is provided, and necessary metadata fields (such as user_id, timestamp, geo_location, etc.) are supported for uploading in JSON / XML format. For real-time video streams, RTMP protocol is used for access and H.264 encoding is used (ensuring a resolution of no less than 720p and a bit rate of approximately 2Mbps) to ensure high-quality transmission.

[0070] For example, captured video is first processed through dynamic keyframe extraction. A dual-threshold trigger mechanism can be used to detect changes in video content, combining image HSV histogram differences and audio energy mutations. When the color histogram difference between adjacent frames exceeds a set threshold or an abnormal peak in the audio signal is detected, a scene change or significant event is detected, and that frame is extracted as a keyframe.

[0071] Example: Keyframe extraction based on HSV histogram differences and audio energy mutations:

[0072] Step (1): Video and audio preprocessing:

[0073] Video frame extraction: Breaks the video into a sequence of frames, typically at a rate of 25 frames per second (25fps).

[0074] Audio signal extraction: Extract audio signals from videos and perform preprocessing such as noise reduction and normalization.

[0075] Step (2): Calculate the HSV histogram difference:

[0076] HSV histogram calculation: For each frame of image, calculate its HSV color space histogram.

[0077] Histogram difference calculation: Calculate the HSV histogram difference between adjacent frames. Common methods include Bhattacharyya distance or histogram cross entropy.

[0078] Step (3): Calculate the audio energy change:

[0079] Short-time energy calculation: Divide the audio signal into frames and calculate the short-time energy of each frame.

[0080] Energy mutation detection: Detects energy changes between adjacent audio frames and identifies energy mutation points.

[0081] Step (4): Dual-threshold trigger mechanism

[0082] Set the threshold:

[0083] Image histogram difference threshold (e.g., 0.5).

[0084] Audio energy change threshold (e.g., 10dB).

[0085] Trigger condition: When the HSV histogram difference between adjacent frames exceeds the set threshold, or the audio energy change exceeds the set threshold, it is determined to be a scene switch or an important event, and the current frame is extracted as a key frame.

[0086] For example, in addition to histograms and energy thresholds, more advanced triggering conditions can be introduced, such as optical flow change detection and semantic segmentation differences. Motion amplitude can be assessed by calculating the optical flow field between the current frame and the previous keyframe, or by comparing changes in object regions before and after semantic segmentation. Keyframe capture is also triggered when the motion vector amplitude or semantic region difference exceeds a predetermined threshold. This multi-condition fusion dynamic frame extraction mechanism improves the sensitivity and accuracy of keyframe extraction, adapting to the frame selection requirements of scenes with intense motion and sudden semantic changes.

[0087] A frame extraction strategy is implemented on video data to extract key frames, which can further generate a spatiotemporal correlation coding structure between frames. Specifically, the extracted key frames are organized into graph-structured data in chronological order, so that each frame retains image content features while also recording the temporal adjacency between frames through graph nodes and edges. For example, with key frames as nodes, directed edges in chronological order are established between adjacent frame nodes; edge weights can be assigned based on inter-frame similarity (for example, calculated using optical flow mean displacement or structural similarity (SSIM)), thereby forming a temporal graph. This graph structure can further integrate spatial relationships: if multiple objects are detected within a frame, the objects can also be used as child nodes, connected by additional edges to represent spatial co-occurrence relationships. The final output spatiotemporal coding data not only contains visual content features but also explicitly depicts the context of event development over time.

[0088] While extracting video keyframes, we can also obtain the audio clips and text descriptions corresponding to each keyframe's timestamp and attach them to the frame node, forming a multimodal joint representation. This way, when analyzing a keyframe later, we can easily access its related audio features (such as Mel-frequency cepstral coefficients (MFCCs) and speech-to-text content) and text descriptions, laying the foundation for multimodal fusion analysis.

[0089] Furthermore, in specific implementations, for example, a hierarchical storage architecture can be employed to improve efficiency for data of varying value and access frequency: recently accessed data is stored in a high-speed cache layer (e.g., a Redis cluster, which uses an LRU strategy to manage memory for fast access to the latest content), while historical data is stored in a distributed object store (e.g., a MinIO cluster) and combined with a search engine index (e.g., Elasticsearch) for retrieval of massive amounts of content. If necessary, structured metadata can also be stored in a graph database (e.g., Neo4j) to correlate event relationships. This collection and storage mechanism provides well-structured, multimodal data with good temporal and spatial correlation for subsequent value assessment.

[0090] S20. Perform semantic analysis on key frames through the predicate logic parsing engine, and generate subject-action-scene triples in combination with text data, and double-verify them through spatial position and semantic association.

[0091] In this step, the input image content can be semantically parsed to extract the "subject-action-scene" predicate triples that can represent the news event. Figure 2 The figure shows the working process of predicate extraction logic, including:

[0092] 1. First, perform subject detection, using a computer vision object detection algorithm to locate the main entities (such as people, vehicles, etc.) in the image. Preferably, a pre-trained object detection model (such as YOLO or ResNet-CNN) is used to obtain preliminary detection results. Then, a non-maximum suppression (NMS) algorithm is used to filter out overlapping bounding boxes, ensuring that only the most relevant main objects are retained.

[0093] 2. Next, action recognition is performed: A video action recognition model is used to analyze the action categories of the detected subject regions. Training combines public datasets (such as COCO and ImageNet videos) with a self-built library of action samples from news events. Deep convolutional neural networks (such as ResNet-50) are used to extract spatial features. Attention mechanisms on time series (such as the introduction of a spatiotemporal attention module) are also incorporated to improve the robustness of action recognition across consecutive frames, enabling accurate determination of subject actions (e.g., "running," "talking," "colliding," etc.).

[0094] 3. Then perform scene semantic analysis: Use image classification or semantic segmentation technology to extract the image's environmental and scene information (such as "fire scene", "sports stadium" or "everyday street", etc.), or combine it with the image's text description to enrich the scene semantics.

[0095] 4. Finally, the above results are integrated to generate a predicate triple representation: subject-action-scene. To ensure the validity and accuracy of this triple, this module introduces a dual spatial and semantic verification mechanism at the symbolic logic layer. First, spatial positional relationships are used to verify the consistency of the subject and scene (for example, requiring the detected subject bounding boxes to have sufficient overlap within the scene area, with an Intersection over Union (IoU) value of 0.6 or higher). Second, semantic matching is calculated (for example, using the CLIP multimodal model to calculate the similarity between the image content and the corresponding text description, with a threshold of at least 0.7) to verify the proper pairing of action and scene. If a candidate triple fails the above spatial or semantic threshold tests, it is discarded or marked as uncertain. A triple passes verification only if both the spatial relationship and semantic relevance meet the threshold conditions. This step ensures that the generated predicate triple is accurate and contextually relevant, avoiding mismatches between subject and action. This predicate logic parsing engine can extract structured semantic information embedded in news images, providing highly explanatory foundational features for subsequent value assessment.

[0096] S30. Calculate the news value score of the verified triplet data based on a three-dimensional dynamic evaluation model; the three-dimensional dynamic evaluation model integrates communication intensity, information entropy compensation, and ontological value;

[0097] The three-dimensional dynamic evaluation model comprehensively considers dimensions such as communication intensity, information entropy compensation, and ontological value to quantitatively evaluate the value of news images. The three-dimensional dynamic evaluation model is expressed as follows:

[0098]

[0099] In the formula, I(t), H(t), and S(t) represent the transmission intensity, information entropy compensation, and ontological value, respectively; It reflects the changing trend of event importance over time; coefficients α, β, and γ are used to balance the contribution of each part to the total value score.

[0100] 1) I(t) represents the intensity of dissemination, reflecting the breadth and speed of news diffusion through various channels (which can be calculated by channel weight, dissemination entropy and time decay function);

[0101] The diffusion intensity I(t) is typically used to measure the breadth and speed of news or information dissemination across various channels. Its calculation takes into account factors such as channel weight, diffusion entropy, and time decay function. The following is a detailed calculation method for these factors:

[0102] 1.1) Channel Weight:

[0103] Different communication channels (such as social media, news websites, television, etc.) have different influences on information diffusion. To quantify this difference, a weight w can be assigned to each channel. i , reflecting its importance in the overall communication. These weights can be determined based on historical data, audience coverage, or expert evaluation.

[0104] The contribution of channel weight to the propagation intensity can be expressed as:

[0105]

[0106] N represents the total number of information dissemination channels; w i represents the weight of the i-th channel; f i (t) represents the dissemination volume of the i-th channel at time t (such as the number of forwarding, reading, etc.); H p (t) represents the propagation entropy, which is used to measure the uniformity of information distribution among channels.

[0107] 1.2) Propagation Entropy:

[0108] Propagation entropy is used to measure how evenly information is distributed across channels. A higher entropy value indicates that information is more evenly distributed across multiple channels and has a wider reach. It can be calculated using Shannon Entropy:

[0109]

[0110] Where: p i (t) represents the propagation ratio of the i-th channel at time t.

[0111]

[0112] f j (t) represents the propagation volume of the jth channel at time t.

[0113] 1.3) Time Decay Function:

[0114] The influence of information usually weakens over time. To reflect this, a time decay function D(t) can be introduced. Commonly used time decay functions include exponential decay and power law decay:

[0115] aExponential decay:

[0116] D(t)=e -δt (5)

[0117] Where δ>0 is the attenuation coefficient, which controls the attenuation speed.

[0118] b Power law decay:

[0119]

[0120] Where ε>0 is the decay exponential.

[0121] The choice of decay function depends on the specific application scenario and data characteristics.

[0122] Taking all the above factors into consideration, the propagation intensity I(t) can be expressed as:

[0123] I(t)=D(t)*I channel (t)*H p (t) (7)

[0124] H p (t) represents the propagation entropy, which is used to measure the uniformity of information distribution among channels.

[0125] 2) H(t) represents information entropy compensation, which measures the rarity or uncertainty of event information (it can be represented by the inverse of event similarity. When the event is too common, the entropy value is high and points are deducted for compensation);

[0126] In the information propagation model, the information entropy compensation term H(t) is used to measure the rarity or uncertainty of event information. Its calculation is usually based on the inverse of event similarity, and the specific process is as follows:

[0127] 2.1) Event similarity calculation

[0128] First, define the current event E t With the historical event set {E1,E2,...,E n}. Similarity Sim(E t ,E i ) can be calculated in a variety of ways, for example:

[0129] Semantic similarity: Use word vector models (such as Word2Vec and BERT) to calculate the cosine similarity between event descriptions.

[0130] Keyword matching: Calculate Jaccard similarity based on the event keyword set.

[0131] Content features: Consider attributes such as time, location, and participants of an event to calculate the similarity of structured data.

[0132] After getting all the similarities, calculate the average similarity:

[0133]

[0134] Sim(·) represents the event similarity calculation.

[0135] 2.2) Information entropy compensation calculation

[0136] The information entropy compensation term H(t) reflects the rarity or uncertainty of an event. When the event is highly similar to historical events, it means that the event is more common and the entropy value should be reduced; conversely, the entropy value should be increased. It can be calculated as follows:

[0137] Countdown form:

[0138]

[0139] Here, ∈ is a small constant that prevents the denominator from being zero.

[0140] Logarithmic form:

[0141]

[0142] This form emphasizes that low similarity (i.e., rare events) corresponds to higher entropy values.

[0143] 2.3) Example Calculation

[0144] Assume that the average similarity between the current event and the historical event is Assume ∈ = 0.01:

[0145] Countdown form:

[0146] Logarithmic form: H(t)=-log(0.2+0.01)≈1.70

[0147] This indicates that the event is relatively rare and the information entropy compensation value is high.

[0148] Through the above method, the information entropy compensation term H(t) can effectively reflect the rarity or uncertainty of the event, thereby adjusting the propagation intensity of the event in the propagation model and improving the accuracy and robustness of the model.

[0149] 3) S(t) represents the ontological value or static importance of an event (calculated from multiple dimensions, such as the status of the individuals involved, economic impact, and social attention). This value measures the inherent influence and importance of an event before it spreads. To quantify S(t), we can comprehensively evaluate it from multiple dimensions, including the status of individuals involved, economic impact, and social attention.

[0150] The following is the specific calculation process:

[0151] 3.1) Dimensional Indicator Definition

[0152] Decompose S(t) into several key dimensions and set quantifiable indicators under each dimension:

[0153]

[0154] 3.2) Quantification and standardization of indicators

[0155] Each indicator is quantified and standardized so that its value is within a uniform scale (such as 0 to 1):

[0156] Normalization method: Min-Max Normalization or Z-score normalization is used.

[0157]

[0158] Example: If there are 500 news reports related to an event, and the minimum number of historical data reports is 100 and the maximum number of historical data reports is 1000, then the normalized value is:

[0159]

[0160] 3.3) Weight distribution and comprehensive scoring

[0161] Assign weight w to each dimension k , reflecting its importance in the overall evaluation. The weight can be determined through methods such as expert scoring and the Analytic Hierarchy Process (AHP).

[0162] For example, based on historical data and expert knowledge, basic weights are assigned to news event elements, and real-time factors are incorporated to calculate the news value score of the video content. First, a basic weighting table is constructed: for different categories of event actions, initial weights are determined jointly by domain experts and large-scale historical news data statistics. For example, high-impact actions with a significant impact on news value (such as "explosion," "collapse," and "collision") are assigned higher weights (e.g., 0.8-0.9), social event actions (such as "assembly," "strike," and "signing") are assigned medium weights (approximately 0.5-0.7), and everyday actions (such as "walking," "talking," and "standing") are assigned lower weights (approximately 0.3). These weights reflect the differences in the amount of information (self-information) and social attention contained in different actions. Through statistical analysis of historical news events, this weighting table can be optimized to better align with objective news value patterns.

[0163] The calculation formula of the comprehensive score S(t) is:

[0164]

[0165] Where w k represents the weight of the k-th event evaluation dimension; ∑w k =1;xk represents the standardized value of the k-th event evaluation dimension; M represents the total number of event evaluation dimensions; the evaluation dimensions are abstracted and analyzed from the mainstream academic definition of news value, including but not limited to: person status, economic impact, social attention, geographical influence range and time type, etc.

[0166] 4) Example calculation

[0167] Assume that the standardized indicator values ​​and weights of an event are as follows:

[0168] Dimensions <![CDATA[Normalized value x k > <![CDATA[Weight w k > Character status 0.8 0.3 Economic impact 0.6 0.25 Social attention 0.7 0.2 Geographical influence 0.5 0.15 Event Type 0.9 0.1

[0169] The comprehensive score is:

[0170] S(t)=(0.3×0.8)+(0.25×0.6)+(0.2×0.7)+(0.15×0.5)+(0.1×0.9)=0.695

[0171] Therefore, the ontological value score of this event is 0.695.

[0172] Through the above method, the ontological value of an event can be systematically evaluated, providing a quantitative basis for news value assessment, resource allocation, etc.

[0173] Through the above model or similar mathematical expressions (such as other differential equations or neural network approximation), the basic weight table (corresponding to the S dimension) and the real-time calculated propagation / entropy information (corresponding to the I and H dimensions) are integrated to calculate the comprehensive news value score of the current image content.

[0174] Furthermore, step S30 further includes: obtaining a real-time public opinion factor P(t) associated with the triple data, and collecting event attention indexes through a social media API interface;

[0175] Dynamically adjust the model weight coefficient according to the public opinion factor P(t):

[0176] When P(t)>θ hot When , increase the propagation intensity weight α;

[0177] When P(t)<θ cold When , increase the information entropy compensation weight β;

[0178] Among them, θ hot is the high thermal threshold, θ cold is the low thermal threshold.

[0179] To support dynamic weight calculations, this system provides real-time public opinion factors and event background information. First, a news event database (historical news value map) is constructed: a large amount of historical news event data is collected, and event clustering analysis is performed using natural language processing and clustering algorithms. For example, the pre-trained language model BERT is used to extract news text features, and the OPTICS density clustering algorithm is used to cluster similar news items into categories, forming event clusters. For each event cluster, its information entropy is calculated: if a category of events contains redundant or common information, the entropy value is high; conversely, if rare events contain a large amount of new information, the entropy value is low (indicating high unique value). This information entropy can be used as a compensation factor for news value: when a video content involves an event in a high-entropy (common) category, its score is appropriately lowered; conversely, for low-entropy (rare) events, the score is increased, encouraging the discovery of unique news. Second, real-time public opinion factors are obtained: Attention metrics related to the current event are obtained through public data interfaces on social media and online platforms (such as the Weibo Hot Search Index and the Baidu Index API). This real-time public opinion data is converted into quantitative factors, such as the "public opinion heat value." By combining public opinion heat with basic weights and information entropy, the weighting network is dynamically adjusted: when an event related to a news image garners strong attention in real-time public opinion, the corresponding evaluation weight is increased; if an event has low real-time popularity but high basic value, the influence of both is balanced to prevent undervaluation of less popular but important news. By combining the event library and public opinion factors, the evaluation model of this invention is sensitive to time evolution: it both references historical patterns and promptly responds to current dynamics, making news value scoring more objective and real-time.

[0180] S40: Optimizing the dynamic evaluation model by visually tuning the evaluation parameters and rules.

[0181] In specific implementation, for example, interactive interfaces and mechanisms can be provided to enable human experience and feedback to be integrated into the model, thereby continuously optimizing the evaluation results.

[0182] First, a parameter visualization sandbox was designed: the effects of key model parameters (such as indicator weights α, β, and γ, action base weights, and public opinion weighting coefficients) were visualized as trajectory graphs or heat maps, visually demonstrating the impact of parameter adjustments on scoring results. Users can drag sliders in the sandbox interface to adjust parameter values ​​and observe the changing trends in news value scores, thereby intuitively understanding the basis for model decisions. This parameter sandbox mechanism ensures that the parameter adjustment process is reversible and transparent: each adjustment is recorded as a parameter evolution trajectory, which can be rolled back or repeated at any time.

[0183] Secondly, the reversible compilation function of rules is realized: users are supported to write evaluation rules in natural language (for example, "If the event category is a disaster, then increase the weight of the relevant action to above 0.8"), and such rules are dynamically compiled into machine-executable constraints and applied to the model; conversely, the decision logic implicit in the current model can also be extracted into a readable rule description for user review and modification. This bidirectional programmable capability ensures the deep integration of human experience and machine models. Through the human-machine collaboration of step S40, the news image value assessment method forms a closed-loop optimization: manual feedback continuously corrects model deviations, and the interpretable output of the model guides manual decision-making, realizing a virtuous cycle of automatic evaluation and manual supervision, and improving the robustness and credibility of the system.

[0184] After the human-computer interaction is complete, the model can be updated and its effectiveness tracked. First, based on the user-confirmed parameter adjustments, the corresponding values ​​are sent to the value assessment module to dynamically update α, β, γ, or related thresholds. Simultaneously, the model's decision logic is updated based on the newly added rules. Since the model is continuously running, feedback control must ensure a smooth update process. This can be achieved through gradual updates, such as linearly interpolating new parameters over several minutes to replace old ones, rather than abruptly changing them. During this period, new rule triggering is suspended until the parameters are in place, preventing the simultaneous activation of new rules and causing uncertainty. Next, the model output is monitored, comparing the changes before and after the adjustments. If significant deviations or anomalies are detected, a warning is logged, or the change may be automatically rolled back (for example, if a new rule causes all scores to drop to 0, the rule is immediately deactivated and an alarm is issued). This step also ensures that the results of manual adjustments are persisted, for example, by writing them to a log database, including the details of the adjustment, who made the adjustment, and the time of the adjustment, for auditing and subsequent analysis.

[0185] Example 2:

[0186] Based on the same inventive concept, an embodiment of the present invention also provides a news image value assessment system based on predicate logic and multimodal fusion. Since the principle of solving the problem by this system is similar to the aforementioned news image value assessment method based on predicate logic and multimodal fusion, the implementation of this system can refer to the implementation of the aforementioned method, and the repeated parts will not be repeated.

[0187] Reference Figure 3 As shown in the figure, the overall system architecture includes:

[0188] The multimodal acquisition module is used to obtain multimodal data of news images, including associated audio, visual and text data, and uses an adaptive trigger mechanism to extract key frames of audio and visual data;

[0189] The predicate logic parsing engine module is used to perform semantic parsing on key frames using the predicate logic parsing engine, generate subject-action-scene triples based on text data, and perform dual verification through spatial location and semantic association;

[0190] A news value scoring module is used to calculate the news value score of the verified triplet data based on a three-dimensional dynamic evaluation model; the three-dimensional dynamic evaluation model integrates communication intensity, information entropy compensation and ontological value;

[0191] The human-machine collaborative optimization module is used to optimize the dynamic evaluation model by visually tuning the evaluation parameters and rules.

[0192] The various modules of the system work together to perform semantic analysis, value calculation and feedback optimization on the input image content.

[0193] Its overall architecture can be divided into three layers: data layer, computing layer, and decision layer. The data layer is responsible for the acquisition and hierarchical storage of multimodal image data, providing structured spatiotemporal coding data; the computing layer includes a historical news value map and a dynamic weight assessment model to implement the core calculation of image value; the decision layer uses the human-computer collaborative module to perform interpretable tuning and feedback optimization on the model, forming a closed-loop adaptive improvement mechanism. The collaborative relationship between the modules at each level is as follows:

[0194] The multimodal acquisition module preprocesses the real-time video, audio and other raw image streams into spatiotemporal coding data and stores them in a layered storage medium; the historical news value map module uses a large amount of historical event data to calculate the benchmark reference for news value assessment (including background data in dimensions such as dissemination intensity, information entropy, and ontological value); the predicate logic parsing engine module and the news value scoring module extract semantic triples from real-time image frames and calculate value scores based on historical maps and real-time public opinion; the human-computer collaborative optimization module obtains manual feedback (adjusting weights, editing rules, etc.) to generate optimized model parameters, which react to the three-dimensional dynamic evaluation model, so that the system can continuously evolve and adapt to environmental changes.

[0195] Among them, the hierarchical storage architecture: adopts hot / warm / cold tiered storage strategy to improve read and write efficiency:

[0196] Hot Data Layer: Deploy a high-speed cache database (such as a Redis in-memory database cluster) to temporarily store recent keyframe data and metadata. The hot layer uses culling strategies such as LRU to ensure the rapid availability of the latest data.

[0197] Warm data layer: Deploy a distributed NoSQL database (such as Cassandra or MongoDB) to store medium-term data and support efficient queries by event or time index.

[0198] Cold data layer: Deploy massive object storage (such as MinIO cluster or HDFS) combined with offline indexing (Elasticsearch, etc.) to store long-term historical data, large-scale video clips, image files, etc. Cold data is archived through regular batch processing to provide a basis for future reference.

[0199] Constructing a Historical News Value Graph: A large amount of historical news footage and related event data is collected and organized into a graph database using knowledge graph technology. Nodes in the graph represent news events or elements (people, locations, event themes, etc.), while edges represent connections between events (e.g., similar subject matter, temporal causality, and transmission paths). To obtain the graph's embedding representation and clustering characteristics, the TransR algorithm is used to vectorize event relationships. Simultaneously, the OPTICS clustering algorithm is used to identify high-density clusters of similar events in the embedding space. This allows for the calculation of event rarity and aggregation: if an event falls within a dense cluster, it indicates a common event type; conversely, if it falls within a sparse region, it indicates a rare event. The graph records the historical dissemination intensity (e.g., the reach and popularity of each event across various media platforms at the time), information entropy (e.g., frequency and uncertainty of the event), and ontological value (e.g., the importance of the event itself, such as casualties and social impact). These statistics will serve as reference benchmarks for subsequent models.

[0200] The system of the present invention: from multimodal data collection and preprocessing, to dynamic model calculation, to semantic analysis and scoring, and finally through human-computer collaborative continuous optimization, forming a complete automatic evaluation and adaptive improvement cycle.

[0201] Specifically, the present invention uses predicate logic parsing of the "subject-action-scene" triple to conduct in-depth semantic analysis of news image content and extract the core elements of news events; combined with the constructed historical news value map and information entropy model, it introduces public opinion factors to dynamically weight the evaluation process; and through human-computer collaboration, it tunes the model parameters to ensure the accuracy and interpretability of the evaluation results. The system can not only automatically identify the main characters (subjects), behaviors (actions), and environmental backgrounds (scenes) in news images, but also adjust the evaluation strategy according to real-time changes in public opinion, thereby making timely and reasonable assessments of the news value of the image content.

[0202] Through the collaborative operation of the above modules, the system has achieved significant technical results in news image value assessment. The system efficiently processes massive amounts of multimodal data, significantly reducing response time to emergencies. By incorporating interpretable logical analysis and human-computer interaction optimization mechanisms, the number of manual review interventions is reduced, improving overall review efficiency. Furthermore, the assessment results are more transparent and reliable. The system outputs the key triples and weights that support the scoring, making it easier for users to understand and trust the model's decisions.

[0203] The present invention has a wide range of applications and practical value. First, for news media organizations, the system can be used to automatically screen user-uploaded news photos and videos, quickly identifying clues with high news value from massive amounts of material, improving editing efficiency and reducing the risk of missing important news. Second, on social media and content community platforms, the present invention can serve as a content recommendation or review tool, grading user-generated content by news value, thereby prioritizing the promotion of content related to significant events and suppressing the dissemination of low-value information that may disrupt public opinion. Third, in the areas of government regulation and public opinion monitoring, the system can help relevant departments promptly capture online public opinion hotspots. By assessing the value of citizen-uploaded event videos, it can assist in determining the severity and potential impact of events, providing data support for emergency response and decision-making. Finally, the present invention can also be expanded to include media big data analysis and intelligent archive management. By quantifying the value of image data, it can enable more diverse media application scenarios. The present invention not only improves the accuracy and efficiency of assessments but also provides an important tool for the intelligent management and dissemination of media content, and is expected to generate significant social and economic benefits.

[0204] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0205] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A news image value assessment method based on predicate logic and multimodal fusion, characterized by: The following steps are involved: S10, acquiring multimodal data of news images, including associated audio, visual, and text data, and extracting key frames of audio and visual data using an adaptive triggering mechanism; S20, semantically analyze the keyframes using a predicate logic parsing engine, and generate subject-action-scene triples based on text data, which are then verified using both spatial location and semantic association. S30. Calculate the news value score of the verified triplet data based on a three-dimensional dynamic evaluation model; the three-dimensional dynamic evaluation model integrates communication intensity, information entropy compensation, and ontological value; S40: Optimizing the dynamic evaluation model by visually tuning the evaluation parameters and rules.

2. The news image value assessment method based on predicate logic and multimodal fusion according to claim 1 is characterized in that: In step S10, the adaptive trigger mechanism extracts the key frames of the audio-visual data, including: Detect changes in visual data features by analyzing image mutations through the color histogram differences between adjacent frames, triggering key frame extraction; For audio data feature change detection, the energy changes between adjacent audio frames are detected through short-time energy analysis to trigger key frame extraction.

3. The news image value assessment method based on predicate logic and multimodal fusion according to claim 1 is characterized in that: The step S20 includes: S201. Use the pre-trained target detection model and the non-maximum suppression strategy to filter redundant detection boxes and perform subject detection. S202. Extract subject motion features through the image sequence optical flow method, feed the visual features into the trained action classification model, and integrate the spatiotemporal attention mechanism to capture the subject's actions; S203, using the pre-trained CNN model to perform semantic label judgment on the entire frame; at the same time, performing instance segmentation on the background area in the frame, extracting scene elements, and generating a brief semantic description of the scene; S204: Verify the spatial position by calculating whether the spatial overlap between the subject and the scene exceeds a first threshold; and verify the semantic association by using cross-modal semantic analysis to calculate whether the CLIP similarity between the triple text description and the key frame image is greater than a second threshold.

4. The news image value assessment method based on predicate logic and multimodal fusion according to claim 1 is characterized in that: In step S30, the three-dimensional dynamic evaluation model is represented as follows: In the formula, I(t), H(t), and S(t) represent the transmission intensity, information entropy compensation, and ontological value, respectively; It reflects the changing trend of event importance over time; coefficients α, β, and γ are used to balance the contribution of each part to the total value score.

5. The news image value assessment method based on predicate logic and multimodal fusion according to claim 4 is characterized in that: The expression formula of the propagation intensity I(t) is as follows: I(t)=D(t)*I channel (t)*H p (t) Where D(t) represents the time decay function of the influence of information; when exponentially decaying, D(t) = e -δt or power law decay δ represents the exponential attenuation coefficient of the influence of information; ε represents the power law attenuation coefficient of the influence of information; t represents time; I channel (t) represents the weighted sum of channel communication volume; N represents the total number of information dissemination channels; w i represents the weight of the i-th channel; f i (t) represents the propagation volume of the i-th channel at time t; H p (t) represents the propagation entropy, which is used to measure the uniformity of information distribution among channels; Among them, p i (t) represents the propagation ratio of the i-th channel at time t, f j (t) represents the propagation volume of the jth channel at time t.

6. The news image value assessment method based on predicate logic and multimodal fusion according to claim 4 is characterized in that: The information entropy compensation H(t) is expressed as follows: Countdown form: Or in logarithmic form: Indicates the current event E t Average similarity to historical events; Sim(·) represents the event similarity calculation, and ∈ represents the zero-proof constant.

7. The news image value assessment method based on predicate logic and multimodal fusion according to claim 4 is characterized in that: The ontological value S(t) is expressed as follows: Where w k represents the weight of the k-th event evaluation dimension; ∑w k =1;x k represents the standardized value of the evaluation dimension of the kth event; M represents the total number of dimensions for event evaluation; the evaluation dimensions include: person status, economic impact, social attention, geographical impact range, and time type.

8. The news image value assessment method based on predicate logic and multimodal fusion according to claim 4 is characterized in that: The step S30 further includes: Obtain the real-time public opinion factor P(t) associated with the triplet data, and collect event attention indicators through the social media API interface; Dynamically adjust the model weight coefficient according to the public opinion factor P(t): When P(t)>θ hot When , increase the propagation intensity weight α; When P(t)<θ cold When , increase the information entropy compensation weight β; Among them, θ hot is the high thermal threshold, θ cold is the low thermal threshold.

9. The news image value assessment method based on predicate logic and multimodal fusion according to claim 4 is characterized in that: The step S40 includes: Build a parameter sandbox to visualize the α, β, and γ weight trajectories, display the historical adjustment path of the model weight coefficients and the real-time changes of the scores; Adjust any weight parameter based on natural language rules, compile into machine constraints, and support reverse parsing; The logical basis and weight contribution ratio of each score adjustment are recorded to optimize the dynamic evaluation model.

10. A news image value assessment system based on predicate logic and multimodal fusion, characterized by: Using the news image value assessment method based on predicate logic and multimodal fusion as described in any one of claims 1 to 9, the system includes: The multimodal acquisition module is used to obtain multimodal data of news images, including associated audio, visual and text data, and uses an adaptive trigger mechanism to extract key frames of audio and visual data; The predicate logic parsing engine module is used to perform semantic parsing on key frames using the predicate logic parsing engine, generate subject-action-scene triples based on text data, and perform dual verification through spatial location and semantic association; A news value scoring module is used to calculate the news value score of the verified triplet data based on a three-dimensional dynamic evaluation model; the three-dimensional dynamic evaluation model integrates communication intensity, information entropy compensation and ontological value; The human-machine collaborative optimization module is used to optimize the dynamic evaluation model by visually tuning the evaluation parameters and rules.

Citation Information

Cited By

  • Live broadcast goods carrying real-time detection method and system based on multi-modal fusion

    CN121330409A

  • News event exposure assessment method and system based on Internet propagation

    CN121705347A

  • A news event exposure amount evaluation method and system based on internet dissemination

    CN121705347B