Large model GEO+SEO content generation closed loop optimization system

CN122287845BActive Publication Date: 2026-08-21HEFEI STAR ARTIFICIAL INTELLIGENCE APPLICATION SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610570961.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-28
Publication Date
2026-08-21
Estimated Expiration
2046-04-28

AI Technical Summary

Technical Problem

[0003]但在生成式引擎输出具有强动态性且呈现形态随平台、终端与交互方式变化的条件下,引用相关质量治理面临更细粒度的工程挑战:其一,同一评估查询集合在多个采样时刻触发的答案文本可能发生措辞与结构变化,引用标识的出现位置、占用面积、呈现形态与可见交互特征也可能改变,导致引用呈现记录的曝光影响难以用统一口径衡量;其二,引用目标页面往往存在版本演化,若缺少与页面版本标识、引用基础快照标识绑定的引用基础沉淀机制,即使保留了引用标识,也难以在事后对照核验当时的证据上下文,从而增加复核成本;其三,现有引用校验更多停留在链接可达性、页面存在性或粗粒度的内容相似性判断,较少在命题分解得到的命题集合层面建立命题与候选证据片段集合之间的证据支撑关系,并进一步计算跨源证据一致性置信度来量化命题与证据的一致程度,因此难以形成可按采样时刻串联、可解释冲突点来源的错引证据链;其四,风险评估在实践中往往侧重正确与否的单点评分,较少将错引风险分量与引用曝光势能指数、以及由命题风险类别、业务影响向量、合规约束向量与品牌损伤向量共同决定的危害谱系权重进行融合,从而难以得到兼顾风险强度与曝光强度的曝光加权错引风险势能;其五,当引用基础变化或生成策略变化导致风险在相邻采样时刻之间快速波动时,若缺少以风险漂移速度刻画变化速率并驱动采样频率调整指令、评估查询集合更新指令、内容生成约束指令与内容发布控制指令的闭环调控机制,则难以及时、可度量地抑制风险扩散并稳定优化结果

Benefits of technology

本发明通过在多个采样时刻对评估查询集合发起生成式引擎查询并结构化采集引用呈现记录,同时基于引用标识在引用基础库中形成可追溯的引用基础,使答案文本中的引用依据能够与采样时刻、页面版本演化建立对应关系,显著降低事后复核与责任定位成本并提升可核验性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122287845B_ABST
    Figure CN122287845B_ABST
Patent Text Reader

Abstract

The present application relates to the field of natural language processing and information retrieval technology, and particularly relates to a large model GEO+SEO content generation closed loop optimization system, which collects answer texts and forms citation basis and citation presentation records, models proposition and citation relationship and establishes evidence support relationship, calculates cross-source evidence consistency confidence to generate mis-citation evidence chain, evaluates exposure weighted mis-citation risk potential and risk drift speed, and outputs sampling frequency adjustment instructions, evaluation query set update instructions, content generation constraint instructions and content release control instructions closed loop optimization. The present application collects citation identifiers to form citation basis and citation presentation records which can be traced back; calculates cross-source evidence consistency confidence by proposition decomposition, citation relationship graph and evidence support relationship and generates mis-citation evidence chain; obtains exposure weighted mis-citation risk potential and risk drift speed by fusing citation exposure potential index, hazard spectrum weight and mis-citation risk component, and drives regulation and control instruction closed loop optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing and information retrieval technology, and more specifically, to a large-scale model GEO+SEO content generation closed-loop optimization system. Background Technology

[0002] In content production practices geared towards generative engine optimization and search engine optimization, common practices involve developing topic selection and writing guidelines around the evaluation query set, generating answer text through generative engines or writing aids, and enhancing searchability and readability through keyword placement, structured paragraphs, heading hierarchy, and internal linking. To enhance credibility and explainability, some practices include adding citations or source links to the answer text and iterating based on search feedback, clicks, and conversion metrics after publication. Other approaches incorporate offline evaluation, sampling verification, and manual review to revise obvious factual errors or inappropriate statements. These methods play a positive role in improving content organization and operational efficiency.

[0003] However, given the highly dynamic nature of generative engine outputs and the variations in presentation across platforms, terminals, and interaction methods, citation-related quality governance faces more granular engineering challenges: First, the wording and structure of answer texts triggered at multiple sampling times for the same evaluation query set may change, and the location, area, presentation, and visible interaction features of citation identifiers may also change, making it difficult to measure the exposure impact of citation presentation records using a uniform standard. Second, citation target pages often undergo version evolution. Without a citation base accumulation mechanism bound to page version identifiers and citation base snapshot identifiers, even if citation identifiers are retained, it is difficult to verify the evidence context at the time of the event, thus increasing the cost of review. Third, existing citation verification focuses more on link reachability, page existence, or coarse-grained content similarity judgments, rarely establishing a connection between propositions and candidate evidence fragment sets at the proposition set level obtained from proposition decomposition. Based on supporting relationships, and further calculating the cross-source evidence consistency confidence level to quantify the consistency between propositions and evidence, it is difficult to form a chain of miscitation evidence that can be linked according to sampling time and explain the source of conflict points; fourth, in practice, risk assessment often focuses on single-point scoring of correctness or incorrectness, and rarely integrates the miscitation risk component with the citation exposure potential index, as well as the harm spectrum weight determined by the proposition risk category, business impact vector, compliance constraint vector, and brand damage vector, thus making it difficult to obtain an exposure-weighted miscitation risk potential that takes into account both risk intensity and exposure intensity; fifth, when changes in the citation basis or generation strategy cause the risk to fluctuate rapidly between adjacent sampling times, if there is a lack of a closed-loop control mechanism that characterizes the rate of change with the risk drift speed and drives the sampling frequency adjustment instruction, the evaluation query set update instruction, the content generation constraint instruction, and the content release control instruction, it is difficult to timely and measurably suppress the spread of risk and stabilize and optimize the results.

[0004] It is evident that existing technologies still have room for improvement in areas such as referencing verifiable records, fine alignment of propositions to evidence, risk potential measurement of exposure and harm weighting, and dynamic closed-loop control that evolves over time. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the purpose of this invention is to provide a closed-loop optimization system for large-scale GEO+SEO content generation.

[0006] To achieve the above objectives, the present invention provides the following technical solution: The large-scale GEO+SEO content generation closed-loop optimization system includes: The data collection module is configured to: initiate generative engine queries on the evaluation query set at multiple sampling times to obtain the answer text; extract citation identifiers from the answer text and form a citation base in the citation base library based on the citation identifiers; and perform structured collection of citation presentation elements related to the citation identifiers in the answer text to generate citation presentation records. The relation modeling module is configured as follows: decompose the answer text into propositions to obtain a set of propositions; construct a reference relationship graph to represent the reference relationship between the set of propositions and the reference identifiers; retrieve a set of candidate evidence fragments from the reference base library based on the reference relationship graph, and establish the evidence support relationship between propositions and candidate evidence fragments; The evidence alignment module is configured as follows: based on the evidence support relationship, calculate the cross-source evidence consistency confidence level for the proposition and candidate evidence fragments; and generate a chain of miscited evidence based on the cross-source evidence consistency confidence level and the citation basis. The risk assessment module is configured as follows: calculate the citation exposure potential energy index based on the citation presentation record; calculate the hazard spectrum weight; determine the miscitation risk component based on the consistency confidence of cross-source evidence; and determine the exposure-weighted miscitation risk potential energy and risk drift velocity based on the miscitation risk component, the citation exposure potential energy index and the hazard spectrum weight. The dynamic control module is configured to generate control instructions and execute closed-loop optimization based on the exposure-weighted misleading risk potential energy and risk drift speed. The control instructions include at least the sampling frequency adjustment instruction, the evaluation query set update instruction, the content generation constraint instruction, and the content publishing control instruction.

[0007] Furthermore, the reference identifier is obtained from the answer text parsing, and the corresponding reference target page identifier is determined; the target page content is obtained based on the reference target page identifier, and a reference base snapshot is generated; the reference base snapshot is written into the reference base library to form the reference base.

[0008] Furthermore, the referenced presentation records include location features, occupancy features, presentation form features, visible interaction features, and terminal features.

[0009] Furthermore, the citation relationship graph includes proposition nodes, citation identifier nodes, and the relationship edges between proposition nodes and citation identifier nodes. The edge attributes of the relationship edges include at least the proposition fragment position, proposition semantic vector, citation presentation record identifier, and sampling time.

[0010] Furthermore, the reference identifier node points to the reference base snapshot in the reference base library through the reference base index relationship. The reference base index relationship includes at least the index mapping between the reference identifier, the reference target page identifier, the page version identifier, and the reference base snapshot identifier, so that the reference relationship can be verified on the reference base.

[0011] Furthermore, the miscited evidence chain is generated by selecting proposition nodes with cross-source evidence consistency confidence below a threshold in the citation relationship graph, locating the citation base snapshot along the citation identifier node through the citation base index relationship, extracting the conflict points between candidate evidence fragments and propositions, and concatenating them according to the sampling time.

[0012] Furthermore, based on the reference presentation record, location features, occupancy features, presentation form features, visible interaction features, and terminal features are extracted. Each feature is normalized and then weighted and fused with platform calibration to obtain the reference exposure potential index.

[0013] Furthermore, the hazard spectrum weights are calculated based on the propositional risk category, business impact vector, compliance constraint vector, and brand damage vector. The hazard spectrum weights are used to weight the risk of miscitation of citation relationships under the citation basis.

[0014] Furthermore, by integrating the miscitation risk component, the citation exposure potential index, and the hazard spectrum weight, the exposure-weighted miscitation risk potential is obtained.

[0015] Furthermore, the risk drift velocity is calculated based on the ratio of the change in exposure-weighted misleading risk potential energy between adjacent sampling times to the corresponding sampling time interval.

[0016] Compared with the prior art, the present invention has the following beneficial effects: This invention initiates generative engine queries on the evaluation query set at multiple sampling times and collects citation presentation records in a structured manner. At the same time, it forms a traceable citation base in the citation base library based on citation identifiers, so that the citation basis in the answer text can establish a correspondence with the sampling time and page version evolution, which significantly reduces the cost of post-event review and responsibility location and improves verifiability. By decomposing propositions to obtain a set of propositions, constructing a citation relationship graph, and establishing the evidence support relationship between propositions and candidate evidence fragments, the consistency confidence of cross-source evidence is further calculated and a chain of miscited evidence is generated, so that the source of miscited evidence can be explained and verified in a conflict point and time series manner, thereby improving the precision, interpretability and consistency of miscited identification. By integrating the miscitation risk component with the citation exposure potential energy index and the hazard spectrum weight, the exposure-weighted miscitation risk potential energy is obtained. The risk drift velocity is then calculated by combining adjacent sampling times. This generates sampling frequency adjustment instructions, evaluation query set update instructions, content generation constraint instructions, and content release control instructions to implement closed-loop optimization. This enables risk management to match the exposure impact and hazard intensity, and achieves more timely and stable dynamic control during the risk acceleration phase. Attached Figure Description

[0017] Figure 1 A schematic diagram of the overall structure of a closed-loop optimization system for GEO+SEO content generation in a large model; Figure 2 A schematic diagram of the closed-loop optimization process for a large-scale GEO+SEO content generation closed-loop optimization system. Figure 3 This is a structural diagram illustrating the reference relationship and the basic index relationship. Figure 4 This is a schematic diagram illustrating the dynamic control of exposure-weighted misleading risk potential energy and risk drift velocity. Detailed Implementation

[0018] Reference Figures 1 to 2 The large-scale GEO+SEO content generation closed-loop optimization system includes: The data acquisition module is configured to: initiate generative engine queries against the evaluation query set at multiple sampling times to obtain answer text; and repeatedly observe the same batch of evaluation query sets at multiple sampling times, enabling the system to capture the output changes of the generative engine queries over time, thereby providing a continuous data foundation for subsequent calculations of risk drift velocity. The evaluation query set is used to cover the core intent and key themes in the target business scenario, and the answer text, as the output carrier of the generative engine queries, carries all the original corpora required for subsequent proposition decomposition, citation extraction, evidence alignment, and risk assessment.

[0019] The data collection module is configured to: extract citation identifiers from the answer text and form a citation base in the citation base library based on these identifiers; explicitly represent the citation information in the answer text as citation identifiers and assign these identifiers to manageable objects in the citation base library, forming a traceable and verifiable citation base. This avoids the unverifiable problem caused by relying solely on the instantaneous presentation of the answer text, enabling any subsequent judgments regarding whether a citation supports the proposition or whether a citation is miscited to be verified retrospectively based on the citation base in the citation base library; and simultaneously creates the prerequisites for establishing a complete link between the proposition set—citation identifier—citation base.

[0020] The data collection module is configured to: structurally collect citation presentation elements related to citation identifiers in the answer text, generating citation presentation records; transforming the presentation of citations in the answer text from an unstructured phenomenon into a computable data object, namely, the citation presentation record. The citation presentation record describes the exposure and reach characteristics of citation identifiers in the answer text, directly serving the calculation of the subsequent citation exposure potential index; it also allows the influence of citation relationships (the likelihood of being seen, clicked, or followed) to be quantified, thereby achieving a weighted coupling between the risk of miscitation and exposure during the risk assessment stage.

[0021] In one specific implementation, a generative engine query is initiated on the evaluation query set at multiple sampling times to obtain answer text. The answer texts of the same evaluation query set at different sampling times are used as comparable samples. Reference identifiers are extracted from the answer texts and a reference base is formed in the reference base library based on the reference identifiers. At the same time, the reference presentation elements related to the reference identifiers in the answer texts are structurally collected to generate reference presentation records. The reference presentation records include location features, occupancy features, presentation form features, visible interaction features, and terminal features. The reference presentation records are used to calculate the reference exposure potential index and characterize the exposure-side features of the reference base and the reference relationship. For example, the evaluation query set includes queries such as the filter replacement cycle of a certain brand of air purifier and weekend family routes in a certain city. In the answer text obtained at sampling time one, the reference identifiers are presented in the form of footnote numbers and occupy the first screen. In sampling time two, the same reference identifiers are presented in the form of cards and can be clicked to jump. The reference presentation records are based on this to solidify the differences in location features, occupancy features, and visible interaction features. The structured collection of reference presentation elements can be based on Implementation of layout information: at sampling time Get the answer page Using the rendering layout tree, locate the bounding box of the rendering block corresponding to each reference identifier. The system analyzes the type of display location (first screen, folded area, or sidebar) and calculates the position ratio and occupancy ratio accordingly. It also analyzes the reference presentation format (footnote number, embedded link, or card display) and visible interactive capabilities (clickable to jump, expandable for preview), and records the terminal type and screen size to form terminal characteristics. Finally, it generates a reference presentation record bound to the reference identifier. When the platform cannot directly obtain... At that time, screenshots were used in conjunction with coordinate mapping rules to locate the bounding box of the referenced block, and then the position, occupancy, shape and interaction features were calculated according to the same criteria to ensure the comparability of referenced presentation records on different platforms. The process involves parsing the answer text to obtain a reference identifier and determining the corresponding target page identifier. Based on this identifier, the target page content is retrieved, and a reference base snapshot is generated. This snapshot is then written into the reference base library to form the reference base, thus anchoring the reference relationship to verifiable page content. For example, the reference identifier is parsed to reveal a target page identifier pointing to a product description page. When retrieving the target page content, the page text and key parameter paragraphs are simultaneously retained to form a reference base snapshot. After being written into the reference base library, the reference identifier at any subsequent sampling time can be traced back to the same reference base snapshot or its updated version through the target page identifier. This enables a traceable comparison between the reference base and the reference relationship over time and provides a consistent presentation context for calculating the reference exposure potential index.

[0022] The relational modeling module is configured to: decompose the answer text into a set of propositions; and break down the answer text into a set of propositions that are verifiable, alignable, and measurable, with minimal semantic units. Proposition decomposition enables subsequent evidence retrieval and consistency judgment to move beyond the level of the entire text and establish a correspondence between propositions and evidence at the propositional granularity. This supports the fine-grained calculation of cross-source evidence consistency confidence and reduces alignment noise caused by mixed expressions in long texts.

[0023] Reference Figure 3 The relationship modeling module is configured as follows: It constructs a citation relationship graph to represent the citation relationships between the proposition set and citation identifiers; the citation relationship graph structures and expresses these relationships, enabling the system to clearly identify which propositions depend on which citation identifiers. As the organizational structure for subsequent retrieval and verification, the citation relationship graph supports both tracing back from citation identifiers to the citation base in the citation base library and locating candidate evidence fragments from the proposition set and establishing evidence support relationships. The citation relationship graph also provides a path framework for generating subsequent miscitation evidence chains, allowing miscitation issues to be located and explained along the graph structure.

[0024] The relationship modeling module is configured as follows: based on the citation relationship graph, it retrieves a set of candidate evidence fragments from the citation base library and establishes an evidentiary support relationship between the proposition and the candidate evidence fragments; using the citation relationship graph as the index entry point, it performs a retrieval in the citation base library targeting the citation base to obtain a set of candidate evidence fragments related to the proposition set, and establishes a clear evidentiary support relationship between the proposition and the candidate evidence fragments. The establishment of this evidentiary support relationship provides a clear alignment object for the subsequent calculation of cross-source evidence consistency confidence, and transforms the miscitation judgment from "subjective comparison" to a structured verification process of "proposition-evidence fragment"; simultaneously, it advances the connection of the citation relationship from the proposition to the citation identifier to a verifiable level from the proposition to the citation base content fragment.

[0025] In one specific implementation, the answer text is decomposed into a proposition set, and the position of the proposition fragment for each proposition is preserved during the proposition decomposition process so that the position can be retrieved in the answer text later. A reference relationship graph is constructed to represent the reference relationship between the proposition set and the reference identifier. The reference relationship graph includes proposition nodes, reference identifier nodes, and relationship edges between proposition nodes and reference identifier nodes. The edge attributes of the relationship edges include at least the proposition fragment position, proposition semantic vector, reference presentation record identifier, and sampling time. This allows the dependency path of different propositions on different reference identifiers within the same sampling time to be expressed in a structured manner, and allows the change of reference relationship across sampling times to be identified by comparison. For example, for the query of the filter replacement cycle of a certain brand of air purifier in the evaluation query set, the answer text is decomposed into a proposition set including propositions such as the filter replacement cycle is six months and the applicable condition is eight hours per day. Relationship edges are established between proposition nodes and corresponding reference identifier nodes, and the proposition fragment position and sampling time are written into the edge attributes so that the coverage difference of the proposition set corresponding to the same reference identifier at different sampling times can be distinguished later. Based on the citation relationship graph, a set of candidate evidence fragments is retrieved from the citation base library, and an evidentiary support relationship is established between propositions and candidate evidence fragments. The citation identifier node points to a citation base snapshot in the citation base library through a citation base index relationship. The citation base index relationship includes at least an index mapping between the citation identifier, the citation target page identifier, the page version identifier, and the citation base snapshot identifier, so that the citation relationship is verifiable on the citation base. In this step, the citation target page identifier is parsed from the citation identifier as the entry point. The citation base snapshot is located based on the page version identifier and the citation base snapshot identifier. A set of candidate evidence fragments matching the proposition semantic vector is extracted from the citation base snapshot. Each proposition is then linked to... The corresponding candidate evidence fragments establish an evidence support relationship so that during subsequent verification, the specific evidence fragment in the reference base snapshot can be traced from the proposition node along the reference relationship graph. For example, when the reference identifier corresponds to the reference target page identifier of the product description page, the page version identifier indicates the updated parameter page, and the reference base snapshot contains paragraphs on the replacement cycle and usage duration conditions, after the proposition and candidate evidence fragments establish an evidence support relationship, it is possible to directly verify whether the proposition is supported by the clauses in the reference base snapshot. Furthermore, when the page version identifier changes due to changes in the sampling time, the index mapping enables traceable comparison of different versions of the reference base snapshot for the same reference identifier, ensuring a consistent verification path between the reference base and the reference relationship.

[0026] The evidence alignment module is configured as follows: based on the evidence support relationship, it calculates the cross-source evidence consistency confidence score for the proposition and candidate evidence fragments; using the evidence support relationship as a boundary, it limits the set of candidate evidence fragments that each proposition should be aligned with, and calculates the cross-source evidence consistency confidence score accordingly. The cross-source evidence consistency confidence score is used to quantify the degree of consistency between the proposition and candidate evidence fragments in terms of semantics and key information, thereby transforming whether something is supported by evidence into a thresholdable, sortable, and aggregateable indicator, providing quantitative input for subsequent miscitation risk components.

[0027] The evidence alignment module is configured to generate a chain of incorrect citations based on cross-source evidence consistency confidence and the citation basis. When the cross-source evidence consistency confidence indicates an inconsistency between the proposition and the candidate evidence fragment, the citation basis is used as a verifiable carrier to connect "proposition—citation identifier—citation basis content—inconsistency point" to form a chain of incorrect citations. The purpose of the chain of incorrect citations is to transform incorrect citations from an alert at the result level into interpretable evidence at the process level. This allows the subsequent dynamic control module to generate targeted content generation constraint instructions based on the chain of incorrect citations and provides a directly verifiable basis for manual review.

[0028] In one specific implementation, based on the evidence support relationship, the cross-source evidence consistency confidence level is calculated for the proposition and candidate evidence fragments. First, each proposition is semantically aligned with the candidate evidence fragments pointed to by its evidence support relationship. Then, considering factors such as the proposition's limiting conditions, numerical range, time conditions, and applicable objects, the degree of matching between the proposition and the candidate evidence fragments in terms of factual points, constraint boundaries, and conclusion consistency is calculated to obtain the cross-source evidence consistency confidence level. This cross-source evidence consistency confidence level is then linked to the sampling time to allow for comparison of changes in the evidence consistency of the same proposition at different sampling times. For example, regarding the query set for the replacement cycle of a certain brand of air purifier filter, the proposition set includes a six-month filter replacement cycle with an applicable condition of eight hours per day. The candidate evidence fragment comes from a product description page paragraph in the cited basis. When the candidate evidence fragment states a replacement cycle of three to six months with pollution level as a condition, the cross-source evidence consistency confidence level will decrease due to inconsistent conditions. Based on the supporting evidence relationship, the cross-source evidence consistency confidence score is calculated for the proposition and candidate evidence fragments. Specifically, this includes generating a proposition semantic vector for the proposition and generating candidate evidence fragment semantic vectors for the candidate evidence fragments, and then using the cosine similarity between the two. Calculate semantically consistent sub-scores Extract elements such as numerical range, time conditions, applicable objects, and limiting conditions from the proposition, and extract corresponding elements from candidate evidence fragments. Calculate the element consistency sub-score based on the ratio of the number of elements that are consistent item by item to the total number of elements. Natural language inference is used to obtain the implication probability of the proposition relative to the candidate evidence fragments. Contradictory probability And calculate the consistent sub-scores. ;Will , and By weight , and Weighted fusion yields single-segment consistency confidence. ,in When the same proposition corresponds to multiple candidate evidence fragments, take... As the cross-source evidence consistency confidence level of this proposition at the sampling time, the identifier of the selected candidate evidence fragment is associated with... Bind records for review; When the same proposition When dealing with multiple candidate evidence fragments, to avoid a single fragment accidentally supporting the conflict and obscuring the truth, the confidence level of cross-source evidence consistency is crucial. Do not take directly Instead, aggregation with contradictory penalties is adopted: ;in To be according to Sort selection Collection of evidence fragments For normalized weights (e.g.) ), Let be the probability of a contradiction between the proposition and the evidence. For contradiction penalty coefficient; when When the preset contradiction threshold is exceeded, Set as Or it can directly trigger the generation of a chain of evidence for misquotation, so as to ensure that it is not covered by highly similar fragments when there is obvious conflict; Based on cross-source evidence consistency confidence and citation basis, a chain of miscited evidence is generated. This chain is generated by selecting proposition nodes with cross-source evidence consistency confidence below a threshold in the citation relationship graph, locating the citation basis snapshot along the citation identifier node via the citation basis index relationship, extracting conflict points between candidate evidence fragments and propositions, and concatenating them according to the sampling time. In this step, proposition nodes requiring explanation are filtered by a threshold. The citation identifier node associated with the proposition node is traced back to the citation basis index relationship. The citation basis snapshot is located based on the index mapping of the citation identifier, the citation target page identifier, the page version identifier, and the citation basis snapshot identifier. Conflicts with the proposition are extracted from the citation basis snapshot. The candidate evidence fragments are labeled with the type and location of the conflict points, and the conflict points at different sampling times are linked together in chronological order. This allows the chain of incorrect citation evidence to simultaneously present the expression of the proposition in the answer text, the associated path of the citation identifier, the original evidence of the citation basis, and the evolution of the conflict points with the sampling time. For example, if the candidate evidence fragment at sampling time one is clearly changed every three months, and the page version identifier at sampling time two is updated to be changed every six months with a usage duration condition, the chain of incorrect citation evidence will link and show the source of the difference between the proposition's absolute periodic expression and its conditional periodic expression, and indicate whether the conflict point comes from the overgeneralization of the proposition or the change of the version of the citation basis, thus forming a verifiable chain of incorrect citation evidence. In one specific implementation, the threshold is denoted as This is used to determine whether the cross-source evidence consistency confidence level triggers the generation of a miscited evidence chain, where... ; The threshold can be either a preset threshold or an adaptive threshold: When it is a preset threshold, it is determined based on historical manually reviewed samples to ensure that the false positive rate and false negative rate of the miscited evidence chain meet the preset target. And when there are no historical samples, the default value is used. When using an adaptive threshold, the distribution of cross-source evidence consistency confidence is statistically analyzed over multiple sampling times in each round, and the threshold is set using quantiles. This prioritizes including low-consistency proposition nodes in the misreferenced evidence chain; for example, it assigns higher values ​​to proposition nodes whose risk category is price commitment and return rules. To improve recall, lower risk levels will be adopted for low-risk topics such as weekend family routes in a certain city. To reduce false alarms.

[0029] The risk assessment module is configured to: calculate the citation exposure potential index based on citation presentation records; and utilize the structured presentation information accumulated in the citation presentation records to calculate the citation exposure potential index, which is used to characterize the potential exposure influence of citation identifiers in the answer text environment. The citation exposure potential index allows the system to focus not only on whether there is a miscitation, but also on the extent of dissemination and reach that a miscitation might cause in actual presentation, providing a weighting source for the exposure dimension of subsequent exposure-weighted miscitation risk potential calculations.

[0030] The risk assessment module is configured to: calculate the hazard spectrum weights; introduce hazard spectrum weights to stratify and quantify the severity of risks for different proposition categories and different business consequences, so that inconsistencies of the same degree will have different weighted results in different hazard contexts, thereby making the risk assessment adjustable at the business and compliance levels, and avoiding the determination of the intensity of handling based solely on the single indicator of textual consistency.

[0031] The risk assessment module is configured as follows: It determines the miscitation risk component based on the cross-source evidence consistency confidence level; it maps the cross-source evidence consistency confidence level to the miscitation risk component, ensuring that the consistency results are absorbed by the risk framework. The miscitation risk component describes the probability and degree of deviation of miscitation, is one of the core components of the exposure-weighted miscitation risk potential energy, and directly affects the triggering and intensity of subsequent control instructions from the dynamic control module. In one specific implementation, the misreference risk component is denoted as... Confidence of consistency of cross-source evidence The mapping is obtained, and the mapping satisfies The lower The higher the value, the more piecewise linear the mapping becomes: when season ,when season ,when season ,in and The preset consistency threshold and satisfy ; and Determined based on historical manually reviewed samples, or by using the default value if no historical samples are available. and Through the above mapping, the miscitation risk component is used to characterize the degree of inconsistency between the proposition and the candidate evidence fragments and is used to calculate the exposure-weighted miscitation risk potential energy.

[0032] The risk assessment module is configured as follows: It determines the exposure-weighted miscitation risk potential energy and risk drift velocity based on the miscitation risk component, the citation exposure potential energy index, and the hazard spectrum weights; it integrates the miscitation risk component, the citation exposure potential energy index, and the hazard spectrum weights to obtain the exposure-weighted miscitation risk potential energy, thus simultaneously reflecting the probability of miscitation, the impact of exposure, and the severity of the hazard; and it obtains the risk drift velocity based on continuous calculations at multiple sampling times, used to characterize the trend and acceleration of risk changes over time. This allows the system to capture both the current risk level and the rate of risk change, providing a more stable decision-making basis for closed-loop optimization. In one specific implementation, the exposure weighted misindex risk potential energy is denoted as... Risk component due to misleading Exposure potential index Harm spectrum weight The fusion is obtained by employing multiplicative gating and normalization: first, the weighted exposure term is calculated. With weighted hazard items ,in and For interval Configurable parameters within; then calculate exposure-weighted misleading risk potential energy. ,in To crop the results to a range The function; through this fusion method, when the miscitation risk component is high and the citation exposure potential energy index or hazard spectrum weight is high, the exposure-weighted miscitation risk potential energy is amplified, thereby achieving priority handling of high-exposure and high-hazard miscitations.

[0033] In one specific implementation, the reference exposure potential index is calculated based on the reference presentation record. Specifically, this involves extracting location features, occupancy features, presentation form features, visible interaction features, and terminal features from the reference presentation record, and normalizing each feature to eliminate differences in feature scale between different sampling times and different terminals. The normalized features are then weighted and fused to form the original exposure potential. The original exposure potential is then calibrated by the generative engine querying the platform's display rules to obtain the reference exposure potential index. This ensures that the reference exposure potential index can stably represent the visibility and accessibility of the reference basis and reference relationship on the exposure side. For example, if the same reference identifier is located on the first screen of the answer text and is presented as a card with visible interaction features at sampling time one, but is reduced to the folded area and only presented as a footnote with no visible interaction features at sampling time two, the reference exposure potential index obtained after normalization and platform calibration will decrease significantly, thus reflecting that although the reference relationship exists, its exposure influence is reduced. The citation presentation record consists of location features, occupancy features, presentation morphology features, visible interaction features, and terminal features. Location features describe the relative position of the citation identifier within the answer text, calculated as the ratio of the paragraph number containing the citation identifier to the total number of paragraphs in the answer text. Occupancy features describe the degree to which the citation identifier occupies the visible area, calculated as the ratio of the height of the corresponding presentation block to the first screen height. Presentation morphology features describe the presentation method of the citation identifier, mapping footnote numbers, embedded links, and card displays to different morphology scores; for example, footnote numbers map to 0.2, embedded links to 0.6, and card displays to 1.0. Visible interaction features describe whether the citation identifier has clickable navigation, expandable preview, or other interactive capabilities, mapping these capabilities to an interaction score. Terminal features describe the terminal type and screen size used by the generative engine query, and are used to perform scale correction across different terminals. Based on this, min-max normalization is performed on each feature to obtain a normalized feature vector. and and by weight and Weighted fusion forms the exposure of the original potential energy The sum of all weights is 1; this is combined with the display rules of the platform where the generative engine query is located. The platform calibration was performed to obtain the reference exposure potential index. calibration factor The display positions, such as the first screen, the folded area, and the sidebar, can be configured with different values. For example, the first screen can be configured with... Folded area This is to reflect the amplification or attenuation effect of different display positions on exposure; Platform calibration factor The display slot type determines the display slot and is maintained in the form of a configuration table. The configuration table includes at least the display slots such as the main screen, the collapsed area, and the sidebar, and their corresponding configuration tables. The value is mapped and can be periodically updated based on historical click-through rates or viewing duration statistics; the exposure potential index is used... To ensure stable values ​​and cross-platform comparability; harming the phylogenetic weights. The default is to use a weighted average mapping: for dimensions of vector ,make ,in Non-negative weights and The risk factors are determined through scoring rules or manually labeled sample statistics. When business strategies or compliance constraints change, the scoring rules and weight tables for risk categories to vector components are updated online, and the version number and effective time are recorded each time an update is made to support audit review and cross-sampling time comparison. The calculation of hazard spectrum weights involves classifying the set of propositions into risk categories based on their risk types, and constructing a business impact vector, a compliance constraint vector, and a brand damage vector for each risk category. The hazard spectrum weights are then calculated according to preset vector combination rules, and are used to weight the risk of miscitation in the citation relationship supported by the citation basis. For example, when the set of propositions involves price commitments and return / exchange rules, the risk categories of propositions are more biased towards the compliance constraint vector and the brand damage vector. When the set of propositions involves safe usage conditions, the weight of the business impact vector is increased, thus forming different hazard spectrum weights under the same cross-source evidence consistency confidence level. After categorizing the set of propositions according to their risk types, a business impact vector, a compliance constraint vector, and a brand damage vector are constructed for each risk type. The business impact vector represents the severity of the business consequences that a proposition error may cause, the compliance constraint vector represents the strength of the compliance constraints that a proposition error may trigger, and the brand damage vector represents the degree of brand damage that a proposition error may cause. Each vector uses the same dimension. The real number vector representation of , whose components take values ​​in the range of . The risk factor scoring rules are used to obtain the risk spectrum weights from pre-defined risk factor scoring rules or manually labeled samples. The pre-defined vector combination rules obtain the hazard spectrum weights by weighted summation and normalization of the above three types of vectors. ,in and These represent the business impact vector, compliance constraint vector, and brand damage vector, respectively. and Non-negative weights and , To map a vector to a scalar and normalize it to an interval The function; where Take the maximum value or weighted average of the vector components; for example, when the propositional risk category involves price commitments and return / exchange rules, increase... and To enhance the contribution of compliance constraint vector and brand damage vector, when the propositional risk category involves safe use conditions, it improves... To enhance the contribution of the business impact vector; The miscitation risk component is determined based on the cross-source evidence consistency confidence level. The miscitation risk component, the citation exposure potential index, and the hazard spectrum weights are then integrated to obtain the exposure-weighted miscitation risk potential. Specifically, the cross-source evidence consistency confidence level is mapped to a miscitation risk component to characterize the degree of inconsistency between the proposition and candidate evidence fragments. The miscitation risk component is then integrated with the citation exposure potential index and the hazard spectrum weights, so that the exposure-weighted miscitation risk potential simultaneously reflects the inconsistency intensity, exposure intensity, and hazard intensity. For example, when an absolute statement appears in the proposition set while the evidence in the citation basis is a conditional statement, the miscitation risk component increases. If the citation identifier also occupies a large feature and is located at the beginning of the mobile terminal features, the citation exposure potential index increases. After adding the hazard spectrum weights, the exposure-weighted miscitation risk potential presents a high risk. The risk drift velocity is calculated based on the ratio of the change in exposure-weighted miscitation risk potential energy between adjacent sampling times to the corresponding sampling time interval. Specifically, the exposure-weighted miscitation risk potential energy output by the same evaluation query set at adjacent sampling times is differentially analyzed to obtain the change, which is then divided by the sampling time interval to obtain the risk drift velocity. This risk drift velocity can characterize the speed at which risk increases or decreases and provide a trend basis for subsequent regulatory instructions. For example, when a page version change causes an update to the citation basis, and the proposition set shows a transition from consistency to inconsistency in two consecutive sampling times, the change in exposure-weighted miscitation risk potential energy increases, and the risk drift velocity increases accordingly. This can indicate that the risk is in a rapidly deteriorating stage and requires more intensive sampling and stricter content generation constraints.

[0034] Reference Figure 4 The dynamic control module is configured to generate control instructions based on the exposure-weighted misleading risk potential energy and risk drift speed, and execute closed-loop optimization. These control instructions include at least sampling frequency adjustment instructions, evaluation query set update instructions, content generation constraint instructions, and content publishing control instructions. The risk assessment results are transformed into executable control instructions, and the execution results are then fed back into the subsequent sampling and generation processes, forming a closed-loop optimization.

[0035] In one specific implementation, the dynamic control module is based on the exposure-weighted misleading risk potential energy. With risk drift speed Generate control instructions, including the risk drift rate The value is calculated based on the ratio of the change in exposure-weighted misleading risk potential energy between adjacent sampling times to the corresponding sampling time interval; to improve stability, the value is adjusted within the sliding window. Take the mean or median to get and to Take the average value ;when or and When a sampling frequency adjustment command is triggered to increase the sampling time density, a content publishing control command is triggered to switch the content from automatic publishing to publishing after manual review; when and When a content generation constraint instruction is triggered to restrict absolute statements and require a one-to-one correspondence between propositions and reference identifiers, an evaluation query set update instruction is triggered to supplement long-tail queries on the same topic; when and The sampling frequency is reduced and some content generation constraints are removed; among them and To make the threshold configurable, the objective is determined based on historical closed-loop optimization results to reduce the potential energy of exposure-weighted misleading risk and slow down the risk drift rate, and a default value is used when there are no historical samples. and .

[0036] The sampling frequency adjustment command is used to increase the observation density when the risk changes drastically or the risk level is high, so that the citation base library, citation presentation record and citation relationship graph are updated more timely, and the risk lag is reduced.

[0037] The evaluation query set update command is used to expand or adjust the coverage of the evaluation query set, enabling the system to more comprehensively monitor and evaluate new sets of propositions and changes in citation relationships.

[0038] Content generation constraint instructions are used to transform evidentiary results such as misquoted evidence chains into constraints that can be executed during the generation phase, thereby reducing the probability of generating inconsistent propositions or inappropriate references again from the source.

[0039] Content release control instructions are used to control the way and pace of content entering the release chain when the exposure-weighted misleading risk potential is high and the risk drift speed indicates that the risk is intensifying, thereby reducing the risk of external dissemination before the closed-loop optimization is stable.

[0040] In one specific implementation, control instructions are generated based on the exposure-weighted miscitation risk potential energy and the risk drift speed. Specifically, the exposure-weighted miscitation risk potential energy is used as the control intensity benchmark, and the risk drift speed is used as the control urgency benchmark. The two are linked and judged to form an executable set of control instructions. The control instructions include at least a sampling frequency adjustment instruction, an evaluation query set update instruction, a content generation constraint instruction, and a content publication control instruction. The causal consistency between the instructions is maintained during the generation process. The sampling frequency adjustment instruction is used to increase or decrease the observation density at multiple sampling times. The evaluation query set update instruction is used to expand or rearrange the coverage and priority of the evaluation query set. The content generation constraint instruction is used to constrain the expression boundary of the answer text, the citation identifier binding method, and the proposition expression granularity. The content publication control instruction is used to control the external exposure path before the risk disposal is completed. For example, when the exposure-weighted miscitation risk potential corresponding to a certain assessment query set reaches a high level and the risk drift speed shows an upward trend, the sampling frequency adjustment instruction will increase the sampling time density to capture risk changes, the assessment query set update instruction will supplement long-tail queries on the same topic to verify whether there is diffusion of similar propositions, the content generation constraint instruction will restrict absolute statements and require a one-to-one correspondence between propositions and citation identifiers, and the content publishing control instruction will switch the content of this topic from automatic publishing to publishing after manual review; Based on the control instructions and perform closed-loop optimization, specifically, the sampling frequency adjustment instruction is applied to the collection rhythm of subsequent sampling times, the evaluation query set update instruction is applied to the input set of the next round of generative engine query, the content generation constraint instruction is applied to the generation stage to reduce the miscitation risk component, and the content publishing control instruction is applied to the publishing link to reduce the risk amplification effect caused by the reference exposure potential energy index. Then, at the new sampling time, the exposure-weighted miscitation risk potential energy and risk drift speed are recalculated and compared with the baseline before execution to form a verifiable closed-loop optimization result. For example, regarding the topic of filter replacement cycle for a certain brand of air purifier, after executing the content generation constraint instruction, the proposition must include applicable conditions and be bound to the reference base with the reference identifier. After executing the content release control instruction, the release of high-exposure positions is temporarily suspended. In addition, with the sampling frequency adjustment instruction to encrypt the sampling time observation, it can be observed in the subsequent sampling time that the exposure weighted misreference risk potential energy decreases and the risk drift speed slows down, thus proving that the control instruction has effectively suppressed the risk diffusion and completed closed-loop optimization.

[0041] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0042] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A large-scale GEO+SEO content generation closed-loop optimization system, characterized in that: include: The data collection module is configured to: initiate generative engine queries on the evaluation query set at multiple sampling times to obtain the answer text; extract citation identifiers from the answer text and form a citation base in the citation base library based on the citation identifiers; and perform structured collection of citation presentation elements related to the citation identifiers in the answer text to generate citation presentation records. The relation modeling module is configured as follows: decompose the answer text into propositions to obtain a set of propositions; construct a reference relationship graph to represent the reference relationship between the set of propositions and the reference identifiers; retrieve a set of candidate evidence fragments from the reference base library based on the reference relationship graph, and establish the evidence support relationship between propositions and candidate evidence fragments; The evidence alignment module is configured as follows: based on the evidence support relationship, calculate the cross-source evidence consistency confidence level for the proposition and candidate evidence fragments; and generate a chain of miscited evidence based on the cross-source evidence consistency confidence level and the citation basis. The risk assessment module is configured to calculate the reference exposure potential index based on the reference presentation record; Calculate the hazard spectrum weights; Determine the miscitation risk component based on the consistency confidence level of cross-source evidence; The component of misquotation risk is denoted as Confidence of consistency of cross-source evidence The mapping is obtained, and the mapping satisfies The lower The higher the value, the more piecewise linear the mapping becomes: when season ,when season ,when season ,in and The preset consistency threshold and satisfying ; and Determined based on historical manually reviewed samples, or by using the default value if no historical samples are available. and ; The exposure-weighted miscitation risk potential energy and risk drift velocity are determined based on the miscitation risk component, the citation exposure potential energy index, and the hazard spectrum weight. Exposure weighted misleading risk potential energy is recorded as Risk component due to misleading Exposure potential index Harm spectrum weight The fusion is obtained by employing multiplicative gating and normalization: first, the weighted exposure term is calculated. With weighted hazard items ,in and For interval Configurable parameters within; then calculate the exposure-weighted misleading risk potential energy. ,in To crop the results to a range The function; The dynamic control module is configured to generate control instructions and execute closed-loop optimization based on the exposure-weighted misleading risk potential energy and risk drift speed. The control instructions include at least the sampling frequency adjustment instruction, the evaluation query set update instruction, the content generation constraint instruction, and the content publishing control instruction.

2. The large-scale model GEO+SEO content generation closed-loop optimization system according to claim 1, characterized in that, The reference identifier is obtained from the answer text parsing, and the corresponding reference target page identifier is determined; the target page content is obtained based on the reference target page identifier and a reference base snapshot is generated; the reference base snapshot is written into the reference base library to form the reference base.

3. The large-scale model GEO+SEO content generation closed-loop optimization system according to claim 1, characterized in that, The referenced presentation record includes location characteristics, occupancy characteristics, presentation form characteristics, visible interaction characteristics, and terminal characteristics.

4. The large-scale model GEO+SEO content generation closed-loop optimization system according to claim 1, characterized in that, The citation relationship graph includes proposition nodes, citation identifier nodes, and the relationship edges between proposition nodes and citation identifier nodes. The edge attributes of the relationship edges include at least the proposition fragment position, proposition semantic vector, citation presentation record identifier, and sampling time.

5. The large-scale GEO+SEO content generation closed-loop optimization system according to claim 4, characterized in that, The reference identifier node points to the reference base snapshot in the reference base library through the reference base index relationship. The reference base index relationship includes at least the index mapping between the reference identifier, the reference target page identifier, the page version identifier, and the reference base snapshot identifier, so that the reference relationship can be verified on the reference base.

6. The large-scale model GEO+SEO content generation closed-loop optimization system according to claim 1, characterized in that, The miscited evidence chain is generated by selecting proposition nodes with cross-source evidence consistency confidence below a threshold in the citation relationship graph, locating the citation base snapshot along the citation identifier node through the citation base index relationship, extracting the conflict points between candidate evidence fragments and propositions, and concatenating them according to the sampling time.

7. The large-scale model GEO+SEO content generation closed-loop optimization system according to claim 1, characterized in that, Based on the location features, occupancy features, presentation form features, visible interaction features and terminal features extracted from the reference presentation record, the reference exposure potential index is obtained by normalizing each feature and then performing weighted fusion and platform calibration.

8. The large-scale model GEO+SEO content generation closed-loop optimization system according to claim 1, characterized in that, The hazard spectrum weights are calculated based on the propositional risk category, business impact vector, compliance constraint vector, and brand damage vector. The hazard spectrum weights are used to weight the risk of miscitation of citation relationships under the citation basis.

9. The large-scale model GEO+SEO content generation closed-loop optimization system according to claim 1, characterized in that, The exposure-weighted miscitation risk potential energy is obtained by integrating the miscitation risk component, the citation exposure potential energy index, and the hazard spectrum weight.

10. The large-scale model GEO+SEO content generation closed-loop optimization system according to claim 9, characterized in that, The risk drift velocity is calculated based on the ratio of the change in exposure-weighted misleading risk potential energy between adjacent sampling times to the corresponding sampling time interval.

Citation Information

Patent Citations

  • Multi-modal event data processing and collaborative circulation method and device, equipment and medium

    CN121032421A

  • Illegal content auditing method and device based on multi-modal data, equipment and medium

    CN121278126A