Digital marketing information collection method and system based on large model multi-round inference verification
By employing a large-scale model, multi-round inference verification, and a multi-agent collaboration framework, the problem of ensuring data quality in digital marketing systems is solved, enabling high-quality data collection and reliable marketing decisions in multi-source heterogeneous environments.
Patent Information
- Application Number
- CN202610523439.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-10
AI Technical Summary
Existing digital marketing systems lack effective data quality verification mechanisms in the face of strict privacy compliance requirements and data fragmentation, leading to biased marketing decision-making results. Furthermore, traditional methods struggle to verify the authenticity and consistency of marketing information across multi-source heterogeneous data.
A method based on large model and multi-round inference verification is adopted. Through a multi-agent collaborative framework, external evidence is retrieved and multi-round inference is performed on candidate marketing information. Combined with the reliability weight of data source, a comprehensive credibility score is calculated to achieve multi-round inference confidence and credibility screening of marketing information.
It improves the credibility and consistency of marketing data, reduces the interference of false information, ensures the reliability and accuracy of marketing decisions, and achieves high-quality data collection and analysis in a multi-channel data environment.
Smart Images

Figure CN122367523A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent digital marketing technology, and in particular to a digital marketing information collection method and system based on large model multi-round inference verification. Background Technology
[0002] With the development of intelligent technologies, digital marketing has evolved from experience-driven to an intelligent decision-making model centered on data and algorithms. To accurately understand user needs and optimize resource allocation, marketing systems typically need to collect and analyze user behavior data from multiple channels. In traditional digital marketing systems, data collection mainly relies on third-party cross-site tracking technologies (such as third-party cookies) to achieve cross-platform user behavior mapping and analysis.
[0003] However, with increasingly stringent standards for cybersecurity and data privacy protection, traditional data acquisition methods face significant technological barriers. Existing privacy compliance frameworks require data collection to be based on explicit user authorization, and major browsers and operating systems have gradually restricted or phased out third-party tracking technologies at the underlying level. This changing technological environment is forcing digital marketing data sources to shift towards first-party authorized data.
[0004] This shift in data collection models has introduced new technical challenges. First, the heavy reliance on user authorization has drastically reduced the amount of effective data the system can acquire. Second, first-party data is typically scattered across multiple heterogeneous systems, including websites, mobile applications, email systems, and offline business terminals, resulting in severe data fragmentation and the formation of "data silos" that are difficult to communicate with each other. Furthermore, when multi-source, heterogeneous marketing data is aggregated, underlying data quality issues such as missing data, format conflicts, duplicate records, and outdated information are commonly encountered.
[0005] Currently, data analysis tools mainly focus on meeting the compliance requirements of data collection through event-level tracking models at the front end. However, under the conditions of limited data volume and fragmented sources, there is a lack of automatic verification mechanisms for the authenticity, consistency, and high-value content of marketing information itself. Low-quality, noisy data that has not been deeply purified and externally cross-verified will be directly input into downstream big data analysis and AI prediction models, ultimately leading to serious deviations in the output results of the marketing decision-making system. Summary of the Invention
[0006] This application provides a digital marketing information collection method, system, storage medium, computer program product, and electronic device based on large model multi-round inference verification, which can at least solve the problem of the difficulty in guaranteeing the quality of marketing data in current related technologies, leading to deviations in intelligent decision-making results.
[0007] In a first aspect, embodiments of this application provide a method for collecting digital marketing information based on multi-round inference verification using a large model. The method includes: acquiring multi-channel event data with the user's authorization and consent, and mapping the multi-channel event data to a unified data model to generate an integrated event stream; performing semantic classification on the event text in the integrated event stream to filter out marketing text, and performing feature matching filtering on the marketing text based on a preset marketing rule base to identify and extract at least one candidate marketing information; for each candidate marketing information, inputting the candidate marketing information into a pre-built multi-agent inference module based on a large language model, whereby the multi-agent inference module performs external evidence retrieval on the candidate marketing information, and performs multi-round inference and authenticity determination based on the retrieved external evidence to output a multi-round inference confidence level for the candidate marketing information; calculating a comprehensive credibility score for the candidate marketing information by combining the multi-round inference confidence level and the source reliability weight of the data source corresponding to the candidate marketing information; comparing and filtering the comprehensive credibility score based on a preset credibility threshold, and outputting candidate marketing information that reaches the credibility threshold as target marketing data.
[0008] Secondly, embodiments of this application provide a digital marketing information collection system based on a large model and multi-round inference verification. The system includes: an authorization collection and integration unit, used to acquire multi-channel event data upon obtaining user authorization and consent, and map the multi-channel event data to a unified data model to generate an integrated event stream; a marketing semantic filtering unit, used to perform semantic classification on the event text in the integrated event stream to filter out marketing text, and perform feature matching filtering on the marketing text based on a preset marketing rule base, thereby identifying and extracting at least one candidate marketing information; and a multi-agent inference verification unit, used to verify each candidate marketing information... The data is input into a pre-built multi-agent inference module based on a large language model. The multi-agent inference module performs external evidence retrieval on the candidate marketing information and performs multiple rounds of inference and truth / falseness determination based on the retrieved external evidence to output the multi-round inference confidence level for the candidate marketing information. A credibility fusion evaluation unit is used to calculate the comprehensive credibility score of the candidate marketing information by combining the multi-round inference confidence level and the source reliability weight of the data source corresponding to the candidate marketing information. A threshold decision output unit is used to compare and filter the comprehensive credibility score based on a preset credibility threshold and output the candidate marketing information that reaches the credibility threshold as the target marketing data.
[0009] Thirdly, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the digital marketing information collection method based on large model multi-round inference verification according to any embodiment of this application.
[0010] Fourthly, embodiments of this application provide a storage medium storing a computer program thereon, characterized in that, when the program is executed by a processor, it implements the steps of the digital marketing information collection method based on large model multi-round inference verification according to any embodiment of this application.
[0011] Fifthly, embodiments of this application provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the digital marketing information collection method based on large model multi-round inference verification according to any embodiment of this application.
[0012] The digital marketing information collection method and system based on large-model multi-round inference and verification provided in this application can produce at least the following technical effects: (1) With user authorization, multi-channel event data are mapped to a unified model and integrated into an event stream. At the same time, the event text is coupled with semantic classification and rule base feature matching to transform the original, scattered, and inconsistent event information into a set of candidate marketing information that can be aligned and filtered. Since semantic classification can focus marketing-related semantic fragments from noisy events, and rule base matching further filters them in a structured manner with reusable business constraints, the system can still stably produce candidate information with clear semantic boundaries and consistent features even when the data sources are diverse, the field definitions are different, and the text expressions are significantly different. This reduces the dependence of subsequent processing on a single data source or a specific tracking format, and significantly enhances the comparability and convergence of candidate information.
[0013] (2) A multi-agent inference mechanism based on a large language model is introduced for candidate marketing information. Confidence is generated through external evidence retrieval and multi-round inference, and further integrated with the reliability weight of the data source to form a comprehensive credibility score. Finally, the target marketing data is output in a thresholded manner. Since multiple agents can strengthen evidence and verify consistency of the same candidate information from different perspectives, multi-round inference can continuously converge and explicitly quantify the stability of the inference conclusion. Therefore, the system can not only give a true or false tendency of candidate information, but also output a confidence profile that can be used for engineering governance. At the same time, the confidence score is integrated with the source reliability, so that the "content-level evidence support strength" and "source-level credibility" are uniformly measured under the same dimension. This enables the output data to have the characteristics of being graded and controllable, thereby reducing the probability of unreliable information being directly included in downstream analysis and prediction, and improving the stability of target marketing data in terms of consistency and timeliness.
[0014] This technical solution employs a fusion data purification paradigm driven by external evidence, involving multi-agent multi-round inference verification and credibility fusion scoring. This elevates candidate marketing information from passively collected factual records to credible data units supported by evidence and verified by inference consistency. Furthermore, it achieves engineerable data quality governance through quantitative scoring and threshold output. This makes marketing data generated under authorized constraints and multi-source heterogeneous environments more verifiable, controllable, and a reliable foundation for intelligent decision-making. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 A flowchart illustrating an example of a digital marketing information collection method based on large model multi-round inference verification according to an embodiment of this application is shown. Figure 2 A flowchart illustrating an example of generating an integrated event stream in a method according to an embodiment of this application is shown. Figure 3 A flowchart illustrating an example of identifying and extracting candidate marketing information in a method according to an embodiment of this application is shown. Figure 4 A structural block diagram of an example of a multi-agent inference module according to an embodiment of this application is shown; Figure 5 This paper illustrates an example of the operational principle mechanism of a digital marketing information collection method based on large model multi-round inference verification according to an embodiment of this application; Figure 6 A schematic diagram of experimental simulation results is shown, illustrating an example of how accuracy and computational cost evolve with the number of inference rounds in a multi-round inference mechanism according to an embodiment of this application. Figure 7 A schematic diagram illustrating the experimental simulation effect of an example of the source reliability weight dynamic update mechanism according to an embodiment of this application in combating data source noise interference is shown. Figure 8 A simulation diagram illustrating the overall performance of different digital marketing information collection methods is shown. Figure 9 A structural block diagram of an example of a digital marketing information collection system based on large model multi-round inference verification according to an embodiment of this application is shown. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] In the exploration of marketing data governance, current technologies primarily attempt to mitigate the negative impacts of browser limitations and data silos by introducing Customer Data Platforms (CDPs) or employing server-side tagging. However, these measures essentially focus on the engineering aggregation and format standardization of underlying data, mainly addressing the physical aggregation of data flows, without delving into the in-depth verification of the authenticity of marketing content and its consistency with business operations. Lacking effective external cross-validation and dynamic quality assessment mechanisms, and facing limited data volume and fragmented information, the system remains highly susceptible to interference from false information or low-quality data sources, failing to guarantee the credibility of the data ultimately entering the analysis system.
[0019] In terms of intelligent data verification, to improve the accuracy of large language models in processing complex information, some studies have proposed multi-round inference optimization techniques. These techniques utilize previous outputs as contextual cues for the next round of inference to enhance the model's self-correction capabilities. However, related research also points out that such purely model-level inference methods face challenges in practical applications, including difficulty in defining fine-grained inference steps, susceptibility to inference drift, and difficulty in verifying the correctness of intermediate inferences. Relying solely on the model's own loops often fails to converge to an absolutely reliable conclusion.
[0020] Furthermore, some scholars have explored applying multi-agent collaborative frameworks to automated fact-checking tasks, attempting to introduce multiple dedicated agent nodes to perform sub-tasks such as input deconstruction, external evidence retrieval, and result determination. While this framework has shown some potential, it currently focuses primarily on plain text question answering in the general news domain, lacking the ability to deeply process complex, structured data in digital marketing scenarios.
[0021] It should be understood that the above description of the relevant technologies is intended only to help the public better understand the inventive spirit and motivation of this application, and is not intended to limit this application. Furthermore, the technical solutions described in the above-mentioned relevant technologies are not prior art, and may also be undisclosed technical solutions, such as those under research or in the laboratory stage.
[0022] The technical solutions in this application, including the collection, storage, use, processing, transmission, provision, and disclosure of users' personal information, comply with relevant laws and regulations and do not violate public order and good morals.
[0023] Figure 1 A flowchart illustrating an example of a digital marketing information collection method based on large model multi-round inference verification according to an embodiment of this application is shown.
[0024] Regarding the execution subject of the method in the embodiments of this application, it can be any controller or processor with computing or processing capabilities, such as a platform controller deployed in a digital marketing data platform or marketing automation platform. By calling the first-party collection interface or SDK of the website / mobile application, accessing server logs and event streams, and connecting to the data synchronization channel (such as message queue, API gateway or batch ETL pipeline) between the email system and offline business terminals, the system can complete the access, parsing, unified modeling and subsequent calculation processing of multi-channel event data.
[0025] In some examples, the execution entity can also be integrated into electronic devices or terminals in a software, hardware, or a combination of both, and the types of terminals or electronic devices can be diverse, such as servers, cloud computing nodes, edge gateways, enterprise workstations, or mobile smart terminals, to adapt to different deployment architectures and computing power conditions.
[0026] like Figure 1 As shown, in step S110, after obtaining the user's authorization and consent, multi-channel event data is acquired and mapped to a unified data model to generate an integrated event stream.
[0027] It should be noted that in the actual digital marketing ecosystem, data collection faces the dual challenges of "strict privacy compliance" and "fragmented data sources." To address these issues, this embodiment employs a "authorize first, then collect" compliance mechanism.
[0028] Specifically, the system first interacts with users through a consent management platform deployed on the front end (such as a web client or app), clearly displaying the purpose and scope of data collection, and verifying the user's privacy authorization status (such as CookieConsent status) in real time. The system only activates the underlying collection probe when it detects an explicit consent signal from the user.
[0029] In some implementations, the multi-channel event data acquired by the underlying acquisition probe includes behavioral event data triggered by client users at business touchpoints such as websites, mobile applications, emails, and offline terminals. Furthermore, the event types corresponding to the behavioral event data include at least one of the following: page browsing, product purchase, form submission, and video viewing. By covering diverse online and offline interaction channels and the aforementioned key conversion nodes, the system can comprehensively capture the complete behavioral trajectory of users throughout the entire marketing funnel, thereby providing rich and granular underlying materials for subsequent commercial intent identification and data verification.
[0030] At the data collection level, to avoid the blocking of third-party cookies by client browsers (such as ITP mechanisms) and the impact of ad-blocking plugins, this embodiment preferably employs server-side tagging technology. The client is only responsible for sending the most basic encrypted event packets through a secure channel; the complex parsing and distribution logic is completed in an isolated server environment.
[0031] Furthermore, the acquired raw multi-channel event data often exhibits significant heterogeneity (e.g., website data is based on HTTP headers, while offline terminal data is based on POS machine logs). Specifically, the system can incorporate a standardized mapping engine that predefines a unified data model, encompassing globally unique timestamps, standardized user identity identifiers (ID Mapping), unified action type definitions, and extended content payloads. The system automatically parses the raw messages from each channel, cleansing and mapping their heterogeneous fields (such as "url", "page_path", and "screen_view") to the corresponding dimensions of the standard model. Through this structured normalization process, the system successfully integrates previously isolated data from different business platforms into a coherent, time-series-based event stream, providing a standardized data foundation for subsequent intelligent analysis.
[0032] In step S120, the event texts in the integrated event stream are semantically classified to filter out marketing texts, and the marketing texts are filtered by feature matching based on a preset marketing rule base to identify and extract at least one candidate marketing information.
[0033] It should be noted that although the integrated event stream has a unified structure, its content has an extremely low signal-to-noise ratio, mixing massive amounts of non-business system logs (such as heartbeat packets and error logs) with routine user interactions that have no commercial value (such as simple page scrolling). Directly inputting the entire dataset into a large model for validation would result in a huge waste of computing power and inference delays. Therefore, this embodiment designs a two-layer filtering funnel mechanism of "coarse screening-fine extraction".
[0034] First, the system uses a lightweight Natural Language Processing (NLP) model or a text classifier fine-tuned for the marketing domain to perform semantic classification of text payloads in the event stream. This classifier can identify contextual intent, automatically remove noisy data irrelevant to commercial conversion, and accurately retain marketing text containing potential commercial intents such as "promotion," "discount," "new product launch," and "limited-time event."
[0035] Subsequently, to further improve the accuracy of intelligence, the system introduced domain prior knowledge. Based on a pre-set marketing rule base, the system performed a second deep feature filtering on the initially selected marketing texts. This rule base is configured with core element templates that constitute independent marketing intelligence, such as the requirement to include "brand entity," "benefit description (e.g., xx% discount)," and "timeliness limitation." Specifically, the system uses regular expression matching and a keyword dictionary to verify whether the text completely possesses the above core elements. Only information fragments that simultaneously satisfy semantic relevance and feature completeness are packaged and extracted by the system and its accompanying metadata (such as source URL and publication time) to generate candidate marketing information. This effectively achieves data dimensionality reduction, ensuring that every piece of data fed into the subsequent large-scale model has verification value.
[0036] In step S130, for each candidate marketing information, the candidate marketing information is input into a pre-built multi-agent inference module based on a large language model. The multi-agent inference module performs external evidence retrieval for the candidate marketing information and performs multiple rounds of inference and truth / falseness determination in combination with the retrieved external evidence to output the multi-round inference confidence for the candidate marketing information.
[0037] It should be noted that traditional single large model calls are prone to "illusions" or logical omissions when faced with complex commercial marketing jargon, obscure promotional rules, or deliberately vague false information. To address this issue, this embodiment constructs a multi-agent collaboration inference architecture, breaking down the complex verification task into sub-tasks such as "understanding," "retrieval," and "determination," each handled collaboratively by a dedicated agent.
[0038] In some implementations, the understanding agent first deconstructs the candidate information and extracts the key facts to be verified; the query generation agent dynamically generates search instructions based on these facts; and the evidence retrieval agent uses these instructions to access the enterprise's authorized internal knowledge base (such as ERP inventory tables) or legitimate public networks (such as brand websites and authoritative news sources) in real time through API interfaces to retrieve external evidence documents from multiple sources.
[0039] Furthermore, the multi-agent inference module introduces a multi-round iterative self-reflection mechanism. In each round of inference, the decision agent not only cross-checks the contradictions between candidate information and current evidence, but also inputs the previous round's decision as the "reflection context" into the model. By adopting this mechanism, the large model is prompted to continuously question itself: "Is the current evidence sufficient?", "Are there logical flaws in the previous round's reasoning?". After multiple rounds of in-depth fact-checking and logical correction, the system finally outputs a quantified multi-round inference confidence score. This confidence score is no longer a simple "true / false" binary label, but a continuous probability value that objectively reflects the credibility of the information at the semantic, logical, and factual levels, greatly enhancing the system's depth of counterfeit detection in information-fragmented environments.
[0040] In step S140, the comprehensive credibility score of the candidate marketing information is calculated by combining the confidence scores of multiple rounds of inference and the source reliability weights of the data sources corresponding to the candidate marketing information.
[0041] It should be noted that in open and complex network environments, there exists an advanced attack method: malicious nodes or marketing trolls can generate fake promotional information in bulk with sound logic, even capable of fooling plain text verification. However, relying solely on content-level inference from large models is insufficient to completely prevent such attacks based on "source contamination."
[0042] Therefore, this embodiment introduces a reputation consideration at the physical link level. The system can track the underlying production path of each candidate marketing message in real time and extract the "source reliability weight" of its corresponding data source. Specifically, this weight is not a static value, but a reputation score dynamically calculated based on the accuracy, stability, and complaint rate of the information provided by the data source during the historical collection period.
[0043] In calculating the final credibility score, the system employs a dimensionality reduction calculation rule that combines probabilistic verification and deterministic evaluation. It weights and fuses the confidence scores of the large-scale model, representing the authenticity of the content, with the source weights, representing the reliability of the channel. Therefore, even if a piece of false information is extremely well-disguised (leading to a high confidence score in the large-scale model), if it originates from a historically discredited "spam farm" or "unknown crawler node" (with extremely low source weight), the final overall credibility score will still be significantly lowered. This builds a low-level security defense line for the system, unaffected by semantic interference, and greatly enhances the system's robustness against interference from noisy data sources.
[0044] In step S150, the overall credibility score is compared and filtered based on a preset credibility threshold, and the candidate marketing information that reaches the credibility threshold is output as target marketing data.
[0045] Here, after obtaining the comprehensive credibility score, the system enters the final decision-making and output stage. To adapt to the differentiated data quality requirements of different business scenarios, the system supports dynamic configuration of the credibility threshold. For example, in high-risk scenarios such as "protecting the reputation of high-value brands" or "redeeming large-value coupons," the system will automatically increase the threshold (e.g., set to 0.9) and implement the strictest admission standards; while in scenarios with high recall requirements, such as "initial screening of market sentiment," the threshold can be appropriately reduced.
[0046] More specifically, the system performs a hard comparison between the overall score of each candidate marketing message and the currently effective threshold. For messages with scores below the threshold, the system classifies them as fraudulent marketing, expired activities, or low-quality noise, and discards, isolates, or manually reviews them according to the configured policy. For candidate marketing messages that successfully cross the threshold, the system confirms them as target marketing data. This high-value data, after multiple rounds of cleaning and deep verification, is then securely pushed to downstream customer data platforms (CDP), marketing automation systems (MA), or business intelligence (BI) dashboards via standard API interfaces. This fundamentally prevents fraudulent data from contaminating the enterprise's decision-making center, achieving a complete closed loop from compliant data collection and intelligent fraud detection to high-quality delivery.
[0047] Figure 2 A flowchart illustrating an example of generating an integrated event flow in a method according to an embodiment of this application is shown.
[0048] like Figure 2As shown, in step S210, in response to the collection instruction triggered by the client based on the authorization and consent status, basic event information is received through a pre-deployed server-side tag, and sensitive fields in the user identifier contained in the basic event information are desensitized and one-way hash encryption is performed in the server-side isolated environment to generate encrypted identifiers and generate multi-channel event data that meets privacy compliance requirements.
[0049] More specifically, traditional client-side direct data collection is highly susceptible to blocking by browser Intelligent Tracking Protection (ITP) mechanisms or ad-blocking plugins, and poses a risk of privacy breaches. Therefore, this embodiment proposes postponing the data collection process, with the client acting only as a lightweight trigger. Upon detecting explicit user authorization, the client sends minimal basic event information (such as timestamps, device fingerprint fragments, and basic interaction types) to a pre-deployed server-side tagging container. This container runs in an isolated server environment, acting as a "security scrubbing chamber" before data flows into the enterprise intranet.
[0050] In this isolated environment, the system first uses an irreversible one-way hash algorithm (such as the salted SHA-256 algorithm) to de-identify personal information (such as plaintext phone numbers, email addresses, and original device IMEIs) contained in the basic event information, converting it into anonymous encrypted identifiers. Simultaneously, the system removes or generalizes sensitive location fields such as IP addresses in real time. This effectively severs direct contact between the original sensitive data and the external environment from both physical and cryptographic perspectives, significantly reducing the front-end data loss rate and ensuring that the generated multi-channel event data strictly meets the data minimization and anonymization requirements of privacy regulations before entering the subsequent analysis process.
[0051] In step S220, heterogeneous fields are extracted from multi-channel event data, mapped to the standard dimensions of a unified data model, and feature interpolation is performed on the mapped missing fields based on the same-source data completion rules and user historical behavior sequences to obtain the filled event data.
[0052] It should be noted that because multi-channel event data originates from probes of different business systems (such as HTTP request parameters from the web and POS machine logs from offline stores), its field structure exhibits significant heterogeneity. The system first calls the built-in data dictionary to parse and map various heterogeneous fields (such as "click_url" and "page_path") to standard dimensions in a unified data model. However, due to cross-channel communication latency or anomalies at the data collection end, the mapped data often suffers from missing key features.
[0053] To ensure the completeness of information for subsequent large-scale model inferences, the system introduces a feature interpolation mechanism based on time-decay to dynamically predict and fill in missing fields. Specifically, for an event data point with a missing target feature, the system extracts the historical behavior sequence of the same user (based on an encrypted identifier) within the current time window. The missing feature to be estimated is defined as... Then, the interpolation fill formula based on the time step can be expressed as:
[0054] Equation (1) In the formula, This represents the total number of available observation samples containing homologous features in the user's historical behavior sequence. For a historic moment Actual observed values of homology characteristics; The time at which the currently missing event occurred; The preset time decay coefficient ( This refers to the pattern that the reference value of historical behavior decreases exponentially over time.
[0055] Therefore, through the mathematical interpolation model that incorporates time decay weights, the system can scientifically restore and complete missing business features without introducing subjective illusion noise, thereby significantly improving the richness and consistency of the underlying data.
[0056] In step S230, the graph feature identity parsing algorithm is used to perform cross-channel association on the encrypted identifiers contained in the filled event data, and the event data belonging to the same entity object are merged into the same user view to generate an integrated event stream with time series characteristics.
[0057] In multi-channel marketing scenarios, the same physical user often generates multiple fragmented encrypted identifiers on different devices or platforms (e.g., App_Hash_ID on mobile devices and Cookie_Hash_ID on desktop devices). To break down these "identity silos," this embodiment introduces a graph-based identity resolution algorithm. The system dynamically constructs an identity resolution graph network G=(V, E) in memory, with various encrypted identifiers as nodes and cross-channel co-occurrence relationships or deterministic mapping rules as edges. Then, by calculating the connectivity and edge weights between nodes, the system uses connected component extraction or community detection algorithms to cluster and merge multiple discrete nodes belonging to the same real user into a unified entity object (One-ID).
[0058] After completing cross-channel identity normalization, the system concatenates and merges the populated event data belonging to the same entity object strictly according to the chronological order of timestamps, ultimately outputting an integrated event stream with clear time-series characteristics. This eliminates record redundancy and identity fragmentation caused by cross-device data collection, providing downstream users with a comprehensive user-level view with complete context and continuous behavioral trajectories, thus improving the accuracy of marketing attribution and intent recognition.
[0059] In step S240, the data source identifiers corresponding to each event data in the integrated event stream are extracted, and combined with the historical verification performance data of the business domain corresponding to the data source identifiers, the basic source reliability weights for each data source are initialized.
[0060] Here, the system performs source weight initialization for each source channel in the integrated event stream. Specifically, the system precisely extracts the data source identifier corresponding to each event (e.g., originating from "self-operated mini-program" or "third-party distribution media") by parsing the message header or metadata of the integrated event stream. Subsequently, the system queries the pre-maintained business domain historical database to obtain the verification performance of the data source in past periods and assigns it an initial weight accordingly.
[0061] To avoid the "cold start" problem for newly connected data sources, the system adopts a weight initialization formula based on prior knowledge: Equation (2) In the formula, Indicates data source Initialize the base source reliability weights; This indicates the data source The pass rate of valid information determined to be real in the historical collection records (0 if it is a purely new data source). This indicates the industry default benchmark reliability of the business domain to which the data source belongs (such as social media or vertical e-commerce). This is a confidence adjustment factor for historical data, and its value is positively correlated with the amount of historical samples accumulated from the data source (the more abundant the historical samples, the higher the confidence level). The closer to 1).
[0062] Equation (2) above anchors an objective initial reputation baseline for various data sources, effectively avoiding the drastic disturbances caused by long-tail or emerging data sources to the downstream inference system, and providing a robust calculation starting point for subsequent source quality closed-loop feedback (i.e., dynamic update of source weights).
[0063] Through the embodiments of this application, firstly, by employing a server-side isolated environment and one-way hash encryption in tandem, strict privacy compliance requirements are thoroughly met at both the physical architecture and cryptographic levels, constructing a secure and compliant data collection moat. Secondly, by introducing a time-decay-based feature interpolation model and a graph feature identity parsing algorithm, the originally multi-source, heterogeneous, feature-incomplete, and identity-fragmented "dirty data" is cleverly transformed into a highly standardized, feature-complete, and coherent temporal logical high-quality panoramic data stream. Finally, combined with a domain-prior source weight initialization mechanism, an objective reputation benchmark is pre-positioned for the downstream multi-agent verification module. This filters out noise and gaps in the original data collection chain, reduces the ambiguity and computational load of large language models when performing multi-round inference, and lays a solid high-quality data foundation for accurate attribution and intelligent decision-making in digital marketing systems.
[0064] Figure 3 A flowchart illustrating an example of identifying and extracting candidate marketing information in a method according to an embodiment of this application is shown.
[0065] like Figure 3 As shown, in step S310, the event text of each event data in the integrated event stream is extracted and input into a pre-trained language model fine-tuned based on marketing domain labeled data for feature mapping to calculate the classification probability of the event text belonging to the marketing category; when the classification probability is greater than the preset semantic classification threshold, the event text is retained and marked as marketing text.
[0066] It should be understood that in actual integrated event flows, the data generated by user interactions is often filled with massive amounts of routine logs (such as "page loaded" or "user changed password") or casual chat text lacking commercial value. To achieve efficient noise reduction, this embodiment introduces a lightweight pre-trained language model (such as a lightweight encoder based on the Transformer architecture), which has been deeply fine-tuned in advance using a large amount of labeled data containing digital marketing corpora (such as various promotional copy and subtle sales pitches).
[0067] Specifically, the model will input the event text. Convert to high-dimensional dense feature vector The probability of classifying the text as belonging to a marketing intent is calculated using a non-linear classification head. Its computational logic can be expressed as:
[0068] Equation (3) In the formula, This represents the vector transpose operation. It is the Sigmoid activation function. and To fine-tune the weight matrix and bias terms of the classifier head, For event text The global semantic representation vector after being mapped by the encoder.
[0069] When this probability Strictly greater than the preset semantic classification threshold When the system determines that the text has potential commercial promotion intent, it is marked as marketing text. By introducing deep semantic representation based on dense vectors, the system can overcome the limitations of traditional keyword matching and accurately capture potential promotional information that uses obscure expressions, variant vocabulary, or new online marketing rhetoric. Thus, by adopting a semantic initial screening mechanism with "high recall," it effectively prevents the omission of valuable marketing intelligence while greatly filtering out invalid background noise.
[0070] In step S320, a preset marketing rule base is obtained. The marketing rule base is configured with a multi-dimensional marketing feature extraction template constructed from regular expressions and domain dictionaries. The multi-dimensional marketing feature extraction template includes at least the brand entity dimension, the price discount dimension, and the time limit dimension.
[0071] It should be noted that after the initial semantic screening, the remaining marketing texts, while possessing promotional intent, still exhibit a loose structure and lack business focus. To transform unstructured natural language into downstream verifiable structured statements, the system synchronously acquires and loads a pre-defined marketing rule base. This rule base integrates a wealth of prior digital marketing knowledge and business expert experience; its core can be a multi-dimensional marketing feature extraction template jointly constructed from a set of regular expressions and a dynamically updated domain dictionary.
[0072] Specifically, these templates are divided into multiple orthogonal business dimensions: for example, the brand entity dimension template uses a pre-built brand thesaurus and named entity recognition (NER) rules to anchor the "subject" in the text; the price discount dimension template uses regular expression combinations of "spend more, save more," "XX% OFF," or specific currency symbols to lock in the "benefit points" of promotions; and the time-limited dimension template is used to extract "time-limited" boundaries such as "limited to 24 hours" or "until a certain date." This establishes a rigid business analysis framework for the originally freely disseminated text data, enabling the system to accurately define and extract the three core business elements (Who, How much, When) of digital marketing intelligence from a unified machine perspective.
[0073] In step S330, multi-dimensional feature recognition is performed on marketing texts based on the multi-dimensional marketing feature extraction template to extract the set of marketing feature items contained in the marketing texts.
[0074] Specifically, after loading the multi-dimensional marketing feature extraction templates, the system's scheduling rule matching engine performs multi-channel sliding scans and feature matching on each data payload labeled as marketing text. The system applies the templates of each dimension to the target text in parallel and extracts the hit segments based on the maximum matching principle or sequence labeling mechanism, thereby decomposing a long string of complex marketing rhetoric into discrete and clearly categorized feature dictionaries.
[0075] For example, given the text "XX brand anniversary celebration, all skincare products up to 50% off, this weekend only," the feature extraction engine will output a structured set of marketing features. , Thus, through precise machine-executed analysis, rhetorical nonsense and emotionally charged words piled up in marketing copy to attract attention (such as "fire sale at a loss" and "miracle water") were completely eliminated, achieving a lossless purification of core factual statements.
[0076] In step S340, when the marketing feature set meets the preset feature completeness matching condition, character cleaning is performed on the marketing text to filter out noisy characters in the marketing text to obtain filtered marketing text. The filtered marketing text and its corresponding context metadata in the integrated event stream are combined and encapsulated to extract at least one candidate marketing information.
[0077] It should be noted that not all extracted feature sets are valuable for multiple rounds of verification in a large model. If a piece of text only contains "It's really cheap today" but lacks "brand entity" and "specific discount," it is essentially an incomplete statement that cannot be verified by facts. Therefore, the system introduces a feature completeness matching condition for final checkpoints. The feature completeness score of the current text is defined as follows: The calculation logic is as follows:
[0078] Equation (4) In the formula, This is a pre-defined set of core marketing dimensions (such as brand, price, time, etc.). In dimension The specific feature terms extracted below; This is an indicator function that takes the value 1 when the feature term is not empty (i.e., an entity has been successfully extracted), and 0 otherwise. Pre-assigned weight coefficients to different business dimensions (e.g., brands or product entities are typically required to be included, which have higher weights). This is reflected in the feature completeness score. When the preset minimum integrity threshold is reached, the system determines that the text has business value for independent verification.
[0079] Subsequently, the system performs low-level character cleaning on the text (such as removing HTML tags, invisible control characters, and meaningless consecutive emojis) to output clean, filtered marketing text. Finally, the system strongly binds and encapsulates this text with its contextual metadata in the integrated event stream (including timestamps, data source tracking identifiers, and anonymized user profile tags) to formally generate candidate marketing information and pushes it into the processing queue. Through quantitative completeness assessment, the system fundamentally eliminates the waste of valuable tokens and inference computing power of downstream multi-agent large models by meaningless incomplete data; at the same time, the encapsulation of metadata ensures the absolute traceability of this information in subsequent cross-module flow.
[0080] Through the embodiments of this application, firstly, by utilizing a pre-trained language model fine-tuned based on the marketing domain, semantic intent interception with extremely high generalization ability and recall rate is achieved in massive, high-noise concurrent event streams, avoiding the loss of implicit business intelligence; secondly, by introducing multi-dimensional regular expressions and dictionary templates that integrate industry prior knowledge, complex marketing copy is deconstructed step by step, supplemented by rigorous feature completeness quantification evaluation. This process forcibly filters out incomplete and invalid data lacking core verification elements. As a result, not only is highly unstructured, free-flowing text perfectly reshaped into high-density, high-purity, and clearly structured "candidate marketing assets," but it also significantly reduces the logical parsing difficulty and computational overhead of subsequent multi-agent inference modules while ensuring the integrity of data traceability. This significantly improves the processing efficiency and business hit accuracy of the entire digital marketing information collection system.
[0081] Figure 4 A structural block diagram of an example of a multi-agent inference module according to an embodiment of this application is shown.
[0082] like Figure 4 As shown, the multi-agent inference module 400 includes an understanding agent 410, a query generation agent 420, an evidence retrieval agent 430, and a decision agent 440.
[0083] Regarding the implementation details of the single-round processing of inferring confidence by calling the multi-agent inference module in step S130, in some examples of embodiments of this application, the multi-agent inference module 400 performs the following operations in each inference round: The understanding agent 410 uses a large language model to semantically deconstruct candidate marketing information and extract structured representation information containing marketing entities, promotion conditions and applicable scope.
[0084] In some implementations, the agent, based on pre-configured prompt engineering, leverages the few-shot learning capabilities of a large language model to parse raw, long text into a standardized set of key-value pairs (such as JSON format). For example, it can precisely break down complex advertising rhetoric into discrete, structured representations such as "marketing entity (e.g., specific product SKU or brand name)," "promotional conditions (e.g., minimum spending requirement, discount percentage)," and "applicable scope (e.g., specific city, limited membership level)." Thus, through the deep semantic understanding capabilities of the large language model, the system can effectively remove emotionally charged words and meaningless modifiers from marketing copy, transforming ambiguous natural language into machine-readable, logically clear, and independently verifiable factual statements.
[0085] The query-generating agent 420 generates query statements for targeted acquisition of external evidence based on structured representation information and the inference context of the current round.
[0086] In some implementations, after obtaining the structured representation information, the system passes it to the query generation agent 420. During the initial inference (first round), this agent directly maps the structured fields to a preset business retrieval template, generating a basic Boolean query string or API call parameters (e.g., "XX brand AND 100 off for every 500 spent AND promotional announcement"). In subsequent rounds of inference (the first round...)... Wheel, and When performing a search, the agent simultaneously reads the inference context of the current round, which includes contradictions or information blind spots discovered in the previous round of verification. Furthermore, the agent can dynamically adjust its search strategy based on this context, such as expanding the generalization of search keywords or adding restrictions based on specific timestamps.
[0087] Thus, the query-generating agent acts as an intermediary between the system's internal understanding and external world data. It can not only transform static structured data into query statements that are easy for search engines or databases to process, but also enable the system's evidence acquisition process to have an adaptive tracking capability of "following the clues" through dynamic response to "inference context", thereby ensuring the targeting and recall rate of external evidence acquisition.
[0088] The evidence retrieval agent 430 uses query statements to perform matching searches in pre-authorized knowledge bases and public network data sources to obtain relevant external evidence documents.
[0089] In some implementations, the evidence retrieval agent 430 is primarily responsible for executing the physical interaction between the system and external data sources. Upon receiving a query, this agent concurrently invokes internal and external retrieval interfaces. On one hand, the evidence retrieval agent 430 can penetrate the enterprise's pre-authorized knowledge base (e.g., internal ERP inventory systems, official CRM activity schedules) via API for precise matching. On the other hand, the evidence retrieval agent 430 can also utilize web crawler technology or search engine interfaces to conduct extensive searches within legitimate public online data sources (e.g., brand websites, details pages of mainstream e-commerce platforms). Furthermore, after obtaining the initial massive amount of returned web pages or document fragments, the evidence retrieval agent will use vector similarity algorithms or the BM25 algorithm for sorting and filtering, extracting the top-K text fragments with the highest relevance scores as external evidence documents.
[0090] Thus, it provides large language models with a real, objective, and current "external fact brain," fundamentally overcoming the "knowledge illusion" problem caused by outdated training data or memory generalization, and ensuring that all truth and falsehood judgments are based on objective factual evidence rather than the model's subjective conjecture.
[0091] The decision agent 440 compares the content of candidate marketing information with external evidence documents, calls the large language model to perform fact verification, outputs the single-round decision probability of the current inference round and the corresponding explanatory text, and feeds the corresponding explanatory text back to update the inference context of the next inference round.
[0092] As the core decision-making hub of the single-round inference chain, the decision agent 440 aggregates the structured information to be proven from the understanding agent and the external evidence documents from the retrieval agent. This agent utilizes the reasoning capabilities of the large language model to perform cross-comparison verification. The model is required to output not only a numerical single-round decision probability indicating whether the information is true or false. Alternatively, a Chain of Thought (CoT) approach can be used to output step-by-step logical deductions and points of conflict in evidence, forming an explanatory text.
[0093] Furthermore, in order to drive the system's deep reflection mechanism, the system will... Explanatory text generated by the wheel As an incremental state, feedback is appended and updated to the next state. In the inference context of the wheel (i.e., the context state is updated to) Thus, the decision agent 440 not only provides the quantitative confidence level for the current round, but also white-boxes the decision process; through the closed-loop feedback of interpreting the text, the large language model can carry the "questions" of the previous round to direct the retrieval and re-decision in the next round of processing, forming a "self-correction and reflection" loop similar to the repeated verification by human experts.
[0094] Then, when the preset inference termination condition is met, the multi-round inference ends, and the single-round decision probabilities output by each inference round are aggregated to calculate the multi-round inference confidence.
[0095] Here, to prevent multi-agent inference from falling into an infinite loop and to balance computational overhead, the system performs a conditional check at the end of each round. When the preset inference termination condition is met (e.g., reaching the system's allowed inference depth limit, or the model being confident that it has collected enough contradictory evidence), the system terminates the iteration loop. At this point, the system extracts each valid inference round before termination. Output single-round decision probability The system then calls a pre-defined mathematical aggregation function to perform time-series fusion calculations. For example, a weighted summation model can be used to converge the discrete single-judgment results from multiple rounds into a unified macro-level verification index, namely the final multi-round inference confidence. Thus, by using a scientific fusion algorithm to reduce and aggregate the results of multiple reflections, not only are the random errors that may exist in single inferences smoothed out, but an extremely stable and high-fidelity measure of factual probability is also provided for the downstream credibility comprehensive scoring module.
[0096] To further clarify the working mechanism of the multi-agent inference module 400 in multi-turn interactions, let's take a candidate marketing message captured by the system as an example: "Famous beauty brand A's anniversary celebration, all star serums are up to 30% off, and 500 exclusive coupons for internal employees will be issued for a limited time. Click the link below to grab them."
[0097] In the first round of inference, the understanding agent first deconstructs the unstructured text into a structured representation: {Brand: Brand A, Product: Star Essence, Promotional Price: 30% off, Additional Condition: Exclusive internal employee coupon}. Then, the query generation agent generates the initial search query "Brand A Star Essence 30% off internal employee coupon". The evidence retrieval agent searches the official online store and authorized database, returning preliminary evidence: "Brand A does have an anniversary sale today, but the highest discount is 20%, and no publicly distributed internal employee coupons were found." Based on this, the judgment agent performs fact verification and outputs the single-round judgment probability for the first round (e.g., ...). The explanatory text states: "The official maximum discount is 20%, which is seriously inconsistent with the claimed 30% discount, and no record of internal coupons has been found. This information is highly suspicious and is likely a fake referral link or phishing link."
[0098] Furthermore, the system feeds back the explanatory text from the first round and updates it to the inference context of the second round. At this point, the query generation agent reads the contextual hint that it is "probably a fake lead generation or phishing link," automatically adjusts its search strategy, and generates a targeted verification query: "Brand A - 30% off - Internal employee coupon - Fake activity - Debunking." The evidence retrieval agent uses this new instruction to conduct a secondary targeted search in a broader network domain or enterprise anti-fraud blacklist database, and successfully obtains a latest official announcement: Recently, criminals have been forging phishing links for "30% off internal employee coupons" to commit online fraud. Consumers are advised not to click on them. After receiving this decisive evidence, the judgment agent outputs a very low single-round judgment probability in the second round (e.g., The final explanation text reads: "We have obtained an official anti-fraud clarification announcement from Brand A, which confirms that the candidate marketing information is a malicious phishing fraud."
[0099] Since the second round of verification has yielded decisive and conclusive evidence, and the confidence gradient between adjacent rounds has converged, the system triggers the inference termination condition to end the loop. It then outputs the final, extremely low multi-round inference confidence by aggregating the single-round decision probabilities of these two rounds. As can be seen from the above example, the multi-agent multi-round inference module in this embodiment perfectly simulates the coherent logical loop of a human domain expert: "initial investigation reveals suspicious points—in-depth targeted review with these suspicious points—obtaining irrefutable evidence and finally determining the nature of the crime." It can extremely accurately identify advanced "deepfake" marketing texts that are logically well-disguised and difficult to distinguish with a single large model call, thus constructing a security firewall with deep thinking capabilities for the digital marketing data collection system.
[0100] Through the embodiments of this application, the aforementioned inference architecture based on multi-agent collaboration and multi-round iterative verification overturns the traditional single-model "black box" verification paradigm. The system decouples the complex task of verifying the authenticity of digital marketing data into four independent sub-tasks: semantic understanding, dynamic verification, evidence retrieval, and factual judgment, achieving professional large-model capability distribution. Simultaneously, it cleverly utilizes the closed-loop feedback of "interpretive text" to construct a multi-round reflection loop with dynamic gap-filling and self-correction capabilities. Thus, by employing a system of "multi-channel external intelligence connected to an external factual brain" combined with "cross-round logical self-reflection," it not only fundamentally suppresses the "information illusion" that large models are prone to when processing complex long-tail marketing information, but also, without any human intervention, greatly approximates the in-depth research level of human domain experts, thereby achieving accurate de-falsification of fragmented and highly disguised marketing data.
[0101] Regarding the implementation details of the inference confidence level in step S130 during multiple iterations, in some examples of embodiments of this application, firstly, during the multi-agent inference module's execution of multiple inferences, the current inference round is monitored in real time. Compared with the previous round of inference The confidence gradient of the single-round decision probability output between.
[0102] It should be noted that in a multi-agent collaborative system, the large language model continuously incorporates new external evidence and engages in self-reflection during multiple rounds of interaction. However, the model's inference results are not always monotonically increasing; they may exhibit probabilistic oscillations or gradually converge in the face of certain controversial evidence. To quantify the evolution trajectory of this inference state, the system incorporates a rigorous real-time monitoring mechanism in the background.
[0103] Specifically, after each round of inference, the system can extract the latest single-round decision probability output by the decision agent. And compared with the probability of judgment in the previous round Perform difference calculations to obtain the gradient of the absolute confidence level change. This change gradient objectively reflects the degree of influence of the newly added reflection rounds on the model's final decision.
[0104] Then, when it is determined that the current inference round has reached the preset maximum inference round... If the confidence gradient remains below a preset convergence threshold, the preset inference termination condition is confirmed, triggering a dynamic early termination mechanism to end the multi-round inference and recording the total number of inference rounds actually occurred. .
[0105] It should be noted that in the practical application of large models, continuously increasing the number of inference rounds will lead to significant increases in computational latency and token costs, and will follow the law of diminishing marginal returns. In order to find the optimal balance between "accuracy" and "computing cost", this embodiment designs a dynamic early termination mechanism with dual blocking.
[0106] More specifically, the first barrier is rigid cost control, that is, when the estimated number of rounds reaches the maximum number of rounds preset by the business. (For example, when set to 5 rounds), the loop is forcibly terminated to prevent it from getting stuck in an infinite loop or timeout; the second layer of blocking is intelligent convergence detection, which detects the confidence change gradient after a preset number of consecutive rounds (e.g., two consecutive rounds). When the convergence threshold is less than a very small value (e.g., 0.02), the system determines that the model's inference has reached an "information saturation" state, and further iteration will not provide any additional effective increments, thus triggering a termination mechanism. This endows the system with adaptive computing power allocation capabilities, avoiding the waste of resources caused by over-inference of clear and unambiguous marketing information, and ensuring that the system maintains excellent response efficiency even under extremely high concurrent requests.
[0107] Then, extract each valid inference round before the dynamic early termination mechanism is triggered. Output single-round decision probability and for each inference round Assign decay weights based on time steps Using attenuation weighting coefficients Probability of a single round of judgment Time-series weighted smoothing is performed, and the multi-round inference confidence is calculated using the following weighted multi-round confidence model. :
[0108] Equation (5) Equation (6) In the formula, It is a natural constant. To deduce the round number and , The total number of inference rounds that actually occurred. For the first The single-round decision probability of the output of a valid inference round. For multi-round inference confidence, For the first The decay weight coefficient corresponding to each inference round; The preset attenuation control parameters and This is used to control the rate at which weights decay as the number of inference rounds increases, in order to suppress inference drift and illusion amplification caused by excessive reflection in large models over multiple rounds.
[0109] A multi-round result fusion model based on time step decay was constructed using the above equations (5) and (6). In the self-correction mechanism of natural language processing, the initial inferences are often directly based on core facts and strongly related retrieval evidence, and the judgment basis is relatively solid; however, as the rounds deepen without limit, the large model may fall into excessive divergence on minor details, thereby producing "reasoning drift" or illusion amplification phenomena.
[0110] To scientifically integrate these discrete probability results, equation (6) introduces an exponential decay function. .when At that time, the weight of the initial round Dominating; with each round The increase in weight It exhibits an exponential, landslide-like decline, with the steepness of the decay determined by the hyperparameter. Strict control. Subsequently, equation (5) uses these calculated attenuation weights to determine the probability of output in each round. A robust global confidence index is obtained by performing weighted summation and normalization smoothing. In this way, a dynamic mathematical balance is established between "affirming the basic reliability of early inferences" and "absorbing the error-correcting gains from subsequent reflections". This not only smooths out the numerical jitter caused by a single abnormal inference, but also mathematically curbs the contamination of the final decision by the illusory noise accumulated in long-term inferences, ensuring that the output verification probability is extremely close to the objective fact.
[0111] This embodiment proposes a complete and highly automated quantitative control mechanism for multi-round iterative inference confidence calculation. It transforms the previously uncontrollable and resource-intensive multi-round reflection process of large language models into an adaptive system with precise mathematical boundaries. By dynamically terminating the system in real-time through monitoring the confidence gradient, redundant computation is perfectly prevented, achieving a significant reduction in computational power and time latency while maintaining judgment accuracy. Simultaneously, an exponentially decaying time-step weighted fusion model solves the inference drift problem easily induced by excessive iteration in large models. Thus, a closed-loop inference control mechanism that balances efficiency and accuracy is achieved, enabling the digital marketing fact-checking system to exhibit industrial-grade robustness and extremely high cost-effectiveness when processing massive amounts of heterogeneous and fragmented information.
[0112] Regarding the implementation details of calculating the comprehensive credibility score of candidate marketing information in step S140, in some examples of embodiments of this application, firstly, a set of marketing context constraint rules matching the business domain to which the candidate marketing information belongs is obtained. The set of marketing context constraint rules covers official promotional periods, regional whitelists, and applicable audience tag configurations.
[0113] In some cases, in the actual operation of digital marketing, whether a marketing message is "effective" depends not only on the truthfulness of its textual expression, but also on strict commercial time and space boundaries. For example, a genuine Double Eleven discount message, if captured in December, is still ineffective noise for current marketing decisions. Therefore, before proceeding to comprehensive scoring, this embodiment first dynamically retrieves a "set of marketing context constraint rules" that matches the business line to which the candidate information being processed belongs from the enterprise's customer relationship management system (CRM) or official event schedule database through a standard data interface.
[0114] More specifically, these constraints constitute the rigid boundaries of marketing activities, encompassing officially defined promotional activation and deactivation times (official promotional periods), permitted physical or IP geographic regions for participation (regional whitelists), and defined user group characteristics (such as applicable audience tag configurations for "newly registered users only" or "specific membership levels"). This introduces an "objective business perspective" beyond pure textual semantics into the system, making the validation of the large language model no longer an isolated fact-checking process, but rather a business logic validation deeply rooted in the actual operational norms of the enterprise.
[0115] Then, the structured representation information parsed from the candidate marketing information in the multi-agent inference module is extracted, and multi-dimensional business conflict detection is performed to verify whether the promotion conditions and applicable scope in the structured representation information meet the marketing context constraint rule set; if they match, a context with a value of 1 is generated via a flag bit. If a rule conflict exists, an invalid marketing warning will be triggered, and a context flag with a value of 0 will be generated. .
[0116] In some implementations, the system can directly reuse the structured representation information (such as extracted key-value pairs like specific discount rates, applicable regions, and expiration dates) that has been precisely extracted by the "understanding agent" in the preceding steps, and perform a multi-dimensional cross-comparison with the aforementioned set of marketing context constraint rules. Since both sides of the comparison are highly structured, this step can be completed with extremely low computational latency.
[0117] In terms of validation logic, the system can employ a strict Boolean gating mechanism. If all conditions of the information fall within the officially allowed rule range (matching), the system assigns a context pass flag with a value of 1. If any boundary is crossed in any dimension (for example, the activity is marked as "available to all," but the official rules limit it to "VIPs only"), the system will immediately determine a rule conflict, trigger an internal marketing invalidity warning log, and generate a context flag with a value of 0. Therefore, through lightweight structured rule comparison, a rigid logical firewall is constructed, capable of quickly identifying and isolating "partially true" marketing fraud caused by outdated information, geographical misalignment, or audience manipulation. Simultaneously, complex business judgments are condensed into a single, extremely simple binary control variable. This laid the foundation for subsequent mathematical dimensionality reduction and integration.
[0118] Then, based on the extraction of the first The event data of each candidate marketing message determines the set of source data sources that provide that candidate marketing message. And query the current source reliability weight of each data source in the source data source set.
[0119] It should be noted that in an open network environment, a highly popular digital marketing message is often forwarded and collected simultaneously by multiple different media channels or terminal nodes. To eliminate the bias of a single data source, the system traces the current event being processed based on the source tags of the underlying event stream. This involves identifying all physical data collection channels for each candidate marketing message, thereby constructing a set of data sources for that message. (For example, the information may come from "social media platform A", "price comparison website B" and "offline store C" at the same time).
[0120] Next, the system accesses the global source quality configuration table and queries it in real time. The source reliability weight of each data source. This weight is an objective reputation indicator that the system continuously maintains based on the historical performance of each data source, extending the verification perspective from "the information content itself" to "the carrier of information dissemination." Therefore, by aggregating the channel reputation of multiple sources of information from the same source, the system can effectively identify false intelligence fabricated in isolation by a single, low-quality "spam farm" or malicious web crawler network, thus providing evidence of the information's authenticity at the physical link level.
[0121] Furthermore, by combining context and using markers Confidence of multi-round inference The overall credibility score is calculated using the following credibility fusion scoring formula, which includes the current source reliability weights of each data source in the source data source set. : Equation (7) In the formula, This is the sequence number of the candidate marketing information currently being processed. This is the index number of a single data source within the collection of source data sources; For the first The overall credibility score of the candidate marketing information For the first The context of each candidate marketing message is conveyed through flags. For data source Current source reliability weights, Indicates the collection of source data. The number of sources in; To adjust the balance coefficient between the source reliability and the model inference confidence ratio and .
[0122] Formula (7) above designs a comprehensive evaluation model that integrates deterministic gating and probabilistic weighting. (The brackets in formula (7) are...) Internally, the system utilizes balance coefficients. Linearly weight the probability indices of the two dimensions: the first half The average reliability weight of all data sources providing this information was calculated, representing "channel-level physical reputation"; the latter half... This refers to the "logical credibility at the content level" derived from semantic cross-validation by the multi-agent large model. The value can be adjusted according to business characteristics. For example, if the system is more vigilant against online troll attacks, then the value can be increased. Increase the weight of channel allocation.
[0123] In addition, the context representing business rules is conveyed through flag bits. It was extracted to the outermost layer of the entire formula and participated in the calculation as a global multiplier. This mathematical expression constructs an absolutely safe "one-vote veto" mechanism: no matter how authoritative the source of a piece of information is ( (Extremely high), and no matter how flawlessly the text content is judged by the large model ( (Extremely high), as long as it violates the company's pre-set rigid promotional context (such as incompatible region, expired promotion), leading to Then all the high scores within the brackets will be instantly reset to zero, resulting in the final score. This forces the divergent probabilities of the large model to be contained within the enterprise's compliance control framework, ensuring that the high-scoring output simultaneously meets the three stringent conditions of "reliable source, logical truth, and business compliance," thus greatly improving the security and availability of the final decision data.
[0124] This application's embodiments overcome the vulnerability of traditional AI verification frameworks that rely solely on textual features to determine authenticity, transforming the hard business context constraints unique to digital marketing, such as "time, space, and audience," into global gating factors in mathematical calculations. By designing a "one-vote veto" mechanism through a formulaic model, the system can not only effectively resist expired and invalid information disseminated by high-weight channels, but also intercept high-scoring misjudgments caused by illusions in large models. This achieves strong oversight of the probabilistic deep learning black box by deterministic business rules, fundamentally ensuring that every piece of marketing data ultimately output to the enterprise's core decision-making system is an objective, fresh, and compliant high-value data asset.
[0125] In some examples of embodiments of this application, after outputting candidate marketing information that has reached the credibility threshold as target marketing data, a source quality closed-loop feedback operation can also be performed.
[0126] Specifically, firstly, according to the preset sliding time window Statistical sliding time window Internal data source The proportion of candidate marketing information whose overall credibility score reaches the credibility threshold is used as the basis for determining the data source. Recent true quality observation ratio : Equation (8) In the formula, Indicates the sliding time window Internal source: data source The total collection of candidate marketing information, Represents the total set The number of samples; As a credibility threshold, For indicator functions, when the condition The value is 1 when it is true, and 0 otherwise.
[0127] It should be noted that in real-world open network environments, the quality of data sources is not constant. For example, some originally high-quality official channels or partner media may be flooded with a large amount of spam data containing false discounts in a short period of time due to hacker attacks, account theft, or adjustments to system crawler strategies. To accurately capture such dynamic changes, the system does not treat a single verification task as the end point after outputting target marketing data. Instead, it treats the verification results as a high-value feedback signal, and applies them according to a preset sliding time window. Conduct periodic reviews.
[0128] Specifically, the system will aggregate specific data sources within that time window. Submit all candidate information and utilize indicator functions. Filter out the "overall credibility score" Successfully reached or exceeded the credibility threshold The high-quality samples are obtained through mathematical calculation using equation (8), which calculates the ratio of the number of high-quality samples to the total base number of samples submitted by the data source. From this ratio, the system derives an objective indicator purely determined by recent actual performance—the proportion of recent true quality observations. This provides the system with an "automated business health monitor" that is free from human subjective experience intervention, enabling the system to accurately and without being misled by historical performance to perceive the latest information inflation and true conversion rate of various underlying data sources.
[0129] Furthermore, based on the exponential smoothing algorithm, the proportion of recent true quality observations is used. Update data source In the next time window Source reliability weight : Equation (9) In the formula, For data source In the time window The current source reliability weight, The updated source reliability weights; The smoothing coefficient and .
[0130] It should be noted that if the original weights are directly replaced with the recent observation ratios, the system will be highly susceptible to short-term data fluctuations, resulting in severe oscillations (for example, an occasional interface failure on a certain day causing data truncation, leading to a sharp drop in the pass rate on that day). To ensure the continuity and stability of the source weight evaluation, this embodiment introduces the exponential smoothing algorithm from time series analysis into the weight update mechanism.
[0131] In equation (9), the smoothing coefficient is used. As a regulatory lever, the current weight representing "historically accumulated credibility" will be used. The proportion of recent observations representing "the latest performance feedback" Perform weighted fusion. Smoothing coefficient. This determines the system's "memory depth" of the data source's historical reputation; if the business scenario requires extreme sensitivity to fraud risk, it can be appropriately lowered. This is to amplify the penalties for recent misconduct.
[0132] Through the embodiments of this application, the high-dimensional verification results of multi-round inference from a large model are used to inversely empower the reputation management of the underlying data source. By introducing a sliding time window and an exponential smoothing algorithm, the system can not only quantify the dynamic quality evolution of each collection channel in real time and objectively, but also automatically update the weights with an extremely stable mathematical model. As a result, the entire digital marketing information collection system possesses a "self-purification" capability similar to a biological immune system—as the running time increases, the system automatically amplifies the influence of high-quality data sources and continuously deprives high-noise and low-quality data sources of their influence, thereby achieving a data collection foundation that becomes more accurate and better with use at the macro-architectural level.
[0133] In some examples of embodiments of this application, after updating the source reliability weight of the data source, a circuit breaker interception operation can also be performed for abnormal data sources.
[0134] More specifically, firstly, if a data source is detected within the preset continuous monitoring period... Updated source reliability weights If the risk level is below the preset risk blocking threshold, a data source will be generated. The circuit breaker command.
[0135] It should be noted that in the complex open network ecosystem, the security threats faced by the system include not only occasional low-quality data, but also malicious and organized attacks (such as data pollution by marketing black market operators using automated scripts). The aforementioned source reliability weight update mechanism (such as exponential smooth decay) is a kind of "soft weight reduction" strategy, which can dynamically reduce the information weight when the quality of the data source fluctuates. However, when a data source is completely compromised or evolves into a pure garbage farm, "soft weight reduction" alone will still allow a massive amount of useless data to flood into the system, resulting in serious network bandwidth consumption and wasted computing power for large models. Therefore, the system introduces a hard "circuit breaker" mechanism.
[0136] Specifically, the system has a continuous monitoring daemon running in the background to smoothly track the time-series weight curves of various data sources. To avoid false positives caused by single network fluctuations or accidental errors, the system has implemented a continuous monitoring cycle. Risk blocking threshold The system employs a dual-determination logic. The logical condition for the system to determine whether to trigger a circuit breaker instruction can be represented by the following Boolean function:
[0137] Equation (10) In the formula, For data sources The circuit breaker trigger flag is set to 1 when a circuit breaker command is generated. The preset continuous monitoring period is used to filter high-frequency transient jitter; The preset risk blocking threshold is (which is usually much lower than the aforementioned credibility threshold). Thus, by introducing a continuous judgment over time and a security baseline with extremely low tolerance, the system can accurately and stably identify malicious high-risk noise sources that have been completely degraded, preventing the soft weight reduction mechanism from failing in extreme attack scenarios.
[0138] Then, based on the circuit breaker command, the data source is extracted. The data source identifier is then written into the dynamic blocking blacklist.
[0139] Specifically, once a data source is generated... Upon receiving the circuit breaker command, the system control plane immediately initiates a security isolation procedure. The system first parses the context registration information of the abnormal data source, extracting its unique and unforgeable data source identifier (e.g., a specific API Key, application layer App ID, or a specific edge node IP range). Subsequently, the system distributes this data source identifier and writes it to a dynamic blocking blacklist configured in the memory of the server-side front-end gateway or Web Application Firewall (WAF). This blacklist employs a high-speed caching architecture based on Redis or a similar in-memory database, supporting microsecond-level retrieval response and dynamic expiration time (TTL) configuration. Thus, the penalty action against abnormal data sources is shifted from "backend algorithmic downgrading" to the "front-end network defense" level, achieving second-level full-network blocking synchronization of high-risk identifiers using a high-speed in-memory blacklist mechanism.
[0140] Furthermore, based on a dynamic blacklist, it blocks and rejects reception originating from the data source. Information on subsequent basic events.
[0141] Here, after the blacklist takes effect, the system's data collection architecture enters a proactive defense state. When the receiving tag (Server-side Container) pre-deployed in the server-side isolation environment faces another data inflow request, it will perform a hash collision comparison between the source identifier carried by the tag and the dynamically blocked blacklist during the outermost protocol parsing stage (such as the HTTP request header verification stage). If the comparison matches the target data source... If the server detects the abnormal data source, it will either discard the data packet at the network access layer or return a denial-of-service status code (such as an HTTP 403 alert) to the sender, thereby completely intercepting and refusing to receive any subsequent basic event information from the abnormal data source, cutting off the injection channel of low-quality marketing information at the lowest physical access link.
[0142] This application's embodiments link time-series-based "algorithm-level reputation assessment (source reliability weight)" with "network-level physical isolation (server-side dynamic blacklist)" across layers. When the quality of the data source deteriorates to the point of breaching the security baseline, the system no longer relies solely on backend probability filtering but decisively triggers a circuit breaker mechanism, directly severing the physical link of data injection at the edge gateway. Thus, by adopting a proactive, hard-line interception design, the malicious consumption of massive amounts of junk data on the system's backend large-scale computing power pool (such as token exhaustion attacks) is completely prevented, endowing the entire intelligent information collection platform with extremely strong anti-attack resilience and industrial-grade computing resource protection capabilities.
[0143] Figure 5The diagram illustrates the operational principle of an example of a digital marketing information collection method based on a large model and multi-round inference verification according to an embodiment of this application. The architecture mainly consists of four collaborative modules: an input area, a core processing area (divided into stage 1 and stage 2), an output area, and a bottom control area.
[0144] like Figure 5 As shown, during the data access process, the system is constrained by the consent management platform in the input area and the compliance scheduling of the privacy and ethics controller in the bottom area to securely receive server-side event data from multiple channels. Subsequently, the data flows into the core processing area in stage 1, where underlying data preprocessing and unified mapping are performed, and candidate extraction is completed through a dual filtering funnel, purifying the massive, disordered raw logs into structured candidate marketing information.
[0145] The extracted candidate marketing information then enters Phase 2 of the core processing area, namely the multi-agent inference loop system. In this loop, the understanding agent, query generation agent, evidence retrieval agent (directly connecting to external data sources), and judgment agent work closely together to perform multi-round fact verification with a closed-loop reflection mechanism. Subsequently, the multi-round inference confidence output by the large model, together with the marketing context constraints provided by the bottom area and the underlying channel reputation, are all incorporated into the "credibility fusion assessment" node for dimensionality reduction and veto calculation. Finally, high-value information that successfully crosses the credibility threshold is securely delivered to the verified marketing database in the output area; simultaneously, based on the results of the fusion assessment, the system generates source reliability feedback in the output area and writes abnormal channels that trigger circuit breaker conditions into the dynamic blacklist at the bottom. This blacklist is looped back to the input area, directly intercepting the subsequent injection of high-risk server-side event data from the physical access link, thereby constructing a perfect technical closed loop in the global architecture from compliant collection, intelligent counterfeit identification to security defense.
[0146] To comprehensively verify the effectiveness of the proposed method and its boundary conditions in complex network environments, this embodiment constructs a semi-synthetic simulation test dataset (DM-VeriBench) based on real marketing data features. This experiment aims to answer three core questions: 1) Can the multi-round inference and source reliability weighting mechanism effectively combat interference from low-quality data sources? 2) Is there an optimal balance between accuracy improvement and computational resource consumption? 3) What is the impact of introducing privacy compliance and marketing context constraints on the final high-value information recall rate?
[0147] Considering the privacy compliance restrictions of collecting real marketing data, the DM-VeriBench test set constructed in this experiment contains 2,500 samples and specifically introduces "context-dependent" test cases. The specific data distribution covers: 40% fact-consistent samples (i.e., explicit promotional information that perfectly matches official announcements), 30% fact-conflicting or forged samples (including tampered prices, incorrect dates, or forged product entities), 20% context-conflicting samples (the information content itself is true, but invalid under a specific compliance or business context, such as a coupon limited to region A being distributed to the data stream of city B; this type of sample is specifically used to verify the veto power of the "marketing context constraint" module), and 10% long-tail noise samples (with messy formatting and containing a large amount of non-marketing text, used to test the robustness of the front-end semantic filtering preprocessing).
[0148] Regarding the experimental control group setup, to objectively evaluate the core technical contributions of this scheme, three parallel baseline comparison schemes were set up: the first group was baseline A (Zero-shot LLM), which only used a large language model for a one-time zero-shot judgment, without any external evidence retrieval or multi-agent collaboration mechanism; the second group was baseline B (ReActAgent), which adopted a standard large model agent framework with external retrieval capabilities, but stripped the "multi-round iterative reflection" mechanism of this invention and the "source reliability weight" evaluation detached from the text level; the third group was the complete method proposed in this application (Ours), which enabled complete decaying multi-round inference based on multi-agent collaboration (setting a maximum number of inference rounds). Attenuation control parameters It also fully implemented a source reliability weight dynamic update mechanism based on exponential smoothing and context constraint gating.
[0149] In constructing the evaluation index system, this experiment comprehensively measured the system from three dimensions: system performance, computing cost, and operational stability. The core performance indicators include the F1 score (the harmonic mean of precision and recall, which comprehensively evaluates the system's accuracy) and trust alignment (a measure of the statistical correlation between the model's multi-turn inference confidence and objective facts). Cost and expense indicators focus on the average large model token consumption (Avg. Token Cost) and end-to-end system processing latency during the processing of a single marketing message. Furthermore, to verify whether the multi-turn reflection mechanism effectively suppresses inference drift, the experiment also introduced variance across runs as a key quantitative indicator for measuring the stability of the system's underlying inference.
[0150] Figure 6The diagram illustrates an example of how accuracy and computational cost evolve with the number of inference rounds in a multi-round inference mechanism according to an embodiment of this application. The dual-axis simulation visually demonstrates the effect of inference rounds... The trend of overall system performance and resource consumption as the value increases from 1 to 5 is shown. The line graph (corresponding to the left vertical axis and including the error band) represents the harmonic mean of the system precision and recall (i.e., the F1 score), and the bar graph (corresponding to the right vertical axis) represents the average processing time of the system.
[0151] like Figure 6 As shown, regarding the evolution of accuracy, when the inference rounds... In stages 1 to 3, the F1 score of this method shows a significant upward trend (rapidly jumping from approximately 0.76 to approximately 0.86), which fully verifies that the multi-agent execution of multiple rounds of "understanding-retrieval-decision" can effectively correct information illusions and errors in early inference. However, when the number of inference rounds... Subsequently, the F1 score curve flattens out and enters a plateau. This aligns with the natural law of information entropy convergence, meaning that for the vast majority of clear marketing information, about three rounds of inference are sufficient to obtain adequate evidence and converge conclusions. Excessive rounds only have a weak corrective effect on a very small number of extremely complex ambiguous samples.
[0152] Meanwhile, the average system latency, representing computational cost, exhibits a strictly linear growth trend. This reveals a key technical trade-off in multi-round inference mechanisms: to achieve only a 1% to 2% increase in accuracy after the third round, the system needs to incur nearly double the computational cost and time. Based on the pattern revealed by these simulation results, Figure 6 The Chinese label is marked The "best dessert area" that balances performance and cost.
[0153] To adaptively capture the sweet spot in actual dynamic deployment, this application introduces the aforementioned dynamic early termination mechanism. The system monitors the confidence gradient between two adjacent rounds in real time, and when the gradient ( Below a very small convergence threshold (e.g.) When the system reaches its peak, it automatically determines that revenue has peaked and terminates the inference loop prematurely. Thus, the digital marketing data collection system can achieve the highest fact-checking accuracy while effectively avoiding unnecessary computational waste and latency spikes.
[0154] Figure 7This diagram illustrates an experimental simulation of the source reliability weight dynamic update mechanism according to an embodiment of this application in combating data source noise interference. The three-dimensional surface plot visually demonstrates the evolution of the system's final judgment accuracy under the influence of the average noise level of the data source and the core algorithm mechanism. The X-axis represents the average noise level of the data source (ranging from 0% to 60%); the Y-axis represents whether the proposed source reliability weight dynamic update mechanism is enabled (where 0 represents baseline B, the baseline control group, and 1 represents the method with the mechanism enabled); and the Z-axis and the corresponding surface color depth represent the final judgment accuracy of the system output.
[0155] By viewing as follows Figure 7 As shown in the distribution trend of the three-dimensional surface, in a low-noise "clean environment" where the average noise level of the data source is less than 10%, the performance of the baseline control group and this method is not significantly different. The accuracy of both is in the same warm color range with high scores, and their performance shows a trend of convergence.
[0156] However, as the average noise level of the data sources climbed above 40% (to simulate a real and complex open network environment), the system performance of the two models diverged significantly. Because the baseline control group (Baseline B) lacked dynamic consideration of the underlying physical link reputation, the large language model could not effectively distinguish between high-quality external evidence and mass-generated spam content, resulting in a "cliff-like drop" in its accuracy curve, as indicated in the figure. In contrast, this method, because the system strictly implements the source reliability weights calculated based on the exponential smoothing formula (…),… The dynamic update mechanism enables the system to sensitively and automatically reduce the weights of low-confidence sources that continuously inject false information. Therefore, even under extreme high-noise interference, this method maintains a smooth and high-performing performance surface. This simulation comparison strongly demonstrates that the source weight fusion mechanism in this method possesses extremely strong anti-interference robustness and system stability in complex business environments characterized by severe data fragmentation and rampant malicious online attacks.
[0157] Figure 8 A simulation diagram illustrating the overall performance of different digital marketing information collection methods is shown. The five-dimensional radar chart visually demonstrates the system-level performance comparison between our method (Ours) and the baseline control groups (Baseline A and Baseline B) across five core evaluation dimensions, including accuracy, contextual compliance, robustness to interference, inference speed, and cost-effectiveness.
[0158] like Figure 8As shown in the radar chart, the outward expansion trend reveals that this method significantly outperforms the baseline control group in three dimensions: "accuracy," "contextual compliance," and "anti-interference capability," demonstrating a remarkable technical advantage. Particularly in the "contextual compliance" dimension, thanks to the innovative marketing context constraint mechanism and veto gating algorithm introduced in this invention, the method significantly reduces the identification error rate by 45% compared to traditional methods when handling complex contextual conflict samples. Furthermore, through multi-agent collaborative multi-round reflection verification and exponentially smoothed source reliability weight updates, this method also builds a strong technical barrier in terms of "accuracy" and "anti-interference capability," effectively addressing advanced misinformation fraud and interference from spam data farms in open networks.
[0159] However, objectively speaking, the performance metrics of this method exhibit a certain degree of inward contraction in terms of "inference speed" and "cost-effectiveness." For example, due to the system performing multiple large language model calls and external evidence cross-referencing, the average inference latency of this method is approximately 3.5 times slower than Baseline A, which only performs a single zero-sample direct inference. This reveals the technical trade-off between the present invention and achieving ultimate decision accuracy and business security by moderately sacrificing underlying computing resources and time efficiency. Based on the performance in the above dimensions, it can be concluded that the digital marketing information collection and verification architecture proposed in this application is most suitable for high-level digital marketing scenarios with extremely stringent requirements for data authenticity and extremely low risk tolerance (such as high-value brand reputation protection, anti-fraud verification for large-scale promotions, and deep attribution analysis). By intercepting inferior decision data in complex scenarios, this system can recover huge resource misallocation losses caused by false data from a business-wide perspective, thereby achieving a high return on investment (ROI) at the macro level.
[0160] In summary, the digital marketing information collection method and system based on multi-round inference verification using a large model proposed in this application solves the core pain points of information fragmentation, difficulty in distinguishing truth from falsehood, and susceptibility to contamination by high-risk nodes in the traditional marketing collection chain by constructing a global architecture of "data compliance mapping, multi-agent reflective inference, business context gating, and source quality physical feedback." A series of simulation experiments fully demonstrate that the proposed method not only exhibits industrial-grade anti-interference robustness in extremely noisy network environments, but also significantly reduces the compliance error rate in complex business scenarios by 45% through an original "one-vote veto" context gating mechanism. At the same time, combined with a dynamic early termination mechanism based on confidence gradient, the system achieves a highly adaptive dynamic balance between pursuing the ultimate factual accuracy and controlling the massive concurrent computing power cost of the large model, building an intelligent digital marketing data foundation for enterprises that becomes more accurate with use.
[0161] In terms of future practical industrial applications and architectural evolution, with the maturity of multimodal large language model technology, the multi-agent collaborative verification architecture revealed in this invention has strong horizontal scalability potential and can be smoothly migrated to the cross-verification tasks of cross-modal marketing materials (such as promotional poster images and short video broadcast streams). Furthermore, in terms of system deployment and computing power optimization, future development can involve introducing model distillation and knowledge compression technologies to deploy some lightweight query and retrieval agents to edge computing nodes. This will further compress the inference latency across the entire chain while maintaining the existing high-dimensional verification confidence, fully empowering a real-time intelligent marketing decision-making ecosystem with millisecond-level response requirements.
[0162] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of combined actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0163] Figure 9 A structural block diagram of an example of a digital marketing information collection system based on large model multi-round inference verification according to an embodiment of this application is shown.
[0164] like Figure 9 As shown, the digital marketing information collection system 900 based on large model multi-round inference verification includes an authorized collection and integration unit 910, a marketing semantic filtering unit 920, a multi-agent inference verification unit 930, a credibility fusion evaluation unit 940, and a threshold decision output unit 950.
[0165] The authorized collection and integration unit 910 is used to acquire multi-channel event data after obtaining the user's authorization and consent, and to map the multi-channel event data to a unified data model to generate an integrated event stream.
[0166] The marketing semantic filtering unit 920 is used to perform semantic classification on the event text in the integrated event stream to filter out marketing text, and to perform feature matching filtering on the marketing text based on a preset marketing rule base, thereby identifying and extracting at least one candidate marketing information.
[0167] The multi-agent inference verification unit 930 is used to input the candidate marketing information into a pre-built multi-agent inference module based on a large language model for each candidate marketing information. The multi-agent inference module performs external evidence retrieval for the candidate marketing information and performs multiple rounds of inference and truth / falseness determination in combination with the retrieved external evidence to output the multi-round inference confidence for the candidate marketing information.
[0168] The credibility fusion evaluation unit 940 is used to combine the multi-round inference confidence and the source reliability weight of the data source corresponding to the candidate marketing information to calculate the comprehensive credibility score of the candidate marketing information.
[0169] The threshold decision output unit 950 is used to compare and filter the comprehensive credibility score based on a preset credibility threshold, and output the candidate marketing information that reaches the credibility threshold as target marketing data.
[0170] In some embodiments, this application provides a non-volatile computer-readable storage medium storing one or more programs including execution instructions. The execution instructions can be read and executed by electronic devices (including but not limited to computers, servers, or network devices) to perform the steps of any of the above-described digital marketing information collection methods based on large model multi-round inference verification of this application.
[0171] In some embodiments, this application also provides a computer program product, the computer program product including a computer program stored on a non-volatile computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the steps of any of the above-described digital marketing information collection methods based on large model multi-round inference verification.
[0172] In some embodiments, this application also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of a digital marketing information collection method based on large model multi-round inference verification.
[0173] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.
[0174] The electronic devices in this application can exist in various forms, including but not limited to: mobile communication devices, ultra-mobile personal computer devices, portable entertainment devices, or other airborne electronic devices with data interaction functions.
[0175] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0176] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A digital marketing information collection method based on large-scale model multi-round inference and verification, characterized in that, The method includes: With the user's authorization and consent, multi-channel event data is acquired and mapped to a unified data model to generate an integrated event stream; The event text in the integrated event stream is semantically classified to filter out marketing text, and the marketing text is filtered by feature matching based on a preset marketing rule base to identify and extract at least one candidate marketing information. For each candidate marketing information, the candidate marketing information is input into a pre-built multi-agent inference module based on a large language model. The multi-agent inference module performs external evidence retrieval for the candidate marketing information and performs multiple rounds of inference and truth / falseness determination based on the retrieved external evidence to output the multi-round inference confidence for the candidate marketing information. By combining the confidence scores from the multiple rounds of inference with the source reliability weights of the data sources corresponding to the candidate marketing information, a comprehensive credibility score for the candidate marketing information is calculated. The comprehensive credibility score is compared and filtered based on a preset credibility threshold, and candidate marketing information that reaches the credibility threshold is output as target marketing data.
2. The method according to claim 1, characterized in that, The step of acquiring multi-channel event data after obtaining user authorization and consent, and mapping the multi-channel event data to a unified data model to generate an integrated event stream, includes: In response to the collection instruction triggered by the client based on the authorized consent status, basic event information is received through a pre-deployed server-side tag. Sensitive fields are desensitized and one-way hash encryption is performed on the user identifier contained in the basic event information in the server-side isolated environment to generate an encrypted identifier and generate the multi-channel event data that meets privacy compliance requirements. Extract heterogeneous fields from the multi-channel event data, map the heterogeneous fields to the standard dimensions of the unified data model, and fill the missing fields after mapping with feature interpolation based on the same source data completion rules and user historical behavior sequences to obtain the filled event data. The encrypted identifier contained in the filled event data is associated across channels using a graph feature identity parsing algorithm, and event data belonging to the same entity object are merged into the same user view to generate an integrated event stream with time series characteristics. Extract the data source identifiers corresponding to each event data in the integrated event stream, and combine them with the historical verification performance data of the business domain corresponding to the data source identifiers to initialize the basic source reliability weights for each data source.
3. The method according to claim 2, characterized in that, The process involves semantically classifying the event text in the integrated event stream to filter out marketing text, and then performing feature matching filtering on the marketing text based on a preset marketing rule base to identify and extract at least one candidate marketing information, including: Extract the event text from each event data in the integrated event stream, and input the event text into a pre-trained language model fine-tuned based on marketing domain labeled data for feature mapping to calculate the classification probability of the event text belonging to the marketing category; when the classification probability is greater than a preset semantic classification threshold, the event text is retained and marked as marketing text; Obtain a preset marketing rule base, which is configured with a multi-dimensional marketing feature extraction template constructed from regular expressions and domain dictionaries. The multi-dimensional marketing feature extraction template includes at least the brand entity dimension, price discount dimension, and time limit dimension. Based on the multi-dimensional marketing feature extraction template, multi-dimensional feature recognition is performed on the marketing text to extract the set of marketing feature items contained in the marketing text; When the set of marketing features meets the preset feature completeness matching condition, character cleaning is performed on the marketing text to filter out noisy characters in the marketing text to obtain filtered marketing text. The filtered marketing text and its corresponding context metadata in the integrated event stream are combined and encapsulated to extract at least one candidate marketing information.
4. The method according to claim 3, characterized in that, The multi-agent inference module includes an understanding agent, a query generation agent, an evidence retrieval agent, and a judgment agent; The multi-agent inference module performs external evidence retrieval for the candidate marketing information and combines the retrieved external evidence to perform multiple rounds of inference and truth / falseness determination to output the multi-round inference confidence level for the candidate marketing information, including: In each inference round, the multi-agent inference module performs the following operations: The understanding agent uses a large language model to semantically deconstruct the candidate marketing information and extract structured representation information containing marketing entities, promotion conditions, and applicable scope. The query-generating agent generates a query statement for targeted acquisition of external evidence based on the structured representation information and the inference context of the current round. The evidence retrieval agent uses the query statement to perform matching searches in a pre-authorized knowledge base and public network data sources to obtain relevant external evidence documents. The decision-making agent compares the content of the candidate marketing information with the external evidence document, calls the large language model to perform fact verification, outputs the single-round decision probability of the current inference round and the corresponding explanatory text, and feeds back the corresponding explanatory text to update the inference context of the next inference round; The multi-round inference ends when the preset inference termination condition is met, and the single-round decision probabilities output from each inference round are aggregated to calculate the multi-round inference confidence.
5. The method according to claim 4, characterized in that, The process of ending multi-round inference when a preset inference termination condition is met, and aggregating the single-round decision probabilities output from each inference round to calculate and generate multi-round inference confidence, includes: During the multi-agent inference module's execution of multiple rounds of inference, the current inference round is monitored in real time. Compared with the previous round of inference The confidence gradient of the single-round decision probability output between; When it is determined that the current inference round has reached the preset maximum inference round. If the confidence gradient remains below a preset convergence threshold, the preset inference termination condition is confirmed to be met, thereby triggering a dynamic early termination mechanism to end multiple rounds of inference, and recording the total number of inference rounds actually occurring. ; Extract each valid inference round before the dynamic early termination mechanism is triggered. Output single-round decision probability and for each inference round Assign decay weights based on time steps ; Using the attenuation weighting coefficient The single-round determination probability Time-series weighted smoothing is performed, and the multi-round inference confidence is calculated using the following weighted multi-round confidence model. : , , In the formula, It is a natural constant. To deduce the round number and , The total number of inference rounds that actually occurred. For the first The single-round decision probability of the output of a valid inference round. For multi-round inference confidence, For the first The decay weight coefficient corresponding to each inference round; The preset attenuation control parameters and This is used to control the rate at which weights decay as the number of inference rounds increases, in order to suppress inference drift and illusion amplification caused by excessive reflection in large models over multiple rounds.
6. The method according to claim 5, characterized in that, The step of combining the confidence scores from the multiple rounds of inference with the source reliability weights of the data sources corresponding to the candidate marketing information to calculate the overall credibility score of the candidate marketing information includes: Obtain a set of marketing context constraint rules that match the business domain of the candidate marketing information. The set of marketing context constraint rules includes official promotional periods, regional whitelists, and applicable audience tag configurations. The candidate marketing information is extracted from the structured representation information parsed in the multi-agent inference module, and multi-dimensional business conflict detection is performed to verify whether the promotion conditions and applicable scope in the structured representation information satisfy the marketing context constraint rule set; if they match, a context with a value of 1 is generated via a flag bit. If a rule conflict exists, an invalid marketing warning will be triggered, and a context flag with a value of 0 will be generated. ; Based on the extraction of the first The event data of each candidate marketing message determines the set of source data sources that provide that candidate marketing message. And query the current source reliability weight of each data source in the source data source set; Based on the aforementioned context, through the flag bit Confidence of multi-round inference The overall credibility score is calculated using the following credibility fusion scoring formula, which includes the current source reliability weights of each data source in the source data source set. : , In the formula, This is the sequence number of the candidate marketing information currently being processed. The index number of a single data source in the set of source data sources; For the first The overall credibility score of the candidate marketing information For the first The context of each candidate marketing message is conveyed through flags. For data source Current source reliability weights, Represents the set of source data. The number of sources in; To adjust the balance coefficient between the source reliability and the model inference confidence ratio and .
7. The method according to claim 6, characterized in that, After outputting candidate marketing information that meets the credibility threshold as target marketing data, the method further includes: According to the preset sliding time window Statistical analysis of the sliding time window Internal data source The proportion of candidate marketing information whose overall credibility score reaches the aforementioned credibility threshold is used as the data source. Recent true quality observation ratio : , In the formula, Indicates the sliding time window Internal source: data source The total collection of candidate marketing information, Represents the total set The number of samples; As a credibility threshold, For indicator functions, when the condition The value is 1 if the condition is met, and 0 otherwise. Based on the exponential smoothing algorithm, using the proportion of recent true quality observations Update data source In the next time window Source reliability weight : , In the formula, For data source In the time window The current source reliability weight, The updated source reliability weights; The smoothing coefficient and .
8. The method according to claim 7, characterized in that, After updating the source reliability weights of the data source, the method further includes: If a data source is detected within the preset continuous monitoring period Updated source reliability weights If the risk level is below a preset risk blocking threshold, a risk blocking mechanism for the data source will be generated. The circuit breaker command; Based on the circuit breaker command, the data source is extracted. The data source identifier is then written into the dynamic blocking blacklist. Based on the dynamic blacklist, intercept and refuse to receive data originating from the data source. Information on subsequent basic events.
9. The method according to claim 1, characterized in that, The multi-channel event data includes behavioral event data triggered by client users on websites, mobile applications, emails, and offline terminals; as well as The event types corresponding to the behavioral event data include at least one of the following: page browsing, product purchase, form submission, and video viewing.
10. A digital marketing information collection system based on large-scale model multi-round inference and verification, characterized in that, The system includes: The authorized collection and integration unit is used to acquire multi-channel event data after obtaining the user's authorization and consent, and to map the multi-channel event data to a unified data model to generate an integrated event stream; The marketing semantic filtering unit is used to perform semantic classification on the event text in the integrated event stream to filter out marketing text, and to perform feature matching filtering on the marketing text based on a preset marketing rule base, thereby identifying and extracting at least one candidate marketing information. The multi-agent inference verification unit is used to input the candidate marketing information into a pre-built multi-agent inference module based on a large language model for each candidate marketing information. The multi-agent inference module performs external evidence retrieval for the candidate marketing information and performs multiple rounds of inference and truth / falseness determination based on the retrieved external evidence to output the multi-round inference confidence for the candidate marketing information. The credibility fusion evaluation unit is used to combine the confidence of the multi-round inference and the source reliability weight of the data source corresponding to the candidate marketing information to calculate the comprehensive credibility score of the candidate marketing information; The threshold decision output unit is used to compare and filter the comprehensive credibility score based on a preset credibility threshold, and output the candidate marketing information that reaches the credibility threshold as target marketing data.