An ai marketing content iteration optimization method and system based on content conversion data

CN122840066APending Publication Date: 2026-09-29KIDSWANT DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611355563.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-09-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,转化链路的构建依赖于预设的归因维度与目标参数,链路结构受限于归因模型的先验设定,缺乏对用户访问时序特征的独立考量,且该方案仅进行转化贡献的量化分析,并未将分析结果直接映射至内容层面的优化操作,即分析与优化之间存在断层,分析结论未能自动驱动内容呈现形式的动态调整,且该方案同样未涉及页面内部的段落级内容优化,优化粒度停留在链路或触点层级,无法定位到具体内容段落的质量问题

Benefits of technology

[0018]本发明的有益效果在于,与现有技术相比,本发明的技术效果如下:本发明通过访问间隔序列的统计学特征量化用户浏览节奏的稳定性,克服了传统路径分析方法仅依赖访问频次或贡献度指标的局限性,使得目标路径的筛选结果能够更真实地反映用户连贯浏览意图与商业转化效率的协同作用;通过末段问句提取与语义匹配机制的配合,实现了跨页面场景下上下文语义的自动承接,有效缩短了用户在内容递进过程中的认知断层与搜索定位时间。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840066A_ABST
    Figure CN122840066A_ABST
Patent Text Reader

Abstract

The application provides an AI marketing content iteration optimization method and system based on content conversion data, and belongs to the technical field of artificial intelligence. The stability of user browsing rhythm is quantified by statistical characteristics of access interval sequences, the limitation of traditional path analysis methods that only rely on access frequency or contribution degree indexes is overcome, the screening result of the target path can more truly reflect the synergistic effect of user coherent browsing intention and commercial conversion efficiency, the automatic connection of context semantics in a cross-page scene is realized through cooperation of the last segment question extraction and the semantic matching mechanism, and the cognitive gap and search positioning time of the user in the content progression process are effectively shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to an AI-based marketing content iteration and optimization method and system based on content conversion data. Background Technology

[0002] With the rapid development of big data and artificial intelligence technologies, content optimization methods based on user access path analysis have been widely researched and applied. Existing technologies typically extract user page access sequences from server logs or advertising monitoring logs, using path mining algorithms to identify frequent access patterns or conversion contribution nodes, thereby providing a basis for decision-making regarding website structure optimization, advertising budget allocation, or content layout adjustments. However, existing solutions still have several technical limitations in terms of the timeliness of path evaluation, the degree of automation in cross-page content semantic connection, and the level of fine-grained optimization.

[0003] For example, patent application CN103823883A discloses a method and system for analyzing website user access paths. It extracts user click data from business systems and text data sources, removes noise and anomalies, constructs a user access path tree based on matching relationships, and introduces node sequence characteristics on top of the Apriori algorithm to mine frequent path graphs. However, this technical solution's path evaluation relies solely on support thresholds, i.e., it only uses the frequency of path occurrence as a metric, failing to incorporate the temporal rhythm characteristics of user access to characterize the coherence and stability of browsing behavior. This leads to high-frequency but discrete paths potentially being misjudged as high-value paths. Furthermore, this solution remains at the page-level path analysis level, without addressing content quality assessment and targeted optimization at the paragraph level within a page. Additionally, this solution does not provide a semantic connection mechanism between pages; users still need to rely on manual scrolling or search to navigate between different content pages, lacking automated contextual connection methods.

[0004] Patent application CN112529634B discloses a conversion link analysis method, system, and computer equipment based on big data. From an advertising and marketing perspective, it uses advertising monitoring log data collected from advertising platforms to convert the log data into conversion links according to preset attribution dimensions and target parameters. It then uses an attribution model to obtain the weights of each touchpoint in the conversion link, thereby obtaining the contribution frequency and degree of each touchpoint in all conversion links. This technical solution achieves post-hoc effect evaluation of marketing campaign touchpoints through attribution analysis, providing data support for subsequent marketing plans. However, the construction of conversion links relies on preset attribution dimensions and target parameters. The link structure is limited by the prior settings of the attribution model, lacking independent consideration of user access timing characteristics. Furthermore, this solution only performs quantitative analysis of conversion contribution and does not directly map the analysis results to content-level optimization operations. That is, there is a gap between analysis and optimization; the analysis conclusions fail to automatically drive dynamic adjustments in content presentation. Also, this solution does not involve paragraph-level content optimization within the page; the optimization granularity remains at the link or touchpoint level, failing to pinpoint quality issues in specific content paragraphs. Summary of the Invention

[0005] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this section, the abstract and title of the invention. Such simplifications or omissions shall not be used to limit the scope of the present invention.

[0006] In view of the aforementioned existing problems, the present invention is proposed.

[0007] Therefore, the technical problem solved by this invention is: how to utilize the historical browsing behavior data of converted users to integrate temporal coherence features in the path analysis stage to improve the reliability of path filtering, and to achieve cross-page contextual semantic continuity and paragraph-level retention optimization in the content optimization stage, thereby improving the overall conversion effect of marketing content.

[0008] To address the aforementioned technical problems, the present invention provides the following technical solution: A method for iterative optimization of AI marketing content based on content conversion data includes: On the server side, acquiring the content browsing sequence of converting users within a preset time window, and recording each consecutive access chain as a candidate path; deriving path weights based on the access times of adjacent content in each candidate path, and filtering target paths and key nodes based on the conversion value of each candidate path and the path weights, wherein the conversion value is the achievement amount of a preset conversion event; On the client side, recording the key node currently accessed by the user as the current node, extracting the last paragraph of the adjacent content before the current node in the target path, matching the last paragraph with the paragraph of the key node to locate the corresponding paragraph, and generating a jump to... The prompt for the corresponding paragraph; when the user triggers the prompt, the key node is loaded, the corresponding paragraph is marked as an exemption optimization paragraph, the visible area of ​​the current viewport when the key node is loaded and scrolled to the corresponding paragraph is defined as the first screen visible area, and the static material of the key node in the first screen visible area is replaced with dynamic material associated with the first screen theme of the first screen visible area; the paragraph-level retention data of each paragraph of the key node is obtained, the paragraphs with low retention data are classified as low retention paragraphs, and for the other low retention paragraphs other than the exemption optimization paragraph, the preview is collapsed or the paragraph is folded and hidden respectively depending on whether it is located in the current viewport.

[0009] In a preferred embodiment of the present invention, the step of deriving path weights based on the access times of adjacent content in each candidate path includes: removing invalid nodes with a dwell time shorter than a preset duration, retaining valid nodes to form a valid browsing sequence; generating an access interval sequence based on the access time difference between adjacent nodes in the valid browsing sequence; calculating the standard deviation of the access interval sequence as a dispersion value; determining path smoothness based on the dispersion value, wherein the dispersion value is negatively correlated with the path smoothness; and deriving the path weights based on the path smoothness, wherein the path weights are positively correlated with the path smoothness.

[0010] In a preferred embodiment of the present invention, after generating an access interval sequence based on the access time difference between adjacent nodes in the effective browsing sequence, the method further includes: obtaining the mean of the access interval sequence; when the mean exceeds a preset duration, generating an attenuation coefficient based on the ratio of the mean to the preset duration, and multiplying the path weight by the attenuation coefficient; the attenuation coefficient is negatively correlated with the ratio.

[0011] As a preferred embodiment of the present invention, the screening of target paths and key nodes includes: obtaining the total conversion value and number of launches associated with each candidate path, generating a comprehensive index based on the ratio of the total conversion value to the number of launches and the path weight, wherein the comprehensive index is the weighted sum or product of the ratio and the path weight. The candidate path with the highest comprehensive index is determined as the target path; nodes in the target path with a transfer probability lower than a first preset level or a single conversion value after passing through the corresponding node higher than a second preset level are determined as candidate key nodes; the transfer probability is the ratio of the number of times the current node accesses the next node to the total number of times the current node is accessed; the single conversion value is the ratio of the number of conversion events achieved after passing through the current node to the total number of sessions passing through the current node; the candidate key nodes are sorted according to the product of the reciprocal of the transfer probability and the single conversion value, and the top N candidate key nodes are determined as the key nodes.

[0012] As a preferred embodiment of the present invention, the extraction of the final question of the adjacent content before the current node in the target path includes: obtaining the unclosed question of the last part of the adjacent content before the current node in the target path; the unclosed question is a sentence that contains an interrogative word at the beginning and does not end with a period or exclamation mark, or the unclosed question is used as the final question; if there is no unclosed question, the core declarative sentence in the last natural paragraph of the adjacent content is extracted, and the core declarative sentence is converted into an interrogative sentence and used as the final question.

[0013] As a preferred embodiment of the present invention, the method of matching the final question with the paragraphs of the key node to locate the corresponding paragraph includes: performing semantic matching between the final question and each natural paragraph of the key node to obtain the semantic matching degree of each natural paragraph; if the highest semantic matching degree is lower than a preset matching threshold, then switching to keyword matching between the core nouns in the final question and the first sentence of each paragraph of the key node, and locating the corresponding paragraph based on the matching result.

[0014] As a preferred embodiment of the present invention, the replacement with dynamic materials associated with the first screen theme includes: extracting the core keywords of the first screen text of the key node within the visible area of ​​the current viewport as the first screen theme; retrieving historical conversion material records associated with the key node in the target path, and selecting several dynamic materials with the highest conversion rates from the dynamic material library as candidate materials; using the TF-IDF method to generate the keyword vector of the first screen theme and the text description tag vector of each candidate material, calculating the cosine similarity between the two, and selecting the candidate material with the highest cosine similarity to replace the first screen static material.

[0015] As a preferred embodiment of the present invention, obtaining the paragraph-level retention data of each paragraph of the key node includes: recording the last fully visible natural paragraph in the visible area of ​​the current viewport when the user stops scrolling as the current reading paragraph; if the user does not continue scrolling and does not perform a jump operation after reading the current reading paragraph, and the current reading paragraph is not the last natural paragraph of the key node, then the corresponding paragraph is counted as a paragraph-level jump; the ratio of the number of paragraph-level jumps to the number of displays for each natural paragraph is calculated as the paragraph-level retention data.

[0016] As a preferred embodiment of the present invention, the following steps are performed: the retention rates in the paragraph-level retention data are sorted from low to high, and the natural paragraphs in the top K positions are divided into low retention paragraphs; the paragraphs in the low retention paragraphs located within the visible area of ​​the current viewport, excluding the exempted optimization paragraphs, are collapsed for preview; and the paragraphs in the low retention paragraphs located outside the visible area of ​​the current viewport, excluding the exempted optimization paragraphs, are collapsed for hiding.

[0017] The present invention also provides an AI marketing content iteration optimization system based on content conversion data, comprising: the server comprising: a browsing sequence acquisition module, used to acquire the content browsing sequence of converted users within a preset time window, and record each consecutive access chain as a candidate path; a path optimization module, used to derive path weights based on the access times of adjacent content in each candidate path, and filter target paths and key nodes based on the conversion value of each candidate path and the path weights, wherein the conversion value is the achievement amount of a preset conversion event; The client includes: a semantic connection module, used to record the key node currently visited by the user as the current node, extract the last paragraph of the question in the adjacent content before the current node in the target path, match the last paragraph with the paragraph of the key node to locate the corresponding paragraph, and generate a prompt to jump to the corresponding paragraph; a visual replacement module, used to replace the static material of the first screen of the key node in the visible area of ​​the current viewport with dynamic material associated with the theme of the first screen in response to the user triggering the prompt; and a retention optimization module, used to obtain the paragraph-level retention data of each paragraph of the key node, classify paragraphs with low retention data as low retention paragraphs, collapse the preview of low retention paragraphs located in the visible area of ​​the current viewport but not the corresponding paragraph, and collapse and hide low retention paragraphs located outside the visible area of ​​the current viewport but not the corresponding paragraph.

[0018] The beneficial effects of this invention are as follows: Compared with the prior art, the technical effects of this invention are as follows: This invention quantifies the stability of user browsing rhythm by using the statistical characteristics of access interval sequences, overcoming the limitations of traditional path analysis methods that rely solely on access frequency or contribution indicators, so that the target path screening results can more realistically reflect the synergistic effect of user's coherent browsing intent and commercial conversion efficiency; through the combination of end-segment question extraction and semantic matching mechanisms, automatic contextual semantic connection is achieved in cross-page scenarios, effectively shortening the cognitive gap and search positioning time for users in the content progression process.

[0019] By further refining the collection of paragraph-level retention data, the granularity of content optimization is reduced from the page level to the paragraph level. Combined with differentiated strategies of collapsing previews within the viewport and folding and hiding outside the viewport, the readability of low-retention paragraphs is significantly improved while releasing rendering resources outside the viewport. In addition, exemption tags are used to protect paragraphs that users actively intend to access, avoiding conflicts between automated optimization strategies and users' real-time operational intentions. Ultimately, the conversion and retention effects of marketing content are improved in a synergistic manner from multiple dimensions such as path guidance, reading fluency, content quality, and user intent protection. Attached Figure Description

[0020] Figure 1 This is a flowchart of the candidate path and target path selection process in this invention.

[0021] Figure 2 This is a flowchart of cross-page semantic connection and visual replacement in this invention.

[0022] Figure 3 This is a structural diagram of an AI marketing content iteration and optimization system based on content conversion data, as described in this invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The embodiments described in this application are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this invention.

[0024] like Figures 1-2 As shown, the present invention provides an AI-powered marketing content iteration optimization method based on content conversion data, comprising: S1: The server obtains the content browsing sequence of the converted user within a preset time window and records each consecutive access chain as a candidate path; S2: Path weights are derived based on the access time of adjacent content in each candidate path. Target paths and key nodes are selected based on the conversion value and path weight of each candidate path. The conversion value is the achievement amount of a preset conversion event; S3: The client records the key node currently accessed by the user as the current node, extracts the last paragraph of the adjacent content before the current node in the target path, matches the last paragraph of the question with the paragraph of the key node to locate the corresponding paragraph, and generates a prompt to jump to the corresponding paragraph. S4: When the user triggers a prompt, the key node is loaded, the corresponding paragraph is marked as an exemption optimization paragraph, and the visible area of ​​the current viewport when the key node is loaded and scrolled to the corresponding paragraph is defined as the first screen visible area. The static material of the key node in the first screen visible area is replaced with dynamic material associated with the first screen theme of the first screen visible area. S5: Obtain the paragraph-level retention data of each paragraph of the key node, classify the paragraphs with low retention data as low retention paragraphs, and for the other low retention paragraphs other than the exemption optimization paragraphs, perform collapse preview or collapse and hide respectively depending on whether they are located in the current viewport.

[0025] In some embodiments, the AI ​​marketing content iteration optimization method based on content conversion data provided by the present invention includes the following steps in step S1: Step S11: The server extracts all user access records within a preset time window from the user behavior log system.

[0026] The preset time window is a backward-looking time window with the current time as the endpoint. The preset time window is dynamically set based on the content update frequency and conversion cycle: a 7-day window is used when the average daily update volume of content pages exceeds 5% of the total number of pages; a 14-day window is used when the daily update volume is between 1% and 5%; and a 30-day window is used when the daily update volume is less than 1%. The daily update volume is obtained from the version logs of the content management system. The default is a 14-day window.

[0027] Step S12: Filter converted users from all user access records. Converted users are those who complete a preset conversion event within a preset time window. The preset conversion events include form submission, order payment completion, target page depth reaching a preset threshold, or preset control being triggered.

[0028] Based on the user identifier of the converted user, extract all access records of the converted user within a preset time window, and arrange them in ascending order of access time to form a content browsing sequence.

[0029] Step S13: Each browsing record in the content browsing sequence includes a content identifier, access time, page dwell time, and source page identifier. The content identifier uniquely points to a content page, and the source page identifier indicates the upstream jump source of the current browsing record. When the source page identifier is empty, it indicates that the current browsing record is the session start node.

[0030] Step S14: Perform a segmentation operation on the content browsing sequence according to the preset session segmentation rules: calculate the time interval between adjacent browsing records. When the time interval is greater than the preset session timeout threshold (e.g., 30 minutes), use the interval as the segmentation boundary and classify the browsing records before and after the segmentation boundary into different sessions. The content identification sequence accessed continuously within the same session is recorded as a continuous access chain, i.e., a candidate path.

[0031] In some embodiments, the default value of the session timeout threshold is 30 minutes. As a preferred embodiment of the present invention, this threshold can be dynamically adjusted according to the average reading time of the content type: when the average reading time is greater than 5 minutes, the timeout threshold is set to 60 minutes; when the average reading time is 1 to 5 minutes, the timeout threshold is set to 30 minutes; when the average reading time is less than 1 minute, the timeout threshold is set to 15 minutes. For ease of implementation, those skilled in the art can also directly use the default value of 30 minutes without dynamic adjustment. The preset session timeout threshold in step S28 has the same meaning as the session timeout threshold in step S14, and uses the same value selection logic.

[0032] Step S15: Each candidate path is stored in the path pool as an ordered node sequence, forming a node sequence. Each node contains attributes such as content identifier, content type, access time, and dwell time. The path depth is the total number of nodes contained in the candidate path. Each candidate path is associated with and stored with its corresponding converted user identifier and conversion event achievement record.

[0033] It should be noted that by using preset time windows and session segmentation rules to structure the raw logs, the discontinuous noise caused by users being offline or accessing the site across days is filtered out. This allows the candidate paths to accurately reflect the complete content progression trajectory of the converting user driven by a single intent, providing a data foundation for subsequent path weight calculation and target path mining.

[0034] In some embodiments, the AI ​​marketing content iterative optimization method based on content conversion data provided by the present invention includes the following steps in step S2: Step S21: For each candidate path, traverse the node sequence and remove invalid nodes whose dwell time is lower than a preset threshold (e.g., 3 seconds). The dwell time is calculated based on the difference between the access time of the current node and the access time of the next node. For the last node in the path, the dwell time is preferentially calculated based on the difference between the trigger time of the page unload event (beforeunload or pagehide) and the access time. If the page unload event cannot be obtained, such as when the user directly closes the browser process or the network is disconnected, try using the hidden timestamp of the page visibility API (visibilitychange) as a substitute. If neither of the above two methods can be obtained, the dwell time of the last node is set to the average dwell time of all non-last nodes in the path. If there are no non-last nodes in the path, it is set to the global default value of 10 seconds. The above three methods are tried in order of priority, and the first successfully obtained timestamp is used.

[0035] Nodes whose dwell time is greater than or equal to a preset duration threshold constitute a valid browsing sequence.

[0036] It should be noted that nodes whose dwell time is less than the preset time threshold indicate that the user has not completed substantive reading or information extraction on the content page. Their corresponding access behavior is accidental touch or quick skipping, and including them in the path analysis would introduce noise, so they are removed.

[0037] Step S22: Generate an access interval sequence based on the access time difference between adjacent nodes in the valid browsing sequence; let the valid browsing sequence be... ,in Indicates the first Valid nodes If the access time is specified, then the access interval sequence is represented as: ; in, , Decrease the total number of nodes in the valid browsing sequence by 1.

[0038] Step S23: Calculate the dispersion value of the access interval sequence. The dispersion value is characterized by calculating the standard deviation of the access interval sequence. The formula for calculating the standard deviation is: ; ; in, This is the mean of the access interval sequence.

[0039] Furthermore, the degree of dispersion is characterized by calculating the coefficient of variation of the access interval sequence. The calculation formula is: ; in, This is the mean of the access interval sequence.

[0040] Step S24: Determine the path smoothness based on the dispersion value. The dispersion value is negatively correlated with path smoothness; that is, the smaller the dispersion value, the more uniform the user's switching rhythm between content nodes, and the higher the path smoothness. The formula for calculating path smoothness is: ; in, For path smoothness, A preset very small positive number (e.g., 0.01) is used to ensure that the denominator is not zero.

[0041] Step S25: Derive path weights based on path smoothness. : ; in, The preset baseline weight value is 1.0. This is a preset scaling factor used to adjust the mapping magnitude of path smoothness to weights. The value can be determined by those skilled in the art through conventional experimental methods based on the actual application scenario. For example, it can be traversed within the range of [0.1, 1.0] with a step size of 0.1 to select the value that makes the comprehensive index most correlated with the actual conversion rate of the path. The specific method of value selection is not an improvement of this invention. Those skilled in the art can flexibly choose fixed empirical values ​​(such as...) according to business needs. =0.5) or a periodic dynamic optimization value, not as a limitation of the present invention.

[0042] It should be noted that by calculating the dispersion of access intervals, the stability of user browsing rhythm is quantified, so that path weights can reflect the behavioral confidence of the path. Paths with more uniform access intervals receive higher weights and have a greater advantage in subsequent path selection, thus introducing the time-domain feature of user behavior into the path evaluation system.

[0043] In some embodiments, after step S22 and before step S23, a duration attenuation correction step is further included: Step S221: Obtain the mean of the access interval sequence. When it exceeds a preset duration threshold... (For example, at 120 seconds), an attenuation coefficient is generated based on the ratio of the mean to a preset duration threshold. : ; in, A preset adjustment factor (e.g., 0.5) is used to control the decay rate; when hour, .

[0044] Step S222, the path weights derived in step S25 are... Multiplying by the attenuation factor yields the corrected path weight after time penalty: ; in, To correct path weights. This will be abbreviated as […] in all subsequent steps. , to represent the final weight value that has been factored into the duration penalty.

[0045] It can be seen that when the average access interval in the path is too large, even if the dispersion of the access interval is low (the path smoothness is high), the path is difficult to reflect the user's true continuous browsing intention. By applying a negative penalty to the path weight based on the decay coefficient of the average interval, pseudo-high confidence paths caused by factors such as cross sessions and users returning after going offline can be effectively filtered out.

[0046] Step S26: Obtain the total conversion value associated with each candidate path. and number of starts The total conversion value is the sum of the achieved amounts of preset conversion events triggered by all nodes in the candidate path. ; in, For path depth, Indicates the first The number of conversion events achieved by each node. The number of launches is the number of times the candidate path is executed by different users or different sessions.

[0047] Average conversion value per launch The calculation is as follows: ; in, A preset minimum positive value (e.g., 0.001) is used to prevent the denominator from being zero.

[0048] Step S27: Generate a comprehensive index based on the average conversion value per launch and path weight, and determine the candidate path with the highest comprehensive index as the target path; comprehensive index The calculation formula is: ; in, A preset weighting coefficient is used to balance the contribution of path behavior confidence and business conversion efficiency to the overall index. The specific value can be determined by those skilled in the art through conventional experimental methods based on business objectives (emphasizing behavior stability or conversion efficiency). For example, when emphasizing behavior stability, γ is greater than 0.5, and when emphasizing conversion efficiency, γ is less than 0.5. The method of determining the value of γ can be achieved through existing technologies such as grid search, manual experience setting, or business rule configuration, and is not intended to limit the invention. For ease of implementation, those skilled in the art can directly choose 0.5 as the balancing configuration.

[0049] In other embodiments, the formula for calculating the comprehensive index is in a multiplicative form: ; It should be noted that the comprehensive index considers both the behavioral confidence of the path and the commercial conversion efficiency. Candidate paths that excel in one dimension but are deficient in the other are unlikely to obtain the highest comprehensive score, thus ensuring that the target path has both behavioral stability and commercial effectiveness.

[0050] Step S28: After the target path is determined, a criticality assessment is performed on each node in the target path, and nodes that meet one of the following conditions are identified as candidate critical nodes: Condition 1: Transition Probability If the value is below the first preset level, the transition probability is defined as the conditional probability of transitioning from the current node to the next node in the target path, and the calculation formula is as follows: ; in, Indicates from node Transfer to node Number of times, Represents a node Total number of visits.

[0051] The first preset level is set to 0.1 by default, and its value range is [0.05, 0.2]. When the quantile is dynamically set, the distribution is calculated based on the transition probability values ​​of all non-last nodes in all candidate paths to determine the lower quartile.

[0052] It should be noted that condition one only applies to non-last nodes in the target path (j < d). For the last node in the target path, since there is no defined behavior to move to the next node, condition one is not applicable to avoid misjudging the user's natural exit behavior as a churn bottleneck that needs optimization.

[0053] In this embodiment of the invention, exiting the target path is defined as the user performing any of the following actions from the last node page: (i) closing the browser tab or browser window, (ii) leaving the page after staying on the page for more than a preset session timeout threshold (e.g., 30 minutes) without any interactive behavior, and (iii) jumping to a third-party domain or external link that does not match any node in the target path; the number of exits is obtained from the page uninstallation event, session timeout event and external jump event in the user behavior log.

[0054] Condition 2: The value of a single conversion after the current node: ; in, Indicates via node The amount of conversion achieved after the event is triggered. Indicates via node Total number of sessions.

[0055] For the last node, the criticality assessment is based solely on condition two. If the single conversion value after passing through the last node is higher than the second preset level (the average single conversion value of the last nodes in all candidate paths is taken by default), then the last node is identified as a candidate critical node. The higher the single conversion value of the last node, the stronger its effectiveness as the conversion closing stage, and it can be retained and strengthened as a positive critical node.

[0056] Step S29, sort candidate key nodes according to product value Sort the nodes from highest to lowest, and determine the top N candidate key nodes as key nodes; the product value is calculated using the following formula: ; in, A preset minimum value, for example, 0.001, is used to prevent the product from diverging due to a zero transition probability.

[0057] It should be noted that this product value scoring formula only applies to non-last nodes in the target path. For last nodes that are identified as candidate key nodes only through condition two, their scores are directly based on the calculated value as the ranking basis, without multiplying by the reciprocal of the transition probability. Last nodes and non-last nodes whose product values ​​have been calculated using the scoring formula are ranked from highest to lowest according to their respective scores and participate in the selection of the top N.

[0058] Where N is a preset positive integer, ranging from 1 to 5, with a default value of N=3; or it can be dynamically determined based on the total number of nodes in the target path: ,in This represents the total number of nodes in the target path.

[0059] It should be noted that when the node migration probability is extremely low, its reciprocal is extremely large, indicating that the node is a bottleneck in the conversion funnel; when the single conversion value of a node is extremely high, it indicates that the node has a significant catalytic effect on downstream conversions. The product of these two factors is then sorted to achieve a balance between points of churn risk and conversion trigger points.

[0060] In some embodiments, the AI ​​marketing content iteration optimization method based on content conversion data provided by the present invention includes the following steps in step S3: Step S31: The client obtains the key node identifier issued by the path optimization module. When the content identifier of the content page currently accessed by the user matches the key node identifier, the currently accessed node is recorded as the current node, and the content data of the adjacent content in the target path that is located before the current node is loaded. The adjacent content is the content page corresponding to the node directly adjacent to the current node in the ordered node sequence in the target path.

[0061] In some embodiments, when the current node is the first node of the target path, i.e. there is no adjacent content before the current node, the execution of step S3 is terminated and no jump prompt is generated.

[0062] Step S32: Perform natural paragraph boundary identification on the content data of adjacent content and extract the content text of the last part of adjacent content; traverse the last part to see if there are any unclosed questions. Unclosed questions refer to sentences that are introduced by interrogative words, end with a question mark, or have a question-like syntactic structure and seek semantic completion even without a question mark; if there are unclosed questions, then take the unclosed question as the last paragraph question.

[0063] The rules for determining unclosed questions include: the sentence begins with any one of the following interrogative words: why, how, how, whether, whether, what, whether, or not, and does not end with a period or exclamation mark; if the sentence ends with a question mark, the existence of a pronoun (such as "this scheme," "the above method," or "this strategy") pointing to the content of the current node in the syntactic dependency tree of the question is used as the basis for determining that the answer to the question points to the content of the current node; if a pronoun exists and the pronoun is the object or subject of the sentence in the dependency syntax relation, the answer to the question is determined to point to the content in the current node.

[0064] Specifically, the interrogative syntactic structure is determined objectively through the following methods: Dependency parsing is performed on the sentence. If the root node of the sentence is an interrogative pronoun (such as who, what, which, how, etc.) or the core predicate of the sentence (is, has, can, can, should, etc.) and the sentence does not begin with a clear declarative subject, then it is determined to be an interrogative syntactic structure. Semantic completion is determined through the following methods: Named entity recognition and referential resolution are performed on the sentence. If the sentence contains at least one referential pronoun (such as this, that, the above, its, etc.) whose antecedent cannot be found within the current sentence, or if the argument position corresponding to the core predicate of the sentence is missing (i.e., lacking a subject or object), then it is determined to be semantically complete. A sentence is considered an unclosed question only if both the interrogative syntactic structure and semantic completion conditions are simultaneously met.

[0065] Step S33: If there are no unclosed questions at the end, extract the core statement sentence from the last paragraph of the adjacent content.

[0066] The core declarative sentence is determined as follows: the last paragraph is segmented into sentences, the sum of the keyword weights of each sentence is calculated, and the sentence with the highest sum of keyword weights is selected as the core declarative sentence. Specifically, after segmenting the sentence, stop words are filtered (the stop word list uses the Harbin Institute of Technology stop word list), retaining words with a length of ≥2; the keyword weight of each retained word is obtained by calculating the TF-IDF value of the word in the industry category of the current content page, and the industry category thesaurus is loaded from the industry corpus pre-stored on the server; the sum of the TF-IDF values ​​of each retained word in the sentence is obtained.

[0067] Transform the core declarative sentence into an interrogative sentence as follows: (a) When the core declarative sentence is a subject-verb-complement structure (such as "X is Y") or a subject + transitive verb + object structure, and does not contain a modal verb, add "whether" at the beginning of the sentence, and keep the predicate and object of the original sentence unchanged, and replace the punctuation at the end of the sentence with a question mark; (b) When the core statement contains modal verbs such as can, can, should, etc., move the modal verbs to the beginning of the sentence and replace the punctuation at the end of the sentence with a question mark. For example, change "this strategy can improve the conversion rate" to "can this strategy improve the conversion rate?" (c) When the core declarative sentence contains a manner or method object (such as through..., using...), add "how" at the beginning of the sentence, and keep the subject and predicate of the original sentence unchanged, and replace the punctuation at the end of the sentence with a question mark; (d) When the core statement is a negative sentence, retain the negative word, add "whether" at the beginning of the sentence and restore the predicate to the affirmative form. For example, if the strategy is not optimal, change it to "whether the strategy is optimal".

[0068] If none of the above rules are matched, then by default, "whether" is added at the beginning of the sentence and the punctuation at the end of the sentence is replaced with a question mark; the resulting interrogative sentence is then used as the final question.

[0069] For example, the core statement is that the strategy can effectively improve the conversion rate, and the final question after conversion is whether the strategy can effectively improve the conversion rate.

[0070] It should be noted that by prioritizing the extraction of unclosed questions as semantic anchors, the writing pattern of content creators guiding users to continue reading by posing questions at the end of paragraphs is fully utilized; when there are no unclosed questions, the automatic construction of cross-content semantic anchors is achieved by actively converting the core declarative sentence into an interrogative sentence, so that the S3 step can adapt to content pages with different writing styles.

[0071] Step S34: Perform semantic matching between the final question and each paragraph of the key node to obtain the semantic matching degree of each paragraph.

[0072] Semantic matching is represented by inputting the final question and the paragraph text into a semantic matching model, and obtaining the semantic vector similarity between the two. Specifically, sample pairs consisting of the final question and the paragraphs of the next page between adjacent content pages in the conversion path are extracted from historical user behavior logs stored on the server. Each sample pair is labeled by a human annotator to indicate whether it matches (match / non-match), constructing a training dataset with a positive-to-negative sample ratio of 1:3. Pre-trained weights from BERT-base-chinese are used as initial parameters, and fine-tuning is performed on the training dataset. The loss function is cross-entropy loss, the optimizer is AdamW, and the learning rate is set to 2×10⁻. 5 The batch size was set to 32, and the training rounds were 3. The model with the highest F1 score on the validation set was selected as the final semantic matching model. The final paragraph question and the natural paragraph text were input into the fine-tuned semantic matching model, and the 768-dimensional vector output by the [CLS] bits was taken as the sentence-level semantic representation. The cosine similarity between the two was calculated as the semantic matching degree. The similarity value ranged from 0 to 1, with values ​​closer to 1 indicating greater semantic relevance.

[0073] Furthermore, when there are insufficient labeled samples, a fixed threshold of 0.65 is used. This value is derived from the weighted average of the optimal threshold calculated on multiple validation sets of different content types (the weight is proportional to the number of pages of each content type). This fixed value ensures that the F1 loss on most content types does not exceed 5% of the dynamic threshold scheme.

[0074] Step S35: Extract the highest semantic matching score from each paragraph and determine whether the highest semantic matching score is lower than the preset matching threshold (default 0.65). If the highest semantic matching score is lower than the preset matching threshold, switch the matching strategy: extract the core nouns in the question in the last paragraph. The core nouns are determined by dependency parsing, including noun entities in the subject and object of the sentence. Perform keyword matching between the core nouns and the first sentence of each paragraph of the key node. Keyword matching is to calculate the number of core nouns contained in the first sentence of the paragraph. The paragraph with the most core nouns is determined as the corresponding paragraph.

[0075] The preset matching threshold is determined by performing ROC curve analysis on the matching accuracy of the semantic matching model on the validation set and selecting the threshold point corresponding to the maximum value of the Youden index; 0.65 is used by default when there is no validation set.

[0076] If the highest semantic match degree is greater than or equal to the preset matching threshold, then the natural paragraph corresponding to the highest semantic match degree will be determined as the corresponding paragraph.

[0077] Step S36: After determining the corresponding paragraph, the client generates a jump prompt at the position of the corresponding paragraph during the current page rendering process of the key node. The jump prompt is an interactive element that is visually distinct from ordinary text. The style of the interactive element includes, but is not limited to, any one or more combinations of highlighted underline, floating button, guide arrow or collapse / expand control. The jump prompt is configured to respond to the user's trigger operation, execute page scrolling or anchor jump, so that the corresponding paragraph enters the visible area of ​​the current viewport.

[0078] In some embodiments, the jump prompt is displayed at the beginning of the corresponding paragraph or at a fixed position on the side of the area where the corresponding paragraph is located.

[0079] In some embodiments, the jump prompt is accompanied by a fade-in / fade-out animation to reduce interference with the user's normal reading.

[0080] Step S37: The client monitors the trigger status of the jump prompt: If the user triggers the jump prompt within the preset response time window (e.g., 30 seconds), then step S4 is executed; if it is not triggered, the jump prompt is canceled, step S4 is not executed, and the process proceeds directly to step S5 to perform paragraph-level retention statistics.

[0081] It should be noted that by semantic matching between the final question and the paragraph, as well as backtracking keyword matching, the unfinished reading intentions of users at the end of the previous content page are accurately anchored to the most relevant paragraph in the current content page. The user is then guided to that paragraph through a jump prompt, thus achieving cross-page contextual continuity. This can effectively shorten the cognitive gaps and search costs for users during the content progression process.

[0082] In some embodiments, the AI ​​marketing content iterative optimization method based on content conversion data provided by the present invention includes the following steps in step S4: In step S41, in response to the user's triggering operation of the redirect prompt, the client first checks whether the content data of the key node has been loaded in the current client and whether there is a valid rendering node in the page document object model: if it has been loaded and the rendering node is valid, the step of sending a data request to the server is skipped, and subsequent operations are performed directly with the currently loaded content data; if it has not been loaded or the rendering node is invalid, a content loading request for the key node is sent to the server.

[0083] The loading request includes the content identifiers of key nodes and the paragraph anchor identifiers of the corresponding paragraphs; the server returns the complete content data and rendering resources of the key nodes based on the content identifiers, and the client receives and parses the content data and performs content rendering in the page viewport.

[0084] In step S42, after the content is loaded, the client locates the corresponding paragraph element in the document object model of the key node according to the paragraph anchor point identifier, and performs page scrolling or anchor point jump operation to place the corresponding paragraph in the visible area of ​​the current viewport; at the same time, an exemption flag is set for the corresponding paragraph element in memory. The exemption flag is stored as a boolean attribute in the extended attribute of the paragraph element. When the attribute value is true, it means that the paragraph is marked as an exemption optimization paragraph. The exemption flag is read in the subsequent retention optimization steps to prevent the paragraph from being incorrectly executed with collapse preview or collapse hide operation.

[0085] The exemption flag is used to skip paragraphs with the exemption flag during step S5 when performing low retention paragraph optimization. This means the paragraph will not be collapsed or hidden due to low retention data. The exemption flag is cleared after the user leaves the current page. When the user visits the page again, steps S3-S4 must be triggered again to regain the exemption flag.

[0086] It should be noted that, by setting exemption markers at the paragraph element level, the subsequent paragraph-level retention optimization steps can identify and jump to the paragraph that the user actively intends to access, thus avoiding conflicts between automated optimization strategies and the user's real-time operational intentions.

[0087] Step S43: Obtain the first-screen text content of the key node within the first-screen visible area of ​​the current viewport. The first-screen visible area is defined as the page area covered by the user's current viewport after the page has finished loading and scrolled to the corresponding paragraph; extract the core keywords from the first-screen text content as the first-screen theme.

[0088] In some embodiments, core keywords are extracted as follows: All text within the visible area of ​​the first screen is segmented and filtered for parts of speech, retaining nouns and gerunds as candidate keywords; the TF-IDF weight of each candidate keyword is calculated using the following formula: ; in, Keywords Frequency of appearance in the first screen text For keywords Number of content pages The total number of content pages; select the candidate keywords ranked Lth in the TF-IDF weight ranking as core keywords, with L being 3. The value of L is determined by statistically analyzing the average number of natural paragraphs on all content pages: L = min(3, total number of sentences in the visible area of ​​the first screen). When the total number of sentences is less than 3, the actual number of sentences shall prevail.

[0089] The core keywords constitute the keyword vector of the theme on the first screen.

[0090] In other embodiments, core keywords are extracted using a pre-trained text semantic encoding model. The first screen text is input into the encoding model, and the semantic vector of the output text is used as the semantic representation of the first screen topic.

[0091] Step S44: Retrieve historical conversion material records associated with key nodes in the target path. These historical conversion material records are stored in a dynamic material library on the server. The dynamic material library is a relational database table or key-value store with the material identifier as the primary key and indexed by the content identifier. Each record contains the material identifier, material type, material file storage path or CDN link, text description tags of the material, and conversion contribution index values ​​for each dimension.

[0092] The dynamic material library can be constructed using conventional data warehouse construction methods in this field. Specifically, when new or updated materials are added to a content page, the server receives material change notifications through the event listener interface of the content management system, triggering the material entry process: For new materials, text description tags are extracted and conversion contribution indicators are initialized (initial values ​​are set to the industry average or the historical average of similar materials); for existing materials, their conversion contribution indicators are updated in batches according to a preset period (e.g., daily). The text description tags for materials are extracted as follows: For image and text materials, the text content of the content area where the material is located is directly extracted, and after word segmentation and stop word filtering, the 5-10 keywords with the highest TF-IDF weights are selected as text description tags; for video materials, video keyframes are first extracted, and OCR text recognition is performed on the keyframes to obtain the on-screen text. At the same time, the audio track of the video is extracted and converted into text through speech recognition. The on-screen text and speech recognition text are merged, and after word segmentation and stop word filtering, the 5-10 keywords with the highest TF-IDF weights are selected as text description tags; for carousel materials, the OCR recognition results of each frame are merged and keywords are extracted in the same way.

[0093] When materials are added to the database, operations personnel can manually review and correct the automatically extracted text description tags to ensure that the tags are consistent with the theme of the material content. Manual review is an optional optimization step, and its presence or absence does not affect the complete implementation of the technical solution of this invention. The text description tag vector is composed of the TF-IDF weight vectors of the aforementioned keywords.

[0094] The conversion contribution metrics include one or more of the following: click-through rate (CTR), conversion rate triggered by creatives (CVR), and retention improvement rate. When multiple metrics are used, the composite score for ranking is the weighted sum of the normalized metrics. The weight coefficients are pre-stored in the server configuration. The normalization of each metric is performed using Min-Max normalization to the [0,1] interval. The weight coefficients of the weighted sum are preset as follows: CTR weight 0.3, CVR weight 0.5, and retention improvement rate weight 0.2. Each weight coefficient is stored in the server configuration file and can be manually adjusted according to the business scenario.

[0095] Furthermore, the historical dynamic materials are sorted in descending order based on the composite score, and the top M dynamic materials are selected as candidate materials, where M is a preset positive integer, and in this embodiment, M=5 is preferred.

[0096] It should be noted that historical conversion material records are stored in a dynamic material library on the server side. Each record contains material identifier, material type, text description tags of the material, and conversion contribution index values ​​for each dimension.

[0097] Step S45: Calculate the TF-IDF cosine similarity between the theme of the first screen and the text description tags of each candidate material.

[0098] It should be noted that the text description tag vectors of the candidate materials use the same encoding method as the keyword vectors of the main theme on the first screen, that is, generated based on TF-IDF weighting, ensuring that the two are in the same vector space. The TF-IDF cosine similarity is obtained by calculating the cosine similarity between two TF-IDF vectors. The formula for calculating cosine similarity is: ; in, The TF-IDF vector of the theme on the first screen. For the first The TF-IDF vectors of the text description tags of each candidate element. The value range is [0,1], and the closer it is to 1, the higher the similarity.

[0099] Step S46: Select the candidate material with the highest TF-IDF cosine similarity as the target dynamic material, and replace the static material of the first screen with the target dynamic material in the first screen's visible area of ​​the key node; the static material of the first screen is any one of the static images, static text and image combinations, or static backgrounds displayed in the current first screen's visible area; the target dynamic material is any one of the video clips, motion graphics, carousels, or interactive elements with animation effects.

[0100] In some embodiments, the replacement operation involves removing or hiding the Document Object Model (DOM) node of the static material from the page document tree and inserting the DOM node of the dynamic material into the same position, while simultaneously triggering the automatic playback of the dynamic material.

[0101] In step S461, after the client completes the dynamic material replacement, it records the material replacement status and replacement timestamp of the current page in memory. If the user does not trigger the jump prompt in step S4, that is, no dynamic material replacement has occurred, the statistical starting point of the paragraph-level retention data is based on the access time when the user first enters the current node in step S3 (i.e., the time when the page is fully loaded).

[0102] Once the key nodes have finished loading, the client starts monitoring user scrolling behavior and paragraph visibility status within the page. If the user triggers the dynamic material replacement in step S4, the client records the replacement timestamp in memory. The statistics of paragraph-level retention data are based on this replacement timestamp as the dividing point. Display behavior data generated before the replacement is completed is deducted from the display count, and only user behavior data generated after the replacement is completed is counted. If the user does not trigger the dynamic material replacement in step S4, the statistics of paragraph-level retention data start from the time the page is fully loaded, and all user behavior data is accumulated normally.

[0103] For paragraphs that have been displayed before the replacement, the following method is used: Before the dynamic material replacement begins, the client saves the current value of the display count counter of each visible paragraph in the current viewport as a snapshot; after the replacement is completed, it iterates through the visible paragraphs in the current viewport, restores their display count counters to the snapshot value, and restarts the timing to ensure that the brief display before the replacement is not counted in the retention statistics.

[0104] As can be seen, by extracting the theme of the first screen after the user triggers the jump and combining it with the dynamic material with the best historical conversion performance for semantic matching and replacement, this invention dynamically associates the theme content that users care about with the visual presentation form with high conversion, thereby improving the immediate attractiveness of the first screen content to users and the conversion driving effect. At the same time, the TF-IDF cosine similarity matching mechanism ensures that the replaced dynamic material is consistent with the user's current reading focus in terms of content theme, avoiding the sense of theme deviation caused by visual replacement.

[0105] In some embodiments, the AI ​​marketing content iteration optimization method based on content conversion data provided by the present invention includes the following steps in step S5: Step S51: After the key nodes are loaded, the client starts listening to the user's scrolling behavior and paragraph visibility status within the key node page, without waiting for the dynamic material replacement to complete.

[0106] The client monitors user scrolling behavior and paragraph visibility status within key node pages; when user scrolling stops, it obtains the visible area boundary coordinates of the current viewport, traverses the document object model elements of each paragraph within the key node page, and obtains the boundary coordinates of each paragraph; the last fully visible paragraph within the visible area boundary is recorded as the current reading paragraph.

[0107] Among them, "finally fully visible" means that the bottom boundary of the paragraph does not exceed the bottom boundary of the current viewport's visible area, and the top boundary of the paragraph is within the visible area; when there is no fully visible paragraph within the visible area, the paragraph with the highest exposure ratio in the visible area is taken as the current reading paragraph.

[0108] It should be noted that after the anchor point jump operation in step S42 is completed and before the user's first scrolling operation begins, the initial display state of the key node, that is, the state in which the corresponding paragraph is located within the viewport, does not trigger the timing window and paragraph-level jump judgment in step S52; the paragraph-level jump statistics only apply to the current reading paragraph when the user stops scrolling actively, so as to exclude false jump data caused by initial jump.

[0109] When executing the anchor jump in step S42, the client sets a boolean flag to true; after scrolling is complete, the flag is set to false. In step S51, when listening to the user's scrolling behavior, the client checks the flag to determine whether the current scrolling was triggered by the program: if the boolean flag is true, all calculations triggered by this scrolling are ignored, the current reading paragraph is not recorded, and the timer window is not started; only when it is false is the recording and subsequent processing of the current reading paragraph executed.

[0110] Step S52: After the current reading paragraph is determined, start the timer window with the preset duration.

[0111] The timing window is activated under the following conditions: no new scroll trigger signal is detected within 500ms of the client scroll listener event (i.e., the user stops scrolling), the page is redrawn, and the boundary coordinates of the currently read paragraph remain stable for two consecutive animation frames (approximately 33ms). After the timing window is activated, if a scroll event or jump event is detected before the window times out, the current timing is immediately terminated and the timer is reset. The timing window is restarted after the user stops scrolling again.

[0112] Furthermore, the duration of the timing window is set based on the average paragraph reading time of the content page. The average paragraph reading time = the global average dwell time corresponding to this content page type / the total number of natural paragraphs on the page.

[0113] The default value of the timing window can be determined by those skilled in the art through conventional statistical methods based on the content type (such as long text / image, short text / image, video page, etc.). As a preferred embodiment of the present invention, 5 seconds can be used as a general default value. However, this default value is not a limitation on the scope of protection of the present invention. Those skilled in the art can adaptively adjust the timing window duration within the range of 3 to 10 seconds based on the average paragraph reading time distribution of the actual content page, or directly use the calculated value as the timing window duration when the calculated value exceeds this range (for example, when the average paragraph reading time is 12 seconds, the timing window can be directly set to 12 seconds instead of truncating it to 10 seconds). The above-mentioned truncation method is only one optional implementation scheme; those skilled in the art can also choose not to perform truncation processing, all of which are equivalent embodiments of the present invention.

[0114] Continuously monitor the following three states within the timing window: State 1: Does the user want to continue scrolling? Status 2: Whether the user has triggered a navigation action within the page. Navigation actions include clicking an anchor link, clicking a navigation prompt, or switching page tabs. Status 3: Is the currently reading paragraph the last paragraph of a key node page?

[0115] If the user does not perform either state one or state two operation after the timer window expires, and the currently read paragraph does not belong to state three, then the current read paragraph will be counted as a paragraph-level exit, and the number of paragraph-level exits for that paragraph will be accumulated.

[0116] In some embodiments, when the currently read paragraph is the last paragraph of a key node page, it is not counted as a paragraph-level jump, in order to prevent the natural completion of the content from being mistakenly judged as a jump.

[0117] Step S53: Count the number of times each paragraph is displayed. The number of times is defined as the cumulative number of times the paragraph enters the visible area of ​​the current viewport and is viewed by the user for at least the preset minimum duration.

[0118] The preset minimum duration is set according to the page type. For example, it is set to 1 second for text and image content pages and 0.5 seconds for video content pages. The default value is 1 second. This duration is used to filter out invalid displays caused by users scrolling quickly through paragraphs.

[0119] Furthermore, the paragraph-level retention data for each natural paragraph is calculated. The paragraph-level retention data is represented by the retention rate, and the formula for calculating the retention rate is as follows: ; in, No. Retention rate of each paragraph For the first The number of paragraph-level jumps per natural paragraph. For the first The number of times each paragraph is displayed. This is a preset minimum value used to prevent the denominator from being zero; The value ranges from [0,1], and a higher value indicates better retention performance.

[0120] In other embodiments, paragraph-level retention data is directly represented by paragraph-level bounce rate, calculated using the following formula: ; in, For the first The bounce rate per paragraph; a higher value indicates poorer retention performance.

[0121] It should be noted that, by defining the exit behavior at the end of a paragraph without further interaction at the paragraph level and excluding the case of reading the last paragraph completely, this embodiment of the invention achieves a fine-grained quantification of the retention capability of each paragraph on the content page, providing fine-grained data support for subsequent targeted optimization. Compared with page-level retention metrics, paragraph-level data can accurately pinpoint the specific content location where users drop off.

[0122] Step S54: Based on the paragraph-level retention data obtained in step S53, sort the paragraphs from lowest to highest retention rate (i.e., the paragraphs with the worst retention performance are ranked higher); classify the paragraphs ranked in the top K as low retention paragraphs, where the K value is dynamically determined based on the total number of paragraphs at the key nodes. ,in This represents the total number of natural segments at the key nodes. For a preset ratio threshold, for example when When ≤3, =1 (i.e., K=1); when 3 < When ≤15, =0.2; when >15 o'clock, =3 / This rule ensures that the value of K is at least 1 and at most 3. At the same time, if there is a natural segment with a retention rate of less than 0.3 and its ranking is outside the top K positions, then the segment is added to the low retention segment set. That is, the low retention segment is the union of the top K positions and the retention rate < 0.3.

[0123] In other embodiments, retention rates below a preset retention threshold (e.g.) Natural paragraphs with a value <0.3 are directly classified as low retention paragraphs, without being restricted by the K value.

[0124] Step S55: Read the exemption flag set in step S42 and determine whether each low retention paragraph is an exemption optimization paragraph. If step S42 has never been executed, that is, the user has not triggered the jump prompt, the exemption flag attribute value of all paragraphs is false by default. If the exemption flag attribute value of a low retention paragraph is true, then the paragraph is removed from the set to be optimized and no subsequent collapse preview or collapse and hide operation is performed on it.

[0125] Step S56: For the remaining low retention paragraphs, i.e. low retention paragraphs other than the exempted optimization paragraphs, obtain the boundary coordinates of the document object model elements of each low retention paragraph and the boundary coordinates of the visible area of ​​the current viewport, and determine whether each low retention paragraph is located within the visible area of ​​the current viewport.

[0126] The phrase "within the visible area of ​​the current viewport" means that at least a portion of the paragraph (the top or bottom boundary) is located within the visible area boundary of the current viewport.

[0127] It should be noted that the optimization operation in step S56 is executed according to the following priority order: (1) If the current user is not a first-time visitor to the page, and the paragraph-level retention data of the page already exists in the server-side cache, then the optimization operations of S56 to S58 are executed immediately after the page is loaded; (2) If the current user is a first-time visitor, but the global bounce rate of the page is higher than 1.2 times the average of the page type, then the optimization operations of S56 to S58 are executed 1 second after the page is loaded; if the user has already scrolled within this 1 second... If the operation is cancelled, the scheduled execution will be cancelled and the operation will be handled according to condition (3); (3) If the current user is visiting for the first time and the global bounce rate is normal, the optimization operation of S56~S58 will be delayed. The client listens to the user's scrolling behavior. When the user finishes scrolling at least once and stops (i.e. no new scrolling trigger signal is detected within 500ms of the scrolling listener event), the optimization operation of S56~S58 will be executed; if the user does not perform any scrolling operation within 30 seconds after the page is loaded, the optimization operation of S56~S58 will be executed automatically after 30 seconds.

[0128] Step S57: Collapse preview is performed on the low-retention paragraphs located in the visible area of ​​the current viewport. Collapse preview means that the main text content of the corresponding paragraph is collapsed, and only the paragraph title or the first sentence of the paragraph is retained as a preview summary, accompanied by the display of the expand control. The expand control is configured to restore the full content display of the paragraph in response to the user's click.

[0129] Collapse preview is achieved through the following front-end operations. For example, the client obtains the DOM element corresponding to the target paragraph, iterates through all child nodes under that element, and excludes the paragraph title node (i.e., the first child node). <h>Set the display property of all child nodes other than the text node containing the tag or the first sentence to none or hidden, or remove these child nodes from the DOM tree and store them in a memory variable for later restoration; insert an expand control after the paragraph title node. <button>or The element is bound to a click event listener. When the user clicks the control, previously hidden or removed child nodes are restored to their original positions in the DOM tree, and the control text is switched from expanded to collapsed.

[0130] In some embodiments, the preview summary when the preview is collapsed may also include word count information or core keyword tags for the paragraph, so that users can quickly understand the main idea of ​​the collapsed content.

[0131] Step S58: Collapse and hide low-retention paragraphs located outside the visible area of ​​the current viewport. Collapse and hide means removing the complete content of the corresponding paragraph from the page document object model or setting it to an invisible state. Collapsed and hidden paragraphs do not retain any preview information in order to maximize the release of rendering resources outside the viewport.

[0132] Specifically, the collapsing and hiding is achieved through the following front-end operations: the client obtains the DOM element corresponding to the target paragraph, removes it entirely from the DOM tree and stores it in the memory cache, and inserts a placeholder node at that position. Placeholder elements (with a height of 0 or 1px) are used to maintain page layout stability. When the user scrolls to a collapsed paragraph that is about to enter the viewport (i.e., the distance between the top edge of the paragraph and the top edge of the current viewport is less than 1.5 times the viewport height), the client retrieves the DOM element of that paragraph from the memory cache, re-inserts it into the placeholder node's position, and sets the paragraph to a collapsed preview state. Setting it to invisible is an alternative implementation method, specifically setting the paragraph's `style.visibility` property to `hidden` or `style.display` to `none`, and recording its original state for later restoration.

[0133] In some embodiments, the collapse / hide operation is automatically triggered when the page scrolls to the point where the hidden paragraph is about to enter the viewport, switching the hidden state to a collapsed preview state to avoid a sense of blank content when the user scrolls. Specifically, the client listens for scroll events, and when it detects that the top edge of any collapsed / hidden paragraph is less than 1.5 times the viewport height from the top edge of the current viewport, it triggers the state switch for that paragraph in advance, restoring it to a placeholder node from the document object model and setting it to a collapsed preview state (only displaying the paragraph title / first sentence and the expand control), but not automatically expanding the full text; after the paragraph enters the viewport, it remains in the collapsed preview state and does not automatically switch to the fully expanded state until the user manually clicks the expand control. The same paragraph can only be in one of the three states—collapse / hide, collapsed preview, or fully expanded—at any given time, and the collapsed preview state has higher priority than the collapsed / hide state.

[0134] This invention distinguishes between collapsing previews and folding hides inside and outside the viewport, respectively. It retains paragraph summaries inside the viewport to reduce the abruptness of reading caused by optimization operations, and completely hides them outside the viewport to minimize the rendering load of low-quality content and negative user experience. In addition, it combines the protection of exempted optimized paragraphs with markup to ensure that the paragraphs that the user intends to jump to are not interfered with by any optimization strategies, thus achieving the synergy between intelligent content optimization and the user's real-time intentions.

[0135] like Figure 3 As shown, the present invention also provides an AI marketing content iteration optimization system based on content conversion data, including: a server including: a browsing sequence acquisition module, used to acquire the content browsing sequence of converted users within a preset time window, and record each consecutive access chain as a candidate path; a path optimization module, used to derive path weights based on the access time of adjacent content in each candidate path, and filter target paths and key nodes based on the conversion value and path weight of each candidate path, wherein the conversion value is the achievement amount of a preset conversion event.

[0136] The client includes: a semantic connection module, which records the key node currently visited by the user as the current node, extracts the last paragraph of the adjacent content before the current node in the target path, matches the last paragraph of the question with the paragraph of the key node to locate the corresponding paragraph, and generates a prompt to jump to the corresponding paragraph; a visual replacement module, which responds to the prompt triggered by the user, replaces the static material of the first screen of the key node in the visible area of ​​the current viewport with dynamic material related to the theme of the first screen; and a retention optimization module, which obtains the paragraph-level retention data of each paragraph of the key node, classifies paragraphs with low retention data as low retention paragraphs, collapses the preview of low retention paragraphs that are located in the visible area of ​​the current viewport but are not corresponding paragraphs, and collapses and hides low retention paragraphs that are located outside the visible area of ​​the current viewport but are not corresponding paragraphs.

[0137] The AI ​​marketing content iteration optimization system based on content conversion data provided in this embodiment of the invention can execute an AI marketing content iteration optimization method based on content conversion data provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0138] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0139] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention. < / button> < / h>

Claims

1. An AI-powered marketing content iteration and optimization method based on content conversion data, characterized in that, include: The server obtains the content browsing sequence of converted users within a preset time window and records each consecutive access chain as a candidate path. The path weight is derived based on the access time of adjacent content in each candidate path. The target path and key nodes are selected based on the conversion value of each candidate path and the path weight. The conversion value is the amount of a preset conversion event achieved. The client records the key node currently accessed by the user as the current node, extracts the last paragraph of the question from the adjacent content before the current node in the target path, matches the last paragraph with the paragraph of the key node to locate the corresponding paragraph, and generates a prompt to jump to the corresponding paragraph. When the user triggers the prompt, the key node is loaded, the corresponding paragraph is marked as an exemption optimization paragraph, the visible area of ​​the current viewport when the key node is loaded and scrolled to the corresponding paragraph is defined as the first screen visible area, and the static material of the key node in the first screen visible area is replaced with dynamic material associated with the first screen theme of the first screen visible area. Obtain paragraph-level retention data for each paragraph of the key node, classify paragraphs with low retention data as low retention paragraphs, and for the remaining low retention paragraphs other than the exempted optimization paragraphs, perform either collapse preview or collapse and hide depending on whether they are located in the current viewport.

2. The AI ​​marketing content iterative optimization method based on content conversion data according to claim 1, characterized in that, The process of deriving path weights based on the access times of adjacent content in each candidate path includes: Remove invalid nodes whose dwell time is less than the preset time, and retain valid nodes to form a valid browsing sequence; A sequence of access intervals is generated based on the access time difference between adjacent nodes in the valid browsing sequence. The standard deviation of the access interval sequence is calculated as the dispersion value. The path smoothness is determined based on the dispersion value, and the dispersion value is negatively correlated with the path smoothness. The path weight is derived based on the path smoothness, and the path weight is positively correlated with the path smoothness.

3. The AI ​​marketing content iterative optimization method based on content conversion data according to claim 2, characterized in that, After generating the access interval sequence based on the access time difference between adjacent nodes in the effective browsing sequence, the method further includes: The mean of the access interval sequence is obtained. When the mean exceeds a preset duration, an attenuation coefficient is generated based on the ratio of the mean to the preset duration, and the path weight is multiplied by the attenuation coefficient. The attenuation coefficient is negatively correlated with the ratio.

4. The AI ​​marketing content iterative optimization method based on content conversion data according to claim 1, characterized in that, The target path and key nodes for screening include: Obtain the total conversion value and number of launches associated with each candidate path, and generate a comprehensive index based on the ratio of the total conversion value to the number of launches and the path weight. The comprehensive index is the weighted sum or product of the ratio and the path weight. The candidate path with the highest comprehensive index is determined as the target path; Nodes in the target path whose transfer probability is lower than the first preset level or whose single conversion value after passing through the corresponding node is higher than the second preset level are identified as candidate key nodes. The transition probability is the ratio of the number of times the current node visits the next node to the total number of times the current node is visited; The single conversion value is the ratio of the number of conversion events achieved after the current node to the total number of sessions through the current node; The candidate key nodes are sorted according to the product of the reciprocal of the transition probability and the single transformation value, and the top N candidate key nodes are determined as the key nodes.

5. The AI ​​marketing content iterative optimization method based on content conversion data according to claim 1, characterized in that, Extracting the last segment of the question from the adjacent content preceding the current node in the target path includes: Obtain the unclosed question at the end of the adjacent content preceding the current node in the target path; The unclosed question is a sentence that begins with an interrogative word and does not end with a period or exclamation mark, or the unclosed question is used as the final question. If there is no unclosed question, then extract the core statement from the last paragraph of the adjacent content, and convert the core statement into an interrogative sentence to serve as the question in the last paragraph.

6. The AI ​​marketing content iterative optimization method based on content conversion data according to claim 5, characterized in that, Matching the final question with the paragraphs of the key nodes to locate the corresponding paragraphs includes: The semantic matching of the final question with each paragraph of the key node is performed to obtain the semantic matching degree of each paragraph. If the highest semantic matching degree is lower than the preset matching threshold, then the process switches to matching the core nouns in the final question with the first sentence of each paragraph of the key node, and locating the corresponding paragraph based on the matching results.

7. The AI ​​marketing content iterative optimization method based on content conversion data according to claim 1, characterized in that, The replacement with dynamic materials associated with the homepage theme includes: Extract the core keywords of the first screen text of the key node within the visible area of ​​the current viewport as the first screen theme; Retrieve historical conversion material records associated with the key nodes in the target path, and select several dynamic materials with the highest conversion rates from the dynamic material library as candidate materials; The TF-IDF method is used to generate keyword vectors for the theme of the first screen and text description tag vectors for each candidate material. The cosine similarity between the two is calculated, and the candidate material with the highest cosine similarity is selected to replace the static material of the first screen.

8. The AI ​​marketing content iterative optimization method based on content conversion data according to claim 1, characterized in that, Obtaining paragraph-level retained data for each paragraph of the key node includes: Record the last fully visible paragraph within the current viewport's visible area when the user stops scrolling as the current reading paragraph; If the user does not continue scrolling or perform a jump operation after reading the current paragraph, and the current paragraph is not the last paragraph of the key node, then the corresponding paragraph will be counted as a paragraph-level jump. The ratio of the number of paragraph-level bounces to the number of times each paragraph is displayed is calculated and used as the paragraph-level retention data.

9. The AI ​​marketing content iterative optimization method based on content conversion data according to claim 8, characterized in that, Sort the retention rates of the paragraph-level retention data from low to high, and divide the natural paragraphs in the top K positions into low retention paragraphs. Collapse previews for all low-retention segments within the visible area of ​​the current viewport, excluding the exempted optimized segments. Collapse and hide paragraphs outside the visible area of ​​the current viewport that have low retention, excluding the exempted optimized paragraphs.

10. An AI marketing content iteration optimization system based on content conversion data, used to execute the AI ​​marketing content iteration optimization method based on content conversion data as described in any one of claims 1 to 9, characterized in that: include: The server includes: The browsing sequence acquisition module is used to acquire the content browsing sequence of converted users within a preset time window and record each consecutive access chain as a candidate path; The path optimization module is used to derive path weights based on the access times of adjacent content in each candidate path, and to filter target paths and key nodes based on the conversion value of each candidate path and the path weights, wherein the conversion value is the amount of a preset conversion event achieved. The client includes: The semantic connection module is used to record the key node currently visited by the user as the current node, extract the last paragraph of the question sentence of the adjacent content before the current node in the target path, match the last paragraph of the question sentence with the paragraph of the key node to locate the corresponding paragraph, and generate a prompt to jump to the corresponding paragraph. The visual displacement module is used to replace the static material of the key node in the visible area of ​​the current viewport with dynamic material associated with the theme of the first screen in response to the prompt triggered by the user. The retention optimization module is used to obtain paragraph-level retention data for each paragraph of the key node, classify paragraphs with low retention data as low retention paragraphs, collapse the preview of low retention paragraphs that are located within the visible area of ​​the current viewport but are not the corresponding paragraphs, and collapse and hide low retention paragraphs that are located outside the visible area of ​​the current viewport but are not the corresponding paragraphs.

Citation Information

Patent Citations

  • Analysis method and system for website user access path

    CN103823883A

  • Conversion link analysis method, system and computer device based on big data

    CN112529634B