A collaborative management method and system for content creation that responds to the needs of interest groups

By employing multi-perspective, multi-modal coding and hierarchical optimization, the problem of insufficient interest-based demand analysis in existing technologies has been solved, thereby improving the stability of candidate content and the utilization of feedback, and forming an efficient collaborative management mechanism for content creation.

CN122412698APending Publication Date: 2026-07-17
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Filing Date
2026-05-18
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies lack joint analysis and collaborative optimization mechanisms that address the overall needs of interest groups, resulting in high homogeneity of candidate content, insufficient stability in the selection process, and inadequate utilization of feedback, making it difficult to accurately reflect the dynamic changes in the needs of interest groups.

Method used

By collecting multi-source heterogeneous data and performing multi-perspective, multi-modal encoding, we generate feature vectors for circle demand and content features, determine the target creation direction, generate candidate creation objects and perform interrelationship modeling and optimization ranking, publish content and collect feedback data, identify touchpoint mismatch types, and execute hierarchical optimization processing strategies.

Benefits of technology

It improves the accuracy and completeness of interest group demand identification, reduces the homogeneity of candidate content, improves the matching effect and optimization efficiency of content creation and group demand, and forms a closed-loop content creation collaborative management mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122412698A_ABST
    Figure CN122412698A_ABST
Patent Text Reader

Abstract

This invention discloses a content creation collaborative management method and system that responds to the needs of interest groups, comprising: collecting multi-source heterogeneous data corresponding to the target interest group; performing multi-view, multi-modal encoding on the multi-source heterogeneous data to obtain a group demand feature vector and a content feature vector; determining at least one target creation direction based on the group demand feature vector; generating at least two candidate creation objects for each target creation direction; performing interrelationship modeling and optimization sorting on the candidate set formed by the candidate creation objects to determine the target content; publishing the target content and collecting content response feedback data corresponding to the target content; identifying touchpoint mismatch types based on the content response feedback data; and determining and executing a hierarchical optimization processing strategy according to the touchpoint mismatch type. This invention can improve the matching degree between content and interest group needs, reduce the homogenization of candidate content, and enhance the dynamic optimization capability of content creation collaborative management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of Internet information processing technology, specifically relating to a content creation collaborative management method and system that responds to the needs of interest groups. Background Technology

[0002] With the rapid development of internet platforms, social media platforms, and content distribution platforms, content production and dissemination targeting different interest groups has become an important part of digital content services. Platforms typically analyze user interests using data such as browsing history, likes, comments, reposts, and tags, and use this information for content recommendations, topic planning, and content delivery management. Simultaneously, some content production systems have introduced technologies such as text analysis, image analysis, user profiling, and dissemination effect statistics to assist in content creation, operation, and effect evaluation. However, existing technologies mostly focus on identifying single user preferences or localized trending content, rarely considering the overall needs of interest groups and conducting joint analysis of multi-source heterogeneous data within those groups. Especially in content creation management scenarios, demand analysis, content generation, candidate solution selection, posting feedback analysis, and subsequent adjustments are often handled independently in different stages, lacking a unified collaborative mechanism centered on the needs of interest groups. This results in content creation direction often relying on static tags, single-event trend statistics, or manual experience judgment, making it difficult to accurately reflect the dynamic changes in the needs of interest groups.

[0003] Furthermore, existing technologies for processing content candidates typically involve scoring and ranking individual candidates individually, with little consideration given to the relationships between different candidates, such as their similarity, substitutability, and differentiation. Therefore, during the selection of content topics, generation of content versions, and determination of publishing targets, issues such as high homogeneity of candidate content, insufficient differentiation, and unstable selection results can easily arise, affecting the final matching effect between the content and the needs of interest groups.

[0004] Furthermore, regarding the utilization of feedback after content publication, most existing technologies only evaluate effectiveness based on outcome metrics such as readership, play count, likes, or shares, and then make simple manual parameter adjustments or subsequent operational adjustments accordingly. They lack fine-grained identification and classification of different types of feedback data, making it difficult to distinguish between different mismatches such as insufficient content appeal, insufficient content comprehension, insufficient interactive response, or insufficient relationship conversion. Due to the lack of a layered processing mechanism based on feedback data, content optimization often only allows for localized modifications to the current content, making it difficult to further link with the front-end processes of determining the creative direction, generating candidate content, and ranking candidates. This results in insufficient utilization of feedback, an incomplete content optimization chain, and affects the overall efficiency and accuracy of collaborative content creation management. Summary of the Invention

[0005] The purpose of this invention is to provide a content creation collaborative management method and system that responds to the needs of interest groups, so as to at least solve the problems in the existing technology that the content creation management process lacks a joint analysis and collaborative optimization mechanism oriented towards the overall needs of interest groups, resulting in high homogeneity of candidate content, insufficient stability of selection, and insufficient utilization of feedback.

[0006] To achieve the above objectives, the present invention provides a content creation collaborative management method that responds to the needs of interest groups, including: collecting multi-source heterogeneous data corresponding to the target interest group; Multi-source heterogeneous data is encoded from multiple perspectives and in multiple modes to obtain the demand feature vector and content feature vector for each circle. Determine at least one target creative direction based on the feature vector of circle demand; Generate at least two candidate creative objects for each target creative direction; The candidate set of candidate creative objects is modeled for mutual relationships and optimized and sorted to determine the target content; Publish the target content and collect the corresponding content response feedback data; Identify touchpoint mismatch types based on content response feedback data; Based on the touchpoint mismatch type, a hierarchical optimization processing strategy is determined and implemented to optimize at least one of the following: target content, circle demand feature vector, target creative direction, candidate creative object or candidate set optimization ranking.

[0007] A further preferred technical solution is that obtaining the circle demand feature vector and content feature vector includes: dividing multi-source heterogeneous data into perspectives based on content semantic perspective, dissemination interaction perspective, and temporal context perspective; extracting and encoding features for text data, visual data, audio data, and behavioral data under each perspective to obtain corresponding single-perspective single-modal features; aligning and fusing the single-perspective single-modal features to obtain joint features; and aggregating the circle-side features and content-side features based on the joint features to obtain the circle demand feature vector and content feature vector.

[0008] A further preferred technical solution, wherein determining at least one target creative direction based on the feature vector of circle demand includes: performing theme clustering on the feature vector of circle demand to obtain several demand themes; combining timestamp data from multi-source heterogeneous data to perform time-series evolution analysis on each demand theme to obtain the demand attention corresponding to each demand theme; statistically analyzing the existing content coverage corresponding to each demand theme based on the content feature vector; calculating the difference between the demand attention of each demand theme and the existing content coverage; and determining the demand theme whose difference value meets the preset condition as the target creative direction.

[0009] A further preferred technical solution includes the following steps for determining the target content: extracting candidate features from each candidate creative object, wherein the candidate features include at least semantic features, content form features, adaptation features, and risk features; determining similarity association parameters between any two candidate creative objects based on the candidate features and constructing a candidate relationship matrix; aggregating the candidate relationship matrix to obtain a candidate relationship score corresponding to each candidate creative object; performing a weighted calculation based on the matching degree between each candidate creative object and the feature vector of the circle's demand, as well as the candidate relationship score, according to a preset weight, to obtain a comprehensive score corresponding to each candidate creative object; and optimizing and ranking the candidate set according to the comprehensive score, and determining the candidate creative object ranked first as the target content.

[0010] In a further preferred embodiment, the content response feedback data includes content touchpoint feedback data, interaction touchpoint feedback data, and relationship touchpoint feedback data. The content touchpoint feedback data represents the user's exposure, access, dwell time, and completion status regarding the target content. The interaction touchpoint feedback data represents the user's actions such as liking, commenting, forwarding, collecting, or creating derivative works based on the target content. The relationship touchpoint feedback data represents the user's formation of relationships such as following, revisiting, joining circles, or spreading based on the target content. The method of identifying touchpoint mismatch types based on content response feedback data includes: evaluating the credibility of content touchpoint feedback data, interaction touchpoint feedback data, and relationship touchpoint feedback data to obtain the feedback credibility of each content response feedback data, determining the content response feedback data whose feedback credibility meets preset conditions as credible feedback data, and identifying touchpoint mismatch types based on credible feedback data.

[0011] In a further preferred embodiment, the identification of touchpoint mismatch types based on trusted feedback data includes: identifying whether the target content has an attraction mismatch type or a comprehension mismatch type based on the content touchpoint feedback data in the trusted feedback data; identifying whether the target content has an interaction mismatch type based on the interaction touchpoint feedback data in the trusted feedback data; and identifying whether the target content has a relationship transformation mismatch type based on the relationship touchpoint feedback data in the trusted feedback data. The attraction mismatch type is insufficient conversion of exposure to visits for the target content; the understanding mismatch type is insufficient dwell time or completion of the target content; the interaction mismatch type is insufficient interaction conversion of the target content; and the relationship conversion mismatch type is insufficient relationship conversion of the target content.

[0012] A further preferred technical solution is that the step of determining and executing a layered optimization processing strategy based on the touchpoint mismatch type includes: when the touchpoint mismatch type is identified as an attraction mismatch type, replacing or adjusting at least one of the title information, cover information, or opening introductory segment of the target content; when the touchpoint mismatch type is identified as a comprehension mismatch type, replacing or adjusting at least one of the content structure, information density, or expression template of the target content; when the touchpoint mismatch type is identified as an interaction mismatch type, replacing or adjusting at least one of the interactive guidance, topic linking method, and publication time of the target content; and when the touchpoint mismatch type is identified as a relationship conversion mismatch type, replacing or adjusting at least one of the attention guidance content, subsequent follow-up content, or circle entry information of the target content.

[0013] A further preferred technical solution includes, in which the step of determining and executing a hierarchical optimization processing strategy based on the touchpoint mismatch type, further comprising: after the candidate set is optimized and sorted, retaining at least one candidate creation object that has not been determined as the target content as a backup candidate creation object; when the touchpoint mismatch type meets the candidate reordering condition, determining the feedback correction parameter based on the content response feedback data corresponding to the target content, and calculating the replacement score corresponding to each backup candidate creation object based on the matching degree between each backup candidate creation object and the feature vector of the circle demand, the candidate relationship score corresponding to each backup candidate creation object, and the feedback correction parameter; reordering the backup candidate creation objects according to the replacement score, and determining the backup candidate creation object ranked first as the updated target content.

[0014] A further preferred technical solution, wherein the step of determining and executing a hierarchical optimization processing strategy based on the touchpoint mismatch type, further includes: when at least two target contents determined by the same target creation direction are consecutively identified as the same touchpoint mismatch type, determining that there is a demand comprehension bias in the current circle demand feature vector; extracting mismatch correction features based on the reliable feedback data corresponding to the same touchpoint mismatch type, mapping the mismatch correction features to the corresponding feature dimension of the circle demand feature vector, adjusting the weight of the corresponding feature dimension according to a preset correction coefficient, and obtaining an updated circle demand feature vector; and re-determining the target creation direction, regenerating candidate creation objects, and re-optimizing and sorting the candidate set based on the updated circle demand feature vector.

[0015] This invention also provides a content creation collaborative management system that responds to the needs of interest groups, including: a data acquisition module, a data encoding module, a creation direction determination module, a candidate creation object generation module, a candidate set modeling and sorting module, a content publishing and feedback acquisition module, a touchpoint mismatch identification module, and a hierarchical optimization processing module. Each module is used to execute the corresponding steps of the above methods.

[0016] The content creation collaborative management method and system of the present invention have the following beneficial effects: 1. This invention uses multi-perspective and multi-modal encoding of multi-source heterogeneous data corresponding to interest circles to extract demand information related to content creation from multiple dimensions such as circle semantic features, propagation and interaction features, and temporal context features, thereby improving the accuracy and completeness of interest circle demand identification.

[0017] 2. This invention generates candidate creative objects and models and sorts the relationships between candidate sets, no longer limiting itself to scoring individual candidate content independently. This helps reduce the homogeneity among candidate content and improves the stability and distinguishability of the target content determination process.

[0018] 3. This invention identifies different touchpoint mismatch types based on content response feedback data and further executes a layered optimization processing strategy. It can not only adjust the current target content, but also optimize the candidate creation objects and the front-end demand understanding results in a coordinated manner, thereby forming a closed-loop content creation collaborative management mechanism oriented towards the needs of interest groups, improving the matching effect between content creation and the needs of the groups and the efficiency of subsequent optimization. Attached Figure Description

[0019] Figure 1 This is a schematic diagram illustrating the operation of a content creation collaborative management method in one embodiment of the present invention; Figure 2 This is a framework diagram of a content creation collaborative management system in one embodiment of the present invention. Detailed Implementation

[0020] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the scope of protection of the present invention.

[0021] In existing content production and distribution scenarios, platforms typically analyze user behavior such as browsing, liking, commenting, and forwarding to understand content interests and use this information for content recommendations, topic planning, or placement management. However, existing solutions often focus on localized analysis of single hot topics, single feedback, or single-modal data, lacking a unified modeling mechanism for the overall needs of interest groups. Furthermore, there is a general lack of a coherent collaborative processing chain between content direction determination, candidate content generation, candidate content selection, utilization of post-publication feedback, and subsequent adjustments. This can easily lead to inaccurate content direction, strong homogeneity among candidate content, and the inability to perform only partial repairs after publication, hindering the integration with the front-end generation process. Therefore, this invention provides a content creation collaborative management method that responds to the needs of interest groups. By incorporating group need identification, candidate content selection, content publication feedback, and hierarchical optimization into a unified processing chain, the content creation process can be dynamically and collaboratively managed around the needs of interest groups.

[0022] This invention provides a collaborative management method for content creation that responds to the needs of interest groups, such as... Figure 1 As shown, the method specifically includes: 1) collecting multi-source heterogeneous data corresponding to the target interest circle; 2) performing multi-view multi-modal encoding on the multi-source heterogeneous data to obtain circle demand feature vectors and content feature vectors; 3) determining at least one target creation direction based on the circle demand feature vectors; 4) generating at least two candidate creation objects for each target creation direction; 5) modeling the interrelationships and prioritizing the candidate set formed by the candidate creation objects to determine the target content; 6) publishing the target content and collecting content response feedback data corresponding to the target content; 7) identifying touchpoint mismatch types based on the content response feedback data; 8) determining and executing a hierarchical optimization processing strategy according to the touchpoint mismatch type to optimize at least one of the target content, circle demand feature vectors, target creation directions, candidate creation objects, or candidate set prioritization.

[0023] Regarding the collection of multi-source heterogeneous data corresponding to target interest groups, this multi-source heterogeneous data refers to a collection of data from different sources, with different structures and modalities, used to collectively reflect the demand status and content response status of the group. This multi-source heterogeneous data can include textual data such as posts, comments, private messages, tags, title text, and long text content published by group members; multimedia data such as cover images, video keyframes, and audio clips; and behavioral data such as likes, dwell times, reposts, favorites, repeat visits, and secondary creations. It can also further include associated information such as timestamps, platform identifiers, content identifiers, and user identifiers. By uniformly collecting this data, we can avoid the information gaps caused by judging group needs based solely on single text content, and also provide a common data foundation for subsequent demand modeling, content modeling, and feedback identification.

[0024] This section discusses multi-perspective, multi-modal encoding of multi-source heterogeneous data. The multi-modal approach does not only process text data but also simultaneously processes text, visual, audio, and behavioral data. The multi-perspective approach does not simply encode the literal meaning of the content but organizes and represents the same batch of data from different observation angles, such as content semantics, dissemination interaction, and temporal context. In practice, the collected data can first be divided according to perspective: for example, post text, title text, tag text, and comment text can be categorized into the content semantic perspective; likes, comments, reposts, favorites, and revisits can be categorized into the dissemination interaction perspective; and timestamps, dissemination order, and the chronological order of content appearance can be categorized into the temporal context perspective. Then, features are extracted modally under each perspective: semantic features are extracted from text data, visual features from images or video frames, audio features from audio clips, and behavioral statistical features or behavioral trajectory features from behavioral sequences. Next, the modal features from each perspective are aligned and fused to obtain joint features; then, based on these joint features, the circle-side features and content-side features are aggregated to obtain the circle demand feature vector and the content feature vector.

[0025] The circle-based demand feature vector is used to characterize the demand status of the target interest circle within the current time window. This vector does not simply describe the theme of a single piece of content, but rather comprehensively reflects the circle's focus, emotional inclination, interaction preferences, and demand change trends at a certain stage. For example, the same circle may simultaneously exhibit high attention to a certain type of topic, high acceptance of a certain type of expression, and sensitivity to a certain type of dissemination rhythm within a certain period; all of these can be encoded into the circle's demand feature vector. The content feature vector, on the other hand, is used to characterize the semantic connotation, presentation, and dissemination attributes of the content to be analyzed or candidate content. In other words, the former focuses on answering what kind of content the circle currently needs, while the latter focuses on answering what content characteristics the current or candidate content possesses. By forming these two types of vectors respectively, subsequent processing can be carried out around the matching relationship between the demand side and the content side.

[0026] Regarding the determination of at least one target creative direction based on the feature vector of circle demand, where the target creative direction is the content direction that is more worthy of priority response in the current interest circle, the specific implementation can be as follows: First, the data corresponding to the circle demand feature vector can be used for theme identification or theme clustering to obtain several demand themes; then, the changes of each demand theme within the current time window can be analyzed in combination with time information to determine which themes are in a state of significant increase in popularity, continued activity, or high attention value; at the same time, the distribution of existing content on each demand theme can be analyzed based on the content feature vector to determine which content directions still have significant supply-demand imbalances under the current demand status. For directions with large supply-demand imbalances and high potential for circle response, they can be identified as target creative directions. Thus, the system does not identify an isolated hot spot, but rather a content direction that directly corresponds to the demand status of the circle and has continued creative value.

[0027] After determining the target creative direction, at least two candidate creative objects are generated for each target creative direction. These candidate creative objects are not limited to a single title or topic, but can be understood as different content candidate schemes formed for the same creative direction. Each candidate creative object can correspond to at least one relatively complete creative expression unit, such as including topic expression, preliminary content structure, presentation method arrangement, and release adaptation information. The reason for generating at least two candidate creative objects for the same target creative direction is that different expression schemes may exist under the same direction, and the degree to which different expression schemes fit the needs of different target audiences is not entirely consistent. If only a single object is generated, subsequent comparison between candidates is impossible, and it is also not conducive to preventing content homogenization.

[0028] Subsequently, the candidate set of candidate creative objects is modeled and ranked to determine the target content. The key here is not to score each candidate creative object individually, but to treat all candidate creative objects in the same set as a whole. In practice, candidate features of each candidate creative object can be extracted first. These candidate features can include at least semantic features, content form features, adaptation features, and risk features. Then, based on the candidate features, similarity correlation parameters between any two candidate creative objects are determined, and a candidate relationship matrix is ​​constructed accordingly. This candidate relationship matrix is ​​used to characterize the similarity, substitutability, or complementarity relationships between candidate creative objects. For example, if two candidate creative objects are highly similar in theme expression, content structure, and presentation, their similarity correlation is high, indicating a strong tendency for substitution between them; if two candidate creative objects are related in theme but have significant differences in presentation angle or form, their similarity correlation is relatively low, indicating a greater likelihood of a complementary relationship. Next, the candidate relationship matrix can be aggregated to obtain the candidate relationship score for each candidate creative object. Then, combined with the matching degree between each candidate creative object and the feature vector of the circle's needs, a comprehensive score for each candidate creative object is calculated according to a preset weight, and the candidate set is optimized and ranked based on the comprehensive score. The candidate creative objects at the top of the ranking results can be used as target content. Through this processing method, the determination of target content does not consider the local score of a single candidate object, but rather the matching relationship between the candidate object and the needs of the circle, as well as the relative position and interrelationship of the candidate object in the entire candidate set, thereby improving the stability of target content determination and reducing the risk of homogenization of candidate content.

[0029] After identifying the target content, it is published, and corresponding content response feedback data is collected. This content response feedback data is not simply a matter of views or clicks, but rather a collection of data reflecting various response states of the target content during its dissemination. This data reflects both whether the target content has been noticed by users and their interactions and subsequent relationship changes after engaging with the content. By continuously collecting content response feedback data, the true response performance of the target content after publication can be obtained, providing a basis for subsequently identifying whether the target content deviates from the needs of the target audience.

[0030] After obtaining content response feedback data, touchpoint mismatch types are identified based on this data. Touchpoints can be understood as key locations or crucial links where the target content connects with users during dissemination and consumption. Touchpoint mismatch types refer to the target content failing to achieve the expected response at a certain key connection point, thus reflecting a deviation between the target content and the needs of the current user group. Specifically, some target content may have gained some exposure, but users did not further access or consume it, indicating a deviation in the attraction level; some target content can guide users to the content, but users stay for a short time or do not complete content consumption, indicating a deviation in the understanding level; some target content completes basic consumption, but users lack further interaction such as comments, reposts, or favorites, indicating a deviation in the interaction level; and some target content can generate some consumption and interaction, but fails to form a relationship of attention, repeat visits, entry into the user group, or spread, indicating a deviation in the relationship building level. By performing structured analysis of the content response feedback data, the current problems of the target content can be categorized into different touchpoint mismatch types. The significance of this step is that it no longer simply interprets all negative feedback as poor content performance, but rather further determines whether the deviation occurs in the content attraction, content comprehension, user interaction, or relationship conversion stages.

[0031] After identifying the touchpoint mismatch type, a tiered optimization strategy is determined and implemented based on the mismatch type to optimize at least one of the following: target content, sub-group demand feature vector, target creative direction, candidate creative objects, or candidate set optimization ranking. Tiered optimization means that the optimization actions are not fixed at the same level, but are applied to different processing levels according to the mismatch type and degree of deviation. Firstly, if the identification results indicate that the deviation mainly occurs at the target content level, the target content can be directly optimized. For example, adjustments can be made to the title, cover image, opening introduction, content structure, information density, or publication time of the target content itself to make the current target content more aligned with the sub-group's needs without changing the overall direction. In this case, the system does not need to go back to the front end to remodel, but rather to make local corrections based on the already determined content. Secondly, if the identification results indicate that the current target content is performing poorly not because the target creative direction itself is wrong, but because the objects selected from the candidate set are no longer optimal, then the candidate creative objects or candidate set optimization ranking can be optimized. For example, the matching of previously unselected candidate content with the current needs of the target audience can be reassessed, and the backup candidates can be rearranged based on feedback and correction information after publication. If a new candidate is in a better position in the updated ranking, it can be replaced with the new target content. In this way, the system does not simply repeat the generation of content, but uses existing candidate results for faster and more targeted replacement. Thirdly, if the identification results show that multiple target contents generated under the same target creation direction show the same type of mismatch, it can be determined that the problem no longer comes from the current target content or the current candidate selection, but from a deviation in the front-end's understanding of the needs of the target audience. In this case, correction information can be extracted from relevant feedback data to update the feature vector of the needs of the target audience, thereby redetermining the target creation direction, regenerating candidate content, and re-ranking the candidate set. In other words, the system will follow the path of feedback—needs understanding—direction determination—candidate generation—candidate ranking, and the back-end feedback will act in reverse on the front-end generation link, thus forming a closed-loop optimization.

[0032] In one specific embodiment, obtaining the circle demand feature vector and content feature vector includes: dividing multi-source heterogeneous data into perspectives based on content semantic perspective, dissemination interaction perspective, and temporal context perspective; extracting and encoding features for text data, visual data, audio data, and behavioral data under each perspective to obtain corresponding single-perspective single-modal features; aligning and fusing the single-perspective single-modal features to obtain joint features; and aggregating the circle-side features and content-side features based on the joint features to obtain the circle demand feature vector and content feature vector.

[0033] In this embodiment, the multi-source heterogeneous data can originate from content platforms, social media platforms, community platforms, or content operation backends. For ease of implementation, the multi-source heterogeneous data can be accessed via a unified data interface, such as JSON, CSV, Parquet, or event stream formats in message queues. Static historical data can be stored in MySQL, PostgreSQL, or Hive data warehouses; log data can be stored in real-time log systems integrated with Elasticsearch, ClickHouse, or Kafka; and image, audio, and video files can be stored in object storage systems. In a common implementation, the processing program can be deployed on a Linux server environment, using Python as the primary development language, combined with PyTorch or TensorFlow to build the coding model. For offline processing of large-scale behavioral logs, Spark can be used; for incremental processing of real-time interactive data, Flink or Kafka Streams can be used.

[0034] The so-called content semantic perspective refers to organizing data from the angles of what the content itself expresses, what its core theme is, its emotional tendency, and its expression method. This perspective primarily includes information that can characterize the semantics of the content, such as the post body, title text, tag text, comment text, cover text, and speech-to-text. The so-called dissemination and interaction perspective organizes data from the angles of how users reach, consume, and participate in the content. This perspective primarily includes data such as exposure records, access records, dwell time, completion rate, number of likes, number of comments, number of shares, number of favorites, number of repeat visits, and secondary creation behavior. The so-called temporal context perspective organizes data from the angles of the time sequence of content and interaction, the evolution of hot topics, time period distribution, and contextual changes. This perspective primarily includes data such as timestamps, content publication time, sequence of behavior, topic activity cycle, and intraday or weekly time period distribution. This perspective means that the same original data is not simply categorized into a unified data pool, but is reorganized according to three observation dimensions: content itself, interaction process, and time evolution. This allows subsequent encoding results to describe not only what the user saw, but also how the user participated and when that participation occurred.

[0035] For text data, preprocessing such as word segmentation, stop word removal, entity extraction, and sentiment annotation can be performed on post text, comment text, title text, and tag text. This data is then input into a pre-trained language model, such as BERT, RoBERTa, ERNIE, or their lightweight counterparts, to obtain text semantic vectors. For short text tags, word vector averaging, TextCNN, or Sentence-BERT can be used to obtain sentence vector representations. For visual data, cover images, accompanying images, and video keyframes can be uniformly scaled, normalized, and deduplicated. This data is then input into common visual encoding models, such as ResNet, Vision Transformer, and CLIP image encoder, to obtain visual feature vectors. If the video content is long, keyframes can be extracted at preset sampling intervals, and keyframe features can be temporally aggregated. For audio data, original audio, dubbing tracks, and background audio can be denoised, framed, and subjected to spectral transformation, such as converting to Mel spectrograms, before being input into models such as CNN, Audio Transformer, or wav2vec to obtain audio feature vectors. If no obvious audio modality exists on the platform, this part can be empty or set to the default vector. Regarding behavioral data, behaviors such as exposure, visits, dwell time, completion, likes, comments, reposts, favorites, repeat visits, and secondary creations can be represented as statistical features, sequence features, or graph structure features. For example, the click-through rate, average dwell time, interaction rate, and repeat visit rate of a certain content within a preset time window can be statistically analyzed; alternatively, user behaviors can be arranged in chronological order and input into sequence models such as GRU, LSTM, and Transformer Encoder for encoding to obtain behavioral feature vectors. In this embodiment, the above four types of data do not uniformly use the same model, but rather feature extraction and encoding are performed separately according to modal characteristics to form single-view, single-modal features applicable to different modalities.

[0036] From a content semantic perspective, textual semantic features, visual semantic features, and audio semantic features can be formed separately. For example, for a short video about a topic within a certain interest group, from a content semantic perspective, the text portion can be encoded by the title and subtitles to form textual semantic features, the cover and keyframes to form visual semantic features, and the voice-over content to form audio semantic features. From a communication and interaction perspective, behavioral statistical features or behavioral sequence features can be formed. For example, the number of exposures, conversion rates, dwell time, completion rates, like rates, comment rates, and forwarding rates of the same content can be encoded to obtain interactive behavioral features. From a temporal context perspective, the changes in popularity, behavior, and content distribution of a topic in different time windows can be encoded as temporal features. For example, time series can be constructed at the hourly, daily, or weekly granularity, and the changing trend within the sliding window can be calculated. Then, contextual features can be obtained through a temporal model. In this way, with the combination of three perspectives and four modalities, the actual system can obtain several sets of single-perspective, single-modal features. Here, single-perspective, single-modal features can be understood as local representation results obtained only for one modality from a certain observation perspective.

[0037] Since the feature dimensions of modalities such as text, vision, audio, and behavior are usually different, and the meanings of features from different perspectives are not entirely consistent, alignment and fusion are necessary before forming a unified result. In a common implementation, features from different modalities can be mapped to a unified dimension using linear mapping layers, fully connected layers, or projection layers, for example, mapping them to 128-dimensional, 256-dimensional, or 512-dimensional vectors to achieve dimension alignment. For data at different time scales, resampling, time window segmentation, or timestamp alignment can be used to ensure that data from different perspectives correspond within the same time window.

[0038] After alignment, joint features can be obtained using methods such as concatenation, weighted summation, attention fusion, or gating fusion. For example, text features can be set as follows: Visual features are Audio characteristics are behavioral characteristics are Then the joint feature F can be expressed conceptually as follows: ;in, The fusion weights for each modality can be preset or learned automatically during training. If an attention mechanism is used, the model can also automatically assign weights based on the importance of the current sample to different modalities, thereby enhancing the contribution of temporal and behavioral modalities in hotspot scenarios and enhancing the contribution of textual and visual modalities in deep content scenarios.

[0039] In this embodiment, the joint features are not directly used as the final output, but are further divided into circle-side features and content-side features, which are then aggregated separately. Circle-side features are primarily used to characterize the overall demand status of an interest circle. For example, the joint features corresponding to multiple pieces of content, comments, interactive behaviors, and temporal changes within the same circle within a certain time window can be aggregated along the circle dimension to obtain the overall demand vector of that circle in the current window. Aggregation methods can include average pooling, weighted pooling, attention pooling, or cluster center extraction. For example, if an interest circle has consistently seen high-frequency discussions, high dwell times, and high reposts around a certain topic in the last three days, the textual semantic features, behavioral features, and temporal features related to that topic will be enhanced during circle-side aggregation, thus reflecting higher weights in the circle's demand feature vector. Content-side features are primarily used to characterize the expressive characteristics of the specific content object itself. For example, the joint features corresponding to a single piece of content, such as the title, body text, cover image, audio, and initial interaction, can be aggregated along the content object dimension to obtain the content feature vector of that content. In other words, the feature vector of a social circle's overall demand reflects the content demand state that the current circle tends to have, while the content feature vector reflects the expressive and dissemination attributes of the current content object. Although both originate from joint features, they are aggregated differently: the former aggregates around the social circle, while the latter aggregates around the content object.

[0040] For example, in a content creation scenario targeting a gaming interest group, the system collects the following data within a certain time window: Part of the data consists of textual data such as post titles, strategy guides, comments, and bullet comments; part consists of visual data such as cover images, video keyframes, and character shots; part consists of audio data such as narration and background audio; and another part consists of behavioral data such as exposure counts, visit duration, completion rate, comment rate, forwarding rate, and repeat visit rate. The system first organizes this data from the perspectives of content semantics, dissemination interaction, and temporal context. Then, it encodes the title text, strategy guide text, and comment text using a pre-trained Chinese language model; encodes the cover image and video keyframes using an image encoding network; extracts audio vectors from the narration track using an audio encoding model; and extracts behavioral features from the behavioral logs such as exposure, completion, comments, and forwarding using sequence models or statistical models. Finally, a projection layer maps all these features to the same dimension, and attention is used to fuse them into joint features. Next, for the community aspect, the joint features of the game community over the past three days are aggregated by theme and time to form a community demand feature vector. For the content aspect, the joint features of a video to be analyzed are aggregated by the video object to form a content feature vector. If the community shows high attention to the theme of analyzing the strength of new characters in the past three days, and the time spent and reposting increases significantly in the evening, then this information will be encoded in the community demand feature vector. At the same time, if the title, cover, and script of a newly released video all revolve around this theme, then its content feature vector will also show high matching features in the corresponding dimensions. The system can then further perform processing such as determining the creation direction, generating candidate content, and sorting based on these two types of vectors.

[0041] In one specific embodiment, determining at least one target creative direction based on the feature vector of circle demand includes: performing theme clustering on the feature vector of circle demand to obtain several demand themes; combining timestamp data from multi-source heterogeneous data to perform time-series evolution analysis on each demand theme to obtain the demand attention corresponding to each demand theme; statistically analyzing the existing content coverage corresponding to each demand theme based on the content feature vector; calculating the difference between the demand attention of each demand theme and the existing content coverage; and determining the demand theme whose difference value meets the preset condition as the target creative direction.

[0042] The inputs in this embodiment mainly include the circle demand feature vector, the content feature vector, and timestamp data from multi-source heterogeneous data. The circle demand feature vector can come from the output of the pre-encoding process and is usually represented as a fixed-dimensional floating-point array, such as a 128-dimensional, 256-dimensional, or 512-dimensional vector; the content feature vector can be represented as a set of vectors corresponding one-to-one with the content object; the timestamp data can be stored using Unix timestamps, ISO8601 format time strings, or datetime fields in a database. The circle demand feature vector and the content feature vector can be stored in a vector database or a relational database, such as in a Milvus, FAISS index, PostgreSQL, or MySQL feature table, while the timestamp data can be stored together with the content identifier and circle identifier in a log table or a detail table.

[0043] In this embodiment, the topic clustering involves merging features representing the demand status of a specific interest group according to topic similarity to form several distinguishable demand topics. Specifically, demand feature vectors for each interest group can be extracted first, using a preset time window, such as the last 24 hours, the last 3 days, or the last 7 days. Vector samples related to the target interest group within this time window are organized into a feature set. Then, K-means, hierarchical clustering, spectral clustering, or density-based DBSCAN algorithms are used to cluster these vector samples. For example, in a short video interest group scenario, demand feature vectors collected within the last three days might aggregate to form several demand topics such as new product reviews, user tutorials, price comparisons, and interpretations of trending events. These demand topics are not manually specified but automatically formed by the clustering algorithm based on feature similarity. The clustering results can be saved in a PostgreSQL table or Redis cache with a structure of topic number + topic center vector + topic sample list for easy and quick subsequent retrieval.

[0044] After obtaining several demand topics, it is necessary to further determine the changing trends of each demand topic over time. To this end, timestamp data from multi-source heterogeneous data can be combined to perform time-series evolution analysis on each demand topic. In practice, data samples belonging to the same demand topic can be arranged in chronological order, and their frequency of occurrence, interaction intensity, or dissemination activity within continuous time windows can be statistically analyzed. For example, time series can be formed at the granularity of hours, days, or weeks, and the topic activity value in each time unit can be calculated. In a common implementation, a sliding window statistical method can be used: let the number of active samples for a certain demand topic in the t-th time window be... The corresponding comprehensive interaction value is The level of demand for this topic within this window can be represented by the following conceptual expression: ,in and Preset weights. To further illustrate the growth trend, the rate of change can be calculated between adjacent windows, for example... This allows for the differentiation between consistently popular topics and rapidly gaining popularity. For example, a new product review topic might maintain high activity over the past three days, while a topic analyzing a sudden trending event might rise rapidly within the last six hours. Both scenarios can be identified through time-series evolution analysis. This part can be completed using Spark, Flink, or a standard Python time-series script, and the results can be saved as a table structure consisting of topic number, time window, and demand level.

[0045] After determining the popularity or attention level of a demand topic, it's also necessary to assess the coverage of existing content on the platform. This can be done by statistically analyzing the existing content coverage for each demand topic based on content feature vectors. Specifically, the content feature vectors of published or databased content objects on the platform can be compared with the topic center vectors of each demand topic to calculate similarity. For example, cosine similarity, the reciprocal of Euclidean distance, or dot product similarity can be used to measure the proximity of a piece of content to a particular demand topic. When the similarity of a piece of content to a demand topic exceeds a preset threshold, it can be considered as covered content for that demand topic. Based on this, the number of covered content items under a given demand topic, the cumulative exposure, cumulative visits, or average completion rate of covered content can be statistically analyzed to determine the existing content coverage for that topic. For example, regarding the topic of usage tutorials, if the platform already has a large amount of highly similar tutorial content, and this content remains active in recent days, then the existing content coverage of this topic is high. Conversely, if a topic is frequently discussed, but the amount of existing content is small, or most of the existing content is old and has weak interaction, then the existing content coverage of this topic is low. To facilitate implementation, a vector index library can be pre-built for content feature vectors, and the matching between topic center vectors and content feature vectors can be completed through nearest neighbor search, thereby improving computational efficiency in large-scale content scenarios.

[0046] After obtaining the demand attention and existing content coverage, the difference value for each demand topic can be further calculated. This difference value reflects the degree of deviation between the intensity of demand and the coverage of supply. It can be understood as the gap between the current level of attention a demand topic receives from its target audience and the extent to which the platform's existing content covers it. In a simplified implementation, the difference value can be calculated as follows: ,in, This represents the difference value of the i-th demand topic. This indicates the level of attention paid to the topic of this demand. This indicates the existing content coverage corresponding to the topic of this demand. This represents a coverage reduction factor. If a topic has high demand but low existing content coverage, the difference value for that topic will be large, indicating that it is more likely to be a priority area for expansion. Conversely, if a topic has high demand but already high existing content coverage, the difference value may be low, suggesting that while the topic is popular, there is less need to continue investing creative resources in it. Furthermore, the difference value can be stored as a floating-point field and written into the analysis results table along with the topic number and time window.

[0047] After obtaining the difference values ​​for each demand theme, the demand themes whose difference values ​​meet preset conditions can be identified as target creative directions. Preset conditions can be set in one or more ways. A common method is the threshold method, for example, when the difference value is greater than a preset threshold τ, the demand theme is considered to have creative priority. Another method is the ranking method, for example, after sorting by difference value from high to low, the top N demand themes are selected as target creative directions. A combination of threshold and ranking can also be used, that is, first eliminating demand themes below a basic threshold, and then selecting the top few from the remaining themes. For example, in a digital enthusiast community, the system analysis yields four demand themes and their difference values: New Product Review 0.82, User Guide 0.73, Price Comparison 0.41, and Hot Topic Interpretation 0.68. If the preset condition is a difference value greater than 0.65, then New Product Review, User Guide, and Hot Topic Interpretation can be identified as target creative directions; if the preset condition is to select the top two, then New Product Review and User Guide can be identified as target creative directions. The target creative direction determined in this way is not simply based on trending topics, but takes into account both the current level of attention in the target audience and the existing content supply on the platform. Therefore, it is more suitable as input for generating subsequent candidate creative objects.

[0048] For example, in a content creation system targeting a fitness enthusiast community, the system collects and stores relevant post texts, short video covers, audio explanations, user dwell logs, and posting timestamps over the past seven days. Pre-processing yields the community's demand feature vector for the past seven days, along with the content feature vector of existing fitness content on the platform. The system first performs thematic clustering on the community's demand feature vectors, resulting in several demand themes such as weight-loss meal preparation, strength training basics, equipment usage techniques, and fitness myth debunking. Then, combining timestamp data, it analyzes the activity level of each demand theme over the past seven days, finding a significant increase in demand for weight-loss meal preparation and fitness myth debunking. Next, based on the content feature vectors, it analyzes the coverage of these themes by existing platform content, finding that while strength training basics have consistently been popular, there is already a large amount of existing content, while fitness myth debunking has seen a significant increase in popularity in recent days, but the platform has relatively little existing content. Finally, it calculates the difference values ​​for each theme and identifies weight-loss meal preparation and fitness myth debunking, with higher difference values, as target creation directions. The system can then continue to generate and sort candidate creative objects around these two directions, instead of treating all popular topics equally.

[0049] In one specific embodiment, the process of determining the target content includes: extracting candidate features of each candidate creative object, wherein the candidate features include at least semantic features, content form features, adaptation features, and risk features; determining similarity association parameters between any two candidate creative objects based on the candidate features and constructing a candidate relationship matrix; aggregating the candidate relationship matrix to obtain the candidate relationship score corresponding to each candidate creative object; performing a weighted calculation based on the matching degree between each candidate creative object and the feature vector of the circle demand and the candidate relationship score according to a preset weight to obtain the comprehensive score corresponding to each candidate creative object; optimizing and sorting the candidate set according to the comprehensive score, and determining the candidate creative object ranked first as the target content.

[0050] In this embodiment, candidate creative objects can be understood as different content scheme units generated around the same creative direction. Each candidate creative object can be represented using a structured data format, such as JSON, dictionary objects, or database records, and includes at least fields such as object identifier, title text, script summary, content format identifier, platform compatibility information, risk markers, and preset material descriptions. For example, in a short video content scenario, a candidate creative object can contain {title + script outline + cover draft + platform tags + risk tags}; in a text and image content scenario, a candidate creative object can contain {topic sentence + paragraph framework + image suggestions + layout tags + publishing end tags}. This data can be stored in MySQL, PostgreSQL, or MongoDB for subsequent feature extraction module calls.

[0051] The candidate features in this embodiment include at least semantic features, content format features, adaptation features, and risk features. Semantic features characterize the content attributes of the candidate creation object in terms of theme expression, semantic center, and information emphasis. Semantic features can be extracted using a pre-trained language model; for example, the title, script outline, and tag text can be concatenated and input into a BERT, RoBERTa, ERNIE, or Sentence-BERT model to obtain a fixed-dimensional semantic vector. Content format features characterize the presentation style of the candidate creation object, such as short videos, long videos, text and images, live stream scripts, question-and-answer content, comparative content, list-style content, etc. This type of feature can be mapped to a numerical vector using one-hot encoding, embedding encoding, or enumerated fields. Adaptation features characterize the compatibility between the candidate creation object and the publishing environment, such as the compatible platform type, video duration range, landscape / portrait format, publishing time period, and the common expression style of the target audience. This type of feature can be generated by a rule engine, tagging system, or simple classification model. Risk features are used to characterize the potential risk levels of candidate content during dissemination and use, such as risks associated with sensitive expressions, escalation of negative emotions, decreased completion rates, or platform review. These features can be obtained through sensitive word detection, sentiment analysis, statistical analysis of historical performance of similar content, or simple risk classification models. The above four types of features can be uniformly represented as floating-point vectors and concatenated to form a candidate feature vector, for example, stored as a NumPy array, a Torch Tensor, or an entry in a vector database.

[0052] After obtaining the candidate features of each candidate creative object, it is necessary to further determine the relationship between any two candidate creative objects. Here, a similarity association parameter is used to quantify the closeness between two candidate creative objects. In a common implementation, the candidate feature vectors of any two candidate creative objects can be denoted as... and The cosine similarity metric is used to calculate the similarity correlation parameters between the two. ,in, This parameter represents the similarity correlation between the i-th and j-th candidate creative objects. If the two candidate creative objects are highly similar in theme expression, content form, and suitability attributes, this parameter is high, indicating a strong tendency for substitution between them. If the two candidate creative objects are related in core theme but differ significantly in expression angle, presentation form, or suitability attributes, this parameter is low, indicating a higher degree of distinguishability between them, making them more suitable as distinct candidates in the candidate set. Besides cosine similarity, dot product similarity, Mahalanobis distance transform values, or relationship scores output by a small multilayer perceptron can also be used as similarity correlation parameters, as long as they ultimately reflect the similarity relationship between any two candidate creative objects.

[0053] In this embodiment, all candidate creative objects can be arranged in a preset order, and a candidate relationship matrix can be constructed using the candidate creative objects as row and column indices. If there are N candidate creative objects in the candidate set, an N×N matrix R can be constructed, where the element in the i-th row and j-th column is the aforementioned similarity association parameter. In this way, each element in the candidate relationship matrix represents the similarity relationship between two candidate creative objects. The diagonal elements of the matrix can be set to 1 or a preset constant to represent the relationship between the candidate creative object and itself. This matrix can be temporarily stored in memory as a two-dimensional floating-point array, or it can be saved as a matrix table in a database, a sparse matrix file, or a serialized object in a caching system. The candidate relationship matrix allows the overall relationships within the candidate set to be explicitly expressed, rather than simply scoring each candidate independently.

[0054] After constructing the candidate relationship matrix, the pairwise relationship information in the matrix needs to be transformed into a candidate relationship score for each candidate creative object. In a simplified implementation, the candidate relationship matrix can be aggregated row by row. For example, summing, averaging, or weighted averaging the elements in the i-th row of the matrix yields the candidate relationship score for the i-th candidate creative object. ,in, This represents the candidate relationship score for the i-th candidate creative object. If an average method is used, the candidate relationship score reflects the overall similarity level between the candidate creative object and other candidate creative objects; if a weighted method is used, higher weight can be assigned to the relationship with high-quality candidate objects. In a practical example, if candidate object A is highly similar to candidates B, C, and D, then A's candidate relationship score will be high, indicating that A is relatively close to other objects in the candidate set; while if candidate object E is only close to a few objects and significantly different from most objects, then E's candidate relationship score will be relatively low, indicating that E has stronger discriminative power in the candidate set. The aggregation process here can be implemented using conventional statistical methods, or it can be implemented using readout operations or Set Pooling methods in graph neural networks, but in this embodiment, row-wise aggregation is sufficient to meet the engineering implementation requirements.

[0055] In this embodiment, relying solely on candidate relationship scores is insufficient to determine the final target content; it is also necessary to consider the matching degree between each candidate creation object and the feature vector of the circle's needs. Specifically, the content representation vector corresponding to the i-th candidate creation object can be denoted as... Let the feature vector of circle demand be denoted as The matching degree between the two was calculated using cosine similarity. Then, the matching degree and candidate relationship score are weighted according to preset weights to obtain the comprehensive score for each candidate creative object. For example, the following conceptual formula can be used. ,in, This represents the overall score of the i-th candidate creation object. and Preset weights. and This can be configured through offline experiments; for example, when the platform places greater emphasis on catering to the needs of specific user groups, it can improve... When the platform places greater emphasis on the discriminativeness and stability within the candidate set, it can improve... Preset weights can be configured in the system configuration file and saved in YAML, JSON, or a database parameter table for easy adjustment during subsequent iterations.

[0056] After obtaining the comprehensive scores for each candidate content, the candidate set can be sorted from highest to lowest comprehensive score. Sorting can be implemented using conventional sorting algorithms, such as quicksort, heapsort, or by directly calling built-in sorting functions in programming languages ​​like Python and Java. After sorting, the candidate content ranked highest is determined as the target content. Here, "ranked highest" can be understood as the candidate content with the highest comprehensive score. In some implementations, a candidate retention pool can be set up to retain the top-ranked candidates, but in this embodiment, only the candidate content with the highest comprehensive score and highest ranking needs to be determined as the target content. For example, in a content scenario targeting a beauty interest group, three candidate content are generated around the theme of choosing summer sunscreen: A is an ingredient comparison explanation, B is a short video evaluating usage scenarios, and C is a price list with images and text. After extracting the candidate features of the three, the system constructs a candidate relationship matrix and obtains their respective candidate relationship scores; simultaneously, it calculates their matching degree with the feature vector of the current interest group's needs. If B best matches the needs of the current social circle and has the highest overall score after weighting candidate relationship score and matching degree, then B is identified as the target content. The target content obtained in this way is not necessarily the one that looks most like a trending topic, but rather an object that maintains a high degree of matching with the needs of the current social circle while being constrained by the overall relationships in the candidate set.

[0057] In one specific embodiment, the content response feedback data includes content touchpoint feedback data, interaction touchpoint feedback data, and relationship touchpoint feedback data. The content touchpoint feedback data represents the user's exposure, access, dwell time, and completion status of the target content. The interaction touchpoint feedback data represents the user's actions of liking, commenting, forwarding, collecting, or creating derivative works based on the target content. The relationship touchpoint feedback data represents the user's formation of relationships such as following, revisiting, joining circles, or spreading based on the target content. Identifying touchpoint mismatch types based on content response feedback data includes: evaluating the credibility of the content touchpoint feedback data, interaction touchpoint feedback data, and relationship touchpoint feedback data to obtain the feedback credibility corresponding to each content response feedback data; determining the content response feedback data whose feedback credibility meets preset conditions as credible feedback data; and identifying touchpoint mismatch types based on credible feedback data.

[0058] In practice, a feedback log table can be created for each target piece of content in the content publishing system. Log data can be collected in JSON or event stream format and saved to MySQL, ClickHouse, or Elasticsearch. Content touchpoint feedback data can directly come from exposure logs, click logs, dwell logs, and playback completion logs, recording fields such as whether it was exposed, whether it was clicked, dwell time, and whether playback completed. Interaction touchpoint feedback data can come from like tables, comment tables, repost tables, favorite tables, and secondary creation record tables. Relationship touchpoint feedback data can come from follow tables, repeat visit tables, community joining tables, and sharing and diffusion link tables. To facilitate unified processing later, the three types of feedback data can be mapped to structured feature records, and then associated with content identifiers, user identifiers, and timestamps for merging.

[0059] During the credibility assessment phase, the credibility of each feedback record can be calculated. A common approach is to extract assessment factors such as consistency of feedback source, stability of time distribution, completeness of behavioral links, and consistency with historical behavioral patterns. These factors are then input into a logistic regression model, random forest model, or lightweight scoring function, outputting a credibility value between 0 and 1. For example, if a piece of content experiences a large number of high-frequency reposts within a very short period but lacks normal dwell and access links, the credibility of that interaction feedback can be determined to be low. Conversely, if a user's exposure-access-dwell-interaction-follow link is complete, the corresponding feedback credibility is high. The system can retain data with a credibility higher than a preset threshold as credible feedback data, while marking data below the threshold as downweighted or isolated data to reduce interference from abnormal noise in subsequent identification. This processing ensures that the data entering subsequent touchpoint mismatch identification has higher usability, providing a more reliable input basis for determining whether the deviation in the target content occurs in content reach, user interaction, or relationship conversion stages.

[0060] In one specific embodiment, the identification of touchpoint mismatch types based on trusted feedback data includes: identifying whether the target content has an attraction mismatch type or a comprehension mismatch type based on content touchpoint feedback data in trusted feedback data; identifying whether the target content has an interaction mismatch type based on interaction touchpoint feedback data in trusted feedback data; and identifying whether the target content has a relationship conversion mismatch type based on relationship touchpoint feedback data in trusted feedback data. The attraction mismatch type is insufficient conversion of exposure to access for the target content; the comprehension mismatch type is insufficient dwell time or completion of the target content; the interaction mismatch type is insufficient interaction conversion of the target content; and the relationship conversion mismatch type is insufficient relationship conversion of the target content.

[0061] In practice, an analysis task can be created for each target piece of content in the content publishing system, using the content identifier as the primary key, to retrieve the data from the feedback database that has undergone credibility assessment. This feedback database can be stored in MySQL, PostgreSQL, ClickHouse, or Elasticsearch. Content touchpoint feedback data, interaction touchpoint feedback data, and relationship touchpoint feedback data can be stored in different detail tables, or they can be distinguished by event type using a unified event table. The system can run on a common Linux server, and the identification program can be implemented in Python, Java, or Go, executing on a scheduled basis (minutes or hours).

[0062] In one embodiment, the system first calculates the exposure, visits, average dwell time, and completion rate of the target content based on content touchpoint feedback data. Then, it determines whether the content exhibits an attraction mismatch or a comprehension mismatch based on preset rules or a simple classification model. Attraction mismatch primarily refers to situations where users do not effectively engage with the content after exposure; therefore, it can be determined based on the conversion rate between exposure and visits. For example, within the same demographic, similar time period, and similar content format, if the current target content's visit-to-conversion rate is significantly lower than a preset baseline, an attraction mismatch can be identified. Comprehension mismatch primarily refers to situations where users have entered the content but do not stay long enough or complete the content effectively; therefore, it can be determined based on average dwell time, effective view rate, complete reading rate, or playback completion rate. If these metrics consistently fall below the reference range for similar content, a comprehension mismatch can be identified.

[0063] Next, the system identifies interaction mismatch types based on the feedback data from interactive touchpoints. Specifically, it can count likes, comments, shares, favorites, and secondary creation triggers to form interaction behavior characteristics. If the exposure and access performance of the target content is normal, and the dwell or completion performance is also within an acceptable range, but the likes, comments, shares, favorites, or secondary creation behaviors are significantly low, it indicates that the content has been seen and partially consumed by users, but has not effectively stimulated further participation. In this case, an interaction mismatch type can be identified. In engineering implementation, this step can use threshold rules, such as an interaction rate lower than a certain percentage of the average for similar content, or it can use common binary classification models, such as logistic regression or lightweight gradient boosting tree models, to output whether it belongs to interaction mismatch based on the input interaction behavior characteristics.

[0064] Then, the system identifies relationship conversion mismatch types based on relationship touchpoint feedback data. This primarily focuses on whether the target content further generates attention, repeat visits, community joining, or dissemination. For example, it can track the number of new followers, repeat visit rate, number of visits to community pages, and the depth of subsequent dissemination links after a user consumes the target content. If the target content has generated some visits and interactions but fails to drive users to form sustained relationship connections, then a relationship conversion mismatch type can be identified. For instance, in an interest-based community scenario, a piece of content receives a high number of comments, but the follower conversion and repeat visit rate after comments are low, indicating that while the content sparked discussion, it failed to build a subsequent relationship.

[0065] To improve the stability of the identification process, this embodiment can execute the above identification process within a unified mismatch identification service. This service aggregates the three types of reliable feedback data corresponding to the same content identifier and outputs identification results for attraction mismatch, comprehension mismatch, interaction mismatch, and relationship transformation mismatch, respectively. The identification results can be saved using Boolean labels, enumerated labels, or status codes. For example, attraction mismatch can be denoted as Type 1, comprehension mismatch as Type 2, interaction mismatch as Type 3, and relationship transformation mismatch as Type 4, and written into an analysis result table for subsequent use by the hierarchical optimization module.

[0066] For example, in a video content scenario targeting digital interest groups, if the system detects that a certain piece of content has high exposure but a significantly low click-through rate, it can be identified as an attraction mismatch. If another piece of content has a normal click-through rate but a short average dwell time and low completion rate, it can be identified as an understanding mismatch. If a piece of content has good access and completion rates but low likes, comments, and shares, it can be identified as an interaction mismatch. If a piece of content receives many comments and favorites but few new followers and repeat visits, it can be identified as a relationship conversion mismatch.

[0067] In one embodiment, the step of determining and executing a hierarchical optimization processing strategy based on the touchpoint mismatch type includes: when the touchpoint mismatch type is identified as an attraction mismatch type, replacing or adjusting at least one of the title information, cover information, or opening introductory segment of the target content; when the touchpoint mismatch type is identified as a comprehension mismatch type, replacing or adjusting at least one of the content structure, information density, or expression template of the target content; when the touchpoint mismatch type is identified as an interaction mismatch type, replacing or adjusting at least one of the interactive guidance, topic linking method, and publication time of the target content; and when the touchpoint mismatch type is identified as a relationship conversion mismatch type, replacing or adjusting at least one of the attention guidance content, subsequent follow-up content, or circle entry information of the target content.

[0068] In a typical implementation, a touchpoint identification and tiered optimization service can be deployed on the backend of the content publishing system. This service can run on a Linux server environment, implemented using Python, Java, or Go, and connect to a log collection system, a content management system, and a publishing scheduling system. Content response feedback data can be stored as JSON event streams, Kafka messages, relational database tables, or Elasticsearch indexed documents. The system uses the content identifier as the primary key, aggregates reliable feedback data such as exposure, visits, dwell time, completion, likes, comments, reposts, favorites, follows, repeat visits, and joining of user communities by time window, and sends this data to the mismatch identification module.

[0069] Attraction mismatch is primarily used to identify situations where target content has been seen by users but has failed to effectively attract them to the content. In practice, the system can extract metrics such as exposure, visits, and exposure-to-visit conversion rate from reliable feedback data. If the exposure of a target piece of content reaches a preset range, but the visits are significantly low, or the exposure-to-visit conversion rate is consistently lower than the benchmark value for the same demographic, content format, and time period, then the target content can be identified as exhibiting attraction mismatch. For example, in short video scenarios, the system can compare the number of times the cover image is displayed with the number of times users click to enter the playback page; in text and image scenarios, it can compare the number of times the title card is displayed with the number of times the body page is visited. The identification process can be implemented using threshold comparison or by using lightweight models such as logistic regression and decision trees to output binary classification results. When the target content is identified as exhibiting attraction mismatch, the system enters the first type of hierarchical optimization processing strategy. Specifically, it can retrieve replacement candidates from a pre-configured title template library, cover material library, or introductory segment library, and replace or adjust at least one of the title information, cover information, or introductory segment. For example, if the target content is a short video reviewing digital products, the system can replace the original explanatory title with a question-based title, replace the cover image with a main visual that highlights the core selling points, or change the original long intro into a shorter, more direct opening segment.

[0070] Understanding mismatch types is primarily used to identify situations where users have entered content but failed to stay or complete content consumption. In practice, the system extracts metrics such as average dwell time, effective viewing ratio, exit position, completion rate, or reading completion rate from reliable feedback data and compares them with historical reference intervals for similar content. If access to a target content is normal, but the dwell time is significantly low, a large number of users drop off at fixed points, or the completion rate consistently falls below a preset threshold, then the target content can be identified as an understanding mismatch type. For example, in video scenarios, the system can record at what second a user exits playback; in text and image scenarios, it can record at what screen a user stops browsing. If a large number of users exit at similar positions, it indicates that there may be problems with the content structure, information density, or expression method corresponding to that position. When a target content is identified as an understanding mismatch type, the system enters the second type of layered optimization processing strategy. At this time, at least one of the content structure, information density, or expression template of the target content can be replaced or adjusted. Specifically, this can involve reorganizing paragraph order, reducing redundant information, adding key summaries, rewriting long narratives into list-style expressions, or replacing existing templates from continuous explanations to question-and-answer, comparison, or step-by-step templates. For example, for a fitness tutorial article, if users tend to drop off after the third paragraph, the system can shorten the introductory content, introduce the steps earlier, and add brief conclusions at key points to improve sustained reading ability.

[0071] Interaction mismatch is primarily used to identify situations where users have completed basic consumption but have not engaged in further interaction. In practice, the system can extract metrics such as like rate, comment rate, forwarding rate, collection rate, and secondary creation rate from reliable feedback data. If the access, dwell time, or completion rate of the target content is within the normal range, but the overall interaction behavior is significantly low, it indicates that the content has not effectively stimulated users to express their opinions, participate in discussions, or spread the word; this can be identified as an interaction mismatch. In a common implementation, the number of visits can be used as the interaction normalization base to calculate the proportion of likes, comments, forwards, and collections per unit of visit; if this proportion consistently falls below a preset benchmark, the interaction mismatch identification result is output. When the target content is identified as an interaction mismatch, the system enters the third type of layered optimization processing strategy. At this point, at least one of the following can be replaced or adjusted: the interactive prompts, the topic linking method, and the publication time of the target content. For example, interactive prompts such as "Which option do you prefer?" or "Have you encountered similar problems?" can be added at the end of the content; general tags can be replaced with topic-based links that are more relevant to the discussion within the community; and the posting time can be adjusted from low-activity hours to high-activity hours in the evening, depending on the community's activity time. For instance, in a beauty interest community, if a piece of content has a high completion rate but a low comment rate, the system can automatically add an interactive prompt asking whether you prefer a sheer or long-lasting finish, and adjust the posting time from weekday mornings to the evening's active window.

[0072] Relationship conversion mismatch is primarily used to identify situations where content has generated some visits or interactions, but hasn't further solidified into engagement, repeat visits, community joining, or wider dissemination. In practice, the system can extract metrics such as engagement conversion rate, repeat visit rate, community joining rate, post-shared link depth, or dissemination level from reliable feedback data. If the target content has generated some views or discussion, but no sustained connection has formed, it can be classified as a relationship conversion mismatch. For example, in an interest community, a piece of content receives many comments, but very few new followers, and the commenting users don't further enter the community page. This indicates that while the content can generate short-term interaction, it hasn't formed a subsequent relationship. When the target content is identified as a relationship conversion mismatch, the system enters the fourth type of layered optimization processing strategy. At this point, at least one of the target content's engagement guidance content, subsequent follow-up content, or community entry information can be replaced or adjusted. Specifically, the system can add prompts at the end of the target content, such as "Continue to view content on the same topic," "Enter the topic page," or "Join the discussion group." It can also add links to subsequent content directly related to the current content. Furthermore, it can move the group's entry point from the bottom of the page to a more easily accessible location. For example, in a photography interest group, a post on night scene shooting techniques might receive many saves but have low conversion rates. The system can automatically add links at the end of the post, such as "View more night scene parameter templates" or "Enter the photography techniques group," and link to a subsequent piece of content to improve relationship building.

[0073] Using the above method, the system first identifies mismatches in the target content based on reliable feedback data, categorizing the problems into four types: attraction mismatch, comprehension mismatch, interaction mismatch, and relationship transformation mismatch. Then, it executes corresponding layered optimization strategies based on the mismatch type. In other words, the system doesn't uniformly change the title or publication time for every instance of poor feedback. Instead, it precisely targets adjustments to different aspects such as the title / cover / opening, content structure / information density / expression template, interactive prompts / topic linking / publication time, attention guidance / follow-up transitions / entry points into specific communities, based on the specific stage at which the mismatch occurs.

[0074] In this embodiment, the step of determining and executing a hierarchical optimization processing strategy based on the touch point mismatch type further includes: after the candidate set is optimized and sorted, retaining at least one candidate creation object that has not been determined as the target content as a backup candidate creation object; when the touch point mismatch type meets the candidate reordering condition, determining the feedback correction parameter based on the content response feedback data corresponding to the target content, and calculating the replacement score corresponding to each backup candidate creation object based on the matching degree between each backup candidate creation object and the feature vector of the circle demand, the candidate relationship score corresponding to each backup candidate creation object, and the feedback correction parameter; reordering the backup candidate creation objects according to the replacement score, and determining the backup candidate creation object ranked first as the updated target content.

[0075] After the preliminary candidate set is sorted, the system does not only retain the top-ranked candidate, but also saves at least one unselected candidate as a backup candidate. In practice, the sorting results can be written to a candidate content table, where the currently selected object is marked as the target content, and the remaining objects are marked as backup candidates. The database can be MySQL, PostgreSQL, or MongoDB; for faster retrieval, the key features of the backup candidates can also be cached in Redis. Each backup candidate can store at least the following information: object identifier, corresponding creative direction, content summary, candidate relationship score, matching degree with the feature vector of the target audience's needs, publication time adaptation tag, and associated material path. This way, when a re-sorting occurs later, there is no need to regenerate content objects; instead, the system can directly select from the retained backup candidates. For example, under the creative direction of summer sunscreen selection, the system initially generated three candidate objects: A is an ingredient comparison graphic, B is a scene evaluation short video, and C is a price list card. The system initially selects B as the target content for publication, while A and C are retained as backup candidates.

[0076] In this embodiment, candidate reordering is not executed immediately after any feedback occurs, but rather triggered when the touchpoint mismatch type meets the candidate reordering conditions. These candidate reordering conditions can be implemented using rule-based judgment. For example, they can be pre-configured in the strategy table: when an attraction mismatch type occurs consecutively more than a preset number of times, and local adjustments to the title, cover, etc., do not improve the situation, candidate reordering is triggered; when an understanding mismatch type occurs consecutively, indicating a significant deviation between the current expression and the needs of the target audience, candidate reordering is triggered; when interaction mismatch or relationship conversion mismatch reaches a preset threshold, indicating that the current target content, while edible, is not conducive to further participation or retention, candidate reordering is triggered. The rule engine can read the touchpoint mismatch identification results and historical optimization record table to determine whether the candidate reordering conditions are met. If met, a reordering instruction is sent to the candidate reordering service.

[0077] Once the candidate reordering conditions are met, the system first determines the feedback correction parameters based on the content response feedback data corresponding to the target content. These feedback correction parameters can be understood as parameters reflecting the direction and degree of deviation exposed in the actual release of the current target content. They do not directly indicate which backup candidate is best, but rather indicate in which type of response stage the current target content encountered a problem, and how significant the problem is. In specific implementation, the following information can be extracted from the content response feedback data: deviation in the attraction stage (the decrease in exposure-to-access conversion rate); deviation in the understanding stage (the decrease in average dwell time and completion rate); deviation in the interaction stage (the degree to which comment rate, forwarding rate, and collection rate deviate from the baseline); and deviation in the relationship conversion stage (the degree to which attention rate, repeat visit rate, and community joining rate deviate from the baseline). Then, this deviation information can be encoded into feedback correction parameters. For example, the feedback correction parameters can be organized into a multi-dimensional vector, with different positions in the vector corresponding to the four types of bias strengths: attraction, understanding, interaction, and relationship transformation. Alternatively, they can be organized into structured fields such as loss_attract, loss_understand, loss_interact, and loss_relation. These parameters can be stored in the intermediate results table of candidate rearrangement for subsequent use by the replacement score calculation module.

[0078] After receiving the feedback correction parameters, the system needs to re-evaluate each backup candidate creation object and calculate its replacement score. In this embodiment, the input of the replacement score includes at least three types of information: 1. The matching degree between each backup candidate creation object and the feature vector of the circle's needs. This value reflects whether a backup candidate still meets the current circle's needs. 2. The candidate relationship score corresponding to each backup candidate creation object. This value reflects the overall relationship position of the backup candidate in the original candidate set, such as whether it is more stable or more distinctive compared to other candidates. 3. Feedback correction parameters. These parameters reflect the reason and degree of failure of the current target content, and are used to adjust the priority of backup candidates. In a common implementation, the replacement score can be calculated in a weighted manner. For example, a basic matching item, a relationship stability item, and a feedback correction item can be calculated separately for each backup candidate creation object. The role of the feedback correction item is: if the current target content performs poorly in a certain stage, the score of the backup candidate that is more advantageous in that stage will be increased. For example, if the current target content exhibits an attraction mismatch, the system can increase the correction weight for alternative candidates with more direct titles, more prominent cover images, and stronger opening hooks; if the current target content exhibits a comprehension mismatch, the system can increase the correction weight for alternative candidates with clearer structures, more reasonable information density, and simpler expression templates. In the specific implementation process, a rule-based weighting method or a lightweight scoring model, such as a linear scoring model or a gradient boosting tree model, can be used. The inputs are the alternative candidate matching degree, candidate relationship score, and feedback correction parameters; the output is the replacement score corresponding to each alternative candidate creation object. The replacement score results can be written to the candidate reordering result table.

[0079] After obtaining replacement scores for all alternative candidate content, the system reorders them based on these scores. This sorting process can be implemented using conventional sorting algorithms, such as quicksort, heapsort, or built-in sorting functions in the programming language. Once sorted, the candidate content that ranks highest and whose replacement scores meet preset criteria is selected as the updated target content. These preset criteria primarily prevent forced replacement when none of the alternative candidates are suitable. For example, the replacement score could be greater than a preset threshold, higher than the reference score for the current target content, or the replacement score could be at least as far apart from the second-ranked candidate. If the criteria are met, replacement is executed; otherwise, the current target content is maintained, and other optimization strategies are implemented.

[0080] For example, in a light-meal, weight-loss interest group, the system generated three candidate content items around the theme of high-protein breakfast combinations: A is a text and image list type of content, B is a short video demonstration type of content, and C is a Q&A type of content. Initially, the system selected B as the target content to publish based on a comprehensive score, while A and C were saved as backup candidates. After publication, the system identified that although B had normal traffic, its comments and favorites were significantly low, indicating an interaction mismatch. This mismatch persisted even after two local optimizations, thus meeting the candidate re-ranking criteria. The system then extracted interaction deviations from the content response feedback data, generating feedback correction parameters, indicating a significant problem with the current target content in terms of insufficient interaction guidance. Next, the system calculated the replacement scores for A and C respectively. C, being a Q&A type of content, was better suited for guiding comments and discussions; therefore, after considering the feedback correction parameters, its replacement score was higher than A's. Thus, the system re-ranked the content based on the replacement scores, identifying C as the updated target content and pushing it to the publishing scheduling module for replacement publishing.

[0081] In this embodiment, instead of simply making repeated local modifications to the current target content, it quickly rearranges and replaces content when an incompatibility is detected, using pre-reserved backup candidate creation objects. This reduces redundant generation and manual rework, while improving the system's content response efficiency in different touchpoint mismatch scenarios.

[0082] In one embodiment, the step of determining and executing a hierarchical optimization processing strategy based on the touchpoint mismatch type further includes: when at least two target contents determined by the same target creation direction are consecutively identified as the same touchpoint mismatch type, determining that there is a demand comprehension bias in the current circle demand feature vector; extracting mismatch correction features based on the reliable feedback data corresponding to the same touchpoint mismatch type, mapping the mismatch correction features to the corresponding feature dimension of the circle demand feature vector, adjusting the weight of the corresponding feature dimension according to a preset correction coefficient, and obtaining an updated circle demand feature vector; and re-determining the target creation direction, regenerating candidate creation objects, and re-optimizing and sorting the candidate set based on the updated circle demand feature vector.

[0083] In practical implementation, a demand correction service can be deployed on the backend of the content management system. This service can run on a common Linux server environment, implemented using Python or Java, and connected to the feature library, feedback database, direction generation module, and candidate ranking module. The feature vectors for different user groups can be stored as fixed-length floating-point vectors, such as 128-dimensional or 256-dimensional arrays, in a vector database, Redis cache, or PostgreSQL feature table. Trusted feedback data can be stored as structured records, such as JSON documents or relational table records, based on content identifiers, creation direction identifiers, and time windows.

[0084] First, the system needs to determine when a demand comprehension mismatch can be identified. In this embodiment, the trigger condition is that at least two target content pieces determined by the same creative direction are consecutively identified as the same touchpoint mismatch type. That is, the system does not immediately correct the demand vector upon seeing a single mismatch, but rather requires that at least two published target content pieces under the same creative direction fall into the same mismatch type after reliable feedback analysis. For example, two pieces of content were published successively on the topic of digital product reviews, and both were identified as a comprehension mismatch type; or two pieces of content were published successively on the topic of summer skincare tips, and both were identified as a relationship transformation mismatch type. At this point, the system considers the problem no longer to be a localized manifestation of a single target content piece, but rather a deviation in the front-end's demand judgment for that creative direction, thus determining that there is a demand comprehension mismatch in the current user segment's demand feature vector.

[0085] After identifying a misinterpretation of requirements, the system extracts mismatch correction features from the credible feedback data corresponding to the same type of mismatch at the touchpoint. These mismatch correction features can be understood as a set of structured features that reflect the reasons for this type of mismatch. For example, when attraction mismatch occurs repeatedly, information related to exposure, visits, cover clicks, and title clicks can be extracted from the credible feedback data; when comprehension mismatch occurs repeatedly, information related to dwell time, content exit position, and completion rate decline range can be extracted; when interaction mismatch occurs repeatedly, information related to low comment rates, forwarding rates, and collection rates can be extracted; when relationship conversion mismatch occurs repeatedly, information related to low attention rates, repeat visit rates, and community joining rates can be extracted. These mismatch correction features can be organized into a set of fields, for example, using a record format of {mismatch type + key indicator + deviation degree + time window}, or encoded into a mapping vector of the same or lower dimension as the requirement vector.

[0086] Subsequently, the system maps the mismatch correction features to the corresponding feature dimensions of the circle demand feature vector. These corresponding feature dimensions can be determined using a pre-established feature dimension mapping table. This mapping table can be formed from offline training results, manually labeled configurations, or a rule base, and stored in a configuration file or database. For example, if some dimensions in the circle demand feature vector are used to represent attraction preferences, then the correction features corresponding to attraction mismatch are mapped to these dimensions; if some dimensions are used to represent understanding preferences or structural acceptability, then the correction features corresponding to understanding mismatch are mapped to these dimensions; if some dimensions are used to represent interaction willingness or relationship retention tendency, then interaction mismatch and relationship transformation mismatch are mapped to their respective dimensions. In this way, the system does not indiscriminately modify the entire circle demand feature vector, but only adjusts the local dimensions directly related to the current mismatch type.

[0087] After mapping, the system adjusts the weights of the corresponding feature dimensions according to preset correction coefficients, resulting in an updated feature vector for the circle's needs. These preset correction coefficients can be set based on the type of mismatch, the degree of deviation, and the number of consecutive mismatches. For example, a smaller correction coefficient is needed for two consecutive minor comprehension mismatches; a larger correction coefficient is needed for three consecutive significant relationship transformation mismatches. In engineering implementation, the basic correction coefficients for different mismatch types can be stored using a rule table, or dynamically distributed using a lightweight parameter configuration service. Adjustments can be made by increasing the weights of certain dimensions, decreasing the weights of others, or performing both increases and decreases simultaneously. After adjustment, the updated feature vector for the circle's needs better reflects the true demand status of the current circle at the corresponding mismatch stage.

[0088] For example, in a content system targeting a fitness enthusiast community, the system published two consecutive pieces of content focusing on home workout techniques. Both were identified as exhibiting a comprehension mismatch. Analysis revealed that a large number of users churned midway through the content. Confidential feedback indicated that the current community prefers short, concise, and minimally detailed presentations, while the original demand vector placed higher weights on dimensions related to in-depth explanations and lengthy descriptions. In this case, the system extracted mismatch correction features from the credible feedback data, such as concentrated mid-section churn, low completion rates, and increased preference for short-step content. These features were mapped to the structural acceptability and information density preference dimensions in the demand vector. Then, the weights of dimensions related to lengthy explanations were reduced and the weights of dimensions related to step-by-step presentations were increased using preset correction coefficients, resulting in an updated demand feature vector for the community.

[0089] After obtaining the updated feature vector of circle demand, the system continues with the following steps: First, it redefines the target creation direction based on the updated feature vector; then, it regenerates candidate creation objects around the newly determined target creation direction; finally, it re-sorts and optimizes the candidate set. In this way, the subsequently generated target content no longer uses the old demand understanding results, but is regenerated based on the revised feature vector of circle demand, thereby reducing the probability of repeated mismatches at the same touchpoints under the same creative direction.

[0090] This invention also provides a content creation collaborative management system that responds to the needs of interest groups, such as... Figure 2As shown, the system includes: a data acquisition module for acquiring multi-source heterogeneous data corresponding to the target interest circle; a data encoding module for performing multi-view, multi-modal encoding on the multi-source heterogeneous data to obtain circle demand feature vectors and content feature vectors; a creation direction determination module for determining at least one target creation direction based on the circle demand feature vectors; a candidate creation object generation module for generating at least two candidate creation objects for each target creation direction; a candidate set modeling and sorting module for modeling the interrelationships and performing optimal sorting on the candidate set formed by the candidate creation objects to determine the target content; a content publishing and feedback acquisition module for publishing the target content and acquiring content response feedback data corresponding to the target content; a touchpoint mismatch identification module for identifying touchpoint mismatch types based on the content response feedback data; and a hierarchical optimization processing module for determining and executing a hierarchical optimization processing strategy according to the touchpoint mismatch type to optimize at least one of the target content, circle demand feature vectors, target creation directions, candidate creation objects, or candidate set optimal sorting.

[0091] Specifically, the system can be deployed in a server cluster, which includes at least a data acquisition server, a computing server, a business server, a storage server, and a publishing server. The servers are connected via a local area network (LAN) or a cloud private network and can communicate using HTTP / HTTPS, gRPC, message queues, or database connections. The data acquisition server can be deployed in the platform's data access layer, and a data acquisition module can be embedded in this server to acquire text, image, audio, and behavioral log data from content platforms, community platforms, log systems, and object storage. The acquired data can be in JSON, CSV, Parquet, or message stream event formats and transmitted to subsequent processing nodes via Kafka or other message queues.

[0092] The data encoding module, creation direction determination module, candidate creation object generation module, and candidate set modeling and ranking module can be deployed on a computing server or a business server. The data encoding module is preferably deployed on a computing server with a GPU or high-performance CPU to perform joint encoding of text, images, audio, and behavioral features. The creation direction determination module, candidate creation object generation module, and candidate set modeling and ranking module can be deployed on the business server and connected to the data encoding module via internal service calls. The business server is connected to a storage server, which can be configured with a relational database, document database, vector database, and object storage to store original logs, candidate object records, feature vectors, ranking results, template resources, and multimedia materials, respectively.

[0093] The content publishing and feedback collection module can be embedded in the publishing server or content management server. It distributes target content to the corresponding content platform and collects feedback data from the platform, including exposure, visits, dwell time, completion, likes, comments, reposts, favorites, follows, and repeat visits. This module collects feedback through platform open interfaces, event tracking SDKs, or server-side log feedback and writes the feedback data to a log database or message queue. The touchpoint mismatch identification module and the hierarchical optimization processing module can be deployed on the business server. Both can process content response feedback data, candidate object records, and feature vector records. The touchpoint mismatch identification module outputs the mismatch type result, and the hierarchical optimization processing module adjusts content parameters, feature vectors, creation direction, candidate objects, or sorting results based on this result.

[0094] In a common deployment, the data acquisition server resides in the data access layer, the computing server in the feature processing layer, the business server in the decision control layer, the publishing server in the content output layer, and the storage server in the data persistence layer. These layers communicate via an intranet, and data can be transmitted between layers using structured or vector records. This ensures that the entire system has a clearly defined hardware structure while maintaining consistency with the databases, file formats, service modules, and processing flows described in the aforementioned embodiments.

[0095] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention; any modifications, equivalent substitutions, improvements or combinations made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A collaborative management method for content creation that responds to the needs of interest groups, characterized in that, include: Collect multi-source heterogeneous data corresponding to the target interest groups; Multi-source heterogeneous data is encoded from multiple perspectives and in multiple modes to obtain the demand feature vector and content feature vector for each circle. Determine at least one target creative direction based on the feature vector of circle demand; Generate at least two candidate creative objects for each target creative direction; The candidate set of candidate creative objects is modeled for mutual relationships and optimized and sorted to determine the target content; Publish the target content and collect the corresponding content response feedback data; Identify touchpoint mismatch types based on content response feedback data; Based on the touchpoint mismatch type, a hierarchical optimization processing strategy is determined and implemented to optimize at least one of the following: target content, circle demand feature vector, target creative direction, candidate creative object or candidate set optimization ranking.

2. The content creation collaborative management method for responding to the needs of interest groups as described in claim 1, characterized in that, The process of obtaining the circle demand feature vector and content feature vector includes: dividing multi-source heterogeneous data into perspectives based on content semantic perspective, dissemination interaction perspective, and temporal context perspective; extracting and encoding features for text data, visual data, audio data, and behavioral data under each perspective to obtain corresponding single-perspective single-modal features; aligning and fusing the single-perspective single-modal features to obtain joint features; and aggregating the circle-side features and content-side features based on the joint features to obtain the circle demand feature vector and content feature vector.

3. The content creation collaborative management method for responding to the needs of interest groups as described in claim 1, characterized in that, The step of determining at least one target creative direction based on the feature vector of circle demand includes: performing theme clustering on the feature vector of circle demand to obtain several demand themes; combining timestamp data from multi-source heterogeneous data to perform time-series evolution analysis on each demand theme to obtain the demand attention corresponding to each demand theme; statistically analyzing the existing content coverage corresponding to each demand theme based on the content feature vector; calculating the difference between the demand attention of each demand theme and the existing content coverage; and determining the demand theme whose difference value meets the preset conditions as the target creative direction.

4. The content creation collaborative management method for responding to the needs of interest groups as described in claim 1, characterized in that, The process of determining the target content includes: extracting candidate features for each candidate creative object, wherein the candidate features include at least semantic features, content form features, adaptation features, and risk features; determining similarity association parameters between any two candidate creative objects based on the candidate features and constructing a candidate relationship matrix; aggregating the candidate relationship matrix to obtain the candidate relationship score corresponding to each candidate creative object; calculating the comprehensive score corresponding to each candidate creative object based on the matching degree between each candidate creative object and the feature vector of the circle demand and the candidate relationship score, according to a preset weight; optimizing and ranking the candidate set according to the comprehensive score, and determining the candidate creative object ranked first as the target content.

5. The content creation collaborative management method for responding to the needs of interest groups as described in claim 1, characterized in that, The content response feedback data includes content touchpoint feedback data, interaction touchpoint feedback data, and relationship touchpoint feedback data. The content touchpoint feedback data represents the user's exposure, access, dwell time, and completion status of the target content. The interaction touchpoint feedback data represents the user's actions of liking, commenting, forwarding, collecting, or creating derivative works based on the target content. The relationship touchpoint feedback data represents the user's formation of relationships such as following, revisiting, joining circles, or spreading based on the target content. The method of identifying touchpoint mismatch types based on content response feedback data includes: evaluating the credibility of content touchpoint feedback data, interaction touchpoint feedback data, and relationship touchpoint feedback data to obtain the feedback credibility of each content response feedback data, determining the content response feedback data whose feedback credibility meets preset conditions as credible feedback data, and identifying touchpoint mismatch types based on credible feedback data.

6. The content creation collaborative management method for responding to the needs of interest groups as described in claim 5, characterized in that, The identification of touchpoint mismatch types based on trusted feedback data includes: identifying whether the target content has an attraction mismatch type or a comprehension mismatch type based on the content touchpoint feedback data in trusted feedback data; identifying whether the target content has an interaction mismatch type based on the interaction touchpoint feedback data in trusted feedback data; and identifying whether the target content has a relationship transformation mismatch type based on the relationship touchpoint feedback data in trusted feedback data. The attraction mismatch type is insufficient conversion of exposure to visits for the target content; the understanding mismatch type is insufficient dwell time or completion of the target content; the interaction mismatch type is insufficient interaction conversion of the target content; and the relationship conversion mismatch type is insufficient relationship conversion of the target content.

7. The content creation collaborative management method for responding to the needs of interest groups as described in claim 6, characterized in that, The layered optimization processing strategy based on touchpoint mismatch type includes: when the touchpoint mismatch type is identified as an attraction mismatch type, replacing or adjusting at least one of the title information, cover information, or opening introductory segment of the target content; when the touchpoint mismatch type is identified as a comprehension mismatch type, replacing or adjusting at least one of the content structure, information density, or expression template of the target content; when the touchpoint mismatch type is identified as an interaction mismatch type, replacing or adjusting at least one of the interactive introductory text, topic linking method, and publication time of the target content; and when the touchpoint mismatch type is identified as a relationship conversion mismatch type, replacing or adjusting at least one of the attention guidance content, subsequent follow-up content, or circle entry information of the target content.

8. The content creation collaborative management method for responding to the needs of interest groups as described in claim 6, characterized in that, The step of determining and executing a hierarchical optimization processing strategy based on touch point mismatch type further includes: after the candidate set is optimized and sorted, retaining at least one candidate creation object that has not been determined as the target content as a backup candidate creation object; when the touch point mismatch type meets the candidate reordering condition, determining the feedback correction parameter based on the content response feedback data corresponding to the target content, and calculating the replacement score corresponding to each backup candidate creation object based on the matching degree between each backup candidate creation object and the circle demand feature vector, the candidate relationship score corresponding to each backup candidate creation object, and the feedback correction parameter; reordering the backup candidate creation objects according to the replacement score, and determining the backup candidate creation object ranked first as the updated target content.

9. The content creation collaborative management method for responding to the needs of interest groups as described in claim 6, characterized in that, The step of determining and executing a hierarchical optimization processing strategy based on touchpoint mismatch type further includes: when at least two target contents determined by the same target creation direction are consecutively identified as the same touchpoint mismatch type, determining that there is a demand comprehension bias in the current circle demand feature vector; extracting mismatch correction features based on the reliable feedback data corresponding to the same touchpoint mismatch type, mapping the mismatch correction features to the corresponding feature dimension of the circle demand feature vector, adjusting the weight of the corresponding feature dimension according to a preset correction coefficient, and obtaining an updated circle demand feature vector; and re-determining the target creation direction, regenerating candidate creation objects, and re-optimizing and sorting the candidate set based on the updated circle demand feature vector.

10. A content creation collaborative management system that responds to the needs of interest groups, characterized in that, include: The data acquisition module is used to collect multi-source heterogeneous data corresponding to the target interest groups; The data encoding module is used to encode multi-source heterogeneous data from multiple perspectives and in multiple modes to obtain the circle demand feature vector and content feature vector. The creative direction determination module is used to determine at least one target creative direction based on the feature vector of circle demand. The candidate creation object generation module is used to generate at least two candidate creation objects for each target creation direction. The candidate set modeling and sorting module is used to model the relationships between candidate creative objects and optimize their sorting to determine the target content. The content publishing and feedback collection module is used to publish target content and collect content response feedback data corresponding to the target content. The touch mismatch identification module is used to identify the type of touch mismatch based on content response feedback data; The hierarchical optimization processing module is used to determine and execute hierarchical optimization processing strategies based on the touch point mismatch type, so as to optimize at least one of the following: target content, circle demand feature vector, target creative direction, candidate creative object or candidate set optimization ranking.