Data processing method, apparatus and electronic device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU NETZHIYI INNOVATION TECH CO LTD
- Filing Date
- 2026-05-29
- Publication Date
- 2026-08-07
AI Technical Summary
本公开提供的一种数据处理方法、装置和电子设备,首先响应于目标应用程序接收到用户发布内容,将用户发布内容输入至预先训练完成的目标模型中,得到用户发布内容对应的结构化数据;其中,结构化数据至少包括:用户发布内容对应的互动导向类型和多维特征数据;目标模型基于预设的教师模型输出的数据进行监督训练得到,目标模型继承教师模型的行为特征;然后基于用户发布内容对应的结构化数据,在目标应用程序中对用户发布内容进行推荐。该方式可以实现对目标应用程序中各个互动导向类型的用户发布内容的高效挖掘与精准推荐;同时,该方式通过轻量级的目标模型识别用户发布内容的结构化数据,提高了数据识别效率和数据推荐效率。
Smart Images

Figure CN122527409A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a data processing method, apparatus, and electronic device. Background Technology
[0002] In related technologies, content recommendation and mining technologies in applications mainly include the following two methods: one is to rely on operators to manually screen, classify and tag the posts in the application, and then push the content that is determined to be of high value to the target users; the other is to recommend content based on large models. However, large models have high computational overhead and cannot be compatible with the real-time processing needs of massive amounts of newly added content. Moreover, manual recommendation requires a lot of human resources and has low efficiency. Summary of the Invention
[0003] In view of this, the purpose of this disclosure is to provide a data processing method, apparatus and electronic device to balance the relationship between labor costs, processing speed and recommendation effectiveness in the data recommendation process.
[0004] In a first aspect, embodiments of this disclosure provide a data processing method applied to an electronic device, wherein the electronic device has a target application configured thereon. The method includes: in response to the target application receiving user-posted content, inputting the user-posted content into a pre-trained target model to obtain structured data corresponding to the user-posted content; wherein the structured data includes at least: interactive guidance type and multi-dimensional feature data corresponding to the user-posted content; the target model is obtained through supervised training based on data output by a preset teacher model, and the target model inherits the behavioral features of the teacher model; and based on the structured data corresponding to the user-posted content, recommending the user-posted content in the target application.
[0005] Secondly, this disclosure also provides a data processing apparatus, which is disposed in an electronic device, and the electronic device is provided with a target application; the apparatus includes: a data output module, used to respond to the target application receiving user-posted content, inputting the user-posted content into a pre-trained target model to obtain structured data corresponding to the user-posted content; wherein, the structured data includes at least: the interaction guidance type and multi-dimensional feature data corresponding to the user-posted content; the target model is obtained by supervised training based on the data output by a preset teacher model, and the target model inherits the behavioral features of the teacher model; and a content recommendation module, used to recommend user-posted content in the target application based on the structured data corresponding to the user-posted content.
[0006] Thirdly, this disclosure provides an electronic device including a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the above-described data processing method.
[0007] Fourthly, this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when invoked and executed by a processor, cause the processor to implement the aforementioned data processing method.
[0008] The embodiments disclosed herein bring the following beneficial effects: This disclosure provides a data processing method, apparatus, and electronic device. First, in response to a target application receiving user-posted content, the user-posted content is input into a pre-trained target model to obtain structured data corresponding to the user-posted content. This structured data includes at least: the interaction-oriented type and multi-dimensional feature data corresponding to the user-posted content. The target model is trained under supervised conditions based on data output by a pre-set teacher model, inheriting the behavioral features of the teacher model. Then, based on the structured data corresponding to the user-posted content, recommendations are made within the target application. This method enables efficient mining and accurate recommendation of user-posted content of various interaction-oriented types within the target application. Furthermore, this method improves data recognition and recommendation efficiency by using a lightweight target model to identify the structured data of user-posted content.
[0009] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objects and other advantages of this disclosure are realized and obtained through the structures particularly pointed out in the description, claims and drawings.
[0010] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0012] Figure 1 A flowchart of a data processing method provided in an embodiment of this disclosure; Figure 2This is a schematic diagram of the structure of a data processing apparatus provided in an embodiment of the present disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0014] Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely to illustrate selected embodiments of the disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0015] In related technologies, content recommendation and mining technologies for community applications mainly fall into two categories, both of which have significant technical defects and application limitations: The first type is the content discovery and recommendation model dominated by manual operations. This model relies on professional operators to manually screen, categorize, and tag the content posted within the application, and then push the content deemed high-value to the target users. The core problem with this model is that manual operation is extremely costly, and the limited human resources mean that the operation work often only covers top users or popular content, while a large amount of high-quality, potentially valuable content posted by ordinary users is missed. At the same time, facing the massive amount of new content added to the application every day, manual processing efficiency is extremely low, completely unable to meet the needs of real-time recommendations, especially in the initial stage of the application, where it is difficult to quickly discover highly interactive content to attract and retain users.
[0016] The second category is purely algorithm-driven recommendation models, which can be further subdivided into traditional machine learning models and large-scale model applications. Traditional machine learning models are mostly built based on users' historical behavior data (such as clicks, favorites, and comments). Their shortcomings lie in their weak understanding of the semantic and visual features of the content itself. They can only judge quality based on posterior recommendation metrics and cannot judge the interactive potential of the content from its essence. This easily leads to the "information cocoon" problem, and the model has poor generalization ability, making it insufficiently adaptable to the diverse types of content added to the application. Large-scale models have a certain content understanding ability, but because they are trained on massive amounts of data and general tasks, their accuracy in judging the interactive orientation of content is relatively low. In addition, some solutions that use large-scale models for real-time inference are incompatible with the real-time processing needs of massive amounts of newly added content due to the large number of parameters and high computational cost of large-scale models. This presents a contradiction between processing speed and recommendation effect.
[0017] To address the aforementioned issues, this disclosure provides a data processing method, apparatus, and electronic device, which can be applied to content recommendation and content mining scenarios within applications.
[0018] To facilitate understanding of this disclosure, a data processing method provided by an embodiment of this disclosure will first be described in detail. This method is applied to an electronic device that has a target application program configured thereon; such as Figure 1 As shown, the method includes the following specific steps: Step S102: In response to the target application receiving user-posted content, the user-posted content is input into the pre-trained target model to obtain structured data corresponding to the user-posted content; wherein, the structured data includes at least: the interaction guidance type and multi-dimensional feature data corresponding to the user-posted content; the target model is obtained by supervised training based on the data output by the preset teacher model, and the target model inherits the behavioral features of the teacher model.
[0019] The target application mentioned above can be any social application. Users can log in to the target application by registering an account. After logging in, users can publish any content on the target application. This content is user-published content, which can be text content, image content, or multimodal content such as text and image content.
[0020] In practical implementation, the trained target model is deployed in the content processing pipeline of the target application to achieve real-time recognition of newly added user-posted content. When the target application receives newly added user-posted content, the system automatically extracts the content and inputs it into the pre-trained target model. This allows the target model to output structured data corresponding to the user-posted content within a short time. This structured data includes, but is not limited to, the interaction guidance type corresponding to the user-posted content and multi-dimensional feature data. The multi-dimensional feature data is used to indicate the basis for determining the user-posted content as the corresponding interaction guidance type.
[0021] In an optional embodiment, the multidimensional feature data includes the following types of feature data: text feature data, visual feature data, and interactive feedback feature data.
[0022] The aforementioned text features may include, but are not limited to, title keywords, body structure, content purpose, and text sentiment. Specifically, title keywords may include terms such as "tutorial," "help," "recommendation," and "controversy"; body structure may include step-by-step instructions, open-ended questions, and statements of viewpoints; and content purpose may include delivering practical information, expressing emotions, and guiding the discussion.
[0023] The aforementioned visual features may include, but are not limited to: image type, the correlation between image and text, and the attractiveness index of visual elements. Image types include infographics, everyday photos, product images, and controversial images; attractiveness indexes include color saturation and subject sharpness.
[0024] The aforementioned interactive feedback characteristics may include, but are not limited to: historical interactive behavior data and the temporal distribution characteristics of interactive behavior. Historical interactive behavior data includes the number of favorites, comments, comment keywords, and reposts; the temporal distribution characteristics of interactive behavior include whether interactions occur in a concentrated period or whether there is a long period of inactivity.
[0025] The target model described above is trained based on the teacher model, meaning that the knowledge from the teacher model is transferred to the smaller target model to improve model performance and generalization ability. Specifically, the teacher model is usually a complex and large model, while the target model is a lightweight neural network model. Through knowledge distillation, the behavioral features of the teacher model can be transferred to the target model, thus giving the target model the same data processing capabilities as the teacher model.
[0026] For example, a powerful and knowledgeable visual-language model can be used as the teacher model. This model is an AI capable of simultaneously processing image visual information and textual language information, possessing strong cross-modal understanding and analysis capabilities. A large amount of content to be labeled can be input into the teacher model, requiring it to provide detailed annotations based on preset requirements. This annotation not only needs to identify the multi-dimensional feature data corresponding to the content but also the corresponding interaction guidance type. Based on the annotation information output by the teacher model, a large and high-quality training database can be established. Then, a lightweight target model can be trained using this database. The training objective of the target model is: given any user-posted content, the target model can output the corresponding interaction guidance type and multi-dimensional feature data.
[0027] After training, the target model can quickly generate the corresponding interactive guidance type and multi-dimensional feature data based on the input user-posted content during real-time inference. This target model perfectly inherits the behavioral characteristics of the teacher model, but with significantly improved running speed, fully meeting the needs of real-time interaction.
[0028] In an optional embodiment, the interaction guidance type corresponding to the user-posted content can be a high-interaction type or a no-interaction type, etc. The high-interaction type can include collection-oriented types and comment-oriented types.
[0029] The aforementioned "collection-oriented" types refer to content types whose core purpose is to provide practical information (such as tutorials, guides, tool recommendations, etc.), which can encourage users to collect them for easy access later. The aforementioned "comment-oriented" types refer to content types whose core purpose is to raise open-ended questions, share controversial viewpoints, or evoke emotional resonance, which can stimulate user participation in comments and discussions.
[0030] Step S104: Based on the structured data corresponding to the user-posted content, recommend the user-posted content in the target application.
[0031] In practice, based on the structured data corresponding to the user-published content, the user-published content with a high degree of matching with the target user or target scenario can be determined, and the user-published content can be sent to the target account corresponding to the target user or target scenario. The target account is also the registration account of the target application.
[0032] The data processing method provided in this disclosure enables efficient mining and accurate recommendation of user-posted content of various interactive types in a target application. Simultaneously, this method improves data identification and recommendation efficiency by identifying structured data of user-posted content through a lightweight target model.
[0033] The following examples describe the training method for the target model.
[0034] Specifically, the target model described above is trained through the following steps 10-14: Step 10: Obtain the first sample set; wherein the first sample set includes: multiple historical published content and the interactive guidance type corresponding to each historical published content; wherein the interactive guidance type is a preset guidance type among multiple preset guidance types.
[0035] In practical implementation, the first step is to construct a first sample set, which contains multiple historical posts. These posts can be content that has already been recommended and distributed within the target application, such as content with varying exposure and interaction frequencies. Alternatively, they can be verified as highly interactive trending content within the target application, selected based on metrics such as the number of favorites and comments. For example, posts by users with a first preset number of favorites are considered trending content, as are posts by users with a second preset number of comments.
[0036] Since the first sample set includes historical published content, the interactive guidance type corresponding to each historical published content is known.
[0037] Step 11: Extract multi-dimensional features for each historical post in the first sample set to obtain multi-dimensional feature data corresponding to each historical post.
[0038] For each historical post in the first sample set, perform the following operation: extract multi-dimensional features from the historical post to obtain multi-dimensional feature data corresponding to the historical post, which includes text features, visual features, and interactive feedback features.
[0039] Step 12: Based on the multi-dimensional feature data corresponding to each historical published content, determine the feature patterns corresponding to multiple preset guidance types.
[0040] In the specific implementation, after obtaining the multidimensional feature data corresponding to each historical published content in the first sample set, the multidimensional feature data can be associated with the interactive guidance type corresponding to each historical published content, so as to obtain the feature patterns corresponding to each interactive guidance type. The interactive guidance type here is the aforementioned preset guidance type.
[0041] For example, the aforementioned preset guidance types can include collection-oriented and comment-oriented types. The characteristics of collection-oriented content include: content containing keywords such as "tutorial" and "guide," presented in a step-by-step structure, and often accompanied by informational images. The characteristics of comment-oriented content include: content often containing open-ended question keywords, the main text introducing viewpoints or controversial topics, and sometimes accompanied by emotionally impactful visual elements.
[0042] Step 13: Input the feature patterns corresponding to the second sample set and multiple preset guidance types into the teacher model, so that the teacher model can automatically annotate each content to be annotated in the second sample set based on the feature patterns corresponding to the multiple preset guidance types, and obtain the annotation data corresponding to each content to be annotated; wherein, the annotation data includes at least: the interactive guidance type and multidimensional feature data corresponding to the content to be annotated.
[0043] In practice, the second sample set contains multiple pieces of content to be labeled. These pieces of content can be unlabeled historical posts in the target application or user-posted content added in real time in the target application.
[0044] In practical applications, the API interface of the teacher model is used to batch call the content to be labeled in the second sample set and the feature patterns corresponding to multiple preset guidance types, respectively, into the teacher model. This enables automated labeling of a large amount of content to be labeled in the second sample set. The labeling results are structured data, containing the interactive guidance type and multi-dimensional feature data corresponding to each piece of content to be labeled.
[0045] The labeled data also includes a confidence score, which indicates the reliability of the interaction guidance type corresponding to the content to be labeled output by the teacher model. Specifically, the higher the confidence score, the higher the reliability of the interaction guidance type corresponding to the content to be labeled output by the teacher model.
[0046] In an optional embodiment, the specific implementation of step 13 above may include: determining target prompt words based on the feature patterns corresponding to multiple preset guidance types; wherein, the target prompt words include task objectives, analysis dimensions, and data output formats; inputting the target prompt words and the second sample set into the teacher model, guiding the teacher model to identify the multi-dimensional feature data corresponding to each content to be labeled in the second sample set through the target prompt words, and determining the interactive guidance type of the content to be labeled based on the multi-dimensional feature data; obtaining the labeled data corresponding to each content to be labeled based on the multi-dimensional feature data and interactive guidance type corresponding to each content to be labeled in the second sample set.
[0047] In practical implementation, based on the characteristic patterns corresponding to multiple preset guidance types, actual multi-dimensional target prompts are generated. These target prompts need to clearly define information such as task objectives, analysis dimensions, and output formats so that they can guide the teacher model to perform accurate content annotation.
[0048] For example, multiple preset guidance types are collection-oriented and comment-oriented, and the target prompt could be "You are a social app content analyst. Please determine whether the note is 'collection-oriented,' 'comment-oriented,' or 'no guidance' based on its content." 1. Analysis dimensions: Title keywords (e.g., "tutorial" / "help"); Body structure (step-by-step / open-ended questions); Content purpose (providing practical information / stimulating discussion); Visual element description (e.g., whether there are infographics / controversial images).
[0049] 2. Output format: {"Type":"Collection / Comment Oriented","Confidence":"0-1","Reason":["Keyword 1","Structural Feature 1"...]}.
[0050] The content to be annotated in the second sample set is combined with the target prompt words and then input into the teacher model. The teacher model automatically annotates each piece of content to be annotated in the second sample set, obtaining the annotation data corresponding to each piece of content to be annotated. This annotation data is structured data, which includes the interaction guidance type of the content to be annotated (this interaction guidance type can be collection guidance type, comment guidance type, or no guidance type), the judgment confidence level (a value between 0 and 1, the closer to 1, the more reliable the judgment), and the judgment reason. This judgment reason is also multi-dimensional feature data. For example, the annotation data of a certain content to be annotated is: {"Type":"Collection guidance","Confidence level":"0.92","Reason":["Keyword: Tutorial","Structural feature: Step-by-step explanation","Visual element: Infographic"]}.
[0051] The aforementioned multi-dimensional target cue word-guided batch annotation method transforms feature analysis results into structured target cue words, guiding the teacher model to achieve efficient and standardized annotation of multimodal content. This solves the problems of high cost and low efficiency in large-scale content annotation, while ensuring high quality of annotated data.
[0052] Step 14: Generate a training dataset based on the labeled data corresponding to each content to be labeled, and train the target model based on the training dataset and the preset loss function to obtain the trained target model.
[0053] In practical implementation, the unlabeled content in the second sample set and its corresponding labeled data can be combined into a training dataset. This training dataset is used to train the target model so that the target model can inherit the feature behavior of the teacher model. The aforementioned preset loss function can include a classification loss function, a regression loss function, or a label loss function, etc.
[0054] In practical applications, the aforementioned target model is a lightweight large model, and the output distribution of the trained target model closely approximates the output distribution of the teacher model. Specifically, in the knowledge distillation process, the output of the teacher model is used as a "teacher signal" to enable the target model to learn the teacher model's judgment logic and feature extraction capabilities regarding content interaction guidance.
[0055] In an optional embodiment, a batch iterative training method is adopted during the training of the target model, and an early stopping mechanism is introduced (based on the accuracy of the annotation on the validation set) to avoid model overfitting. The final target model maintains a judgment accuracy similar to that of the teacher model (error controlled within 5%), while the parameter size is only 1 / 36 of that of the teacher model (72b->2b), and the inference speed is improved by 20 times, which can meet the requirements of real-time processing.
[0056] In an optional embodiment, the specific process of training the target model based on the training dataset and a preset loss function to obtain the trained target model may include: obtaining target training samples from the training dataset, wherein the target training samples include content to be labeled and the corresponding labeled data; inputting the content to be labeled in the target training samples into the target model to obtain the output result; determining the loss value based on the output result, the labeled data in the target training samples, and the preset loss function; adjusting the model parameters of the target model based on the loss value to obtain the trained target model; and continuing to execute the step of obtaining target training samples from the training dataset until the loss value reaches a preset condition or a preset number of training iterations is reached.
[0057] In practical implementation, the aforementioned training dataset includes multiple training samples, each containing content to be labeled and corresponding labeled data. The target training sample can be any training sample in the training dataset, and each training sample in the training dataset can be used as a target training sample to train the target model. After determining the target training sample, the content to be labeled in the target training sample needs to be input into the target model to obtain the output result. Then, based on a preset loss function, the difference between the output result and the labeled data in the target training sample is calculated, and the loss value is determined based on this difference. If the loss value meets a preset condition, the target model at this point is determined as the trained target model. If the loss value does not meet the preset condition, the model parameters of the target model are adjusted based on the loss value to obtain the trained target model. Then, the trained target model is used as the new target model, and the steps of obtaining target training samples from the training dataset are repeated until the loss value reaches the preset condition or the preset number of training iterations is reached.
[0058] The target model trained in the above manner breaks through the limitations of traditional single-dimensional feature analysis, establishes a three-in-one feature analysis system of "text-visual-interactive feedback", accurately captures the core feature patterns of different interactive content, and provides a solid foundation for subsequent annotation and recognition.
[0059] In addition, to meet the need for real-time processing of user-published content, a targeted preset loss function was designed to enable efficient transfer of teacher model knowledge to the small model. While ensuring the accuracy of judgment, the model inference speed was greatly improved, thus resolving the contradiction between processing speed and effectiveness.
[0060] The following examples illustrate a method of content recommendation.
[0061] Specifically, the structured data mentioned above includes confidence scores, which indicate the reliability of the interaction guidance type corresponding to the user-posted content output by the target model. Based on this, the specific process of recommending user-posted content in the target application based on the structured data corresponding to user-posted content can be achieved through the following steps 20-22: Step 20: If the structured data corresponding to the user's published content indicates that the interaction guidance type corresponding to the user's published content is a specified guidance type, and the confidence level is greater than or equal to the preset confidence threshold, then the user's published content is stored in the candidate recommendation set.
[0062] The aforementioned designated guidance type is also known as the high-interaction type, which is a guidance type with a relatively high interaction frequency. This interaction frequency is related to factors such as exposure, number of favorites, and number of comments. Specifically, the aforementioned designated guidance type can be a favorites-oriented type or a comment-oriented type. The aforementioned preset reliability threshold can be 0.7 or 0.8, etc.
[0063] In the recall phase of the recommendation system, a new recall channel for highly interactive, guided content (i.e., content with a specified guidance type) has been added: user-posted content identified by the target model as belonging to the specified guidance type and with a confidence score ≥ 0.7 is given priority for inclusion in the candidate recommendation set. The user-posted content in this candidate recommendation set is the content that can be recommended to users of the target application.
[0064] Step 21: Set category tags for user-posted content in the candidate recommendation set according to the interaction guidance type corresponding to the user-posted content; different interaction guidance types correspond to different category tags.
[0065] In practice, when the interaction guidance type corresponding to a user's published content is a specified guidance type, a category tag is set for that user's published content. This category tag matches the interaction guidance type corresponding to the user's published content. For example, if the interaction guidance type corresponding to a user's published content is a collection guidance type, a "collection guidance tag" is set for the user's published content; if the interaction guidance type corresponding to a user's published content is a comment guidance type, a "comment guidance tag" is set for the user's published content.
[0066] Step 22: Based on the category tags set by the user-posted content, recommend the user-posted content in the target application.
[0067] In practice, based on the category tags set by each user's published content in the candidate recommendation set, the target account corresponding to the target user in the target application that matches the published content of each user can be determined, and the published content of the corresponding user can be recommended to the target account.
[0068] In an optional embodiment, the specific implementation process of step 22 above may include: obtaining the user interaction characteristics of the target account corresponding to the target application; sorting the user-posted content with category tags in the candidate recommendation set based on the user interaction characteristics to obtain the sorting result; and recommending a preset number of user-posted content ranked at the top of the sorting result to the target account.
[0069] In the subsequent sorting process, the high-interaction characteristics of content added based on users' historical preferences are used as user interaction features to personalize the sorting of user-published content with different interaction orientations, and finally push it to the target users.
[0070] In an optional embodiment, the system periodically collects the matching results of the target model's recognition results (equivalent to the structured data corresponding to the user-posted content mentioned above) with the actual user interaction data, feeds back misjudged cases to the training set, and incrementally updates and optimizes the target model to continuously improve the recognition accuracy. Misjudged cases can be user-posted content that the target model classifies as having a high-interaction orientation (equivalent to the specified orientation mentioned above), but whose actual interaction rate is low.
[0071] The above approach deeply integrates the target model's identification results with the recall and ranking stages of the recommendation system. Through a dedicated recall channel guided by "high interaction" and a personalized ranking strategy, it achieves accurate delivery of highly interactive content and improves recommendation effectiveness.
[0072] Corresponding to the above method embodiments, this disclosure also provides a data processing apparatus, which is disposed in an electronic device, the electronic device having a target application configured therein; such as Figure 2 As shown, the device includes: The data output module 20 is used to respond to the target application receiving user-published content, input the user-published content into the pre-trained target model, and obtain the structured data corresponding to the user-published content; wherein, the structured data includes at least: the interaction guidance type and multi-dimensional feature data corresponding to the user-published content; the target model is obtained by supervised training based on the data output by the preset teacher model, and the target model inherits the behavioral features of the teacher model.
[0073] The content recommendation module 21 is used to recommend user-published content in the target application based on the structured data corresponding to the user-published content.
[0074] The aforementioned data processing device enables efficient mining and accurate recommendation of user-posted content of various interactive types within the target application. Simultaneously, this method improves data identification and recommendation efficiency by identifying structured data from user-posted content using a lightweight target model.
[0075] Furthermore, the aforementioned device also includes a model training module, used for: acquiring a first sample set; wherein the first sample set includes: multiple historical published content and an interactive guidance type corresponding to each historical published content; wherein the interactive guidance type is a preset guidance type among multiple preset guidance types; performing multi-dimensional feature extraction on each historical published content in the first sample set to obtain multi-dimensional feature data corresponding to each historical published content; determining the feature patterns corresponding to the multiple preset guidance types based on the multi-dimensional feature data corresponding to each historical published content; inputting the second sample set and the feature patterns corresponding to the multiple preset guidance types into the teacher model, so that the teacher model automatically annotates each content to be annotated in the second sample set based on the feature patterns corresponding to the multiple preset guidance types, to obtain the annotation data corresponding to each content to be annotated; wherein the annotation data includes at least: the interactive guidance type and multi-dimensional feature data corresponding to the content to be annotated; generating a training dataset based on the annotation data corresponding to each content to be annotated, and training the target model based on the training dataset and a preset loss function to obtain the trained target model.
[0076] Furthermore, the target model described above is a lightweight large model, and the output distribution of the trained target model approximates the output distribution of the teacher model.
[0077] Furthermore, multidimensional feature data includes the following types of feature data: text feature data, visual feature data, and interactive feedback feature data.
[0078] Furthermore, the aforementioned model training module is also used to: determine target prompt words based on the feature patterns corresponding to multiple preset guidance types; wherein, the target prompt words include task objectives, analysis dimensions, and data output formats; input the target prompt words and the second sample set into the teacher model, guide the teacher model to identify the multi-dimensional feature data corresponding to each content to be labeled in the second sample set through the target prompt words, and determine the interactive guidance type of the content to be labeled based on the multi-dimensional feature data; and obtain the labeled data corresponding to each content to be labeled based on the multi-dimensional feature data and interactive guidance type corresponding to each content to be labeled in the second sample set.
[0079] Furthermore, the aforementioned labeled data also includes confidence scores, which indicate the reliability of the interactive guidance type corresponding to the content to be labeled output by the teacher model.
[0080] Furthermore, the aforementioned structured data includes a confidence level, which indicates the reliability of the interaction guidance type corresponding to the user-published content output by the target model. Based on this, the content recommendation module 21 is configured to: store the user-published content in a candidate recommendation set if the structured data corresponding to the user-published content indicates that the interaction guidance type is a specified guidance type, and the confidence level is greater than or equal to a preset confidence threshold; set category labels for the user-published content in the candidate recommendation set according to the interaction guidance type corresponding to the user-published content; wherein different interaction guidance types correspond to different category labels; and recommend the user-published content in the target application based on the category labels set for the user-published content.
[0081] Furthermore, the content recommendation module 21 described above is also used to: obtain user interaction features of the target account corresponding to the target application; sort the user-posted content with category tags in the candidate recommendation set based on the user interaction features to obtain the sorting result; and recommend a preset number of user-posted content that ranks first in the sorting result to the target account.
[0082] The data processing apparatus provided in this disclosure has the same implementation principle and technical effects as the aforementioned method embodiments. For the sake of brevity, any parts not mentioned in the apparatus embodiments can be referred to the corresponding content in the aforementioned method embodiments.
[0083] This embodiment also provides an electronic device, such as... Figure 3 As shown, the electronic device includes a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the aforementioned data processing method. This electronic device can be a server or a terminal device.
[0084] Specifically, the above data processing method is applied to an electronic device, which has a target application installed. The method includes: in response to the target application receiving user-posted content, inputting the user-posted content into a pre-trained target model to obtain structured data corresponding to the user-posted content; wherein, the structured data includes at least: the interaction guidance type and multi-dimensional feature data corresponding to the user-posted content; the target model is obtained through supervised training based on the data output by a pre-set teacher model, and the target model inherits the behavioral features of the teacher model; based on the structured data corresponding to the user-posted content, the target application recommends the user-posted content.
[0085] The data processing method described above enables efficient mining and accurate recommendation of user-posted content of various interactive types within the target application. Furthermore, this approach improves data identification and recommendation efficiency by using a lightweight target model to identify structured data from user-posted content.
[0086] In an optional embodiment, the target model is trained as follows: A first sample set is obtained; wherein the first sample set includes: multiple historical published content and the interactive guidance type corresponding to each historical published content; wherein the interactive guidance type is a preset guidance type among multiple preset guidance types; multi-dimensional feature extraction is performed on each historical published content in the first sample set to obtain multi-dimensional feature data corresponding to each historical published content; based on the multi-dimensional feature data corresponding to each historical published content, the feature patterns corresponding to the multiple preset guidance types are determined; the second sample set and the feature patterns corresponding to the multiple preset guidance types are input into the teacher model, so that the teacher model automatically annotates each content to be annotated in the second sample set based on the feature patterns corresponding to the multiple preset guidance types, obtaining the annotation data corresponding to each content to be annotated; wherein the annotation data includes at least: the interactive guidance type and multi-dimensional feature data corresponding to the content to be annotated; a training dataset is generated based on the annotation data corresponding to each content to be annotated, and the target model is trained based on the training dataset and a preset loss function to obtain the trained target model.
[0087] In an optional embodiment, the target model is a lightweight large model, and the output distribution of the trained target model approximates the output distribution of the teacher model.
[0088] In an optional embodiment, the multidimensional feature data includes the following types of feature data: text feature data, visual feature data, and interactive feedback feature data.
[0089] In an optional embodiment, the step of inputting the feature patterns corresponding to the second sample set and multiple preset guidance types into the teacher model, so that the teacher model can automatically annotate each content to be annotated in the second sample set based on the feature patterns corresponding to the multiple preset guidance types, and obtain the annotated data corresponding to each content to be annotated, includes: determining target prompt words based on the feature patterns corresponding to the multiple preset guidance types; wherein, the target prompt words include task objectives, analysis dimensions, and data output formats; inputting the target prompt words and the second sample set into the teacher model, guiding the teacher model to identify the multi-dimensional feature data corresponding to each content to be annotated in the second sample set through the target prompt words, and determining the interaction guidance type of the content to be annotated based on the multi-dimensional feature data; and obtaining the annotated data corresponding to each content to be annotated based on the multi-dimensional feature data and interaction guidance type corresponding to each content to be annotated in the second sample set.
[0090] In an optional embodiment, the above-mentioned labeled data also includes confidence scores, which are used to indicate the reliability of the interactive guidance type corresponding to the content to be labeled output by the teacher model.
[0091] In an optional embodiment, the structured data includes a confidence score, which indicates the reliability of the interaction guidance type corresponding to the user-posted content output by the target model. Based on this, the step of recommending user-posted content in the target application based on the structured data corresponding to the user-posted content includes: if the structured data corresponding to the user-posted content indicates that the interaction guidance type corresponding to the user-posted content is a specified guidance type, and the confidence score is greater than or equal to a preset confidence threshold, storing the user-posted content in a candidate recommendation set; setting classification labels for the user-posted content in the candidate recommendation set according to the interaction guidance type corresponding to the user-posted content; wherein different interaction guidance types correspond to different classification labels; and recommending the user-posted content in the target application based on the classification labels set for the user-posted content.
[0092] In an optional embodiment, the step of recommending user-posted content in a target application based on the category tags set for user-posted content includes: obtaining user interaction features of the target account corresponding to the target application; sorting the user-posted content with category tags in the candidate recommendation set based on the user interaction features to obtain a sorting result; and recommending a preset number of user-posted content items ranked first in the sorting result to the target account.
[0093] Furthermore, Figure 3 The electronic device shown also includes a bus 102 and a communication interface 103, with the processor 101, the communication interface 103 and the memory 100 connected via the bus 102.
[0094] The memory 100 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 102 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0095] Processor 101 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 101 or by instructions in software form. The processor 101 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 100, and processor 101 reads information from memory 100 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.
[0096] This disclosure also provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are invoked and executed by a processor, they cause the processor to implement the above-described data processing method. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0097] Specifically, the above data processing method is applied to an electronic device, which has a target application installed. The method includes: in response to the target application receiving user-posted content, inputting the user-posted content into a pre-trained target model to obtain structured data corresponding to the user-posted content; wherein, the structured data includes at least: the interaction guidance type and multi-dimensional feature data corresponding to the user-posted content; the target model is obtained through supervised training based on the data output by a pre-set teacher model, and the target model inherits the behavioral features of the teacher model; based on the structured data corresponding to the user-posted content, the target application recommends the user-posted content.
[0098] The data processing method described above enables efficient mining and accurate recommendation of user-posted content of various interactive types within the target application. Furthermore, this approach improves data identification and recommendation efficiency by using a lightweight target model to identify structured data from user-posted content.
[0099] In an optional embodiment, the target model is trained as follows: A first sample set is obtained; wherein the first sample set includes: multiple historical published content and the interactive guidance type corresponding to each historical published content; wherein the interactive guidance type is a preset guidance type among multiple preset guidance types; multi-dimensional feature extraction is performed on each historical published content in the first sample set to obtain multi-dimensional feature data corresponding to each historical published content; based on the multi-dimensional feature data corresponding to each historical published content, the feature patterns corresponding to the multiple preset guidance types are determined; the second sample set and the feature patterns corresponding to the multiple preset guidance types are input into the teacher model, so that the teacher model automatically annotates each content to be annotated in the second sample set based on the feature patterns corresponding to the multiple preset guidance types, obtaining the annotation data corresponding to each content to be annotated; wherein the annotation data includes at least: the interactive guidance type and multi-dimensional feature data corresponding to the content to be annotated; a training dataset is generated based on the annotation data corresponding to each content to be annotated, and the target model is trained based on the training dataset and a preset loss function to obtain the trained target model.
[0100] In an optional embodiment, the target model is a lightweight large model, and the output distribution of the trained target model approximates the output distribution of the teacher model.
[0101] In an optional embodiment, the multidimensional feature data includes the following types of feature data: text feature data, visual feature data, and interactive feedback feature data.
[0102] In an optional embodiment, the step of inputting the feature patterns corresponding to the second sample set and multiple preset guidance types into the teacher model, so that the teacher model can automatically annotate each content to be annotated in the second sample set based on the feature patterns corresponding to the multiple preset guidance types, and obtain the annotated data corresponding to each content to be annotated, includes: determining target prompt words based on the feature patterns corresponding to the multiple preset guidance types; wherein, the target prompt words include task objectives, analysis dimensions, and data output formats; inputting the target prompt words and the second sample set into the teacher model, guiding the teacher model to identify the multi-dimensional feature data corresponding to each content to be annotated in the second sample set through the target prompt words, and determining the interaction guidance type of the content to be annotated based on the multi-dimensional feature data; and obtaining the annotated data corresponding to each content to be annotated based on the multi-dimensional feature data and interaction guidance type corresponding to each content to be annotated in the second sample set.
[0103] In an optional embodiment, the above-mentioned labeled data also includes confidence scores, which are used to indicate the reliability of the interactive guidance type corresponding to the content to be labeled output by the teacher model.
[0104] In an optional embodiment, the structured data includes a confidence score, which indicates the reliability of the interaction guidance type corresponding to the user-posted content output by the target model. Based on this, the step of recommending user-posted content in the target application based on the structured data corresponding to the user-posted content includes: if the structured data corresponding to the user-posted content indicates that the interaction guidance type corresponding to the user-posted content is a specified guidance type, and the confidence score is greater than or equal to a preset confidence threshold, storing the user-posted content in a candidate recommendation set; setting classification labels for the user-posted content in the candidate recommendation set according to the interaction guidance type corresponding to the user-posted content; wherein different interaction guidance types correspond to different classification labels; and recommending the user-posted content in the target application based on the classification labels set for the user-posted content.
[0105] In an optional embodiment, the step of recommending user-posted content in a target application based on the category tags set for user-posted content includes: obtaining user interaction features of the target account corresponding to the target application; sorting the user-posted content with category tags in the candidate recommendation set based on the user interaction features to obtain a sorting result; and recommending a preset number of user-posted content items ranked first in the sorting result to the target account.
[0106] The aforementioned functions, when implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of this disclosure, essentially, or the parts that contribute to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal device, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0107] In the description of this disclosure, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this disclosure. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0108] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A data processing method, characterized in that, The method is applied to an electronic device, wherein a target application is configured thereon; the method includes: In response to the target application receiving user-posted content, the user-posted content is input into a pre-trained target model to obtain structured data corresponding to the user-posted content; wherein, the structured data includes at least: the interaction guidance type and multi-dimensional feature data corresponding to the user-posted content; the target model is obtained through supervised training based on data output by a preset teacher model, and the target model inherits the behavioral features of the teacher model; Based on the structured data corresponding to the user-posted content, the target application makes recommendations for the user-posted content.
2. The method according to claim 1, characterized in that, The target model is trained in the following manner: Obtain a first sample set; wherein the first sample set includes: multiple historical published content and an interactive guidance type corresponding to each of the historical published content; wherein the interactive guidance type is a preset guidance type among multiple preset guidance types; Multi-dimensional feature extraction is performed on each of the historical published contents in the first sample set to obtain multi-dimensional feature data corresponding to each of the historical published contents. Based on the multidimensional feature data corresponding to each of the historical published contents, the feature patterns corresponding to multiple preset guidance types are determined respectively. The second sample set and the feature patterns corresponding to the multiple preset guidance types are input into the teacher model, so that the teacher model automatically annotates each content to be annotated in the second sample set based on the feature patterns corresponding to the multiple preset guidance types, and obtains the annotation data corresponding to each content to be annotated; wherein, the annotation data includes at least: the interactive guidance type and multidimensional feature data corresponding to the content to be annotated; A training dataset is generated based on the labeled data corresponding to each content to be labeled, and the target model is trained based on the training dataset and a preset loss function to obtain the trained target model.
3. The method according to claim 2, characterized in that, The target model is a lightweight large model, and the output distribution of the trained target model approximates the output distribution of the teacher model.
4. The method according to claim 2, characterized in that, The step of inputting the feature patterns corresponding to the second sample set and the multiple preset guidance types into the teacher model, so that the teacher model can automatically annotate each content to be annotated in the second sample set based on the feature patterns corresponding to the multiple preset guidance types, and obtain the annotated data corresponding to each content to be annotated, includes: Based on the characteristic patterns corresponding to the multiple preset guidance types, target prompt words are determined; wherein, the target prompt words include task objectives, analysis dimensions, and data output formats; The target prompt word and the second sample set are input into the teacher model. The teacher model is guided by the target prompt word to identify the multi-dimensional feature data corresponding to each content to be labeled in the second sample set, and the interactive guidance type of the content to be labeled is determined based on the multi-dimensional feature data. Based on the multidimensional feature data corresponding to each content to be annotated in the second sample set and the interactive guidance type, the annotation data corresponding to each content to be annotated is obtained.
5. The method according to claim 4, characterized in that, The labeled data also includes a confidence level, which indicates the reliability of the interactive guidance type corresponding to the content to be labeled output by the teacher model.
6. The method according to claim 1, characterized in that, The structured data includes confidence scores, which are used to indicate the reliability of the interaction guidance type corresponding to the user-posted content output by the target model. The step of recommending user-posted content in the target application based on the structured data corresponding to the user-posted content includes: If the structured data corresponding to the user-posted content indicates that the interaction guidance type corresponding to the user-posted content is a specified guidance type, and the confidence level is greater than or equal to a preset confidence threshold, the user-posted content is stored in the candidate recommendation set. Based on the interaction guidance type corresponding to the user-posted content, category tags are set for the user-posted content in the candidate recommendation set; wherein, different interaction guidance types correspond to different category tags; Based on the category tags set for the user-posted content, the user-posted content is recommended in the target application.
7. The method according to claim 6, characterized in that, The step of recommending user-posted content in the target application based on the category tags set by the user includes: Obtain the user interaction characteristics of the target account corresponding to the target application; Based on the user interaction features, the user-posted content with category tags in the candidate recommendation set is sorted to obtain the sorting result; The top-ranked user-posted content from the sorting results will be recommended to the target account.
8. A data processing apparatus, characterized in that, The device is installed in an electronic device, which contains a target application; the device includes: The data output module is used to respond to the target application receiving user-published content, input the user-published content into a pre-trained target model, and obtain structured data corresponding to the user-published content; wherein, the structured data includes at least: the interaction guidance type and multi-dimensional feature data corresponding to the user-published content; the target model is obtained by supervised training based on the data output by the preset teacher model, and the target model inherits the behavioral features of the teacher model; The content recommendation module is used to recommend user-published content in the target application based on the structured data corresponding to the user-published content.
9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the data processing method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the data processing method according to any one of claims 1-7.