Data processing method and device, storage medium and electronic equipment
By dividing the material collection into high-interaction and low-interaction categories and using the target model for multi-level verification, the problem of insufficient accuracy in multimodal material analysis is solved, and more accurate creative analysis and optimization are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies struggle to effectively analyze multimodal materials, failing to capture the interactions between and within different modalities, resulting in insufficient accuracy and practicality of the analysis results.
The material set is divided into two categories: Category 1 materials with high interactive data volume and Category 2 materials with low interactive data volume. Attribution analysis is performed using the target model to generate initial analysis results, and multi-level validation is conducted to generate target analysis results.
It improves the comprehensiveness and accuracy of creative analysis, identifies the true creative drivers, and promotes the optimization of creative content.
Smart Images

Figure CN121858796A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a data processing method and apparatus, a storage medium and an electronic device. Background Technology
[0002] In the field of creative content analytics, with the explosive growth of internet media and the rapid evolution of consumer preferences, accurately understanding and predicting which creative elements will attract the target audience has become a major challenge. Although the industry is actively exploring data-driven analytical methods to analyze the effectiveness of content, most existing technologies rely on statistical models to analyze single-dimensional data. However, when faced with multimodal content containing complex combinations of images, videos, and text, these models struggle to capture the interactions between and within different modalities, resulting in significant deficiencies in the accuracy and practicality of the analytical results.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This application provides a data processing method and apparatus, a storage medium and an electronic device to at least solve the technical problem in the related art that incomplete analysis of materials leads to relatively low accuracy of analysis results.
[0005] According to one aspect of the embodiments of this application, a data processing method is provided, comprising: acquiring a material set, and dividing the materials in the material set into a first type of material and a second type of material, wherein the interaction data volume of the first type of material is higher than that of the second type of material; performing attribution analysis on the first type of material and the second type of material through a target model to obtain an initial analysis result, wherein the initial analysis result includes at least: feature labels of differences between the first type of material and the second type of material; performing analysis and verification based on the initial analysis result to obtain a verification result, and generating a target analysis result corresponding to the material set based on the verification result and the initial analysis result, wherein the target analysis result is used to provide a reference for generating target materials.
[0006] Further, dividing the materials in the material set into a first category and a second category includes: obtaining interactive data information of the materials in the material set; calculating the evaluation value corresponding to the materials in the material set based on the interactive data information; and dividing the materials in the material set according to the evaluation value to obtain the first category and the second category of materials.
[0007] Furthermore, before performing attribution analysis on the first type of material and the second type of material through the target model, the method further includes: clustering the first type of material to obtain multiple first content type clusters; and clustering the second type of material to obtain multiple second content type clusters.
[0008] Furthermore, attribution analysis is performed on the first type of material and the second type of material using the target model to obtain initial analysis results, including: performing attribution analysis on multiple first type materials corresponding to content type clusters in the multiple first content type clusters to obtain a first analysis result, and performing attribution analysis on multiple second type materials corresponding to content type clusters in the multiple second content type clusters to obtain a second analysis result; performing difference analysis on the first type of material and the second type of material to obtain a third analysis result; and performing inductive analysis on the first analysis result, the second analysis result, the third analysis result, the first type of material, and the second type of material to obtain the initial analysis result.
[0009] Further, attribution analysis is performed on multiple first-category materials corresponding to the multiple first content type clusters to obtain a first analysis result, including: for a target content type cluster, intra-category feature analysis is performed on multiple first-category materials in the target content type cluster to obtain a first analysis sub-result; a preset number of target materials are randomly selected from the multiple first-category materials, and the target materials are subjected to in-depth analysis to obtain a second analysis sub-result; based on the first analysis sub-result and the analysis sub-result, the first analysis result is obtained.
[0010] Furthermore, the initial analysis results are obtained by summarizing and analyzing the first analysis result, the second analysis result, the third analysis result, the first type of material, and the second type of material. This includes: merging the first type of material and the second type of material into categories to obtain merged material; and obtaining the initial analysis results based on the merged material, the first analysis result, the second analysis result, and the third analysis result.
[0011] Further, the initial analysis results are analyzed and verified to obtain the verification results, including: performing single-factor attribution verification on the initial analysis results to obtain a first initial verification result; performing full-factor attribution verification on the initial analysis results to obtain a second initial verification result; and obtaining the verification result based on the first initial verification result and the second initial verification result.
[0012] Further, performing single-factor attribution validation on the initial analysis results to obtain a first initial validation result includes: calculating a first proportion value of the feature labels in the initial analysis results in the first type of materials and a second proportion value in the second type of materials, and obtaining a first validation sub-result based on the first proportion value and the second proportion value; drawing a box plot on the feature labels in the initial analysis results based on the evaluation values corresponding to the materials in the material set, and obtaining a second validation sub-result based on the box plot; performing a significance analysis on the feature labels in the initial analysis results based on the material set to obtain a third validation sub-result; and obtaining the first initial validation result based on the first validation sub-result, the second validation sub-result, and the third validation sub-result.
[0013] Further, performing full-factor attribution verification on the initial analysis results to obtain a second initial verification result includes: obtaining a prediction model with feature tags in the initial analysis results as features and evaluation values corresponding to the materials in the material set as targets; obtaining the content influence value corresponding to the feature tags in the initial analysis results through the prediction model, and obtaining the second initial verification result based on the content influence value.
[0014] According to another aspect of the embodiments of this application, a data processing method is also provided, comprising: acquiring a set of materials uploaded by a client; dividing the materials in the set into a first type of materials and a second type of materials in a cloud server, wherein the interaction data volume of the first type of materials is higher than that of the second type of materials; performing attribution analysis on the first type of materials and the second type of materials through a target model to obtain an initial analysis result, wherein the initial analysis result includes at least: feature labels of differences between the first type of materials and the second type of materials; performing analysis and verification based on the initial analysis result to obtain a verification result, and generating a target analysis result corresponding to the set of materials based on the verification result and the initial analysis result, wherein the target analysis result is used to provide a reference for generating target materials; and returning the target analysis result to the client.
[0015] According to another aspect of the embodiments of this application, a data processing apparatus is also provided, comprising: an acquisition unit, configured to acquire a material set and divide the materials in the material set into a first type of material and a second type of material, wherein the amount of interactive data of the first type of material is higher than the amount of interactive data of the second type of material; an analysis unit, configured to perform attribution analysis on the first type of material and the second type of material through a target model to obtain an initial analysis result, wherein the initial analysis result includes at least: feature labels of differences between the first type of material and the second type of material; and a verification unit, configured to perform analysis and verification based on the initial analysis result to obtain a verification result, and generate a target analysis result corresponding to the material set based on the verification result and the initial analysis result, wherein the target analysis result is used to provide a reference for generating target materials.
[0016] Furthermore, the acquisition unit includes: an acquisition subunit for acquiring interactive data information of materials in the material set; a calculation subunit for performing calculations based on the interactive data information to obtain evaluation values corresponding to the materials in the material set; and a division subunit for dividing the materials in the material set based on the evaluation values to obtain the first type of materials and the second type of materials.
[0017] Furthermore, the device further includes: a first clustering unit, used to cluster the first type of material to obtain multiple first content type clusters before performing attribution analysis on the first type of material and the second type of material through the target model; and a second clustering unit, used to cluster the second type of material to obtain multiple second content type clusters.
[0018] Further, the analysis unit includes: a first analysis subunit, used to perform attribution analysis on multiple first-type materials corresponding to the content type clusters in the plurality of first content type clusters to obtain a first analysis result, and to perform attribution analysis on multiple second-type materials corresponding to the content type clusters in the plurality of second content type clusters to obtain a second analysis result; a second analysis subunit, used to perform difference analysis on the first-type materials and the second-type materials to obtain a third analysis result; and a third analysis subunit, used to perform inductive analysis on the first analysis result, the second analysis result, the third analysis result, the first-type materials, and the second-type materials to obtain the initial analysis result.
[0019] Further, the first analysis subunit includes: a first analysis module, used to perform intra-class feature analysis on multiple first-type materials in the target content type cluster to obtain a first analysis sub-result; a second analysis module, used to randomly select a preset number of target materials from the multiple first-type materials and perform deep analysis on the target materials to obtain a second analysis sub-result; and a first determination module, used to obtain the first analysis result based on the first analysis sub-result and the analysis sub-result.
[0020] Furthermore, the third analysis subunit includes: a merging module, used to merge the first type of material and the second type of material to obtain merged material; and a second determining module, used to obtain the initial analysis result based on the merged material, the first analysis result, the second analysis result, and the third analysis result.
[0021] Further, the verification unit includes: a first verification subunit, used to perform single-factor attribution verification on the initial analysis results to obtain a first initial verification result; a second verification subunit, used to perform full-factor attribution verification on the initial analysis results to obtain a second initial verification result; and a determination subunit, used to obtain the verification result based on the first initial verification result and the second initial verification result.
[0022] Further, the first verification subunit includes: a calculation module, used to calculate the first proportion value of the feature labels in the initial analysis result in the first type of material and the second proportion value in the second type of material, and obtain a first verification sub-result based on the first proportion value and the second proportion value; a processing module, used to draw a box plot on the feature labels in the initial analysis result based on the evaluation values corresponding to the materials in the material set, and obtain a second verification sub-result based on the box plot; a third analysis module, used to perform saliency analysis on the feature labels in the initial analysis result based on the material set, and obtain a third verification sub-result; and a third determination module, used to obtain the first initial verification result based on the first verification sub-result, the second verification sub-result, and the third verification sub-result.
[0023] Furthermore, the second verification subunit includes: a first acquisition module, used to acquire a prediction model with feature tags in the initial analysis results as features and evaluation values corresponding to the materials in the material set as targets; and a second acquisition module, used to acquire the content influence degree value corresponding to the feature tags in the initial analysis results through the prediction model, and obtain the second initial verification result based on the content influence degree value.
[0024] According to another aspect of the present invention, an electronic device is also provided, comprising: a memory storing an executable program; and a processor for running the program, wherein the program executes the data processing method described above during runtime.
[0025] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein the storage medium stores a program, wherein the program controls the device where the storage medium is located to execute the data processing method described above during runtime.
[0026] According to another aspect of the present invention, a computer program product is also provided, including a computer program or instructions, which, when executed by a processor, implement the data processing method described above.
[0027] In this embodiment, the following steps are adopted: obtaining a material set and dividing the materials in the material set into a first type of material and a second type of material, wherein the interaction data volume of the first type of material is higher than that of the second type of material; performing attribution analysis on the first type of material and the second type of material through a target model to obtain initial analysis results, wherein the initial analysis results include at least: feature labels of the differences between the first type of material and the second type of material; performing analysis and verification based on the initial analysis results to obtain verification results, and generating target analysis results corresponding to the material set based on the verification results and the initial analysis results, wherein the target analysis results are used to provide a reference for generating target materials, thereby solving the technical problem in related technologies where incomplete material analysis leads to low accuracy of analysis results.
[0028] In this application, materials are explicitly categorized into two types based on the amount of interactive data. This categorization helps focus on the most effective materials for in-depth analysis, while also allowing for comparison with less effective materials to identify the true creative drivers. Attribution analysis is performed on these two types of materials using a target model, generating initial analysis results including feature labels. To ensure the high reliability and effectiveness of the insights generated by the target model, the initial analysis results are validated to produce more accurate and scientific verification results. Finally, based on the initial and validation results, the target analysis results are obtained. Through innovative material categorization, target model analysis, multi-level validation, and comprehensive report generation, the shortcomings of existing technologies in material analysis are overcome, improving the comprehensiveness and accuracy of creative analysis, thereby promoting the optimization of creative content. Attached Figure Description
[0029] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0030] Figure 1 This is a hardware structure block diagram of a computer terminal provided according to Embodiment 1 of this application;
[0031] Figure 2 This is a flowchart of the data processing method provided according to Embodiment 1 of this application;
[0032] Figure 3 This is a schematic diagram of the data processing method provided according to Embodiment 1 of this application;
[0033] Figure 4 This is a flowchart of the data processing method provided according to Embodiment 2 of this application;
[0034] Figure 5 This is a schematic diagram of a data processing apparatus provided according to Embodiment 3 of this application;
[0035] Figure 6 This is a structural block diagram of an electronic device provided according to Embodiment 4 of this application. Detailed Implementation
[0036] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0037] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0038] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0039] Example 1
[0040] According to an embodiment of this application, a data processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0041] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method is shown. Figure 1 As shown, the computer terminal (or mobile device) 10 may include a processor set 102 (the processor set 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA, and the processor set 102 may include a processor set, Figure 1 The data is illustrated using 102a, 102b, ..., 102n. A memory 104 is used for storing data, and a transmission module 106 is used for communication functions. In addition, it may include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0042] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0043] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the data processing method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the aforementioned data processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0044] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0045] The display may be, for example, a touchscreen LCD display that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0046] Under the aforementioned operating environment, this application provides the following: Figure 2 The data processing method shown. Figure 2 This is a flowchart of a data processing method according to Embodiment 1 of this application. The data processing method includes:
[0047] Step S201: Obtain the material set and divide the materials in the material set into a first category of materials and a second category of materials, wherein the amount of interactive data of the first category of materials is higher than that of the second category of materials.
[0048] Optionally, materials can be collected from various channels such as databases, social media, and advertising platforms, including but not limited to images, videos, text, and audio. These materials constitute the initial material set. Each piece of material is quantified, for example, by the number of user clicks, viewing time, comments, likes, shares, and conversion rate (i.e., the percentage of users who complete the expected behavior). By statistically analyzing these metrics, the amount of interaction data for the materials can be obtained, which is an important indicator of user interest and engagement with the creative content.
[0049] In an optional embodiment, materials are categorized into a first category and a second category based on the amount of interactive data they generate. It should be noted that the first category of materials generates more interactive data than the second category. Based on established criteria, the materials in the set are classified. Materials that perform better are classified as first-category materials, while the rest are classified as second-category materials. This classification process ensures that subsequent analysis focuses on effective and impactful content.
[0050] Step S202: Attribution analysis is performed on the first type of material and the second type of material using the target model to obtain initial analysis results. The initial analysis results include at least the feature labels of the differences between the first type of material and the second type of material.
[0051] Optionally, attribution analysis can be performed on the categorized first and second categories of materials using a target model (Language Model, LLM). This analysis aims to identify the differences in creative elements between the two categories of materials, thereby providing preliminary insights for subsequent analysis, verification, and creative optimization. It should be noted that the target model can be a Language Model (LLM), which can handle multimodal data, such as images, videos, text, and audio.
[0052] For example, features are extracted from the content of the first and second categories of materials, including color, style, characters, scenes, and copywriting style. Then, natural language processing (NLP) techniques are used to convert these features into vector representations for LLM analysis. For image and video materials, pre-trained visual models can be used for feature extraction; for audio, speech recognition technology can be used to transcribe it into text; and for text content, text vectorization is performed directly. The feature vectors of the two categories of materials are compared and analyzed to identify the unique features of the first category compared to the second, as well as potential differences. This comparison process helps to understand what creative elements drive the high interactivity of the first category of materials. Based on the above analysis, LLM generates initial analysis results, which may include differences in feature labels between the first and second categories of materials. Feature labels can involve visual elements (such as color and composition), textual elements (such as copywriting style and emotional tone), sound characteristics (such as background music and sound effects), and overall creative methods (such as storytelling and visual presentation techniques).
[0053] Step S203: Analyze and verify the initial analysis results to obtain the verification results. Based on the verification results and the initial analysis results, generate the target analysis results corresponding to the material set. The target analysis results are used to provide a reference for generating the target materials.
[0054] Optionally, after obtaining the initial analysis results generated by the target model (LLM), the initial analysis results are analyzed and validated. For example, key individual features, such as color, persona, or copywriting style, are selected from the feature labels proposed by the LLM. Statistical methods are used to verify whether the difference between the first category (high interaction) and the second category (low interaction) of the feature is significant, thus obtaining the target analysis results. The target analysis results may also include details of data analysis combined with creative practices, as well as actionable insights and recommendations, such as "using a combination of X-style copywriting and Y-tone images can significantly improve the interaction rate of the material." It should be noted that target material refers to the new material to be generated.
[0055] In summary, based on the amount of interactive data, the materials are clearly divided into two categories: Category 1 and Category 2. This classification helps focus on effective materials for in-depth analysis, while also allowing for comparison with mediocre materials to identify the true creative drivers. Attribution analysis is performed on these two categories of materials using a target model, generating initial analysis results including feature labels. To ensure the high reliability and effectiveness of the insights generated by the target model, the initial analysis results are validated to produce more accurate and scientific verification results. Finally, based on the initial and validation results, the target analysis results are obtained. Through innovative material classification, target model analysis, multi-level validation, and comprehensive report generation, the shortcomings of existing technologies in material analysis are overcome, improving the comprehensiveness and accuracy of creative analysis, thereby promoting the optimization of creative content.
[0056] To improve the accuracy of classification, the data processing method provided in Embodiment 1 of this application divides the materials in the material set into a first category of materials and a second category of materials, including: obtaining the interaction data information of the materials in the material set; calculating the evaluation value corresponding to the materials in the material set based on the interaction data information; and dividing the materials in the material set according to the evaluation value to obtain the first category of materials and the second category of materials.
[0057] Optionally, interaction data for the materials can be collected from multiple sources, including but not limited to user click-through rate, view completion rate, likes and comments ratio, and sharing and saving behavior. This data reflects the degree of user interaction with the materials and is a direct indicator of the materials' attractiveness and influence. Based on the interaction data, an evaluation value is calculated for each material in the material set. This evaluation value can be a comprehensive score, calculated based on a weighted average or composite indicators, used to quantify the overall performance of the materials.
[0058] Finally, based on the established classification criteria, the material collection is divided into Category 1 (high interaction volume) and Category 2 (low interaction volume) materials according to their evaluation scores. For example, content with a popularity score (i.e., the evaluation score mentioned above) higher than "75th percentile + 1.5 times IQR (Interquartile Range, a statistical measure of data dispersion, equal to the difference between the upper quartile (Q3) and the lower quartile (Q1))" is defined as high-quality content (i.e., Category 1 materials), while the rest are considered ordinary content (i.e., Category 2 materials).
[0059] In an optional embodiment, the selected high-quality content (i.e., the first type of material) and its fine-grained tags can be stored in a database. The database supports multi-dimensional retrieval, allowing content creators to find high-quality examples that have been proven in the market based on specific creative directions (such as "summer coolness" or "tech style"), providing "direct references" for content production.
[0060] The above steps effectively categorize the materials in the material collection according to the amount of interaction data, providing accurate basic data for subsequent attribution analysis and creative optimization.
[0061] To improve the comprehensiveness of subsequent analysis, in the data processing method provided in Embodiment 1 of this application, before performing attribution analysis on the first type of material and the second type of material through the target model, the method further includes: clustering the first type of material to obtain multiple first content type clusters; and clustering the second type of material to obtain multiple second content type clusters.
[0062] Optionally, key features are extracted from the materials, including but not limited to color, composition, facial expressions, musical rhythm, and textual emotion. Then, natural language processing techniques and computer vision models are used to convert these features into semantic vectors, facilitating clustering algorithms. The feature vectors of the first and second categories of materials are then clustered separately to obtain multiple first content type clusters and multiple second content type clusters.
[0063] Clustering groups materials with similar attributes into the same cluster, with materials in different clusters exhibiting significant differences in characteristics. After clustering, it's crucial to check for outlier clusters with insufficient sample sizes or features that deviate significantly from other clusters. These outlier clusters may consist of noisy data or a very small number of exceptional cases, and if left untreated, they could interfere with subsequent analysis. In such cases, re-clustering or removal should be performed depending on the specific circumstances. After clustering, clusters can be defined and named based on common characteristics of the materials within each cluster. For example, a cluster of "bright colors + dynamic music" could be named "energetic party style," while a cluster of "soft colors + narrative text" could be named "emotional story style."
[0064] Cluster analysis of materials allows for a deeper understanding of the specific attributes of different types of materials, rather than simply making broad comparisons based on total interaction volume. When conducting attribution analysis, segmenting and comparing different cluster types allows for a more precise identification of which specific creative elements (such as composition or facial expressions) have a significant impact on the material's performance.
[0065] To improve the accuracy of the analysis, the data processing method provided in Embodiment 1 of this application performs attribution analysis on the first type of material and the second type of material using a target model to obtain initial analysis results. This includes: performing attribution analysis on multiple first type materials corresponding to content type clusters in multiple first content type clusters to obtain a first analysis result; performing attribution analysis on multiple second type materials corresponding to content type clusters in multiple second content type clusters to obtain a second analysis result; performing difference analysis on the first type of material and the second type of material to obtain a third analysis result; and performing inductive analysis on the first analysis result, the second analysis result, the third analysis result, the first type of material, and the second type of material to obtain the initial analysis result.
[0066] Optionally, intra-cluster analysis is performed on the first content type cluster within the first category of materials (high-interaction materials) to obtain the first analysis results. This allows for a deeper exploration of how the materials within each cluster attract and engage users, resulting in first analysis results for different clusters and revealing the commonalities and differences among high-interaction materials across various types.
[0067] Attribution analysis was performed on multiple second-content-type clusters within the second category of materials (low-interaction materials), yielding the second analysis result. Based on the first and second analysis results, a cross-category cluster comparative analysis was conducted to explore the differences between high-interaction and low-interaction materials in terms of creative elements, presentation format, and audience targeting, resulting in the third analysis result. This analysis helps identify which creative features are significantly positively correlated with user interaction and which features may be factors leading to low interaction rates. Finally, the first three analysis results (first, second, and third analysis results) and specific examples from both categories of materials were comprehensively summarized and analyzed to obtain the initial analysis result.
[0068] Through the detailed implementation steps described above, the method of this application can not only deeply explore the relationship between the intrinsic attributes of materials and market performance, but also ensure the comprehensiveness and accuracy of the analysis process, effectively avoiding the bias that may be caused by a single analysis dimension, providing content creators with data-based and highly actionable creative insights, and improving the efficiency of content strategy formulation and implementation.
[0069] To improve the comprehensiveness of the material analysis, in the data processing method provided in Embodiment 1 of this application, attribution analysis is performed on multiple first-type materials corresponding to multiple first-content-type clusters to obtain a first analysis result, including: for a target content-type cluster, intra-class feature analysis is performed on multiple first-type materials in the target content-type cluster to obtain a first analysis sub-result; a preset number of target materials are randomly selected from multiple first-type materials, and in-depth analysis is performed on the target materials to obtain a second analysis sub-result; based on the first analysis sub-result and the analysis sub-result, a first analysis result is obtained.
[0070] Optionally, a detailed intra-category feature analysis is performed on the target content type clusters (i.e., subcategories within the first type of material). This analysis focuses on identifying common creative elements and expressive forms among the materials within the cluster, including but not limited to color usage, composition style, textual emotion, and musical rhythm. Through this process, the system can extract the feature profile of each cluster, obtaining the first analysis sub-result. A preset number of target materials are randomly selected from the target content type clusters for further in-depth analysis. For example, by comprehensively analyzing multi-dimensional data such as creative elements and user interaction patterns of the selected materials using the target model, a deeper understanding is gained of which specific elements and strategies play a key role within a particular type cluster, resulting in the second analysis sub-result.
[0071] Then, the first analysis sub-result and the analysis sub-result are integrated, and the advanced semantic understanding and inductive reasoning capabilities of LLM are used to conduct a comprehensive analysis of the intra-class features and deep analysis results. Finally, a comprehensive analysis conclusion for the first target content type cluster is extracted, which is the first analysis result.
[0072] Through the multi-layered analysis described above, this solution ensures that the attribution analysis of materials not only covers a wide range of content types but also delves into the details of specific creative elements, thereby improving the comprehensiveness and depth of the analysis. This hierarchical and clustered attribution analysis method can more accurately identify the relationship between creative elements and user interactions, providing content creators with more specific and effective optimization directions.
[0073] For further analysis, in the data processing method provided in Embodiment 1 of this application, the first analysis result, the second analysis result, the third analysis result, the first type of material and the second type of material are summarized and analyzed to obtain the initial analysis result, which includes: merging the first type of material and the second type of material to obtain the merged material; and obtaining the initial analysis result based on the merged material, the first analysis result, the second analysis result and the third analysis result.
[0074] Optionally, the first analysis results (including intra-category feature analysis results and in-depth analysis results of randomly selected target materials) obtained from multiple first content type clusters and the second analysis results obtained from the second content type clusters reflect the internal analysis results of high-interaction materials and low-interaction materials, respectively. The third analysis results, by comparing the first and second categories of materials, reveal the differences in creative elements and presentation styles between the two categories. Based on the first, second, and third analysis results, category merging can integrate insights from different content type clusters, forming a more holistic understanding. Through category merging, common creative elements and unique differentiation strategies across type clusters can be identified. Based on the merged materials, the first, second, and third analysis results, a comprehensive inductive analysis is performed, for example, identifying which creative elements stand out in all materials, which elements are additional contributors to specific type clusters, and which elements lead to differences in user interaction. Finally, initial analysis results based on type cluster characteristics, differences between materials, and market feedback are obtained, providing a comprehensive and in-depth basis for the formulation of creative strategies.
[0075] In an alternative embodiment, it is assumed that in a first target content type cluster (e.g., "tech-themed ads"), clear visuals, concise copy, and tech-inspired background music are common characteristics of highly interactive content. In a second target content type cluster (e.g., "lifestyle ads"), authentic personal stories and emotional resonance are found to be key to attracting users. Cross-category difference analysis also reveals that using "3D animation" in "tech-themed ads" significantly increases user interaction, while this element does not show the same effect in "lifestyle ads." Finally, through category merging and comprehensive inductive analysis, it can be concluded that "clear visuals + concise copy + tech-inspired background music + 3D animation" is a key combination strategy for "tech-themed ads," while "authentic personal stories + emotional resonance" is the core element of "lifestyle ads."
[0076] The above analysis process not only enhances the comprehensiveness and depth of material analysis but also ensures the accuracy and practicality of the final results. The initial analysis not only includes a detailed analysis of single-type clusters but also covers the general patterns and unique differences across cross-type clusters, providing content creators with comprehensive guidance from macro-strategies to specific creative elements.
[0077] To improve the reliability of the initial analysis results, in the data processing method provided in Embodiment 1 of this application, the analysis and verification based on the initial analysis results are performed to obtain the verification results, including: performing single-factor attribution verification on the initial analysis results to obtain a first initial verification result; performing full-factor attribution verification on the initial analysis results to obtain a second initial verification result; and obtaining the verification result based on the first initial verification result and the second initial verification result.
[0078] Optionally, the initial analysis results can be validated using single-factor attribution. For example, if the initial analysis indicates that "real people appearing on camera" is an important factor in increasing click-through rates, statistical methods can be used to test this element separately, assessing whether its performance varies significantly across different materials, and whether this variation can be quantified. This type of validation ensures that each mentioned feature has been rigorously tested, eliminating randomness and external confounding factors, and yielding the first initial validation results.
[0079] Considering the potential interactions and synergistic effects among creative elements, a full-factor attribution validation is performed on the initial analysis results. For example, a second initial validation result is obtained by constructing a multiple linear regression model, a decision tree model, or a neural network model for full-factor attribution validation. The first and second initial validation results are then combined and analyzed to obtain the final validation result.
[0080] The above verification process ensures the depth and accuracy of creative attribution analysis. Single-factor and full-factor attribution verification complement each other; the former focuses on examining the direct impact of individual creative elements, while the latter delves into the interactions and synergistic effects between features, together forming a comprehensive verification framework. This method reduces the subjectivity and uncertainty of the analysis, enhancing its practical value and scientific rigor in guiding creative strategies.
[0081] To improve the reliability of single-factor validation, the data processing method provided in Embodiment 1 of this application performs single-factor attribution validation on the initial analysis results to obtain a first initial validation result, including: calculating the first proportion value of the feature labels in the initial analysis results in the first type of materials and the second proportion value in the second type of materials, and obtaining a first validation sub-result based on the first and second proportion values; drawing a box plot on the feature labels in the initial analysis results based on the evaluation values corresponding to the materials in the material set, and obtaining a second validation sub-result based on the box plot; performing a significance analysis on the feature labels in the initial analysis results based on the material set to obtain a third validation sub-result; and obtaining the first initial validation result based on the first validation sub-result, the second validation sub-result, and the third validation sub-result.
[0082] Optionally, statistical processing is performed on the feature tags mentioned in the initial analysis results. The first percentage value of each feature tag in the first type of content (high-interaction content) and the second percentage value in the second type of content (low-interaction content) are calculated. By comparing these percentages, the difference in the feature tag between the two types of content can be quantified. If the percentage of the feature tag in high-interaction content is significantly higher than that in low-interaction content, this indicates that the feature may be an important factor leading to high interaction, and vice versa. The first validation sub-result preliminarily verifies the correlation between feature tags and content interaction volume through intuitive percentage differences.
[0083] Then, exploratory data analysis methods are used to visualize the evaluation values (such as click-through rate, viewing time, etc.) corresponding to the materials in the material set. For example, box plots are drawn on the feature labels in the initial analysis results to visually show the distribution of user interaction volume under different feature values. Box plots can clearly show the relationship between feature values and interaction volume, such as the position of the median in the box and the distribution of outliers. The second validation sub-result provides intuitive evidence of the influence of feature labels, helping to identify which feature values may lead to abnormally high interaction volume.
[0084] Statistical methods were used to analyze the dataset to examine whether the values of feature tags had a significant impact on user interaction volume. Significance analysis ensured that the results were not only based on observed differences but also that these differences were statistically significant and not accidental. The third validation sub-result improved the confidence level of the analysis conclusions, ensuring the reliability of the association between feature tags and high / low interaction volume content.
[0085] Finally, the first verification sub-result, the second verification sub-result, and the third verification sub-result are determined as the first initial verification result.
[0086] For example, the initial analysis indicated that "high color saturation" might be a key feature for improving click-through rate (CTR). By calculating the difference in feature tag proportions, if it was found that high-saturation colors accounted for a significantly higher proportion in high-interaction content than in low-interaction content, then the first validation sub-result initially supported this hypothesis. Next, a box plot visually showed that CTR distribution for high-saturation color content tended to be higher, further confirming the positive correlation between color saturation and CTR. Finally, statistical analysis determined that the impact of color saturation on CTR was statistically significant, solidifying the reliability of this finding. Combining the three validation sub-results, the first initial validation result can be concluded: high-saturation colors are one of the effective strategies for improving CTR. This conclusion is fully validated by the data, providing reliable guidance for subsequent content creation.
[0087] The above verification steps significantly improved the reliability of the initial analysis results. The first initial verification results not only validated the relationship between feature tags and material market performance, but also explored the nature and strength of these relationships through box plots and saliency analysis, ensuring that creative decisions are based on accurate and reliable data insights.
[0088] To improve the reliability of full-factor validation, in the data processing method provided in Embodiment 1 of this application, full-factor attribution validation is performed on the initial analysis results to obtain a second initial validation result. To improve the reliability of single-factor validation, the second initial validation result includes: obtaining a prediction model with feature labels in the initial analysis results as features and evaluation values corresponding to materials in the material set as targets; obtaining the content influence value corresponding to the feature labels in the initial analysis results through the prediction model; and obtaining the second initial validation result based on the content influence value.
[0089] Optionally, a predictive model can be constructed, which takes all feature labels identified in the initial analysis results as input features and the evaluation values (such as click-through rate, conversion rate, etc.) corresponding to the materials in the material set as the target output. It should be noted that the predictive model includes, but is not limited to, XG Boost or neural network models.
[0090] After the predictive model is trained, feature importance information (i.e., the content influence value mentioned above) is extracted from the model. Feature importance is an indicator that measures the degree of influence of a specific feature on the model's prediction results (i.e., the evaluation value of the material). In tree models such as XG Boost, feature importance is usually measured by calculating the number of times the feature is used at the model's split points or by reducing impurity; in neural network models, feature importance can be evaluated by observing the absolute value of the feature weights or by using gradient attribution methods (such as Gradient-based Feature Attribution). Finally, a second initial validation result is obtained based on the extracted content influence value.
[0091] For example, the initial analysis identified three feature tags—"technological elements," "bright colors," and "emotional expression"—that significantly impact the market performance of "technology product advertising" creatives. After training the constructed XGBoost model, feature importance analysis revealed that "technological elements" had a significantly higher content influence value than the other two features, and also contributed substantially to the prediction errors of click-through rate and conversion rate. In other words, the presence or absence of "technological elements" in technology product advertising has a decisive impact on the market performance of the creative. This result, serving as the second initial validation result, not only verified the importance of "technological elements" as a feature tag but also quantified its actual impact on the effectiveness of the creative, providing a reliable basis for subsequent creative decisions.
[0092] In an alternative embodiment, it can be achieved through, as follows: Figure 3 The diagram illustrates a comprehensive analysis of the materials, including: Step 1: Data Preprocessing and Construction of a High-Quality Content Library. Data Access: Integrating multimodal materials (videos, images, etc.), structured tags (elements describing the content, style, scenes, etc.), and performance data (click-through rate, conversion rate, etc.). Definition and Selection of High-Quality Content: Defining "high-quality content" using statistical standards. For example, content with a popularity score higher than "75th percentile + 1.5 times IQR" is defined as high-quality content, while the rest are considered ordinary content. Construction of a High-Quality Content Library: Storing high-quality content and its fine-grained tags in a database. This library supports multi-dimensional retrieval, allowing content creators to find market-proven high-quality examples based on specific creative directions (such as "summer coolness" or "tech style"), providing "direct references" for content production.
[0093] Step 2: Combines the "black-box" reasoning capabilities of LLM with the "white-box" verification capabilities of statistics / machine learning. Specifically, this includes: the LLM macro-level insight module:
[0094] The intelligent content clustering submodule vectorizes the tag sets and text descriptions of header and regular content, converting them into high-dimensional semantic vector representations. Based on these semantic vectors, it uses methods such as K-means to perform clustering analysis, automatically identifying different content type clusters. It filters out outlier categories with insufficient sample sizes to ensure the statistical validity of subsequent analyses.
[0095] The hierarchical insight generation submodule uses LLM to summarize intra-class features for valid categories and randomly selects typical examples for in-depth analysis. A carefully designed prompt guides LLM to compare and summarize the core differences between the two sets of content from a macro and semantic perspective. The analysis results of all categories are input into LLM for global summarization, automatically distinguishing different content types, merging categories with high similarity, and generating categorical inductive insights. For example, LLM might conclude: "Viral content tends to use fast-paced editing and live-action appearances, while ordinary content is more often presented with static images and text."
[0096] Multidimensional Validation Module: This module aims to rigorously validate the insights and hypotheses generated by the LLM, enhancing the confidence of the conclusions. One-Factor Attribution Analysis: Percentage Difference Analysis: For features mentioned in the LLM (such as "fast-paced editing"), calculate their percentage appearance in top-tier and regular content. Significant differences in percentage can corroborate their importance. Box plots are generated for key features to visually demonstrate the distribution differences in content popularity under different feature values. One-Factor Analysis of Variance: For categorical features (such as "scene type"), ANOVA (Analysis of Variance) is used to test whether different values of these features have a significant impact on content popularity.
[0097] Global Attribution Analysis: Build a predictive model (XG Boost, Light GBM, etc.) that uses all tags as features and aims at content popularity. Extract global feature importance ranking from the trained model to identify key factors that significantly impact content performance, thereby supplementing and validating the insights gained from LLM.
[0098] Step 3: Automatic Generation of a Comprehensive Attribution Report. This step integrates the macro-level insights generated by the LLM with the data evidence (such as significance p-values, percentage difference plots, and feature importance rankings) produced by the multi-dimensional validation module. A comprehensive analysis report, rich in visuals and data support, is automatically generated, providing strategic "indirect reference" for the task party. The report clearly identifies which creative elements are validated "growth points" and which are "risk points."
[0099] In the data processing method provided in Embodiment 1 of this application, a material set is obtained, and the materials in the material set are divided into a first type of material and a second type of material, wherein the amount of interactive data of the first type of material is higher than that of the second type of material; attribution analysis is performed on the first type of material and the second type of material through a target model to obtain initial analysis results, wherein the initial analysis results include at least: feature labels of the differences between the first type of material and the second type of material; analysis and verification are performed based on the initial analysis results to obtain verification results, and target analysis results corresponding to the material set are generated based on the verification results and the initial analysis results, wherein the target analysis results are used to provide a reference for generating target materials, thereby solving the technical problem in related technologies where incomplete material analysis leads to low accuracy of analysis results.
[0100] In this application, materials are explicitly categorized into two types based on the amount of interactive data. This categorization helps focus on effective materials for in-depth analysis, while also allowing for comparison with mediocre materials to identify the true creative drivers. Attribution analysis is performed on these two types of materials using a target model, generating initial analysis results including feature labels. To ensure the high reliability and effectiveness of the insights generated by the target model, the initial analysis results are validated to produce more accurate and scientific verification results. Finally, based on the initial and validation results, the target analysis results are obtained. Through innovative material categorization, target model analysis, multi-level validation, and comprehensive report generation, the shortcomings of existing technologies in material analysis are overcome, improving the comprehensiveness and accuracy of creative analysis, thereby promoting the optimization of creative content.
[0101] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0102] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0103] Example 2
[0104] According to embodiments of this application, a data processing method is also provided, such as... Figure 4 As shown, the data processing method includes:
[0105] Step S401: Obtain the collection of materials uploaded by the client;
[0106] Step S402: In the cloud server, the materials in the material set are divided into a first category and a second category, wherein the interaction data volume of the first category is higher than that of the second category; attribution analysis is performed on the first and second categories of materials using the target model to obtain initial analysis results, wherein the initial analysis results include at least the feature labels of the differences between the first and second categories of materials; analysis and verification are performed based on the initial analysis results to obtain verification results, and target analysis results corresponding to the material set are generated based on the verification results and the initial analysis results, wherein the target analysis results are used to provide a reference for generating target materials;
[0107] Step S403: Return the target analysis results to the client.
[0108] It should be noted that the specific processing procedure for the materials on the cloud server is the same as in Example 1, and will not be repeated here.
[0109] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0110] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0111] Example 3
[0112] According to embodiments of this application, a data processing apparatus for implementing the above-described data processing method is also provided, such as... Figure 5 As shown, the device includes: an acquisition unit 501, an analysis unit 502, and a verification unit 503.
[0113] The acquisition unit 501 is used to acquire a set of materials and divide the materials in the set into a first category of materials and a second category of materials, wherein the amount of interactive data of the first category of materials is higher than that of the second category of materials.
[0114] Analysis unit 502 is used to perform attribution analysis on the first type of material and the second type of material through the target model to obtain initial analysis results, wherein the initial analysis results include at least: feature labels of the differences between the first type of material and the second type of material;
[0115] The verification unit 503 is used to perform analysis and verification based on the initial analysis results, obtain the verification results, and generate the target analysis results corresponding to the material set based on the verification results and the initial analysis results. The target analysis results are used to provide a reference for generating the target materials.
[0116] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the acquisition unit includes: an acquisition subunit for acquiring interactive data information of materials in the material set; a calculation subunit for performing calculations based on the interactive data information to obtain the evaluation values corresponding to the materials in the material set; and a division subunit for dividing the materials in the material set based on the evaluation values to obtain a first type of material and a second type of material.
[0117] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the apparatus further includes: a first clustering unit, used to cluster the first type of material to obtain multiple first content type clusters before performing attribution analysis on the first type of material and the second type of material through the target model; and a second clustering unit, used to cluster the second type of material to obtain multiple second content type clusters.
[0118] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the analysis unit includes: a first analysis subunit, used to perform attribution analysis on multiple first-type materials corresponding to multiple content type clusters in multiple first content type clusters to obtain a first analysis result, and to perform attribution analysis on multiple second-type materials corresponding to multiple content type clusters in multiple second content type clusters to obtain a second analysis result; a second analysis subunit, used to perform difference analysis on the first-type materials and the second-type materials to obtain a third analysis result; and a third analysis subunit, used to perform inductive analysis on the first analysis result, the second analysis result, the third analysis result, the first-type materials, and the second-type materials to obtain an initial analysis result.
[0119] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the first analysis subunit includes: a first analysis module, used to perform intra-class feature analysis on multiple first-type materials in the target content type cluster to obtain a first analysis sub-result; a second analysis module, used to randomly select a preset number of target materials from the multiple first-type materials and perform in-depth analysis on the target materials to obtain a second analysis sub-result; and a first determination module, used to obtain a first analysis result based on the first analysis sub-result and the analysis sub-result.
[0120] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the third analysis subunit includes: a merging module, used to merge the first type of material and the second type of material to obtain the merged material; and a second determining module, used to obtain the initial analysis result based on the merged material, the first analysis result, the second analysis result, and the third analysis result.
[0121] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the verification unit includes: a first verification subunit, used to perform single-factor attribution verification on the initial analysis results to obtain a first initial verification result; a second verification subunit, used to perform full-factor attribution verification on the initial analysis results to obtain a second initial verification result; and a determination subunit, used to obtain a verification result based on the first initial verification result and the second initial verification result.
[0122] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the first verification subunit includes: a calculation module, used to calculate the first proportion value of the feature labels in the initial analysis result in the first type of material and the second proportion value in the second type of material, and obtain a first verification sub-result based on the first proportion value and the second proportion value; a processing module, used to draw a box plot on the feature labels in the initial analysis result based on the evaluation values corresponding to the materials in the material set, and obtain a second verification sub-result based on the box plot; a third analysis module, used to perform saliency analysis on the feature labels in the initial analysis result based on the material set, and obtain a third verification sub-result; and a third determination module, used to obtain a first initial verification result based on the first verification sub-result, the second verification sub-result, and the third verification sub-result.
[0123] Optionally, in the data processing apparatus provided in Embodiment 3 of this application, the second verification subunit includes: a first acquisition module, used to acquire a prediction model with feature tags in the initial analysis results as features and evaluation values corresponding to materials in the material set as targets; and a second acquisition module, used to acquire the content influence degree value corresponding to the feature tags in the initial analysis results through the prediction model, and obtain a second initial verification result based on the content influence degree value.
[0124] It should be noted that the acquisition unit 501, analysis unit 502, and verification unit 503 mentioned above correspond to steps S201 to S203 in Embodiment 1. The three units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above units, as part of the device, can run on the computer terminal 10 provided in Embodiment 1.
[0125] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0126] Example 4
[0127] Embodiments of this application may provide an electronic device, which may be any one of a group of electronic device terminals. Optionally, in this embodiment, the aforementioned electronic device may also be replaced by a terminal device such as a mobile terminal.
[0128] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0129] In this embodiment, the aforementioned electronic device can execute the program code for the following steps in the data processing method: acquiring a material set and dividing the materials in the material set into a first type of material and a second type of material, wherein the amount of interactive data of the first type of material is higher than that of the second type of material; performing attribution analysis on the first type of material and the second type of material through a target model to obtain initial analysis results, wherein the initial analysis results include at least: feature labels of the differences between the first type of material and the second type of material; performing analysis and verification based on the initial analysis results to obtain verification results, and generating target analysis results corresponding to the material set based on the verification results and the initial analysis results, wherein the target analysis results are used to provide a reference for generating target materials.
[0130] The aforementioned electronic device can execute the program code for the following steps in the data processing method: dividing the materials in the material set into a first category of materials and a second category of materials, including: obtaining interactive data information of the materials in the material set; calculating based on the interactive data information to obtain the evaluation value corresponding to the materials in the material set; and dividing the materials in the material set into a first category of materials and a second category of materials based on the evaluation value.
[0131] The aforementioned electronic device can execute the program code for the following steps in the data processing method: before performing attribution analysis on the first type of material and the second type of material through the target model, the method further includes: clustering the first type of material to obtain multiple first content type clusters; and clustering the second type of material to obtain multiple second content type clusters.
[0132] The aforementioned electronic device can execute the program code for the following steps in the data processing method: performing attribution analysis on the first type of material and the second type of material through the target model to obtain initial analysis results, including: performing attribution analysis on multiple first type materials corresponding to content type clusters in multiple first content type clusters to obtain first analysis results, and performing attribution analysis on multiple second type materials corresponding to content type clusters in multiple second content type clusters to obtain second analysis results; performing difference analysis on the first type of material and the second type of material to obtain third analysis results; and performing inductive analysis on the first analysis results, the second analysis results, the third analysis results, the first type of material and the second type of material to obtain initial analysis results.
[0133] The aforementioned electronic device can execute the program code for the following steps in the data processing method: performing attribution analysis on multiple first-type materials corresponding to multiple first-type content clusters to obtain a first analysis result, including: for a target content cluster, performing intra-class feature analysis on multiple first-type materials in the target content cluster to obtain a first analysis sub-result; randomly selecting a preset number of target materials from multiple first-type materials and performing in-depth analysis on the target materials to obtain a second analysis sub-result; and obtaining a first analysis result based on the first analysis sub-result and the analysis sub-result.
[0134] The aforementioned electronic device can execute the program code for the following steps in the data processing method: summarizing and analyzing the first analysis result, the second analysis result, the third analysis result, the first type of material, and the second type of material to obtain the initial analysis result, including: merging the first type of material and the second type of material to obtain the merged material; and obtaining the initial analysis result based on the merged material, the first analysis result, the second analysis result, and the third analysis result.
[0135] The aforementioned electronic device can execute the program code for the following steps in the data processing method: performing analysis and verification based on the initial analysis results to obtain verification results, including: performing single-factor attribution verification on the initial analysis results to obtain a first initial verification result; performing full-factor attribution verification on the initial analysis results to obtain a second initial verification result; and obtaining a verification result based on the first and second initial verification results.
[0136] The aforementioned electronic device can execute the program code for the following steps in the data processing method: performing single-factor attribution verification on the initial analysis results to obtain a first initial verification result, including: calculating the first proportion value of the feature labels in the initial analysis results in the first category of materials and the second proportion value in the second category of materials, and obtaining a first verification sub-result based on the first proportion value and the second proportion value; drawing a box plot on the feature labels in the initial analysis results based on the evaluation values corresponding to the materials in the material set, and obtaining a second verification sub-result based on the box plot; performing significance analysis on the feature labels in the initial analysis results based on the material set to obtain a third verification sub-result; and obtaining the first initial verification result based on the first verification sub-result, the second verification sub-result, and the third verification sub-result.
[0137] The aforementioned electronic device can execute the program code for the following steps in the data processing method: performing full-factor attribution verification on the initial analysis results to obtain a second initial verification result, including: obtaining a prediction model with the feature labels in the initial analysis results as features and the evaluation values corresponding to the materials in the material set as targets; obtaining the content influence value corresponding to the feature labels in the initial analysis results through the prediction model, and obtaining the second initial verification result based on the content influence value.
[0138] Optionally, Figure 6 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 6 As shown, the electronic device 60 may include: one or more ( Figure 6 (Only one is shown in the image) Processor 602 and memory 604. The electronic device 60 may also include a memory controller to control and manage the memory 604; the electronic device 60 may also include a peripheral interface to connect to a radio frequency module, an audio module, and a display screen, etc.
[0139] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the data processing method and apparatus in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned data processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the electronic device 60 via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0140] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: acquire a set of materials and divide the materials in the set into a first category and a second category, wherein the interaction data volume of the first category of materials is higher than that of the second category of materials; perform attribution analysis on the first and second categories of materials using a target model to obtain initial analysis results, wherein the initial analysis results include at least feature labels of the differences between the first and second categories of materials; perform analysis and verification based on the initial analysis results to obtain verification results, and generate target analysis results corresponding to the set of materials based on the verification results and the initial analysis results, wherein the target analysis results are used to provide a reference for generating target materials.
[0141] The processor can invoke information and applications stored in the memory through the transmission device to perform the following steps: dividing the materials in the material set into a first category and a second category, including: acquiring interactive data information of the materials in the material set; calculating based on the interactive data information to obtain the evaluation value corresponding to the materials in the material set; and dividing the materials in the material set into a first category and a second category based on the evaluation value.
[0142] The processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: before performing attribution analysis on the first type of material and the second type of material through the target model, the method further includes: clustering the first type of material to obtain multiple first content type clusters; and clustering the second type of material to obtain multiple second content type clusters.
[0143] The processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: performing attribution analysis on first-type and second-type materials using a target model to obtain initial analysis results, including: performing attribution analysis on multiple first-type materials corresponding to content type clusters in multiple first content type clusters to obtain first analysis results, and performing attribution analysis on multiple second-type materials corresponding to content type clusters in multiple second content type clusters to obtain second analysis results; performing difference analysis on first-type and second-type materials to obtain third analysis results; and performing inductive analysis on the first analysis results, second analysis results, third analysis results, first-type materials, and second-type materials to obtain initial analysis results.
[0144] The processor can invoke information and applications stored in the memory through the transmission device to perform the following steps: perform attribution analysis on multiple first-type materials corresponding to multiple first-type content clusters to obtain a first analysis result, including: for a target content cluster, perform intra-class feature analysis on multiple first-type materials in the target content cluster to obtain a first analysis sub-result; randomly select a preset number of target materials from multiple first-type materials and perform deep analysis on the target materials to obtain a second analysis sub-result; and obtain a first analysis result based on the first analysis sub-result and the analysis sub-result.
[0145] The processor can call the information and application program stored in the memory through the transmission device to perform the following steps: summarizing and analyzing the first analysis result, the second analysis result, the third analysis result, the first type of material and the second type of material to obtain the initial analysis result, including: merging the first type of material and the second type of material to obtain the merged material; and obtaining the initial analysis result based on the merged material, the first analysis result, the second analysis result and the third analysis result.
[0146] The processor can call the information and application program stored in the memory through the transmission device to perform the following steps: perform analysis and verification based on the initial analysis results, and obtain the verification results including: perform single-factor attribution verification on the initial analysis results to obtain a first initial verification result; perform full-factor attribution verification on the initial analysis results to obtain a second initial verification result; and obtain the verification result based on the first initial verification result and the second initial verification result.
[0147] The processor can invoke information and application programs stored in the memory via a transmission device to perform the following steps: performing single-factor attribution validation on the initial analysis results to obtain a first initial validation result, including: calculating the first proportion value of the feature labels in the initial analysis results in the first category of materials and the second proportion value in the second category of materials, and obtaining a first validation sub-result based on the first and second proportion values; drawing a box plot on the feature labels in the initial analysis results based on the evaluation values corresponding to the materials in the material set, and obtaining a second validation sub-result based on the box plot; performing a significance analysis on the feature labels in the initial analysis results based on the material set to obtain a third validation sub-result; and obtaining the first initial validation result based on the first validation sub-result, the second validation sub-result, and the third validation sub-result.
[0148] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: perform full factorial attribution verification on the initial analysis results to obtain a second initial verification result, including: obtaining a prediction model with the feature labels in the initial analysis results as features and the evaluation values corresponding to the materials in the material set as targets; obtaining the content influence value corresponding to the feature labels in the initial analysis results through the prediction model, and obtaining the second initial verification result based on the content influence value.
[0149] Those skilled in the art will understand that Figure 6 The structure shown is for illustrative purposes only. Electronic device 60 can also be a smartphone, tablet, handheld computer, mobile internet device (MID), PAD and other terminal devices. Figure 6 This does not limit the structure of the aforementioned electronic device. For example, electronic device 60 may also include components that are more... Figure 6 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 6 The different configurations shown.
[0150] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0151] Example 4
[0152] Embodiments of this application also provide a computer program product. Optionally, in this embodiment, the computer program product can be used to store the program code executed by the data processing method provided in Embodiment 1.
[0153] Optionally, in this embodiment, the computer program product may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0154] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0155] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0156] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0157] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0158] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0159] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0160] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A data processing method, characterized in that, include: Obtain a set of materials and divide the materials in the set into a first category and a second category, wherein the amount of interactive data of the first category of materials is higher than the amount of interactive data of the second category of materials; Attribution analysis is performed on the first type of material and the second type of material using the target model to obtain initial analysis results, wherein the initial analysis results include at least: feature labels of the differences between the first type of material and the second type of material; The initial analysis results are analyzed and verified to obtain verification results. Based on the verification results and the initial analysis results, the target analysis results corresponding to the material set are generated. The target analysis results are used to provide a reference for generating target materials.
2. The method according to claim 1, characterized in that, Dividing the materials in the aforementioned material set into a first category and a second category includes: Obtain the interactive data information of the materials in the material collection; The evaluation values corresponding to the materials in the material set are obtained by calculating based on the interactive data information. Based on the evaluation values, the materials in the material set are divided into the first type of materials and the second type of materials.
3. The method according to claim 1, characterized in that, Before performing attribution analysis on the first type of material and the second type of material using the target model, the method further includes: Cluster the first type of material to obtain multiple first content type clusters; Clustering the second type of material yields multiple second content type clusters.
4. The method according to claim 3, characterized in that, Attribution analysis was performed on the first type of material and the second type of material using the target model, and the initial analysis results included: Attribution analysis is performed on multiple first-type materials corresponding to the content type clusters in the multiple first content type clusters to obtain a first analysis result, and attribution analysis is performed on multiple second-type materials corresponding to the content type clusters in the multiple second content type clusters to obtain a second analysis result; A difference analysis was performed on the first type of materials and the second type of materials to obtain the third analysis result; The first analysis result, the second analysis result, the third analysis result, the first type of material, and the second type of material are summarized and analyzed to obtain the initial analysis result.
5. The method according to claim 4, characterized in that, Attribution analysis is performed on multiple first-type materials corresponding to the content type clusters in the multiple first content type clusters, and the first analysis results include: For a target content type cluster, intra-class feature analysis is performed on multiple first-type materials in the target content type cluster to obtain a first analysis sub-result; A preset number of target materials are randomly selected from the plurality of first-class materials, and the target materials are subjected to in-depth analysis to obtain a second analysis sub-result; Based on the first analysis sub-result and the analysis sub-result, the first analysis result is obtained.
6. The method according to claim 4, characterized in that, The initial analysis results are obtained by summarizing and analyzing the first analysis result, the second analysis result, the third analysis result, the first type of material, and the second type of material, including: The first type of material and the second type of material are merged to obtain the merged material; The initial analysis result is obtained based on the merged materials, the first analysis result, the second analysis result, and the third analysis result.
7. The method according to claim 1, characterized in that, Based on the initial analysis results, further analysis and verification were performed, and the verification results include: The initial analysis results were subjected to univariate attribution validation to obtain the first initial validation result; The initial analysis results were subjected to full factorial attribution validation to obtain a second initial validation result; The verification result is obtained based on the first initial verification result and the second initial verification result.
8. The method according to claim 7, characterized in that, The initial analysis results were subjected to univariate attribution validation, and the first initial validation results included: Calculate the first proportion value of the feature tags in the initial analysis results in the first type of material and the second proportion value in the second type of material, and obtain the first verification sub-result based on the first proportion value and the second proportion value; Based on the evaluation values corresponding to the materials in the material set, a box plot is drawn for the feature labels in the initial analysis results, and a second verification sub-result is obtained based on the box plot; Based on the aforementioned material set, a saliency analysis is performed on the feature labels in the initial analysis results to obtain a third verification sub-result; Based on the first verification sub-result, the second verification sub-result, and the third verification sub-result, the first initial verification result is obtained.
9. The method according to claim 7, characterized in that, The initial analysis results were subjected to full factorial attribution validation, resulting in a second initial validation result including: Obtain a prediction model that uses the feature labels in the initial analysis results as features and the evaluation values corresponding to the materials in the material set as targets; The prediction model is used to obtain the content influence value corresponding to the feature label in the initial analysis result, and the second initial verification result is obtained based on the content influence value.
10. A data processing method, characterized in that, include: Get the collection of materials uploaded by the client; In the cloud server, the materials in the material set are divided into a first category and a second category, wherein the interaction data volume of the first category is higher than that of the second category. Attribution analysis is performed on the first and second categories of materials using a target model to obtain initial analysis results, wherein the initial analysis results include at least feature labels indicating the differences between the first and second categories of materials. Analysis and verification are performed based on the initial analysis results to obtain verification results. Based on the verification results and the initial analysis results, target analysis results corresponding to the material set are generated, wherein the target analysis results are used to provide a reference for generating target materials. The target analysis results are returned to the client.
11. A data processing apparatus, characterized in that, include: The acquisition unit is used to acquire a set of materials and divide the materials in the set into a first category of materials and a second category of materials, wherein the amount of interactive data of the first category of materials is higher than the amount of interactive data of the second category of materials; The analysis unit is used to perform attribution analysis on the first type of material and the second type of material through the target model to obtain initial analysis results, wherein the initial analysis results include at least: feature labels of the differences between the first type of material and the second type of material; The verification unit is used to perform analysis and verification based on the initial analysis results, obtain verification results, and generate target analysis results corresponding to the material set based on the verification results and the initial analysis results, wherein the target analysis results are used to provide a reference for generating target materials.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the storage medium is located to perform the data processing method according to any one of claims 1 to 10.
13. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the data processing method according to any one of claims 1 to 10.
14. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processor, implement the data processing method according to any one of claims 1 to 10.