Intelligent Analysis System and Method for Media Information Based on Big Data

By integrating and analyzing media data in multiple formats, identifying differences in user group characteristics, evaluating advertising delivery results and adjusting strategies, the problem of incomplete advertising performance evaluation in the existing technology is solved, and more accurate and efficient advertising delivery is achieved.

CN119444325BActive Publication Date: 2025-06-27BEIJING SANRENXING TIMES DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510036980.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-27
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The existing media information analysis system focuses on basic indicators in terms of advertising effectiveness evaluation, such as click-through rate and exposure rate, which cannot fully reflect the advertising effect, limiting the space for optimization of delivery strategies.

Method used

By obtaining media data in multiple formats from different data channels, format identification and integration, extracting user behavior characteristics, clustering and identifying user groups using the K-means algorithm, combining the differences between the characteristics of the target user group and the characteristics of the actual user group, evaluating the effectiveness of the advertising delivery strategy, and calculating the optimization demand coefficient based on the evaluation index, and adjusting the delivery strategy.

Benefits of technology

It has achieved a more comprehensive evaluation of the advertising delivery effect, improved the accuracy and effectiveness of advertising delivery, and ensured that advertising delivery is always in the optimal state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119444325B_ABST
    Figure CN119444325B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent analysis system and method for media information based on big data, relating to the technical field of data analysis. The present invention obtains media data related to target advertisement placement from multiple channels, performs format recognition and processing to form data to be analyzed in a unified format; extracts the placement records of the target advertisement, analyzes its placement strategy, identifies the characteristics of the actual user group, and compares them with the characteristics of the target user group to evaluate the advertisement placement effect; uses the K-means algorithm to perform clustering analysis on user behavior data to identify the actual user group from it, calculates the feature differences and correlations between the two; according to the placement effect evaluation index, analyzes and optimizes the advertisement placement strategy, and proposes adjustment suggestions. The present invention not only improves the efficiency and effect of advertisement placement, but also makes the decision-making process more scientific and reasonable through a data-driven method, can better adapt to market changes, and meet user needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and particularly to an intelligent media information analysis system and method based on big data. Background Art

[0002] In the current era of information explosion, media advertising placement has become an indispensable part of enterprise marketing strategies. With the popularization of the Internet and mobile devices, the channels for advertising placement have become increasingly diversified, including social media, search engines, video platforms, etc. In order to improve the efficiency and effectiveness of advertising placement, enterprises urgently need to use big data technology to intelligently analyze media information. The media data involved in advertising placement comes from a wide range of sources, including online advertisements, social platform interactions, user behavior logs, etc. These data usually exist in different formats and structures, posing challenges to data integration and analysis.

[0003] However, although existing media information analysis systems have basic data collection and processing capabilities, there are still some deficiencies, specifically in terms of advertising placement effect evaluation. For example, existing media information analysis systems focus more on basic indicators such as click-through rate and exposure rate in advertising effect evaluation. Such a single evaluation method cannot comprehensively reflect the advertising effect and limits the optimization space of placement strategies. Summary of the Invention

[0004] The purpose of the present invention is to provide an intelligent media information analysis system and method based on big data to solve the problems raised in the above background art.

[0005] To solve the above technical problems, the present invention provides the following technical solutions:

[0006] An intelligent media information analysis method based on big data, the method comprising the following steps:

[0007] Step S100. Obtain media data related to the target advertising placement in different formats from different data channels, perform format recognition on the obtained media data, and perform corresponding processing on the media data according to the results of the format recognition to integrate the media data into a unified format, thereby obtaining data to be analyzed;

[0008] Step S200. Obtain the placement record corresponding to the target advertisement, extract the placement strategy of the target advertisement from the placement record, and obtain the target user group characteristics according to the placement strategy of the target advertisement; analyze the data to be analyzed to identify the actual user group;

[0009] Step S300. Extract the characteristics of the actual user group and compare and analyze them with the characteristics of the target user group to obtain the differences between the two; Combine the differences between the corresponding characteristics of the target user group and the actual user group to evaluate the delivery effect of the target advertising delivery strategy;

[0010] Step S400. Analyze the optimization demand coefficient of the delivery strategy corresponding to the target advertisement according to the delivery effect evaluation index of the target advertisement; Analyze the corresponding delivery strategy according to the optimization demand coefficient and output the corresponding prompt information based on the analysis result.

[0011] Further, step S100 includes:

[0012] S101. Obtain media data related to the target advertisement delivery in different formats from different data channels. The formats of the media data include text, image, audio, and video formats, and the data format corresponding to each piece of media data is not unique; For example, a social media post may have text descriptions, pictures, audio, and video; For each piece of media data Di, perform format recognition, where i represents the media data number and takes a positive integer; According to the result of format recognition, divide the current media data, so that the media data Di can be expressed as: Di = {d1, d2,..., dn}i, where d1 represents the first data segment of the media data Di, d2 represents the second data segment of the media data Di, and so on. dn represents the nth data segment of the media data Di, where n represents the data format number corresponding to the media data Di and takes values from 1 to 4, representing text, image, audio, and video formats respectively;

[0013] S102. For the data segments in the image, audio, and video formats in each piece of media data Di, convert them all into text format, and replace the corresponding data segments in the image, audio, and video formats with the converted text format data, so as to represent the media data as Di'. Calculate the data proportion b of each data segment of the media data Di', and b = N_dj / N_Di', where N_dj represents the number of bytes of the jth data segment in the media data Di', and N_Di' represents the number of bytes of the media data Di'; Perform semantic analysis on each data segment of the media data Di' to form the corresponding semantic feature vector set Vi, and Vi = {v1, v2,..., vn}i. Similarly, v1 represents the semantic feature vector of the first data segment, v2 represents the semantic feature vector of the second data segment, vn represents the semantic feature vector of the nth data segment, and each semantic feature vector contains multi-dimensional information, such as keywords, themes, etc.;

[0014] S103. Calculate the semantic association index L between each data segment according to the semantic feature vector set Vi of the media data Di'. The specific calculation formula is: Lj k = (vj · vk) / |vj||vk|, where Lj k represents the semantic association index between the j-th data segment and the k-th data segment, vj represents the semantic feature vector corresponding to the j-th data segment, vk represents the semantic feature vector corresponding to the k-th data segment, and both j and k take values from 1 to n, j ≠ k; for each data segment di, calculate the corresponding semantic association strength G by combining the data proportion and the semantic association index, and For each data segment of each media data Di, sort them according to the magnitude relationship of the semantic association strength G. Take the data segment with the largest semantic association strength G as the main data segment, and take the data segments with the semantic association strength G not equal to 0 as the secondary data segments; take the data segments with the semantic association strength G equal to 0 as the irrelevant data segments and filter them out, so as to obtain the media data D; organize the media data D into structured data for storage, so as to obtain the data to be analyzed.

[0015] Further, step S200 includes:

[0016] S201. Obtain the placement records corresponding to the target advertisement from the database, and organize them to obtain the target advertisement placement record data set F, and F{f1, f2,..., fm}, where f1 represents the first placement record of the target advertisement, f2 represents the second placement record of the target advertisement, and so on, fm represents the m-th placement record of the target advertisement; each placement record of the target advertisement corresponds to only one placement strategy, and each placement record contains multiple features, such as placement time, placement platform, placement cost, placement area, user interaction data, etc.; for each placement record fe, extract the corresponding placement strategy from it, and define the extracted placement strategy as the strategy data set Se, so as to obtain the target user group characteristics, and Se = {s1, s2, s3, s4}, where s1 represents the numerical representation of the placement time, s2 represents the geographical code, s3 represents the target user group characteristics, and s4 represents the budget;

[0017] S202. Obtain the user behavior data from the data to be analyzed, extract features from the user behavior data, and construct a feature matrix U. The dimension of the feature matrix U is q × r, where q represents the number of users and r represents the number of features; use the K-means algorithm to cluster the users, so as to identify different user groups, and represent the identified user groups as Pu, where u represents the number of the user group; for each user group, obtain the corresponding number of users, and define the user group with the largest number of users as the actual user group.

[0018] Further, the specific process of using the K-means algorithm to cluster the users is as follows:

[0019] Perform standardization processing on the feature matrix U, select k0 clustering centers, and denote the clustering centers as c. For each user q, calculate the distances from each user q to all clustering centers, and assign it to the nearest cluster. The calculation formula for the distances from each user q to all clustering centers is as follows:

[0020]

[0021] where uh represents the feature vector of the h-th user, cp represents the p-th clustering center, uh_w represents the w-th eigenvalue of the h-th user, and cp_w represents the w-th eigenvalue of the p-th clustering center; after all users are assigned, recalculate the center C of each cluster, and the update formula is:

[0022] Cp = (1 / |Qp|)Σ uh∈Qp uh,

[0023] where Qp represents all user feature vectors assigned to cluster p, and |Qp| is the number of users in cluster p;

[0024] Repeat the above content until the preset number of iterations is reached.

[0025] Furthermore, step S300 includes:

[0026] S301. Extract the actual user group features and the target user extraction features, perform standardization processing on the actual user group features and the target user group features, perform correlation analysis on the actual user group features and the target user group features respectively, calculate the Pearson correlation coefficient ρ between the actual user group features and the target user group features and the preset advertising placement effect index respectively, compare the absolute value of the calculated Pearson correlation coefficient ρ with the threshold ρ0, and screen out the actual user group features and the target user group features that are greater than the absolute value of the threshold ρ0;

[0027] S302. Calculate the difference metric index CY between the screened actual user group features and the target user group features. The specific calculation formula is:

[0028]

[0029] where x_b and y_b respectively represent the feature values of the b-th actual user group and the target user group, and g represents the maximum value of the number of the screened actual user group features and the target user group features;

[0030] S303. Evaluate the delivery effect of the target advertising delivery strategy by combining the difference index between the corresponding characteristics of the target user group and the actual user group, and calculate the delivery effect evaluation index TF of the target advertising delivery strategy. The specific calculation formula is: TF = CY × R, where R represents the original advertising effect evaluation function, which can be indicators such as click-through rate and conversion rate, depending on the specific situation.

[0031] Further, step S400 includes:

[0032] S401. Summarize the delivery effect evaluation indexes TF corresponding to all delivery strategies of the target advertisement to form a delivery effect evaluation index set M, and M = {TF1, TF2,...., TFm}, where TF1 represents the delivery effect evaluation index corresponding to the first delivery strategy of the target advertisement, TF2 represents the delivery effect evaluation index corresponding to the second delivery strategy of the target advertisement, and so on. TFm represents the delivery effect evaluation index corresponding to the mth delivery strategy of the target advertisement; According to the delivery effect evaluation index set M, calculate the average value TF0 and standard deviation σ of the delivery effect evaluation indexes of all delivery strategies. For each delivery strategy of the target advertisement, calculate the corresponding optimization requirement coefficient O. The specific calculation formula is: O = |TF - TF0| / σ;

[0033] S402. Compare the size relationship between the optimization requirement coefficient O corresponding to each delivery strategy and the threshold T. When O > T, mark the corresponding delivery strategy as the delivery strategy to be adjusted, and output the number corresponding to the delivery strategy to be adjusted to the relevant personnel for corresponding adjustment; When O ≤ T, mark the corresponding delivery strategy as the target delivery strategy for storage.

[0034] An intelligent analysis system for media information based on big data. The system includes: a data collection and processing module, a user group characteristic analysis module, a characteristic comparison and difference analysis module, a delivery effect evaluation module, and an optimization suggestion module;

[0035] The data collection and processing module obtains media data related to target advertisement placement from different data channels, identifies and processes the formats of the obtained media data, and integrates them into a unified format; the user group characteristic analysis module obtains the placement records of target advertisements, extracts placement strategies and characteristics, and forms a target user group characteristic data set; extracts user behavior characteristics from the data to be analyzed, constructs a characteristic matrix, uses the K-means algorithm to cluster users, identifies the actual user group and extracts its characteristics; the feature comparison and difference analysis module standardizes the characteristics of the actual user group and the target user group, calculates the correlation between the characteristics of the actual user group and the target user group, and screens out important features; calculates the difference measure between the screened features, and evaluates the feature differences between the actual user group and the target user group; the placement effect evaluation module calculates the placement effect evaluation index according to the feature difference index and the original advertisement effect evaluation function, summarizes the effect evaluation indexes of all placement strategies, calculates the average value and standard deviation; and calculates the optimization requirement coefficient of each placement strategy to evaluate the optimization requirement of the placement strategy; the optimization suggestion module compares the optimization requirement coefficient with the set threshold, marks the placement strategies to be adjusted and the target placement strategies, outputs the numbers of the strategies to be adjusted to relevant personnel, and stores the target placement strategies.

[0036] Furthermore, the data collection and processing module includes a data acquisition unit, a data format identification unit, and a data integration and storage unit;

[0037] The data acquisition unit is responsible for obtaining media data related to target advertisement placement from different data channels, including text, image, audio, and video formats; the data format identification unit identifies the formats of the obtained media data, analyzes the types of each data segment, and converts them into a unified format; the data integration and storage unit organizes the processed media data into structured data, stores it, and forms a data set to be analyzed.

[0038] Furthermore, the user group characteristic analysis module includes a placement record extraction unit and a user behavior data analysis unit;

[0039] The placement record extraction unit obtains the placement records of target advertisements from the database, extracts the placement strategies and the characteristics of the target user group, and forms a target characteristic data set; the user behavior data analysis unit analyzes the user behavior data in the data to be analyzed, constructs a characteristic matrix, uses the K-means algorithm to cluster users, and identifies the actual user group;

[0040] The feature comparison and difference analysis module includes a correlation analysis unit and a difference measure calculation unit;

[0041] The correlation analysis unit standardizes the characteristics of the actual user group and the target user group, and calculates the Pearson correlation coefficient between them and the advertising placement effect indicators to screen out important characteristics; the difference metric calculation unit calculates the difference metric indicators between the characteristics of the actual user group and the target user group after screening.

[0042] Furthermore, the placement effect evaluation module includes an effect evaluation index calculation unit and an optimization requirement analysis unit;

[0043] The effect evaluation index calculation unit calculates the corresponding effect evaluation index according to the difference metric indicators, summarizes the effect evaluation indexes of each placement strategy, and calculates the average value and standard deviation; the optimization requirement analysis unit calculates the optimization requirement coefficient of each placement strategy according to the comparison between the effect evaluation index and the threshold, and marks the placement strategies that need to be adjusted:

[0044] The optimization suggestion module includes an adjustment strategy suggestion unit and a strategy storage unit;

[0045] The adjustment strategy suggestion unit outputs the label of the placement strategy to be adjusted to the relevant personnel according to the result of the optimization requirement analysis; the strategy storage unit stores the records marked as the target placement strategy.

[0046] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention can process media data in various formats (text, image, audio, and video) and integrate them into a unified format. This multi-dimensional data integration can more comprehensively capture the context information of advertising placement, enhancing the depth and breadth of analysis; by performing semantic analysis on the media data and calculating the semantic association strength, the present invention can identify the semantic connections between data segments, thereby better understanding the true reactions of users to the advertising content. This semantic association analysis can provide empirical evidence for optimizing the advertising content; using the K-means algorithm to cluster users can effectively identify different user groups, enabling the advertising placement strategy to be more targeted to match the actual users, thereby improving the precise placement effect of the advertisement; the present invention can identify user behavior differences by extracting the characteristics of the actual user group and comparing them with the characteristics of the target user group, providing a more comprehensive evaluation of the advertising effect; the present invention provides a scientific evaluation model by calculating the difference metric indicators and the placement effect evaluation index, which can clearly reflect the effects of each placement strategy and provide a basis for subsequent adjustment of the placement strategy; by calculating the optimization requirement coefficient, the placement strategies that need to be adjusted can be identified in a timely manner to ensure that the advertising placement always remains in the optimal state. This dynamic adjustment mechanism enhances the flexibility and effectiveness of advertising placement. Description of the Drawings

[0047] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings:

[0048] Figure 1 It is a schematic diagram of the modules of the intelligent analysis system for media information based on big data of the present invention. Detailed implementation manners

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] Please refer to Figure 1 , the present invention provides a technical solution:

[0051] An intelligent analysis system for media information based on big data, the system includes: a data collection and processing module, a user group feature analysis module, a feature comparison and difference analysis module, a placement effect evaluation module, and an optimization suggestion module;

[0052] The data collection and processing module obtains media data related to the target advertisement placement from different data channels, performs format recognition and processing on the obtained media data, and integrates it into a unified format; the user group feature analysis module obtains the placement records of the target advertisement, extracts the placement strategies and features, and forms a target user group feature data set; extracts user behavior features from the data to be analyzed, constructs a feature matrix, uses the K-means algorithm to cluster users, identifies the actual user group and extracts features; the feature comparison and difference analysis module performs standardization processing on the actual user group features and the target user group features, calculates the correlation between the actual user group features and the target user group features, and screens out important features; calculates the difference metric between the screened features, and evaluates the feature differences between the actual user group and the target user group; the placement effect evaluation module calculates the placement effect evaluation index according to the feature difference index and the original advertisement effect evaluation function, summarizes the effect evaluation indexes of all placement strategies, calculates the average value and the standard deviation; and calculates the optimization requirement coefficient of each placement strategy, and evaluates the optimization requirement of the placement strategy; the optimization suggestion module compares the optimization requirement coefficient with the set threshold, marks the placement strategies to be adjusted and the target placement strategies, outputs the numbers of the strategies to be adjusted to relevant personnel, and stores the target placement strategies.

[0053] The data collection and processing module includes a data acquisition unit, a data format recognition unit, and a data integration and storage unit;

[0054] The data acquisition unit is responsible for obtaining media data related to target advertisement placement from different data channels, including text, image, audio, and video formats; the data format recognition unit performs format recognition on the obtained media data, analyzes the types of each data segment, and converts them into a unified format; the data integration and storage unit organizes the processed media data into structured data, stores it, and forms a dataset to be analyzed.

[0055] The user group characteristic analysis module includes a placement record extraction unit and a user behavior data analysis unit;

[0056] The placement record extraction unit obtains the placement records of the target advertisement from the database, extracts the placement strategy and the characteristics of the target user group, and forms a target characteristic dataset; the user behavior data analysis unit analyzes the user behavior data in the data to be analyzed, constructs a characteristic matrix, uses the K-means algorithm to cluster users, and identifies the actual user group;

[0057] The feature comparison and difference analysis module includes a correlation analysis unit and a difference metric calculation unit;

[0058] The correlation analysis unit performs standardization processing on the characteristics of the actual user group and the target user group, and calculates the Pearson correlation coefficient between them and the advertisement placement effect indicators to screen out important features; the difference metric calculation unit calculates the difference metric indicators between the characteristics of the actual user group and the target user group after screening.

[0059] The placement effect evaluation module includes an effect evaluation index calculation unit and an optimization requirement analysis unit;

[0060] The effect evaluation index calculation unit calculates the corresponding effect evaluation index according to the difference metric indicators, summarizes the effect evaluation indexes of each placement strategy, and calculates the average value and standard deviation; the optimization requirement analysis unit calculates the optimization requirement coefficient of each placement strategy according to the comparison between the effect evaluation index and the threshold, and marks the placement strategies that need to be adjusted:

[0061] The optimization suggestion module includes an adjustment strategy suggestion unit and a strategy storage unit;

[0062] The adjustment strategy suggestion unit outputs the label of the placement strategy to be adjusted to the relevant personnel according to the optimization requirement analysis result; the strategy storage unit stores the records marked as the target placement strategy.

[0063] An intelligent analysis method for media information based on big data, the method includes the following steps:

[0064] Step S100. Obtain media data related to target advertisement placement in different formats from different data channels, perform format recognition on the obtained media data, and process the media data accordingly according to the results of format recognition, and integrate the media data into a unified format to obtain the data to be analyzed;

[0065] Step S200. Obtain the placement records corresponding to the target advertisement, extract the placement strategy of the target advertisement from the placement records, and obtain the characteristics of the target user group according to the placement strategy of the target advertisement; analyze the data to be analyzed to identify the actual user group;

[0066] Step S300. Extract the characteristics of the actual user group and compare and analyze them with the characteristics of the target user group to obtain the differences between the two; combine the differences between the corresponding characteristics of the target user group and the actual user group to evaluate the placement effect of the target advertisement placement strategy;

[0067] Step S400. Analyze the optimization requirement coefficient of the placement strategy corresponding to the target advertisement according to the placement effect evaluation index of the target advertisement; analyze the corresponding placement strategy according to the optimization requirement coefficient, and output the corresponding prompt information based on the analysis results.

[0068] Step S100 includes:

[0069] S101. Obtain media data related to target advertisement placement in different formats from different data channels. The formats of the media data include text, image, audio, and video formats, and the data format corresponding to each piece of media data is not unique; for example, a social media post may have text description, pictures, audio, and video; for each piece of media data Di, where i represents the media data number and takes a positive integer, perform format recognition; according to the results of format recognition, divide the current media data, so that the media data Di can be expressed as: Di = {d1, d2,..., dn}i, where d1 represents the first data segment of the media data Di, d2 represents the second data segment of the media data Di, and so on, dn represents the nth data segment of the media data Di, where n represents the data format number corresponding to the media data Di and takes values from 1 to 4, representing text, image, audio, and video formats respectively;

[0070] S102. For each of the image, audio, and video format data segments in the media data Di, convert them into text format, and replace the corresponding image, audio, and video format data segments with the converted text format data, so as to represent the media data as Di'. Calculate the data proportion b of each data segment in the media data Di', and b = N_dj / N_Di', where N_dj represents the number of bytes of the j-th data segment in the media data Di', and N_Di' represents the number of bytes of the media data Di'; perform semantic analysis on each data segment of the media data Di' to form the corresponding semantic feature vector set Vi, and Vi = {v1, v2,..., vn}i. Similarly, v1 represents the semantic feature vector of the first data segment, v2 represents the semantic feature vector of the second data segment, vn represents the semantic feature vector of the n-th data segment, and each semantic feature vector contains multi-dimensional information, such as keywords, topics, etc.;

[0071] S103. According to the semantic feature vector set Vi of the media data Di', calculate the semantic association index L between each data segment. The specific calculation formula is: Lj k = (vj · vk) / |vj||vk|, where Lj k represents the semantic association index between the j-th data segment and the k-th data segment, vj represents the semantic feature vector corresponding to the j-th data segment, vk represents the semantic feature vector corresponding to the k-th data segment, and both j and k take values from 1 to n, j ≠ k; for each data segment di, combine the data proportion and the semantic association index to calculate the corresponding semantic association strength G, and For each data segment of each media data Di, sort them according to the magnitude relationship of the semantic association strength G. Take the data segment with the largest semantic association strength G as the main data segment, and take the data segments with the semantic association strength G not equal to 0 as the secondary data segments; take the data segments with the semantic association strength G equal to 0 as the irrelevant data segments and perform screening to obtain the media data D; organize the media data D into structured data for storage to obtain the data to be analyzed.

[0072] In this embodiment, an advertising company hopes to analyze the advertising effect on the social media platform, and the obtained data includes text posts, pictures, audio comments, and video ads.

[0073] Scrape advertising-related data from multiple social media channels (such as Twitter, Instagram, Facebook). The format of each media data Di may include:

[0074] Text: Descriptive text, Image: Advertising picture, Audio: User comment recording, Video: Advertising video clip;

[0075] Analyze each piece of data through a format recognition tool to identify its respective data segments. For example: D1 = {d1, d2, d3}1;

[0076] Convert the image and audio data segments into text format to obtain D1', and convert d2 into descriptive text and d3 into transcribed text;

[0077] Calculate the number of bytes of each data segment and calculate the proportion b of each data segment:

[0078] For example, b1 = N_d1 / N_D1';

[0079] Perform semantic analysis on each data segment to generate a feature vector set V1, and V1 = {v1, v2, v3}1;

[0080] Calculate the semantic association index L. For example: L12 = (v1 · v2) / |v1||v2|;

[0081] Calculate the semantic association strength G:

[0082] Combine the proportion and the association index to calculate the strength of each data segment. For example:

[0083] G1 = b1 × (∣L11∣ + ∣L12∣ + ∣L13∣).

[0084] Sort all data segments according to the G value: Determine the main data segments, secondary data segments, and irrelevant data segments.

[0085] Organize the main and secondary data segments into structured data and save it as a database table or a CSV file to obtain the data to be analyzed.

[0086] Step S200 includes:

[0087] S201. Obtain the placement records corresponding to the target advertisement from the database, organize them to obtain the target advertisement placement record data set F, and F{f1, f2,..., fm}, where f1 represents the first placement record of the target advertisement, f2 represents the second placement record of the target advertisement, and so on, fm represents the mth placement record of the target advertisement; each placement record of the target advertisement corresponds to only one placement strategy, and each placement record contains multiple features, such as placement time, placement platform, placement cost, placement area, user interaction data, etc.; for each placement record fe, extract the corresponding placement strategy from it, and define the extracted placement strategy as the strategy data set Se, so as to obtain the target user group characteristics, and Se = {s1, s2, s3, s4}, where s1 represents the numerical representation of the placement time, s2 represents the geographical code, s3 represents the target user group characteristics, and s4 represents the budget;

[0088] S202. Obtain user behavior data from the data to be analyzed, and extract features from the user behavior data to construct a feature matrix U, and the dimension of the feature matrix U is q×r, where q represents the number of users and r represents the number of features; use the K-means algorithm to cluster users to identify different user groups, and represent the identified user groups as Pu, where u represents the number of the user group; for each user group, obtain the corresponding number of users, and define the user group with the largest number of users as the actual user group.

[0089] The specific process of clustering users using the K-means algorithm is as follows:

[0090] The feature matrix U is standardized, k0 cluster centers are selected, and the cluster center is recorded as c. For each user q, the distance from each user q to all cluster centers is calculated, and it is assigned to the nearest cluster. The calculation formula for the distance from each user q to all cluster centers is:

[0091]

[0092] Among them, uh represents the feature vector of the h-th user, cp represents the p-th cluster center, uh_w represents the w-th feature value of the h-th user, and cp_w represents the w-th feature value of the p-th cluster center. After all users are assigned, the center C of each cluster is recalculated, and the update formula is:

[0093] Cp=(1 / |Qp|)Σ uh∈Qp uh,

[0094] Where Qp represents the feature vectors of all users assigned to cluster p, |Qp| is the number of users in cluster p;

[0095] Repeat the above until the preset number of iterations is reached.

[0096] Step S300 includes:

[0097] S301. extracting actual user group characteristics and target user group characteristics, standardizing the actual user group characteristics and target user group characteristics, performing correlation analysis on the actual user group characteristics and target user group characteristics, respectively calculating the Pearson correlation coefficient ρ between the actual user group characteristics and the target user group characteristics and the preset advertising delivery effect indicators, comparing the calculated Pearson correlation coefficient ρ with the absolute value of the threshold ρ0, and screening out the actual user group characteristics and the target user group characteristics whose absolute values ​​are greater than the threshold ρ0;

[0098] S302. Calculate the difference measurement index CY between the actual user group characteristics after screening and the target user group characteristics. The specific calculation formula is:

[0099]

[0100] Where x_b and y_b represent the feature values ​​of the bth actual user group and the target user group respectively, and g represents the maximum number of the actual user group features and the target user group features after screening;

[0101] S303. Based on the difference indicators between the corresponding characteristics of the target user group and the actual user group, the delivery effect of the target advertising delivery strategy is evaluated, and the delivery effect evaluation index TF of the target advertising delivery strategy is calculated. The specific calculation formula is: TF=CY×R, where R represents the original advertising effect evaluation function, which can be indicators such as click-through rate and conversion rate, depending on the specific situation.

[0102] In this embodiment, it is assumed that the original advertising effect evaluation function is composed of the following indicators: click-through rate CTR, conversion rate CVR and return on investment ROI. According to the above indicators, a weighted comprehensive effect evaluation function R is defined, and R=w1×CTR+w2×CVR+w3×ROI, where w1, w2 and w3 represent the weights of each indicator. Therefore, the calculation formula of the delivery effect evaluation index TF of the target advertising delivery strategy is: TF=CY×R=CY×(w1×CTR+w2×CVR+w3×ROI).

[0103] Step S400 includes:

[0104] S401. Summarize the delivery effect evaluation index TF corresponding to all delivery strategies of the target advertisement to form a delivery effect evaluation index set M, where M = {TF1, TF2, ...., TFm}, where TF1 represents the delivery effect evaluation index corresponding to the first delivery strategy of the target advertisement, TF2 represents the delivery effect evaluation index corresponding to the second delivery strategy of the target advertisement, and so on, TFm represents the delivery effect evaluation index corresponding to the mth delivery strategy of the target advertisement; according to the delivery effect evaluation index set M, calculate the average value TF0 and standard deviation σ of the delivery effect evaluation index of all delivery strategies, and for each delivery strategy of the target advertisement, calculate the corresponding optimization demand coefficient O, and the specific calculation formula is: O = |TF-TF0| / σ;

[0105] S402. Compare the size relationship between the optimization demand coefficient O corresponding to each delivery strategy and the threshold T. When O>T, mark the corresponding delivery strategy as the delivery strategy to be adjusted, and output the number corresponding to the delivery strategy to be adjusted to the relevant personnel, who will make corresponding adjustments; when O≤T, mark the corresponding delivery strategy as the target delivery strategy for storage.

[0106] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0107] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent analysis method of media information based on big data, characterized by: The method comprises the following steps: Step S100. Acquire media data related to target advertisement delivery in different formats from different data channels, perform format recognition on the acquired media data, and process the media data accordingly according to the result of the format recognition, so as to integrate the media data into a unified format, thereby obtaining data to be analyzed; Step S200. Obtain the delivery record corresponding to the target advertisement, extract the delivery strategy of the target advertisement from the delivery record, and obtain the target user group characteristics according to the delivery strategy of the target advertisement; analyze the data to be analyzed, thereby identifying the actual user group; Step S300: extract the characteristics of the actual user group and compare and analyze them with the characteristics of the target user group to obtain the difference between the two; combine the differences between the corresponding characteristics of the target user group and the actual user group to evaluate the delivery effect of the target advertising delivery strategy; Step S400. Analyze the optimization demand coefficient of the delivery strategy corresponding to the target advertisement according to the delivery effect evaluation index of the target advertisement; analyze the corresponding delivery strategy according to the optimization demand coefficient, and output corresponding prompt information based on the analysis result.

2. The method for intelligent analysis of media information based on big data according to claim 1, characterized in that: The step S100 includes: S101. Obtain media data related to target advertising in different formats from different data channels, wherein the formats of the media data include text, image, audio and video formats, and the data format corresponding to each piece of media data is not unique; perform format identification on each piece of media data Di, wherein i represents the media data number, which is a positive integer; divide the current media data according to the result of the format identification, so that the media data Di is represented as: Di = {d1, d2, ..., dn}i, wherein d1 represents the first data segment of the media data Di, d2 represents the second data segment of the media data Di, and so on, dn represents the nth data segment of the media data Di, wherein n represents the data format number corresponding to the media data Di, which is 1 to 4, representing text, image, audio and video formats respectively; S102. For each data segment in the image, audio and video format in the media data Di, convert them into text format, replace the corresponding data segment in the image, audio and video format with the converted text format data, so as to represent the media data as Di', calculate the data proportion b of each data segment of the media data Di', and b = N_dj / N_Di', wherein N_dj represents the number of bytes of the j-th data segment in the media data Di', and N_Di' represents the number of bytes of the media data Di'; perform semantic analysis on each data segment of the media data Di', so as to form a corresponding semantic feature vector set Vi, and Vi = {v1, v2, ..., vn}i, similarly, v1 represents the semantic feature vector of the first data segment, v2 represents the semantic feature vector of the second data segment, vn represents the semantic feature vector of the n-th data segment, and each semantic feature vector contains multi-dimensional information; S103. According to the semantic feature vector set Vi of the media data Di', the semantic association index L between each data segment is calculated. The specific calculation formula is: Ljk = (vj·vk) / |vj||vk|, where Ljk represents the semantic association index between the j-th data segment and the k-th data segment, vj represents the semantic feature vector corresponding to the j-th data segment, vk represents the semantic feature vector corresponding to the k-th data segment, and j and k are both 1 to n, j≠k; for each data segment di, the corresponding semantic association strength G is calculated in combination with the data proportion and the semantic association index, and For each data segment of media data Di, sort them according to the size relationship of semantic association strength G, take the data segment with the largest semantic association strength G as the main data segment, and take the data segment with semantic association strength G not equal to 0 as the secondary data segment; take the data segment with semantic association strength G equal to 0 as irrelevant data segment and filter it out to obtain media data D; organize the media data D into structured data for storage to obtain the data to be analyzed.

3. The method for intelligent analysis of media information based on big data according to claim 1, characterized in that: The step S200 includes: S201. Obtain the delivery records corresponding to the target advertisement from the database, and sort out the target advertisement delivery record data set F, and F{f1,f2,...,fm}, where f1 represents the first delivery record of the target advertisement, f2 represents the second delivery record of the target advertisement, and so on, fm represents the mth delivery record of the target advertisement; each delivery record of the target advertisement corresponds to only one delivery strategy, and each delivery record contains multiple features; for each delivery record fe, extract the corresponding delivery strategy therefrom, and define the extracted delivery strategy as the strategy data set Se, so as to obtain the target user group characteristics, and Se={s1,s2,s3,s4}, where s1 represents the numerical representation of the delivery time, s2 represents the regional code, s3 represents the target user group characteristics, and s4 represents the budget; S202. Obtain user behavior data from the data to be analyzed, and extract features from the user behavior data to construct a feature matrix U, and the dimension of the feature matrix U is q×r, where q represents the number of users and r represents the number of features; use the K-means algorithm to cluster users to identify different user groups, and represent the identified user groups as Pu, where u represents the number of the user group; for each user group, obtain the corresponding number of users, and define the user group with the largest number of users as the actual user group.

4. The method for intelligent analysis of media information based on big data according to claim 3 is characterized in that: The specific process of clustering users using the K-means algorithm is as follows: The feature matrix U is standardized, k0 cluster centers are selected, and the cluster center is recorded as c. For each user q, the distance from each user q to all cluster centers is calculated, and it is assigned to the nearest cluster. The calculation formula for the distance from each user q to all cluster centers is: Among them, uh represents the feature vector of the h-th user, cp represents the p-th cluster center, uh_w represents the w-th feature value of the h-th user, and cp_w represents the w-th feature value of the p-th cluster center. After all users are assigned, the center C of each cluster is recalculated, and the update formula is: Cp=(1 / |Qp|)Σ uh∈Qp uh, Where Qp represents the feature vectors of all users assigned to cluster p, |Qp| is the number of users in cluster p; Repeat the above until the preset number of iterations is reached.

5. The method for intelligent analysis of media information based on big data according to claim 1, characterized in that: The step S300 includes: S301. extracting actual user group characteristics and target user group characteristics, standardizing the actual user group characteristics and target user group characteristics, performing correlation analysis on the actual user group characteristics and target user group characteristics, respectively calculating the Pearson correlation coefficient ρ between the actual user group characteristics and the target user group characteristics and the preset advertising delivery effect indicators, comparing the calculated Pearson correlation coefficient ρ with the absolute value of the threshold ρ0, and screening out the actual user group characteristics and the target user group characteristics whose absolute values ​​are greater than the threshold ρ0; S302. Calculate the difference measurement index CY between the actual user group characteristics after screening and the target user group characteristics. The specific calculation formula is: Where x_b and y_b represent the feature values ​​of the bth actual user group and the target user group respectively, and g represents the maximum number of the actual user group features and the target user group features after screening; S303. Based on the difference indicators between the corresponding characteristics of the target user group and the actual user group, the delivery effect of the target advertising delivery strategy is evaluated, and the delivery effect evaluation index TF of the target advertising delivery strategy is calculated. The specific calculation formula is: TF=CY×R, where R represents the original advertising effect evaluation function.

6. The method for intelligent analysis of media information based on big data according to claim 5, characterized in that: The step S400 includes: S401. Summarize the delivery effect evaluation index TF corresponding to all delivery strategies of the target advertisement to form a delivery effect evaluation index set M, where M = {TF1, TF2, ...., TFm}, where TF1 represents the delivery effect evaluation index corresponding to the first delivery strategy of the target advertisement, TF2 represents the delivery effect evaluation index corresponding to the second delivery strategy of the target advertisement, and so on, TFm represents the delivery effect evaluation index corresponding to the mth delivery strategy of the target advertisement; according to the delivery effect evaluation index set M, calculate the average value TF0 and standard deviation σ of the delivery effect evaluation index of all delivery strategies, and for each delivery strategy of the target advertisement, calculate the corresponding optimization demand coefficient O, and the specific calculation formula is: O = |TF-TF0| / σ; S402. Compare the size relationship between the optimization demand coefficient O corresponding to each delivery strategy and the threshold T. When O>T, mark the corresponding delivery strategy as the delivery strategy to be adjusted, and output the number corresponding to the delivery strategy to be adjusted to the relevant personnel, who will make corresponding adjustments; when O≤T, mark the corresponding delivery strategy as the target delivery strategy for storage.

7. A media information intelligent analysis system based on big data, applied to the media information intelligent analysis method based on big data according to any one of claims 1 to 6, characterized in that: The system includes: a data collection and processing module, a user group feature analysis module, a feature comparison and difference analysis module, an advertising effect evaluation module, and an optimization suggestion module; The data acquisition and processing module obtains media data related to the target advertisement delivery from different data channels, performs format recognition and processing on the obtained media data, and integrates them into a unified format; the user group feature analysis module obtains the delivery record of the target advertisement, extracts the delivery strategy and features, and forms a target user group feature data set; extracts user behavior features from the data to be analyzed, constructs a feature matrix, clusters users using the K-means algorithm, identifies the actual user group and performs feature extraction; the feature comparison and difference analysis module performs standardization processing on the actual user group features and the target user group features, calculates the correlation between the actual user group features and the target user group features, and screens out important features; calculates the difference measure between the screened features, and evaluates the feature difference between the actual user group and the target user group; the delivery effect evaluation module calculates the delivery effect evaluation index according to the feature difference index and the original advertisement effect evaluation function, summarizes the effect evaluation index of all delivery strategies, calculates the average value and standard deviation; and calculates the optimization demand coefficient of each delivery strategy, and evaluates the optimization demand of the delivery strategy; the optimization suggestion module compares the optimization demand coefficient with the set threshold, marks the delivery strategy to be adjusted and the target delivery strategy, outputs the number of the strategy to be adjusted to the relevant personnel, and stores the target delivery strategy.

8. The intelligent media information analysis system based on big data according to claim 7 is characterized by: The data acquisition and processing module includes a data acquisition unit, a data format recognition unit, and a data integration and storage unit; The data acquisition unit is responsible for acquiring media data related to target advertising from different data channels, including text, image, audio and video formats; the data format identification unit identifies the format of the acquired media data, analyzes the type of each data segment, and converts it into a unified format; the data integration and storage unit organizes the processed media data into structured data, stores it, and forms a data set to be analyzed.

9. The intelligent media information analysis system based on big data according to claim 7, characterized in that: The user group feature analysis module includes a delivery record extraction unit and a user behavior data analysis unit; The delivery record extraction unit obtains the delivery record of the target advertisement from the database, extracts the delivery strategy and the target user group characteristics, and forms a target characteristic data set; the user behavior data analysis unit analyzes the user behavior data in the data to be analyzed, constructs a characteristic matrix, and uses the K-means algorithm to cluster the users to identify the actual user group; The feature comparison and difference analysis module includes a correlation analysis unit and a difference metric calculation unit; The correlation analysis unit normalizes the actual user group characteristics and the target user group characteristics, and calculates the Pearson correlation coefficient between them and the advertising delivery effect index to screen out important characteristics; the difference measurement calculation unit calculates the difference measurement index between the screened actual user group characteristics and the target user group characteristics.

10. The media information intelligent analysis system based on big data according to claim 7, characterized in that: The delivery effect evaluation module includes an effect evaluation index calculation unit and an optimization demand analysis unit; The effect evaluation index calculation unit calculates the corresponding effect evaluation index according to the difference measurement index, summarizes the effect evaluation index of each delivery strategy, and calculates the average value and standard deviation; the optimization demand analysis unit calculates the optimization demand coefficient of each delivery strategy according to the comparison between the effect evaluation index and the threshold value, and marks the delivery strategy that needs to be adjusted: The optimization suggestion module includes an adjustment strategy suggestion unit and a strategy storage unit; The adjustment strategy suggestion unit outputs the label of the delivery strategy to be adjusted to relevant personnel according to the optimization demand analysis result; and the strategy storage unit stores the record marked as the target delivery strategy.

Citation Information

Patent Citations

  • Delivery object screening method and device, equipment and storage medium

    CN117009631A

  • Analysis method and system for realizing advertisement putting effect based on clustering algorithm

    CN118429017A