New media advertisement putting method and system based on big data

Through big data analysis and similarity matching, the new media advertising delivery system has solved the problems of platform fragmentation and user resistance, achieving precise targeting and efficient conversion, thereby improving advertising effectiveness and user experience.

CN121921064APending Publication Date: 2026-04-24XIAMEN SOFTWARE VOCATIONAL & TECH COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN SOFTWARE VOCATIONAL & TECH COLLEGE
Filing Date
2026-01-21
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

New media advertising platforms are fragmented and complex to manage, resulting in high advertising costs. User resistance and blocking behaviors also negatively impact the effectiveness of advertising.

Method used

The big data-driven new media advertising system extracts and categorizes page and ad feature data through content page analysis, user preference analysis, and ad content analysis modules. It uses similarity matching algorithms to ensure that ads are highly relevant to page content and combines user browsing history for precise targeting.

Benefits of technology

It improved the accuracy and conversion rate of ad placement, enhanced user experience, increased the targeting and user acceptance of ads, and achieved precision marketing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921064A_ABST
    Figure CN121921064A_ABST
Patent Text Reader

Abstract

The invention provides a new media advertisement putting method and system based on big data, and the system is composed of four core modules: a content page analysis module, a user preference analysis module, an advertisement content analysis module and an advertisement putting control module, and comprises an advertisement putting strategy: firstly analyzing a preset content page through a content feature extraction algorithm, and generating page feature data; then classifying the page feature data and determining category features by using a feature classification method; meanwhile, the system also performs similar feature extraction and classification processing on the advertisement content to generate advertisement feature data and category feature data. And finally, the system adopts a similarity matching algorithm to match the page category features with the advertisement category features, and finds out the most suitable advertisement content to be put. The method has the advantages that reasonable matching of the advertisement content and the page content can be achieved, the advertisement putting accuracy and effect are improved, the advertisement putting efficiency is improved, and the manual operation cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of advertising, and in particular to a new media advertising delivery method and system based on big data. Background Technology

[0002] With the rapid development of internet technology and mobile devices, new media advertising has become an important part of modern marketing communication. New media advertising refers to advertising through digital platforms such as the internet and mobile devices. It breaks through the time and space limitations of traditional advertising, enabling precise targeting and real-time interaction.

[0003] Currently, mainstream new media advertising channels include social media platforms, search engines, short video platforms, and mobile applications. On social media platforms, advertisers can reach their target audience through in-feed ads, WeChat Moments ads, and other formats. Search engine advertising utilizes keyword targeting, displaying ads when users actively search for relevant content. Splash screen ads and in-feed ads on short video platforms are also becoming increasingly popular. Furthermore, programmatic advertising, through big data analytics and artificial intelligence, has achieved automation and intelligence in ad delivery.

[0004] New media advertising has significant advantages, including precise audience targeting, delivering ads based on users' age, gender, interests, and other attributes to improve advertising effectiveness.

[0005] However, new media advertising also faces some challenges. The numerous and fragmented advertising platforms significantly increase the complexity and management costs of ad placement. User resistance to advertising and ad-blocking behavior also affect ad performance. Summary of the Invention

[0006] This application provides a new media advertising delivery system based on big data, including: Content page analysis module; User preference analysis module; Advertising content analysis module; and Advertising delivery control module; The content page analysis module, the user preference analysis module, and the advertising content analysis module are all connected to the advertising delivery control module. The advertising delivery control module includes an advertising delivery strategy, which includes the following steps: A1. Based on each preset content page, the content page analysis module generates page feature data for each content page using a preset content feature extraction algorithm. A2, Generate page classification data and determine the corresponding page category feature data based on the page feature data of all content pages using a preset feature classification method; A3, based on each preset advertisement content, the advertisement content analysis module generates advertisement feature data for each advertisement content using a content feature extraction algorithm; A4. Based on the advertising feature data of all advertising content, generate advertising classification data and determine the corresponding advertising category feature data using the feature classification method. A5. Based on the page category feature data corresponding to the content page and the ad category feature data corresponding to each ad category, a preset similarity matching algorithm is used to determine the approximate ad category data. A6, through the advertising delivery control module, delivers various advertising contents from the approximate advertising category data to the content page.

[0007] By adopting the above technical solution, the big data-based new media advertising delivery system can extract and classify page feature data for ads that need to be delivered, and extract and classify ad feature data for content to be delivered, and ensure that the delivered ads are highly relevant to the page content through similarity matching. This not only improves the accuracy and conversion effect of ad delivery, but also improves the user experience.

[0008] Optionally, the big data-based new media advertising delivery system further includes the following steps: A7 retrieves user identification information and corresponding content page access history from each content page; A8 generates user classification information based on each user's identification information using a preset user classification method; A9 generates corresponding category user access history based on the combination of content page access history records corresponding to each user identification information in the user category information. A10, based on the user access history records and the statistics of each page category data, determine the page category data with the highest access volume and define it as the preferred access page category data; A11, determine the corresponding approximate ad category data based on the page category feature data corresponding to the preference access page category data and define it as preference ad category data; A12, which displays various ad content from the preferred ad category data on the content page.

[0009] By adopting the above technical solution, the big data-based new media advertising delivery system can collect and analyze users' identification information and corresponding access history, classify users to determine the categories of content they prefer to browse, and match and deliver similar categories of advertisements based on the categories of preferred content. This not only improves the targeting and user acceptance of advertisements, but also increases the conversion rate of advertisements and achieves precision marketing.

[0010] Optionally, the content feature extraction algorithm includes the following steps: B1, extract the corresponding object text content, object image content and object video content based on the preset content features; B2 generates corresponding image description text based on the content of the object image using a pre-trained image description model; B3: Obtain the corresponding audio content based on the video content of the object and generate the corresponding audio text using a pre-trained audio-to-text conversion model; B4: Obtain several video frames from the object's video content and generate corresponding video frame description text using an image description model; B5 generates a comprehensive description text for the object based on a combination of the object's text content, image description text, audio text, and video frame description text. B6. Generate a set of corresponding object keywords based on the object's comprehensive description text using a preset keyword extraction algorithm; B7, Generates a corresponding set of object keyword vectors based on the set of object keywords using a preset word vector model; B8, calculate the corresponding set vector based on the object keyword vector set and define it as content feature data.

[0011] By adopting the above technical solution, the big data-based new media advertising delivery system can extract and decompose content pages or advertising content into different forms of content, and convert each form of content into standardized text descriptions. Finally, by extracting keywords and generating unified feature data through word vector models, it can comprehensively capture the semantic information of the content and improve the accuracy of subsequent content matching.

[0012] Optionally, the feature classification method includes the following steps: C1: For all content feature extraction objects, the content feature data are used to obtain multiple corresponding feature vector clusters using a preset clustering algorithm; C2 classifies all content feature extraction objects based on each feature vector cluster to generate content classification data; C3 calculates the corresponding cluster center vector for each feature vector cluster and defines it as content category feature data.

[0013] By adopting the above technical solution, the big data-based new media advertising delivery system can classify the content feature data of the content feature extraction object through clustering algorithm, and obtain the typical feature data of each category by calculating the cluster center vector. This not only avoids the time-consuming and labor-intensive manual classification, but also, by determining the typical feature data of each category, it can be compared with the typical feature data of each category when matching content in the future, which can greatly reduce the computational overhead.

[0014] Optionally, the similarity matching algorithm includes the following steps: D1, calculate the cosine value of the corresponding vector angle based on the page category feature data and the feature data of each advertisement category; D2, determine the maximum value among all the cosine values ​​of the angle between the vectors; D3 defines the advertising category data corresponding to the advertising category feature data corresponding to the maximum value of the cosine of the included angle between vectors as the approximate advertising category data.

[0015] By adopting the above technical solution, the big data-based new media advertising delivery system can achieve efficient and accurate advertising type matching by calculating the cosine similarity between the page category feature vector and the advertising category feature vector and selecting the matching result with the highest similarity. The matching method based on the feature data of content categories can avoid performing calculation matching between specific content pages and advertising content, which can not only improve matching efficiency, but also avoid the discomfort caused to browsing users due to excessively high matching between advertising content and page content.

[0016] Optionally, the big data-based new media advertising delivery system further includes the following steps for delivering advertising content: E1, if the approximate ad category data and the preferred ad category data are the same ad category data, then generate the corresponding ad pool based on the combination of all ad content in the approximate ad category data; E2, if the approximate ad category data and the preferred ad category data are not the same ad category data, then combine the ad content in the approximate ad category data and the ad content in the preferred ad category data to generate the corresponding ad pool; E3, when a user visits a content page, randomly selects advertising content from the advertising pool and displays it in the corresponding advertising slot on the content page; E4, when the ad content is displayed, calculate the corresponding ad display duration.

[0017] By adopting the above technical solutions, the big data-based new media advertising delivery system can ensure the diversity of advertising display by using random delivery. At the same time, by statistically analyzing the advertising display duration, it can not only meet the interests of different users, but also balance the effect of advertising delivery and avoid the advertising display being too monotonous.

[0018] Optionally, the big data-based new media advertising delivery system further includes the following steps for monitoring advertising content: F1 generates a set of supervised keyword vectors based on a pre-defined set of supervised keywords using a word vector model. F2, calculate the corresponding set vector based on the set of supervised keyword vectors and define it as the supervised word set vector; F3, calculates the cosine value of the angle between the corresponding vectors based on the set vector of supervised words and the advertising feature data corresponding to each advertising content, and defines it as the similarity of supervised words; F4. If the similarity of the monitored words is greater than the preset warning threshold, the corresponding advertising content is defined as suspected non-compliant advertising content.

[0019] By adopting the above technical solution, the big data-based new media advertising delivery system can identify and screen potentially risky advertising content through preset supervisory keyword vectors and similarity calculations. By setting reasonable warning thresholds, the system can proactively discover suspected non-compliant advertising content, realizing risk warning and quality control for advertising delivery. It can not only carry out preventive control before advertising delivery, but also ensure the compliance of advertising content and reduce the workload of manual review.

[0020] This application also provides a new media advertising delivery method based on big data, including the following steps: Page feature data for each content page is generated based on a preset content feature extraction algorithm. Based on the page feature data of all content pages, page classification data is generated and the corresponding page category feature data is determined using a preset feature classification method. Based on the preset advertising content, the advertising feature data of each advertising content is generated using a content feature extraction algorithm; Based on the advertising feature data of all advertising content, advertising classification data is generated separately using feature classification methods, and the corresponding advertising category feature data is determined. Based on the page category feature data corresponding to the content page and the feature data of each ad category, a preset similarity matching algorithm is used to determine the approximate ad classification data; Serving ad content from similar ad category data on the content page.

[0021] By adopting the above technical solution, the new media advertising delivery method based on big data can extract and classify the page feature data of the page to be advertised, and extract and classify the advertising feature data of the page content to be advertised, and ensure that the advertised ads are highly relevant to the page content through similarity matching. This not only improves the accuracy and conversion effect of advertising delivery, but also improves the user experience.

[0022] In summary, this application includes at least one of the following beneficial technical effects: 1. By extracting and classifying the feature data of pages that need to be targeted with ads, and extracting and classifying the feature data of the content to be targeted with ads, and by using similarity matching to ensure that the ads are highly relevant to the page content, the accuracy of ad targeting and conversion effect can be improved, as well as the user experience.

[0023] 2. By collecting and analyzing user identification information and corresponding access history, and classifying users, we can determine the categories of content that each user category prefers to browse. Based on the categories of preferred browsing content, we can match and deliver similar categories of advertisements, which not only improves the targeting and user acceptance of advertisements, but also increases the conversion rate of advertisements and achieves precision marketing.

[0024] 3. By extracting and breaking down content pages or advertisements into different forms of content, and converting each form of content into standardized text descriptions, and finally generating unified feature data through keyword extraction and word vector models, the semantic information of the content can be fully captured, thereby improving the accuracy of subsequent content matching. Attached Figure Description

[0025] Figure 1 This is a schematic diagram illustrating the principle of a new media advertising delivery system based on big data, as described in this invention.

[0026] Figure 2 This is a schematic diagram illustrating the process of a new media advertising delivery method based on big data according to the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0028] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.

[0029] refer to Figure 1 This invention provides a new media advertising delivery system based on big data, which is used to deliver appropriate advertising content on various content pages accessible to users, so as to improve the attractiveness of advertising content to users and increase the click-through rate of advertising content.

[0030] The big data-based new media advertising delivery system includes: Content page analysis module 10; User preference analysis module 20; Advertising content analysis module 30; and Advertising delivery control module 40; The content page analysis module 10, the user preference analysis module 20, and the advertising content analysis module 30 are respectively connected to the advertising delivery control module 40.

[0031] The content page analysis module 10 is mainly used to analyze the categories of content on a content page in order to match appropriate categories of advertising content. The content page can be a traditional webpage, an application on a mobile device, or a page displayed on other interactive devices.

[0032] The user preference analysis module 20 is mainly used to analyze the access preferences of users and determine the advertising categories that are suitable for user preferences.

[0033] The advertising content analysis module 30 is mainly used to analyze advertising content and classify each piece of advertising content. The advertising content can be in text, image, audio, or video format, or any combination thereof.

[0034] The advertising delivery control module 40 is mainly used to control various modules and the way advertisements are delivered. By analyzing content pages and advertising content, advertising content can be matched and delivered to appropriate content pages, thereby increasing the number of times users view advertising content and improving the click-through rate of advertisements when browsing content pages.

[0035] Furthermore, the advertising delivery control module 40 includes an advertising delivery strategy, which includes the following steps: A1, based on each preset content page, the content page analysis module 10 generates page feature data for each content page using a preset content feature extraction algorithm; Content pages are various content formats available for users to access and browse; The content feature extraction algorithm is a pre-defined algorithm used to extract corresponding feature data from selected content in order to reduce the computational overhead during subsequent data analysis. Page feature data refers to the feature data obtained from the content page through a content feature extraction algorithm.

[0036] A2, Generate page classification data and determine the corresponding page category feature data based on the page feature data of all content pages using a preset feature classification method; The feature classification method is a preset method that classifies all content pages based on their page feature data. The feature classification method can be a pre-trained classification model or a classification algorithm, such as deep learning, logistic regression, K-nearest neighbors (KNN), decision tree, and other algorithms and models. The page category data consists of a set of subcategories generated after the page feature data of all content pages have been classified using a feature classification method. Page category feature data are representative feature data corresponding to page classification data. They can be obtained based on all page feature data in the page classification data through a specific algorithm. For example, feature data can be further extracted from all page feature data in the page classification data, or the central feature vector can be determined by clustering algorithm.

[0037] A3, based on each preset advertisement content, the advertisement content analysis module 30 generates advertisement feature data for each advertisement content using a content feature extraction algorithm; The advertising content is the pre-determined data of the advertisements to be placed. Advertising feature data refers to the feature data corresponding to the advertising content.

[0038] A4. Based on the advertising feature data of all advertising content, generate advertising classification data and determine the corresponding advertising category feature data using the feature classification method. The advertising classification data is a set of subclasses generated after the advertising feature data of all advertising content has been classified using a feature classification method; The advertising category feature data are representative feature data corresponding to the advertising classification data.

[0039] A5. Based on the page category feature data corresponding to the content page and the ad category feature data corresponding to each ad category, a preset similarity matching algorithm is used to determine the approximate ad category data. The similarity matching algorithm is a pre-defined algorithm used to compare the page category feature data corresponding to each content page with the ad category feature data corresponding to each ad category data to find ad category data that is similar to the content page in terms of content, that is, to find ad subcategories with similar content based on the content of each content page. Approximate ad classification data refers to ad classification data that corresponds to ad category feature data that has similarity to the page category feature data corresponding to the content page.

[0040] A6, through the advertising delivery control module 40, each advertisement content in the approximate advertisement category data is delivered to the content page; Obtain the corresponding ad content from the approximate ad category data of the content page and place it on the content page for display.

[0041] Through the above technical solutions, the big data-based new media advertising delivery system can extract and classify page feature data for ads that need to be delivered, and extract and classify ad feature data for content to be delivered. By using similarity matching, it can ensure that the delivered ads are highly relevant to the page content, thereby improving the accuracy and conversion effect of ad delivery and enhancing the user experience.

[0042] Furthermore, the advertising strategy also includes the following steps: A7 retrieves user identification information and corresponding content page access history from each content page; User identification information is non-privacy information that can be used to identify users, such as browser cookies, login information of registered users, or browser fingerprints, based on the identification information generated when users visit content pages. Content page access history is the historical access data of all content pages for each user corresponding to their user identification information.

[0043] A8 generates user classification information based on each user's identification information using a preset user classification method; The user classification method is a pre-defined algorithm used to classify all user identification information. User identification information obtained using the same acquisition method usually has the same data format, making classification relatively convenient. For example, users can be classified based on their IP address, system time zone, system language, and other data. User classification information is a set of subclasses generated by the user classification method based on each user's identification information.

[0044] A9 generates corresponding category user access history based on the combination of content page access history records corresponding to each user identification information in the user category information. The category user access history is the collection of content page access history corresponding to each user identification information in the user category information, that is, the collection of content page access history of all users in the user category information.

[0045] A10, based on the user access history records and the statistics of each page category data, determine the page category data with the highest access volume and define it as the preferred access page category data; The preferred page category data consists of the page categories that have been visited most frequently in the user's browsing history, i.e., the page categories with the highest user visits in the user category information.

[0046] A11, determine the corresponding approximate ad category data based on the page category feature data corresponding to the preference access page category data and define it as preference ad category data; The preferred ad category data is ad category data that is similar to the page category feature data corresponding to the preferred access page category data. The corresponding approximate ad category data can be determined by the same method as in step A5.

[0047] A12, placing each ad content from the preferred ad category data on the content page; The corresponding ad content is retrieved from the preferred ad category data of the content page and displayed on the content page.

[0048] Through the above steps, the big data-based new media advertising delivery system can collect and analyze user identification information and corresponding access history, classify users to determine the categories of content they prefer to browse, and match and deliver similar categories of advertisements based on the categories of preferred content. This not only improves the targeting and user acceptance of advertisements, but also increases the conversion rate of advertisements and achieves precision marketing.

[0049] Furthermore, the content feature extraction algorithm includes the following steps: B1, extract the corresponding object text content, object image content and object video content based on the preset content features; The content feature extraction target is the pre-selected content that needs feature extraction, such as the content page or advertising content mentioned above; The text content of the object is the content data of the text portion obtained from the content feature extraction object; The object image content is the content data of the image portion obtained from the content feature extraction object; The video content of the object is the content data of the video portion obtained from the content feature extraction object.

[0050] B2 generates corresponding image description text based on the content of the object image using a pre-trained image description model; Image description models are pre-selected or pre-trained existing recognition models used to recognize images and generate corresponding descriptive text content. For example, existing large commercial models such as ChatGPT and Claude, or existing open-source models such as LLaVA, CogVLM, BLIP / BLIP-2, can all recognize the content in images and generate corresponding text description content well. Image description text is the text content generated by recognizing the content of an object image through an image description model.

[0051] B3: Obtain the corresponding audio content based on the video content of the object and generate the corresponding audio text using a pre-trained audio-to-text conversion model; The audio content is the audio data in the object video content, which can be obtained by extracting the audio track from the video file using applicable software; Audio-to-text conversion models are pre-selected or pre-trained models used to identify language content from audio files and convert it into text content. For example, open-source models such as Whisper, Project DeepSerch, and Kaldi can effectively identify text content in audio and generate corresponding text content. Audio text refers to the text content corresponding to the audio data in the object's video content.

[0052] B4: Obtain several video frames from the object's video content and generate corresponding video frame description text using an image description model; Video frame description text is the descriptive text for the image content corresponding to each video frame in the video content.

[0053] B5 generates a comprehensive description text for the object based on a combination of the object's text content, image description text, audio text, and video frame description text. The comprehensive description text of an object is a collection of the object's text content, image description text, audio text, and video frame description text, which can comprehensively reflect the content features and extract the content information contained in the object.

[0054] B6. Generate a set of corresponding object keywords based on the object's comprehensive description text using a preset keyword extraction algorithm; The keyword extraction algorithm is a pre-defined algorithm used to generate corresponding keywords based on the comprehensive descriptive text of the object; The object keyword set is a collection of multiple keywords extracted from the comprehensive description text of the object.

[0055] B7, Generates a corresponding set of object keyword vectors based on the set of object keywords using a preset word vector model; A word vector model is a pre-defined model used to convert keywords into corresponding word vectors, forming feature data that is easy to process later. Commonly used word vector models include Tencent Word Vectors and Baidu's Chinese Word Vectors, or you can train your own word vector model according to different needs. The object keyword vector set is the set of word vectors corresponding to each keyword in the object keyword set.

[0056] B8, calculate the corresponding set vector based on the object keyword vector set and define it as content feature data; A set vector is a vector obtained by combining the word vectors corresponding to each keyword in the set of keyword vectors for an object. A set vector can be obtained by averaging the word vectors of all keywords, or by combining certain weights and averaging them. Content feature data refers to the feature data corresponding to the content feature extraction object. For content pages, it corresponds to page feature data, and for advertising content, it corresponds to advertising feature data. For example, keyword vector set A contains keywords (artificial intelligence, machine learning, deep learning), and their corresponding word vectors are: Artificial intelligence vector: [0.2, 0.5, -0.3, ..., 0.1]; Machine learning vector: [0.3, 0.4, -0.2, ..., 0.2]; Deep learning vector: [0.25, 0.45, -0.25, ..., 0.15]; Then, summing the results for each dimension and dividing by the number of keywords yields the value for that dimension of the set vector. First dimension: (0.2 + 0.3 + 0.25) ÷ 3 = 0.25; Second dimension: (0.5 + 0.4 + 0.45) ÷ 3 = 0.45; The set vector can be obtained by analogy.

[0057] Through the above steps, the big data-based new media advertising delivery system can extract and decompose content pages or advertising content into different forms of content, and convert each form of content into standardized text descriptions. Finally, by extracting keywords and using word vector models to generate unified feature data, it can comprehensively capture the semantic information of the content and improve the accuracy of subsequent content matching.

[0058] Furthermore, the feature classification method includes the following steps: C1: For all content feature extraction objects, the content feature data are used to obtain multiple corresponding feature vector clusters using a preset clustering algorithm; Clustering algorithms are predefined algorithms used to classify objects by extracting content feature data based on content characteristics; for example, K-means algorithm, mean shift clustering algorithm, DBSCAN clustering algorithm, etc. The feature vector cluster is a subset of the content feature data after being classified by the clustering algorithm. The content feature data of content feature extraction objects with similar content are divided into the same feature vector cluster.

[0059] C2 classifies all content feature extraction objects based on each feature vector cluster to generate content classification data; By classifying the content feature extraction objects corresponding to each content feature data in each feature vector family, the content category data corresponding to the content feature extraction objects can be determined; the content category data corresponding to the content page is the page category data, and the content category data corresponding to the advertising content is the advertising category data.

[0060] C3, calculate the corresponding cluster center vector for each feature vector cluster and define it as content category feature data; The cluster center vector is the vector that represents the location of the cluster center of each eigenvector. Content category feature data refers to the cluster center vector of the feature vector cluster corresponding to each content category data.

[0061] Through the above steps, the big data-based new media advertising delivery system can classify the content feature data of the content feature extraction object through clustering algorithm, and obtain the typical feature data of each category by calculating the cluster center vector. This not only avoids the time-consuming and labor-intensive manual classification, but also, by determining the typical feature data of each category, allows for comparison of the typical feature data of each category when matching content in the future, which can greatly reduce the computational overhead.

[0062] Furthermore, the similarity matching algorithm includes the following steps: D1, calculate the cosine value of the corresponding vector angle based on the page category feature data and the feature data of each advertisement category; Calculate the cosine of the vector angle between the page category feature data and the feature data of each advertisement category; The cosine value of the angle between vectors is called cosine similarity.

[0063] D2, determine the maximum value among all the cosine values ​​of the angle between the vectors; Typically, the cosine value of the angle between vectors ranges from -1 to 1. The closer it is to 1, the more similar the two vectors are. Taking the maximum value indicates that the most similar page category feature data and advertising category feature data are paired.

[0064] D3 defines the advertising category feature data corresponding to the maximum value of the cosine of the included angle between vectors as the approximate advertising category data; Approximate advertising classification data refers to the advertising category feature data corresponding to the maximum value of the cosine of the included angle between vectors.

[0065] Through the above steps, the big data-based new media advertising delivery system can achieve efficient and accurate advertising type matching by calculating the cosine similarity between the page category feature vector and the advertising category feature vector and selecting the matching result with the highest similarity. The matching method based on the feature data of content categories can avoid performing calculations and matching between specific content pages and advertising content, which can improve matching efficiency and avoid causing discomfort to browsing users due to excessively high matching between advertising content and page content.

[0066] Furthermore, the big data-based new media advertising delivery system also includes the following steps for delivering advertising content: E1, if the approximate ad category data and the preferred ad category data are the same ad category data, then generate the corresponding ad pool based on the combination of all ad content in the approximate ad category data; E2, if the approximate ad category data and the preferred ad category data are not the same ad category data, then combine the ad content in the approximate ad category data and the ad content in the preferred ad category data to generate the corresponding ad pool; The ad pool is a collection of ad content from approximate ad category data and preferred ad category data.

[0067] E3, when a user visits a content page, randomly selects advertising content from the advertising pool and displays it in the corresponding advertising slot on the content page; The users who access the content pages are unspecified users.

[0068] E4, when the ad content is displayed, the corresponding ad display duration is calculated; The ad display duration is the length of time each ad appears on the screen when a user browses a content page. It can be used to analyze user ad preferences and adjust the frequency of ad display based on the display duration of each ad.

[0069] Through the above steps, the big data-based new media advertising delivery system can ensure the diversity of advertising displays by adopting a random delivery method. At the same time, by statistically analyzing the duration of advertising displays, it can not only meet the interests of different users, but also balance the effectiveness of advertising delivery and avoid the advertising display being too monotonous.

[0070] Furthermore, the big data-based new media advertising delivery system also includes the following steps for monitoring advertising content: F1 generates a set of supervised keyword vectors based on a pre-defined set of supervised keywords using a word vector model. The set of monitored keywords is a collection of keywords that are not suitable for use in advertising content; The supervised keyword vector set is the set of vectors for each keyword in the supervised keyword set.

[0071] F2, calculate the corresponding set vector based on the set of supervised keyword vectors and define it as the supervised word set vector; The set vector of supervised words is the set vector of the set of supervised keyword vectors.

[0072] F3, calculates the cosine value of the angle between the corresponding vectors based on the set vector of supervised words and the advertising feature data corresponding to each advertising content, and defines it as the similarity of supervised words; Supervised word similarity is the cosine similarity between the ad feature data corresponding to each ad content and the vector of the supervised word set.

[0073] F4. If the similarity of the monitored words is greater than the preset warning threshold, the corresponding advertising content is defined as suspected non-compliant advertising content.

[0074] The warning threshold is a pre-set reference value used to determine whether the similarity of the supervised words has reached a level that requires a warning. For suspected non-compliant advertising content, i.e., advertising content with a similarity of monitoring words exceeding the warning threshold, staff can be notified for further manual review.

[0075] Through the above steps, the big data-based new media advertising delivery system can identify and screen potentially risky advertising content through preset supervisory keyword vectors and similarity calculations. By setting reasonable warning thresholds, the system can proactively discover suspected non-compliant advertising content, achieving risk warning and quality control for advertising delivery. This not only enables preventative management before advertising delivery, ensuring the compliance of advertising content, but also reduces the workload of manual review.

[0076] refer to Figure 2 This application also provides a new media advertising delivery method based on big data, including the following steps: G1 generates page feature data for each content page based on a preset content feature extraction algorithm. G2 generates page classification data and determines the corresponding page category feature data based on the page feature data of all content pages using a preset feature classification method. G3 generates advertising feature data for each pre-set advertising content using a content feature extraction algorithm; G4 generates ad category data and determines the corresponding ad category feature data based on the ad feature data of all ad content using a feature classification method. G5 determines approximate ad classification data based on the page category feature data corresponding to the content page and the feature data of each ad category using a preset similarity matching algorithm; G6 displays ad content from similar ad category data on the content page.

[0077] Through the above steps, the big data-based new media advertising delivery method can extract and classify page feature data for ads that need to be delivered, and extract and classify ad feature data for the content to be delivered, and ensure that the delivered ads are highly relevant to the page content through similarity matching. This not only improves the accuracy and conversion effect of ad delivery, but also improves the user experience.

[0078] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.

Claims

1. A new media advertising delivery system based on big data, characterized in that, include: Content page analysis module; User preference analysis module; Advertising content analysis module; and Advertising delivery control module; The content page analysis module, the user preference analysis module, and the advertising content analysis module are all connected to the advertising delivery control module. The advertising delivery control module includes an advertising delivery strategy, which includes the following steps: A1. Based on each preset content page, the content page analysis module generates page feature data for each content page using a preset content feature extraction algorithm. A2, Generate page category data and determine the corresponding page category feature data based on the page feature data of all content pages using a preset feature classification method; A3, based on each preset advertisement content, the advertisement content analysis module generates advertisement feature data for each advertisement content using a content feature extraction algorithm; A4. Based on the advertising feature data of all advertising content, generate advertising classification data and determine the corresponding advertising category feature data using the feature classification method. A5. Based on the page category feature data corresponding to the content page and the ad category feature data corresponding to each ad category, a preset similarity matching algorithm is used to determine the approximate ad category data. A6, through the advertising delivery control module, delivers various advertising contents from the approximate advertising category data to the content page.

2. The new media advertising delivery system based on big data according to claim 1, characterized in that, Further steps include: A7 retrieves user identification information and corresponding content page access history from each content page; A8 generates user classification information based on each user's identification information using a preset user classification method; A9 generates corresponding category user access history based on the combination of content page access history records corresponding to each user identification information in the user category information. A10, based on the user access history records and the statistics of each page category data, determine the page category data with the highest access volume and define it as the preferred access page category data; A11, determine the corresponding approximate ad category data based on the page category feature data corresponding to the preference access page category data and define it as preference ad category data; A12, which displays various ad content from the preferred ad category data on the content page.

3. The new media advertising delivery system based on big data according to claim 2, characterized in that, The content feature extraction algorithm includes the following steps: B1, extract the corresponding object text content, object image content and object video content based on the preset content features; B2 generates corresponding image description text based on the content of the object image using a pre-trained image description model; B3: Obtain the corresponding audio content based on the video content of the object and generate the corresponding audio text using a pre-trained audio-to-text conversion model; B4: Obtain several video frames from the object's video content and generate corresponding video frame description text using an image description model; B5 generates a comprehensive description text for the object based on a combination of the object's text content, image description text, audio text, and video frame description text. B6. Generate a set of corresponding object keywords based on the object's comprehensive description text using a preset keyword extraction algorithm; B7, Generates a corresponding set of object keyword vectors based on the set of object keywords using a preset word vector model; B8, calculate the corresponding set vector based on the object keyword vector set and define it as content feature data.

4. The new media advertising delivery system based on big data according to claim 3, characterized in that, The feature classification method includes the following steps: C1: For all content feature extraction objects, the content feature data are used to obtain multiple corresponding feature vector clusters using a preset clustering algorithm; C2 classifies all content feature extraction objects based on each feature vector cluster to generate content classification data; C3 calculates the corresponding cluster center vector for each feature vector cluster and defines it as content category feature data.

5. The new media advertising delivery system based on big data according to claim 4, characterized in that, The similarity matching algorithm includes the following steps: D1, calculate the cosine value of the corresponding vector angle based on the page category feature data and the feature data of each advertisement category; D2, determine the maximum value among all the cosine values ​​of the angle between the vectors; D3 defines the advertising category data corresponding to the advertising category feature data corresponding to the maximum value of the cosine of the included angle between vectors as the approximate advertising category data.

6. The new media advertising delivery system based on big data according to claim 5, characterized in that, This further includes the following steps for delivering advertising content: E1, if the approximate ad category data and the preferred ad category data are the same ad category data, then generate the corresponding ad pool based on the combination of all ad content in the approximate ad category data; E2, if the approximate ad category data and the preferred ad category data are not the same ad category data, then combine the ad content in the approximate ad category data and the ad content in the preferred ad category data to generate the corresponding ad pool; E3, when a user visits a content page, randomly selects advertising content from the advertising pool and displays it in the corresponding advertising slot on the content page; E4, when the ad content is displayed, calculate the corresponding ad display duration.

7. The new media advertising delivery system based on big data according to claim 6, characterized in that, Further steps are included for monitoring advertising content: F1 generates a set of supervised keyword vectors based on a pre-defined set of supervised keywords using a word vector model. F2, calculate the corresponding set vector based on the set of supervised keyword vectors and define it as the supervised word set vector; F3, calculates the cosine value of the angle between the corresponding vectors based on the set vector of supervised words and the advertising feature data corresponding to each advertising content, and defines it as the similarity of supervised words; F4. If the similarity of the monitored words is greater than the preset warning threshold, the corresponding advertising content is defined as suspected non-compliant advertising content.

8. A new media advertising delivery method based on big data, characterized in that, Includes the following steps: Page feature data for each content page is generated based on a preset content feature extraction algorithm. Based on the page feature data of all content pages, page classification data is generated and the corresponding page category feature data is determined using a preset feature classification method. Based on the preset advertising content, the advertising feature data of each advertising content is generated using a content feature extraction algorithm; Based on the advertising feature data of all advertising content, advertising classification data is generated separately using feature classification methods, and the corresponding advertising category feature data is determined. Based on the page category feature data corresponding to the content page and the feature data of each ad category, a preset similarity matching algorithm is used to determine the approximate ad classification data; Serving ad content from similar ad category data on the content page.