Media content intelligent label generation and target classification method

Through multimodal feature extraction and analysis, combined with advanced NLP, CV and ML technologies, the efficiency and accuracy of integrated media content label generation and classification are solved, and efficient and accurate content management and personalized recommendation are achieved.

CN119939314APending Publication Date: 2025-05-06SPACE VISION (CHONGQING) TECH CO LTD
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510068154.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has problems such as inefficiency, poor consistency, poor flexibility and insufficient semantic understanding in the generation and classification of integrated media content, and it is difficult to cope with the need for rapid processing of large-scale content.

Method used

Multimodal feature extraction and analysis methods are used, combined with NLP, CV and ML technologies, key features of integrated media content are automatically identified and extracted to generate accurate labeling and classification results. Specific steps include data collection and preprocessing, feature extraction and analysis, intelligent label generation and target classification, as well as model training and optimization.

Benefits of technology

It improves the accuracy and efficiency of labeling and classification, reduces the workload of manual annotation and classification, and realizes the efficiency of content management and the accuracy of personalized recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a convergent media content intelligent label generation and target classification method, and belongs to the technical field of artificial intelligence and convergent media. In order to solve the problems of low labor efficiency, poor rule label flexibility and the like of traditional fusion media content label generation, the deep learning and multi-modal fusion technology is innovatively adopted. In a data processing layer, through deep integration of NLP and CV technologies, feature extraction of multi-source data such as texts, pictures, videos, audios and the like is realized. In a label generation layer, an intelligent label engine is constructed based on a CNN and RNN hybrid model, and a dynamic label library is established in cooperation with a rule matching strategy. And in the classification optimization level, the K-means clustering and hierarchical classification methods are combined to realize adaptive classification and optimization of the labels. Through continuous data training and a user feedback mechanism, the label generation accuracy, the classification efficiency and the system adaptability are remarkably improved, and an innovative solution is provided for intelligent management of the convergence media content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent processing of converged media content, and specifically to a method for generating intelligent tags and classifying targets for converged media content. Background Art

[0002] With the rapid development of information technology, converged media content has become one of the main forms of information dissemination. Converged media content not only integrates information from traditional media (such as newspapers, magazines, radio, and television), but also integrates multiple dissemination forms and content from emerging media (such as the Internet, mobile terminals, social media, etc.). This integration brings about the diversity, interactivity, and timeliness of content, but also raises new management and utilization challenges. Especially in application scenarios such as content retrieval and personalized recommendations, how to efficiently and accurately classify and annotate converged media content has become a key issue. Existing technologies have the following limitations: (1) The traditional manual tag generation method is inefficient and inconsistent: it mainly relies on manual operation, with editors manually adding tags to integrated media content. Although this method can ensure the accuracy and relevance of the tags, manual operation is time-consuming and labor-intensive, and cannot meet the needs of rapid processing of large-scale integrated media content. In addition, different editors may have different understandings of the same content, resulting in low consistency and standardization of tags; (2) Rule-based automatic tag generation methods have poor flexibility and insufficient semantic understanding: Rule-based automatic tag generation methods automatically generate tags for converged media content through predefined rules and pattern matching technology. This method improves the efficiency of tag generation to a certain extent, but the predefined rules and patterns are difficult to adapt to the diversity and complexity of converged media content, and the processing effect of non-standard content is poor; rule-based methods mainly rely on surface features and keyword matching, lack of understanding of the deep semantics of the content, resulting in low relevance and accuracy of the generated tags; In recent years, with the development of technologies such as NLP (natural language processing), CV (computer vision) and ML (machine learning), intelligent label generation and target classification methods have gradually become research hotspots. These methods can automatically identify and extract key features in integrated media content, and generate accurate labels and classification results based on these features, greatly improving work efficiency and accuracy. Summary of the invention

[0003] The present invention provides a method for intelligent tag generation and target classification of converged media content, which aims to accurately generate tags and reasonably classify targets for converged media content by using advanced technical means, so as to improve content management efficiency, promote personalized recommendations and precision marketing, etc. The present invention is implemented by the following technical solution, and the specific steps are as follows: S1: Collection and preprocessing of converged media content data. Collect converged media content data from multiple sources, including text, pictures, audio, video and other information. Synchronously record metadata such as content source, release time, and dissemination channel. When preprocessing the collected content, perform operations such as denoising (removing irrelevant characters), word segmentation (splitting text into words), and part-of-speech tagging (marking the part of speech of words) on text information; unify the format and adjust the resolution of pictures and videos; and perform noise reduction and format conversion on audio to ensure that the data is accurate and usable. S2: Feature extraction and analysis. Use NLP technology to obtain keywords, topics, sentiment tendencies and other features from text content. Use CV technology to analyze visual features such as color, shape, object recognition, etc. of pictures and videos. For audio, extract features such as spectrum, volume, and speech speed. Combine content metadata, including content source, release time, etc., to build a multimodal feature set to deeply analyze content semantics, sentiment and style; S3: Intelligent label generation. Based on the extracted features, the CNN, RNN and its variants are used to generate labels for the converged media content. By learning from a large amount of annotated data, the model can automatically generate relevant labels based on the characteristics of the new content. The labels are descriptive and accurate, and can reflect the core theme, emotional tendency, audience group and other information of the content. At the same time, it supports cross-media association analysis, integrates the feature information of multiple media types to generate more comprehensive and accurate labels, and combines context understanding technology to generate more accurate and comprehensive relevant labels; S4: Target classification. Based on the preset classification system, the news categories are divided into specific types such as politics, economy, entertainment, sports, etc., and the social media content is classified into categories such as life sharing, opinion discussion, and product promotion. At the same time, the K-means clustering analysis method (which clusters similar data) is used to fully consider the multimodal characteristics of the content, classify the integrated media content, and ensure that the classification results are reasonable and accurate; S5: Model training and optimization. Continuously collect new integrated media content data and corresponding labels and classification results for incremental learning and optimization of the model. Evaluate the performance of the model with the help of accuracy (the proportion of correctly predicted samples to total samples), recall (the proportion of correctly predicted positive samples to actual positive samples), and F1 value (an indicator of comprehensive accuracy and recall). For problems found in the evaluation results, such as inaccurate labels and classification errors, improve the performance and generalization ability of the model by adjusting the model parameters, improving the algorithm structure, or increasing the training data.

[0004] Preferably, in said S2, it also includes: When processing natural language, Word2Vec (word vector model) is used to convert text words into vectors to facilitate capturing semantic relationships and improve the accuracy of keyword extraction and topic analysis. In terms of computer vision, VGG (pre-trained model) is used for feature extraction, combined with target detection algorithms to identify key objects. When extracting audio features, STFT (short-time Fourier transform) technology is used to convert audio into a spectrogram and then extract features. Different modal features (text, vision, audio) are then fused through splicing or weighted summation to form a more representative comprehensive feature vector, providing more powerful data support for subsequent processing; Preferably, in said S3, it also includes: Build a tag library that includes general tags (common category tags) and field-specific tags (depending on the specific field), and dynamically update it based on industry progress and user needs. In the tag generation process, in addition to model prediction, rule matching is also used to directly assign corresponding tags to some content with obvious characteristics. For example, news with specific keywords in the title can be directly marked as related topics. For new vocabulary or concepts, update the model in a timely manner through online learning to ensure the timeliness and adaptability of the tags, so that the generated tags can better fit the characteristics and development changes of the integrated media content; Preferably, in said S4, it also includes: The classification system adopts a hierarchical structure, which covers both coarse-grained categories such as news, entertainment, and life services, and more detailed subcategories such as politics, economy, entertainment, and sports under the news category, so as to facilitate flexible classification according to different needs. At the same time, the K-means clustering analysis method is used to deeply consider the multimodal characteristics of the content (that is, the characteristics displayed by various forms such as comprehensive text, images, audio, and video) to classify the integrated media content. On this basis, an adaptive classification strategy is introduced to automatically adjust the clustering parameters and classification rules according to the integrated media content of different sources and types, so as to further improve the accuracy and adaptability of the classification. For news content with authoritative sources and strong professionalism, the weight of text features is appropriately increased, and it is more accurately divided into corresponding categories; for social media content with strong entertainment, the consideration of image and video features is enhanced.

[0005] Compared with the prior art, the method for generating intelligent tags and classifying targets for integrated media content disclosed in the present invention has the following beneficial effects: 1. Improve the accuracy of labels and classifications. Through multimodal data processing and advanced algorithm models, multiple information sources are comprehensively considered to generate more accurate, comprehensive and relevant labels and classification results, which better reflect the essential characteristics of converged media content; 2. Improve efficiency and reduce costs, reduce the workload of manual labeling and classification, and improve the degree of automation of content processing, which has a high cost-effectiveness advantage in large-scale integrated media content processing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative work: Figure 1 It is a flow chart of the collection and preprocessing of converged media content data of the present invention; Figure 2 is a flow chart of feature extraction and analysis of the present invention; Figure 3 is a flow chart of smart label generation of the present invention; Figure 4 is a flow chart of the classification of the objects of the present invention; Figure 5 It is a flow chart of model training and optimization of the present invention. DETAILED DESCRIPTION

[0007] Step 1: Collection and preprocessing of integrated media content data S11: Collecting media content data and metadata. For different types of media content, use the corresponding collection technology to carry out the collection work. For text content, such as news articles, social media texts, etc., crawl from relevant platforms with the help of web crawler technology. At the same time, record the source of each article (including the specific website domain name), release time and other information in detail to ensure the traceability of the data. For image and video content, obtain it through the API interface provided by the relevant platform, and record the resolution, file format and other related metadata of the image and video at the same time. For audio content, use professional audio recording and conversion tools for collection, use Audacity to record audio, set the sampling rate to 44.1kHz, and the number of channels to 2 to ensure audio quality. In the collection process, strictly ensure the integrity of the data, covering the format specifications of the text (uniformly using UTF-8 encoding), the clarity of the pictures and videos (the picture resolution is not less than 800×600 pixels, the video bit rate is not less than 1Mbps), and the audio quality is excellent (the signal-to-noise ratio is not less than 30dB), so as to provide a good data foundation for subsequent processing; S12: Data cleaning and format standardization. For the collected text data, html2text and emoji (text cleaning tools in natural language processing) are used to remove irrelevant information such as HTML tags and emoticons. Then, duplicate text content is removed and a hash algorithm is used to check for duplicate text. Specifically, the MD5 hash value of the text is calculated and the hash value is compared to determine whether the text is duplicated. After comparison, after cleaning and deduplication of 1,000 news articles in a certain system, the duplication rate was significantly reduced from 10% to less than 4%. For image and video data, we use deep learning-based image recognition technology (using a pre-trained ResNet model, with model parameters that are the default parameters trained on the ImageNet dataset) and video analysis algorithms (using the OpenCV library) to identify and remove blurry (using the Laplacian operator to calculate image clarity, with a threshold of 50), low-quality (based on image contrast, brightness and other indicators, contrast below 0.3 and brightness not in the range of [0.4, 0.6] is considered low quality) and repeated (using the hash algorithm to calculate the hash value of the key frames of the image or video for comparison). After processing, the effective data volume of images and videos increased by 20%. For audio data, we use the librosa library to remove abnormal sounds such as background noise and noise. By setting the noise threshold to 0.05, audio signals below this threshold are considered noise and removed, which significantly improves the signal-to-noise ratio of the audio. At the same time, the text encoding format was unified to UTF-8, the image and video sizes were adjusted to the standard format, the audio sampling rate was unified to 44.1kHz, and the number of channels was unified to 2, and standardized processing was performed to ensure that the data could smoothly enter the subsequent analysis process; S13: Metadata extraction and conversion. From the content management system, with the help of the data extraction interface, obtain various metadata in the integrated media content, including the title, author, and keywords of the text content (if provided by the platform, directly obtain them; if not provided, extract them in subsequent text analysis), the size of the picture and video (in pixels), format (including JPEG, PNG, MP4, etc.), shooting location (if included in the EXIF ​​information of the picture or video, directly extract it), the duration of the audio (accurate to milliseconds), format (including MP3, WAV, etc.), etc. For some unstructured metadata, use data parsing technology to convert it into structured data for subsequent processing. With the help of OCR (optical character recognition) technology (using the Tesseract OCR engine, the training data is its own English and Chinese language packages), extract the text in the picture and convert it into text data. Using speech recognition technology, convert the voice content in the audio into text information to further enrich data resources; S14: Fine preprocessing and data quality improvement. When preprocessing text information, regular expressions and the natural language processing library NLTK are used to perform more detailed text cleaning, restore abbreviations (replace them through predefined abbreviation dictionaries), unify date formats, etc. Then, word segmentation is performed, using the Jieba word segmentation tool, and stop words are removed (using the English stop word list in the NLTK library and the self-built Chinese stop word list). For the metadata of pictures and videos, data verification algorithms are used to carefully check the accuracy and consistency of the data, including checking whether the aspect ratio of the picture is in line with the norm (the common aspect ratio of pictures is 4:3 or 16:9, and the allowable error is ±0.1) and whether the duration of the video is accurate (compared with the actual duration of the video file, the error does not exceed 10 milliseconds). For audio metadata, the compatibility check of the audio format is performed (to ensure that the audio format can be supported by subsequent processing tools, and to check whether it is a common audio format such as MP3, WAV, etc.) to ensure the availability of the data in subsequent processing. Through the above series of operations, the converged media content data meets the requirements of high quality and structure before entering the feature extraction stage. After data preprocessing, compared with the original data, the accuracy increased by 35% and the usability increased by 50%, laying the foundation for subsequent feature extraction and analysis. Compared with traditional data processing methods, this method has obvious advantages in improving data quality, can effectively reduce the interference of noise data and invalid data on subsequent analysis, and improve the performance of the entire intelligent label generation and target classification system.

[0008] Step 2: Feature extraction and analysis S21: Text content feature extraction. For text content, we first use stem extraction and lemmatization techniques in natural language processing technology to restore words to their basic form and improve the accuracy of keyword extraction. When performing stem extraction and lemmatization on 1,000 news reports, we used the Snowball Stemmer stem extraction algorithm (with default parameters), and used WordNetLemmatizer (based on the WordNet lexicon) for lemmatization, which increased the accuracy of keyword extraction by 15%. Then, we used part-of-speech tagging technology and the part-of-speech tagger in the NLTK library (the training data is its own large-scale corpus) to identify nouns, verbs, adjectives and other parts of speech in the text, focusing on representative content words as keyword candidates. We used Stanford NLP (text analysis tool) to perform syntactic analysis. Its syntactic analysis model was trained based on a large-scale syntactic tree library. By analyzing sentence structure, extracting sentence structure and grammatical relations, it assisted in understanding the semantics of the text and further screened out more representative keywords. S22: Extraction of visual features of image and video content. For image and video content, the target detection algorithm Faster R-CNN in deep learning is used (pre-trained model is used, and the model parameters are the default parameters obtained by training on the COCO dataset) to identify key objects, characters, scenes and other elements, and extract their position (expressed in coordinates in the image coordinate system, accurate to pixels), size (expressed in the width and height of the bounding box, in pixels), category (such as people, cars, animals, etc.) and other features. Using image segmentation technology (using the U-Net model based on deep learning, the training data is the Pascal VOC dataset), different areas in the image or video are segmented, and the color, texture, shape (calculating the perimeter, area and other geometric features of the area) and other features of each area are analyzed. For video content, the motion information between frames is also analyzed, including the object's motion trajectory, speed, etc., to capture dynamic features. After processing 500 video clips, the accuracy of key object recognition reached 85%, and the accuracy of scene classification reached 80%; (1) For the extraction of inter-frame motion information of video content, a combination of optical flow method and target tracking algorithm is adopted. The optical flow method is used to calculate the motion vector of pixel points between adjacent frames. The pyramid-based Lucas-Kanade optical flow algorithm is adopted. When constructing the image pyramid, the number of pyramid layers is set to 5, and Gaussian kernel is used for image smoothing with smoothing parameter σ = 1.5. By calculating the optical flow at different resolutions, the accuracy and robustness of motion vector calculation are improved. The target tracking algorithm is based on a deep learning target tracker (Siamese, using a pre-trained model and training data from a large-scale video tracking dataset), which can accurately track the position changes of the target object in the video sequence; (2) When calculating the optical flow, assume that the video sequence is ,(in represents the t-th frame image), and the motion vector field between adjacent frames is calculated by the optical flow method (where m(x,y)=(u(x,y),v(x,y)) represents the displacement of pixel (x,y) between adjacent frames). The Lucas-Kanade optical flow algorithm minimizes the error function ,(in Represents the gradient operator. Here, the Gaussian weight function is used. The weight of the central pixel is 1, and the weights of the surrounding pixels decrease according to the Gaussian distribution. is the time derivative of the image at (x,y), is a regularization parameter with a value of 0.01) to solve the motion vector (u(x,y),v(x,y)). The motion trajectory of an object can be obtained by tracking the changes in the motion vector in consecutive frames. The speed can be calculated based on the size of the motion vector and the time interval between frames, such as the speed (in is the time interval between adjacent frames). When the target tracking algorithm is initialized, the target object is detected in the first frame by the target detection algorithm, and then the target object is tracked in subsequent frames based on its appearance features, which can accurately obtain the dynamic features such as the object's motion trajectory and speed in the video; S23: Audio content feature extraction. For audio content, audio feature extraction technology is used, including Mel frequency cepstral coefficient (MFCC, using 12-dimensional MFCC features, frame length of 25ms, frame shift of 10ms, Hamming window function), linear prediction coding (LPC, using 10th order LPC coefficients), to extract audio spectrum features, pitch (obtained by calculating the fundamental frequency of the audio signal, using the autocorrelation method, the search range is 50Hz - 500Hz), volume (calculating the root mean square value of the audio signal), speech rate (through speech endpoint detection, calculating the number of syllables per unit time, speech endpoint detection uses a double threshold method based on energy and zero crossing rate, with an energy threshold of 0.01 and a zero crossing rate threshold of 5) and other information. Use speech recognition technology to convert the speech in the audio into text, and then use text analysis technology (using the TextBlob library for sentiment analysis, the LDA topic model for topic extraction, and the training data is a large-scale news corpus) to perform sentiment analysis (divided into three categories: positive, negative, and neutral) and topic extraction (extracting the top 5 topics) on the converted text, combining the semantic features of the audio with the spectral features to form a more comprehensive audio feature representation; S24: Multimodal feature fusion. Integrate various features extracted from text, pictures, videos, and audio. For text features, use word vector representation, including Word2Vec (using the Skip-gram model, the word vector dimension is 300, and the training data is a large-scale text corpus such as news, novels, and blogs), GloVe (the word vector dimension is 300, and the training data is a large-scale corpus containing 6 billion words), to convert keywords, topics, emotions, etc. into vector form. For visual features, object recognition results (using one-hot encoding to represent categories, such as people [1,0,0,…], cars [0,1,0,…], etc., and the vector length is 80 categories), color features (quantizing color histograms into vectors with a length of 16), texture features (joining gray-level co-occurrence matrix statistics into vectors with a length of 4), etc. are encoded and converted into vectors; for audio features, spectral features (MFCC feature vector length is 12, LPC coefficient vector length is 10, and the length after concatenation is 22), semantic features (encoding sentiment analysis results and topic extraction results, such as positive [1,0,0], negative [0,1,0], neutral [0,0,1], and the topic is one-hot encoded, the vector length is 5 topics, and the length after concatenation is 8) are quantized. Feature fusion techniques, such as feature concatenation and weighted summation methods, are used to merge feature vectors of different modalities into a comprehensive feature vector, making full use of multimodal information and improving the representativeness and discrimination of features. In the actual fusion process, different weights can be assigned according to the importance of different modal features to the overall features. For example, for integrated media content that is mainly textual, the weight of the text feature vector can be appropriately increased so that the integrated feature vector after fusion can better reflect the essence of the content. In the feature concatenation method, the text feature vector, visual feature vector, and audio feature vector are directly concatenated in order to form a comprehensive feature vector. In the weighted summation method, different weights are assigned according to the importance of different modal features to the overall feature. For integrated media content that is mainly textual, such as news articles, the weight of the text feature vector is set to 0.6, the weight of the visual feature vector is set to 0.2, and the weight of the audio feature vector is set to 0.2. The comprehensive feature vector is calculated by weighted summation. (in , , are text, visual, and audio feature vectors, respectively. , , is the corresponding weight). Through these feature fusion technologies, we can make full use of multimodal information, improve the representativeness and discrimination of features, and make the fused comprehensive feature vector more able to reflect the essence of the content.

[0009] Step 3: Smart label generation S31: Label generation based on text features. The label generation model was constructed using the deep learning framework PyTorch (version 1.8.0). The model uses a bidirectional long short-term memory network (BiLSTM) structure, with the number of hidden layers set to 2 layers, each layer containing 128 neurons. The number of neurons in the input layer is determined according to the dimension of the text feature vector (for example, when a 300-dimensional Word2Vec word vector is used, the number of neurons in the input layer is 300), and the number of neurons in the output layer is equal to the number of label categories (for example, for news text, the label categories include 20 categories such as technology, entertainment, and sports, and the number of neurons in the output layer is 20). For text content, the extracted feature vectors such as keywords, topics, and emotions are used as inputs to the BiLSTM model. The Adam optimizer is used for model training (the initial value of the learning rate is set to 0.001, β1 = 0.9, β2 = 0.999, = 1e - 8), and the loss function uses the multi-label classification cross entropy loss function (Multi-Label Cross-Entropy Loss). During the training process, a large amount of labeled text data was used for supervised learning, and the model parameters were continuously adjusted to improve the accuracy of label prediction. The feature vector obtained after feature extraction of the text was input into the BiLSTM model for training, and the number of training rounds was set to 50. After the training was completed, label prediction was performed on another 1,000 unlabeled news articles. After manual verification, the accuracy of label prediction reached 80%. The model can effectively generate accurate labels based on text features; S32: Label generation by integrating visual features. For image and video content, the visual feature vector is input into the convolutional neural network (CNN), and the VGG pre-trained model is used for feature extraction, and then the fully connected layer is connected for label prediction. Multimodal fusion technology is used to integrate text features with visual features. The text feature vector and the visual feature vector are concatenated or weighted summed in the middle layer of the model, and then the label is predicted through the subsequent layers, so that the label generation comprehensively considers a variety of media information and improves the accuracy and comprehensiveness of the label. Taking short video content as an example, the VGG model is first used to extract visual features from the key frames in the video, and the text feature vector is extracted from the text description related to the video (including title, introduction, etc., and the text feature vector is obtained by the same method as text feature extraction). The two features are then fused and input into the model for label prediction. The prediction results show that after integrating visual and text features, the accuracy of the label description of the video content is 20% higher than that of using only visual features, and it can more accurately cover the theme, style and other information of the video; S33: Label generation combined with audio features. For audio content, the audio feature vector is input into the audio classification model based on deep learning, using a convolutional recurrent neural network (CRNN) structure, where the convolution layer contains 3 convolution kernels with a size of 3×3 and a step size of 1, the pooling layer uses maximum pooling with a pooling kernel size of 2×2, and the recurrent layer uses GRU (gated recurrent unit), which contains 2 layers, 128 neurons in each layer, and the number of neurons in the output layer is equal to the number of label categories. Classify the audio and generate corresponding labels. Combine the labels generated by the text information converted from the audio content (the text after speech recognition, the text feature vector is obtained using the same method as the text feature extraction) with the labels generated by the audio features themselves to further optimize the label generation results of the audio content. The CRNN model is used to analyze the audio's spectral features (using 12-dimensional MFCC features), volume, speech speed and other features to generate initial labels. At the same time, the text after speech recognition is subjected to topic analysis to generate auxiliary labels (using the TextBlob library for topic extraction, extracting the first three topics). When the two are fused, a weighted sum method is used (the weight of the audio feature label is set to 0.6, and the weight of the text feature label is set to 0.4) to obtain the final label. According to evaluation, the fused label has improved the accuracy of program type classification by 18% compared with the label generated by a single audio feature, and can better reflect the characteristics of the audio content; S34: Label library construction and optimization. A label library is established, which contains general labels and domain-specific labels. General labels include "technology", "entertainment", "sports", etc. Domain-specific labels are based on the specific fields of the integrated media content, such as "football game" and "basketball star" in the sports field. In the label generation process, the model matches and predicts the labels in the label library based on the input feature vector, and at the same time, combined with the contextual information of the content, uses the contextual understanding technology based on the attention mechanism (the attention mechanism uses Bahdanau Attention, and the model parameters are obtained by pre-training on a large-scale text corpus) to adjust and optimize the labels to ensure that the generated labels are accurate, relevant and of high quality, and can accurately reflect the key features and themes of the integrated media content. Based on the preliminary matching of the feature vector, the general label "entertainment" and the domain-specific label "movie release" may be obtained. Then, the model uses the contextual understanding technology to further analyze the contextual information such as the types of movies mentioned in the article (such as science fiction movies) and the starring lineup (including the names of well-known actors). After comprehensive judgment, the labels are finally labeled as "entertainment, film and television, science fiction, [starring name]". Such labels not only clarify the entertainment field to which the movie belongs, but also include the key features of the movie, such as genre, starring actors and features, making the labels more accurate and informative. At the same time, with the development of the industry, such as the emergence of the concept of the emerging industry "metaverse", relevant labels are added to the label library in a timely manner (reviewed and added by experts in the professional field), and through online learning (using incremental learning algorithms, such as gradient-based online learning algorithms, with a learning rate set to 0.01), the model can recognize and apply new labels to ensure the timeliness of the label library.

[0010] Step 4: Target Classification S41: Initial classification based on the preset classification system. According to the preset classification system, the classification system adopts a multi-level structure. The first-level classification includes news, entertainment, education, technology, life services and other major categories. The second-level classification is subdivided into current affairs news, financial news, sports news, entertainment news, social news, etc. under news, and each second-level classification can be further subdivided, such as current affairs news can be divided into domestic current affairs, international current affairs, local current affairs, etc. The construction of the classification system is based on the analysis of a large amount of integrated media content as well as industry standards and actual application needs to ensure comprehensiveness and rationality. When matching the generated smart tags with the categories in the classification system, the smart tags are first preprocessed, including removing stop words (using a general stop word list, such as the English nltk stop word list and the Chinese Harbin Institute of Technology stop word list), stem extraction (using the Porter Stemmer algorithm) and other operations to improve the accuracy of the match. Then, a matching method based on the keyword vector space model is used to convert the tag keywords into vectors (using the Word2Vec model with a vector dimension of 300), and the cosine similarity with the keyword vectors of each category in the classification system is calculated. The similarity threshold is set to 0.6. When the similarity is greater than or equal to the threshold, it is determined to be a successful match, thereby determining the initial category to which the integrated media content belongs; S42: Classification prediction based on machine learning algorithm. Using the classification algorithm support vector machine (SVM) in machine learning, the radial basis function (RBF) kernel function ( , Where = 0.1, penalty parameter C = 1.0, and the optimization algorithm uses the sequential minimum optimization (SMO) algorithm. The extracted feature vector (the feature vector is a fusion of text feature vector, visual feature vector, and audio feature vector. For example, the text feature vector uses a TF-IDF weighted vector with a dimension of 500; the visual feature vector is extracted by the VGG model and then reduced to a dimension of 200; the audio feature vector is composed of audio features such as MFCC, with a dimension of 100, and the fusion method is weighted summation, with weights of 0.5, 0.3, and 0.2 respectively) is used as input to train the classification model. The training set data is collected from converged media platforms in multiple fields, including 8,000 accurately classified converged media content data, covering news, entertainment, technology, education and other categories, and the data distribution of each category is relatively balanced. During the training process, the 10-fold cross-validation technique was used, that is, the training set was divided into 10 parts, 9 of which were used as training data and 1 as validation data in turn, and the accuracy, recall rate and other indicators of each validation were calculated. The average value was taken as the basis for model performance evaluation to ensure the accuracy and generalization ability of the model. After the training was completed, another 2,000 unclassified converged media content (the source was similar to the training set to ensure the consistency of data distribution) was predicted. After evaluation, the classification accuracy reached 82%, indicating that the model has a high effectiveness in classification prediction. For new converged media content, the trained model is used to predict its category; S43: Cluster analysis optimizes classification results. Cluster analysis algorithms K-means clustering (the initial cluster center is randomly selected, and the number of clusters is initially set based on experience and data distribution, such as 10 categories, which are subsequently adjusted based on clustering effects) and DBSCAN (neighborhood distance e = 0.5, minimum number of sample points MinPts = 5) are introduced to perform unsupervised clustering of converged media content. According to the feature similarity of the content (feature similarity calculation is based on the fused multimodal feature vector, using the Euclidean distance metric), it is divided into different clusters. In the news content, 5,000 news articles were collected (collected from different sections of major news media websites, including current affairs, finance, technology, entertainment and other sections to ensure that the news types are comprehensive). The K-means clustering algorithm was used to cluster the articles based on multimodal features such as text features (keywords, topics, etc., keywords were extracted by TF-IDF algorithm, topics were analyzed by LDA topic model, and the number of topics was set to 5), visual features (if there are pictures, image feature extraction methods are used, such as extracting color histograms, edge features, etc., and then normalizing them). After clustering, the average distance between samples within the cluster and the average distance between samples between clusters are calculated, and the classification results are further optimized by analyzing the feature differences within and between clusters (including comparing the keyword distribution, topic distribution, and visual feature distribution of different clusters). Potential new categories or subcategories are discovered. If after clustering, it is found that some of the news articles have obvious differences in content features from other news, after manual analysis, these articles mainly focus on specific technological breakthroughs in the field of emerging science and technology, and they can be used as a new subcategory "Emerging Science and Technology Dynamics" under the news category. Through cluster analysis, we can dig out the potential classification structure in the content and improve the refinement of classification. At the same time, we can dynamically adjust and supplement the classification system according to the clustering results to make it more perfect and adapt to actual needs.

[0011] Step 5: Model training and optimization S51: Construction and division of training sets. Collect a large amount of labeled media content data as training sets. Data labeling is performed by professionals to ensure the accuracy of labels and classifications. Use the training subset to train the label generation model and classification model. During the training process, select the appropriate loss function according to the type of model. For the classification model, the cross entropy loss function is used to measure the difference between the model prediction result and the true label, so as to encourage the model to continuously adjust parameters during the training process to reduce this difference. (1) For the label generation model, the multi-label classification cross entropy loss function (Multi-LabelCross-Entropy Loss) is used. Since the integrated media content may have multiple related labels, for example, a news report may involve multiple fields such as science and technology, economy, and society at the same time, this loss function can effectively handle the multi-label classification problem. Suppose the label probability distribution predicted by the model is , the true label is , then the loss function (where C is the total number of label categories). By minimizing this loss function, the model can learn the ability to accurately predict multiple labels; (2) For the classification model, the cross-entropy loss function is used as mentioned above. In multi-classification problems, this loss function can measure the difference between the model prediction result and the true label. Suppose the category probability distribution of the model output is p(y|x) (where x is the input data and y is the category label), and the one-hot encoding of the true label is , then the loss function The reason for choosing this loss function is that it has good mathematical properties in classification problems, which can prompt the model to continuously adjust parameters during training, making the predicted probability distribution closer to the true label distribution; S52: Model performance monitoring and early stopping application. During the training process, the validation subset is used to monitor the performance of the model. The accuracy, recall, F1 value and other evaluation indicators of the model on the validation subset are calculated regularly. In the early stage of training, as the number of training rounds (epochs) increases, the accuracy of the model gradually increases and the loss function value gradually decreases. When the performance of the model on the validation subset no longer improves or there are signs of overfitting, such as the accuracy no longer increases and the loss function value begins to increase, stop training and save the current model parameters. In one experiment, when training to the 50th epoch, it was found that the accuracy of the model on the validation set reached the highest value of 85%. Although there were slight fluctuations afterwards, it no longer increased overall. At the same time, the loss function value began to increase. At this time, the early stopping method was used to stop training, which effectively prevented the model from overtraining, improved the generalization ability of the model, and ensured that the model could perform well on unseen data; S53: Implementation of model optimization strategy. Carry out targeted optimization for the problems found in the model evaluation process. If the model has the problem of low accuracy, analyze the possible reasons, such as insufficient feature extraction, unreasonable model structure, etc. For feature extraction problems, improve the feature extraction algorithm or increase the feature dimension. For model structure problems, adjust the model architecture parameters such as the number of network layers and the number of nodes. For example, for classification models, adjust the number of nodes in each layer and observe the changes in model performance. At the same time, adjust the model's hyperparameters, including learning rate, regularization parameters, etc., according to actual application scenarios and user feedback. After a series of optimization measures, re-evaluate the model and effectively improve the performance of the model; S54: Continuous model updating and feedback-driven optimization. With the continuous generation of new converged media content data, collect the newly generated converged media content data every week (the source is similar to the initial training set to ensure data consistency and diversity), add it to the training set after annotation, and restart the model training process. Continuously collect user feedback on the model output results, including the following specific methods: set up a user feedback portal on the application platform, where users can evaluate the label and classification results (satisfied or dissatisfied), and provide specific opinions and suggestions (such as if a certain label is inaccurate and should be modified to a certain label). Conduct user questionnaires regularly to collect user satisfaction with recommended content and evaluation of the accuracy of labels and classifications (including correct, partially correct, wrong, etc.), and use user feedback as an important basis for optimizing the model. After continuous optimization, the accuracy of labels and classifications in subsequent tests of the model has increased by about 10% compared with the previous ones, providing users with better services and further improving the practicality and adaptability of the model. At the same time, according to user satisfaction feedback on recommended content, adjust the parameters of the recommendation algorithm (increase the recommendation weight of user-interested areas, etc.) to improve user experience.

Claims

1. A method for generating intelligent tags and classifying targets for converged media content, characterized in that: The method comprises the following steps: S1: Collection and preprocessing of converged media content data, including collecting converged media content data from multiple sources, synchronously recording metadata such as content source, release time, and dissemination channel; preprocessing the collected content to ensure that the data is accurate and usable; S2: Feature extraction and analysis, including using natural language processing technology to obtain keywords, topics, sentiment tendencies and other features from text content, using computer vision technology to analyze the visual features of pictures and videos, extracting features such as spectrum, volume, and speech rate from audio, and building a multimodal feature set in combination with content metadata; S3: Intelligent tag generation, including generating tags based on extracted features using deep learning models, supporting cross-media association analysis, and combining contextual understanding technology to generate more accurate, comprehensive and relevant tags; S4: Target classification, including classifying the converged media content based on a preset classification system, and using cluster analysis methods to fully consider the multimodal characteristics of the content to ensure that the classification results are reasonable and accurate; S5: Model training and optimization, including continuous collection of new data for incremental learning and optimization of the model, evaluating model performance with the help of evaluation indicators, improving the model for problems, and improving model performance and generalization capabilities.

2. The method according to claim 1, characterized in that: The S2 further includes: In natural language processing, Word2Vec is used to convert text words into vectors to capture semantic relationships and improve the accuracy of keyword extraction and topic analysis. Stem extraction and word form restoration techniques are used to restore words to their basic form, and part-of-speech tagging techniques are used to identify parts of speech. Syntactic analysis is used to extract sentence structure and grammatical relationships to assist in understanding text semantics and screen out more representative keywords. In terms of computer vision, VGG is used for feature extraction, combined with the target detection algorithm to identify key objects, people, scenes and other elements and extract their location, size, category and other features. Image segmentation technology is used to analyze the color, texture, shape and other features of each area. For video content, the inter-frame motion information is also analyzed. When extracting features from audio, the STFT technology is used to convert the audio into a spectrogram and then extract the spectral features. The Mel-frequency cepstral coefficients, linear predictive coding and other technologies are used to extract the pitch, volume, speaking speed and other information of the audio. The speech recognition technology is used to convert the speech in the audio into text. The text analysis technology is then used to perform sentiment analysis and topic extraction on the converted text. The semantic features of the audio are combined with the spectral features to form a more comprehensive audio feature representation. Finally, the different modal features are fused through splicing or weighted summation to form a more representative comprehensive feature vector.

3. The method according to claim 1, characterized in that: The S3 further includes: When building a label generation model based on text features, we use the deep learning framework PyTorch, adopt a bidirectional long short-term memory network structure, set the number of hidden layers and neurons, determine the number of neurons in the input layer and output layer, and use the Adam optimizer and multi-label classification cross entropy loss function for the training model. We use a large amount of labeled text data for supervised learning. When fusing visual features to generate labels, the visual feature vector is input into the convolutional neural network, and the VGG pre-trained model is used for feature extraction and then connected to the fully connected layer for label prediction. The multimodal fusion technology is used to concatenate or weighted sum the text features and visual features in the middle layer and then pass them through the subsequent layers for label prediction. When generating labels based on audio features, the audio feature vectors are input into an audio classification model based on deep learning. A convolutional recurrent neural network structure is used to classify the audio and generate corresponding labels. The labels generated by the audio features themselves are fused with the labels generated by the text information converted from the audio content. The final labels are obtained by weighted summation. At the same time, a label library is constructed, which includes general labels and domain-specific labels. The model matches and predicts the labels in the label library based on the feature vectors, and optimizes the labels in combination with context understanding technology to ensure that the labels are accurate, relevant and of high quality, and the label library is updated in a timely manner to ensure timeliness.

4. The method according to claim 1, characterized in that: The S4 further includes: When the initial classification is based on the preset classification system, the classification system adopts a multi-level structure. After preprocessing the intelligent tags, a matching method based on the keyword vector space model is used to calculate the cosine similarity between the tag keywords and the keyword vectors of each category in the classification system to determine the initial category to which the integrated media content belongs; When making classification predictions based on machine learning algorithms, we use the support vector machine algorithm, select the radial basis function kernel function, determine the penalty parameters and optimization algorithm, use the fused feature vector as input to train the classification model, and use the 10-fold cross-validation technology to ensure the accuracy and generalization ability of the model to predict unclassified integrated media content; When clustering analysis is used to optimize classification results, clustering analysis algorithms such as K-means clustering and DBSCAN are introduced to divide the content into different clusters according to the feature similarity of the content, calculate the relevant distances and feature differences of samples within and between clusters, explore the potential classification structure, optimize the classification results, and dynamically adjust and supplement the classification system based on the clustering results.

Citation Information

Cited By

  • Visualization method and device for digital twin basin animation and server

    CN120182450A

  • Intelligent data labeling method and system, electronic equipment and storage medium

    CN120375375A

  • Big data management system and method based on hierarchical label system

    CN120596475A

  • A big data governance system and method based on a hierarchical label system

    CN120596475B

  • Label setting method and system for unstructured data

    CN120950691A