Social media video analysis method and system
By performing multimodal semantic disambiguation analysis and cross-platform unified popularity ranking on social media videos, and generating hierarchical tags, this solution addresses the issues of multimodal feature semantic ambiguity, cross-platform analysis distortion, and adaptation to trending topics in vertical fields in social media video analysis, thus achieving an efficient and accurate video content creation solution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG HENGQIN SHUSHUSHUO STORY INFORMATION TECH CO LTD
- Filing Date
- 2025-11-03
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies cannot achieve accurate semantic matching of multimodal features in social media videos, have distorted cross-platform analysis results, are difficult to adapt to changes in hot topics in vertical fields, and cannot guarantee the completeness of analysis. There is a contradiction between excessive and insufficient data anonymization.
By collecting and preprocessing multi-source data from social media videos, performing multimodal semantic disambiguation analysis, generating hierarchical tags for vertical fields, and conducting cross-platform unified popularity ranking, video content creation analysis is performed in conjunction with user interaction data and compliance attribute characteristics.
It achieves accurate semantic matching of multi-source features of social media videos, eliminates distortion of cross-platform analysis results, adapts to changes in hot topics in vertical fields, improves the timeliness, accuracy and adaptability of analysis to vertical fields, and balances compliance and analysis precision.
Smart Images

Figure CN121963012A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of video analysis, and in particular to a social media video analysis method and system. Background Technology
[0002] With the rapid development of social media platforms such as short videos and social communities, social media videos have become a core carrier for users to obtain information and interact. Their characteristics of multi-platform distribution, high-frequency content iteration, and vertical field segmentation have spurred a strong demand for batch analysis, popularity ranking, and user demand insights of social media videos.
[0003] Currently, there are technical solutions for video structuring and content distribution in the industry. However, existing video structuring technologies rely on text-generated tags, which cannot distinguish between different meanings for the same feature, making it impossible to achieve accurate semantic matching of multimodal features in social media videos. Secondly, existing content distribution technologies struggle to achieve unified analysis of interaction data and popularity rankings for the same video across multiple platforms, resulting in distorted cross-platform analysis results. Furthermore, the statically preset tag system of video structuring technologies cannot cover emerging hot topics in vertical social media fields and is difficult to adapt to changes in these trends. Moreover, social media videos contain a large amount of user privacy information, and core analysis relies on detailed visual features. Existing technologies cannot balance these two aspects, resulting in a contradiction between excessive and insufficient data anonymization when analyzing social media videos, which makes it impossible to guarantee the completeness of the analysis. Summary of the Invention
[0004] To address the problems of existing technologies, such as difficulty in semantic matching of video features, distortion of cross-platform analysis results, difficulty in adapting to changes in trending topics in vertical fields, and inability to guarantee the completeness of analysis, this invention proposes a social media video analysis method and system that can achieve accurate semantic matching of video features, eliminate the problem of distortion of cross-platform analysis results, adapt to changes in trending topics in vertical fields, and thus guarantee the completeness of analysis.
[0005] To achieve the above-mentioned technical effects, the technical solution of the present invention is as follows: A social media video analysis method includes the following steps: S1. Collect multi-source data of social media videos, and preprocess the multi-source data of social media videos to obtain standardized video data; S2. Perform multimodal semantic disambiguation analysis on the standardized video data to obtain the multimodal semantic disambiguation results; S3. Based on the multimodal semantic disambiguation results, generate hierarchical labels for vertical domains; S4. Based on the hierarchical tags, perform cross-platform unified popularity ranking on the standardized video data to obtain popularity data; S5. Based on the popularity data and the hierarchical tags, perform video content creation analysis and output a video content creation plan.
[0006] Preferably, the preprocessing of the multi-source social media video data to obtain standardized video data includes: S11. The social media video multi-source data is classified using a preset two-dimensional classification evaluation model to obtain two-dimensional classification data; S12. Perform dynamic desensitization on the dual-dimensional hierarchical data to obtain the data desensitization results; S13. Perform a quality score assessment on the data anonymization results to obtain a quality score; S14. Determine whether the quality score is greater than a preset threshold. If so, standardize the format and structure of the data desensitization result to obtain standardized video data. S15. Determine whether the standardized video data belongs to a new field or a new platform. If so, perform cold start cross-domain adaptation on the standardized video data to obtain standardized video data adapted to the new field or platform. If not, directly output the original standardized video data.
[0007] Preferably, the dual-dimensional hierarchical evaluation model includes a parallel privacy risk level layer and an analysis value level layer; the privacy risk level layer is used to classify the social media video multi-source data into four levels: extremely high privacy, high privacy, medium privacy, and low privacy; the analysis value level layer is used to classify the social media video multi-source data into four levels: extremely high value, high value, medium value, and low value.
[0008] Preferably, the step of performing multimodal semantic disambiguation analysis on the standardized video data to obtain multimodal semantic disambiguation results includes: S21. Extract multimodal basic features and contextual features from the standardized video data to obtain a multimodal feature set containing multimodal basic features and contextual features; S22. Use a pre-defined corpus in the vertical domain to perform semantic disambiguation on the multimodal feature set to obtain an initial multimodal semantic disambiguation result; S23. Introduce user interaction data to correct the initial multimodal semantic disambiguation results and obtain the final multimodal semantic disambiguation results.
[0009] Preferably, generating hierarchical labels for vertical domains based on the multimodal semantic disambiguation results includes: S31. Construct dynamic knowledge graphs for vertical domains; S32. Based on the dynamic knowledge graph, automatically label and classify emerging terms in the multimodal semantic disambiguation results, and classify the emerging terms into the parent category of the existing terms that match them in the dynamic knowledge graph. S33. Generate hierarchical labels based on the aforementioned superior classification to obtain hierarchical labels.
[0010] Preferably, the dynamic knowledge graph includes an interconnected core knowledge layer and a hot knowledge layer. The hot knowledge layer captures emerging terms in real time and automatically associates the emerging terms with the corresponding basic terms in the core knowledge layer.
[0011] Preferably, the step of performing cross-platform unified popularity ranking on the standardized video data based on the hierarchical tags to obtain popularity data includes: S41. Perform cross-platform indicator equivalence calibration on the interactive data in the standardized video data to obtain calibrated cross-platform interactive data; S42. Perform source identification on the cross-platform interactive data to obtain source-related video data; S43. Calculate the incremental popularity score for the source video data to obtain the popularity score. Q The expression is as follows:
[0012] in, P This represents a basic interaction metric for video data from the same source. Indicates configurable weights. This indicates the topic popularity adjustment value for layered tags; S44. Rank the data according to the popularity score to obtain the popularity data.
[0013] Preferably, the step of performing video content creation analysis based on the popularity data and the hierarchical tags, and outputting a video content creation plan, includes: S51. Obtain user behavior characteristics and compliance attribute characteristics; S52. The user behavior characteristics, compliance attribute characteristics, and hierarchical tag preference characteristics are fused to obtain a user profile; S53. Establish a mapping relationship between the user profile and the creative elements of the popular data with high popularity scores under the hierarchical tags of the user profile, and output the video content creation scheme according to the mapping relationship.
[0014] This invention also proposes a social media video analysis system, which, based on the aforementioned social media video analysis method, includes: The multi-source data preprocessing module is used to collect multi-source social media video data and preprocess the multi-source social media video data to obtain standardized video data. The multimodal semantic disambiguation analysis module is used to perform multimodal semantic disambiguation analysis on the standardized video data to obtain multimodal semantic disambiguation results; The hierarchical label generation module is used to generate hierarchical labels for vertical domains based on the multimodal semantic disambiguation results. The cross-platform popularity ranking module is used to perform cross-platform unified popularity ranking of the standardized video data based on the hierarchical tags to obtain popularity data; The creation analysis module is used to perform video content creation analysis based on the popularity data and the hierarchical tags, and output a video content creation plan.
[0015] Preferably, it also includes a compliance and containerized deployment support module, used to ensure compliance and scalability of the data output by the multi-source data preprocessing module, the multimodal semantic disambiguation analysis module, the hierarchical tag generation module, the cross-platform popularity ranking module, and the creation analysis module.
[0016] Compared with the prior art, the beneficial effects of the technical solution of the present invention are: This invention proposes a social media video analysis method and system. The method first preprocesses multi-source social media video data and performs multimodal semantic disambiguation analysis on standardized video data to avoid semantic ambiguity of multimodal features, achieving accurate semantic matching of multi-source features and ensuring the integrity of the analysis process. Secondly, based on the multimodal semantic disambiguation results, hierarchical tags for vertical domains are generated, ensuring that the hierarchical tag system can quickly adapt to changes in vertical domain trends. Thirdly, based on the hierarchical tags, the standardized video data is ranked across platforms to obtain trending data, achieving unified cross-platform trending ranking of social media videos and eliminating the distortion problem of cross-platform analysis results. The final output video content creation scheme accurately reflects social media trending topics and significantly improves the timeliness, accuracy, and vertical domain adaptability of social media video analysis. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a social media video analysis method proposed in an embodiment of the present invention. Figure 2 This is another flowchart illustrating a social media video analysis method proposed in an embodiment of the present invention; Figure 3 This diagram illustrates the structure of a social media video analysis system proposed in an embodiment of the present invention. Detailed Implementation
[0018] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent. It is understandable to those skilled in the art that some well-known details may be omitted from the accompanying drawings; To facilitate understanding of this embodiment, the prior art information of this embodiment is first introduced as follows: Prior art 1 discloses a long video structured tag generation scheme, which focuses on semantic recognition and structured tag generation of long videos. The core technical process includes: Data Acquisition and Preprocessing: Acquire long video files, perform preprocessing such as format conversion and redundant data removal on the videos, and generate text content with timestamps (such as speech-to-text results). Video segmentation: Based on the preprocessed text content, the semantic relationships of the text are analyzed through a Large Language Model (LLM) to generate video segmentation suggestions and divide long videos into segments (such as splitting them according to "chapter logic"). Three-level tag generation: For each video segment, extract text keywords to generate a three-level structured tag system of "domain-topic-detail", and integrate it into the overall tag system of the long video through a tag fusion algorithm; Tag application: Based on the generated structured tags, it enables functions such as content retrieval and segment location in long videos.
[0019] The core objective of this solution is to address the problem of "difficulty in structuring and parsing long video content," and it is primarily applied to long video fields such as film and television, and education.
[0020] Existing technology 2 also discloses a VBR-encoded video content distribution scheme based on a hybrid network. This scheme focuses on optimizing video distribution efficiency, and its core technical process includes: Data Acquisition and Modeling: Obtain bitstream information of VBR (Variable Bit Rate) encoded video and statistical data on the popularity of online videos (such as user clicks and playback counts) to establish a user click model; Video segmentation and distribution strategy design: Determine the segment length and number based on the characteristics of the video bitstream, and construct a cost model and storage usage model in combination with real-time network conditions (such as bandwidth and latency); Linking popularity with user profiles: Generate video popularity rankings based on click data, and construct simple user profiles (such as "user groups who frequently click on a certain type of video") by combining user click behavior. Dynamic distribution scheduling: Based on the cost model and popularity ranking, the distribution path of video segments is dynamically selected to achieve "low-cost and high-smoothness" video distribution.
[0021] The core objective of this solution is to address the issue of "network adaptation and cost optimization for video distribution," and it is primarily applied to content transmission scenarios on video platforms.
[0022] While the aforementioned combination of existing technologies and common knowledge can achieve the basic functions of video analytics, the following problems exist when adapting to the core scenarios of social media videos that are multimodal, cross-platform, highly vertical, highly dynamic, and compliant: (1) The semantic ambiguity problem of multimodal features has not been resolved, and the label accuracy cannot meet the needs of social media; Existing technology 1 generates tags based solely on text content, failing to consider the semantic relationships between visual and speech features in social media videos, as well as ambiguous scenarios where the same features have different meanings, for example: In beauty videos, the vocal characteristic of "drying" could correspond to either "the texture of the lipstick is dry" or "the skin is dry after applying makeup." The visual feature of crying in mother-infant videos may correspond to different semantics such as hunger, gas, and drowsiness.
[0023] Existing technology 1 cannot accurately distinguish the semantics of these multimodal features, resulting in a discrepancy between the tags and the actual content of the social media video, which in turn affects the accuracy of subsequent user profiling and business guidance.
[0024] (2) Spatiotemporal bias and index heterogeneity of cross-platform data lead to distortion of analysis results across multiple platforms. Social media videos commonly feature the distribution of the same content across multiple platforms, but existing technologies 2 and their combined solutions have not resolved the issue of cross-platform data equivalence. Heterogeneous metrics: Different platforms have different definitions of "interaction metrics" (e.g., "play count" on X-Sound = 1 full play, "play count" on Y-Book = counted when click to play, "play count" on Z-Station = counted when video page is opened). Existing technology directly uses the original metrics for ranking, resulting in "contradictory popularity rankings of the same video on different platforms". Spatiotemporal discrepancy: The update cycle of interactive data on different platforms is different (A-site updates every 5 minutes, Z-site updates every 15 minutes). Existing technology does not perform timestamp calibration, resulting in time gaps when splicing cross-platform data. For example, a video may have garnered 1,000 likes on Weibo, but the data on Z-site has not yet been updated, leading to a misjudgment of low popularity.
[0025] The aforementioned issues prevent existing technologies from achieving unified cross-platform analysis of social media videos, thus failing to meet actual industry needs.
[0026] (3) Vertical domain tags are dynamically lagging and cannot cover emerging hot topics. Social media trends in specific sectors evolve rapidly. For example, in the beauty industry, the concept of "morning vitamin C, evening vitamin A" quickly spawned "morning vitamin P, evening vitamin R" and "morning vitamin C, evening vitamin B," while in the maternal and infant industry, "milk powder formula" evolved into "comparison of the new national standard for lactoferrin." However, existing technologies have significant limitations. The existing three-level tagging system of technology 1 is a static preset without a real-time iteration mechanism. When new terms or hot topics emerge in the vertical field, the tagging system cannot adapt quickly. For example, after "early P, late R" appears, the existing tags are still "early C, late A", which makes it impossible to accurately classify related videos. Existing technology combination solutions lack a "hotspot verification and elimination" mechanism, and the addition of new tags is prone to "redundancy and invalidity" issues.
[0027] Ultimately, this resulted in the existing tagging system being unable to support "real-time hot topic insights" in the social media vertical field, leading to a significant decline in business value.
[0028] (4) There is a contradiction between compliance and anonymization and analytical accuracy, making it impossible to balance "compliance" and "practicality". Social media videos contain a large amount of user privacy information, while core analysis relies on detailed visual features; current technology cannot balance these two aspects. Either over-desensitization, using simple methods such as blurring the whole face and cropping the background, leads to the loss of core analytical features. For example, after blurring the baby's face, it is impossible to determine whether the crying is due to an allergic reaction to complementary food. Either the desensitization is insufficient, only filtering obvious private information and omitting hidden or dynamic private information in the comments section, which poses a compliance risk.
[0029] This contradiction means that existing technologies either cannot be implemented due to compliance issues or lose their analytical value due to insufficient accuracy.
[0030] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0031] Example 1 like Figure 1 and Figure 2 As shown, this embodiment proposes a social media video analysis method, including the following steps: S1. Collect multi-source data of social media videos, and preprocess the multi-source data of social media videos to obtain standardized video data; In S1, a multi-source cross-platform intelligent acquisition mechanism is used to collect multi-source data of social media videos. The multi-source cross-platform intelligent acquisition mechanism is as follows: The data collection architecture employs a configurable API gateway and an intelligent anti-crawling distributed crawler cluster. The API gateway has a built-in self-learning sub-unit for platform compliance rules, which can synchronize the privacy policies and data permission updates of various social media platforms in real time and automatically adjust the collection scope. For example, when a platform adds a rule to "prohibit the collection of avatars in the comment section", the gateway will immediately filter such data. It supports configuring the collection cycle according to the needs of vertical fields, such as collecting once every 30 minutes in hot fields and once every 2 hours in regular fields. The distributed crawler cluster is equipped with anti-crawling strategy self-learning and rate limiting adaptation subunits. It can dynamically adjust the request interval and crawler strategy based on the rate limiting frequency and CAPTCHA type of the target platform, while strictly following the robots protocol of the target platform to ensure compliance of data collection.
[0032] The preprocessing of the multi-source social media video data to obtain standardized video data includes: S11. The social media video multi-source data is classified using a preset two-dimensional classification evaluation model to obtain two-dimensional classification data; The dual-dimensional hierarchical assessment model described in S11 includes a parallel privacy risk level layer and an analysis value level layer; the privacy risk level layer is used to classify the social media video multi-source data into four levels: extremely high privacy, high privacy, medium privacy, and low privacy; the analysis value level layer is used to classify the social media video multi-source data into four levels: extremely high value, high value, medium value, and low value.
[0033] S12. Dynamically desensitize the dual-dimensional hierarchical data to obtain the data desensitization results; for example, local feature retention desensitization is used for "extremely high privacy - extremely high value" data, such as covering the face with a sticker to retain facial expression features for sentiment analysis; "extremely high privacy - low value" data is directly deleted; "low privacy - extremely high value" data is not desensitized to ensure analysis accuracy; S13. Perform a quality score assessment on the data anonymization results to obtain a quality score; In S13, the data anonymization results are scored using an AI model, with scoring dimensions including video clarity, authenticity of interactive data, and content completeness.
[0034] S14. Determine whether the quality score is greater than a preset threshold. If so, standardize the format and structure of the data desensitization result to obtain standardized video data. In S14, the data anonymization results are standardized in terms of format and structure, including: unifying the video encoding format and resolution, such as H.265, and configuring the resolution according to the needs of the vertical industry, such as 720P for short video platforms and 1080P for long video platforms; and storing the interactive data in a structured manner according to "platform-video ID-timestamp" in a time-series database to provide a unified data format for subsequent cross-platform analysis.
[0035] S15. Determine whether the standardized video data belongs to a new field or a new platform. If so, perform cold start cross-domain adaptation on the standardized video data to obtain standardized video data adapted to the new field or platform. If not, directly output the original standardized video data.
[0036] In S15, for new fields or new platforms, the core technical solution for building rapid analysis capabilities comprises three major sets of technical features: First, a domain similarity assessment mechanism Construct a domain similarity evaluation matrix and screen transferable historical models: Evaluation dimensions: content type similarity, user group similarity; Similarity calculation: The cosine similarity algorithm is used to calculate the similarity between the new domain and the existing domain. If the similarity is ≥70%, the model of the existing domain (label rules, popularity weight, and profile features) is used as the basic transfer model.
[0037] Second, cross-domain model transfer and small-sample fine-tuning mechanisms. Used to reduce reliance on data in new fields and quickly adapt to analytical needs; Model parameter transfer retains the core parameters of the base transfer model and only adjusts parameters that are strongly related to the domain. Small-sample fine-tuning requires only 50-100 labeled data points from the new domain. The transfer model is fine-tuned using a small-sample learning algorithm, and model adaptation is completed within 72 hours. Automatic annotation assistance is used to automatically annotate unlabeled data in new fields using a finely tuned model, while manual review is only required for data with a confidence level of less than 80%, thus reducing annotation costs.
[0038] Thirdly, a dynamic iterative optimization mechanism As data from new fields accumulates, the model accuracy is continuously optimized: once the data reaches a preset threshold (such as 500 records), the model is automatically fine-tuned a second time. Every 14 days, based on interaction data and user feedback in new fields, we adjust the tagging rules and gradually improve the analysis accuracy to be consistent with that of mature fields.
[0039] S2. Perform multimodal semantic disambiguation analysis on the standardized video data to obtain the multimodal semantic disambiguation results; In S2, the multimodal semantic disambiguation analysis performed on the standardized video data to obtain the multimodal semantic disambiguation result includes: S21. Extract multimodal basic features and contextual features from the standardized video data to obtain a multimodal feature set containing multimodal basic features and contextual features; In S21, context-related feature extraction breaks through the limitations of single feature extraction in existing technologies. It simultaneously extracts context-related features while extracting video visual features and speech semantic features. Contextual features include visual context features, speech context features, and text context features. Visual context features include: when extracting a "crying scene", simultaneously capturing related elements such as "baby bottle, diaper, and abdominal distension" in the scene; speech context features include: when extracting the keyword "dry", simultaneously capturing the sentences before and after it, such as "this lipstick is drying" vs. "my skin is dry after using it"; text context features include: when extracting a comment that is suitable for sensitive skin, simultaneously capturing the number of likes and replies to that comment.
[0040] S22. Use a pre-defined corpus in the vertical domain to perform semantic disambiguation on the multimodal feature set to obtain an initial multimodal semantic disambiguation result; In S22, a dedicated semantic disambiguation corpus is built for each vertical domain. The corpus includes: multi-scenario definitions of domain-specific terms and multi-modal feature semantic mapping tables. For example, in the beauty domain, "drying" corresponds to two semantic categories: "dry lipstick texture" and "dry skin," along with their corresponding features; in the maternal and infant domain, "bloating" corresponds to two semantic categories: "improper feeding" and "colic," along with their corresponding features. Multimodal feature semantic mapping tables, such as "dry lipstick texture" corresponding to the visual feature "obvious lip lines" and the voice feature "dry on the lips"; "dry skin" corresponding to the visual feature "facial makeup cakes" and the voice feature "dry and flaky face"; This step's corpus supports dynamic updates, supplementing the corpus every 7 days based on the multimodal features of newly added videos and user feedback data to ensure the timeliness of semantic disambiguation.
[0041] S23. Introduce user interaction data to correct the initial multimodal semantic disambiguation results and obtain the final multimodal semantic disambiguation results.
[0042] In S23, user interaction data such as comments, replies, and likes are introduced as the basis for semantic disambiguation verification. If multimodal analysis initially determines that a video tag is "lipstick is drying", but the comment section frequently mentions "my face is so dry after using it", then the tag semantics are automatically corrected to "skin is drying". The feature "drying" keyword + facial powder caking + comment feedback are added to the semantic disambiguation corpus. For videos with semantic ambiguity ≥50%, they are pushed to the manual review end, where reviewers combine the corpus and interaction data to determine the final semantics, ensuring the accuracy of the analysis.
[0043] S3. Based on the multimodal semantic disambiguation results, generate hierarchical labels for vertical domains; In S3, generating hierarchical labels for vertical domains based on the multimodal semantic disambiguation results includes: S31. Construct a dynamic knowledge graph for the vertical domain; the dynamic knowledge graph includes an interconnected core knowledge layer and a hot topic knowledge layer. The hot topic knowledge layer captures emerging terms in real time and automatically associates these emerging terms with the corresponding basic terms in the core knowledge layer. The core knowledge layer includes basic domain classifications, basic terms, and their relationships; the hot topic knowledge layer captures emerging terms in real time by connecting to the vertical domain's real-time public opinion interface and automatically associates them with the core knowledge layer; the iteration cycle of this dynamic knowledge graph is configurable, iterating every 12 hours for hot topic domains and every 24 hours for regular domains to ensure coverage of the latest hot topics.
[0044] S32. Based on the dynamic knowledge graph, automatically label and classify emerging terms in the multimodal semantic disambiguation results, and classify the emerging terms into the parent category of the existing terms that match them in the dynamic knowledge graph. In S32, the automatic labeling and classification of emerging terms is achieved based on semantic similarity algorithms and dynamic knowledge graph association rules as follows: Semantic similarity calculation: The semantic similarity of emerging terms with existing terms in the knowledge graph is compared. If the similarity is ≥60%, the emerging term is classified into the parent category of the existing term that matches it in the dynamic knowledge graph, such as "beauty - skin care process". S33. Generate hierarchical labels based on the aforementioned superior classification to obtain hierarchical labels.
[0045] In S33, based on the knowledge graph association, the upper-level classification is inherited, and "level 1-level-level-level 3-level" hierarchical labels are assigned to emerging terms, and detailed sub-labels are automatically generated.
[0046] Set a "72-hour validity verification period" for newly added tags: Statistically calculate the percentage of interaction volume of related videos within the verification period (e.g., the percentage of views of videos with the "Morning P Evening R" tag to the total views of the "Skincare Routine" category). If the percentage is ≥15%, the tag will be converted into an "official tag"; if the percentage is <5%, it will be automatically eliminated to avoid tag redundancy. Supports custom tag rules for vertical fields: Users can manually add tags (such as "suitable for sensitive skin") and bind trigger conditions (such as voice containing "sensitive skin", video containing "sensitive skin test report", and comments containing "suitable for sensitive skin"). After the rules take effect, videos that meet the conditions will be automatically matched.
[0047] S4. Based on the hierarchical tags, perform cross-platform unified popularity ranking on the standardized video data to obtain popularity data; S5. Based on the popularity data and the hierarchical tags, perform video content creation analysis and output a video content creation plan.
[0048] In this embodiment, firstly, preprocessing of multi-source social media video data and multimodal semantic disambiguation analysis of standardized video data are performed to avoid semantic ambiguity of multimodal features, achieving accurate semantic matching of multi-source features of social media videos and ensuring the integrity of the analysis process. Secondly, based on the results of multimodal semantic disambiguation, hierarchical tags for vertical fields are generated to ensure that the hierarchical tag system can quickly adapt to changes in hot topics in vertical fields. Thirdly, the standardized video data is ranked across platforms based on the hierarchical tags to obtain popularity data, achieving unified popularity ranking of social media videos across platforms and eliminating the problem of distortion in cross-platform analysis results. The final output video content creation scheme can accurately reflect social media hot topic trends and significantly improve the timeliness, accuracy, and vertical field adaptability of social media video analysis.
[0049] Example 2 This embodiment further explains S4, which involves performing a cross-platform unified popularity ranking on the standardized video data based on the hierarchical tags to obtain popularity data, including: S41. Perform cross-platform indicator equivalence calibration on the interactive data in the standardized video data to obtain calibrated cross-platform interactive data; In S41, a platform indicator mapping rule base and a timestamp synchronization calibration algorithm are constructed to solve the problems of cross-platform data heterogeneity and spatiotemporal deviation, and eliminate the problem of distortion of cross-platform analysis results. The platform metric mapping rule base establishes an equivalent conversion model for each platform metric by analyzing the historical interaction data of the same video on multiple platforms (e.g., Y-book play count × 0.3 ≈ X-video play count, X-site comment count × 1.2 ≈ Weibo comment count). The conversion model is updated every 7 days based on new data. The timestamp synchronization calibration is based on the "first release time of the video". The interactive data of each platform is aligned according to the "configurable time slice" (such as one slice every 10 minutes). For the data of the platform that has not been updated, the "linear interpolation method" is used to temporarily fill in the gaps. For example, if a platform updates once every 15 minutes, only one ranking record is retained to avoid data redundancy.
[0050] S43. Calculate the incremental popularity score for the source video data to obtain the popularity score. Q The expression is as follows:
[0051] in, P This represents a basic interaction metric for video data from the same source. This indicates configurable weights that can be adjusted based on vertical sectors. This represents the topic popularity correction value for the hierarchical tag, calculated based on the average interaction volume of all videos under that hierarchical tag. In S43, the incremental popularity score calculation combines incremental calculation and hot and cold data separation to improve ranking efficiency, while incorporating multimodal features and tag popularity correction. Incremental calculation recalculates the popularity score only for videos whose interaction data changes by more than or equal to a preset threshold (e.g., ≥50 new likes, configurable) or whose tag topic popularity is updated; videos with no changes retain their historical scores to avoid a full recalculation; hot and cold data separation uses a streaming engine to update the ranking of videos with active interaction within the past 24 hours, while videos with cold data that have been active for more than 24 hours and whose interaction tends to be stable are updated in ranking by an offline calculation engine on an hourly basis. S44. Rank the data according to the popularity score to obtain the popularity data.
[0052] This embodiment further explains S5, which involves performing video content creation analysis based on the popularity data and the layered tags, and outputting a video content creation scheme, including: S51. Obtain user behavior characteristics and compliance attribute characteristics; S52. The user behavior characteristics, compliance attribute characteristics, and hierarchical tag preference characteristics are fused to obtain a user profile; The user behavior features are constructed based on the user's video viewing behavior (percentage of effective viewing time, viewing frequency, and viewing time period) and interaction behavior (frequency of comments / likes / shares, and keywords of interactive content); The preference features of the hierarchical tags are calculated based on the tag distribution of the videos watched by users (such as the proportion of time spent watching videos with the tag "Beauty - Lipstick Review") (the higher the proportion of time spent watching, the higher the preference). The preference is also strengthened by combining interaction data (such as high-frequency comments on videos with a certain tag, which increases the preference for that tag). The compliance attribute features are obtained by introducing third-party compliance data (such as user region, age range, and spending power level) and cross-validating it with behavioral / preference features (such as "frequent viewing of maternal and infant tags + first-tier cities" → inferred to be the "young mother" group). User profiles are updated weekly. If a user’s recent behavior / preference changes by more than or equal to a preset threshold, an immediate update is triggered to ensure the timeliness of the user profile.
[0053] S53. Establish a mapping relationship between the user profile and the creative elements of the popular data with high popularity scores under the hierarchical tags of the user profile, and output the video content creation scheme according to the mapping relationship.
[0054] In S53, the creative elements include video theme (such as "Sensitive Skin Morning Vitamin C and Evening Vitamin A Adaptation Solution"), visual design (such as "Practical Demonstration + Subtitle Annotation"), voice style (such as "Gentle Explanation + Precautions Tips"), and release time (such as "8-10 PM" for the maternal and infant field, which is during parents' free time). The video content creation scheme generation process is as follows: Based on the core preference tags of the user profile, combined with the creative elements of high-popularity videos under the tags, 3-5 video content creation schemes are automatically generated, and manual adjustment and optimization are supported.
[0055] The social media video analysis method proposed in this embodiment also includes compliance and scalability assurance for the data output in each step of S1-S5. The compliance assurance adopts a full-link data compliance assurance mechanism, and the scalability assurance adopts a containerized elastic deployment mechanism. The full-chain data compliance assurance mechanism covers compliance control throughout the entire process of "collection stage - storage stage - analysis stage - output stage". In the collection stage, all collection activities are subject to a platform authorization agreement (API collection) or follow the robots protocol (web crawling), and the collection source and time are recorded synchronously. In the storage stage, sensitive data (such as user profile data) is encrypted and stored using encryption algorithms such as AES-256. All data operations (collection, analysis, export) generate tamper-proof audit logs, which are retained for ≥6 months (in compliance with domestic and international data regulations). In the output stage, the scope of user data export is restricted, allowing only the export of non-privacy statistical data (such as tag distribution, profile group proportion), and prohibiting the export of personally identifiable information (PII).
[0056] The containerized elastic deployment mechanism is based on containerized platforms such as Kubernetes to achieve functional deployment and resource scheduling; all functional modules are packaged as independent containers, and functions support individual deployment and upgrades; resource scheduling is based on resource utilization (such as CPU utilization ≥70%, memory utilization ≥80%) to automatically expand / shrink, and during peak periods in hot fields (such as "618" and "Double 11"), the number of container instances for collection and analysis is automatically increased; when adding a new social media platform or vertical field, it can be quickly integrated by developing "collection plugins" and "tag rule plugins" without modifying the core code (for example, when adding "Y Book New Section", a dedicated collection plugin can be developed).
[0057] Addressing the six core challenges in social media video analytics—multimodal semantic ambiguity, cross-platform data heterogeneity, tagging lag, compliance-accuracy contradictions, cold start deficiencies, and lack of business closed loops—this embodiment, combining the aforementioned embodiments, employs six core technical solutions: multimodal semantic disambiguation, cross-platform calibration, dynamic tagging, tiered desensitization, cold start migration, and business closed loops. These solutions significantly improve five dimensions: technical accuracy, analytical efficiency, compliance, scenario adaptability, and commercial value conversion. Each beneficial effect has been verified through comparison with existing technologies, as detailed below. I. Multimodal label accuracy is significantly improved, resolving analytical biases caused by different meanings for the same features. Existing technology 1 relies solely on text content to generate tags, failing to distinguish between "semantic ambiguities in visual-voice features" in social media videos. For example, the accuracy rate of the "drying" tag in beauty videos is only 60%-65% (mistakenly classifying "dry skin" as "dry lipstick texture"), and the accuracy rate of the "reason for crying" tag in maternal and infant videos is less than 55% (failing to distinguish between "hunger," "gas," and "drowsiness"), directly leading to deviations in subsequent user profiling and business guidance.
[0058] This invention achieves a breakthrough improvement in multimodal tag accuracy through a three-dimensional mechanism: context-related feature extraction, a corpus for vertical domain semantic disambiguation, and user interaction feedback verification. It enables accurate data verification and enhances error correction capabilities. In the beauty industry, the accuracy of tags for ambiguous features such as "drying" and "cakey makeup" has increased from 65% to over 92% in the existing technology. In the maternal and infant industry, the accuracy of the "reason for crying" tag has increased from 55% to over 88%. Regarding error correction capabilities, user interaction feedback verification can correct initial semantic ambiguities (such as misjudging "lipstick is drying") by up to 90%, and the corpus, updated every 7 days, allows tag accuracy to continuously improve with data accumulation (monthly accuracy fluctuation ≤3%). This invention is the first to achieve the linkage of multimodal context, domain knowledge and user feedback, upgrading from single text dependency to multi-dimensional semantic verification, and completely solving the industry pain point of different meanings for the same features.
[0059] Second, cross-platform analysis consistency has been significantly improved, eliminating ranking distortion caused by indicator heterogeneity and spatiotemporal bias. Existing technology 2 can only achieve cross-platform data format unification, but cannot handle differences in interactive indicator definitions and asynchronous update cycles. For example, the popularity ranking of the same beauty video on Douyin and Xiaohongshu differs by 30%-40%. Douyin has a higher weight for play counts and a higher weight for comment counts. The time stamp gap in cross-platform data leads to a popularity misjudgment rate of over 25%, which cannot support unified user insights across multiple platforms.
[0060] This invention achieves equivalence and consistency in cross-platform analysis through platform metric mapping, timestamp synchronization calibration, and multimodal source verification. Regarding metric equivalence, the conversion error of interactive metrics across platforms is reduced from 30% in existing technologies to below 8%. For example, the model where a certain number of views × 0.3 ≈ Douyin views achieves an accuracy of 92%, trained on 50,000 source videos. Regarding spatiotemporal consistency, after timestamp calibration, the misjudgment rate of time discontinuities in cross-platform data is reduced from 25% to below 5%. For instance, data updated every 5 minutes on one blog and every 15 minutes on another site can be synchronously compared within a 10-minute slice. Regarding source deduplication efficiency, the accuracy of multimodal feature fingerprint recognition for source videos reaches 95%, avoiding data redundancy caused by duplicate rankings and improving cross-platform analysis efficiency by 40%. This invention is the first to construct a complete cross-platform processing link that achieves equivalence of indicators, time synchronization, and deduplication of sources, thus meeting the industry requirement of unified analysis of social media videos across multiple platforms.
[0061] Third, the response speed of vertical domain tags has been improved, adapting to the social media characteristic of "high-frequency iteration of hot topics". The existing labeling system of technology 1 is static and preset. The generation cycle of labels for emerging hot topics in vertical fields is as long as 7-10 days (for example, in the beauty industry, it takes 1 week from the emergence of "morning P and evening R" to its inclusion in the label). Moreover, there is no redundant label elimination mechanism. For example, after 3 months of use, the redundant label ratio of a certain labeling system in the maternal and infant field reached 35%, resulting in a 20% decrease in analysis efficiency.
[0062] This invention achieves rapid iteration and precise weight reduction of layered tags through real-time public opinion capture, automatic hotspot classification, and 72-hour validity verification. Regarding hotspot response speed, the generation cycle for emerging tags in vertical fields is shortened from 7 days in existing technologies to less than 12 hours. For example, in the beauty industry in Q2 2024, the "morning P, evening R" hotspot was resolved within 12 hours using this invention, while existing technologies require 7 days. Regarding tag validity, the 72-hour verification mechanism can eliminate 90% of "low-value tags" (e.g., the "milk powder packaging design" tag in a certain maternal and infant field was automatically eliminated because its interaction rate was <5%), keeping the tag system redundancy rate below 8% and improving analysis efficiency by 30%. This invention upgrades layered tagging from static maintenance to automatic hotspot following, perfectly adapting to the short, frequent, and fast-paced characteristics of social media hotspots.
[0063] IV. Achieving a balanced approach between compliance and analytical accuracy to resolve the contradiction between "over-desensitization resulting in feature loss and under-desensitization posing risks". Existing desensitization solutions are one-size-fits-all, either over-desensitizing (e.g., blurring the baby's face in a mother-infant video while mistakenly blurring the key feature of "feeding posture," causing the posture recognition rate to drop from 85% to 40%), or under-desensitizing (e.g., omitting hidden privacy information such as "names" in the comments section, resulting in a compliance risk rate exceeding 15%), failing to meet the dual requirements of accuracy in social media analysis.
[0064] This invention achieves a synergistic balance between "compliance anonymization" and "analysis accuracy" through a "privacy-value dual-dimensional assessment model": Regarding accuracy preservation, this invention employs "local feature preservation anonymization" for "high privacy-high value" data (such as breastfeeding postures in mother-infant videos and facial effects in beauty videos), maintaining a recognition rate of over 90% for key features (e.g., the recognition rate for feeding postures has rebounded from 40% in existing technologies to 92%); Regarding compliance, this invention achieves a 100% compliance rate for privacy anonymization (based on testing of 10,000 social media videos by a third-party compliance testing agency, with no "privacy information leakage" issues), and the end-to-end audit logs meet the "traceability" requirements of domestic and international regulations; This invention introduces the dimension of analytical value for the first time, achieving a dynamic balance of "protecting what needs to be protected and de-identifying what needs to be de-identified," filling the technical gap in "compliance and accuracy synergy."
[0065] V. The efficiency of accessing new scenarios is improved exponentially, breaking the vicious cycle of "cold start data scarcity". Existing cold start solutions rely on thousands of labeled data points plus a 1-2 month training period. For example, when accessing emerging vertical fields such as "outdoor frisbee", the cold start period can be as long as 45 days due to insufficient data accumulation, missing the 1-2 week hot spot explosion period and causing the analysis to lose its business value.
[0066] This invention achieves a "rapid cold start" in new domains / platforms through domain similarity assessment, cross-domain model transfer, and small-sample fine-tuning. This invention reduces data dependency, requiring only 50-100 labeled data points for new fields (compared to 5000+ in existing technologies), thus reducing data collection costs by 90%. The time efficiency is improved, and the cold start cycle is shortened from 45 days in the existing technology to less than 72 hours (for example, in the hot field of "urban cycling" in 2024, the present invention can complete the construction of the tag system and the deployment of the heat model within 3 days, while the existing technology requires 1.5 months). With guaranteed accuracy, the model, after fine-tuning with small samples, achieves over 85% accuracy in mature fields (e.g., the label accuracy for the "outdoor frisbee" field reaches 86%, close to 92% in the mature beauty field), and can be rapidly improved to 90%+ with data accumulation; This invention upgrades from "data-driven" to "knowledge transfer + small data optimization" through "cross-domain transfer + few-sample learning", completely breaking the "data scarcity vicious cycle" of cold start.
[0067] VI. Implementing a closed-loop technology-business model to transform analytical results into commercial value. Existing technologies only output basic analysis results such as tag statistics and popularity rankings, without providing guidance for business implementation. For example, a beauty brand may know that "users prefer matte lipsticks" through existing technologies, but cannot obtain specific creative suggestions such as "video theme, visual design, and release time". This leads to a disconnect between technical analysis and business application, resulting in a commercial value conversion rate of less than 30%.
[0068] This invention, through a closed loop from multi-dimensional user profiling to content creation guidance, directly transforms technical analysis into commercial value, improves content creation efficiency and interaction, and optimizes business conversion, as detailed below: Regarding content creation efficiency, the automatically generated "topic + visuals + audio + release time" creation suggestions reduce the content production cycle from 72 hours with existing technology to 24 hours, improving efficiency by 67%. To boost engagement, precise content creation guidance based on user profiles has increased the average engagement (likes + comments + shares) of social media videos by 40% (e.g., a beauty brand's "Sensitive Skin Morning C Evening A" video saw engagement increase from 12,000 to 20,000). To optimize business conversion, the "comment Q&A conversion rate" for parenting-related content in the maternal and infant field has increased from 15% to 25% (users actively participate in the interaction by solving practical problems through guidance content). This invention is the first to construct a business closed loop of analysis, guidance, creation, and re-analysis, directly transforming technological value into commercial value such as improved content efficiency, increased user interaction, and optimized business conversion, filling the industry gap of the technology-business gap.
[0069] This embodiment, combining the above embodiments, achieves a systematic upgrade of social media video analytics from basic functional fulfillment to high precision, high efficiency, high compliance, high adaptability, and high value through the synergy of six core technologies. Each effect of this invention is supported by clear technical technology and verified by data, solving long-standing industry problems and realizing a closed-loop transformation of technological precision into commercial value.
[0070] Example 3 See Figure 3 This embodiment also proposes a social media video analysis system, wherein the social media video analysis method described in the above embodiment is implemented by the system including: The multi-source data preprocessing module is used to collect multi-source social media video data and preprocess the multi-source social media video data to obtain standardized video data. The multimodal semantic disambiguation analysis module is used to perform multimodal semantic disambiguation analysis on the standardized video data to obtain multimodal semantic disambiguation results; The hierarchical label generation module is used to generate hierarchical labels for vertical domains based on the multimodal semantic disambiguation results. The cross-platform popularity ranking module is used to perform cross-platform unified popularity ranking of the standardized video data based on the hierarchical tags to obtain popularity data; The creation analysis module is used to perform video content creation analysis based on the popularity data and the hierarchical tags, and output a video content creation plan.
[0071] It also includes a compliance and containerized deployment support module, which is used to ensure compliance and scalability of the data output by the multi-source data preprocessing module, the multimodal semantic disambiguation analysis module, the hierarchical tag generation module, the cross-platform popularity ranking module, and the creation analysis module.
[0072] It also includes a cold start cross-domain adaptation module, which is used to perform cold start cross-domain adaptation on standardized video data to obtain standardized video data adapted to new domains or new platforms; In this embodiment, the modules collaborate through the following logic to form a closed loop from technology to business: The multi-source data preprocessing module outputs compliant standardized video data, the multimodal semantic disambiguation module outputs multimodal semantic disambiguation results, the hierarchical tag generation module outputs vertical domain hierarchical tags, and the cross-platform popularity ranking module performs cross-platform unified popularity ranking of the standardized video data based on the hierarchical tags. The creation analysis module outputs accurate user profiles based on user behavior characteristics, compliance attribute characteristics, and hierarchical tag preference characteristics. Based on establishing a mapping relationship between the user profiles and the creation elements of popular data with high popularity scores under the hierarchical tags of the user profiles, the module outputs the video content creation plan based on the mapping relationship. New videos after the creation plan is implemented can re-enter the system for analysis by collecting data through the multi-source data preprocessing module, forming a business closed loop of analysis, guidance, creation, and re-analysis. The cold start cross-domain adaptation module provides rapid access capabilities for new domains and platforms, and the compliance and containerization module provides compliance and scalability guarantees for the entire process, ensuring the stable operation of the system in different scenarios. This system is not limited to a specific vertical domain or social media platform. As long as it conforms to the core logic of "social media video multimodal analysis, cross-platform adaptation, and dynamic adjustment of vertical domains," it falls within the protection scope of this solution.
[0073] In summary, this embodiment constructs a closed-loop social media video analysis system encompassing compliant data collection, precise multimodal analysis, dynamic tag generation, cross-platform popularity ranking, multi-dimensional user profiling, and business-oriented content guidance. This system provides high-precision, real-time, and compliant technical support for social media operations, content creation, and user insights in vertical sectors such as beauty, maternal and infant products, and outdoor sports, filling gaps in existing technologies for social media applications. The system uses multimodal semantic disambiguation as its core analysis method, cross-platform data equivalence as its ranking basis, dynamic knowledge graphs for tag support, cold start migration for scenario adaptation, and tiered anonymization as its compliance baseline. Seven functional modules are logically linked in the order of data input, data processing, feature analysis, tag generation, popularity calculation, user profile construction, and business output, while compliance and containerization support modules provide end-to-end assurance. Data interoperability between modules is achieved through standardized data interfaces, forming a closed loop of collection, analysis, application, and feedback, ensuring the integrity and practicality of the technical solution.
[0074] Example 4 This embodiment targets beauty videos on three major platforms: Douyin, Xiaohongshu, and Bilibili. The collaborative execution method of the various modules of the system described in the above embodiment is as follows: Multi-source data preprocessing module: Acquisition end: Collects video data from Douyin beauty topic page, Douyin beauty community, and Douyin beauty section every 30 minutes through "API gateway + distributed crawler".
[0075] Desensitization and standardization: For videos containing "blogger's facial close-up + product on-face effect", they are evaluated according to "high privacy (face) - extremely high value (product effect analysis)" and desensitized by "partially covering the face with stickers (preserving lip / eye features)"; the videos are uniformly encoded in H.265 with a resolution of 720P, and interactive data such as play count, likes, and comment keywords are stored in a structured manner.
[0076] Cold start cross-domain adaptation module: Domain similarity assessment: The "Morning P Evening R" (skincare routine of sun protection + repair) is compared with the existing "Morning C Evening A" domain (skincare routine of whitening + anti-aging). The content type similarity is 85% and the user group similarity is 90%, which determines that it is transferable.
[0077] Model transfer and few-shot fine-tuning: The "multimodal semantic disambiguation model" and "dynamic knowledge graph structure" in the "morning C evening A" domain were transferred. Only 50 videos labeled "morning P (sunscreen action / product)" and "evening R (repair action / product)" were used to fine-tune the model through a few-shot learning algorithm, and the adaptation was completed within 48 hours.
[0078] Multimodal semantic disambiguation analysis module: Context feature capture: Extract visual features related to "early P" (sunscreen product display, application action, outdoor scene), voice features ("Sunscreen must be sufficient", "Apply sunscreen before going out"), and text features (comment section "Sunscreen steps are so detailed"); visual features related to "evening R" (repair product texture, nighttime skincare scene), voice features ("Use repair essence at night"), and text features (comment section "Skin is so soft after repair").
[0079] Semantic disambiguation: The corpus of semantic disambiguation in the beauty field (including semantic mappings of "morning P → daytime sun protection" and "evening R → nighttime repair") is called. Combined with feedback from the comments section such as "This sun protection is enough for commuting" and "Repair essence is a must-have for staying up late", the initial feature ambiguity is corrected (such as avoiding misjudging "anti-tanning" as "anti-UV aging").
[0080] Layered tag generation module: Real-time public opinion capture: Connects with beauty KOLs' Weibo and Xiaohongshu posts, and captures topics related to "morning P and evening R" with over 500,000 views within 2 hours, triggering hot topic capture.
[0081] Tag generation and verification: Automatically categorize "Morning P, Evening R" into the secondary tag of "Beauty - Skincare Routine" and generate tertiary sub-tags "Morning P (Sunscreen)" and "Evening R (Repair)"; if videos with this tag account for 22% of the total views in the "Skincare Routine" category within 72 hours, they are verified as valid tags and the dynamic knowledge graph is updated.
[0082] Cross-platform popularity ranking module: Indicator equivalence calibration: Convert the number of plays on Douyin (TikTok) × 0.3, the number of likes on Douyin (Book of Books) × 1.2, and the number of favorites on Douyin (Site of Books) × 1.5 into a unified popularity value; Align the data of each platform by 10-minute segments according to the video's first release time.
[0083] Same Source Verification and Ranking: Identify videos with the same source by using "visual features (sunscreen product packaging) + voice keywords ('morning P evening R')" (such as the same video being distributed on Douyin and Douyin). After merging the interaction data, rank the videos under the "morning P evening R" tag based on incremental popularity.
[0084] Creation Analysis Module: User Profile Construction: By integrating user behavior (average viewing time of "Morning P and Evening R" videos is 85 seconds, and comment keywords are "sunscreen recommendation" and "repair steps"), tag preferences ("Morning P and Evening R" tag preference is 80%), and compliance attributes (women aged 25-35 in first-tier cities), a profile of "refined skincare, focusing on the details of daytime sun protection and nighttime repair" is generated.
[0085] Creative guidance output: Based on the characteristics of high-traffic videos ("real-person demonstration + step-by-step subtitles + product comparison"), the following suggestions are generated: "Theme: Morning P and Evening R nanny-level tutorial; Visuals: show the sun protection (commuting scenario) and repair (nighttime scenario) practical operation in different time periods; Audio: explain the principle and common mistakes of each step; Release time: 8-10 pm (target users' leisure time)."
[0086] Compliance and containerized deployment support module: end-to-end compliance: during data collection, the platform's robots protocol and privacy policy are followed; during storage, information such as user nicknames is encrypted; and during output, only non-privacy statistical data such as "tag distribution and profile group proportion" are exported.
[0087] Containerized deployment: Each module is deployed in the form of a container. During peak periods such as "early P and late R" when access volume surges, the container instances of the "multi-source collection" and "semantic analysis" modules are automatically expanded to ensure system stability.
[0088] This embodiment achieves rapid capture of emerging beauty trends through the collaboration of various modules, accurately analyzes the multimodal semantics of videos under the trend, avoids ambiguity, uniformly assesses popularity across platforms, identifies core high-quality content, generates feasible creation solutions based on user preferences, and ensures compliance and system flexibility throughout the entire process.
[0089] Tag generation time has been reduced from 7 days to 12 hours; multimodal semantic disambiguation accuracy has reached 93%; and cross-platform popularity ranking error has been reduced from 30% to 7%.
[0090] A beauty brand's "Morning P, Evening R" video, created based on the guidelines of this invention, saw a 45% increase in views compared to similar videos, and a 28% increase in "request for product links" interactions in the comments section. The accuracy of user profiling improved to 89%, providing support for subsequent targeted advertising.
[0091] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A social media video analysis method, characterized in that, Includes the following steps: S1. Collect multi-source data of social media videos, and preprocess the multi-source data of social media videos to obtain standardized video data; S2. Perform multimodal semantic disambiguation analysis on the standardized video data to obtain the multimodal semantic disambiguation results; S3. Based on the multimodal semantic disambiguation results, generate hierarchical labels for vertical domains; S4. Based on the hierarchical tags, perform cross-platform unified popularity ranking on the standardized video data to obtain popularity data; S5. Based on the popularity data and the hierarchical tags, perform video content creation analysis and output a video content creation plan.
2. The social media video analysis method according to claim 1, characterized in that, The preprocessing of the multi-source social media video data to obtain standardized video data includes: S11. The social media video multi-source data is classified using a preset two-dimensional classification evaluation model to obtain two-dimensional classification data; S12. Perform dynamic desensitization on the dual-dimensional hierarchical data to obtain the data desensitization results; S13. Perform a quality score assessment on the data anonymization results to obtain a quality score; S14. Determine whether the quality score is greater than a preset threshold. If so, standardize the format and structure of the data desensitization result to obtain standardized video data. S15. Determine whether the standardized video data belongs to a new field or a new platform. If so, perform cold start cross-domain adaptation on the standardized video data to obtain standardized video data adapted to the new field or platform. If not, directly output the original standardized video data.
3. The social media video analysis method according to claim 2, characterized in that, The dual-dimensional hierarchical assessment model includes a privacy risk level layer and an analytical value level layer. The privacy risk level layer is used to classify the social media video multi-source data into four levels: extremely high privacy, high privacy, medium privacy, and low privacy. The analytical value level layer is used to classify the social media video multi-source data into four levels: extremely high value, high value, medium value, and low value.
4. The social media video analysis method according to claim 1, characterized in that, The process of performing multimodal semantic disambiguation analysis on the standardized video data to obtain multimodal semantic disambiguation results includes: S21. Extract multimodal basic features and contextual features from the standardized video data to obtain a multimodal feature set containing multimodal basic features and contextual features; S22. Use a pre-defined corpus in the vertical domain to perform semantic disambiguation on the multimodal feature set to obtain an initial multimodal semantic disambiguation result; S23. Introduce user interaction data to correct the initial multimodal semantic disambiguation results and obtain the final multimodal semantic disambiguation results.
5. The social media video analysis method according to claim 1, characterized in that, The generation of hierarchical labels for vertical domains based on the multimodal semantic disambiguation results includes: S31. Construct dynamic knowledge graphs for vertical domains; S32. Based on the dynamic knowledge graph, automatically label and classify emerging terms in the multimodal semantic disambiguation results, and classify the emerging terms into the parent category of the existing terms that match them in the dynamic knowledge graph. S33. Generate hierarchical labels based on the aforementioned superior classification to obtain hierarchical labels.
6. The social media video analysis method according to claim 5, characterized in that, The dynamic knowledge graph includes an interconnected core knowledge layer and a hot knowledge layer. The hot knowledge layer captures emerging terms in real time and automatically associates these emerging terms with the corresponding basic terms in the core knowledge layer.
7. The social media video analysis method according to any one of claims 1-6, characterized in that, The process of performing a cross-platform unified popularity ranking of the standardized video data based on the hierarchical tags to obtain popularity data includes: S41. Perform cross-platform indicator equivalence calibration on the interactive data in the standardized video data to obtain calibrated cross-platform interactive data; S42. Perform source identification on the cross-platform interactive data to obtain source-related video data; S43. Calculate the incremental popularity score for the source video data to obtain the popularity score. Q The expression is as follows: in, P This represents a basic interaction metric for video data from the same source. Indicates configurable weights. This indicates the topic popularity adjustment value for layered tags; S44. Rank the data according to the popularity score to obtain the popularity data.
8. The social media video analysis method according to any one of claims 1-6, characterized in that, The process of analyzing video content creation based on the popularity data and the hierarchical tags, and outputting a video content creation plan, includes: S51. Obtain user behavior characteristics and compliance attribute characteristics; S52. The user behavior characteristics, compliance attribute characteristics, and hierarchical tag preference characteristics are fused to obtain a user profile; S53. Establish a mapping relationship between the user profile and the creative elements of the popular data with high popularity scores under the hierarchical tags of the user profile, and output the video content creation scheme according to the mapping relationship.
9. A social media video analysis system, said system being implemented based on the social media video analysis method as described in any one of claims 1-8, characterized in that, include: The multi-source data preprocessing module is used to collect multi-source social media video data and preprocess the multi-source social media video data to obtain standardized video data. The multimodal semantic disambiguation analysis module is used to perform multimodal semantic disambiguation analysis on the standardized video data to obtain multimodal semantic disambiguation results; The hierarchical label generation module is used to generate hierarchical labels for vertical domains based on the multimodal semantic disambiguation results. The cross-platform popularity ranking module is used to perform cross-platform unified popularity ranking of the standardized video data based on the hierarchical tags to obtain popularity data; The creation analysis module is used to perform video content creation analysis based on the popularity data and the hierarchical tags, and output a video content creation plan.
10. The social media video analysis system according to claim 9, characterized in that, It also includes a compliance and containerized deployment support module, which is used to ensure compliance and scalability of the data output by the multi-source data preprocessing module, the multimodal semantic disambiguation analysis module, the hierarchical tag generation module, the cross-platform popularity ranking module, and the creation analysis module.