Content propagation material structure integrity analysis method, medium and equipment
By introducing a pre-defined tagged corpus and a semantic similarity recall mechanism, missing links in intelligent marketing content are analyzed and identified, solving the problem of insufficient accuracy of large language models in products with scarce training data, and achieving efficient content structure completion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies for intelligent marketing content generation rely on large language models for end-to-end analysis, which makes it difficult to accurately identify missing links in products with scarce training corpora. This results in the generated supplementary content deviating from the user's actual communication goals, reducing the practicality and reliability of automated marketing tools.
By acquiring the user's original multimedia data and auxiliary text information, and using a pre-set tagged corpus and semantic similarity recall mechanism, the multimedia data is parsed into initial content description text segments, and tags are recalled from historical content based on semantic similarity to identify missing content elements.
It improves the robustness and accuracy of identifying missing links in scenarios with high product diversity and scarce training data, ensuring that the generated content meets the user's communication goals and enhancing the accuracy and reliability of marketing tools.
Smart Images

Figure CN121809547A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of content distribution material detection, and in particular to a method, medium, and equipment for analyzing the structural integrity of content distribution materials. Background Technology
[0002] In current intelligent marketing content generation scenarios, users often need to quickly generate a complete and logically clear 60-second or 120-second promotional video based on a short original promotional material (such as a 10-second product video or a single product image). Such promotional videos typically need to include several key content components, such as attention guidance, demand stimulation, solution presentation, value proposition, credibility enhancement, behavioral guidance, and brand recall points. However, the original materials provided by users often only cover some of these elements, with the remaining aspects missing, thus limiting the effectiveness of the communication. Existing technologies generally rely on Large Language Models (LLMs) to directly perform end-to-end analysis of the original materials to identify missing elements. However, this method is highly dependent on the corpus coverage encountered by the model during the pre-training or fine-tuning phase.
[0003] Due to the vast variety of products and the continuous emergence of new categories (such as new smart home devices and niche health foods), large language models struggle to fully learn the typical expression patterns of all products from the training data. Furthermore, the inherent subjectivity of video content interpretation, with each person interpreting it differently, further exacerbates the instability of the model's judgments. When the original material involves products with limited training data, the model is highly susceptible to errors in identifying missing or incorrect links due to its inability to accurately correlate semantics with standard marketing structures. This ultimately leads to the generated supplementary content deviating from the user's true communication goals, reducing the practicality and reliability of automated marketing tools. Summary of the Invention
[0004] To address one of the aforementioned technical problems, the present invention adopts the following technical solution: According to one aspect of the present invention, a method for analyzing the structural integrity of content propagation materials is provided, the method comprising the following steps: Acquire raw multimedia data and auxiliary text information input by the user; the auxiliary text information is used to characterize the content dissemination target category that the user expects the raw multimedia data to achieve; Based on auxiliary text information, determine the set of target content constituent elements corresponding to the content dissemination target category; the set of target content constituent elements includes multiple content constituent element tags necessary for the content dissemination target category; The original multimedia data is parsed to generate multiple initial content description text segments corresponding to the content to be represented by the original multimedia data. For each initial content description text segment, based on text semantic similarity, the top few historical content description text segments with the highest semantic similarity are recalled from a pre-set tagged corpus; each historical content description text segment is associated with a content component tag; Based on the recall results, determine the set of existing content components corresponding to the original multimedia data; The content elements corresponding to different tags in the target content element set and the existing content element set are identified as the missing content elements in the original multimedia data.
[0005] According to a second aspect of the present invention, a non-transitory computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the above-described method for analyzing the structural integrity of content propagation materials.
[0006] According to a third aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described content propagation material structure integrity analysis method.
[0007] This invention has at least one of the following beneficial effects: The method described in this invention effectively overcomes the accuracy problem of relying solely on the capabilities of large language models to identify missing content elements in existing technologies by introducing a "pre-defined labeled corpus" and a "semantic similarity-based recall mechanism." Specifically, this method no longer relies solely on the end-to-end understanding of the original multimedia data by a large model. Instead, it first parses the original content into multiple initial content description text segments, and then recalls the most semantically similar historical texts from a large-scale historical corpus containing manually labeled content elements for each segment. Since the recall is based on semantic similarity rather than generalization of model parameters, even with emerging products where training data is scarce, as long as their expression logic is similar to historical cases, the reliable labels carried by similar historical texts can accurately map the content element type of the current segment. Thus, the system can determine the set of elements covered by the original material in a structured and interpretable manner, and accurately identify missing items by comparing them with the target element set. This mechanism significantly reduces the reliance on the generalization capabilities of large models and improves the robustness and accuracy of missing element identification in scenarios with high product diversity and scarce training data. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 A flowchart of a content propagation material structure integrity analysis method provided in an embodiment of the present invention. Detailed Implementation
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] As one possible embodiment of the present invention, such as Figure 1 As shown, a method for analyzing the structural integrity of content dissemination materials is provided, the method including the following steps: S1: Obtain the raw multimedia data and supplementary text information input by the user. The supplementary text information is used to characterize the content dissemination target category that the user expects the raw multimedia data to achieve. The raw multimedia data includes video and / or images.
[0012] The core of this invention is a content structure integrity analysis method based on semantic standardization, tagged corpus retrieval, and multi-dimensional verification. While it has typical applications in marketing and communication (such as short video ad generation), its underlying logic is highly universal and can be transferred to multiple fields requiring "structured content evaluation and missing content identification." Examples include: quality inspection of educational and training content, review of medical and health education materials, review of government / public service information releases, and generation of product manuals / operation guides. For ease of understanding, this embodiment only describes the use case in the marketing and communication field.
[0013] In practical applications, users are often non-professional marketers (such as small and medium-sized businesses and individual entrepreneurs) who lack a systematic understanding of the elements a complete marketing video should include. Therefore, when implementing S1 of this method, structured input guidance can be provided on the front-end interface: For raw multimedia data, the system supports uploading videos (such as MP4 and MOV formats) or one or more images (such as JPG and PNG formats), and prompts "It is recommended to upload materials that demonstrate the core functions of the product or usage scenarios"; it also provides examples, such as "a 10-second video demonstrating the cleaning process of the detergent" or "a main product image + a user image".
[0014] The auxiliary text information is mainly used to determine the specific marketing objectives (i.e., content dissemination target categories) that users want to achieve through the final generated video. This allows for two input modes to improve accuracy: (1) Free input mode: Users can fill in a brief description, such as "want to highlight quick stain removal, suitable for mothers"; in this mode, certain input prompts need to be set to help users understand what the information of the assistant to be input specifically includes.
[0015] (2) Option-guided mode (preferred): The system pre-sets several auxiliary text information corresponding to the content dissemination target categories for users to select, for example: "Encourage conversion" (corresponding supplementary text such as: "We hope users will place an order immediately after reading this"); "New Product Launch" (corresponding supplementary text such as: "Introducing the core functions of the new product"); "Brand exposure" (corresponding supplementary text such as: "Let more people know about our brand"); "User education" (corresponding supplementary text such as: "Teach users how to use the product correctly").
[0016] Due to users' limited expertise, even with supplementary text or a selected target category, the uploaded original videos or images often fail to cover all the essential elements required for the communication objective. For example, a user might select "promote conversion" and upload a product appearance video, but lack pain point scenarios, authoritative evidence, or a clear call to action, resulting in an incomplete original material structure. This invention addresses this issue by automatically identifying missing elements in subsequent steps to compensate for deficiencies in user input.
[0017] The content dissemination target categories mentioned above are merely examples and can be expanded through configuration in a real system. Each category is associated with a standardized "target content component set" in the backend to ensure that subsequent analysis is based on solid evidence.
[0018] Although the input data acquired in stage S1 may have structural gaps, as long as the auxiliary text information (whether freely input or selected options) can roughly reflect the user's communication intent, the system can accurately map the set of content components that should be included in stage S2, thus providing a benchmark for the missing data analysis in stages S3–S6. This "front-end guidance + back-end intelligent completion" design lowers the user's barrier to entry while ensuring the reliability of the analysis results, demonstrating the robustness and practicality of this invention in real-world applications.
[0019] S2: Based on auxiliary text information, determine the set of target content components corresponding to the content dissemination target category. The target content component set includes multiple content component tags necessary for the content dissemination target category.
[0020] In this invention, the content dissemination target category serves as a bridge connecting user intent with a standardized content structure. Through step S2, user-provided auxiliary text information (such as "suitable for mothers to quickly remove stains" or the user-selected "promote conversion") is transformed into a structured set of target content components, which serves as the benchmark for subsequent missing data analysis.
[0021] S2 includes: S2.1: Based on the auxiliary text information, determine the content dissemination target category corresponding to the auxiliary text information.
[0022] Specifically, S2.1 can be implemented in the following three ways: Firstly, based on the semantic vector corresponding to the auxiliary text information, the content dissemination target category corresponding to the auxiliary text information is determined by matching the semantic vectors of multiple preset content dissemination target categories.
[0023] This method is based on semantic vector similarity matching. It requires the pre-construction of a semantic library of content dissemination target categories, which contains N preset categories (such as "promote conversion", "brand exposure", "new product launch", "user education", etc.). Each category is equipped with a set of representative descriptive sentences (for example, "promote conversion" corresponds to "guide users to buy immediately", "highlight limited-time offers", "emphasize call to action").
[0024] For each set of descriptive sentences, average pooling or weighted fusion is performed to generate a standard semantic vector V for that category. k And store it in the database.
[0025] After the user inputs auxiliary text, its semantic vector V is generated using the same semantic encoding model (such as text2vec-large or BGE). T .
[0026] Calculate V T With all categories V k The cosine similarity is used to select the category corresponding to the one with the highest similarity as the output.
[0027] For example, if a user inputs "I want mothers to place an order after seeing this", its vector has a similarity of 0.87 with the standard vector of the "promote conversion" category, which is higher than other categories (such as "brand exposure" which has a similarity of 0.52). Therefore, it is determined that the target category is "promote conversion".
[0028] Secondly, it directly calls the large language model to generate the content dissemination target category corresponding to the auxiliary text information.
[0029] Embed the auxiliary text T into a preset prompt template, for example: "Please select the content dissemination target category that best matches the description from the following options: [Drive conversion, Brand exposure, New product launch, User education]. Description: '{T}'. Output only the category name." Call the deployed large language model (such as Qwen, Llama3, or API interface) to perform inference and return the category name.
[0030] To further improve stability, confidence filtering (such as requiring the model output probability > 0.7) or multiple sampling voting can be set.
[0031] Thirdly, based on the classification results of the semantic vectors corresponding to the auxiliary text information, the content dissemination target category corresponding to the auxiliary text information is determined.
[0032] Train a lightweight text classifier (such as BERT-base + fully connected layers) with auxiliary text as input and a probability distribution of preset categories as output. The model is fine-tuned on historical user data, with labels provided by operations personnel to ensure alignment with business scenarios. During inference, the category with the highest probability is selected as the result.
[0033] S2.2: Query the preset content component set mapping table to determine the target content component set corresponding to the content dissemination target category.
[0034] Once the target category for content dissemination is determined, the content component set mapping table (hereinafter referred to as the "mapping table") can be queried to obtain the corresponding target content component set. This mapping table contains the target content dissemination category and all tags for the content components within each category. Each "content component tag" is a predefined enumeration type with a unique identifier and semantic definition (e.g., "behavioral guidance element" is defined as containing explicit action instructions, such as "click to buy" or "claim immediately"). This mapping table supports dynamic updates; operators can add or delete categories or adjust element combinations through the management backend.
[0035] This mapping table is stored in the form of key-value pairs in a database or configuration file. A specific structure example is shown in Table 1 below: Table 1 S3: Perform content parsing on the original multimedia data to generate multiple initial content description text segments corresponding to the content to be represented by the original multimedia data.
[0036] Since the original materials provided by users may contain a variety of information units (such as product display, usage scenarios, effect comparison, etc.), and these units are often distributed in different time segments (videos) or different images, they need to be decoupled into several semantically independent text segments in order to analyze the content components they carry segment by segment.
[0037] S3 includes: S3.1: If the original multimedia data is video, the video is divided into multiple sub-video segments according to the video's camera transition points and / or audio pause positions.
[0038] Firstly, existing methods can be used to analyze the multimodal signals contained in the video, and the video can be coarsely segmented: For example, using lens boundary detection algorithms (such as histogram difference-based or deep learning-based Shot Detection models) to identify abrupt changes in the scene.
[0039] Simultaneously, by combining the results of Voice Activity Detection (VAD) and Automatic Speech Recognition (ASR), the end position of sentences (such as punctuation marks or silence intervals > 0.8 seconds) is identified. The video is then divided into several sub-video segments.
[0040] S3.2: For each sub-video segment, generate an initial content description text segment corresponding to the sub-video segment based on its visual content and / or audio content and / or on-screen text.
[0041] For each sub-video segment, a multimodal understanding model (such as Video-LLM or CLIP+Whisper fusion model) can be invoked to generate a descriptive text T by integrating the following information. raw : Visual content (subject, action, background) and / or speech-to-text (if any) and / or on-screen text (recognized via OCR).
[0042] Furthermore, since the initial segmentation may be too fine (e.g., an action is split into two segments) or too coarse (e.g., a segment contains "pain point + solution"), semantic aggregation and resegmentation need to be performed on the results of S3.2: All T raw Concatenate the complete context text T in chronological order. full .
[0043] Call the Large Language Model (LLM) and input prompts. For example: "Please divide the following marketing video description into several paragraphs based on semantic themes. Each paragraph should focus on a core information point (such as product demonstration, pain point revelation, effect proof, etc.)." Original text: {T full}".
[0044] The LLM (Large Language Model) outputs re-divided text segments T1, T2, ..., Tm, each corresponding to a semantically coherent information unit. T1, T2, ..., Tm are used as the final initial content description text segments.
[0045] Therefore, by semantic aggregation and resegmentation, we can ensure that each initial content description text segment expresses only one content component as much as possible, avoiding mixed semantic interference with subsequent tag recall.
[0046] S3.3: If the original multimedia data consists of at least one image, then the large language model is used to parse the content of each image and generate an initial content description text segment corresponding to each image.
[0047] For each image, a multimodal large model (such as Qwen-VL, LLaVA) is used to generate descriptive text. If there is only one image, it is directly used as the initial text segment. If there are multiple images (such as product image + scene image + comparison image), then: all descriptive texts are concatenated into context text. Then, LLM is called to perform semantic resegmentation and output the re-segmented text segments, so that image descriptions with similar semantics are merged (such as merging two "usage scenario" images into one segment), while those with large differences are separated.
[0048] This step employs a two-stage strategy of "physical segmentation + semantic resegmentation" to address the issues of information mixing and blurred boundaries in the original materials. Whether it's video or images, the final output initial content description text segments possess high cohesion and low coupling semantic characteristics, ensuring that each text segment accurately corresponds to a single content component. This design directly supports reliable recall and judgment based on tag statistics in S4–S5, fundamentally improving the granularity and robustness of missing link identification, making it particularly suitable for real-world application scenarios where user-uploaded materials vary in quality. Thus, after completing S3, the system obtains multiple initial content description text segments, each with focused semantics and a clear structure.
[0049] S4: For each initial content description text segment, based on text semantic similarity, recall the top few historical content description text segments with the highest semantic similarity from a pre-defined tagged corpus. Each historical content description text segment is associated with a content component tag.
[0050] The recall process can be achieved through the RAG (Retrieval-Augmented Generation) mechanism: First, a semantic vector is generated for each initial content description text segment using a pre-defined semantic encoding model (such as BGE or Sentence-BERT); then, the vector is input into a pre-built vector database (such as FAISS or Milvus) to perform an approximate nearest neighbor search, recalling the Top-K (e.g., K=10) historical content description text segments with the highest semantic similarity; each recall result is associated with a pre-stored content component tag for subsequent statistical analysis.
[0051] The tagged corpus is constructed as follows: a large number of short text fragments are collected from real marketing videos or advertising scripts, ensuring that each fragment focuses on a single information point; each text is labeled by humans according to predefined content component classification standards (such as "demand-stimulating elements" and "behavioral-guiding elements"); then, the same semantic encoding model as the recall stage is used to generate corresponding vectors, and the (text, vector, tag) triples are stored in the vector database to form a searchable knowledge base.
[0052] S5: Based on the recall results, determine the set of existing content components corresponding to the original multimedia data.
[0053] S5 includes: S5.1: For the set of historical text segments recalled corresponding to the semantic vector of the i-th initial content description text segment, count the occurrence frequency of each content component element label to obtain multiple count values NUM. i,1 NUM i,2, NUM i,j, NUM i,z .
[0054] Where i ranges from 1 to n, where n is the number of initial content description text segments. j ranges from 1 to z, where z is the total number of content component element tags. NUM i,j The number of tags for the j-th type of content component element in the historical text segment set recalled for the i-th initial content description text segment.
[0055] S5.2: If NUM i,j / ∑ i NUM >Y 1 NUM Then determine NUM i,j The corresponding content component element label is the target content component element label corresponding to the i-th initial content description text segment.
[0056] Where, ∑ i NUMY represents the total number of historical content description text segments in the recalled historical text segment set corresponding to the i-th initial content description text segment. 1 NUM This is a preset threshold for the percentage of quantities.
[0057] In this step, an initial content description text segment (or its source sub-video segment) is allowed to be matched by multiple content component element tags (when the proportion of multiple tags is all > Y). 1 NUM (At the time); when constructing the "set of existing content components", the union of sets is used instead of mutually exclusive allocation.
[0058] Additionally, if the original multimedia data only includes video, the method further includes the following after S5: S5.3: From the original multimedia data, obtain the total video duration and the total number of sub-video segments corresponding to each content element tag in the existing content element collection.
[0059] In the S3 stage, semantic aggregation and resegmentation of the initially generated sub-video segment descriptions through LLM may lead to a complex "one-to-many" or "many-to-one" mapping relationship between the initial content description text segment and the original video segment. Therefore, this invention needs to explicitly record and maintain the original sub-video segment index and its timestamp information that each initial text segment depends on during the resegmentation process to ensure the integrity of the traceability chain.
[0060] Specifically, when an initial content description text segment is generated by fusing multiple original sub-video segments, all associated original segment IDs and their corresponding durations are saved together. Conversely, if multiple initial content description text segments originate from the same original sub-video segment (e.g., due to semantic splitting), that segment can be referenced multiple times. In the S5.3 statistical phase, for each content component element tag, all its associated initial content description text segments are traversed, and the set of associated original sub-video segments (allowing cross-element sharing) is aggregated. This allows for the accurate summation of the total number of sub-video segments corresponding to that element (duplicate count) and the total video duration (sum of the durations of each segment). This mechanism effectively solves the problem of discontinuity in temporal information caused by semantic recombination, ensuring the data reliability of subsequent sufficiency verification (S5.4).
[0061] Alternatively, the original video can be re-segmented based on the final initial content description text segment obtained after semantic aggregation and re-segmentation, and the correspondence between the initial content description text segment and the re-segmented sub-video segments can be saved to obtain the total video duration and the total number of sub-video segments corresponding to each content component element tag.
[0062] S5.4: If the total video duration (Time) corresponds to the nth content element tag in the existing content element set...n and the total number of sub-video segments ∑ n NUM ∑ n NUM <Y 2 NUM And Time n <Y Time Then, the content elements in the target content element set that have the same label as the nth type of content element are identified as the missing content elements in the original multimedia data. 2 NUM Y is the preset threshold for the number of video segments; Time This is a preset video segment duration threshold.
[0063] The main objective of S5.1–S5.2 is to, based on the frequency of occurrence of each content component tag in the recall results, filter out the content components most likely to be expressed by each initial content description text segment by setting a percentage threshold, thereby initially constructing a "set of existing content components". This scheme ensures that tag determination is based on statistical evidence and avoids misjudgments caused by individual noise in the recall.
[0064] The main objective of S5.3–S5.4 is to further verify the sufficiency of each element in the initially determined set of existing content components when the user-provided original multimedia data consists only of video. Elements are only removed if the number and cumulative duration of the corresponding sub-video segments are both below a preset threshold. This scheme can effectively eliminate weak coverage situations such as "fleeting moments" or "single mentions."
[0065] The synergy of the above schemes can not only determine whether an element is "mentioned", but also assess whether it is "fully presented", thereby making the final set of existing content components more accurate, reliable and more in line with actual communication needs.
[0066] S6: Identify the missing content elements in the original multimedia data as the content elements corresponding to different tags in the target content element set and the existing content element set.
[0067] After S6, relevant prompts can be generated based on missing content elements and input into the LLM (Local Management Library). The LLM then generates the corresponding spoken text for each missing marketing step, and generates supplementary videos and / or images for each spoken text, thus completing the full structure of the content delivery materials.
[0068] As another possible embodiment of the invention, prior to the recall, a standardization process is also included for the initial content description text segment: S4.1: Replace entity information in the initial content description text segment with preset generalized placeholders to generate a standardized text segment. Entity information includes at least one of product name, user role, functional characteristics, or usage scenario. A generalized placeholder is a generic identifier generated after semantically abstracting entity information; it does not carry specific entity attributes but retains the semantic role of the entity.
[0069] S4.2: Recall is based on standardized text segment execution.
[0070] This embodiment can standardize the initial content description text segment before executing the recall to improve the generalization ability and accuracy of semantic matching.
[0071] Specifically, identifying specific entity information in the initial content description text segment can include product name (such as "XX brand fruit and vegetable washing machine"), user role (such as "stay-at-home mom" or "office worker"), functional features (such as "99% pesticide residue removal rate") or usage scenario (such as "kitchen" or "office"). This information carries the specific attributes of the entity. In subsequent semantic matching, the accuracy of semantic matching may decrease due to the difference in specific attributes.
[0072] In this embodiment, these are replaced with preset generalized placeholders, such as replacing "stay-at-home mom" with "user A", "fruit and vegetable washing machine" with "product B", and "kitchen" with "scene C". These placeholders do not carry specific brand, model, or regional attributes, but retain the original entity's role in the semantic structure (such as "user", "operated object", "environment"). This setting, on the one hand, eliminates semantic bias caused by product specificity, allowing niche, emerging, or unseen products to be matched with historical cases of general expression patterns; on the other hand, it improves the reusability efficiency of tagged corpora, avoiding the inability to recall texts with similar logic due to differences in entity vocabulary.
[0073] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0074] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0075] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.
[0076] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: entirely in hardware, entirely in software (including firmware, microcode, etc.), or in a combination of hardware and software, collectively referred to herein as “circuit,” “module,” or “system.”
[0077] An electronic device according to this embodiment of the invention. The electronic device is merely an example and should not be construed as limiting the functionality or scope of the embodiments of the invention.
[0078] Electronic devices are manifested in the form of general-purpose computing devices. Components of an electronic device may include, but are not limited to: at least one processor, at least one memory, and buses connecting different system components (including memory and processor).
[0079] The memory stores program code that can be executed by a processor, causing the processor to perform the steps described in the "Exemplary Methods" section above, according to various exemplary embodiments of the present invention.
[0080] The storage may include readable media in the form of volatile storage, such as random access memory (RAM) and / or cache memory, and may further include read-only memory (ROM).
[0081] The storage may also include programs / utilities having a set (at least one) of program modules, including but not limited to: an operating system, one or more applications, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0082] A bus can represent one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus that uses any of the various bus architectures.
[0083] The electronic device can also communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., routers, modems, etc.). This communication can be performed via input / output (I / O) interfaces. Furthermore, the electronic device can communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter. The network adapter communicates with other modules of the electronic device via a bus. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0084] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of the present invention may also be implemented as a program product comprising program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of the present invention described in the "Exemplary Methods" section above.
[0085] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0086] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0087] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0088] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0089] Furthermore, the accompanying drawings are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes shown in the above drawings do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0090] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0091] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for analyzing the structural integrity of content dissemination materials, characterized in that, The method includes the following steps: Obtain raw multimedia data and auxiliary text information input by the user; the auxiliary text information is used to characterize the content dissemination target category that the user expects the raw multimedia data to achieve; Based on the auxiliary text information, a set of target content constituent elements corresponding to the content dissemination target category is determined; the set of target content constituent elements includes multiple content constituent element tags necessary for the content dissemination target category; The original multimedia data is parsed to generate multiple initial content description text segments corresponding to the content to be represented by the original multimedia data; For each initial content description text segment, based on text semantic similarity, the top few historical content description text segments with the highest semantic similarity are recalled from a preset tagged corpus; each historical content description text segment is associated with a content component tag; Based on the recall results, determine the set of existing content components corresponding to the original multimedia data; The content elements in the target content element set that correspond to different tags in the existing content element set are identified as the missing content elements in the original multimedia data.
2. The method according to claim 1, characterized in that, Based on the auxiliary text information, a set of target content constituent elements corresponding to the content dissemination target category is determined, including: Based on the auxiliary text information, determine the content dissemination target category corresponding to the auxiliary text information; Query the preset content component set mapping table to determine the target content component set corresponding to the content dissemination target category.
3. The method according to claim 2, characterized in that, Based on the auxiliary text information, the content dissemination target category corresponding to the auxiliary text information is determined, including: Based on the semantic vector corresponding to the auxiliary text information, and the matching results with the semantic vectors of multiple preset content dissemination target categories, the content dissemination target category corresponding to the auxiliary text information is determined; or Directly call the large language model to generate the content propagation target category corresponding to the auxiliary text information; or Based on the classification results of the semantic vectors corresponding to the auxiliary text information, the content dissemination target category corresponding to the auxiliary text information is determined.
4. The method according to claim 1, characterized in that, The original multimedia data includes video and / or images; The original multimedia data is parsed to generate multiple initial content description text segments corresponding to the content to be represented by the original multimedia data, including: If the original multimedia data is video, then the video is divided into multiple sub-video segments according to the camera transition points and / or voice pause positions of the video; For each sub-video segment, an initial content description text segment corresponding to the sub-video segment is generated based on its visual content and / or audio content and / or on-screen text.
5. The method according to claim 1, characterized in that, The original multimedia data includes video and / or images; The original multimedia data is parsed to generate multiple initial content description text segments corresponding to the content to be represented by the original multimedia data, including: If the original multimedia data is at least one image, then a large language model is used to parse the content of each image to generate an initial content description text segment corresponding to each image.
6. The method according to claim 4 or 5, characterized in that, Prior to the recall, the process also includes standardization of the initial content description text segment: The entity information in the initial content description text segment is replaced with a preset generalized placeholder to generate a standardized text segment; the entity information includes at least one of product name, user role, functional feature or usage scenario; the generalized placeholder is a general identifier generated after semantic abstraction of the entity information, which does not carry specific entity attributes, but retains the semantic role of the entity. The recall is performed based on the standardized text segment.
7. The method according to claim 1, characterized in that, Based on the recall results, the set of existing content components corresponding to the original multimedia data is determined, including: For the set of historical text segments recalled by the semantic vector corresponding to the i-th initial content description text segment, count the occurrence frequency of each content component element tag to obtain multiple count values NUM. i,1 NUM i,2 , …,NUM i,j , ..., NUM i,z ; Where i ranges from 1 to n, where n is the number of initial content description text segments; j ranges from 1 to z, where z is the total number of content component element tags; NUM i,j The number of tags for the j-th type of content component element in the set of historical text segments recalled for the i-th initial content description text segment; If NUM i,j / ∑ i NUM >Y 1 NUM Then determine NUM i,j The corresponding content component element label is the target content component element label corresponding to the i-th initial content description text segment; Where, ∑ i NUM Y represents the total number of historical content description text segments in the recalled historical text segment set corresponding to the i-th initial content description text segment. 1 NUM This is a preset threshold for the percentage of quantities.
8. The method according to claim 4, characterized in that, The original multimedia data includes only video; After identifying the content elements in the target content element set that correspond to different tags in the existing content element set as the missing content elements of the original multimedia data, the method further includes: From the original multimedia data, obtain the total video duration and the total number of sub-video segments corresponding to each content component tag in the existing content component set; If the total video duration (Time) corresponds to the nth content element tag in the existing content element set... n and the total number of sub-video segments ∑ n NUM ∑ n NUM <Y 2 NUM And Time n <Y Time Then, the content elements in the target content element set that have the same tag as the nth type of content element are identified as the missing content elements in the original multimedia data. Y 2 NUM Y is the preset threshold for the number of video segments; Time This is a preset video segment duration threshold.
9. A non-transitory computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a content propagation material structure integrity analysis method as described in any one of claims 1 to 8.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a content propagation material structure integrity analysis method as described in any one of claims 1 to 8.