Content batch labeling method, apparatus, equipment and media based on large models
By using a content batch tagging method based on a large model, the problem of low efficiency in traditional manual tagging methods is solved. This method enables automated, standardized, and scenario-based batch tagging of multi-format content, thereby improving tagging efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 特赞(上海)信息科技有限公司
- Filing Date
- 2026-03-18
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional manual content tagging methods are time-consuming, labor-intensive, and costly, making it difficult to meet the efficiency needs of large-scale content tagging.
A content batch labeling method based on a large model is adopted. This method involves preprocessing multi-format content files, including paragraph splitting, keyword extraction, entity recognition, and topic information extraction, and combining them with a large language model for semantic feature matching to create a multi-dimensional tag library, thereby achieving automated, standardized, and scenario-based batch labeling.
It improves the efficiency of content tagging, breaks down format barriers, supports batch standardized processing of multiple content formats such as text, images, and videos, enhances the accuracy and flexibility of tagging, and reduces the cost of manual intervention.
Smart Images

Figure CN122132432A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to the field of artificial intelligence technology, and more specifically, to a method, apparatus, device, and medium suitable for batch content tagging based on a large model. Background Technology
[0002] With the continuous emergence of massive amounts of content such as text, images, audio, and video, quickly and accurately classifying and labeling this content has become a key issue in the field of information processing.
[0003] In related technologies, content tagging is mainly achieved through manual tagging. However, traditional manual tagging methods are not only time-consuming and labor-intensive but also costly, and inefficient when dealing with large-scale content tagging needs. Summary of the Invention
[0004] The embodiments described herein provide a method, apparatus, device, and medium for batch content tagging based on a large model, overcoming the aforementioned problems.
[0005] Firstly, based on the content of this disclosure, a method for batch content tagging based on a large model is provided, including: Obtain batch import of multi-format content files, wherein the multi-format content files include: content to be tagged in different storage formats and business scenario information corresponding to the content to be tagged; For text-based content to be tagged, paragraph splitting, keyword extraction, and entity recognition are performed to obtain the first preprocessed content. For image-based content to be tagged, the theme information, object information, and text information in the content are identified to obtain the second preprocessed content. For video-based content to be tagged, keyframe extraction, image content recognition, and speech conversion are performed to obtain the third preprocessed content. The first preprocessed content, the second preprocessed content, and the third preprocessed content are subjected to structured transformation to obtain the corresponding target structured content; Create a multi-dimensional tag library; the multi-dimensional tag library includes multiple sub-tag libraries, and different sub-tag libraries correspond to different business scenario rules; Using a large language model, based on the business scenario information corresponding to the content to be tagged, the target structured content is semantically matched with the multi-dimensional tag library to obtain the target content tag corresponding to each content to be tagged; and the content to be tagged is associated with the corresponding target content tag.
[0006] Secondly, according to the present disclosure, a content batch tagging device based on a large model is provided, comprising: The acquisition module is used to acquire multi-format content files imported in batches. The multi-format content files include: content to be tagged in different storage formats and business scenario information corresponding to the content to be tagged. The processing module is used to perform paragraph splitting, keyword extraction, and entity recognition on the text-based content to be tagged, to obtain the first preprocessed content corresponding to the content to be tagged; to identify the theme information, object information, and text information in the image-based content to be tagged, to obtain the second preprocessed content corresponding to the content to be tagged; and to extract keyframes, recognize the image content, and convert the video-based content to speech, to obtain the third preprocessed content corresponding to the content to be tagged. The conversion module is used to perform structured conversion on the first preprocessed content, the second preprocessed content, and the third preprocessed content to obtain the corresponding target structured content. A creation module is used to create a multi-dimensional tag library; the multi-dimensional tag library includes multiple sub-tag libraries, and different sub-tag libraries correspond to different business scenario rules; The matching module is used to perform semantic feature matching between the target structured content and the multi-dimensional tag library based on the business scenario information corresponding to the content to be tagged using a large language model, so as to obtain the target content tag corresponding to each content to be tagged; and associate the content to be tagged with the corresponding target content tag.
[0007] Thirdly, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the content batch labeling method based on a large model as described in any of the above embodiments.
[0008] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the content batch labeling method based on a large model as described in any of the above embodiments.
[0009] The content batch labeling method based on a large model provided in this application embodiment obtains batch-imported multi-format content files, which include: content to be labeled in different storage formats and business scenario information corresponding to the content to be labeled; for content to be labeled in text format, paragraph splitting, keyword extraction, and entity recognition are performed on the content to be labeled to obtain the first preprocessed content corresponding to the content to be labeled; for content to be labeled in image format, the theme information, object information, and text information in the content to be labeled are identified to obtain the second preprocessed content corresponding to the content to be labeled; for content to be labeled in video format, the content to be labeled is... The process involves keyframe extraction, image content recognition, and speech conversion to obtain the third preprocessed content corresponding to the content to be tagged. The first, second, and third preprocessed content are then structurally transformed to obtain the corresponding target structured content. A multi-dimensional tag library is created, comprising multiple sub-tag libraries, each corresponding to different business scenario rules. Using a large language model, based on the business scenario information corresponding to the content to be tagged, semantic feature matching is performed between the target structured content and the multi-dimensional tag library to obtain the target content tag for each piece of content to be tagged. The content to be tagged is then associated with its corresponding target content tag. Thus, by preprocessing and structuring the content to be tagged, and matching the processed content with a pre-created multi-dimensional tag library, automated, standardized, and scenario-based batch tagged content is achieved, effectively improving content tagged efficiency.
[0010] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more obvious and understandable, specific implementation methods of this application are described below. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. It should be understood that the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure, wherein: Figure 1 This is a flowchart illustrating a method for batch content tagging based on a large model, which is disclosed in this publication.
[0012] Figure 2 This is a schematic diagram of a content batch labeling device based on a large model, which is disclosed in this publication.
[0013] Figure 3 This is a schematic diagram of the structure of a computer device provided in this disclosure.
[0014] It should be noted that the elements in the attached diagram are schematic and not drawn to scale. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are also within the scope of protection of this disclosure.
[0016] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the specification and in the relevant art, and shall not be interpreted in an idealized or overly formal form unless otherwise explicitly defined herein. As used herein, the statement of “connecting” or “coupling” two or more parts together shall mean that these parts are directly joined together or joined through one or more intermediate components.
[0017] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of the phrase "embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0018] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists, A and B exist simultaneously, or B exists. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Terms such as "first" and "second" are only used to distinguish one component (or part of a component) from another component (or another part of a component).
[0019] In the description of this application, unless otherwise stated, "multiple" means two or more (including two), and similarly, "multiple groups" means two or more (including two groups).
[0020] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0021] Figure 1 This is a flowchart illustrating a method for batch content tagging based on a large model, as provided in this embodiment of the disclosure. Figure 1 As shown, the specific process of the content batch tagging method based on large models includes: S110. Obtain multi-format content files for batch import. The multi-format content files include: content to be tagged in different storage formats and business scenario information corresponding to the content to be tagged.
[0022] The import of content files in various formats supports folders, compressed files, etc. Storage formats include, but are not limited to, text (docx / pdf / txt), images (jpg / png), and videos (mp4). The business scenario information corresponding to the content to be tagged can be such as "marketing scenario" or "compliance scenario."
[0023] In some embodiments, before acquiring the batch-imported multi-format content files, the method further includes: detecting whether the current time has reached a preset processing time; or, detecting whether the upload task of the multi-format content files has been triggered.
[0024] For example, users can set up scheduled tasks, such as processing newly added content every day at midnight, and then retrieving multi-format content files for batch import at midnight. Alternatively, the task can automatically initiate tagging operations and perform batch import of multi-format content files after content upload is detected. This allows users to schedule content tagging or trigger the tagging process immediately after content upload, effectively improving the flexibility of content tagging.
[0025] S120. For text-based content to be tagged, perform paragraph splitting, keyword extraction, and entity recognition to obtain the first preprocessed content corresponding to the content to be tagged; for image-based content to be tagged, identify the theme information, object information, and text information in the content to be tagged to obtain the second preprocessed content corresponding to the content to be tagged; for video-based content to be tagged, perform keyframe extraction, image content recognition, and speech conversion to obtain the third preprocessed content corresponding to the content to be tagged.
[0026] By preprocessing content to be tagged in different storage formats, different types of content can be transformed into analyzable intermediate data.
[0027] In some embodiments, the content to be tagged is segmented into paragraphs, keywords are extracted, and entities are identified to obtain the first preprocessed content corresponding to the content to be tagged. This includes: splitting the content to be tagged into multiple independent content paragraphs according to a preset paragraph separator; extracting corresponding core keywords from each independent content paragraph using a preset keyword extraction model; identifying and extracting specific entity information contained in each independent content paragraph using a preset entity recognition model; and determining the first preprocessed content corresponding to the content to be tagged based on the core keywords corresponding to each independent content paragraph and the specific entity information contained in each independent content paragraph.
[0028] The preset paragraph separators can include line breaks, periods, and semicolons. The preset keyword extraction model uses a comprehensive calculation of word frequency, importance weight, and contextual semantic relationships to accurately extract keywords that reflect the core meaning of a paragraph; for example, it extracts key information such as the event subject, time, and location as core keywords from independent content paragraphs in news text. The preset entity recognition model uses named entity recognition technology to identify specific entity information such as product names and brand names contained in each independent content paragraph. By associating and integrating the core keywords of each independent content paragraph with specific entity information, the first preprocessed content is formed. This decomposes the originally complex text content into key elements with clear semantic meaning, improving the processability of the text content.
[0029] In some embodiments, identifying the theme information, object information, and text information in the content to be tagged to obtain the second preprocessed content corresponding to the content to be tagged includes: performing theme recognition on the corresponding image of the content to be tagged to obtain theme information corresponding to the content to be tagged; performing object detection on the corresponding image of the content to be tagged to obtain object information corresponding to the content to be tagged; performing text extraction on the corresponding image of the content to be tagged to obtain text information corresponding to the content to be tagged; and determining the second preprocessed content corresponding to the content to be tagged based on the theme information, object information, and text information corresponding to the content to be tagged.
[0030] This process involves several steps. First, a deep learning-based image classification model can be used to identify the overall visual features of the image corresponding to the content to be tagged, thereby determining the theme information, such as "product display" or "personal activity." Second, object detection is performed on the image corresponding to the content to be tagged to locate and identify the bounding boxes and category information of various objects, such as "car" or "sign." Third, optical character recognition (OCR) technology can be used to extract text from printed or handwritten characters in the image corresponding to the content to be tagged. Finally, by integrating object and text information, a second preprocessing step is formed. This transforms the originally abstract image content into clear data containing theme direction, specific objects, and text information, improving the processability of the image content.
[0031] In some embodiments, keyframe capture, image content recognition, and speech conversion are performed on the content to be marked to obtain the third preprocessed content corresponding to the content to be marked. This includes: extracting keyframe images from the video stream corresponding to the content to be marked at preset time intervals; performing image content recognition on the keyframe images to obtain the visual content contained in the keyframe images; performing speech recognition conversion on the speech signal in the video stream corresponding to the content to be marked to convert the speech information into corresponding text information; and determining the third preprocessed content corresponding to the content to be marked based on the visual content contained in the keyframe images and the text information obtained after speech conversion.
[0032] The preset time interval can be adaptively set according to the dynamic changes of the video content. For video content with rapid changes, the preset time interval can be set to 0.5s-2s to ensure that more key actions and scene changes are captured; for video content with slow changes, the preset time interval can be set to 5s-10s to reduce data processing while ensuring that key information is not lost. The visual content contained in the keyframe images may include clothing features, such as color and style, or the type of object, such as vehicles, tables, and chairs.
[0033] The visual content contained in the keyframe image and the text information obtained after speech conversion can be fused in a multimodal manner, which makes it easier to construct a third preprocessed content that has both visual details of the image and semantic supplementation of speech.
[0034] S130. Perform structured transformation on the first preprocessed content, the second preprocessed content, and the third preprocessed content to obtain the corresponding target structured content.
[0035] In the process of structuring, the first preprocessed content can be converted into structured data containing fields such as paragraph numbers, sentence sequences, and keyword lists; the second preprocessed content can be converted into data stored in structured forms such as feature vector arrays, image dimensions, and key region coordinates; and the third preprocessed content can be converted into structured data containing multi-dimensional information such as timestamps, visual content tag sequences, and speech-text sentence segments.
[0036] S140. Create a multi-dimensional tag library.
[0037] The multi-dimensional tag library includes multiple sub-tag libraries, each corresponding to different business scenario rules. Sub-tag libraries might include a "Marketing Scenario" library containing "product type" and "target audience," or a "Compliance Scenario" library containing "sensitive words" and "copyright identifiers." Business scenario rules might include keyword matching thresholds and matching priorities.
[0038] In addition, you can also set hierarchical associations for sub-tags (such as "product type → electronic products → mobile phones") and weight settings (such as core tags having higher weight than secondary tags).
[0039] S150. Using a large language model, based on the business scenario information corresponding to the content to be tagged, the target structured content is semantically matched with a multi-dimensional tag library to obtain the target content tag corresponding to each content to be tagged; and the content to be tagged is associated with the corresponding target content tag.
[0040] In some embodiments, a large language model is used to perform semantic feature matching between the target structured content and a multi-dimensional tag library based on the business scenario information corresponding to the content to be tagged, to obtain the target content tag corresponding to each content to be tagged. This includes: using a large language model, matching sub-tag libraries corresponding to each sub-content in the target structured content from the multi-dimensional tag library based on the business scenario information corresponding to the content to be tagged; and performing semantic feature matching between each sub-tag library and its corresponding sub-content based on the business scenario rules corresponding to each matched sub-tag library, to obtain the target content tag corresponding to each content to be tagged.
[0041] This involves calling preset business scenario rules from the corresponding sub-tag library, combining the context of the sub-content with the frequency and position of keywords to match content with the tag library. Combined with weight settings, this ensures that core tags are matched preferentially and accurately.
[0042] This allows for the precise identification of sub-tag libraries that closely align with business scenarios, ensuring the relevance and effectiveness of tag matching, avoiding interference from irrelevant tags, and ultimately improving the accuracy and relevance of tag association.
[0043] In addition, during the content tagging process, it also supports task progress visualization (processing / completed / failed) and batch export of tagging results. The tagging results can include content ID, tag list, and matching confidence.
[0044] In this embodiment, a batch of imported multi-format content files are acquired. These files include content to be tagged in different storage formats and corresponding business scenario information. For text-based content, paragraph splitting, keyword extraction, and entity recognition are performed to obtain the first preprocessed content. For image-based content, thematic information, object information, and text information are identified to obtain the second preprocessed content. For video-based content, keyframes are extracted from the content. The process involves capturing, recognizing, and converting the content of the image to speech, resulting in the third preprocessed content corresponding to the content to be tagged. The first, second, and third preprocessed content are then structurally transformed to obtain the corresponding target structured content. A multi-dimensional tag library is created, comprising multiple sub-tag libraries, each corresponding to different business scenario rules. Using a large language model, based on the business scenario information corresponding to the content to be tagged, the target structured content is semantically matched with the multi-dimensional tag library to obtain the target content tag for each piece of content to be tagged. The content to be tagged is then associated with the corresponding target content tag. Thus, by preprocessing and structuring the content to be tagged, and matching the processed content with the pre-created multi-dimensional tag library, automated, standardized, and scenario-based batch tagged content is achieved, effectively improving content tagged efficiency.
[0045] In some embodiments, before associating the content to be tagged with the corresponding target content tag, the method further includes: obtaining the historical tagged content of the target content tag corresponding to the content to be tagged; and performing correlation detection between the content to be tagged and the corresponding target content tag based on the correlation index between the content to be tagged and the historical tagged content corresponding to the target content tag.
[0046] The correlation index can be determined by calculating the similarity between the target structured content of the content to be tagged and the structured data of the historical tagged content. The higher the similarity between the two, the larger the correlation index, and the lower the similarity between the two, the smaller the correlation index.
[0047] When determining the correlation detection result between the content to be tagged and the corresponding target content tag, a threshold can be set for comparison. If the correlation index between the content to be tagged and the historical tagged content is greater than or equal to the set threshold, the correlation detection result between the content to be tagged and the corresponding target content tag is determined to be correlated. If the correlation index between the content to be tagged and the historical tagged content is less than the set threshold, the correlation detection result between the content to be tagged and the corresponding target content tag is determined to be uncorrelated.
[0048] Therefore, by adding a correlation detection before associating the content to be tagged with the corresponding target content tag, it is possible to effectively filter out the content to be tagged that has a low correlation with the target content tag, and avoid incorrectly labeling irrelevant or weakly related content with the target tag, thereby significantly improving the accuracy and reliability of content tagging.
[0049] In addition, this embodiment also supports users to add, delete, and modify tags on automatically tagged content, and can apply or delete tags in batches. After tagging is completed, the system can automatically collect calibration data to optimize the keyword extraction accuracy and semantic model of tag matching in content preprocessing, thereby improving the accuracy of subsequent tagging.
[0050] In summary, the content batch tagging method based on a large model provided in this embodiment can support batch uploading and unified tagging of multiple content formats such as text, images, and videos, breaking down format barriers and achieving batch standardized processing of text / images / videos; it can maintain a custom tag library, and can create and modify tag trees with one click through prompt words; it supports preset tagging rules according to scenarios, quickly responding to the tagging needs of different business scenarios; it achieves automatic tagging through content preprocessing (splitting, keyword extraction) + tag library matching, while supporting automatic application and manual calibration and application, improving tagging efficiency, and manual calibration continuously optimizes model accuracy; it supports scheduled / triggered batch tagging tasks to reduce the cost of manual intervention, adapts to large-scale content management scenarios, and provides task progress monitoring and result export capabilities.
[0051] Figure 2 This embodiment provides a schematic diagram of a content batch labeling device based on a large model. The content batch labeling device based on a large model may include: The acquisition module 210 is used to acquire multi-format content files imported in batches. The multi-format content files include: content to be tagged in different storage formats and business scenario information corresponding to the content to be tagged.
[0052] The processing module 220 is used to perform paragraph splitting, keyword extraction, and entity recognition on the text-based content to be tagged, to obtain the first preprocessed content corresponding to the content to be tagged; to identify the theme information, object information, and text information in the image-based content to be tagged, to obtain the second preprocessed content corresponding to the content to be tagged; and to extract keyframes, recognize the image content, and convert speech to the video-based content to obtain the third preprocessed content corresponding to the content to be tagged.
[0053] The conversion module 230 is used to perform structured conversion on the first preprocessed content, the second preprocessed content and the third preprocessed content to obtain the corresponding target structured content.
[0054] Create module 240 to create a multi-dimensional tag library; the multi-dimensional tag library includes multiple sub-tag libraries, and different sub-tag libraries correspond to different business scenario rules.
[0055] The matching module 250 is used to perform semantic feature matching between the target structured content and the multi-dimensional tag library based on the business scenario information corresponding to the content to be tagged using a large language model, so as to obtain the target content tag corresponding to each content to be tagged; and associate the content to be tagged with the corresponding target content tag.
[0056] In this embodiment, optionally, the processing module 220 is specifically used for: The content to be tagged is divided into multiple independent content paragraphs according to the preset paragraph separators; the corresponding core keywords are extracted from each independent content paragraph using the preset keyword extraction model; and the specific entity information contained in each independent content paragraph is identified and extracted using the preset entity recognition model; based on the core keywords corresponding to each independent content paragraph and the specific entity information contained in each independent content paragraph, the first preprocessing content corresponding to the content to be tagged is determined.
[0057] In this embodiment, optionally, the processing module 220 is specifically used for: The topic information corresponding to the content to be labeled is obtained by performing topic recognition on the corresponding image of the content to be labeled; the object information corresponding to the content to be labeled is obtained by performing object detection on the corresponding image of the content to be labeled; the text information corresponding to the content to be labeled is obtained by performing text extraction on the corresponding image of the content to be labeled; and the second preprocessing content corresponding to the content to be labeled is determined based on the topic information, object information, and text information of the content to be labeled.
[0058] In this embodiment, optionally, the processing module 220 is specifically used for: According to a preset time interval, key frame images are extracted from the video stream corresponding to the content to be labeled; the key frame images are subjected to image content recognition to obtain the visual content contained in the key frame images; and the speech signal in the video stream corresponding to the content to be labeled is subjected to speech recognition conversion to convert the speech information into the corresponding text information; based on the visual content contained in the key frame images and the text information obtained after speech conversion, the third preprocessing content corresponding to the content to be labeled is determined.
[0059] In this embodiment, optionally, the matching module 250 is specifically used for: Using a large language model, based on the business scenario information corresponding to the content to be tagged, sub-tag libraries corresponding to each sub-content in the target structured content are matched from a multi-dimensional tag library; based on the business scenario rules corresponding to each matched sub-tag library, semantic feature matching is performed between each sub-tag library and the corresponding sub-content to obtain the target content tag corresponding to each content to be tagged.
[0060] In this embodiment, optionally, a detection module may also be included.
[0061] The detection module is used to obtain the historical tagged content of the target content tag corresponding to the content to be tagged; and to perform correlation detection between the content to be tagged and the corresponding target content tag based on the correlation index between the content to be tagged and the historical tagged content corresponding to the target content tag.
[0062] In this embodiment, optionally, the detection module is also used to detect whether the current time has reached the preset processing time; or, to detect whether the upload task of multi-format content files has been triggered.
[0063] The content batch labeling device based on large models provided in this disclosure can execute the above-described method embodiments. For its specific implementation principle and technical effects, please refer to the above-described method embodiments, which will not be repeated here.
[0064] This application also provides a computer device. Please refer to the following for details. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.
[0065] The computer device includes a memory 310 and a processor 320 that are interconnected via a system bus. It should be noted that only a computer device with memory 310 and processor 320 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0066] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0067] The memory 310 includes at least one type of readable storage medium, including non-volatile memory or volatile memory, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. RAM may include static RAM or dynamic RAM. In some embodiments, the memory 310 may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory 310 may also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, or flash card equipped on the computer device. Of course, the memory 310 may include both internal storage units and external storage devices of the computer device. In this embodiment, the memory 310 is typically used to store the operating system and various application software installed on the computer device, such as the program code of the method described above. In addition, the memory 310 can also be used to temporarily store various types of data that have been output or will be output.
[0068] Processor 320 is typically used to perform overall operations of a computer device. In this embodiment, memory 310 is used to store program code or instructions, including computer operation instructions, and processor 320 is used to execute the program code or instructions stored in memory 310 or process data, such as program code that runs the methods described above.
[0069] In this article, the bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus system can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0070] Another embodiment of this application also provides a computer-readable medium, which may be a computer-readable signal medium or a computer-readable medium. A processor in a computer reads computer-readable program code stored in the computer-readable medium, enabling the processor to execute the functional actions specified in each step or combination of steps in the above method; and to generate means for implementing the functional actions specified in each block or combination of blocks in the block diagram.
[0071] Computer-readable media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared memory or semiconductor systems, devices or apparatuses, or any suitable combination thereof, wherein the memory is used to store program code or instructions, the program code including computer operation instructions, and the processor is used to execute the program code or instructions of the above-described methods stored in the memory.
[0072] The definitions of memory and processor can be found in the description of the foregoing computer device embodiments, and will not be repeated here.
[0073] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0074] In the various embodiments of this application, the functional units or modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0075] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0076] In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" as described in this application does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims listing several means, several units of these means may be embodied by the same item of hardware. The use of "first," "second," and "third," etc., does not indicate any order and these words should be interpreted as names. Unless otherwise specified, the steps in the above embodiments should not be construed as limiting the order of execution.
[0077] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for batch content tagging based on a large model, characterized in that, include: Obtain batch import of multi-format content files, wherein the multi-format content files include: content to be tagged in different storage formats and business scenario information corresponding to the content to be tagged; For text-based content to be tagged, paragraph splitting, keyword extraction, and entity recognition are performed to obtain the first preprocessed content. For image-based content to be tagged, the theme information, object information, and text information in the content are identified to obtain the second preprocessed content. For video-based content to be tagged, keyframe extraction, image content recognition, and speech conversion are performed to obtain the third preprocessed content. The first preprocessed content, the second preprocessed content, and the third preprocessed content are subjected to structured transformation to obtain the corresponding target structured content; Create a multi-dimensional tag library; the multi-dimensional tag library includes multiple sub-tag libraries, and different sub-tag libraries correspond to different business scenario rules; Using a large language model, based on the business scenario information corresponding to the content to be tagged, the target structured content is semantically matched with the multi-dimensional tag library to obtain the target content tag corresponding to each content to be tagged; and the content to be tagged is associated with the corresponding target content tag.
2. The method according to claim 1, characterized in that, The content to be labeled is segmented into paragraphs, keywords are extracted, and entities are recognized to obtain the first preprocessed content corresponding to the content to be labeled, including: The content to be tagged is split into multiple independent content paragraphs according to the preset paragraph separators; Using a preset keyword extraction model, the corresponding core keywords are extracted from each of the independent content paragraphs; and using a preset entity recognition model, specific entity information contained in each of the independent content paragraphs is identified and extracted. Based on the core keywords corresponding to each independent content paragraph and the specific entity information contained in each independent content paragraph, the first preprocessing content corresponding to the content to be tagged is determined.
3. The method according to claim 1, characterized in that, Identify the theme information, object information, and text information in the content to be tagged to obtain the second preprocessed content corresponding to the content to be tagged, including: The topic information corresponding to the content to be tagged is obtained by performing topic recognition on the corresponding image of the content to be tagged; the object information corresponding to the content to be tagged is obtained by performing target detection on the corresponding image of the content to be tagged; and the text information corresponding to the content to be tagged is obtained by performing text extraction on the corresponding image of the content to be tagged. Based on the theme information, object information, and text information corresponding to the content to be tagged, the second preprocessing content corresponding to the content to be tagged is determined.
4. The method according to claim 1, characterized in that, The content to be labeled is subjected to keyframe extraction, image content recognition, and speech conversion to obtain the third preprocessed content corresponding to the content to be labeled, including: Extract keyframe images from the video stream corresponding to the content to be labeled at preset time intervals; The keyframe image is subjected to image content recognition to obtain the visual content contained in the keyframe image; and the audio signal in the video stream corresponding to the content to be labeled is subjected to speech recognition conversion to convert the audio information into the corresponding text information. Based on the visual content contained in the keyframe image and the text information obtained after speech conversion, the third preprocessed content corresponding to the content to be labeled is determined.
5. The method according to claim 1, characterized in that, Using a large language model, based on the business scenario information corresponding to the content to be tagged, the target structured content is semantically matched with the multi-dimensional tag library to obtain target content tags corresponding to each content to be tagged, including: Using a large language model, based on the business scenario information corresponding to the content to be tagged, the sub-tag library corresponding to each sub-content in the target structured content is matched from the multi-dimensional tag library; Based on the business scenario rules corresponding to each matched sub-tag library, semantic feature matching is performed between each sub-tag library and its corresponding sub-content to obtain the target content tag corresponding to each content to be tagged.
6. The method according to claim 1, characterized in that, Before associating the content to be tagged with the corresponding target content tag, it also includes: Obtain the historical tagged content of the target content tag corresponding to the content to be tagged; Based on the correlation index between the content to be tagged corresponding to the target content tag and the historical tagged content, the correlation detection is performed between the content to be tagged and the corresponding target content tag.
7. The method according to claim 1, characterized in that, Before obtaining the batch import of multi-format content files, the following steps are also included: Check if the preset processing time has been reached at the current time; Alternatively, detect whether the upload task for the multi-format content file has been triggered.
8. A batch content tagging device based on a large model, characterized in that, include: The acquisition module is used to acquire multi-format content files imported in batches. The multi-format content files include: content to be tagged in different storage formats and business scenario information corresponding to the content to be tagged. The processing module is used to perform paragraph splitting, keyword extraction, and entity recognition on the text-based content to be tagged, to obtain the first preprocessed content corresponding to the content to be tagged; to identify the theme information, object information, and text information in the image-based content to be tagged, to obtain the second preprocessed content corresponding to the content to be tagged; and to extract keyframes, recognize the image content, and convert the video-based content to speech, to obtain the third preprocessed content corresponding to the content to be tagged. The conversion module is used to perform structured conversion on the first preprocessed content, the second preprocessed content, and the third preprocessed content to obtain the corresponding target structured content. A creation module is used to create a multi-dimensional tag library; the multi-dimensional tag library includes multiple sub-tag libraries, and different sub-tag libraries correspond to different business scenario rules; The matching module is used to perform semantic feature matching between the target structured content and the multi-dimensional tag library based on the business scenario information corresponding to the content to be tagged using a large language model, so as to obtain the target content tag corresponding to each content to be tagged; and associate the content to be tagged with the corresponding target content tag.
9. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the content batch labeling method based on a large model as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the content batch labeling method based on a large model as described in any one of claims 1 to 7.