New media material content collection management method and system based on image recognition
By combining OCR and a large vision-language model, explicit and implicit information of new media materials is extracted and high-quality labels are generated. This solves the problems of low efficiency and insufficient label accuracy in existing technologies, and enables efficient management and retrieval of material libraries.
Patent Information
- Application Number
- CN202510912616.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-09-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing new media material management method relies on manual labeling, which is inefficient and has inconsistent standards. Traditional automation technology has difficulty in deeply understanding image content, resulting in insufficient labeling accuracy, affecting the management efficiency and retrieval effect of the material library.
Optical character recognition (OCR) technology is used to extract explicit information from images, and combined with a pre-trained vision-language model (VLM) for visual narrative analysis. A double verification mechanism is used to generate the most relevant label set, improving label accuracy and reliability.
It has achieved automated and refined indexing of new media materials, significantly improved the management efficiency and retrieval experience of the material library, and ensured high quality and comprehensive coverage of labels.
Smart Images

Figure CN120611059A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent management, and more specifically, to a new media material content acquisition and management method and system based on image recognition. Background Art
[0002] With the advent of the digital age, the new media industry has experienced explosive growth, with the daily output of visual materials such as images and videos reaching enormous volumes. For content creators, media organizations, and platforms, efficiently collecting, storing, categorizing, and retrieving this vast amount of material has become a core challenge for improving content production efficiency and quality. A poorly managed and inefficient material library not only leads to the accumulation and waste of valuable resources but also directly hinders the agility and innovation of the creative process. Therefore, building an automated, intelligent new media material collection and management solution that enables refined, multi-dimensional indexing and rapid access to material has significant industry value and practical significance.
[0003] However, existing technologies still face significant limitations in achieving this goal. Currently, the management of new media material primarily relies on manual labeling or traditional computer vision technology. While manual methods offer advantages in understanding, they are plagued by inefficiencies, high costs, and inconsistent standards when faced with massive amounts of material. Deviations in the understanding of the same material by different operators lead to inconsistent labeling quality, making scalable application difficult. Traditional automated solutions, such as those based solely on simple keywords or basic object recognition, struggle to address the widespread complexity and diversity of new media image material. A single image often contains multiple objects, complex scene interactions, subtle emotional expressions, and even unique artistic styles—deep-level information that traditional technologies cannot capture. This makes it difficult for existing solutions to achieve a deep and accurate understanding of image content and automatically generate multi-dimensional, high-precision labels that fully reflect its meaning, severely hindering subsequent efficient retrieval and refined management.
[0004] Therefore, an optimized new media material content acquisition and management method based on image recognition is expected. Summary of the Invention
[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a new media material content collection and management method and system based on image recognition, which first uses optical character recognition (OCR) technology to accurately extract explicit information from the image, and uses a pre-trained visual-language large model (VLM) to perform a holistic visual narrative analysis of the image, thereby deeply exploring the implicit connotation of the image; then, the two information sources are merged, and candidate tags are identified and extracted from them; further, the confidence of the matching degree between the candidate tags and the original image is evaluated through the large model, and the most relevant final tag set is screened based on this. In this way, not only can explicit text and implicit visual narratives be fully covered, but also through a unique dual verification mechanism, the accuracy and reliability of the tags are significantly improved, and the automated, refined and high-quality indexing of massive new media materials is realized, thereby fundamentally improving the management efficiency and retrieval experience of the material library.
[0006] According to one aspect of the present application, a method for collecting and managing new media material content based on image recognition is provided, which includes: Obtain image material to be processed; Perform OCR recognition on the image material to be processed to obtain the text information displayed on the image material; Pass the image material to be processed through the pre-trained vision-language model to obtain a natural language description of the image material; Performing text merging on the text information displayed on the image material and the natural language description of the image material to obtain a comprehensive text; The comprehensive text is input into the BERT-based named entity recognition and keyword extraction model to obtain a list of candidate tags.
[0007] Based on the pre-trained visual-language model, calculate the confidence of each candidate label in the candidate label list; Based on the confidence of each candidate label, the candidate label list is filtered to obtain the final label set.
[0008] According to another aspect of the present application, a new media material content acquisition and management system based on image recognition is provided, which includes: An image material acquisition module is used to acquire image materials to be processed; An image material information extraction module is used to perform OCR recognition on the image material to be processed to obtain the text information displayed on the image material; An image description generation module is used to pass the image material to be processed through a pre-trained vision-language model to obtain a natural language description of the image material; A text merging module is used to merge the text information displayed on the image material and the natural language description of the image material to obtain a comprehensive text; The candidate tag success module is used to input comprehensive text into the BERT-based named entity recognition and keyword extraction model to obtain a candidate tag list.
[0009] The confidence calculation module is used to calculate the confidence of each candidate label in the candidate label list based on the pre-trained vision-language model; The tag filtering module is used to filter the candidate tag list based on the confidence of each candidate tag to obtain the final tag set.
[0010] Compared with the existing technology, the present application provides a new media material content collection and management method and system based on image recognition. It first uses optical character recognition (OCR) technology to accurately extract explicit information from the image, and uses a pre-trained visual-language model (VLM) to perform a holistic visual narrative analysis of the image, thereby deeply exploring the implicit connotation of the image; then, the two information sources are merged, and candidate tags are identified and extracted from them; further, the confidence level of the matching degree between the candidate tags and the original image is evaluated through the large model, and the most relevant final tag set is screened based on this. In this way, not only can explicit text and implicit visual narratives be fully covered, but also the accuracy and reliability of tags are significantly improved through a unique dual verification mechanism, and the automated, refined and high-quality indexing of massive new media materials is achieved, thereby fundamentally improving the management efficiency and retrieval experience of the material library. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0012] Figure 1 Flowchart of a method for collecting and managing new media material content based on image recognition according to an embodiment of the present application; Figure 2 A data flow diagram of a new media material content acquisition and management method based on image recognition according to an embodiment of the present application; Figure 3 Flowchart of sub-step S4 of the new media material content acquisition and management method based on image recognition according to an embodiment of the present application; Figure 4 4 is a block diagram of a new media material content acquisition and management system based on image recognition according to an embodiment of the present application. DETAILED DESCRIPTION
[0013] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0014] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0015] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.
[0016] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0017] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0018] To address the problems of existing technologies in which new media material management relies on manual labeling, which is inefficient and has inconsistent standards, as well as the superficial understanding of images and insufficient labeling accuracy of traditional automated technologies, a new media material content collection and management method and system based on image recognition is provided. This application aims to achieve automated, in-depth, and precise content analysis and indexing of image materials through an innovative technical approach, thereby significantly improving the management and retrieval efficiency of new media material libraries.
[0019] Specifically, in the technical solution of the present application, a new media material content acquisition and management method based on image recognition is proposed. Figure 1 Flowchart of a new media material content collection and management method based on image recognition according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of the new media material content collection and management method based on image recognition according to an embodiment of the present application. Figure 1 and Figure 2As shown, the new media material content acquisition and management method based on image recognition according to the embodiment of the present application includes the following steps: S1, obtaining the image material to be processed; S2, performing OCR recognition on the image material to be processed to obtain the image material display text information; S3, passing the image material to be processed through a pre-trained visual-language model to obtain a natural language description of the image material; S4, performing text merging on the image material display text information and the image material natural language description to obtain a comprehensive text; S5, inputting the comprehensive text into a BERT-based named entity recognition and keyword extraction model to obtain a candidate tag list; S6, calculating the confidence of each candidate tag in the candidate tag list based on the pre-trained visual-language model; S7, based on the confidence of each candidate tag, screening the candidate tag list to obtain a final tag set.
[0020] In particular, step S1, obtaining image material to be processed, can be achieved in a variety of ways in a specific example of this application. For example, the system can automatically retrieve images from the content library of a new media platform via an application programming interface (API); alternatively, it can receive local image files uploaded by an operator through a human-computer interface; the system can also be configured to monitor a specific server directory or cloud storage space in real time, and once a new image file is stored, it will automatically identify it as image material to be processed and load it into the processing queue.
[0021] It is worth mentioning that the range of image materials to be processed is wide, especially covering the visual materials combining pictures and texts commonly seen in the new media environment, such as commercial posters, illustrations for social media posts, online advertising banners, and product display pictures.
[0022] In particular, the S2 performs OCR recognition on the image material to be processed to obtain the text information displayed by the image material. It should be understood that new media image materials, such as posters or advertising pictures, usually contain key slogans, titles, brand names or explanatory texts, which are direct clues to understanding the core intention and theme of the image. Therefore, in order to accurately capture and digitize all text content directly presented in the image, in the technical solution of the present application, OCR recognition is performed on the image material to be processed to obtain the text information displayed by the image material. Here, the text information extracted by OCR is objective and literal, and is one of the cornerstones for subsequent semantic fusion and verification.
[0023] In a specific example of this application, OCR recognition of an image to be processed includes: first, performing image normalization on the image to be processed to obtain a standardized image to be processed. It should be understood that in practice, the original image may have issues such as uneven size, uneven lighting, noise, or a complex background, all of which can seriously interfere with the OCR engine's ability to determine character outlines. Therefore, normalization cleans and regularizes the image through a series of image processing algorithms. In practice, this may include resizing the image to a preset standard value, such as 1024x1024 pixels, to ensure consistent processing; performing grayscale conversion, i.e., converting a color image to grayscale, to reduce interference from color information and allow the algorithm to focus more on pixel brightness differences, thereby highlighting the edges and shape of text; and applying a filtering algorithm, such as a Gaussian filter, to smooth the image by calculating the weighted average of each pixel and its neighboring pixels, effectively removing random minor noise and significantly enhancing the clarity and contrast of the text area. After this series of processing, the output standardized image to be processed is more conducive to machine recognition.
[0024] Then, the Tesseract OCR model is used to perform OCR recognition on the standardized image material to be processed to obtain the text information displayed on the image material. Among them, Tesseract, as a mature open source OCR engine, its internal algorithm can analyze the input image. First, it performs layout analysis to locate the areas that may contain text, then performs character segmentation on these areas, and finally uses pattern matching and feature extraction technologies to identify each character and combine them into words and sentences. Finally, the model outputs a text string, that is, the text information displayed on the image material. Here, the text information displayed on the image material refers to the text data directly extracted from the image through OCR technology and is completely consistent with the visible text on the image. It is a structured, machine-readable string that is a transcription of the explicit information contained in the image.
[0025] Taking the solution of this application as an example, first, the image material to be processed (promotional poster) is standardized. For example, its size is uniformly adjusted to 1024x1024 pixels, grayscale conversion is performed, and a Gaussian filtering algorithm is used to remove slight noise in the image to enhance the clarity and contrast of the text area, thereby obtaining a standardized image material to be processed; then, the Tesseract OCR model is called to perform optical character recognition on the above-mentioned standardized image material to be processed. The model recognizes the artistic words on the poster and outputs the image material showing the text information: "Summer Limited Edition, Cool and Cool is Coming." This example vividly demonstrates the complete process of successfully and accurately extracting the core slogan from a complex poster image containing artistic words after standardized preprocessing and then recognition by the Tesseract model, providing high-quality input for subsequent text merging and analysis.
[0026] In particular, the S3 passes the image material to be processed through a pre-trained visual-language model to obtain a natural language description of the image material. It should be understood that new media images often carry richer stories and emotions than the text on the surface. For example, the composition, color, character expression, background environment, etc. of a poster together create a specific feeling or convey some implicit information. The pre-trained visual-language model, with its powerful cross-modal understanding ability, can integrate these complex visual elements and translate them in the form of natural language to obtain a rich and detailed semantic description of the image visual content itself. This description can capture non-text information such as image scenes, objects, atmosphere, actions, relationships and potential themes that OCR cannot reach.
[0027] In a specific example of this application, the pre-trained vision-language model is GPT-4V. In practice, the image material to be processed is first input into the GPT-4V model. The model then uses its complex internal neural network architecture (typically comprising a powerful visual encoder and language decoder, pre-trained on a large number of image-text pairs) to perform in-depth analysis and understanding of the image. This process involves identifying various objects and scene elements in the image (such as the sky, ocean, and beach chairs), understanding the spatial relationships and interactions between these objects and scenes, and even perceiving the overall style and emotional orientation of the image. The model internally converts these visual features into an internal representation. Its language generation module then translates this understanding into a coherent and natural textual description, known as a natural language description of the image material. Here, a natural language description of the image material is a model-generated text written in human language (such as Chinese or English) that describes the visual content of the input image. This description focuses on visual aspects of the image, such as the scene, objects, action, and atmosphere, contrasting with and complementing the textual information extracted by OCR.
[0028] Taking the solution of this application as an example, when the image material to be processed (a promotional poster) is fed into a pre-trained vision-language model (specifically, the GPT-4V model), the GPT-4V model performs in-depth analysis and understanding of the image, generating a natural language description of the image material. For example, the generated description reads: "This is a promotional image with a summer beach theme. The image features a beach chair covered with a beach towel, and a glass of lemon tea with ice and a lemon slice on a small table nearby. The background is a vast beach, a calm blue ocean, and a clear sky." In other words, when the original promotional poster (including elements such as the beach, beach chair, and lemon tea) is fed into the GPT-4V model, the model is able to output a detailed and accurate natural language description that captures the core visual content and atmosphere of the scene. This provides invaluable semantic information for subsequent comprehensive text generation and label extraction.
[0029] In particular, the S4 performs a text merging of the text information displayed on the image material and the natural language description of the image material to obtain a comprehensive text. It should be understood that although the text information displayed on the image material directly and accurately reflects the text on the image, it often lacks context and scene sense; and although the natural language description of the image material has a rich depiction of the visual scene, it may ignore the key marketing slogans or titles with high commercial value or core themes in the image. If the two are simply spliced or juxtaposed, an organic and coherent whole cannot be formed. Therefore, in order to overcome the limitations of a single information source and achieve complementary advantages and verification enhancement at the semantic level, this application uses a complex fusion mechanism to embed the focus information of the OCR text into the rich context of the visual description, and at the same time use the context of the visual description to verify and enrich the connotation of the OCR text, and finally produce a semantically coherent and comprehensive comprehensive text that contains core keywords and has a complete scene narrative, providing the most optimized text input for subsequent candidate tag generation.
[0030] In a specific example of this application, Figure 3 As shown, the S4 includes: S41, performing multi-source semantic mutual verification and compensation on the text information displayed by the image material and the natural language description of the image material to obtain the implicit coding features of the comprehensive text semantic enhancement; S42, obtaining the comprehensive text based on the implicit coding features of the comprehensive text semantic enhancement.
[0031] Specifically, the S41 performs multi-source semantic mutual verification and compensation on the image material display text information and the image material natural language description to obtain comprehensive text semantics enhanced implicit coding features. In an embodiment of the present application, first, the image material display text information is semantically embedded and encoded to obtain the image material display text information semantic embedding coding vector. That is, through semantic embedding coding, human-readable, discrete text strings are converted into continuous numerical representations that machines can understand and perform complex mathematical operations. In a specific example of the present application, the image material display text information can be semantically embedded and encoded by using a pre-trained BERT model. Specifically, after receiving the image material display text information output by the previous step (OCR recognition), the text will first be segmented, that is, the text will be divided into individual tokens (Tokens) according to the vocabulary preset by the model. For example, the text "Summer limited, cool and refreshing coming" may be divided into units such as "[CLS]", "summer", "day", "limited", "fixed", "," "ice", "cool", "coming", "attack", and "[SEP]", among which "[CLS]" and "[SEP]" are special markers used by the BERT model to indicate the beginning and separation of sequences; then, the word sequence is fed into the encoder of the BERT model; the multi-layer bidirectional Transformer structure within the BERT model can simultaneously consider the contextual information on both sides of a word to calculate its representation. As information flows through the multi-layer network structure, the model continuously refines and integrates the semantic information of the entire sequence. Ultimately, the model generates a high-dimensional vector for each input word, that is, the semantic embedding encoding vector of the text information displayed by the image material.
[0032] Next, semantic embedding encoding is performed on the natural language description of the image material to obtain a semantic embedding encoding vector of the natural language description of the image material. That is, through semantic embedding encoding, the natural language paragraph generated by the visual-language model that describes the rich scene and context of the image is also converted into a structured mathematical form that the machine can perform deep semantic operations. In a specific example of this application, semantic embedding encoding can be performed on the natural language description of the image material by using a pre-trained BERT model. Specifically, after receiving the natural language description of the image material output by the previous step (visual-language model), the text is first segmented. For example, a descriptive text is decomposed into a sequence of tokens. Then, this token sequence and its corresponding position information are input into the core architecture of the BERT model - a multi-layer bidirectional Transformer encoder. In each layer of the Transformer, the self-attention mechanism allows the model to fully consider the contextual information of all other tokens in the entire sentence when calculating the representation of each token. After layer-by-layer abstraction and information distillation of the multi-layer network, the model will eventually generate a context-related vector representation for each token in the sequence, namely the semantic embedding encoding vector of the natural language description of the image material.
[0033] Furthermore, semantic mutual verification and compensation are performed on the semantic embedding coding vector of the text information displayed in the image material and the semantic embedding coding vector of the natural language description of the image material to obtain a comprehensive text semantic enhancement implicit coding vector as a comprehensive text semantic enhancement implicit coding feature. It should be understood that the text recognized by OCR may be incomplete or incorrect due to image deformation, complex background or artistic fonts (for example, the inner page of a fashion magazine may be recognized as "Cool New Season" instead of the correct title "Colorful New Season"); while the description generated by the visual language model can capture the elements of the picture, it may ignore key text information (for example, a promotional poster with the text "Limited Time Offer" is only described as "a shopping mall scene with many people gathered"). This fragmented processing will cause the final label to lose core commercial information or artistic intent. Therefore, in the technical solution of this application, through deep interaction, the two types of information complement and correct each other to break through the bottleneck of traditional single-channel information processing. Specifically, text information provides precise named entity anchors for visual descriptions, while visual semantics gives the text a scene-based interpretation. In this way, not only the textual and visual essence of the image is highly concentrated, but also the unified and enhanced semantic connotation formed after the compensation of the two is integrated, thereby enhancing the credibility of the label.
[0034] Specifically, the semantic embedding coding vectors of the text information displayed on the image material and the semantic embedding coding vectors of the natural language description of the image material are first segmented into image material feature segments to obtain the sequence distribution of the local semantic feature coding vectors of the text information displayed on the image material and the sequence distribution of the local semantic feature coding vectors of the natural language description of the image material. It should be understood that although the semantic embedding coding vectors of the text information displayed on the image material and the semantic embedding coding vectors of the natural language description of the image material point to the same image content, they may produce semantic gaps or redundancy due to different information focuses (for example, text emphasizes brand information, while visual description focuses on scene atmosphere). If the overall vector is directly matched, it is easy to obscure key local semantic associations due to macro-feature alignment deviations. Therefore, in the technical solution of this application, by segmenting the semantic embedding vectors of the two modalities (OCR text vector and visual description vector) into a continuous, fine-grained sequence of feature segments, the system is able to break through the chaotic state of macro-semantics and enter into a precise analysis of local information structure. Ultimately, this fine-grained semantic deconstruction and reconstruction capability can achieve a deep portrayal and consistent expression of the overall meaning and context of the image content, providing a core driving force for subsequent high-precision tag construction and intelligent retrieval.
[0035] In a specific example of the present application, the semantic embedding coding vector of the text information displayed by the image material and the semantic embedding coding vector of the natural language description of the image material are segmented into image material feature segments to obtain the sequence distribution of the local semantic feature coding vector of the text information displayed by the image material and the sequence distribution of the local semantic feature coding vector of the natural language description of the image material; wherein the formula is: , , , , in, is the semantic embedding coding vector of the text information displayed by the image material, Indicates the feature segmentation operation, They are the first, second and third in the sequence distribution of the local semantic feature encoding vector of the text information displayed by the image material. and The image material shows the local semantic feature encoding vector of the text information, is the sequence distribution of the local semantic feature encoding vectors of the text information displayed by the image material, is the semantic embedding coding vector of the natural language description of the image material, They are the first, second and third in the sequence distribution of the local semantic feature encoding vector of the natural language description of the image material. and The local semantic feature encoding vector of the natural language description of the image material, It is the sequence distribution of the local semantic feature encoding vectors described in the natural language of the image material.
[0036] Next, semantic concept extraction is performed on each local semantic feature coding vector of the image material display text information in the sequence distribution of the local semantic feature coding vector of the image material display text information to obtain the sequence distribution of the local feature semantic concept coding vector of the image material display text information. It should be understood that after the system obtains the local sequence of the OCR text through feature segmentation, these low-level vector representations have grammatical relevance, but have not yet carried sufficient semantic depth and conceptual unity. Therefore, in the technical solution of the present application, the local feature vector is upgraded to an abstract semantic concept space through nonlinear mapping, so that each segment no longer stays at the surface symbol level, but is transformed into a conceptual unit with a clear semantic orientation. This abstraction process clears the ambiguity barrier for subsequent cross-modal semantic interaction, so that semantic units from different sources can be accurately paired and logically combined in an isomorphic conceptual space.
[0037] In a specific example of the present application, semantic concept extraction is performed on each of the local semantic feature encoding vectors of the image material display text information in the sequence distribution of the local semantic feature encoding vectors of the image material display text information to obtain a sequence distribution of the local feature semantic concept encoding vectors of the image material display text information; wherein the formula is: , in, and represents the activation function, and Respectively represent the local semantic distillation weight matrix of the text information displayed by the first and second image materials, represents the first bias vector, Indicates point multiplication by position, is the first in the sequence distribution of the semantic concept encoding vector of the local feature of the text information displayed by the image material The image material shows the local feature semantic concept encoding vector of text information.
[0038] Similarly, semantic concept extraction is performed on each local semantic feature coding vector of the natural language description of the image material in the sequence distribution of the local semantic feature coding vector of the natural language description of the image material to obtain the sequence distribution of the local feature semantic concept coding vector of the natural language description of the image material. It should be understood that although the visual language description carries the overall semantics of the image content, it is difficult to achieve precise structural alignment with the OCR text information in the original feature space due to the inherent continuity and ambiguity of natural language. Therefore, in the technical solution of the present application, the continuous language description is upgraded to the abstract concept space through nonlinear transformation, so that the local feature vector jumps from the linguistic surface structure (such as modification relationship, grammatical dependency) to the standardized concept entity. In this way, not only the difference in the degree of freedom of language expression is stripped away, but also a semantic coordinate system is established that is isomorphic to the OCR text concept unit, so that the two heterogeneous information sources can achieve point-to-point precise interaction in a unified dimension.
[0039] In a specific example of the present application, semantic concept extraction is performed on each of the local semantic feature coding vectors of the natural language description of the image material in the sequence distribution of the local semantic feature coding vectors of the natural language description of the image material to obtain a sequence distribution of the semantic concept coding vectors of the local feature of the natural language description of the image material; wherein the formula is: , in, and Represent the local semantic distillation weight matrix of the natural language description of the first and second image materials respectively, represents the second bias vector, is the first in the sequence distribution of the semantic concept encoding vector of the local feature of the natural language description of the image material The natural language description of the local features of the image material is encoded into a semantic concept vector.
[0040] Furthermore, semantic concept blending is performed on each corresponding pair of semantic concept encoding vectors for the local features of the text information displayed by the image material and the local features of the natural language description of the image material within the sequence distribution of the semantic concept encoding vectors for the local features of the text information displayed by the image material and the sequence distribution of the semantic concept encoding vectors for the local features of the natural language description of the image material, to obtain a sequence distribution of interaction encoding vectors for the local features of the semantic concept units of the comprehensive text. It should be understood that when the text concept units extracted by OCR and the description concept units generated by the visual model form two parallel sequences, although they point to the same material, they present a fragmented perspective. If only simple feature concatenation is performed, this heterogeneity will lead to the loss of deep semantic associations. Therefore, in the technical solution of this application, by projecting pairs of atomic concepts into the same interaction space for nonlinear fusion, the system can capture three core relationships while preserving their respective characteristics: semantic complementarity (e.g., seasonal elements not mentioned in the text are supplemented by visuals), mutual confirmation (e.g., "cherry blossoms" reinforce the credibility of "spring"), and even contradiction verification (e.g., promotional text conflicts with depressing images). This atomic interaction is like establishing connecting lines between nodes in a multidimensional semantic map, weaving fragmented concepts into an organic network, and establishing a prototype framework of cognitive structure for subsequent sequence coding.
[0041] That is, in the technical solution of the present application, the process of performing semantic concept mixing on each corresponding group of the sequence distribution of the semantic concept coding vectors of the local features of the text information displayed by the image material and the sequence distribution of the semantic concept coding vectors of the local features of the natural language description of the image material to obtain the sequence distribution of the interaction coding vectors of the local features of the semantic concept unit of the comprehensive text includes: First, the basic semantic fusion association vector of each corresponding pair of the sequence distribution of the semantic concept coding vectors of the local features of the text information displayed by the image material and the sequence distribution of the semantic concept coding vectors of the local features of the natural language description of the image material is calculated. The process is expressed as follows: , in, represents the activation function, is a linear interaction term, is the learnable weight matrix, for and The corresponding basic semantic fusion association vector; Next, the basic semantic fusion gate modulation vector of each corresponding group of the sequence distribution of the semantic concept coding vector of the local feature of the text information displayed by the image material and the sequence distribution of the semantic concept coding vector of the local feature of the natural language description of the image material is calculated. The process is expressed by the formula: , in, is a nonlinear gating term, Represent the learnable weight matrix, represents vector multiplication, for and The corresponding basic semantic fusion gate modulation vector; Then, the comprehensive interaction vector of the basic semantic fusion association vector and the basic semantic fusion gate modulation vector is calculated to obtain the comprehensive text local feature semantic concept unit interaction coding vector. The process is expressed by the formula: ,, , in, They are the first, second and third in the sequence distribution of the interaction encoding vector of the local feature semantic concept unit of the comprehensive text. and A comprehensive text local feature semantic concept unit interaction encoding vector, It is the sequence distribution of the interaction encoding vectors of the local feature semantic concept units of the comprehensive text.
[0042] Subsequently, the sequence distribution of the interaction encoding vector of the local feature semantic concept unit of the comprehensive text is subjected to concept unit sequence hybrid encoding based on the forward LSTM model to obtain the comprehensive text semantic enhanced implicit encoding vector. It should be understood that after the system generates a local interaction vector sequence through semantic concept mixing, these atomic concepts contain cross-modal associations, but are still discrete. If the fragmented concept units are directly aggregated without sequence modeling, the inherent semantic evolution structure and dynamic context of the material will be lost, resulting in the labeling system being able to only identify isolated elements but unable to reconstruct the main content. Therefore, in the technical solution of the present application, through the LSTM's unique gated memory mechanism (input gate to screen key concepts, forget gate to dilute background noise, output gate to organize semantic focus), the causal, transitional or parallel relationships between concepts are gradually established during the traversal of the sequence, and the static feature interaction unit is converted into a dynamic semantic flow. This ability to capture dynamic semantic structures ultimately enables the system to generate scenario-based intelligent tags that go beyond element stacking.
[0043] In a specific example of the present application, the sequence distribution of the interaction encoding vector of the semantic concept unit of the comprehensive text local feature is subjected to the concept unit sequence hybrid encoding based on the forward LSTM model to obtain the comprehensive text semantic enhancement implicit encoding vector; wherein, the formula is: , in, represents sequence encoding, is the comprehensive text semantics enhanced implicit coding vector.
[0044] Specifically, S42 is based on the semantic enhancement of the implicit coding features of the comprehensive text to obtain the comprehensive text. In the technical solution of this application, the semantic feature decoding of the semantic enhancement implicit coding vector of the comprehensive text is performed to obtain the comprehensive text. In other words, the high-dimensional semantic representation within the machine is converted into natural language that can be understood by humans, and then used for subsequent label extraction, manual review, or interactive use.
[0045] In specific implementation, the implicit encoding vector of the integrated text semantics enhancement is first input into the decoder's initial state. The decoder then maps the encoding vector to an initial word prediction through an internal loop iteration. Then, based on the previously generated word and encoding information, positional encoding and a multi-head attention mechanism are used to predict the next most likely word. This process is repeated until an end symbol is generated or the maximum length limit is reached. During decoding, the model optimizes a loss function (such as cross-entropy loss) to balance the accuracy of word order generation and the fluency of the text. This entire mechanism effectively captures the context, semantic hierarchy, and syntactic structure information implied in the implicit encoding vector of the integrated text semantics enhancement. The resulting integrated text not only covers the original image information but also integrates the multi-layer semantic associations between text and vision.
[0046] In particular, in S5, the comprehensive text is input into a BERT-based named entity recognition and keyword extraction model to obtain a candidate tag list. It should be understood that although the comprehensive text is semantically complete and rich in content, it is essentially a natural language paragraph and is not convenient for direct indexing, classification or search. Therefore, the comprehensive text is input into a BERT-based named entity recognition and keyword extraction model to automatically and efficiently extract a set of discrete entities and concepts with high information value from the semantically enhanced coherent text description, laying the foundation for subsequent tag screening and final material management.
[0047] In a specific example of this application, the BERT model first analyzes each token in the "synthetic text" one by one, and uses its powerful contextual understanding capabilities to determine whether these tokens constitute pre-defined entity categories. For scenarios with new media materials, these entity categories may include product names (such as "lemon tea"), events (such as "summer limited"), locations (such as "beach"), specific items (such as "beach chair"), or abstract concepts (such as "summer"). The model outputs the identified entities and their corresponding categories, for example, labeling "lemon tea" as "drinks."
[0048] Then, keyword extraction is performed on the comprehensive text to identify words in the text that are not related to specific named entities but are crucial for expressing the theme, emotion, or scene. Keyword extraction can be implemented based on a variety of algorithms. For example, the attention weights within the BERT model can be utilized. Those words that receive higher scores in the self-attention mechanism are usually the focus of the sentence. When these two tasks are completed, the system merges the entity list from the named entity recognition and the keyword list from the keyword extraction, removes duplicates, and finally forms an unfiltered, comprehensive list of candidate labels. Here, the candidate label list is a preliminary, potentially noisy set of labels that aims to comprehensively cover all potential related concepts and serve as a raw material library for subsequent verification and screening.
[0049] Taking the solution of this application as an example, a comprehensive text is input into a BERT-based named entity recognition and keyword extraction model. The model analyzes the comprehensive text, performing named entity recognition (NER) and keyword extraction tasks simultaneously, and generates a list of candidate tags. For example, if the input comprehensive text is "This is a promotional poster for 'Summer Limited Edition, Cool and Cool'...", the model generates the following "candidate tag list" after analysis: "["Summer Limited Edition", "Cool and Cool", "Poster", "Beach", "Beach Chair", "Lemon Tea", "Summer", "Drink", "Sky", "Ocean", "Car"]". While this process successfully extracts highly relevant core tags such as "Summer Limited Edition", "Beach", and "Lemon Tea", it may also introduce some irrelevant words due to training data bias or limited generalization capabilities. For example, the word "car" may be a misclassification caused by the model's training data bias.
[0050] In particular, S6 calculates the confidence level of each candidate label in the candidate label list based on a pre-trained visual-language model. It should be understood that the label extraction in the previous step is entirely based on the fused text. Although this text already contains visual descriptions, it may still lead to misjudgments due to limitations of the language model or bias in the training data (such as the "car" label in the embodiment). Therefore, in the technical solution of this application, the visual-language model is used to determine whether the concept represented by the label actually exists in the image. This approach can greatly improve the accuracy and credibility of the label.
[0051] In a specific example of the present application, the confidence of each candidate tag in the candidate tag list can be calculated by the following steps: First, for each candidate tag in the candidate tag list, a confidence question prompt word is constructed for each candidate tag, and the confidence question prompt word is "Does the image material to be processed contain the candidate tag?" At this stage, the system will traverse each tag in the "candidate tag list". For each tag, the system will embed it into a preset, closed question template, thereby dynamically generating a natural language question for the tag that can be answered by the visual-language model. For example, if one of the candidate tags in the list is "beach", the system will automatically construct a question prompt word like this: "Does the image material to be processed contain a beach?" This fixed template ensures the consistency and clarity of the question, allowing the model to focus on judging the existence or non-existence of specific concepts.
[0052] Then, the confidence question prompts for each candidate tag are input into the pre-trained visual-language model to obtain the confidence of each candidate tag. At this stage, the system combines the image material to be processed and the specific "confidence question prompts" constructed in the previous step as the complete input of a query, and submits it to the pre-trained visual-language model. After receiving the image and question, the model uses its powerful multimodal understanding capabilities to first analyze the visual content of the image, then understand the intent of the question, and finally search for evidence related to the question in the image. Based on its internal judgment logic, the model outputs a probability value indicating that the answer is "yes", which is the confidence of the candidate tag. This process is repeated for each tag in the candidate tag list until each tag obtains a quantified confidence score.
[0053] The confidence question prompt is a structured, natural language sentence template used for querying. It transforms the abstract verification task into a concrete question that can be answered by a language model. Confidence is a numerical value (typically between 0 and 1) that quantifies the degree to which the vision-language model believes a candidate label matches the original image content, or its credibility. Essentially, it represents the probability that the model will answer "yes" to the question.
[0054] Taking the solution of this application as an example, for the label "beach", the question constructed is: "Does the image material to be processed contain a beach?", and then the original poster image and this question are input into GPT-4V, and the confidence returned by the model is as high as 0.99. For "lemon tea" and "summer", which are highly relevant to the image content, the confidence levels also reached 0.98 and 0.95 respectively. As for the label "car" that was misjudged by the text model, when constructing and asking the question "Does the image material to be processed contain a car?", GPT-4V found no trace of a car in the poster image after examining it, so it returned an extremely low confidence level of 0.02. In this way, it is possible to accurately distinguish between labels that are highly relevant, relevant, and completely irrelevant to the image content, and assign each label an accurate quantitative indicator that can be used for subsequent screening.
[0055] In particular, the S7, based on the confidence of each candidate tag, screens the candidate tag list to obtain the final tag set. It should be understood that although the previous step has calculated a quantitative, visual fact-based confidence for each candidate tag, these tags with scores are still mixed together. In the technical solution of the present application, these confidence scores are further utilized to establish a clear admission standard, so that those tags that have low relevance to the image content or are even completely wrong (such as tags generated by text model misjudgment) are completely eliminated, and only those high-quality tags that have undergone double verification (text analysis and visual verification) are retained. This ensures that the final output tag set is not only comprehensive in content, but also accurate and reliable, which can greatly improve the efficiency and accuracy of subsequent material management, retrieval and recommendation systems.
[0056] During specific implementation, first, the system needs to set or load a predefined confidence threshold. This threshold is a key hyperparameter, representing the system's minimum requirement for the "credibility" of the label. Then, the system will traverse the data set generated in the previous step, which contains each candidate label and its corresponding confidence. For each candidate label in the list, the system will perform a numerical comparison of its confidence score with the preset confidence threshold. In response to a candidate label's confidence being greater than or equal to the threshold, the label is considered "qualified" and will be retained; conversely, in response to its confidence being lower than the threshold, the label is judged to be "unqualified" and will be discarded. The system collects all qualified labels that have passed this screening to form a new set, namely the final label set.
[0057] Among them, the confidence threshold is a floating point number between 0 and 1, which acts as the threshold for accepting or rejecting a candidate label. The setting of this threshold can balance the precision and recall of the final label set to a certain extent.
[0058] Taking the solution of the present application as an example, a confidence threshold is set, such as 0.80. Based on the confidence of each candidate tag calculated in step six, the candidate tag list is screened. Tags such as "beach" with a confidence of 0.99, "lemon tea" with a confidence of 0.98, and "summer" with a confidence of 0.95 will be retained, while "car" with a confidence of only 0.02 will be effectively eliminated. In the end, the final tag set obtained is: ["Summer Limited", "Ice", "Poster", "Beach", "Beach Chair", "Lemon Tea", "Summer", "Drink", "Ocean"]".
[0059] In summary, the new media material content collection and management method based on image recognition according to the embodiment of the present application is explained. It first uses optical character recognition (OCR) technology to accurately extract explicit information from the image, and uses a pre-trained visual-language model (VLM) to perform a holistic visual narrative analysis of the image, thereby deeply exploring the implicit connotation of the image; then, the two information sources are merged, and candidate tags are identified and extracted from them; further, the confidence of the matching degree between the candidate tags and the original image is evaluated through the large model, and the most relevant final tag set is screened based on this. In this way, not only can explicit text and implicit visual narratives be fully covered, but also the accuracy and reliability of tags are significantly improved through a unique dual verification mechanism, and the automated, refined and high-quality indexing of massive new media materials is achieved, thereby fundamentally improving the management efficiency and retrieval experience of the material library.
[0060] Furthermore, a new media material content acquisition and management system based on image recognition is also provided.
[0061] Figure 4 FIG is a block diagram of a new media material content collection and management system based on image recognition according to an embodiment of the present application. Figure 4As shown, the new media material content acquisition and management system 300 based on image recognition according to the embodiment of the present application includes: an image material acquisition module 310 for acquiring the image material to be processed; an image material information extraction module 320 for performing OCR recognition on the image material to be processed to obtain the image material display text information; an image description generation module 330 for passing the image material to be processed through a pre-trained visual-language model to obtain a natural language description of the image material; a text merging module 340 for performing text merging on the image material display text information and the image material natural language description to obtain a comprehensive text; a candidate tag success module 350 for inputting the comprehensive text into a BERT-based named entity recognition and keyword extraction model to obtain a candidate tag list. A confidence calculation module 360 is used to calculate the confidence of each candidate tag in the candidate tag list based on the pre-trained visual-language model; a tag filtering module 370 is used to filter the candidate tag list based on the confidence of each candidate tag to obtain a final tag set.
[0062] As described above, the new media material content acquisition and management system 300 based on image recognition according to the embodiment of the present application can be implemented in various wireless terminals, such as a server with a new media material content acquisition and management algorithm based on image recognition. In one possible implementation, the new media material content acquisition and management system 300 based on image recognition according to the embodiment of the present application can be integrated into the wireless terminal as a software module and / or hardware module. For example, the new media material content acquisition and management system 300 based on image recognition can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the new media material content acquisition and management system 300 based on image recognition can also be one of the many hardware modules of the wireless terminal.
[0063] Alternatively, in another example, the new media material content acquisition and management system 300 based on image recognition and the wireless terminal may also be separate devices, and the new media material content acquisition and management system 300 based on image recognition may be connected to the wireless terminal via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.
[0064] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A new media material content collection and management method based on image recognition, characterized in that: include: Obtain image material to be processed; Perform OCR recognition on the image material to be processed to obtain the text information displayed on the image material; Pass the image material to be processed through the pre-trained vision-language model to obtain a natural language description of the image material; Performing text merging on the text information displayed on the image material and the natural language description of the image material to obtain a comprehensive text; Input the comprehensive text into the BERT-based named entity recognition and keyword extraction model to obtain a list of candidate tags; Based on the pre-trained visual-language model, calculate the confidence of each candidate label in the candidate label list; Based on the confidence of each candidate label, the candidate label list is filtered to obtain the final label set.
2. The new media material content acquisition and management method based on image recognition according to claim 1 is characterized in that: Perform OCR on the image material to be processed to obtain the text information displayed on the image material, including: Performing image standardization processing on the image material to be processed to obtain a standardized image material to be processed; Use the Tesseract OCR model to perform OCR recognition on the standardized image material to be processed to obtain the text information displayed on the image material.
3. The new media material content acquisition and management method based on image recognition according to claim 1 is characterized in that: The pre-trained vision-language model is GPT-4V.
4. The new media material content acquisition and management method based on image recognition according to claim 1 is characterized in that: Based on the pre-trained vision-language model, the confidence of each candidate label in the candidate label list is calculated, including: For each candidate tag in the candidate tag list, construct a confidence question prompt word for each candidate tag, wherein the confidence question prompt word is "whether the image material to be processed contains the candidate tag"; The confidence question prompt words of each candidate tag are input into the pre-trained visual-language model to obtain the confidence of each candidate tag.
5. The new media material content acquisition and management method based on image recognition according to claim 1 is characterized in that: Performing text merging on the text information displayed on the image material and the natural language description of the image material to obtain a comprehensive text, including: Perform multi-source semantic cross-verification and compensation on the text information displayed on the image material and the natural language description of the image material to obtain the implicit coding features of the comprehensive text semantic enhancement; The implicit coding features are enhanced based on the semantics of the comprehensive text to obtain the comprehensive text.
6. The new media material content collection and management method based on image recognition according to claim 5 is characterized in that: Perform multi-source semantic cross-verification and compensation on the text information displayed on the image material and the natural language description of the image material to obtain comprehensive text semantic enhancement implicit coding features, including: Performing semantic embedding coding on the text information displayed on the image material to obtain a semantic embedding coding vector of the text information displayed on the image material; Performing semantic embedding coding on the natural language description of the image material to obtain a semantic embedding coding vector of the natural language description of the image material; Semantic mutual verification and compensation are performed on the semantic embedding coding vector of the text information displayed in the image material and the semantic embedding coding vector of the natural language description of the image material to obtain a comprehensive text semantic enhancement implicit coding vector as a comprehensive text semantic enhancement implicit coding feature.
7. The new media material content collection and management method based on image recognition according to claim 6 is characterized in that: Performing semantic mutual verification and compensation on the semantic embedding coding vector of the text information displayed on the image material and the semantic embedding coding vector of the natural language description of the image material to obtain a comprehensive text semantic enhancement implicit coding vector, including: Perform semantic concept fusion coding on the semantic embedding coding vector of the text information displayed on the image material and the semantic embedding coding vector of the natural language description of the image material to obtain the sequence distribution of the interaction coding vector of the semantic concept unit of the comprehensive text local feature; The sequence distribution of the interaction encoding vector of the semantic concept unit of the comprehensive text local feature is mixedly encoded based on the concept unit sequence of the forward LSTM model to obtain the semantic enhanced implicit encoding vector of the comprehensive text.
8. The new media material content collection and management method based on image recognition according to claim 5 is characterized in that: Based on the semantic enhancement of comprehensive text, implicit coding features are obtained to obtain comprehensive text, including: The semantic feature decoding of the semantically enhanced latent coding vector of the comprehensive text is performed to obtain the comprehensive text.
9. A new media material content acquisition and management system based on image recognition, characterized in that: include: An image material acquisition module is used to acquire image materials to be processed; An image material information extraction module is used to perform OCR recognition on the image material to be processed to obtain the text information displayed on the image material; An image description generation module is used to pass the image material to be processed through a pre-trained vision-language model to obtain a natural language description of the image material; A text merging module is used to merge the text information displayed on the image material and the natural language description of the image material to obtain a comprehensive text; The candidate tag success module is used to input the comprehensive text into the BERT-based named entity recognition and keyword extraction model to obtain a list of candidate tags; The confidence calculation module is used to calculate the confidence of each candidate label in the candidate label list based on the pre-trained vision-language model; The tag filtering module is used to filter the candidate tag list based on the confidence of each candidate tag to obtain the final tag set.