Poster Logo automatic extraction method and terminal based on multi-modal model
By employing multimodal model analysis and quality control mechanisms, the problems of low efficiency and unstable results in poster logo extraction have been solved, achieving efficient and accurate automatic logo extraction and intelligent verification, which is applicable to the fields of film and television promotion and distribution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-03
AI Technical Summary
In existing technologies, the extraction efficiency of poster logos is low, the results are unstable, and there is a lack of intelligent verification, making it difficult to meet the needs of short-cycle, large-scale application scenarios.
An automatic extraction method based on a multimodal model is adopted. The poster image is analyzed by a multimodal large model to generate prompt words to guide the image editing model to extract the logo. A multimodal quality inspection mechanism is used to conduct multi-dimensional checks to ensure the accuracy and completeness of the extraction results.
It enables efficient and accurate extraction of logos from posters, outputting standardized transparent logos suitable for various advertising and distribution scenarios, thus improving processing efficiency and the reliability of results.
Smart Images

Figure CN121789232A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, smart terminal, and storage medium for automatic extraction of poster logos based on a multimodal model. Background Technology
[0002] In film and television promotion and distribution, the logo is a crucial visual element, widely used in promotional posters, opening and closing credits, media releases, and other scenarios. Currently, common methods for logo extraction mainly include manual image cutout and traditional image segmentation / detection. Manual image cutout is time-consuming, labor-intensive, and inefficient, making it unsuitable for large-scale batch operations. Traditional image segmentation / detection methods often result in incomplete extraction results or the inclusion of redundant elements due to the complexity of poster designs, where the logo may be intertwined with the background, figures, or decorative elements. Furthermore, they lack accurate recognition of text style.
[0003] Therefore, existing technologies have the following problems: 1) Low extraction efficiency: Manual methods are difficult to meet the needs of short-cycle, large-scale application scenarios. 2) Unstable results: Traditional segmentation models are prone to misextracting people or backgrounds, making it difficult to guarantee that only the logo is retained. 3) Lack of intelligent verification: Existing methods cannot automatically determine whether the extracted logo is complete, whether the text is correct, and whether the style is consistent.
[0004] Therefore, existing technologies still need improvement and development. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a method, apparatus, smart terminal, and storage medium for automatic logo extraction from posters based on a multimodal model. This invention can automatically extract logos from film and television posters, and combined with an intelligent quality inspection mechanism, ensures the accuracy and usability of the extraction results, thereby improving extraction efficiency.
[0006] The technical solution of this application is as follows: A method for automatically extracting poster logos based on a multimodal model, comprising: Get movie poster images and movie titles; The poster image is analyzed using a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and its matching relationship with the movie title; and prompts suitable for image editing models are automatically generated to extract the logo. Based on the prompt words and poster image, a preset image editing model is called to extract the logo, remove irrelevant elements, and obtain the preliminary extracted logo; Through a multimodal quality inspection mechanism, the initially extracted logo is checked in multiple dimensions, including text content, style and handwriting, and element cleanliness. When the quality inspection is qualified, a standardized transparent logo is output.
[0007] The method for automatically extracting a poster logo based on a multimodal model includes the following steps: analyzing the poster image using a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and its matching relationship with the movie title. The poster image is analyzed using a multimodal large model to determine the position, color, and font style information of the logo, and outputs the location area of the logo, the main color of the logo, the font style and decorative element features, and the matching relationship between the logo and the movie title.
[0008] The aforementioned method for automatically extracting a poster logo based on a multimodal model, wherein the step of automatically generating prompts suitable for an image editing model and containing prompts for logo extraction includes: Based on the determined location of the logo, its main color, font style, decorative elements, and the matching relationship between the logo and the movie title, combined with the poster image, a prompt word suitable for the image editing model is automatically generated, which includes prompts for extracting the logo. The prompt word content includes the area of the logo to be retained, the elements to be removed, and color / style information.
[0009] The aforementioned method for automatically extracting a poster logo based on a multimodal model includes the following steps: Based on the prompt words and the poster image, a preset image editing model is invoked to extract the logo, remove irrelevant elements, and obtain a preliminary extracted logo. Input the poster image and the generated prompts into a preset image editing model; The image editing model performs targeted image cutout based on the prompt words, removes irrelevant elements, retains only the Logo area, and outputs the preliminary extracted Logo.
[0010] The aforementioned method for automatically extracting poster logos based on a multimodal model includes a multimodal quality inspection mechanism that performs multi-dimensional checks on the initially extracted logo, including text content, style and handwriting, and element cleanliness. The step of outputting a standardized, transparent logo when the quality inspection is passed includes: The initially extracted Logo is input into a multimodal model for quality inspection, which includes: Perform a text consistency check to determine whether the extracted logo matches the movie title; Perform style and handwriting detection to check whether the extracted logo retains the original font style and stroke characteristics; In addition, an element cleanliness test is conducted to confirm the presence of any figures, decorations, or superfluous elements; If the quality inspection fails, the prompt words are updated and the image editing model is called again for iterative extraction; Once the quality inspection is passed, a standardized transparent logo will be output.
[0011] The aforementioned method for automatically extracting a poster logo based on a multimodal model, wherein the steps of analyzing the poster image using a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and its matching relationship with the movie title; and automatically generating prompts suitable for image editing models with prompts for logo extraction, further include: The poster image is analyzed using a multimodal large model. When multiple languages are identified in the poster, a preliminary style analysis is performed on the logo area for each language. The generated prompt will contain the expected description of the logo in the target language; if the target is known to be to extract the English logo, the prompt will contain the logo with the English title preserved.
[0012] The aforementioned method for automatically extracting poster logos based on a multimodal model includes a multimodal quality inspection mechanism that performs multi-dimensional checks on the initially extracted logo, including text content, style and handwriting, and element cleanliness. The step of outputting a standardized, transparent logo when the quality inspection is passed includes: The system receives the initially extracted Logo, the original film title, and the target language film title through a cross-language semantic and style mapping module. A cross-lingual text-image embedding model is pre-set to understand the semantic associations of the same concepts in different languages and to map text descriptions to a visual style feature space. The original film title and the target language film title are input into a text encoder to obtain the semantic vectors of the original film title and the target language film title, respectively; at the same time, the initially extracted logo is passed through an image encoder to obtain the visual style vector of the initially extracted logo. Compare the semantic vector of the original movie title with the semantic vector of the target language movie title to determine whether they represent the same movie. During style mapping, the degree of matching between the visual style vector of the initially extracted logo and the semantic vector of the target language film name in the style feature space is compared. When performing style and handwriting detection, it is determined whether the style of the initially extracted logo conforms to the reasonable evolution of the original logo in the target language and cultural context as determined by the cross-language semantic and style mapping module. If the initially extracted logo is semantically consistent with the target language film title, and its stylistic changes fall within the reasonable localization range as determined by the cross-language semantic and style mapping module, then the quality inspection passes. Otherwise, update the prompt words and iterate again to extract them; When the quality inspection is passed, a standardized transparent logo will be output.
[0013] An automatic poster logo extraction device based on a multimodal model, wherein the device comprises: The acquisition module is used to acquire movie and TV poster images and movie titles; The positioning and prompt word generation module is used to analyze the poster image through a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and the matching relationship with the movie title; and automatically generate prompt words suitable for the image editing model to extract the logo prompt content. The multimodal model extraction module is used to extract the logo based on the prompt words and the poster image by calling a preset image editing model, removing irrelevant elements, and obtaining the preliminary extracted logo. The multimodal quality inspection and output module is used to perform multi-dimensional checks on the initially extracted logo through a multimodal quality inspection mechanism, including text content, style and handwriting, and element cleanliness. When the quality inspection is qualified, a standardized transparent logo is output.
[0014] A smart terminal includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors, the one or more programs comprising the method for performing any one of the methods.
[0015] A computer-readable storage medium, wherein, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the methods described above.
[0016] As described above, this application provides a method, apparatus, smart terminal, and storage medium for automatic logo extraction from posters based on a multimodal model. This invention first analyzes the poster using a multimodal model to determine the logo's location, color, font style, and other attributes, and automatically generates prompts suitable for an image editing model. Then, the image editing model is called to extract the logo. The extraction results undergo quality control by the multimodal model, ultimately outputting a standardized transparent logo. Furthermore, this invention has the following advantages: 1) Multimodal large model-assisted localization and analysis: Before image editing, the large model is used to determine the position, color, font style and other information of the logo, so as to improve the accuracy and robustness of extraction.
[0017] 2) Automatic prompt word generation: Based on the poster name and the analysis results of the large model, input prompt words for editing the model are dynamically generated to achieve targeted logo extraction.
[0018] 3) Image editing model targeted cutout: Under the guidance of prompts, accurately remove irrelevant elements such as people and backgrounds, keeping only the logo.
[0019] 4) Multimodal quality inspection mechanism: The integrity and correctness of the logo are ensured through triple quality inspection of text consistency, style detection and cleanliness detection. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart illustrating the automatic poster logo extraction method based on a multimodal model according to Embodiment 1 of the present invention.
[0022] Figure 2 This is a flowchart illustrating the automatic poster logo extraction method based on a multimodal model according to Embodiment 2 of the present invention.
[0023] Figure 3 The principle block diagram of an embodiment of the poster logo automatic extraction device based on a multimodal model provided by the present invention.
[0024] Figure 4 This is a block diagram illustrating the internal structure of a smart terminal provided in an embodiment of the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0026] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.
[0027] In film and television promotion and distribution, the logo is a crucial visual element, widely used in promotional posters, opening and closing credits, media releases, and other scenarios. Currently, common methods for logo extraction mainly include manual image cutout and traditional image segmentation / detection. Manual image cutout is time-consuming, labor-intensive, and inefficient, making it unsuitable for large-scale batch operations. Traditional image segmentation / detection methods often result in incomplete extraction results or the inclusion of redundant elements due to the complexity of poster designs, where the logo may be intertwined with the background, figures, or decorative elements. Furthermore, they lack accurate recognition of text style.
[0028] Therefore, existing technologies have the following problems: 1. Low extraction efficiency: Manual methods are difficult to meet the needs of short-cycle, large-scale application scenarios. 2. Unstable results: Traditional segmentation models are prone to mistakenly extracting people or backgrounds, making it difficult to guarantee that only the logo is retained. 3. Lack of intelligent verification: Existing methods cannot automatically determine whether the extracted logo is complete, whether the text is correct, and whether the style is consistent.
[0029] To address the aforementioned technical problems, this invention provides a method for automatically extracting logos from posters based on a multimodal model. This method is primarily applied in the field of film and television promotion and distribution, especially in scenarios requiring efficient and accurate logo extraction from a large number of film and television posters. For example, when film distribution companies promote new films, they need to quickly generate standardized logo materials for various online and offline promotional channels (such as social media, cinema advertisements, and merchandise design). Alternatively, in media content management platforms, it is necessary to archive and classify massive amounts of film and television posters; where the logo, as an important visual identifier, can be automatically extracted, greatly improving the efficiency and accuracy of content management. Furthermore, for developers of film and television merchandise, quickly acquiring high-quality logo materials is a key prerequisite for product design and production. The specific embodiments of this invention are described below.
[0030] Example 1 like Figure 1 As shown in the figure, an automatic poster logo extraction method based on a multimodal model according to an embodiment of the present invention includes the following steps: Step S100: Obtain the movie / TV show poster image and the movie title; The film and television poster images in this embodiment refer to visual materials used for the promotion and publicity of films and television dramas, including electronic files (such as images for media releases and online promotional images) and scanned copies of paper files, which serve as the carrier of the logo. The film title in this embodiment refers to the standard official title of the film or television drama (including complete information such as Chinese and English names, series numbers, etc.), which is the core reference for subsequent logo matching. In the specific implementation of step S100, the original images of the film and television posters to be processed are first collected, ensuring image clarity to meet the requirements of subsequent analysis; simultaneously, the standard name of the film and television series corresponding to the poster is obtained, which can be obtained through manual input, database association, etc., ensuring the accuracy of the name. After the two pieces of information are collected, they will serve as the basic input data for subsequent logo extraction. This step-by-step embodiment can improve subsequent efficiency because the present invention obtains the standard movie name in advance, reducing the information retrieval cost when matching the logo with the movie, making the matching process more accurate and faster; in addition, the present invention can also ensure data integrity because the present invention collects two types of key information, namely images and names, to avoid interruption of subsequent steps or deviation of extraction results due to missing information.
[0031] Step S200: Analyze the poster image using a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and its matching relationship with the movie title; and automatically generate prompts suitable for the image editing model to extract the logo content. The multimodal large model used in this embodiment refers to an artificial intelligence model that can simultaneously process multiple different types of information such as images and text, and perform cross-modal analysis. For example, a model that combines visual recognition and natural language understanding capabilities has stronger scene understanding and information integration capabilities.
[0032] In this embodiment, the key feature attributes of the logo body refer to the core identification information of the logo, including the logo's main color, font style, and decorative element features, which are unique attributes that distinguish it from the background and other elements. In this embodiment, the prompts refer to the instruction text used to guide the artificial intelligence model to perform specific tasks. The task requirements need to be conveyed accurately. For example, the prompts here are extracted from the specified logo so that the model can clearly understand the operation goal. The image editing model used in this embodiment refers to an artificial intelligence model specifically designed for image segmentation, cutout, element extraction, and other editing operations, which can perform targeted image processing based on prompts. In the specific implementation of step S200, the poster image is analyzed using a multimodal large model, including: 1) Visual analysis: scanning the poster image, accurately locating the logo's specific coordinates in the image, and distinguishing it from irrelevant parts such as people, background, and decorative elements; 2) Feature extraction: parsing the key feature attributes of the logo to form a structured feature description, such as circular outline, red background with white text, Song typeface, and minimalist style; 3) Matching verification: performing cross-modal matching between the extracted logo features and known movie names to confirm whether the logo is the official logo of the target movie, avoiding misidentification of other irrelevant logos; 4) Instruction generation: based on the above analysis results, automatically generating prompts that conform to the image editing model's recognition logic, such as extracting the movie logo with red background and white text in Song typeface style within the image coordinates (x1, y1) - (x2, y2) area, retaining the complete outline and text, and removing the background and surrounding decorative elements, providing clear guidance for subsequent extraction operations. This embodiment of the steps enables precise logo positioning because the multimodal large model used in this invention can accurately distinguish the logo from complex backgrounds, people, and other elements, solving the problem of ambiguous positioning in traditional methods. Furthermore, this invention can completely extract the logo's key attributes, providing a basis for subsequent quality inspection and standardization, while also facilitating accurate matching of the logo with movie titles. This step-by-step embodiment achieves high efficiency because the invention uses automatically generated targeted prompts, eliminating the need for manual instruction writing and enabling seamless integration of the analysis and extraction stages, thus significantly improving overall process efficiency.
[0033] Step S300: Based on the prompt words and poster image, call the preset image editing model to extract the Logo, remove irrelevant elements, and obtain the preliminary extracted Logo; The preset image editing model in this embodiment refers to a dedicated image editing model that has been pre-trained and adapted to the logo extraction scenario. It has been optimized for the diversity and complexity of poster logos, and has stronger logo extraction targeting capabilities. The irrelevant elements mentioned in this embodiment refer to all other parts of the poster except the target logo, including people, landscapes, decorative patterns, background textures, promotional slogans, etc. The initial logo extracted in this embodiment refers to the first draft of the logo that has been processed by the image editing model and most irrelevant elements have been removed, but has not yet undergone quality verification. In the specific implementation of step S300, based on the precise prompts generated in step S200, the original poster image is input into a preset dedicated image editing model. The model performs targeted segmentation and cutout operations on the Logo according to the location information and feature requirements in the prompts, accurately preserving the complete main body of the Logo while completely removing irrelevant elements such as background, figures, and decorations, ultimately outputting a preliminary extracted Logo file. This embodiment of the process achieves high extraction efficiency because the image editing model used in this invention automates the process, eliminating the need for manual image cutout and significantly increasing processing speed compared to manual methods, thus meeting the needs of large-scale batch operations. Furthermore, this invention boasts high extraction accuracy because it relies on precise pre-defined prompts, allowing the model to extract only the target logo, avoiding the common problem of traditional methods that mistakenly extract people or backgrounds, thus greatly improving the purity of the extraction results. Moreover, based on the guidance of key logo features, the model can completely preserve the logo's outline, text, colors, and other core elements, solving the problem of incomplete extraction in traditional methods.
[0034] Step S400: Through a multimodal quality inspection mechanism, the initially extracted Logo is inspected in multiple dimensions, including text content, style and handwriting, and element cleanliness. When the quality inspection is qualified, a standardized transparent Logo is output.
[0035] The multimodal quality inspection mechanism in this embodiment refers to a quality inspection mechanism that combines multiple modal capabilities such as text recognition and image analysis, which can comprehensively verify the extraction results from different dimensions, rather than judging from a single dimension. The text content inspection mentioned in this embodiment refers to verifying, through text recognition technology, whether the extracted text in the logo is consistent with the movie title, whether there are any typos, and whether the text is complete, such as without missing or redundant characters. The style and handwriting check in this embodiment refers to verifying whether the extracted logo style, font, and handwriting are consistent with the style of the logo in the original poster, such as no distortion or style deviation, to ensure the brand recognition of the logo. The element cleanliness check in this embodiment refers to checking whether there are any irrelevant impurities such as background or decoration remaining in the initially extracted logo, whether the edges are clear and free of blurry jagged edges, and whether the overall logo is clean. The standardized transparent logo described in this embodiment refers to a standardized logo file that has passed quality inspection, has a uniform format (such as PNG transparent format), accurate content, consistent style, and no impurities. It can be directly applied to various scenarios such as promotional posters, opening and closing credits, and media releases. Specifically, step S400 performs a three-fold verification of the initially extracted Logo using a multimodal quality inspection mechanism. This includes: first, text content verification, ensuring the text in the Logo perfectly matches the movie title without errors or omissions; second, style and handwriting verification, ensuring the Logo's style, font, etc., are consistent with the original poster without distortion or deviation; and third, element cleanliness verification, checking for any residual impurities and ensuring clear, clean edges. If all three verifications pass, the Logo is considered quality-approved and a standardized transparent Logo is directly output. If any non-compliance exists (such as missing text or residual impurities), it can be fed back to the previous steps for secondary optimization (such as regenerating prompts or re-extracting) until it passes quality inspection and is output. As can be seen, the embodiments of the present invention can make up for the shortcomings of traditional methods and solve the problem of lack of intelligent verification in existing technologies through the above steps. The reliability of the extraction results is ensured through multi-dimensional checks. Furthermore, the present invention can also guarantee the output quality by controlling the three core dimensions of text, style and purity to ensure that the final output logo is accurate, complete and consistent, and can be directly applied to various promotional and distribution scenarios. Moreover, this invention can reduce rework costs because it can detect and correct defects in the initial extraction in advance, preventing substandard logos from flowing into subsequent use stages and reducing advertising accidents or rework caused by logo issues. Furthermore, the invention standardizes the output format and quality standards, which is also conducive to rapid application in different scenarios, improving the reusability and efficiency of logos.
[0036] In a further embodiment of the present invention, the detailed step S100 of the present invention further includes: Retrieves film and television poster images from the media resource library and obtains the film title as supplementary information; For example, retrieve the poster image of "A Certain Wave Earth 2" from the media resource library and obtain the film title "A Certain Wave Earth 2" as supplementary information.
[0037] This invention, by standardizing the acquisition of poster images and movie titles, provides accurate and complete raw data for subsequent multimodal large-scale model analysis. The movie title, as auxiliary information, can significantly improve the accuracy of multimodal large-scale model recognition of logo text content, especially when the logo text is integrated with the background or has a high degree of artistic font, effectively avoiding misidentification and thus improving the overall accuracy and robustness of extraction.
[0038] In a further embodiment of the present invention, the step of analyzing the poster image using a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and the matching relationship with the movie title includes: The poster image is analyzed using a multimodal large model to determine the position, color, and font style information of the logo, and outputs the location area of the logo, the main color of the logo, the font style and decorative element features, and the matching relationship between the logo and the movie title.
[0039] The steps of automatically generating prompts suitable for image editing models and containing prompts for logo extraction include: Based on the determined location of the logo, its main color, font style, decorative elements, and the matching relationship between the logo and the movie title, combined with the poster image, a prompt word suitable for the image editing model is automatically generated, which includes prompts for extracting the logo. The prompt word content includes the area of the logo to be retained, the elements to be removed, and color / style information.
[0040] The above two steps are the detailed steps of Logo positioning and prompt word generation in this embodiment of the invention. For example, suppose we need to automatically extract the logo from the poster of the movie "A Certain Wave Earth 2".
[0041] This invention can retrieve the poster image of "A Certain Wave Earth 2" from a media resource library and obtain the movie title "A Certain Wave Earth 2" as auxiliary information.
[0042] When performing logo positioning and prompt generation, the poster image and the movie title "A Certain Wave Earth 2" are input into the multimodal large model. The multimodal large model analyzes the poster and outputs the following information: The approximate location of the logo: for example, centered at the bottom of the poster.
[0043] Logo features include its main color, font style, and decorative elements. For example, the logo is white, with a sci-fi bold font style, metallic texture, and earth-themed decorations.
[0044] Matching relationship with the movie title: The text content of the logo is highly matched with "Mou Lang Earth 2".
[0045] Based on the above information, the present invention can automatically generate prompts suitable for the image editing model. These prompts may include: retaining the white, bold, metallic logo centered at the bottom of the poster; removing figures, background buildings, and starry sky decorations around the logo; and ensuring that the extracted logo retains its original white color and sci-fi style.
[0046] In a further embodiment of the present invention, the step of extracting the logo by calling a preset image editing model based on the prompt word and the poster image, removing irrelevant elements, and obtaining a preliminary extracted logo includes: Input the poster image and the generated prompts into a preset image editing model; The image editing model performs targeted image cutout based on the prompt words, removes irrelevant elements, retains only the Logo area, and outputs the preliminary extracted Logo.
[0047] In specific implementation of this refinement step, for example, following the above embodiment, the generated prompt word and the poster image of "A Certain Wave Earth 2" are input into the image editing model. Then, this invention uses the editing model to perform targeted image matting based on the prompt word, accurately removing irrelevant elements such as people, background buildings, and starry skies around the logo in the poster, retaining only the "A Certain Wave Earth 2" logo area. The system outputs the initially extracted logo, which can be in black background format. Thus, the image editing model of this invention, through the refinement step, performs targeted image matting on the poster based on this precise prompt word, thereby accurately extracting the logo.
[0048] In a further embodiment of the present invention, the step of performing multi-dimensional checks on the initially extracted Logo through a multi-modal quality inspection mechanism, including text content, style and handwriting, and element cleanliness, and outputting a standardized transparent Logo when the quality inspection is qualified, includes: The initially extracted Logo is input into a multimodal model for quality inspection, which includes: Perform a text consistency check to determine whether the extracted logo matches the movie title; Perform style and handwriting detection to check whether the extracted logo retains the original font style and stroke characteristics; In addition, an element cleanliness test is conducted to confirm the presence of any figures, decorations, or superfluous elements; If the quality inspection fails, the prompt words are updated and the image editing model is called again for iterative extraction; Once the quality inspection is passed, a standardized transparent logo will be output.
[0049] The above steps of this invention are a detailed refinement of the multimodal quality inspection process. In specific implementation, we will still use a poster of "A Certain Wave Earth 2" as an example. In this refined step, a text consistency check is first performed. The system of this invention determines whether the initially extracted logo text is consistent with "A Certain Wave Earth 2". Then, style and handwriting detection is performed to check whether the extracted logo retains the original sci-fi bold font style and metallic stroke characteristics.
[0050] Next, an element cleanliness check is performed to confirm whether there are any residual figures, backgrounds, or other redundant elements in the extracted logo. If the quality check finds errors in the logo text, style inconsistencies, or redundant elements, the system will update the prompts (e.g., emphasizing the removal of background details) and re-invoke the image editing model for iterative extraction until the logo fully meets the requirements.
[0051] In its specific implementation, this invention also involves removing the background from the quality-inspected logo and converting it into a transparent PNG format, allowing it to be seamlessly overlaid on various promotional materials without additional manual processing, thus improving the logo's versatility. Then, the logo's resolution and aspect ratio are adjusted according to application requirements. Edge noise is removed, font clarity is enhanced, and a final, qualified transparent logo for "Motorway Earth 2" is generated and stored, ensuring high-quality presentation of the logo on different sizes and display media, enhancing its visual appeal and professionalism.
[0052] In this embodiment of the invention, to ensure the quality of the extraction results, the system will activate a multimodal quality inspection mechanism to perform multi-dimensional checks on the extracted Logo, including text content, style and handwriting, and element cleanliness. If the quality inspection finds any non-compliance, the system will intelligently update the prompt words and call the image editing model again for iterative extraction until a high-quality, standardized transparent Logo is obtained.
[0053] This invention ultimately generates and stores a qualified logo, providing a ready-to-use, high-quality visual asset for subsequent promotional and distribution work, reducing manual intervention and post-production costs.
[0054] In another embodiment of the present invention, to address the technical problem faced during the promotion of multilingual international films and television dramas—the challenge of verifying the semantic and stylistic consistency of multilingual logos—the present invention further proposes introducing a cross-language semantic and style mapping module into the existing multimodal quality inspection mechanism to achieve intelligent recognition and quality inspection of localized logos. Specifically, the step of analyzing the poster image using a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and its matching relationship with the film title, and automatically generating prompts suitable for image editing models to extract logo prompts, further includes: S101. Analyze the poster image using a multimodal large model. When multiple languages are identified in the poster, perform a preliminary style analysis on the logo area of each language. That is, based on the above embodiments, the multimodal large model-assisted localization and analysis remain unchanged, but its output information will increase the preliminary identification of potential localization style features. For example, it can identify multiple language texts that may exist in the poster and perform preliminary style analysis on the logo area of each language.
[0055] S102. The generated prompt will contain the expected description of the logo in the target language; if the target is known to be to extract the English logo, the prompt will contain the English title logo.
[0056] In this embodiment, the automatic generation of prompts remains unchanged, but the generated prompts will contain the expected description of the target language logo. For example, if the target is known to be to extract the English logo, the prompts will include retaining the English title logo, whose style should be semantically related to the original Chinese logo but may have localization design differences.
[0057] In a further embodiment of the multilingual method of the present invention, the automatic poster logo extraction method based on a multimodal model includes the step of performing multi-dimensional checks on the initially extracted logo through a multimodal quality inspection mechanism, including text content, style and handwriting, and element cleanliness, and outputting a standardized transparent logo when the quality inspection is qualified. S201. Receive the initially extracted Logo, original film title, and target language film title through a cross-language semantic and style mapping module. For example, through a cross-language semantic and style mapping module, the system first receives the extracted logo image, the original film title (e.g., "The Wandering Earth 2"), and the target language film title (e.g., "The Wandering Earth 2"). S202. A cross-lingual text-image embedding model is pre-set to understand the semantic associations of the same concepts in different languages and to map text descriptions to a visual style feature space. The pre-trained cross-language text-image embedding model in this embodiment of the invention can understand the semantic associations of the same concepts in different languages and can map text descriptions to a visual style feature space.
[0058] S203. Input the original film title and the target language film title into the text encoder to obtain the semantic vector of the original film title and the semantic vector of the target language film title respectively; at the same time, pass the initially extracted logo through the image encoder to obtain the visual style vector of the initially extracted logo. In this embodiment, the original film title and the target language film title are input into a text encoder to obtain their semantic vectors. Simultaneously, the extracted logo image is also processed by an image encoder to obtain its visual style vector.
[0059] S204. Compare the semantic vector of the original movie title with the semantic vector of the target language movie title to confirm whether they represent the same movie. In this embodiment, when determining semantic relevance, the semantic vector of the original movie title is compared with the semantic vector of the target language movie title to confirm whether they represent the same movie.
[0060] S205. During style mapping, compare the degree of matching between the visual style vector of the initially extracted Logo and the semantic vector of the target language film name in the style feature space. In this embodiment, style mapping is achieved by comparing the degree of matching between the extracted visual style vector of the logo and the semantic vector of the target language film title in the style feature space. For example, if the original Chinese logo has a sci-fi, bold, metallic feel, while the target English logo is designed with a futuristic, streamlined style, this invention can identify this as a reasonable localization style evolution, rather than a simple inconsistency. Thus, this invention achieves its goal by analyzing numerous cross-language logo design cases to learn common mapping patterns of logo styles across different cultural backgrounds.
[0061] S206. When performing style and handwriting detection, determine whether the style of the initially extracted Logo conforms to the reasonable evolution of the original Logo in the target language and cultural background as determined by the cross-language semantic and style mapping module. In this embodiment, style and handwriting detection is no longer simply about "whether the original style is maintained," but rather about determining whether the extracted logo style conforms to the "cross-language semantic and style mapping module's" assessment of a reasonable evolution of the original logo within the target language and cultural context. For example, if the module determines that the English logo should have a "futuristic" feel rather than a strictly "metallic" feel, then quality control will use this as the standard. S207. If the initially extracted logo is semantically consistent with the target language film title, and its style changes fall within the reasonable localization range under the judgment of the cross-language semantic and style mapping module, then the quality inspection passes. S208. Otherwise, update the prompt words and iterate again to extract them. S209. When the quality inspection is qualified, output a standardized transparent logo.
[0062] In this embodiment of the invention, if the extracted logo is semantically consistent with the target language film title, and its stylistic variations fall within the reasonable localization range as determined by the cross-language semantic and style mapping module, then the quality check passes. Otherwise, the system will update the prompt words based on the module's feedback, for example, more explicitly instructing the image editing model to "retain the futuristic English title logo," and then iterate and extract again.
[0063] As can be seen from the above, by introducing a cross-language semantic and style mapping module, this embodiment can intelligently identify and accept reasonable semantic and style changes generated by multilingual logos during the localization process, thereby solving the problem that existing quality inspection mechanisms may misjudge localized logos in cross-cultural promotion scenarios, and greatly improving the accuracy and applicability of automated extraction.
[0064] The present invention will be further described in detail below through specific application examples: like Figure 2 As shown in the second specific application embodiment, a method for automatically extracting poster logos based on a multimodal model is provided, which includes the following steps: Step S10: Begin, proceed to S11; Step S11: Retrieve the movie title and poster from the media resource library, then proceed to S12; In this embodiment, when acquiring poster information, film and television poster images can be read from the media resource library; and the film name can be obtained as auxiliary information.
[0065] Step S12: Combine the movie title to obtain extraction prompts, and proceed to step S13; In this specific embodiment, during the location and prompt word generation, the poster image and movie title are input into the multimodal big data model; the multimodal big data model analyzes the poster and outputs the following information: the approximate location area of the logo (such as top center, bottom left corner, etc.); the main color, font style and decorative element characteristics of the logo; and the matching relationship with the movie title.
[0066] Then, based on the above information, prompts suitable for the image editing model are automatically generated. The prompts include: the area to be retained (the approximate area of the logo); the elements to be removed (people, background decorations, etc.); and color / style information to ensure that the extracted style is consistent with the original style.
[0067] Step S13: Extract the Logo using the multimodal image editing model and proceed to step S14; In this specific embodiment, when extracting the logo using a multimodal model, the generated prompt words and poster image are first input into an image editing model; the editing model performs targeted image matting based on the prompt words, removing irrelevant elements and retaining only the logo area; then the initially extracted logo is output, which is generally in black background format.
[0068] Step S14: Multimodal large model quality inspection: whether the characters are correct, whether the resolution meets the standard, and whether other elements have been introduced. If any one of these conditions is met, proceed to step S15; if all conditions are met, proceed to step S16. In this specific embodiment, the text consistency detection determines whether the extracted logo is consistent with the movie title; the style and handwriting detection checks whether the extracted logo retains the original font style and stroke characteristics; and the element cleanliness detection confirms whether there are any characters, decorations, or other redundant elements.
[0069] Step S15: If the quality inspection fails, update the prompt word and re-call the image editing model for iterative extraction, and determine whether the number of attempts is greater than N. If not, proceed to step S13; otherwise, proceed to step S18. If the quality inspection fails, update the prompt words and re-initiate the image editing model for iterative extraction. This can be attempted N times, with N being optimally 3. Step S16: Extract the corresponding Logo using the Logo segmentation model, and then proceed to S17; Step S17: Obtain the final logo, then proceed to step S18; In this embodiment of the invention, the extracted Logo will also undergo post-processing, including: removing the background by using a Logo segmentation model to process the output image and remove pure black, pure white, or complex backgrounds; making the background transparent by converting the Logo to a transparent background PNG format; scaling the size, such as adjusting the Logo's resolution and aspect ratio according to application requirements; image optimization by removing edge noise and enhancing font clarity and overall appearance; and then generating and storing the final qualified Logo.
[0070] Step S18, End.
[0071] As can be seen from the above, this embodiment of the invention employs a multimodal large model-assisted localization and analysis. Before image editing, the large model is used to determine information such as the logo's position, color, and font style, improving the accuracy and robustness of the extraction. Furthermore, this invention also uses automatic prompt word generation, dynamically generating input prompt words for the editing model by combining the poster name and the large model analysis results, achieving targeted logo extraction.
[0072] Furthermore, this invention employs an image editing model for targeted image matting, accurately removing irrelevant elements such as people and backgrounds under the guidance of prompts, retaining only the logo. Finally, a multimodal quality inspection mechanism is used, employing triple checks of text consistency, style detection, and cleanliness detection to ensure the integrity and correctness of the logo.
[0073] Exemplary device like Figure 3 As shown, this embodiment of the invention provides a poster logo automatic extraction device based on a multimodal model, the device comprising: The acquisition module 310 is used to acquire movie poster images and movie titles; The positioning and prompt word generation module 320 is used to analyze the poster image through a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and the matching relationship with the movie title; and automatically generate prompt words suitable for the image editing model to extract the logo prompt content. The multimodal model extraction module 330 is used to extract the logo based on the prompt words and the poster image by calling a preset image editing model, removing irrelevant elements, and obtaining the initially extracted logo. The multimodal quality inspection and output module 340 is used to perform multi-dimensional checks on the initially extracted Logo through a multimodal quality inspection mechanism, including text content, style and handwriting, and element cleanliness. When the quality inspection is qualified, a standardized transparent Logo is output, as described above.
[0074] Based on the above embodiments, the present invention also provides a smart terminal, the principle block diagram of which can be as follows: Figure 4 As shown. The intelligent terminal includes a processor, memory, network interface, display screen, and database connected via a system bus.
[0075] The memory stores one or more programs configured to be executed by a processor to implement the poster logo automatic extraction method based on a multimodal model as described in the above embodiments.
[0076] Here, "intelligent terminal" refers to a smart computer or similar device with data processing capabilities. The memory can be internal memory, flash memory, hard disk, or cloud storage, used to store program code, preset image editing models, and various data such as preset multimodal quality inspection mechanisms. The processor can be a central processing unit (CPU), used to execute the algorithmic logic within the program. The program includes a method for automatically extracting poster logos based on a multimodal model.
[0077] In a further embodiment, a smart terminal of this embodiment includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: Get movie poster images and movie titles; The poster image is analyzed using a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and its matching relationship with the movie title; and prompts suitable for image editing models are automatically generated to extract the logo. Based on the prompt words and poster image, a preset image editing model is called to extract the logo, remove irrelevant elements, and obtain the preliminary extracted logo; The initial logo is inspected from multiple dimensions using a multimodal quality control mechanism, including text content, style and handwriting, and element cleanliness. Once the quality control is passed, a standardized transparent logo is output, as described above.
[0078] The step of analyzing the poster image using a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and its matching relationship with the movie title includes: The poster image is analyzed using a multimodal large model to determine the position, color, and font style information of the logo, and outputs the location area of the logo, the main color of the logo, the font style and decorative element features, and the matching relationship between the logo and the movie title.
[0079] The step of automatically generating prompts suitable for the image editing model and containing prompts for logo extraction includes: Based on the determined location of the logo, its main color, font style, decorative elements, and the matching relationship between the logo and the movie title, combined with the poster image, a prompt word suitable for the image editing model is automatically generated, which includes prompts for extracting the logo. The prompt word content includes the area of the logo to be retained, the elements to be removed, and color / style information.
[0080] The step of extracting the logo based on the prompt word and the poster image, and removing irrelevant elements to obtain the preliminary extracted logo includes: Input the poster image and the generated prompts into a preset image editing model; The image editing model performs targeted image cutout based on the prompt words, removes irrelevant elements, retains only the Logo area, and outputs the preliminary extracted Logo.
[0081] The step of using a multimodal quality inspection mechanism to perform multi-dimensional checks on the initially extracted logo, including text content, style and handwriting, and element cleanliness, and outputting a standardized transparent logo when the quality inspection is passed, includes: The initially extracted Logo is input into a multimodal model for quality inspection, which includes: Perform a text consistency check to determine whether the extracted logo matches the movie title; Perform style and handwriting detection to check whether the extracted logo retains the original font style and stroke characteristics; In addition, an element cleanliness test is conducted to confirm the presence of any figures, decorations, or superfluous elements; If the quality inspection fails, the prompt words are updated and the image editing model is called again for iterative extraction; Once the quality inspection is passed, a standardized transparent logo will be output.
[0082] The step of analyzing the poster image using a multimodal large model to determine the location of the logo, its key feature attributes, and its matching relationship with the movie title, and automatically generating prompts suitable for image editing models to extract logo information, further includes: The poster image is analyzed using a multimodal large model. When multiple languages are identified in the poster, a preliminary style analysis is performed on the logo area for each language. The generated prompt will contain the expected description of the logo in the target language; if the target is known to be to extract the English logo, the prompt will contain the logo with the English title preserved.
[0083] The step of using a multimodal quality inspection mechanism to perform multi-dimensional checks on the initially extracted logo, including text content, style and handwriting, and element cleanliness, and outputting a standardized transparent logo when the quality inspection is passed, includes: The system receives the initially extracted Logo, the original film title, and the target language film title through a cross-language semantic and style mapping module. A cross-lingual text-image embedding model is pre-set to understand the semantic associations of the same concepts in different languages and to map text descriptions to a visual style feature space. The original film title and the target language film title are input into a text encoder to obtain the semantic vectors of the original film title and the target language film title, respectively; at the same time, the initially extracted logo is passed through an image encoder to obtain the visual style vector of the initially extracted logo. Compare the semantic vector of the original movie title with the semantic vector of the target language movie title to determine whether they represent the same movie. During style mapping, the degree of matching between the visual style vector of the initially extracted logo and the semantic vector of the target language film name in the style feature space is compared. When performing style and handwriting detection, it is determined whether the style of the initially extracted logo conforms to the reasonable evolution of the original logo in the target language and cultural context as determined by the cross-language semantic and style mapping module. If the initially extracted logo is semantically consistent with the target language film title, and its stylistic changes fall within the reasonable localization range as determined by the cross-language semantic and style mapping module, then the quality inspection passes. Otherwise, update the prompt words and iterate again to extract them; When the quality inspection is passed, a standardized transparent logo will be output, as described above.
[0084] This embodiment also provides a computer-readable storage medium, wherein when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is able to perform the method described in any of the above embodiments.
[0085] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0086] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for automatically extracting poster logos based on a multimodal model, characterized in that, include: Get movie poster images and movie titles; The poster image is analyzed using a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and its matching relationship with the movie title; and prompts suitable for image editing models are automatically generated to extract the logo. Based on the prompt words and poster image, a preset image editing model is called to extract the logo, remove irrelevant elements, and obtain the preliminary extracted logo; The initial logo is inspected from multiple dimensions using a multimodal quality control mechanism, including text content, style and handwriting, and element cleanliness. Once the quality control is passed, the logo is output.
2. The method for automatic poster logo extraction based on a multimodal model according to claim 1, characterized in that, The steps of analyzing the poster image using a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and its matching relationship with the movie title include: The poster image is analyzed using a multimodal large model to determine the position, color, and font style information of the logo, and outputs the location area of the logo, the main color of the logo, the font style and decorative element features, and the matching relationship between the logo and the movie title.
3. The method for automatic poster logo extraction based on a multimodal model according to claim 2, characterized in that, The steps of automatically generating prompts suitable for image editing models and containing prompts for logo extraction include: Based on the determined location of the logo, its main color, font style, decorative elements, and the matching relationship between the logo and the movie title, combined with the poster image, a prompt word suitable for the image editing model is automatically generated, which includes prompts for extracting the logo. The prompt word content includes the area of the logo to be retained, the elements to be removed, and color and style information.
4. The method for automatic poster logo extraction based on a multimodal model according to claim 1, characterized in that, The steps of extracting the logo based on the prompt words and poster image, calling a preset image editing model, removing irrelevant elements, and obtaining the preliminary extracted logo include: Input the poster image and the generated prompts into a preset image editing model; The image editing model performs targeted image cutout based on the prompt words, removes irrelevant elements, retains only the Logo area, and outputs the preliminary extracted Logo.
5. The method for automatic poster logo extraction based on a multimodal model according to claim 1, characterized in that, The process of using a multimodal quality inspection mechanism to perform multi-dimensional checks on the initially extracted Logo, including text content, style and handwriting, and element cleanliness, and then outputting the Logo when the quality inspection is passed, includes the following steps: The initially extracted Logo is input into a multimodal model for quality inspection, which includes: Perform a text consistency check to determine whether the extracted logo matches the movie title; Perform style and handwriting detection to check whether the extracted logo retains the original font style and stroke characteristics; In addition, an element cleanliness test is conducted to confirm the presence of any figures, decorations, or superfluous elements; If the quality inspection fails, the prompt words are updated and the image editing model is called again for iterative extraction; Once the quality inspection is passed, a standardized transparent logo will be output.
6. The method for automatic poster logo extraction based on a multimodal model according to claim 1, characterized in that, The step of analyzing the poster image using a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and its matching relationship with the movie title, and automatically generating prompts suitable for image editing models to extract logo information, further includes: The poster image is analyzed using a multimodal large model. When multiple languages are identified in the poster, a preliminary style analysis is performed on the logo area for each language. The generated prompt will contain the expected description of the logo in the target language; if the target is known to be to extract the English logo, the prompt will contain the logo with the English title preserved.
7. The method for automatic poster logo extraction based on a multimodal model according to claim 6, characterized in that, The process of using a multimodal quality inspection mechanism to perform multi-dimensional checks on the initially extracted Logo, including text content, style and handwriting, and element cleanliness, and then outputting the Logo when the quality inspection is passed, includes the following steps: The system receives the initially extracted Logo, the original film title, and the target language film title through a cross-language semantic and style mapping module. A cross-lingual text-image embedding model is pre-set to understand the semantic associations of the same concepts in different languages and to map text descriptions to a visual style feature space. The original film title and the target language film title are input into a text encoder to obtain the semantic vectors of the original film title and the target language film title, respectively; at the same time, the initially extracted logo is passed through an image encoder to obtain the visual style vector of the initially extracted logo. Compare the semantic vector of the original movie title with the semantic vector of the target language movie title to determine whether they represent the same movie. During style mapping, the degree of matching between the visual style vector of the initially extracted logo and the semantic vector of the target language film name in the style feature space is compared. When performing style and handwriting detection, it is determined whether the style of the initially extracted logo conforms to the reasonable evolution of the original logo in the target language and cultural context as determined by the cross-language semantic and style mapping module. If the initially extracted logo is semantically consistent with the target language film title, and its stylistic changes fall within the reasonable localization range as determined by the cross-language semantic and style mapping module, then the quality inspection passes. Otherwise, update the prompt words and iterate again to extract them; When the quality inspection is passed, a standardized transparent logo will be output.
8. A poster logo automatic extraction device based on a multimodal model, characterized in that, The device includes: The acquisition module is used to acquire movie and TV poster images and movie titles; The positioning and prompt word generation module is used to analyze the poster image through a multimodal large model to determine the location of the logo, the key feature attributes of the logo, and the matching relationship with the movie title; and automatically generate prompt words suitable for the image editing model to extract the logo prompt content. The multimodal model extraction module is used to extract the logo based on the prompt words and the poster image by calling a preset image editing model, removing irrelevant elements, and obtaining the preliminary extracted logo. The multimodal quality inspection and output module is used to perform multi-dimensional checks on the initially extracted logo through a multimodal quality inspection mechanism, including text content, style and handwriting, and element cleanliness. When the quality inspection is qualified, a standardized transparent logo is output.
9. A smart terminal, characterized in that, It includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors, wherein the one or more programs include methods for performing any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1-7.