Automatic search keyword generation method and device based on large-scale model content recognition and terminal
Through a large-scale deep learning model combining natural language processing and semantic understanding, search keywords are automatically generated, which solves the problem of insufficient understanding of image content in the existing technology, and realizes efficient and accurate information recognition and keyword extraction.
Patent Information
- Application Number
- CN202510404224.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-08-12
AI Technical Summary
When processing complex and variable images, it is difficult to accurately identify key information, resulting in inaccurate extraction of search keywords and poor adaptability.
A large-scale deep learning model is adopted to automatically generate search keywords through character, scene and text recognition technology, combined with natural language processing and semantic understanding models.
It realizes efficient and accurate identification of characters, scenes and text information in the image, generates high-quality search keywords, and improves the efficiency and accuracy of image search and content recommendation.
Smart Images

Figure CN120470141A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and natural language processing, and in particular to a method, device, intelligent terminal and storage medium for automatically generating search keywords based on large-scale model content recognition. Background Art
[0002] With the explosive growth of digital image data, quickly and accurately extracting valuable information from massive amounts of images has become a pressing challenge. Traditional methods often rely on handcrafted features and template matching, but their ability to understand complex and varied image content is limited. For example, when processing images with complex backgrounds, multiple objects, or people interacting, traditional methods often struggle to accurately identify all key information.
[0003] In this way, due to insufficient understanding of the image content, the search keywords generated by existing technologies often deviate from the image content; this deviation may come from misunderstanding or omission of the characters, scenes or text information in the image, resulting in inaccurate search keyword extraction.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that, in response to the problems and defects of the above-mentioned prior art, a method, device, intelligent terminal and storage medium for automatically generating search keywords based on large-scale model content recognition are provided. The present invention provides a system and method that utilizes a large-scale deep learning model to realize efficient and accurate recognition of characters, scene information and text in pictures, and automatically generate high-quality search keywords, so as to solve the problems of insufficient understanding of picture content and inaccurate extraction of search keywords in the prior art.
[0006] The technical solutions adopted by the present invention to solve the problem are as follows: A method for automatically generating search keywords based on large-scale model content recognition, comprising: Receive and obtain pictures or video frames uploaded by users; Performing person recognition on the image or video frame to identify the person in the image or video frame and extracting the person's facial features and / or identity information; Performing scene recognition on the picture or video frame to identify the scene type in the picture or video frame; Performing text recognition on the picture or video frame to identify text information in the picture or video frame; Based on the facial features and / or identity information, scene type, and text information of the people in the identified pictures or video frames, combined with natural language processing technology and semantic understanding models, search keywords reflecting the content of the pictures or video frames are automatically generated.
[0007] The method for automatically generating search keywords based on large-scale model content recognition, wherein the scene types include outdoor, natural scenery, and urban landscape.
[0008] In the method for automatically generating search keywords based on large-scale model content recognition, the step of receiving and obtaining pictures or video frames uploaded by users includes: Receive and obtain pictures or video frames uploaded by users; Preprocess the image or video frame.
[0009] In the method for automatically generating search keywords based on large-scale model content recognition, the steps of performing person recognition on the picture or video frame, identifying the person in the picture or video frame, and extracting the person's facial features and / or identity information include: Performing person recognition on the picture or video frame, identifying the person in the picture or video frame, and extracting the facial features of the person; When the extracted facial features of the person are identified as a public figure or a historical figure, the name of the public figure or historical figure is identified; When the extracted facial features of a person are identified as an ordinary person or a person not well known to the public, the identity details are automatically ignored in the recognition process.
[0010] In the method for automatically generating search keywords based on large-scale model content recognition, the step of performing scene recognition on the picture or video frame and identifying the scene type in the picture or video frame includes: Collect and annotate datasets containing different scene types; Automatically extract image features using convolutional neural network deep learning models; Using the labeled data set to train the convolutional neural network deep learning model to obtain a trained convolutional neural network deep learning model; A trained convolutional neural network deep learning model is used to perform scene recognition on the image or video frame, and the image or video frame is framed in a specified scene type. In the method for automatically generating search keywords based on large-scale model content recognition, the step of performing text recognition on the picture or video frame to identify text information in the picture or video frame includes: Performing text recognition on the picture or video frame to extract text information from the picture or video frame; The extracted text information is deeply analyzed to extract key information including character names and film and television work titles.
[0011] The method for automatically generating search keywords based on large-scale model content recognition, wherein the step of automatically generating search keywords reflecting the content of the image or video frame based on the facial features and / or identity information, scene type, and text information of the person in the image or video frame recognized, combined with natural language processing technology and semantic understanding models, includes: Based on natural language processing technology and semantic understanding models, keyword filtering is performed on the facial features and / or identity information of the person in the identified image or video frame to obtain filtered facial features and / or identity information keywords; Based on natural language processing technology and semantic understanding models, keyword filtering is performed on the scene types in the identified pictures or video frames to obtain filtered scene type keywords; Based on natural language processing technology and semantic understanding models, keyword filtering is performed on the text information in the identified pictures or video frames to obtain filtered text information keywords; Based on the obtained facial features and / or identity information keywords, scene type keywords, and text information keywords, search keywords reflecting the content of the picture or video frame are automatically generated.
[0012] A device for automatically generating search keywords based on large-scale model content recognition, wherein the device comprises: The acquisition module is used to receive and acquire pictures or video frames uploaded by users; A person recognition module is used to perform person recognition on the picture or video frame, identify the person in the picture or video frame, and extract the facial features and / or identity information of the person; A scene recognition module, configured to perform scene recognition on the image or video frame and identify the scene type in the image or video frame; A text recognition module, configured to perform text recognition on the image or video frame, and identify text information in the image or video frame; The search keyword generation module is used to automatically generate search keywords reflecting the content of the picture or video frame based on the facial features and / or identity information of the characters, scene type, and text information in the identified picture or video frame, combined with natural language processing technology and semantic understanding model.
[0013] An intelligent terminal includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, including the method for executing any one of the methods described above.
[0014] A computer-readable storage medium, wherein when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any one of the methods described above.
[0015] Beneficial effects of the present invention: The present invention provides a method, device, intelligent terminal, and storage medium for automatically generating search keywords based on large-scale model content recognition. The present invention also provides a system and method for utilizing a large-scale deep learning model and integrating the large-scale deep learning model to achieve efficient and accurate recognition of people, scene information, and text in images, and automatically generate search keywords based on the recognition. The present invention can achieve a system and method for efficiently and accurately identifying people, scene information, and text in images, and automatically generate high-quality search keywords, which can significantly improve the efficiency and accuracy of image search, content recommendation, and intelligent analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 This is a flow chart of the method for automatically generating search keywords based on large-scale model content recognition provided in Example 1 of the present invention.
[0018] Figure 2 This is a flow chart of the method for automatically generating search keywords based on large-scale model content recognition provided in Example 2 of the present invention.
[0019] Figure 3 A principle block diagram of an embodiment of a device for automatically generating search keywords based on large-scale model content recognition provided by the present invention.
[0020] Figure 4 This is a block diagram of the internal structure of the smart terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0022] With the explosive growth of digital image data, how to quickly and accurately extract valuable information from massive images has become an urgent problem to be solved. Existing technologies for extracting valuable information from massive images have the following shortcomings: 1) Limited understanding of image content: Traditional methods rely primarily on hand-crafted features and template matching, which have limited understanding of complex and changing image content. For example, when processing images with complex backgrounds, multiple objects, or people interacting, traditional methods often struggle to accurately identify all key information.
[0023] 2) Inaccurate search keyword extraction: Due to insufficient understanding of image content, search keywords generated by existing technologies often deviate from the image content. This deviation may be due to misunderstanding or omission of the characters, scenes, or text information in the image.
[0024] 3) Poor adaptability: Existing technologies are often optimized for specific image types or application scenarios, making them difficult to adapt to a wide range of image types and complexities. For example, some technologies may perform well on natural landscape images but struggle with cityscapes or portraits.
[0025] The rapid development of deep learning technology in recent years, particularly the remarkable achievements of large-scale pre-trained models (such as BERT and the GPT series) in image recognition and text understanding, has opened up new possibilities for in-depth analysis of image content. However, how to effectively integrate these technologies to achieve comprehensive recognition of people, scenes, and text information in images, and automatically generate keywords that match user search intent, remains a hot topic and a challenge in current research.
[0026] To address the technical problems of the above-mentioned prior art, the present invention provides a system and method that utilizes and integrates large-scale deep learning models to achieve efficient and accurate recognition of people, scene information, and text in images, and automatically generates search keywords based on this recognition. The present invention can achieve a system and method that can achieve efficient and accurate recognition of people, scene information, and text in images, and automatically generate high-quality search keywords, which can significantly improve the efficiency and accuracy of image search, content recommendation, and intelligent analysis.
[0027] like Figure 1 As shown, a method for automatically generating search keywords based on large-scale model content recognition in embodiment 1 of the present invention includes the following steps: Step S100: receiving and obtaining pictures or video frames uploaded by users; In the embodiment of the present invention, in order to quickly and accurately extract valuable information from a large number of pictures, a user can select pictures or video frames from which valuable information needs to be extracted.
[0028] In specific implementations, images or video frames uploaded by users can be received and preprocessed. For example, these images can be denoised to eliminate interference such as sensor noise and compression artifacts. Gaussian filtering or median filtering (suitable for salt and pepper noise) can be used. This preprocessing improves subsequent recognition accuracy.
[0029] Specifically, in this embodiment of the present invention, before performing the subsequent step of person recognition, the image can be subjected to image denoising to ensure the accuracy of subsequent analysis. Denoising can eliminate artifacts caused by sensor noise or compression, improve image clarity, and ensure the recognizability of facial information.
[0030] Gaussian filtering: Smoothes the image using a Gaussian function, suitable for general high-frequency noise removal.
[0031] Median filter: effectively processes salt and pepper noise and preserves edge information.
[0032] Non-local mean denoising (NL-Means): denoising is achieved by searching for similar blocks across the entire image, preserving more detailed information.
[0033] Deep learning denoising model (such as DnCNN): A denoising model trained using a deep learning network can more effectively restore image details by learning from a large number of high-noise images.
[0034] Based on the above analysis, by performing person recognition on images or video frames, more intelligent applications can be realized in the real world, and the accuracy and efficiency of recognition can be improved through denoising technology.
[0035] Step S200: performing person recognition on the picture or video frame, identifying the person in the picture or video frame, and extracting the person's facial features and / or identity information; This step is to accurately extract the presence of people from images or videos and obtain their facial feature information (such as facial contour, eyes, mouth, etc.) and identity information (such as name, age, etc.).
[0036] Specifically, for person recognition, algorithms can analyze various parts of an image to locate and identify human faces. For example, deep learning techniques such as convolutional neural networks (CNNs) can be used to detect facial regions in images.
[0037] Regarding facial feature extraction, the present invention can further extract facial features after recognizing a face. These features may include facial geometry, relative positions of each facial part, and texture features.
[0038] Regarding identity information extraction, the present invention can determine the identity information of the face by comparing it with facial features in an existing database, thereby achieving identity verification or recognition. This step can be achieved by combining a facial image database with a machine learning algorithm to determine the user's identity and user name.
[0039] In a further embodiment, step S200 specifically includes: S201, performing person recognition on the picture or video frame, identifying the person in the picture or video frame, and extracting the facial features of the person; In this step, the image or video frame is subjected to character recognition, facial feature information (such as facial contour, eyes, mouth, etc.) in the image or video frame is identified, and facial features of the character are extracted.
[0040] S202: When the extracted facial features of the person are identified as a public figure or a historical figure, the name of the public figure or historical figure is identified; In this embodiment of the present invention, after facial features are extracted, they are compared with features of public figures (such as actors, politicians, and athletes) or historical figures (such as scientists and leaders) stored in a database. This process utilizes a person matching algorithm to quickly identify public or historical figures that match the extracted features by calculating the similarity of facial features.
[0041] Regarding obtaining names and related information: Once a matching public or historical figure is identified, this embodiment of the present invention returns the person's name. Furthermore, the present invention can also extract other information related to the person, such as their life story, achievements, and social influence. This facilitates determining the search keyword location in subsequent steps.
[0042] For example, in a social media app, when a user uploads a photo and the app identifies the face in the photo as a famous actor, the user can then see information about the actor, including his name, latest movies, news reports, etc. This immediate feedback enhances the user experience and encourages users to share and discuss.
[0043] In a history learning application, when students take a photo from a history book, the method of the present invention can instantly identify Li Bai and then provide information such as his name, life, major works, and influence. This makes students' learning more vivid and interesting, helping to improve learning efficiency.
[0044] When a visitor takes a photo of a portrait of a historical figure in a museum during an exhibition, the method of the present invention can identify that the portrait is a celebrity and automatically display his name, biography and important achievements, thereby enhancing the visitor's learning experience and interactive fun.
[0045] This step in the embodiment of the present invention demonstrates the enormous potential of facial recognition technology in promoting the dissemination of knowledge and information acquisition. It not only makes individual information processing more intelligent, but also brings greater convenience to users' learning and daily lives. Of course, the specific implementation of this technology application also requires attention to privacy protection and data security to ensure that facial recognition and recognition results are used in a compliant manner.
[0046] In the embodiments of the present invention, by identifying the public and historical figures, ordinary users can more conveniently acquire relevant knowledge, and identifying historical figures helps to inherit and carry forward culture. When learning history, users can more easily connect different people and events.
[0047] S203: When the extracted facial features of the person are identified as an ordinary person or a person unknown to the public, the identity details are automatically ignored during the identification process.
[0048] In the embodiment of the present invention, when the extracted facial features of a person are identified as an ordinary person or a person unknown to the public, the identity details are automatically ignored during the recognition process, that is, the name of the ordinary person or the person unknown to the public is no longer analyzed.
[0049] Step S300: performing scene recognition on the picture or video frame to identify the scene type in the picture or video frame; Scene recognition refers to the use of computer vision technology to perform semantic analysis on the content of input images or video frames, and automatically determine the scene category to which they belong (such as "city streets", "forests", "indoor meeting rooms", etc.); of course, the scene types can also include outdoor, natural scenery, and urban landscapes.
[0050] In specific implementations, deep learning models (such as ResNet and Vision Transformer) can be used to extract multi-level image features from the image or video frame, including color, texture, object distribution, and spatial relationships. These features are then input into a classifier (such as a fully connected layer + Softmax) and mapped to a set of preset scene labels. For example, a convolutional neural network (CNN) can be used to capture key scene features such as building outlines and vegetation density. The model outputs a probability distribution, determining that the image has an 85% probability of being a "beach" scene and a 12% probability of being a "mountainous" scene.
[0051] Of course, the present invention can also be dynamically optimized, and can combine time sequence information (for video frames) to improve recognition robustness through optical flow analysis or inter-frame feature fusion. The present invention can narrow the recognition range through scene prior knowledge.
[0052] In a further embodiment, step S300 specifically includes the following steps: S301. Collect and label data sets containing different scene types; In an embodiment of the present invention, a neural network model can be used to achieve scene recognition. Specifically, a dataset containing different scene types can be collected and annotated. For example, a hierarchical labeling system can be defined based on the target application scenario. Raw images can then be acquired through multi-source acquisition, and long-tail scene samples can be expanded using a synthetic data engine. A semi-automatic annotation process is then used, combining AI pre-annotation with manual verification to generate annotation results, thereby constructing a dataset containing different scene types.
[0053] For example, collect and label data sets containing different scene types, such as urban, rural, beach, mountainous, etc. These labeled data will be used to train the model.
[0054] S302, automatically extracting image features using a convolutional neural network deep learning model; In this embodiment, after the dataset is labeled, deep learning models such as convolutional neural networks (CNNs) are used to automatically extract image features. CNNs effectively capture important information in images, such as shape, color, and texture, thereby generating high-dimensional feature vectors. These vectors serve as input for subsequent tasks, such as classification, retrieval, or recognition.
[0055] S303: Using the labeled data set to train the convolutional neural network deep learning model to obtain a trained convolutional neural network deep learning model; In this embodiment, a labeled dataset may be used to train a model, wherein the model includes ResNet, VGG, Inception, etc. During the training process, the model will learn the characteristics of different scene types.
[0056] S304: Use the trained convolutional neural network deep learning model to perform scene recognition on the picture or video frame, and frame the picture or video frame into a specified scene type. That is, in this embodiment of the present invention, a trained model is applied to classify a new image or video frame. The model will frame the image into a specific scene type based on the previously learned features.
[0057] Step S400: performing text recognition on the picture or video frame to identify text information in the picture or video frame; This step automatically extracts text from images or videos to facilitate subsequent data processing and information analysis. Regarding the text recognition process in this step: Before text recognition, images typically require preprocessing to improve text clarity and readability. This includes noise removal, contrast adjustment, grayscaling, and image rotation.
[0058] In the embodiment of the present invention, when performing character detection, computer vision technology can be used to first determine the text area in the image and identify which parts contain text. Image processing methods such as edge detection and region growing can be used.
[0059] Then, during character recognition, after detecting the text area, an OCR algorithm can be used to identify the characters within it. This is usually accomplished using a trained deep learning model (such as a convolutional neural network), which can match the characters in the image with a predefined character set and output the corresponding text information.
[0060] Finally, after the above steps, the system will output the recognized text information. This information can be directly used for archiving, indexing, searching or other subsequent processing to provide convenience for users.
[0061] The automated text recognition process of the present invention significantly improves processing speed, saving time spent on manual text entry and proofreading, especially when processing large amounts of documents and data. Furthermore, text recognition makes it possible to extract key information from images and videos, facilitating analysis and utilization, such as extracting invoice information, book content, and advertising information.
[0062] Furthermore, the step S400 specifically includes: S401, performing text recognition on the picture or video frame to extract text information in the picture or video frame; In the embodiment of the present invention, text recognition is performed on the picture or video frame to extract text information from the picture or video frame.
[0063] S402: Perform in-depth analysis on the extracted text information and extract key information including character names and film and television work titles.
[0064] In this step, after the text is extracted, it can be analyzed using natural language processing algorithms, including lexical analysis, syntactic analysis, and semantic analysis. This process breaks the text down into basic language elements, such as words and phrases, for further processing.
[0065] By analyzing the structure and content of text, the NLP model can identify key entities in the text, such as names of people, works, time, place, etc. The model is based on trained named entity recognition (NER) technology to accomplish this task.
[0066] Then, key information extraction is performed. The present invention extracts specific key information from the analysis results, such as the names of people (such as actors and directors) and film and television works (such as the titles of movies or TV series). This extracted information can be stored, categorized, or directly used in subsequent applications.
[0067] In the embodiment of the present invention, text can be converted into structured data to facilitate subsequent analysis, retrieval and storage. Structured data can be quickly queried and analyzed, improving the efficiency of data use.
[0068] By extracting key information, such as names of characters and titles of film and television works, it is convenient to quickly generate search keywords in the subsequent steps.
[0069] Step S500: Based on the facial features and / or identity information, scene type, and text information of the characters in the identified pictures or video frames, combined with natural language processing technology and semantic understanding models, automatically generate search keywords reflecting the content of the pictures or video frames.
[0070] In the embodiment of the present invention, through this step, the most representative keywords can be extracted from multiple information, thereby improving the search and indexing efficiency of the content.
[0071] Specifically, the present invention first comprehensively analyzes different information sources in the identified image or video frame, including: Facial features and identity information: Recognized faces can be classified as well-known people (public figures), ordinary users or characters, etc.
[0072] Scene type: By analyzing the environment of the image or video, such as indoor, outdoor, natural scenery, urban landscape, etc.
[0073] Text information: extracted text content, including names of public figures, slogans, titles, dialogues, etc.
[0074] The present invention then uses natural language processing technology to perform semantic analysis and processing on the above information. For example, named entity recognition (NER) technology is applied to identify specific nouns related to the characters, and the intent and meaning of the text information is understood by using the context.
[0075] The present invention combines the various information provided to automatically generate search keywords that accurately reflect the content of the picture or video. These keywords can be specific names of people, works of art, or general descriptions related to the scene.
[0076] Furthermore, the present invention can continuously optimize the accuracy of keyword generation through the training of machine learning models to better meet user search needs and content characteristics.
[0077] For example, if a user uploads a photo of a group photo with friends on a summer beach, the present invention can identify the friends, the beach scene, and the text (e.g., "Beach Day Fun") in the photo. It can then automatically generate keywords such as "beach," "summer," "friends," "group photo," and "travel." These keywords will help other users more easily find similar content when searching.
[0078] For example, when a user uploads an image for search, the present invention identifies the buildings, skylines, and text content in the image, and then extracts keywords such as "urban architecture," "modern," and "skyline." This greatly improves the efficiency of finding relevant images and supports diverse visual searches.
[0079] The benefit of this invention is that automatically generated search keywords allow users to find the information they need more quickly when searching for content, reducing the burden and time cost of manual input. Generating search keywords also helps the platform better organize and categorize uploaded images and videos, simplifying the content management process.
[0080] In a further embodiment, step S500 specifically includes: S501, based on natural language processing technology and semantic understanding model, filtering the facial features and / or identity information of the person in the identified picture or video frame by keywords to obtain filtered facial features and / or identity information keywords; In an embodiment of the present invention, natural language processing technology is used to associate text information (such as labels and descriptions) in images or videos with facial feature data through algorithms and models to facilitate understanding and classification.
[0081] The semantic understanding model used in this invention can analyze and understand the specific meaning of facial features and convert them into structured information, such as keywords. These keywords may include facial features (such as eye color, hairstyle, expression, etc.) and identity information (such as name, occupation, etc.).
[0082] Keyword filtering is then performed, specifically removing redundant or irrelevant information and retaining only the important keywords related to the person's facial features and identity. This can be achieved by setting filtering rules and conditions to ensure that the output information is valid and relevant.
[0083] This method improves recognition accuracy: through natural language processing and semantic understanding, facial features can be more accurately identified and classified, reducing the probability of misidentification and thus improving overall system performance. Automated keyword filtering also reduces the need for manual intervention and improves work efficiency, especially when processing real-time video streams, enabling rapid acquisition of required information.
[0084] S502, based on natural language processing technology and semantic understanding model, performing keyword filtering on the scene types in the identified pictures or video frames to obtain filtered scene type keywords; In this embodiment of the present invention, based on natural language processing technology and a semantic understanding model, keyword filtering is performed on the scene types in the identified images or video frames, extracting relevant keywords from the identified scenes while removing redundant or irrelevant information. This can be achieved by setting specific thresholds or matching rules to ensure that the obtained keywords truly reflect the characteristics of the current scene.
[0085] In this way, the present invention can effectively improve the accuracy of scene recognition through keyword filtering, ensure that the extracted scene information is highly consistent with the actual situation, and reduce the possibility of misrecognition.
[0086] S503, based on natural language processing technology and semantic understanding model, performing keyword filtering on the text information in the identified image or video frame to obtain filtered text information keywords; In this embodiment, the extracted text information is processed using natural language processing technology to analyze and understand the text, including operations such as word segmentation, part-of-speech tagging, and syntactic analysis to obtain deeper semantic information.
[0087] The semantic understanding model is used to analyze the meaning of text information and identify keywords. The present invention uses the semantic understanding model to understand the context of the text and find the most relevant keywords for a specific task based on semantic similarity and importance scores.
[0088] Then, based on the set standards and rules, important keywords are extracted from the recognized text information, while invalid or redundant information is removed. This ensures that the obtained keywords accurately reflect the main information and theme of the text.
[0089] In the embodiment of this step, through natural language processing and semantic understanding, important keywords can be identified and filtered more accurately, thereby improving the relevance of text information.
[0090] S504: Automatically generate search keywords reflecting the content of the picture or video frame based on the obtained facial features and / or identity information keywords, scene type keywords, and text information keywords.
[0091] In this embodiment, relevant information is integrated from the three different keywords previously extracted. For example, facial features may include details such as "male" and "black eyes"; scene type keywords may include "city street" and "indoor restaurant"; and text information keywords may include descriptions such as "birthday party" and "concert."
[0092] Then, through semantic analysis, we understand the relationships between these keywords. For example, "city streets" may be associated with "men" and "concerts," thus forming a comprehensive description of the characters' actions, background, and events in a specific scene.
[0093] Then, keyword generation is performed: Finally, the present invention automatically generates one or more comprehensive search keywords based on the integrated keywords. These search keywords can not only accurately reflect the content of the image or video frame, but also have both richness and flexibility, suitable for search and classification.
[0094] For example, consider a video frame containing a young woman celebrating her birthday on a city street, surrounded by balloons and friends. After analysis, the following keywords were extracted: Character's facial features and identity information: female, young; Scene type keywords: city streets, party; Text message keywords: birthday, celebration; Based on the above keywords, the present invention can automatically generate search keywords such as: "young women, birthday party, city streets" or "women celebrating birthdays on city streets", so that users can quickly find relevant content when searching for similar videos or photos.
[0095] The search keywords generated by the embodiments of the present invention are highly relevant, enabling users to quickly locate similar content when searching, thereby improving information retrieval efficiency. Furthermore, the automatically generated keywords can help the system understand the context and content of an image or video, thereby better supporting subsequent intelligent analysis and classification. Furthermore, by integrating information from multiple sources, keyword generation can cover different information dimensions and provide a more comprehensive description of the content. Furthermore, using the generated keywords, a personalized recommendation system can be established to help users discover relevant content and topics, thereby enhancing the user experience.
[0096] Moreover, the generated keywords can automatically tag images or videos, making management and classification easier and reducing the workload of manual labeling.
[0097] In further embodiments of the present invention, the ability to analyze a person's facial expressions and gestures can be added to images or video frames. For example, by combining sentiment analysis models, this can add an emotional dimension to generated keywords, such as "happy young woman celebrating her birthday." Such keywords not only reflect the visual content but also capture the emotional atmosphere of the moment.
[0098] In a further embodiment of the present invention, a multilingual search keyword function can be automatically generated for images or video frames based on the user's language preference to enhance the experience for international users. For example, English keywords such as "Birthday party in city street" can be generated for the same video to meet the needs of users of different languages.
[0099] In a further embodiment of the present invention, a user feedback mechanism can be introduced to allow users to manually add or modify generated keywords, thereby improving the accuracy of subsequent keyword generation and user satisfaction. In this way, the present invention can continuously self-optimize and learn based on user feedback.
[0100] In a further embodiment of the present invention, voice recognition technology can be integrated to allow users to upload pictures or videos through voice commands, and automatically generate descriptions and search keywords, further improving ease of use.
[0101] As can be seen above, the embodiments of the present invention automatically generate search keywords reflecting the content of images or video frames by integrating facial features, scene types, and textual keywords. This not only improves search efficiency and content understanding, but also creates rich possibilities for personalized recommendations and user experience. Further inventive features such as sentiment analysis, multilingual support, and user feedback mechanisms will further enhance the system's intelligence and practicality, further unlocking the value of visual content.
[0102] The present invention is further described in detail below through another specific application embodiment. Figure 2 As shown, a method for automatically generating search keywords based on large-scale model content recognition in this specific application embodiment includes the following steps: S10. Receive the image (or video frame) uploaded by the user; pre-process the image to improve subsequent recognition accuracy, and then enter S11, S13, and S15 for parallel processing; S11. Use a general large-scale image to perform person recognition, identify the person in the uploaded image, extract the facial features or identity information of the identified person (if available), and proceed to S12.
[0103] In this step embodiment, when faced with a picture containing people, it is necessary to identify the people therein, with particular attention paid to the recognition of the Chinese names of public figures and historical figures. For ordinary people or individuals who are not well known to the public, their identity details are automatically ignored during the recognition process.
[0104] The instruction template is as follows: ### instraction: {instruction} ### Input: {Example} ### Response .
[0105] S12. Obtain the person recognition result, such as Liu Mouhua, and proceed to S17.
[0106] S13. Use a general large-scale image recognition algorithm to identify the scene type (e.g., indoor, outdoor, natural scenery, urban landscape, etc.) in the image, and proceed to S14.
[0107] S14. Obtain scene recognition results: such as the movie "Mr. Hong", and proceed to S17.
[0108] In the embodiment of this step, when faced with a picture, it is necessary to confirm the source or briefly describe the scene to ensure that the information is accurate and valid.
[0109] The instruction template is as follows: ### instraction: {instruction} ### Input: {Example} ### Response .
[0110] S15. Use a general large-scale text recognition: recognize the text information in the image and proceed to step S16.
[0111] In this step, when processing images rich in text information, the core responsibility is to accurately extract the text and perform in-depth analysis to comprehensively and efficiently identify and parse the text information in the image, including organizing key information such as character names and film and television work titles through a list structure; The instruction template is as follows: ### instraction: {instruction} ### Input: {Example} ### Response .
[0112] S16. Extract text information from the image and obtain text recognition results, such as "Mr. Hong Moumou released the film Ning Mou Liu Mouhua nationwide on the first day of the Lunar New Year."
[0113] As deep learning technology continues to advance, new model architectures (such as more efficient convolutional neural networks and more advanced Transformer models) may emerge. The present invention can adopt these new models in a timely manner to improve the accuracy and efficiency of person recognition, scene recognition, and text recognition.
[0114] In the embodiments of the present invention, algorithm fusion or improvement can be performed to further enhance the overall performance of the system. For example, a face detection and recognition algorithm can be further integrated with a scene recognition algorithm to enhance the recognition capability of complex scenes.
[0115] S17. Based on the above recognition results, combined with natural language processing technology and semantic understanding model, keyword filtering is performed according to the set keyword filtering rules, and then the process proceeds to step S18; S18. Generate search keywords that accurately reflect the image content, such as keyword 1: Liu Mouhua; keyword 2: Mr. Hong; and keyword 3: Ning Mou. The present invention can then output an optimized keyword list for use in applications such as image search and content recommendation. The process ends when the recognition process is complete.
[0116] As can be seen from the above, the present invention utilizes a large-scale deep learning model to realize efficient and accurate recognition of people, scene information and text in pictures, and automatically generates a system and method for high-quality search keywords, so as to solve the problems of insufficient understanding of picture content and inaccurate search keyword extraction in the prior art.
[0117] In further embodiments of the present invention, new functional modules can be added to the system. For example, an object recognition module can be introduced to identify specific objects in an image and extract relevant information; or a sentiment analysis module can be introduced to analyze the emotional tone of an image and generate more personalized search keywords based on this information.
[0118] Exemplary devices like Figure 3 As shown, an embodiment of the present invention provides a device for automatically generating search keywords based on large-scale model content recognition, the device comprising: The acquisition module 310 is used to receive and acquire pictures or video frames uploaded by users; A person recognition module 320 is configured to perform person recognition on the image or video frame, identify the person in the image or video frame, and extract the person's facial features and / or identity information; A scene recognition module 330 is configured to perform scene recognition on the picture or video frame and identify the scene type in the picture or video frame; A text recognition module 340 is configured to perform text recognition on the image or video frame, and identify text information in the image or video frame; The search keyword generation module 350 is used to automatically generate search keywords reflecting the content of the picture or video frame based on the facial features and / or identity information of the characters, scene type, and text information in the identified picture or video frame, combined with natural language processing technology and semantic understanding model, as described above.
[0119] In a further embodiment of the present invention, the preprocessing function of the acquisition module can be optimized and more advanced image enhancement technology can be adopted to improve the accuracy of subsequent recognition; or the search keyword generation module can be optimized and more advanced natural language processing technology and semantic understanding models can be adopted to generate more accurate, concise and attractive search keywords.
[0120] In another embodiment, the present invention upgrades the system architecture to a distributed architecture, and by introducing technologies such as load balancing, distributed storage, and distributed computing, the scalability and stability of the system can be improved.
[0121] In another embodiment, the embodiment of the present invention also adopts the method of migrating some computing tasks to edge devices for execution to reduce data transmission delay and improve system response speed.
[0122] In another embodiment, the present invention can also implement personalized recommendations, specifically by combining the user's historical search behavior and interest preferences to provide users with more personalized content recommendations and search keyword suggestions. This can not only improve the user experience, but also increase user stickiness and activity.
[0123] In another embodiment, the present invention also provides a more user-friendly and intuitive interactive interface to facilitate users to upload images, view recognition results, and search for keywords. At the same time, it can provide a user feedback channel so that users can raise questions and suggestions in a timely manner, promoting the continuous optimization and improvement of the system.
[0124] Based on the above embodiment, the present invention also provides an intelligent terminal, whose principle block diagram can be shown as follows: Figure 4 As shown. The intelligent terminal includes a processor, a memory, a network interface, a display screen, and a database connected via a system bus. The processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for automatically generating search keywords based on large-scale model content recognition is implemented. The database of the intelligent terminal is used to store a program for automatically generating search keywords based on large-scale model content recognition.
[0125] Those skilled in the art will understand that Figure 4 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the solution of the present invention and does not constitute a limitation on the smart terminal to which the solution of the present invention is applied. The specific smart terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0126] In one embodiment, a smart terminal is provided, comprising a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: Receive and obtain pictures or video frames uploaded by users; Performing person recognition on the image or video frame to identify the person in the image or video frame and extracting the person's facial features and / or identity information; Performing scene recognition on the picture or video frame to identify the scene type in the picture or video frame; Performing text recognition on the picture or video frame to identify text information in the picture or video frame; Based on the facial features and / or identity information, scene type, and text information of the characters in the identified pictures or video frames, combined with natural language processing technology and semantic understanding models, search keywords reflecting the content of the pictures or video frames are automatically generated, as described above.
[0127] The scene types include outdoor, natural scenery, and urban landscape.
[0128] In the method for automatically generating search keywords based on large-scale model content recognition, the step of receiving and obtaining pictures or video frames uploaded by users includes: Receive and obtain pictures or video frames uploaded by users; Preprocess the image or video frame.
[0129] The step of performing person recognition on the picture or video frame, identifying the person in the picture or video frame, and extracting the facial features and / or identity information of the person includes: Performing person recognition on the picture or video frame, identifying the person in the picture or video frame, and extracting the facial features of the person; When the extracted facial features of the person are identified as a public figure or a historical figure, the name of the public figure or historical figure is identified; When the extracted facial features of a person are identified as an ordinary person or a person not well known to the public, the identity details are automatically ignored in the recognition process.
[0130] The step of performing scene recognition on the picture or video frame to identify the scene type in the picture or video frame includes: Collect and annotate datasets containing different scene types; Automatically extract image features using convolutional neural network deep learning models; Using the labeled data set to train the convolutional neural network deep learning model to obtain a trained convolutional neural network deep learning model; A trained convolutional neural network deep learning model is used to perform scene recognition on the image or video frame, and the image or video frame is framed in a specified scene type. The step of performing text recognition on the picture or video frame to recognize text information in the picture or video frame includes: Performing text recognition on the picture or video frame to extract text information from the picture or video frame; The extracted text information is deeply analyzed to extract key information including character names and film and television work titles.
[0131] The step of automatically generating search keywords reflecting the content of the picture or video frame based on the facial features and / or identity information of the person, scene type, and text information in the identified picture or video frame, combined with natural language processing technology and semantic understanding models, includes: Based on natural language processing technology and semantic understanding models, keyword filtering is performed on the facial features and / or identity information of the person in the identified image or video frame to obtain filtered facial features and / or identity information keywords; Based on natural language processing technology and semantic understanding models, keyword filtering is performed on the scene types in the identified pictures or video frames to obtain filtered scene type keywords; Based on natural language processing technology and semantic understanding models, keyword filtering is performed on the text information in the identified pictures or video frames to obtain filtered text information keywords; Based on the obtained facial features and / or identity information keywords, scene type keywords, and text information keywords, search keywords reflecting the content of the picture or video frame are automatically generated, as described above.
[0132] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).
[0133] In summary, the present invention provides a method, device, intelligent terminal, and storage medium for automatically generating search keywords based on large-scale model content recognition. The present invention also provides a system and method for utilizing a large-scale deep learning model to integrate large-scale deep learning models to achieve efficient and accurate recognition of people, scene information, and text in images, and automatically generate search keywords based on this. The present invention can achieve a system and method for efficiently and accurately identifying people, scene information, and text in images, and automatically generating high-quality search keywords, which can significantly improve the efficiency and accuracy of image search, content recommendation, and intelligent analysis.
Claims
1. A method for automatically generating search keywords based on large-scale model content recognition, characterized in that: include: Receive and obtain pictures or video frames uploaded by users; Performing person recognition on the image or video frame to identify the person in the image or video frame and extracting the person's facial features and / or identity information; Performing scene recognition on the picture or video frame to identify the scene type in the picture or video frame; Performing text recognition on the picture or video frame to identify text information in the picture or video frame; Based on the facial features and / or identity information, scene type, and text information of the people in the identified pictures or video frames, combined with natural language processing technology and semantic understanding models, search keywords reflecting the content of the pictures or video frames are automatically generated.
2. The method for automatically generating search keywords based on large-scale model content recognition according to claim 1, characterized in that: The scene types include outdoor, natural scenery, and urban landscape.
3. The method for automatically generating search keywords based on large-scale model content recognition according to claim 1, characterized in that: The step of receiving and obtaining the picture or video frame uploaded by the user includes: Receive and obtain pictures or video frames uploaded by users; Preprocess the image or video frame.
4. The method for automatically generating search keywords based on large-scale model content recognition according to claim 1, characterized in that: The step of performing person recognition on the picture or video frame, identifying the person in the picture or video frame, and extracting the facial features and / or identity information of the person includes: Performing person recognition on the picture or video frame, identifying the person in the picture or video frame, and extracting the facial features of the person; When the extracted facial features of the person are identified as a public figure or a historical figure, the name of the public figure or historical figure is identified; When the extracted facial features of a person are identified as an ordinary person or a person not well known to the public, the identity details are automatically ignored in the recognition process.
5. The method for automatically generating search keywords based on large-scale model content recognition according to claim 1, characterized in that: The step of performing scene recognition on the picture or video frame to identify the scene type in the picture or video frame includes: Collect and annotate datasets containing different scene types; Automatically extract image features using convolutional neural network deep learning models; Using the labeled data set to train the convolutional neural network deep learning model to obtain a trained convolutional neural network deep learning model; A trained convolutional neural network deep learning model is used to perform scene recognition on the image or video frame, and the image or video frame is framed in a specified scene type.
6. The method for automatically generating search keywords based on large-scale model content recognition according to claim 1, characterized in that: The step of performing text recognition on the picture or video frame to recognize text information in the picture or video frame includes: Performing text recognition on the picture or video frame to extract text information from the picture or video frame; The extracted text information is deeply analyzed to extract key information including character names and film and television work titles.
7. The method for automatically generating search keywords based on large-scale model content recognition according to claim 1, characterized in that: The step of automatically generating search keywords reflecting the content of the picture or video frame based on the facial features and / or identity information of the person, scene type, and text information in the identified picture or video frame in combination with natural language processing technology and semantic understanding model includes: Based on natural language processing technology and semantic understanding models, keyword filtering is performed on the facial features and / or identity information of the person in the identified image or video frame to obtain filtered facial features and / or identity information keywords; Based on natural language processing technology and semantic understanding models, keyword filtering is performed on the scene types in the identified pictures or video frames to obtain filtered scene type keywords; Based on natural language processing technology and semantic understanding models, keyword filtering is performed on the text information in the identified pictures or video frames to obtain filtered text information keywords; Based on the obtained facial features and / or identity information keywords, scene type keywords, and text information keywords, search keywords reflecting the content of the picture or video frame are automatically generated.
8. A device for automatically generating search keywords based on large-scale model content recognition, characterized in that: The device comprises: The acquisition module is used to receive and acquire pictures or video frames uploaded by users; A person recognition module is used to perform person recognition on the picture or video frame, identify the person in the picture or video frame, and extract the facial features and / or identity information of the person; A scene recognition module, configured to perform scene recognition on the image or video frame and identify the scene type in the image or video frame; A text recognition module, configured to perform text recognition on the image or video frame, and identify text information in the image or video frame; The search keyword generation module is used to automatically generate search keywords reflecting the content of the picture or video frame based on the facial features and / or identity information of the characters, scene type, and text information in the identified picture or video frame, combined with natural language processing technology and semantic understanding model.
9. An intelligent terminal, characterized in that: The device comprises a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, and the one or more programs include being used to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Image element extraction method and device fusing large language model and neural network, equipment and storage medium
CN121095958A