Picture intelligent labeling processing method and device, intelligent terminal and storage medium

By automatically generating image annotations through deep learning and natural language processing technologies, and combining user interaction and index optimization, the problem of low efficiency and poor accuracy of image annotation in existing technologies is solved, achieving efficient and accurate image annotation and management.

CN120877285APending Publication Date: 2025-10-31SHENZHEN COOCAA NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510975100.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In existing technologies, image annotation relies on manual operation, which is inefficient and prone to errors. Automated tools have low recognition accuracy in complex scenarios and struggle to understand the deep semantics of images and adapt to personalized needs.

Method used

Deep learning algorithms are used for multi-dimensional analysis, combined with natural language processing technology to automatically generate labeled text, and optimized labeling strategies are established through user interaction and indexing. Convolutional neural networks are used to identify elements, and an attention mechanism is introduced to improve labeling accuracy. The system also learns user preferences to adjust the labeling strategy.

Benefits of technology

It achieves efficient and accurate image annotation, meets the needs of diverse application scenarios, improves annotation efficiency and accuracy, supports personalized needs, and facilitates image management and retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877285A_ABST
    Figure CN120877285A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent picture annotation processing method and device, an intelligent terminal and a storage medium, and belongs to the technical field of the Internet, and the method comprises the steps: collecting and obtaining pictures from different channels, and carrying out the preprocessing of the pictures; performing multi-dimensional analysis on the preprocessed picture by adopting a deep learning algorithm, identifying various elements in the picture, and extracting key information corresponding to the various elements in the picture; on the basis of key information corresponding to various elements in the extracted pictures, a natural language processing technology is adopted, picture semantic understanding is combined, and a corresponding annotation text is automatically generated for each picture; and automatically classifying and storing each picture for generating the annotated text, and establishing an index. According to the method, annotations can be efficiently, accurately and intelligently added to the pictures, and the requirements of diversified application scenes are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet technology, and in particular to a method, apparatus, smart terminal, and storage medium for intelligent image annotation processing. Background Technology

[0002] In today's era of digital information explosion, images, as a crucial carrier of information, occupy a pivotal position in various fields. From sharing everyday moments on social media platforms to image data processing in professional fields, the number of images is growing exponentially. However, facing massive amounts of image data, how to quickly and accurately add annotations to make them easier to understand and retrieve has become a pressing problem to be solved.

[0003] Current traditional image annotation methods mainly rely on manual operation, which is not only inefficient, but also prone to inaccurate or inconsistent annotations due to human factors.

[0004] Therefore, existing technologies still need improvement and development. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides an intelligent image annotation processing method, apparatus, smart terminal, and storage medium. This invention can efficiently, accurately, and intelligently add annotations to images, meeting the needs of diverse application scenarios.

[0006] This application provides a method for intelligent image annotation processing, the technical solution of which is as follows: An intelligent image annotation processing method, comprising: Images from different sources are collected and preprocessed. Deep learning algorithms are used to perform multi-dimensional analysis on preprocessed images, identify various elements in the images, and extract key information corresponding to each element. Based on the key information corresponding to various elements in the extracted images, natural language processing technology is used, combined with image semantic understanding, to automatically generate corresponding annotation text for each image; For each image that generates labeled text, it is automatically categorized, stored, and indexed.

[0007] The image intelligent annotation processing method, wherein the steps of collecting and acquiring images from different channels and preprocessing the images include: The data collection includes images uploaded from local storage, images crawled from the web, and images captured by a camera. The acquired images are preprocessed to generate preprocessed images; the preprocessing operations include resizing the images, converting them to grayscale, and removing noise to optimize image quality.

[0008] The image intelligent annotation processing method, wherein the steps of employing a deep learning algorithm to perform multi-dimensional analysis on the preprocessed image, identify various elements in the image, and extract key information corresponding to various elements in the image include: A predetermined deep learning algorithm is used to perform multi-dimensional analysis on the preprocessed images; A convolutional neural network model is used to identify various elements in an image and extract key information corresponding to each element. These elements include object elements, scene elements, and person elements. The key information includes category information and location information.

[0009] The image intelligent annotation processing method, wherein the step of generating corresponding annotation text for each image based on the key information corresponding to various elements in the extracted image, using natural language processing technology combined with image semantic understanding, includes: Based on the key information corresponding to various elements in the extracted images, natural language processing technology is used, combined with image semantic understanding, to generate corresponding matching annotation text for each image; Furthermore, an attention mechanism is introduced to focus on key areas in the image, further refining the labeled text to improve labeling accuracy and intelligence.

[0010] The image intelligent annotation processing method further includes, before the step of automatically classifying, storing, and indexing each image for which annotation text is generated: The automatically generated annotation text for each image is displayed in the preset user interaction module. Receive user operation instructions to view and / or modify the automatically generated annotation text, and store the viewed and / or modified annotation text.

[0011] The image intelligent annotation processing method further includes the step of receiving user operation instructions to view and / or modify the automatically generated annotation text, and storing the viewed and / or modified annotation text: Monitor user actions to modify and edit automatically generated annotation text, learn user habits and preferences for annotation modification, and optimize subsequent annotation strategies.

[0012] The image intelligent annotation processing method, wherein the step of automatically classifying, storing, and indexing each image for which annotation text is generated includes: Receive image search commands with text input from the user; The text-based image search command is parsed, and the images stored in the category are searched using the established index to find and display the images that match the text-based image search command.

[0013] An intelligent image annotation processing device, wherein the device comprises: The image acquisition and preprocessing module is used to acquire images from different channels and preprocess the images. The intelligent analysis module uses deep learning algorithms to perform multi-dimensional analysis on preprocessed images, identify various elements in the images, and extract key information corresponding to each element. The annotation module is used to automatically generate corresponding annotation text for each image based on the key information corresponding to various elements in the extracted images, using natural language processing technology combined with image semantic understanding. The user interaction module is used to feed back the automatically generated annotation text of each image to the preset user interaction module for display; receive user operation instructions to view and / or modify the automatically generated annotation text, and store the viewed and / or modified annotation text; The image data storage module is used to automatically classify, store, and index each image used to generate labeled text. The annotation optimization module is used to monitor users' modifications and edits to the automatically generated annotation text, learn users' habits and preferences in annotation modification, and optimize subsequent annotation strategies.

[0014] A smart terminal includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors, the one or more programs comprising the method for performing any one of the methods.

[0015] A computer-readable storage medium, wherein, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the methods described above.

[0016] As can be seen from the above, the image intelligent annotation processing method, device, intelligent terminal and storage medium provided in this application can efficiently, accurately and intelligently add annotations to images, meet the needs of diverse application scenarios, meet the image annotation needs of different users in diverse scenarios, bring innovative breakthroughs to the management and utilization of image data, and provide convenience for users. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the intelligent image annotation processing method provided in Embodiment 1 of the present invention.

[0019] Figure 2 This is a flowchart illustrating the intelligent image annotation processing method provided in Embodiment 2 of the present invention.

[0020] Figure 3 A schematic diagram of an embodiment of the intelligent image annotation processing device provided by the present invention.

[0021] Figure 4 This is a block diagram illustrating the internal structure of a smart terminal provided in an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0023] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.

[0024] Current traditional image annotation methods mainly rely on manual operation, which is not only inefficient but also prone to inaccuracy or inconsistency due to human factors. While some existing automated annotation tools can perform simple recognition and annotation, their annotation results are often unsatisfactory in complex scenarios, suffering from low recognition accuracy, inability to understand the deep semantics of images, and difficulty in adapting to the personalized needs of different users. These problems severely restrict the in-depth application and development of image annotation technology in various fields.

[0025] To address the aforementioned issues, this application proposes an intelligent image annotation processing method. This method can efficiently, accurately, and intelligently add annotations to images, meeting the needs of diverse application scenarios and the image annotation requirements of different users in various scenarios. It brings a revolutionary breakthrough to the management and utilization of image data and provides convenience for users.

[0026] Example 1 like Figure 1 As shown, an intelligent image annotation processing method according to Embodiment 1 of the present invention includes the following steps: Step S100: Collect and acquire images from different channels, and preprocess the images; In this embodiment, images are collected from different channels to ensure the comprehensiveness and diversity of image data by gathering image data from diverse information sources. These different channels include, but are not limited to, the following types: Images directly obtained from local upload; Images obtained from online platforms: including images legally downloaded from various image sharing websites, social media platforms, and professional databases (such as medical imaging databases, geographic information libraries, etc.).

[0027] Images imported from storage media: Saved images read from storage media such as computer hard drives, portable storage devices (such as USB flash drives, portable hard drives), and optical discs; Images obtained through other legal means, such as images provided by partners or uploaded by users.

[0028] This invention preprocesses the acquired images. After image acquisition, a series of standardized processing operations are performed on the collected images to eliminate interference and standardize the data format, laying the foundation for subsequent image analysis, recognition, or applications. The preprocessing operations include, but are not limited to, the following: Size adjustment: Adjust the image to the preset size by scaling, cropping, etc., to reduce the amount of data processing and ensure data consistency; Noise removal: Filtering algorithms (such as mean filtering, median filtering, Gaussian filtering, etc.) are used to remove interference signals such as salt and pepper noise and Gaussian noise from the image; Contrast enhancement: By using methods such as histogram equalization and gamma correction, the brightness and contrast of the image are improved, thereby enhancing the visual quality and information recognition of the image. Redundant information removal: Remove areas in the image that are irrelevant to the target task (such as borders, watermarks, etc.) and focus on the effective information. This embodiment enables multi-channel image acquisition and preprocessing to optimize image data quality and reduce the difficulty of subsequent annotation. Original images may suffer from variations in size, noise interference, and information blurring, directly impacting the accuracy and efficiency of subsequent analysis. Preprocessing effectively addresses these issues, standardizing and clarifying the image data to provide high-quality input for subsequent element analysis and other operations, significantly improving recognition accuracy.

[0029] Step S200: Use deep learning algorithms to perform multi-dimensional analysis on the preprocessed image, identify various elements in the image, and extract key information corresponding to various elements in the image; In this embodiment, deep learning technology is used to perform detailed multi-dimensional analysis on the preprocessed image, thereby identifying various elements in the image and extracting the key feature information of these elements.

[0030] Specifically, regarding deep learning algorithms, this invention can employ models such as convolutional neural networks (CNNs), which can automatically learn complex features of images, thereby achieving accurate element detection, classification, and understanding.

[0031] Regarding multidimensional analysis, multidimensionality refers to analysis from different angles and levels, including: Object detection: Identifying specific objects (such as people, vehicles, animals, etc.) in an image and their locations (bounding boxes).

[0032] Semantic understanding: understanding the categories, attributes (color, size, state, etc.) of objects and the relationships between them.

[0033] Scene analysis: Identify the scene depicted in the image (such as beach, city, indoors, etc.).

[0034] These multi-dimensional analyses help the system obtain rich semantic information about the images.

[0035] Regarding element recognition and key information extraction, through a trained deep learning model, the system can identify various elements in an image (such as people, objects, backgrounds, scenes, etc.) and extract key information, such as: the category and location of objects; attributes such as color, shape, and size; and the relationships between elements (such as "a person is sitting in a chair"). This information provides the foundation for subsequent image content understanding and the generation of labeled text.

[0036] The deep learning model employed in this embodiment of the invention, after extensive training, achieves high-accuracy target detection and element recognition, demonstrating greater intelligence and reliability than traditional methods. Furthermore, the multi-dimensional analysis of this invention allows the system not only to identify individual elements but also to understand the relationships between them and scene attributes, enhancing the depth of understanding. Moreover, the extracted key elements and information can be used for various purposes such as generating richer descriptions, categorized storage, content retrieval, and intelligent recommendation. It is evident that this step, through deep learning technology, achieves a deep and multi-faceted understanding of image content, laying a solid foundation for intelligent image recognition, automatic annotation, and application development.

[0037] Step S300: Based on the key information corresponding to various elements in the extracted images, natural language processing technology is used in combination with image semantic understanding to automatically generate corresponding annotation text for each image; In this embodiment, based on image content analysis and combined with Natural Language Processing (NLP) technology, descriptive annotation text is automatically generated for each image. Specifically, it includes two main parts: first, extracting various elements and their key information from the image; and second, using NLP technology to transform this information into natural and fluent text descriptions.

[0038] In step S200 above, key information is extracted from the image using image recognition, object detection, and semantic segmentation techniques to automatically identify various elements in the image (such as people, objects, backgrounds, scenes, etc.) and their relationships and attributes (such as color, position, actions, etc.). These elements constitute the "semantic content" of the image.

[0039] Then, combined with image semantic understanding, that is, in addition to simply recognizing elements, it also performs scene understanding, judges the theme and intention expressed by the image, and understands the relationship between elements, such as "a person is running" or "a cat is on the sofa".

[0040] Then, the annotated text is automatically generated. This involves using natural language processing techniques to transform the extracted key information into natural, continuous text descriptions that conform to common expression habits. This is typically achieved using image-to-text generation models (such as image description generation models).

[0041] In this embodiment of the invention, the automatically generated text helps users quickly understand the content of the images, improves the efficiency of annotated text generation, and eliminates the tedious step of manually describing each image, greatly saving time and manpower costs, making it particularly suitable for processing large numbers of images.

[0042] This invention enhances multimodal understanding by combining image and text analysis, enabling support for more complex application scenarios, such as image indexing, content recommendation, and intelligent monitoring in search engines. Furthermore, the generated labeled text can serve as tags or keywords for images, facilitating subsequent retrieval and classification.

[0043] Step S400: Automatically classify and store each image for which annotation text is generated, and create an index.

[0044] In this embodiment, after automatically generating the annotation text for each image, each image needs to be categorized and stored, and an index needs to be created to achieve efficient management and retrieval. Specifically, this includes two main aspects: first, automatically classifying the images according to a certain classification system; and second, creating an index for the images to facilitate quick retrieval and access.

[0045] Regarding automatic classification and storage, this invention utilizes keywords or topic information from previously generated labeled text to categorize images into different categories. For example, images can be classified into categories such as "animals," "landscapes," "people," and "food" based on their content. This classification is typically achieved through natural language processing (such as keyword extraction and topic recognition) or machine learning models (such as classifiers).

[0046] Regarding indexing, this invention can establish an index system for each image and its associated tags (or categories), linking the image's storage location, related tags, and descriptive information. This creates a structured database or index library, allowing for quick location of the corresponding image by simply entering keywords or categories during subsequent searches.

[0047] The benefit of this step is that categorization and indexing allow users to quickly find target images using keywords, categories, or tags, greatly improving the efficiency of finding specific content, such as quickly locating "cat pictures" or "landscape pictures" in a massive image library. Furthermore, storing images by category helps the system and users manage large amounts of image resources more systematically, avoiding chaos and facilitating maintenance and expansion. Combined with index information, more intelligent content recommendations can be achieved, such as recommending images of similar categories to users or automatically generating similar content based on tags.

[0048] In this embodiment of the invention, the image intelligent annotation processing method further includes step S100, which comprises: S101. Acquire images including those uploaded from the local machine, images crawled from the web, and images captured by a camera; S102. Perform preprocessing operations on the acquired images to generate preprocessed images; wherein, the preprocessing operations include resizing the images, converting them to grayscale, and removing noise to optimize image quality.

[0049] In this specific embodiment, image acquisition is a crucial step in obtaining raw image data, mainly including: Local image upload: This refers to selecting and uploading existing image files from the user's local storage media such as computers or mobile devices. These images may be previously saved photographs, screenshots, scans, etc. The advantage of this method is that it can directly utilize existing image resources without the need for an additional real-time acquisition process.

[0050] Image web scraping: Using specific web crawling tools or programs, images are automatically retrieved from various websites on the internet (such as image libraries, social media platforms, news websites, etc.) according to predefined rules and keywords. It can quickly and massively collect image data related to a specific topic, making it particularly suitable for scenarios requiring a large sample size.

[0051] Camera-captured images: Using the device's built-in camera or an external camera, images are captured in real time. This method can capture dynamic, on-site image information and is suitable for scenarios requiring instant image data, such as real-time monitoring, facial recognition, and live streaming.

[0052] The present invention then preprocesses the acquired images to optimize image quality and lay a good foundation for subsequent image analysis, recognition, and processing operations. This preprocessing includes three steps: Resizing: Adjust the width and height of the image to a specific size according to actual needs. For example, some image recognition models usually require input images of a fixed size. Resizing can standardize images from different sources and of different sizes, facilitating subsequent processing and calculations.

[0053] Grayscale conversion: Converting a color image to a grayscale image, meaning that each pixel in the image contains only brightness information and no color information. Grayscale conversion can reduce the amount of image data, reduce computational complexity, and in many scenarios, important image features (such as contours and textures) can still be clearly represented in grayscale images.

[0054] Noise Removal Processing: During image acquisition and transmission, images may be subject to various interferences that generate noise (such as lighting interference during shooting, sensor noise, network transmission errors, etc.), leading to a decrease in image quality. Noise removal processing uses specific algorithms (such as mean filtering, median filtering, Gaussian filtering, etc.) to reduce or eliminate this noise, making images clearer and highlighting useful information.

[0055] In this embodiment of the invention, image data of different scenes, types, and times can be obtained through three methods: local upload, web crawling, and camera shooting. This provides sufficient and diverse original materials for subsequent image annotation work, ensuring the diversity and richness of data sources.

[0056] Furthermore, the resizing in the preprocessing operation ensures uniform image specifications, facilitating standardized processing; grayscale conversion simplifies image data and improves processing efficiency; and noise removal reduces interference, making the key features of the image more prominent, thereby improving the accuracy and reliability of subsequent image analysis, recognition, and other operations.

[0057] In this embodiment of the application, the image intelligent annotation processing method, wherein step S200 specifically includes: S201. Using a predetermined deep learning algorithm, perform multi-dimensional analysis on the preprocessed image; In this embodiment, the predetermined deep learning algorithm refers to the appropriate deep learning algorithm selected in advance based on the specific image annotation task requirements, image type, and other factors, such as recurrent neural networks and deep belief networks. Preprocessed images, which have undergone resizing, grayscale conversion, and noise removal, have more uniform quality and specifications, making them easier to analyze. Multi-dimensional analysis refers to dissecting an image from multiple perspectives, encompassing dimensions such as color distribution, texture features, spatial structure, and semantic relationships. For example, analyzing color saturation variations in different areas of an image, the texture coarseness of object surfaces, the spatial arrangement of elements within the image, and the relationships between elements allows for a comprehensive extraction of the information contained within the image, facilitating subsequent image annotation.

[0058] S202. A convolutional neural network model is used to identify various elements in the image and extract key information corresponding to each element. The various elements include: object elements, scene elements, and person elements; the key information includes category information and location information.

[0059] In this embodiment, a convolutional neural network model is used to automatically extract and combine local features of an image in image recognition in order to process complex visual information in the image. The identified elements include object elements (such as cars, trees, furniture, etc.), scene elements (such as streets, forests, living rooms, etc.), and people elements. Among the key information extracted, category information refers to determining the specific type of an element, such as clarifying that an object is a "car" rather than a "truck", or that a scene is an "office" rather than a "classroom"; location information is usually represented in the form of bounding box coordinates, etc., and is used to accurately locate the position of an element in the image. As can be seen, the multi-dimensional analysis using deep learning algorithms in this invention can extract image information from multiple angles, providing rich feature support for subsequent recognition; while convolutional neural network models are good at processing image features. The combination of the two can more accurately identify various elements such as objects, scenes, and people in the image, reduce missed or misidentified cases, and improve the accuracy and comprehensiveness of recognition. This step specifically extracts category and location information, which are the core content of image annotation. Accurate category information clarifies the attributes of elements, while precise location information determines the spatial distribution of elements, providing a reliable data foundation for subsequent image applications (such as object tracking and image retrieval) and enabling the accurate extraction of key information. Furthermore, by leveraging deep learning algorithms and convolutional neural network models, machine annotation can replace a significant amount of manual annotation work, rapidly processing massive amounts of images. Compared to manual annotation, machine annotation is not only faster but also avoids the efficiency decline caused by subjective errors and fatigue from manual operations, significantly improving the overall efficiency of image annotation and enhancing its automation level.

[0060] Furthermore, the image intelligent annotation processing method of the present invention includes the step of generating corresponding annotation text for each image based on the key information corresponding to various elements in the extracted image, using natural language processing technology combined with image semantic understanding: S301. Based on the key information corresponding to various elements in the extracted images, natural language processing technology is used, combined with image semantic understanding, to generate corresponding matching annotation text for each image. In this embodiment, key information includes the category and location information of various elements extracted in the previous step S202. For example, if key information such as "car (category information) on the left side of the street (location information)" and "pedestrian (category information) in the middle of the zebra crossing (location information)" is identified in the image, this information will become the basic material for generating labeled text. This invention employs natural language processing technology to transform the aforementioned structured key information into natural language text that conforms to human language habits. It involves semantic analysis and the application of grammatical rules to ensure that the generated text is fluent, accurate, and clearly expresses the content of the image. This invention employs image semantic understanding, focusing not only on the elements themselves but also analyzing the relationships between elements and the deeper meanings of the overall context presented by the image. For example, by combining semantic understanding, it can determine that "cars and pedestrians in a road scene may be in the state of driving and crossing the road," thus making the generated annotation text more logical and complete, rather than simply listing elements.

[0061] S302 introduces an attention mechanism to focus on key areas in the image and further correct the labeled text to improve the accuracy and intelligence of the labeling.

[0062] In this embodiment, an attention mechanism is used to simulate the characteristics of human visual attention, automatically focusing on more important and critical areas in the image. In image annotation, key areas may include the main object, people performing specific actions, or parts that embody the core features of the scene. For example, in an image containing a "car accident scene," the attention mechanism will focus more on key areas such as the colliding vehicles and injured individuals. In this embodiment of the invention, the initial annotation text generated in S301 is checked and adjusted based on the key area focused by the attention mechanism. This may involve supplementing detailed information about the key area; for example, the original annotation "There are vehicles on the road" may be corrected to "There are two vehicles that have collided on the road, and the front of the vehicles is severely damaged." It may also correct inaccuracies in the initial annotation regarding the key area, such as correcting "A person wearing red clothes is standing on the roadside" to "A person wearing red clothes is lying on the roadside, suspected of being injured." In this embodiment of the invention, by focusing on key areas and correcting the text, it avoids over-description of secondary information or omission and misjudgment of key information, making the labeled text more accurately reflect the core content of the image. At the same time, this method of dynamically adjusting labels based on importance reflects the intelligence of the labeling process, better aligns with human understanding and expression habits of image information, and improves the accuracy and intelligence of labeling.

[0063] Furthermore, in the aforementioned intelligent image annotation processing method, before the step of automatically classifying, storing, and indexing each image for which annotation text is generated, the method further includes: The system automatically generates annotation text for each image and displays it to a pre-defined user interaction module; it also receives user commands to view and / or edit the automatically generated annotation text and stores the viewed and / or edited annotation text.

[0064] Specifically, in this step, the automatically generated annotation text is fed back to the preset user interaction module for display: the automatically generated annotation text (i.e., the text generated by S301 and corrected by S302) is transmitted to the pre-set user interaction module and presented to the user in a visual manner. The user interaction module can be in the form of a graphical interface, a web platform, client software, etc., so that users can intuitively see the annotation content corresponding to each image, such as displaying the annotation text synchronously next to the image, or listing the annotation information in a separate text box. Regarding receiving user commands to view and / or edit automatically generated annotation text, the interactive module allows users to view the annotation text and then issue commands based on their own judgment. Viewing refers to browsing the text content to confirm whether the annotations meet expectations; while editing includes adding or deleting text, correcting wording, and adjusting descriptive logic. For example, if a user finds that the annotation text has omitted a key object, they can manually add it, or if they feel a description is inaccurate, they can rewrite it. The interactive module will receive and respond to these commands in real time, completing the adjustments to the text. In this embodiment of the invention, the viewed and / or modified annotation text will be stored. That is, regardless of whether the user modifies the annotation text, the final version after viewing and confirmation or editing will be saved by the system. The storage method may include local database storage, cloud storage, etc. At the same time, it will be associated with the corresponding image to ensure that each image has its matching annotation text confirmed by the user, which facilitates subsequent retrieval, management and traceability. In this embodiment of the invention, although the automatic generation and correction process has improved the annotation quality, machine processing may still have omissions or be inconsistent with actual needs. Through user interaction, humans can perform secondary verification and optimization of the annotated text, compensating for the limitations of machine intelligence, ensuring that the annotation results are more in line with the requirements of actual application scenarios, and guaranteeing the final accuracy of the annotated text. In this embodiment of the invention, users can directly view and intervene in the annotation process, adjusting the text according to their own professional knowledge or specific needs, avoiding passive acceptance caused by complete reliance on automated processes. This interactive mode gives users control over the annotation process, increasing their acceptance of the annotation results. In a further embodiment, the image intelligent annotation processing method, wherein the step of receiving user operation instructions to view and / or modify the automatically generated annotation text, and storing the viewed and / or modified annotation text, further includes: Monitor user actions to modify and edit automatically generated annotation text, learn user habits and preferences for annotation modification, and optimize subsequent annotation strategies.

[0065] In this embodiment of the invention, the monitoring of user modifications to automatically generated labeled text specifically involves the system tracking and recording all user modifications to the automatically labeled text in the interaction module in real time. This includes details such as the location of the modification (e.g., adding a description, deleting a word, adjusting sentence order), a comparison of the content before and after the modification, the frequency of modification (e.g., the number of times a certain type of element's label is modified), and the duration of the modification. For example, if a user frequently corrects "car" to more specific vehicle types such as "sedan" or "truck," or frequently adds information such as the color or state of an object, these actions will be accurately captured by the system. Then, the system learns from users' habits and preferences in annotation and modification operations. Based on the monitored modification data, the system uses machine learning algorithms (such as association rule analysis and cluster analysis) to uncover users' operational habits and preferences. For example, it might analyze whether users tend to add detailed descriptions such as age and actions to character elements, emphasize environmental features (e.g., "street on a rainy day" rather than simply "street") to scene elements, or habitually use specific terminology in specific industry scenarios (e.g., medical image annotation). This learning process continuously accumulates users' personalized annotation styles and professional needs, forming a user-specific annotation preference model. Based on learned user habits and preferences, the system dynamically adjusts subsequent automatic annotation logic and rules. For example, if it detects that users frequently add color information to objects, the subsequent annotation strategy will prioritize extracting and adding color features; if users frequently modify the description of category names (such as refining "animals" to "canines" or "felines"), the system will optimize the granularity of element recognition category segmentation, making the automatically generated annotations more consistent with the user's commonly used expressions. Meanwhile, for annotation types that users modify less frequently, the system will maintain the existing strategy to ensure reasonable resource allocation. As can be seen, in this embodiment of the invention, by learning users' modification habits, the system can generate annotation text that better meets users' expectations and professional needs, reducing the workload of users in making modifications. For example, users in professional fields (such as architects and doctors) have specific terminology preferences, and after system optimization, annotations that conform to their industry standards can be directly generated, significantly reducing the cost of manual adjustments. Furthermore, different users may have different annotation habits and focuses (e.g., ordinary users pay more attention to the overall object, while professional users pay more attention to detailed features). The system learns to implement a "personalized" annotation strategy, avoiding the problem of insufficient adaptability caused by uniform annotation, demonstrating stronger intelligent adaptability, and enhancing the system's personalization and intelligence. Furthermore, as the system gains a deeper understanding of user preferences, the automatic annotation text becomes increasingly aligned with user needs, gradually reducing the frequency of manual modifications required by users. This leads to a continuous improvement in the efficiency of the entire annotation process, especially in large-scale image annotation scenarios, where it can save significant manpower and time costs and continuously enhance annotation efficiency.

[0066] In a further embodiment, the image intelligent annotation processing method, after the step of automatically classifying, storing, and indexing each image for which annotation text is generated, includes: The system receives a text-based image search command from the user; parses the text-based image search command, and searches the categorized and stored images using an established index to find and display images that match the text-based image search command.

[0067] In this embodiment of the invention, regarding receiving image search instructions with text input by the user, specifically, the user can input image search requests containing text information through the system's interactive interface (such as a search box, voice-to-text input, etc.), such as "image of a red car driving on a rainy street" or "image of a doctor in a white coat in an operating room," etc. The system of this embodiment of the invention will receive these text instructions in real time as the basis for subsequent search operations. The system then parses the text-based image search command. Specifically, it uses natural language processing (NLP) technology to analyze and deconstruct the received text command, extracting key information. This includes identifying the core elements of the command (such as objects, scenes, people, etc.), the attributes of the elements (such as color, state, action, etc.), and the relationships between elements (such as positional relationships, behavioral associations, etc.). For example, for the command "image of a red car driving on a rainy street," the system will parse out key information such as "car" (object element), "red" (color attribute), "street" (scene element), "rainy day" (scene state), and "driving" (action attribute). This invention uses an established index to search for categorized images, finding and displaying images that match the text-based image search command. Specifically, the system pre-creates an index for the categorized images, typically based on the image's labeled text information (such as element category, attributes, and relationships) for easy and rapid retrieval. After parsing the key information of the search command, the system compares this information with the index, filtering out matching images from the categorized image library. Finally, the matched images are sorted according to relevance and other rules before being displayed to the user.

[0068] Thus, the text-based search instructions used in this embodiment of the invention can more accurately express the user's intent. By parsing the instructions and performing precise index-based searches, irrelevant images can be effectively filtered out, allowing users to quickly find content that meets their expectations, avoiding blindly searching through a massive number of images, and satisfying users' precise image search needs. Furthermore, it employs natural language processing technology for precise command parsing, combined with a pre-built index, significantly shortening search time and reducing invalid matches. Compared to traditional image feature-based searches (such as pixels and color histograms), searches based on text commands and labeled indexes are better able to understand users' semantic needs, resulting in higher accuracy and improved search efficiency and accuracy. Furthermore, users don't need complex search skills; they can simply describe the image they want using natural language to complete the search, lowering the barrier to entry. At the same time, quickly finding matching images reduces user waiting time and operational costs, improving user satisfaction and optimizing the user search experience.

[0069] The present invention will be further described in detail below through another specific application embodiment.

[0070] Example 2 like Figure 2 As shown in the second embodiment, an intelligent image annotation processing method includes: S10, Image Acquisition and Preprocessing, corresponding to S11 and S12; S11. Collect images from local storage, network, and camera, then proceed to S12; In this embodiment of the invention, image acquisition will collect image data from different channels, including but not limited to local uploads, web crawling, and camera captures.

[0071] S12. Preprocess the acquired images and proceed to S13; In this embodiment, the acquired images are preprocessed, such as resizing, grayscale conversion, and noise removal, to optimize image quality and improve the efficiency and accuracy of subsequent processing.

[0072] S13. Perform intelligent analysis and annotation on the preprocessed image, then proceed to S14 and 15; S14. Analyze the preprocessed image; S15. Generate annotations for the analyzed images and proceed to S16; In this embodiment of the invention, regarding intelligent analysis and annotation, advanced deep learning algorithms are used to perform multi-dimensional analysis on the preprocessed images; and convolutional neural networks (CNN) and other models are used to accurately identify various elements in the images, such as objects, scenes, and people, to obtain key information such as their categories and locations; then, with the help of natural language processing (NLP) technology and combined with image semantic understanding, appropriate annotation text is generated. At the same time, attention mechanisms are introduced to enable the system to focus on key areas in the images, further improving the accuracy and intelligence of annotation.

[0073] S16, Enter the user interaction steps, then proceed to S17 and S18; S17. Receive user feedback on annotations; S18, Store the data and proceed to S19; In this embodiment of the invention, regarding user interaction, the automatically generated annotation information will be fed back to the user interaction and customization module for users to view, edit and supplement; in this way, users can adjust the annotations according to actual needs, and the system can learn the user's operating habits and preferences to optimize subsequent annotation strategies.

[0074] S19. Store image data.

[0075] In this embodiment of the invention, when storing image data, labeled image data and labeling information can be stored; specifically, an efficient database management system can be used to classify and store the data and establish indexes, which facilitates subsequent fast retrieval and data management.

[0076] Exemplary device like Figure 3 As shown, this embodiment of the invention provides an intelligent image annotation processing device, which includes: Image acquisition and preprocessing module 310 is used to acquire images from different channels and preprocess the images; The intelligent analysis module 320 is used to perform multi-dimensional analysis on the preprocessed image using deep learning algorithms, identify various elements in the image, and extract key information corresponding to various elements in the image. The annotation module 330 is used to automatically generate corresponding annotation text for each image based on the key information corresponding to various elements in the extracted image, using natural language processing technology and combined with image semantic understanding. User interaction module 340 is used to feed back the automatically generated annotation text of each image to the preset user interaction module for display; receive user operation instructions to view and / or modify the automatically generated annotation text, and store the viewed and / or modified annotation text; The image data storage module 350 is used to automatically classify, store, and index each image for which annotation text is generated. The annotation optimization module 360 ​​is used to monitor users' actions in modifying and editing automatically generated annotation text, learn users' habits and preferences in annotation modification, and optimize subsequent annotation strategies.

[0077] Based on the above embodiments, the present invention also provides a smart terminal, the principle block diagram of which can be as follows: Figure 4 As shown. The intelligent terminal includes a processor, memory, network interface, display screen, and database connected via a system bus.

[0078] One type of smart terminal according to this application further includes one or more programs, wherein one or more programs are stored in a memory and configured to be executed by one or more processors, wherein the one or more programs include a method for performing intelligent image annotation processing.

[0079] In this context, the memory refers to the hardware device used to store program code and data, specifically implemented using solid-state drives (SSDs) or flash memory chips. Its function is to provide storage space for the processor to execute instructions. The processor is the arithmetic unit used to execute program instructions, specifically implemented using a central processing unit (CPU) or a graphics processing unit (GPU). Its function is to control the process of intelligent image annotation processing by running program code. The program includes methods and steps such as image acquisition and preprocessing, intelligent analysis and annotation, user interaction, and image data storage. Its purpose is to integrate intelligent image annotation processing technology into the terminal device to realize intelligent image annotation processing.

[0080] Specifically, when the program is executed by the processor, it first continuously collects and acquires images from different channels and preprocesses the images; then, it uses a deep learning algorithm to perform multi-dimensional analysis on the preprocessed images, identify various elements in the images, and extract key information corresponding to each element; based on the extracted key information corresponding to each element in the images, it uses natural language processing technology, combined with image semantic understanding, to automatically generate corresponding labeled text for each image; and for each image for which labeled text is generated, it automatically classifies, stores, and indexes the images, as described in the above method embodiment.

[0081] Through the above technical solution, this application can add annotations to images efficiently, accurately and intelligently, meeting the needs of diverse application scenarios and the image annotation needs of different users in diverse scenarios, bringing a revolutionary breakthrough to the management and utilization of image data.

[0082] In addition, this application proposes a computer-readable storage medium that enables an electronic device to perform an image intelligent annotation processing method when the instructions in the storage medium are executed by the processor of the electronic device.

[0083] Computer-readable storage media refers to the physical carrier used to store executable program code. Specifically, it can be implemented using storage devices such as solid-state drives, flash memory chips, or optical discs. Its function is to carry the program instructions that implement the intelligent image annotation processing method. Processor execution instructions refer to the process by which the central processing unit reads and processes the program code in the storage medium. Specifically, it can be implemented using a multi-core processor or a distributed computing architecture. Its function is to drive electronic devices to perform operations such as image acquisition and preprocessing, intelligent analysis and annotation, and user interaction.

[0084] Specifically, when the instructions in the storage medium are executed by the processor, the electronic device first acquires images from different channels and preprocesses the images; then, it uses a deep learning algorithm to perform multi-dimensional analysis on the preprocessed images, identifies various elements in the images, and extracts key information corresponding to each element; based on the extracted key information corresponding to each element in the images, it uses natural language processing technology, combined with image semantic understanding, to automatically generate corresponding labeled text for each image; and for each image for which labeled text is generated, it automatically classifies, stores, and indexes the images, as described in the above method embodiment.

[0085] This application can efficiently, accurately and intelligently add annotations to images, meeting the needs of diverse application scenarios and the image annotation needs of different users in diverse scenarios, bringing a revolutionary breakthrough to the management and utilization of image data.

[0086] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for intelligent image annotation processing, characterized in that, include: Images from different sources are collected and preprocessed. Deep learning algorithms are used to perform multi-dimensional analysis on preprocessed images, identify various elements in the images, and extract key information corresponding to each element. Based on the key information corresponding to various elements in the extracted images, natural language processing technology is used, combined with image semantic understanding, to automatically generate corresponding annotation text for each image; For each image that generates labeled text, it is automatically categorized, stored, and indexed.

2. The image intelligent annotation processing method according to claim 1, characterized in that, The steps of acquiring images from different channels and preprocessing the images include: The data collection includes images uploaded from local storage, images crawled from the web, and images captured by a camera. The acquired images are preprocessed to generate preprocessed images; the preprocessing operations include resizing the images, converting them to grayscale, and removing noise to optimize image quality.

3. The image intelligent annotation processing method according to claim 1, characterized in that, The steps of using deep learning algorithms to perform multi-dimensional analysis on the preprocessed image, identify various elements in the image, and extract key information corresponding to each element include: A predetermined deep learning algorithm is used to perform multi-dimensional analysis on the preprocessed images; A convolutional neural network model is used to identify various elements in an image and extract key information corresponding to each element. These elements include object elements, scene elements, and person elements. The key information includes category information and location information.

4. The image intelligent annotation processing method according to claim 1, characterized in that, The step of generating corresponding labeled text for each image based on the key information corresponding to various elements extracted from the image, using natural language processing technology and combined with image semantic understanding, includes: Based on the key information corresponding to various elements in the extracted images, natural language processing technology is used, combined with image semantic understanding, to generate corresponding matching annotation text for each image; Furthermore, an attention mechanism is introduced to focus on key areas in the image, further refining the labeled text to improve labeling accuracy and intelligence.

5. The image intelligent annotation processing method according to claim 1, characterized in that, Before the step of automatically classifying, storing, and indexing each image for which annotated text is generated, the following steps are also included: The automatically generated annotation text for each image is displayed in the preset user interaction module. Receive user operation instructions to view and / or modify the automatically generated annotation text, and store the viewed and / or modified annotation text.

6. The image intelligent annotation processing method according to claim 5, characterized in that, The step of receiving user operation instructions to view and / or modify the automatically generated annotation text, and storing the viewed and / or modified annotation text, further includes: Monitor user actions to modify and edit automatically generated annotation text, learn user habits and preferences for annotation modification, and optimize subsequent annotation strategies.

7. The image intelligent annotation processing method according to claim 1, characterized in that, The step of automatically classifying, storing, and indexing each image for which annotation text is generated includes: Receive image search commands with text input from the user; The text-based image search command is parsed, and the images stored in the category are searched using the established index to find and display the images that match the text-based image search command.

8. An intelligent image annotation processing device, characterized in that, The device includes: The image acquisition and preprocessing module is used to acquire images from different channels and preprocess the images. The intelligent analysis module uses deep learning algorithms to perform multi-dimensional analysis on preprocessed images, identify various elements in the images, and extract key information corresponding to each element. The annotation module is used to automatically generate corresponding annotation text for each image based on the key information corresponding to various elements in the extracted images, using natural language processing technology combined with image semantic understanding. The user interaction module is used to feed back the automatically generated annotation text of each image to the preset user interaction module for display; receive user operation instructions to view and / or modify the automatically generated annotation text, and store the viewed and / or modified annotation text; The image data storage module is used to automatically classify, store, and index each image used to generate labeled text. The annotation optimization module is used to monitor users' modifications and edits to the automatically generated annotation text, learn users' habits and preferences in annotation modification, and optimize subsequent annotation strategies.

9. A smart terminal, characterized in that, It includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors, wherein the one or more programs include methods for performing any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Picture labeling method and device, equipment and storage medium

    CN117746183A

  • Scene understanding information generation method and device, equipment and medium

    CN119251657A