Label Recognition Method, Object Processing Method, Computing Device, Storage Medium, and Program Product

By hierarchically grouping and feature extraction of multiple tags and matching features with the features of the objects to be identified, the problem of inaccurate tag recognition in the prior art is solved, and more accurate tag recognition and personalized services are achieved.

CN118552985BActive Publication Date: 2025-07-04ALIBABA CLOUD COMPUTING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411016900.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2025-07-04
Estimated Expiration
2044-07-26

AI Technical Summary

Technical Problem

Existing object recognition services cannot meet users' refined label needs, resulting in inaccurate label recognition results.

Method used

By obtaining multiple tags, grouping them according to hierarchical relationships, label features are extracted, and matching the object features of the object to be identified based on the tag features, label identification is performed in order from high to low.

Benefits of technology

It improves the accuracy and flexibility of label recognition, reduces the probability of misidentification, enhances the user experience, and broadens the possibility of personalized services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118552985B_ABST
    Figure CN118552985B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a label recognition method, an object processing method, a computing device, a storage medium, and a program product. Among them, a plurality of labels are obtained; the plurality of labels are grouped according to a hierarchical relationship to obtain multiple groups of label combinations; label features respectively corresponding to the labels in the multiple groups of label combinations are extracted; wherein, the multiple groups of label combinations are used to match with a to-be-recognized object in descending order of hierarchy based on the label features and the image features (object features) of the to-be-recognized object to obtain at least one target label of the to-be-recognized object. The technical solution provided by the embodiment of the present application improves the accuracy of label recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of data processing, and in particular, to a method for label recognition, a method for object processing, a computing device, a storage medium, and a program product. Background Art

[0002] With the rapid development of artificial intelligence technology, object recognition and automatic annotation technology have increasingly become key components in the fields of information processing and data analysis. For example, in recent years, image recognition algorithms based on artificial intelligence (AI) have become mature and are widely used in various industries, becoming a bridge connecting the physical world and the digital world. Many service providers at home and abroad have launched object recognition services, which can provide users with convenient label generation capabilities. When users provide objects to be recognized, such as images, they can generate labels corresponding to the objects to be recognized with the help of the object recognition models of the service providers.

[0003] However, although existing object recognition services provide rich label libraries and can meet most general needs, the mechanism for matching labels corresponding to objects to be recognized is still to select matching labels from the label library, and it cannot meet the refined label needs of users. Therefore, how to accurately recognize the labels matching the objects to be recognized has become a technical problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] Embodiments of the present application provide a method for label recognition, a method for object processing, a computing device, a storage medium, and a program product, so as to solve the technical problem of inaccurate label recognition results in the prior art.

[0005] In a first aspect, embodiments of the present application provide a method for label processing, including:

[0006] Obtain a plurality of labels;

[0007] Group the plurality of labels according to a hierarchical relationship to obtain multiple groups of label combinations;

[0008] Extract the label features corresponding to the labels in the multiple groups of label combinations respectively;

[0009] Among them, the multiple groups of label combinations are used to match the object to be recognized based on the label features and the object features of the object to be recognized in descending order of hierarchy to obtain at least one target label of the object to be recognized.

[0010] Optionally, the grouping the plurality of labels according to a hierarchical relationship to obtain multiple groups of label combinations includes:

[0011] Identify the hierarchical levels to which the plurality of labels belong;

[0012] Combine labels at the same level to obtain at least one set of label combinations.

[0013] Optionally, the identifying the levels to which the multiple tags belong includes:

[0014] The hypernyms of the multiple tags are searched in a vocabulary database until a root node is reached, and the level to which the tag belongs is determined according to the path length from the tag to the root node.

[0015] Optionally, searching for hypernyms of the plurality of tags in a vocabulary database until a root node is reached comprises:

[0016] determining a synonym set of the plurality of tags in the vocabulary database;

[0017] The hypernyms of the synonym set in the vocabulary database are searched level by level until a root node is reached.

[0018] Optionally, after acquiring the plurality of tags, the method further includes:

[0019] Tags that meet a filtering condition are filtered out from the plurality of tags to update the plurality of tags.

[0020] Optionally, acquiring multiple tags includes:

[0021] Get a tag list including multiple tags provided by the user;

[0022] The step of extracting label features corresponding to the labels in the plurality of label combinations comprises:

[0023] A multimodal model is used to extract label features corresponding to the labels in the multiple groups of label combinations; the multimodal model is a pre-trained model.

[0024] In a second aspect, an embodiment of the present application provides an object processing method, including:

[0025] Get the object to be identified;

[0026] Extracting object features of the object to be identified;

[0027] Determine the label features corresponding to the labels in the extracted multiple label combinations; the multiple label combinations are obtained by grouping the acquired multiple labels according to a hierarchical relationship;

[0028] Based on the object features and the label features, the multiple groups of label combinations are matched with the objects to be identified respectively to obtain at least one target label corresponding to the objects to be identified.

[0029] Optionally, based on the object features and the label features, matching each of the multiple groups of label combinations with the object to be recognized to obtain at least one target label corresponding to the object features, including:

[0030] Based on the label features and the object features, matching the object to be recognized with the labels in the highest-level label combination to obtain successfully matched labels;

[0031] Starting from the next-level label combination of the highest-level label combination, determining at least one label in the current-level label combination that has a hyponymy relationship with the successfully matched label in the previous-level label combination;

[0032] Matching the at least one label with the object to be recognized to obtain successfully matched labels;

[0033] Taking the at least one successfully matched label corresponding to the multiple groups of label combinations as the at least one target label corresponding to the object to be recognized.

[0034] Optionally, based on the object features and the label features, matching each of the multiple groups of label combinations with the object to be recognized to obtain at least one target label corresponding to the object to be recognized, including:

[0035] Calculating the feature similarity between the label features corresponding to the labels in the multiple groups of label combinations and the object features of the object to be recognized respectively;

[0036] Determining at least one target label that meets the similarity condition and has similar features.

[0037] Optionally, extracting the object features of the object to be recognized includes:

[0038] Using a multimodal model to extract the object features of the object to be recognized;

[0039] Determining the label features corresponding to the labels in the multiple groups of label combinations extracted includes:

[0040] Determining the label features corresponding to the labels in the multiple groups of label combinations extracted using the multimodal model; the multiple obtained labels include a label list including multiple labels provided by the user.

[0041] Optionally, when the object to be recognized is an image to be recognized, the method further includes:

[0042] Based on the at least one target label, determining the image category of the image to be recognized;

[0043] Or;

[0044] Search for a target image that matches the image to be recognized based on the at least one target tag;

[0045] Or;

[0046] Verify whether the image to be recognized meets the image requirements based on the at least one target tag.

[0047] In a third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the tag processing method as described in the first aspect above, or the object processing method as described in the second aspect above.

[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processing component, it implements the tag processing method as described in the first aspect above, or the object processing method as described in the second aspect above.

[0049] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processing component, they implement the tag processing method as described in the first aspect above, or the object processing method as described in the second aspect above.

[0050] In the embodiments of the present application, by obtaining multiple tags; grouping the multiple tags according to the hierarchical relationship to obtain multiple groups of tag combinations; extracting the tag features corresponding to the tags in the multiple groups of tag combinations respectively; wherein, the multiple groups of tag combinations are used to match the object to be recognized based on the tag features and the object features of the object to be recognized in the order from high to low in the hierarchy to obtain at least one target tag of the object to be recognized. In the embodiments of the present application, by grouping multiple tags according to the hierarchical relationship, it is possible to better understand and utilize the context and semantic associations between tags to ensure that the subsequent tag recognition process is more meticulous and accurate. Matching the target tag of the object to be recognized based on the tag text features and the visual features of the image to be recognized helps to improve the accuracy of recognition. Matching in the order from high to low in the tag hierarchy to obtain at least one target tag of the object to be recognized can further reduce the probability of misrecognition, improve the tag coverage, and further ensure the accuracy of tag recognition.

[0051] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0053] Figure 1 Fig. shows a flowchart of an embodiment of a tag processing method provided by the present application;

[0054] Figure 2 Fig. shows a flowchart of an embodiment of an object processing method provided by the present application;

[0055] Figure 3 Fig. shows a schematic diagram of scenario interaction in an actual application of the technical solution of the embodiment of the present application;

[0056] Figure 4 Fig. shows a schematic structural diagram of an embodiment of a tag processing device provided by the present application;

[0057] Figure 5 Fig. shows a schematic structural diagram of an embodiment of an object processing device provided by the present application;

[0058] Figure 6 Fig. shows a schematic structural diagram of a computing device provided by the present application. Detailed implementation manners

[0059] To enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application.

[0060] In some processes described in the specification, claims and the above accompanying drawings of the present application, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this text or may be executed in parallel. The operation numbers such as 101, 102, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this text are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0061] The technical solution of the embodiment of the present application can be applied to object recognition scenarios, such as commodity classification on e-commerce platforms, disease diagnosis of medical image analysis, object recognition in smart homes, and customized content recommendation systems. In a practical application, it can be applied to image recognition scenarios, such as e-commerce scenarios, where it is necessary to accurately identify clothing pictures uploaded by users to match the corresponding commodity categories.

[0062] As described in the background technology, the traditional method usually relies on artificial intelligence (AI) recognition models generated by some service providers to perform object recognition. The AI ​​model is usually obtained based on pre-label training. For example, in the clothing image recognition scenario, due to the large number of clothing styles and rapid updates, the preset labels are difficult to cover all details and emerging trend elements, so the recognition results may not meet user needs.

[0063] In order to overcome the limitations of traditional preset labels and generate more accurate labels, the inventors thought that users can retrain the model by customizing labels, but this solution requires users to collect sample images and customize labels by themselves, and then use the model provided by the service provider for customized training, and then deploy it to the production environment to realize personalized label recognition services. This method has a cumbersome and complicated process, which not only involves a lot of data preparation work and professional algorithm knowledge, but also needs to cross the technical threshold of model training and deployment, and if multi-label recognition is involved, the complexity will be higher. Therefore, the inventors have proposed the technical solution of the embodiment of the present application after a series of in-depth studies. By obtaining multiple expressions and grouping the labels in multiple labels according to the hierarchical relationship, multiple groups of label combinations are obtained to obtain multiple labels, that is, the system automatically performs intelligent grouping according to the hierarchical relationship of the labels, so as to better understand and utilize the context and semantic associations between the labels to ensure that the subsequent label recognition process is more detailed and accurate. Matching the target label of the object to be identified based on the text features of the label and the visual features of the image to be identified helps to improve the accuracy of recognition. Matching is performed in descending order of the label levels to obtain at least one target label of the object to be identified, which can not only reduce the probability of misidentification and improve label coverage, but also ensure label recognition accuracy.

[0064] The technical solution of the embodiment of the present application not only simplifies the processing flow of tag recognition and lowers the technical threshold, so that users without AI expertise can easily apply it, but also significantly improves the flexibility and accuracy of object recognition, enhances user experience, and broadens the possibility of personalized services, bringing revolutionary changes to the tag recognition process in various industries.

[0065] It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data can be used in the solutions described herein within the scope permitted by applicable laws and regulations of the country where it is located (for example, with the explicit consent of the user, giving the user a practical notice, etc.).

[0066] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0067] It should be noted that the technical solutions of the embodiments of this application are applicable to a network virtual environment. The users described generally refer to "virtual users". Real users can register user accounts on the server through registration to obtain a user identity in the network environment.

[0068] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope protected by this application.

[0069] Figure 1 This is a flowchart of an embodiment of a tag processing method provided for an embodiment of this application. The technical solution of this embodiment can be executed by the server. In a practical application, the technical solution of the embodiment of this application can be applied to a system architecture composed of a user terminal and a server. The user terminal can interact with the server through the network to receive or send messages, etc.

[0070] The user terminal can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5) application, or a light application (also known as a mini program, a lightweight application program), or a cloud application, etc. The user terminal can be deployed in an electronic device and needs to rely on the device or certain apps in the device to run, etc. The electronic device can, for example, have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, a desktop computer, a smart speaker, a smart watch, and so on.

[0071] The server side may include servers that provide various services, such as servers for tag processing, and servers for object recognition. It should be noted that the server side can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0072] Figure 1 The method of the illustrated embodiment may include the following steps:

[0073] 101: Get multiple tags.

[0074] The multiple tags obtained may be tags directly provided by the user, or tags selected or configured by the user from system-provided tags according to needs. The specific forms of the multiple tags obtained may include but are not limited to: a tag list, a tag set.

[0075] "Getting multiple tags" is the starting point of the entire tag processing process, and its core is to collect multiple tags configured by the user for the object to be identified. "Tag" can refer to a keyword or phrase that can be used to describe or classify the object to be identified. The object to be identified described in this article can be an image, text, audio or video, etc., and this application does not limit this. The tags provided by the user can reflect their understanding and classification intention of the object to be identified, which increases the retrievability of the object to be identified.

[0076] In this step, obtaining multiple tags includes: obtaining a tag list provided by the user including multiple tags. This step emphasizes that the user can play an active role in the annotation process. The server allows the user to customize the tags according to personal understanding or specific needs, thereby improving the personalization and pertinence of the tags. In other words, this application can solve the problem of poor flexibility in tag identification caused by matching the tags corresponding to the objects to be identified through the general tag library in the existing technology by obtaining the tags provided by the user. It can realize the identification of customized tags, meet the personalized needs of users, and improve the user experience.

[0077] Among them, the multiple tags provided by the user are not restricted by the preset tag library and can cover a wider range of topics and details, especially suitable for scenarios with specific industry backgrounds or personalized needs. The multiple tags provided by the user can be very diverse, including but not limited to keywords, phrases, numbers, and even self-created tags. In addition, the user can provide multiple tags through the user side. For example, an input bar for tags is displayed on the display interface of the user side for the user to input tags. Taking the specific case where the multiple tags can be a list of tags submitted by the user as an example, the form of the tag list can include strings.

[0078] In the embodiments of the present application, the scenario of obtaining multiple tags provided by the user can be the scenario of posting pictures on social media. For example, when a user uploads a photo on social media and wants to customarily add tags such as "travel", "food", "pet", etc. according to needs, so that these pictures can be more easily discovered by interested friends and are also convenient for the user to review and search in the future. It can also be the scenario where a merchant in an e-commerce platform needs to label new fashion items with trendy tags. For example, when a merchant shelves new clothing, they can manually add tags such as "dopamine", "retro style", "sustainable materials", etc. according to fashion trends and product features to help consumers quickly locate their favorite products and increase the exposure rate of the products. In addition, it can also include other scenarios. The purpose of obtaining multiple tags provided by the user in the present application is to enable at least one target tag in the tag list to be marked for the subsequent to-be-identified object according to the user's needs, rather than using the system default tags for marking.

[0079] Optionally, after obtaining the multiple tags, it can also include: screening out the tags that meet the screening conditions from the multiple tags to update the multiple tags.

[0080] In order to further improve the accuracy and practicality of the tags, after obtaining the multiple tags provided by the user, the multiple tags can be verified and cleaned, such as removing some meaningless tags such as garbled tags. The screening condition can be the definition rule of meaningless tags, such as regular expressions, etc., so that tags that meet the definition rule can be screened out.

[0081] In addition, a classification model can also be used to identify meaningless tags from the multiple tags, and the screening condition is also the classification result of the classification model being a meaningless tag. The classification model can be obtained through pre-training. For example, it can be trained according to sample tags and training tags indicating whether the sample tags are meaningless tags (for example, meaningless tags are 0, and other cases are 1).

[0082] The named entities in the multiple tags can also be identified through the NER (Name Entity Recognition) algorithm, so that the tags that are not named entities can be removed.

[0083] By performing label verification and cleaning, intelligent means can be used to remove those labels that may be inaccurate, duplicate, or have a weak relevance to the object to be recognized, ensuring that the multiple labels ultimately used for matching are more refined and highly relevant. This helps improve the efficiency and accuracy of subsequent label matching.

[0084] 102: Group multiple labels according to their hierarchical relationships to obtain multiple groups of label combinations.

[0085] Step 102 involves classifying and integrating the labels in the multiple labels provided by the user according to the hierarchical or relevant relationships between the labels. The principle of label grouping is usually based on the universality, specificity, and semantic inclusion relationships of the labels, forming a semantic structure with one or more levels, where each level represents a more specific or generalized concept category. Through this grouping, a well-structured label system that is easy to understand and retrieve can be constructed.

[0086] It can be understood that the hierarchical relationship is a semantic structure determined according to the hyponymy relationship between the labels in the multiple labels.

[0087] Among them, the construction of the hierarchical relationship is based on the internal logic of the labels. The internal logic of the labels refers to the hyponymy relationship determined according to the structure and / or semantics. The construction of the hierarchical relationship in this application is essentially to organize the logical information of the labels into a well-structured form. This process mainly includes: First, determine the broadest and most abstract top-level label as the root node of the hierarchical structure; Second, starting from the top-level label, gradually identify and define more specific and detailed lower-level concepts to form a hierarchical progressive relationship. Assuming that "animal" is used as the top-level label, for example, it can include lower-level labels such as "mammal", and under "mammal", it can be further divided into "felidae", "canidae", etc.

[0088] It should be noted that label grouping is not limited to a single dimension, but can also be divided into multiple levels according to different dimensions, such as by function, by theme, by time, etc., to form a multi-dimensional label tree structure.

[0089] The following embodiments are all exemplified by using the multiple labels obtained as the label list provided by the user. However, it should be noted that this application is not limited to the multiple labels obtained being all provided by the user, nor is it limited to the multiple labels only being in the form of a label list.

[0090] For example, taking the list of tags provided by the user including: "animal", "person", "cat", "dog", "woman", "man" as an example, according to the requirements of step 102, it is necessary to group the tags in the list of tags input by the user according to the hierarchical relationship. For the given list of tags ["animal", "person", "woman", "man", "cat", "dog"], the following grouping structure can be constructed to reflect their hierarchical relationship:

[0091] The first level includes: animal, person, where animal and person belong to different dimensions;

[0092] The second level includes: "cat", "dog", "man", "woman" (where "cat" and "dog" are subordinate tags of "animal", and "man" and "woman" are subordinate tags of "person").

[0093] In the above grouping, "animal" and "person" are the broadest classifications, while "cat", "dog", "woman", and "man" are more specific sub-classifications. Among them, "cat" and "dog" are sub-classes of "animal", and "woman" and "man" are sub-classes of "person", reflecting the inclusion relationship between the tags.

[0094] By such grouping, the tags can be effectively organized and managed, so that in the subsequent process of determining the target tags corresponding to the object to be recognized, the matching target tags can be more accurately located, thereby improving the efficiency of tag processing and the accuracy of tag recognition.

[0095] Furthermore, in step 102, after grouping the tags in the multiple tags according to the hierarchical relationship, multiple groups of tag combinations can be obtained based on the grouping results.

[0096] 103: Extract the tag features corresponding to the tags in the multiple groups of tag combinations respectively.

[0097] Among them, the above multiple groups of tag combinations are used to match the object to be recognized in the order from high to low in the hierarchy, based on the tag features corresponding to the tags in the multiple groups of tag combinations and the object features of the object to be recognized, to obtain at least one target tag of the object to be recognized.

[0098] In this step, the method of extracting the label features corresponding to the labels in the multiple groups of label combinations includes: using a multimodal model to extract the label features corresponding to the labels in the multiple groups of label combinations, wherein the multimodal model is a machine learning model that can process and integrate multiple types of data (such as images, text, audio, video, sensor data, etc.). Among them, the multimodal model is a pre-trained model, which can be a pre-trained large model. The pre-trained model solves the problem of not needing data calibration and model training, and reduces the recognition complexity. In this case, the pre-trained multimodal model can extract useful features more efficiently, without the need to train from scratch, greatly shortening the training time and improving the recognition accuracy. In this step, the multimodal model can be used to extract the label features corresponding to the labels in the multiple groups of label combinations, thereby ensuring the accuracy of subsequent label recognition. The multimodal model has zero-shot learning capabilities, which enables the multimodal model to classify or identify new categories that have never been seen without retraining. In an embodiment of the present application, the multimodal model can be a multimodal algorithm model that realizes text and object matching, which can understand and process data in two different forms of text and objects. For example, in the image recognition scenario, the multimodal model can be CLIP (Contrastive Language-Image Pre-training, a cross-modal pre-training model based on text-image contrast learning), which is pre-trained on large-scale unlabeled image-text pairs and can learn the joint embedding representation of images and texts. The label features extracted by the multimodal model can include the alignment relationship with the object, making the label features more accurate.

[0099] Optionally, the object features of the object to be identified can be extracted using the multimodal model.

[0100] The multimodal model can combine the extracted label features with the object features of the object to be identified and match them at different levels. From the highest level (more abstract labels) to the lower levels (specific and detailed labels), the matching range is gradually narrowed until at least one target label that best matches the object to be identified is found. This process utilizes a hierarchical combination of labels to ensure the accuracy and efficiency of recognition.

[0101] Optionally, label features corresponding to the labels in multiple label combinations can be saved according to a hierarchical relationship, so that when there are objects to be identified, label features corresponding to the labels in multiple label combinations can be obtained from the saved data, and matched with the labels in each label combination in order from high to low levels.

[0102] In this embodiment, by obtaining a plurality of tags; grouping the plurality of tags according to a hierarchical relationship to obtain multiple groups of tag combinations; extracting the tag features corresponding to the tags in the multiple groups of tag combinations respectively; wherein, the multiple groups of tag combinations are used to match with the object to be recognized based on the tag features and the object features of the object to be recognized in the order from high to low in the hierarchy to obtain at least one target tag of the object to be recognized. The embodiment of the present application obtains a plurality of tags to meet personalized needs. By grouping the plurality of tags according to the hierarchical relationship, the context and semantic association between the tags can be better understood and utilized to ensure that the subsequent tag recognition process is more meticulous and accurate. Matching the target tag of the object to be recognized based on the tag text features and the visual features of the image to be recognized helps to improve the accuracy of recognition. Matching in the order from high to low in the tag hierarchy to obtain at least one target tag of the object to be recognized can not only reduce the probability of misrecognition and improve the tag coverage, but also ensure the accuracy of tag recognition.

[0103] In some embodiments, grouping the tags in the plurality of tags according to a hierarchical relationship to obtain multiple groups of tag combinations may include: identifying the hierarchical levels to which the tags in the plurality of tags belong; combining the tags at the same hierarchical level to obtain at least one group of tag combinations.

[0104] First, clarify the hierarchical levels to which each tag belongs in the semantic structure. The hierarchical relationship of the tags defines the inclusion and being-included relationships between them. Subsequently, the tags belonging to the same hierarchical level can be combined together to form at least one group of tag combinations. Such an operation helps to distinguish different clusters of tags at different abstraction levels and facilitates the subsequent tag processing flow.

[0105] As a possible implementation manner, the tags at the same hierarchical level can be combined to obtain one group of tag combinations. It can be understood that there is only one group of tag combinations for one level of tags, and the corresponding number of tag combinations can be obtained according to the number of hierarchical levels.

[0106] For example, taking the second layer in the above example including: "cat", "dog", "man", "woman" as an example, since "cat", "dog", "man", "woman" all belong to the same hierarchical level, "cat", "dog", "man", "woman" can be combined to obtain a group of tag combinations of "cat + dog + man + woman".

[0107] As another possible implementation manner, the tags at the same hierarchical level and in the same dimension can also be combined to obtain one or more groups of tag combinations.

[0108] Specifically, when combining tags at the same level, tags in the same dimension at the same level can be combined to obtain multiple groups of tag combinations, which can greatly expand the flexibility and creativity of tag applications. This means that when combining tags at the same level, it is also necessary to consider the combination of tags within a single dimension at the same level to capture more complex and delicate correlations and user needs. The following are supplementary explanations and specific examples of this concept:

[0109] The free permutation and combination between tags in the same dimension at the same level refers to freely matching tags belonging to the same dimension within the same level. For example, taking the second layer in the above example including "cat", "dog", "man", and "woman" as an example, among them, "cat" and "dog" both belong to the dimension of "animal", and "cat" and "dog" can be combined to obtain a group of tag combinations of "cat + dog", while "man" and "woman" both belong to the dimension of "person", and "man" and "woman" can be combined to obtain a group of tag combinations of "man + woman". At this time, two groups of tag combinations can be obtained, namely the tag combination of "cat + dog" and the tag combination of "man + woman".

[0110] For the above solution of "combining tags in the same dimension at the same level to obtain multiple groups of tag combinations", a detailed embodiment is provided:

[0111] For example, taking the tag list including A, B, C, D, E, F, G, H, I, 1, 2, 3, 4, 5, 6, 7, 8, 9, 0 as an example, after identifying the levels to which the tags in the tag list belong and grouping them according to the hierarchical relationship, the following three levels can be obtained:

[0112] First layer: A, B, C; among them, A, B, and C belong to different dimensions;

[0113] Second layer: D, E, F, G, H, I (where D and E are lower-level tags of A, F and G are lower-level tags of B, and H and I are lower-level tags of C); DE, FG, and HI belong to different dimensions;

[0114] Third layer: 1, 2, 3, 4, 5, 6, 7, 8, 9, 0 (where 1 and 2 are lower-level tags of D, 3 and 4 are lower-level tags of E, 5 and 6 are lower-level tags of F, 7 and 8 are lower-level tags of G, 9 is a lower-level tag of H, and 0 is a lower-level tag of I); 12, 34, 56, 78, 9, and 0 belong to different dimensions;

[0115] Further, once the hierarchy of each tag is identified, the next step is to combine the tags at the same level together to form different tag combinations. The advantage of doing this is that it enables quick positioning and manipulation of the tag sets at the same level, facilitating subsequent processing and application.

[0116] Combining the tags at the same level and in the same dimension, the following multiple groups of tag combinations can be obtained as follows:

[0117] Multiple groups of tag combinations in the first layer:

[0118] Combination 11: {A};

[0119] Combination 12: {B};

[0120] Combination 13: {C};

[0121] Multiple groups of tag combinations in the second layer:

[0122] Combination 21: {D, E};

[0123] Combination 22: {F, G};

[0124] Combination 23: {H, I};

[0125] Multiple groups of tag combinations in the third layer:

[0126] Combination 31: {1, 2};

[0127] Combination 32: {3, 4}

[0128] Combination 33: {5, 6};

[0129] Combination 34: {7, 8};

[0130] Combination 35: {9};

[0131] Combination 36: {0};

[0132] In the above grouping, it reflects that the way to combine the tags at the same level can be the combination between the tags at the same level and in the same dimension. Through such grouping, the hierarchical structure and attribution relationship between the tags can be clearly seen, providing a convenient basis for subsequent tag management and content processing.

[0133] Among them, there are multiple implementation ways to identify the hierarchical belonging of the tags in the above multiple tags.

[0134] In a possible implementation way, identifying the hierarchical belonging of the tags in multiple tags may include: querying the hypernyms of the tags in multiple tags in the lexical database until reaching the root node, and determining the hierarchical belonging of the tags according to the path length from the tags to the root node;

[0135] Among them, the above hierarchical relationship is a semantic structure determined according to the hierarchical relationship between tags in multiple tags. The above root node can be the highest-level vocabulary corresponding to the tags in the lexical database, which is usually an entity or an event, and is a relatively broad category.

[0136] In this embodiment, the purpose of querying the hypernyms of the tags in multiple tags until reaching the root node is to determine the hierarchical position of the tags provided by the user in the semantic structure. First, the hypernym concepts of these tags need to be found until reaching the highest level of the entire semantic structure - the root node. The root node is the starting point of concepts in different dimensions (i.e., there can be multiple root nodes). Through this process, the hierarchical relationship between tags can be established.

[0137] Among them, the lexical database is a dictionary that organizes vocabulary according to semantic relationships to form a semantic network, and this semantic relationship includes the hierarchical relationship. For the tags in multiple tags, in the lexical database, the hypernyms of the tag can be searched level by level according to the hierarchical relationship until the root node is found. Tracing back from the root node to this tag can determine the level where the tag is located.

[0138] For example, the tag list input by the user contains two tags, "A" and "B". At this time, it is not clear which levels "A" and "B" belong to. Therefore, it is necessary to query through the lexical database. Suppose it is found in the lexical database that "A" is already the root node, and for "B", its hypernym "C" is found, and then further through "C" to find the root node "D". At this time, it can be determined that the level where "A" is located is the first level (the highest level), which can also be called a first-level tag. And the level where "B" is located is the third level (i.e., D > C > B), which can also be called a third-level tag.

[0139] Among them, in the lexical database, the vocabulary of each part of speech is organized into synsets, and each synset contains words with similar meanings. The synsets are linked to each other through semantic relationships to form a semantic network. Therefore, in some embodiments, querying the hypernyms of the tags in the tag list in the lexical database until reaching the root node may include: determining the synsets of the tags in multiple tags in the lexical database; searching for the hypernyms of the synsets in the lexical database until reaching the root node.

[0140] That is, for the tags in multiple tags, the synsets that are hit can be searched from the lexical database. The synset may include the tag or a synonym of the tag, and then the hypernym sets of the synsets are searched level by level until reaching the root node.

[0141] Optionally, the lexical database can be, for example, WordNet (a dictionary based on cognitive linguistics), which contains a large number of lexical entries and the semantic relationships between them, including synsets, hypernyms, hyponyms, etc. For each label input by the user, it can be queried in WordNet to find its corresponding hypernym, and then the hypernym is further searched for its hypernym, repeating this process until the corresponding root node is reached.

[0142] In this embodiment, it is considered that if the label input by the user does not exist in the lexical database, this will hinder the determination of the hierarchical relationship. To solve this problem, the synsets of the labels in the multiple labels are determined; the hypernyms of the synsets in the lexical database are searched level by level until the root node is reached. Among them, the synset contains words that are similar in meaning or the same as the input label, and these words are likely to exist in the database and can represent the meaning of the input label. By searching for the set of words with similar meanings to the input label in the lexical database, even if the input label itself is not directly recorded, its hyponym and hypernym relationships can be indirectly determined through the records of its synonyms in the lexical database.

[0143] For example, the label list input by the user contains the label "sports car". When the server assigns a suitable classification level to this label, in the lexical database (such as WordNet or other professional databases), the specific label "sports car" is not directly recorded.

[0144] At this time, the server first identifies the possible synsets or near-synonym sets of "sports car". These synsets may include "racing car", "high-performance car", "supercar", etc. These words have corresponding entries in the database and can represent the basic meaning of "sports car".

[0145] Next, the server searches for the hypernyms of these synsets in the lexical database. For example, the possible hypernym of "racing car" is "car", and "high-performance car" and "supercar" may also point to "car". This process will continue until the root node is found, such as "vehicle".

[0146] Through the above steps, the server can construct the hierarchical relationship of "sports car". Although "sports car" itself does not directly appear in the database, through its synset, the system can indirectly determine the level of "sports car". For example, the finally determined semantic structure is: "vehicle" > "car" > "sports car", and thus it can be indirectly determined that the level of "sports car" is the third level, which can also be called a three-level label.

[0147] Furthermore, by finding the hypernyms of a tag or its synset up to the root node, the hierarchical level of the tag can be determined by the path length from the tag to the root node. The farther a tag is from the root node, the lower its level in the semantic structure, that is, the more specific it is; conversely, the more abstract or generalized it is.

[0148] Starting from the root node, the server traces down along the hypernym relationship and records the path length from each tag to the root node. Based on this length, the server can assign the tags to the corresponding levels to form a well-defined tag structure. Specifically, the server needs to determine one or more of the broadest categories as the root node, which cover all possible hyponym concepts. For example, in a product classification system, "goods" can be used as a root node. Further, the server uses an existing lexical database (such as WordNet, etc.) to query the hypernyms of each tag until it reaches the root node. For example, for the tag "laptop", its hypernyms may be "computer", "electronic device" until "goods". During the query process, the server records the path length from each tag to the root node, that is, the number of hypernyms. According to the number of hypernyms from each tag to the root node, its corresponding level is determined. For example, for the tag "laptop", its hypernyms include "computer", "electronic device" until the root node "goods". When "goods" is regarded as the first level, through the hypernym data of "laptop", it can be determined that "laptop" is located at the fourth level.

[0149] Through the above steps, even if the tags entered by the user are not directly recorded in the lexical database, the system can still determine the semantic levels of these tags by finding synsets and tracing back to the root node, and construct a multi-level and ordered tag system. This not only improves the flexibility and adaptability of tag processing, but also enhances the accuracy and efficiency of tag applications based on semantic understanding, such as content classification, information retrieval, and personalized recommendation scenarios.

[0150] Optionally, the path length corresponding to each level can be set according to requirements. For example, when the path length of each level is set to 1, for "short-sleeved T-shirt", querying in the lexical database, it is found that its hypernym is "T-shirt", and further tracing up to "top", and finally reaching the root node "clothing". The path is: "clothing" > "top" > "T-shirt" > "short-sleeved T-shirt", then the path length of "short-sleeved T-shirt" is determined to be 4 (starting from the root node "clothing", passing through "top", "T-shirt" to "short-sleeved T-shirt"), then the level of short-sleeved T-shirt can be determined to be the fourth level. By analogy, the path length of "T-shirt" is 3, then the level of "T-shirt" can be determined to be the third level. Among them, the path length reflects the degree of abstraction of the tag, and it can be understood that the larger the value of the path length, the more specific the tag is.

[0151] In yet another possible implementation, identifying the hierarchical levels to which the labels in multiple labels belong may include: using a large model to identify the hierarchical levels to which different labels in the multiple labels belong.

[0152] That is, the large model can be used to hierarchically classify the labels in the multiple labels according to the superordinate and subordinate relationships, so as to determine the hierarchical level to which each label belongs and the labels belonging to the same hierarchical level, etc.

[0153] Optionally, the large model can be a large language model (LLM for short). A large language model refers to a natural language processing model with an extremely large number of parameters and computing power, which can be used for various natural language processing tasks such as text generation, machine translation, and question answering systems. The large language model is driven by AI algorithms using deep learning technology and uses a large dataset to evaluate, standardize, and generate relevant content, as well as make accurate predictions. Examples of large language models can be GPT-3 (Generative Pre-Trained Transformer -3, the third-generation generative pre-trained model), GPT-4 (Generative Pre-Trained Transformer -4, the fourth-generation generative pre-trained model), BERT (Bidirectional Encoder Representation from Transformers, a bidirectional encoder model based on Transformers), Turing NLG (Turing Natural language Generation), etc. The present application does not limit this.

[0154] In some embodiments, the method may further include: obtaining an object to be recognized; extracting the image feature object feature of the object to be recognized; based on the image feature object feature and the label feature, respectively matching multiple groups of label combinations with the object to be recognized to obtain at least one target label corresponding to the object to be recognized.

[0155] That is, in the embodiments of the present application, it is also possible to obtain at least one target label that successfully matches the object to be recognized from the multiple labels obtained in real time and the object to be recognized.

[0156] In some embodiments, a multimodal model can be used to extract the label features corresponding to the labels in multiple groups of label combinations and the image feature object feature of the object to be recognized. Among them, the multimodal model is usually a large model. For the convenience of model recognition, the process of using the multimodal model to extract the label features respectively corresponding to the labels in multiple groups of label combinations may include:

[0157] Generate a prompt message for the labels and auxiliary text information in any set of label combinations; the auxiliary text information is used to assist the multimodal model in understanding the labels so that the multimodal model can extract the label features corresponding to the labels.

[0158] Use the multimodal model to extract the label features corresponding to the labels based on the prompt message.

[0159] In this embodiment, in addition to the labels, auxiliary text information will also be added. These auxiliary text information are intended to provide context for the multimodal model to help the model better grasp the context and meaning of the labels. For example, for the label "mountain peak", the auxiliary text information may be "a natural terrain feature towering into the clouds with a sharp top".

[0160] Furthermore, based on the labels and auxiliary text information, generate a "prompt" or "query" form of information. This step is like preparing an accurate question for the multimodal model: for example, "Please extract the features related to'mountain peak' according to the content of this picture in combination with the description 'a natural terrain feature towering into the clouds with a sharp top'." Such prompt information contains both visual elements (through the labels) and in-depth language descriptions (auxiliary text information), providing rich input for the processing of the multimodal model.

[0161] The multimodal model is an AI model that can simultaneously process and integrate multiple types of data (such as images, text, etc.). After receiving the input containing image information and prompt information, it will try to fuse these different modal data to understand and analyze the image content at a deeper level. The model will analyze the image and specifically identify and extract the features that match the labels and descriptions in the prompt information. For example, for the prompt of "mountain peak", the model may look for features such as edge contours and light and shadow changes in the image to confirm and accurately locate the position and shape of the mountain peak. Such feature extraction is not limited to intuitive visual features but may also include abstract semantic understanding to ensure that the extracted features match the deep meaning of the labels.

[0162] In summary, this method of generating prompt information and using the multimodal model to extract label features improves the accuracy and depth of image analysis, enabling the model to better grasp the nuances and context information in complex scenes, and thus playing an important role in fields such as image classification, content generation, and retrieval.

[0163] Of course, as can be seen from the previous description, the label features extracted using the multimodal model can also be pre-grouped and saved, so that for any object to be recognized, based on the saved label features, the matching of the object and the label can be performed without having to re-group and calculate again. As Figure 2 shown in the object processing method, Figure 2The technical solution of the illustrated embodiment can be executed by the server, and the method may include the following steps:

[0164] 201: Obtain the object to be recognized;

[0165] First, obtain the object for which the label needs to be recognized. The object to be recognized can be multimedia content such as a picture, a video, an audio, or a text. For example, on an e-commerce platform, the object to be recognized may be a picture of a product.

[0166] 202: Extract the object features of the object to be recognized;

[0167] Use a multimodal model to analyze and extract the key features of the object to be recognized. The multimodal model can process information in different modalities (such as vision, audition, text, etc.) simultaneously, and can extract corresponding key features for information in different modalities. For example, taking the extraction of visual content (such as a picture) as an example, the multimodal model extracts visual features such as color, shape, texture, and object category. Specifically, for a picture of a piece of clothing, the model may extract features such as "blue", "long sleeves", and "round neck".

[0168] 203: Determine the label features corresponding to the labels in multiple groups of label combinations extracted; the multiple groups of label combinations are obtained by grouping the labels in the multiple obtained labels according to the hierarchical relationship;

[0169] According to Figure 1 the description of the embodiment, the multiple obtained labels are organized into multiple groups of label combinations with a hierarchical relationship. These labels are grouped according to their superordinate and subordinate relationships to generate multiple groups of label combinations. For example, under "clothing" there is "top", and under "top" there are "T-shirt" and "shirt". Further, the model will generate corresponding label features for the labels in the multiple groups of label combinations to facilitate matching with the object features.

[0170] 204: Based on the object features and label features, respectively match the multiple groups of label combinations with the object to be recognized to obtain at least one target label corresponding to the object to be recognized. In this step, based on the extracted object features and label features, the server matches each multiple group of label combinations to find the label combination that best matches the object to be recognized and the labels in this label combination. The matching algorithm will consider the similarity between the label features of the successfully matched labels and the object features, and select the target label that can best describe the object content.

[0171] In some embodiments, the process of matching with the object to be recognized based on the label features and the object features of the object to be recognized to obtain at least one target label of the object to be recognized may include:

[0172] Based on the label features and the object features, match the object to be recognized with the labels in the highest-level label combination to obtain the successfully matched labels; starting from the next-level label combination of the highest-level label combination, determine at least one label in the current-level label combination that has a hyponymy or hypernymy relationship with the successfully matched label in the previous-level label combination; match the at least one label with the object to be recognized to obtain the successfully matched labels; use the at least one successfully matched label corresponding to the multiple groups of label combinations as the at least one target label corresponding to the object to be recognized.

[0173] This embodiment describes a hierarchical label matching method for automatically assigning appropriate labels to an object to be recognized from a series of predefined label sets. This method improves the label matching accuracy by gradually refining from the broadest categories to more specific details.

[0174] Taking the object to be recognized as the image to be recognized as an example, the process will be explained in detail as follows:

[0175] Highest-level matching: First, the server extracts object features from the object to be recognized using a multimodal model. Second, compare these object features with the label features in the highest-level label combination. The highest-level labels are usually the broadest and most abstract classifications, such as "animals", "plants", "vehicles", etc. The matching algorithm will identify the highest-level label that best matches the object features as the successfully matched label.

[0176] Gradual refinement matching: Once a successfully matched label is found in the highest-level label combination, the server will drill down to the next level to find labels that have a hyponymy or hypernymy relationship with the successfully matched label in the previous level. For example, if "animal" is matched at the highest level, the next step may consider more specific classifications such as "mammals", "birds", etc. This step is based on the logic of hyponymy or hypernymy relationships, ensuring the hierarchy and accuracy of the labels.

[0177] Continue the matching process: For each next level, the server will repeat the matching process to find labels that have a hyponymy or hypernymy relationship with the previously successfully matched label in the current level. In this way, the system gradually refines from macro classifications to more specific details, ensuring the matching degree of the selected labels with the object features at each step.

[0178] Determine the target labels: Finally, all the labels that are successfully matched during this step-by-step matching process, regardless of which level they come from, will be aggregated as the "at least one target label" of the image to be recognized. These labels comprehensively reflect the main content and details of the image and can be used for various purposes such as image classification, search, and recommendation.

[0179] In an embodiment of the present application, taking the object to be recognized including an image to be recognized as an example, imagine a simple image recognition scenario where the purpose is to assign appropriate tags to a landscape picture containing "mountains and sunset".

[0180] First, the input tag list (such as the tag list includes: A, B, C, D, E, F, G, H, I, J, K, L, M, 1, 2, 3, 4, 5, 6, 7, 8, 9, etc.) needs to be grouped according to the hierarchical relationship to obtain the following groups:

[0181] First level: A, B, C, D, where C represents "mountains" and D represents "sky";

[0182] Second level: E, F, G, H, I, J, K, L, M, where F and G represent the next-level tags of C mountains, F represents "peaks", and H and I represent the next-level tags of D sky, I represents "sunset";

[0183] Third level: 1, 2, 3, 4, 5, 6, 7, 8, 9, where 8 and 9 represent the next-level tags of I sunset, and 9 represents "evening glow".

[0184] The matching process is as follows:

[0185] First-level matching: Analyze the image to be recognized and find that the image to be recognized contains the image features "mountains" and "sky", and these image features match successfully with the tags C mountains and D sky in the first level. Therefore, the successfully matched tags are obtained: C, D. It should be noted that in the case of successful matching, the next-level matching needs to be continued until the most accurate target tag is found.

[0186] Second-level matching: According to the matching of C mountains, find the tags with a hierarchical relationship with C mountains from the next level, that is, find the tags related to C mountains among F and G, and find that the peak feature in the image matches F peaks. Similarly, according to the matching of D sky, find the tags with a hierarchical relationship with D sky from the next level, among H and I, and identify that the sunset scene in the image matches I sunset. Therefore, the successfully matched tags at the second level are F, I.

[0187] Third-level matching: According to the matching of F peaks, find the tags with a hierarchical relationship with F peaks from the next level, and no successfully matched tags are found; while according to the matching of I sunset, find the tags with a hierarchical relationship with I sunset from the next level as 8 and 9, and find that the evening glow feature above the peaks at sunset in the image matches 9 (evening glow). Therefore, the successfully matched tag at the third level is 9.

[0188] Through the above process, at least one target label sequence finally assigned to the image to be recognized is CD -> FI -> 9, namely "mountain range", "mountain peak and sunset", and "evening glow", which reflects the multi-level characteristics of the image content, from the macroscopic natural landscape to the specific visual elements. By refining the labels step by step, not only the coverage of the recognized labels is improved, but also the accuracy and richness of the description are enhanced.

[0189] In the above embodiment, the label combinations generated by the first level (A, B, C, D) include: {A, B, C, D}. Through the multi-modal model, according to the label features and the object features of the object to be recognized, the matching labels can be determined as C and D. When the object to be recognized includes multiple matching labels, more matching labels can be covered by the feature matching method, thereby improving the coverage of the recognized labels. In addition, when performing the next-level matching labels, the present application only needs to consider the labels having a context relationship with labels C and D, and there is no need to traverse whether the label features of other irrelevant labels match the image features of the image, thereby reducing the computational amount in the label recognition process and improving the efficiency of label recognition.

[0190] In some embodiments, matching the label features corresponding to the labels in the multiple groups of label combinations with the object features of the object to be recognized to obtain at least one target label corresponding to the object to be recognized includes: calculating the feature similarity between the label features corresponding to the labels in the multiple groups of label combinations and the object features of the object to be recognized respectively; determining at least one target label that satisfies the similarity condition and has similar features.

[0191] The object features extracted from the object to be recognized and the label features extracted from the labels in the label combination both need to be converted into a numerical form that can be understood by a computer, usually a vector form. The feature extraction method can be implemented by methods such as word embedding and deep learning feature extraction. Further, by calculating the feature similarity between the label features and the object features, the successfully matched labels are determined. The similarity condition can be, for example, the highest feature similarity or the feature similarity being greater than the similarity threshold. That is, the label with the highest feature similarity or the feature similarity greater than the similarity threshold can be determined as the successfully matched label. Among them, the feature similarity can be calculated, for example, by cosine similarity, Euclidean distance, etc., and the present application does not specifically limit this. The higher the similarity value, the better the matching degree between the label features and the object features. Finally, according to the preset similarity threshold or ranking rule, the labels with similarity higher than the similarity threshold are selected as the successfully matched labels. This may include directly selecting the label with the highest feature similarity, or selecting all or part of the labels on the premise of meeting a certain similarity standard.

[0192] In an embodiment of the present application, assume there is an image recognition task with the goal of tagging an image containing a "mountain bike". The label list provided by the user includes: "outdoor sports", "rock climbing", "transportation means", "mountain bike". Through server-side preprocessing, the determined label combination corresponding to this label list is: First layer: outdoor sports, transportation means; Second layer: mountain bike, rock climbing, where the mountain bike belongs to the next-level label of transportation means, and rock climbing belongs to the next-level label of outdoor sports.

[0193] After obtaining the image features and label features, calculate the feature similarity for "outdoor sports" and "transportation means" in the first layer in sequence. Assume that for the label feature of "outdoor sports", calculate its similarity with the picture feature vector and get 0.2. For the label feature of "transportation means", the similarity is 0.7. Then determine that the successfully matched label is "transportation means". Therefore, it is necessary to further calculate the similarity between the next-level label "mountain bike" of "transportation means" and the image features. For example, for the "mountain bike" label, since this label is more specific and highly relevant to the picture content, the calculated similarity is 0.9. Therefore, "transportation means - mountain bike" is used as the target label for this image.

[0194] Through this process, the system not only identifies the highly relevant label "mountain bike", ensuring the accuracy and pertinence of the label, but also only screens the next-level label "mountain bike" of "transportation means" whose similarity meets the requirements for subsequent matching, without calculating the matching between the next-level label "rock climbing" of "outdoor sports" and the image features, thereby reducing the computational amount in the label recognition process and improving the efficiency of label recognition.

[0195] In a practical application, the object to be recognized can be an image to be recognized. For at least one target label of the obtained image to be recognized, there can be multiple application scenarios.

[0196] As an optional method, the method further includes: determining the image category of the image to be recognized based on at least one target label of the image to be recognized.

[0197] The purpose of determining the image category is: according to at least one target label on the image to be recognized, classify it into one or more predefined image categories. For example, if the target labels are "dog" and "outdoor", the system may classify this picture into the category of "pet outdoor activities". Usually, a trained image classification model is required to achieve this function, and this model can output one or more most likely categories based on the content of the image and the associated label information.

[0198] As another optional method, the method further includes: searching for target images that match the image to be recognized based on the at least one target label.

[0199] The purpose of searching for matching target images is to use at least one target tag as a query condition to find other images from an image database that are similar or related to the content of the image to be identified. This is very useful in scenarios such as image retrieval, content recommendation, and copyright detection. For example, if the target tag is "Eiffel Tower", the system will search for other images containing the Eiffel Tower. Implementing this function usually involves image feature extraction and similarity calculation, as well as efficient image indexing and retrieval algorithms.

[0200] As another optional manner, the method further includes: verifying whether the image to be identified meets image requirements based on the at least one target tag.

[0201] The purpose of verifying whether an image meets the image requirements is to check whether the image to be identified meets specific standards or requirements. These standards may be quality (such as resolution, clarity), content (such as must contain specific elements), or compliance (such as no sensitive content). Depending on the target label, the system may define a set of verification rules or use a machine learning model to evaluate whether the image meets the standards. For example, if the target label is "product display", the system will verify whether the image clearly shows the full picture of the product without occlusion or blur.

[0202] In summary, through these three operations (classification, search, and verification) based on target labels, this method can effectively manage and process image data, which not only improves the degree of automation of image processing, but also enhances the intelligence level of applications, such as playing an important role in content management, search engines, intelligent monitoring and other fields.

[0203] The following is an example of taking the image to be identified as the object to be identified. Figure 3 The scene interaction schematic diagram shown describes the technical solution of the embodiment of the present application.

[0204] In this scenario, the user end 301 and the server end 302 achieve accurate label matching of the "image to be identified" through a series of interactions. The following is a detailed description of the specific interaction process:

[0205] Step 30: The user inputs multiple tags in the user terminal 301.

[0206] Step 31: The client 301 provides the multiple tags to the server 302. Among them, the user of the client 301 (which can be a mobile application, a web interface, or other client applications) manually inputs or selects a series of tags according to their needs or the observed object features, such as "animal", "canine", "golden retriever". These tags represent the user's basic understanding or expected classification of the image to be recognized. In technical implementation, a user interface can be set up to allow the user to input free text in the user interface.

[0207] Step 32: The server 302 can group the tags in the multiple obtained tags according to the hierarchical relationship to obtain multiple groups of tag combinations. Among them, after receiving the multiple tags submitted by the client 301, the server 302 groups them according to the hierarchical relationship. For example, "animal" is the first layer, "canine" is the second layer, and "golden retriever" is the third layer.

[0208] Step 33: The server 302 extracts the tag features corresponding to the tags in the multiple groups of tag combinations respectively. Among them, the server 302 processes the grouped tag combinations and extracts the feature vectors of each tag. This step forms a deep understanding of the tags by analyzing the text description of the tags, relevant images, or other modal data.

[0209] Step 34: The server 302 is also used to obtain the image to be recognized and extract the image features of the image to be recognized. The server 302 also extracts the feature vector of the image to be recognized from these data, and this feature vector reflects the visual or other modal information of the image.

[0210] Among them, the sources of the image to be recognized can at least include: provided by the user through the client 301, automatically crawled by the server 302, integrated with a third-party API, or called from a historical database, etc.

[0211] Specifically, provided by the user through the client 301 can be that the user uploads an image through a client application (such as a mobile application, a web form, etc.). Automatically crawled by the server 302 can be that the server 302 automatically crawls the content on the network, such as pictures and video shares on social media platforms, media content on news websites, etc., for automatic recognition and processing to obtain the image to be recognized. Integrated with a third-party API can be that the server 302 can obtain the image to be recognized by integrating the API of a third-party service, such as downloading a file shared with user authorization from a cloud storage service, or receiving a data stream from a content provider. Called from a historical database can be that the server 302 selects data from its own historical database as the image to be recognized. These data may have been uploaded by the user before or collected through other channels, and this application does not make any limitations on this.

[0212] Step 35: The server 302 matches each of the multiple groups of label combinations with the image to be recognized based on the image features and the label features, so as to obtain at least one target label corresponding to the image to be recognized. In this process, the server 302 uses the previously extracted label features and the features of the image to be recognized for the matching operation. This process may include calculating the similarity between feature vectors and finding out the label set that best matches the image features from the multiple groups of label combinations. Based on the matching result, the server 302 determines at least one target label most relevant to the image to be recognized, such as "Golden Retriever", and feeds this information back to the client 301 (such as displaying "Golden Retriever" on the display interface of the client 301). In this way, the user can obtain accurate classification information about the image to be recognized, meeting their recognition, search, or classification needs.

[0213] Step 36: When the image to be recognized is provided by the user through the client 301, the server 302 may also feed back at least one target label corresponding to the image to be recognized to the client 301.

[0214] It should be noted that the user who provides the image to be recognized and the user who provides multiple labels can be the same user, sharing the same user account. Of course, they can also be different users, etc.

[0215] Of course, the server 302 can also perform other processing on the image to be recognized in combination with at least one target label, such as determining the image category of the image to be recognized, searching for target images matching the image to be recognized, or verifying whether the image to be recognized meets the image requirements, etc., and can feed back the processing results to relevant personnel.

[0216] Figure 4 The following is a schematic structural diagram of an embodiment of a label processing device provided by the present application, as Figure 4 shown. The device includes:

[0217] A first acquisition module 41, configured to acquire multiple labels; group the labels in the multiple labels according to the hierarchical relationship to obtain multiple groups of label combinations;

[0218] A first extraction module 42, configured to extract the label features respectively corresponding to the labels in the multiple groups of label combinations;

[0219] Wherein, the multiple groups of label combinations are used to match with the object to be recognized in descending order of hierarchy based on the label features and the object features of the object to be recognized, so as to obtain at least one target label of the object to be recognized.

[0220] Optionally, in an embodiment of the present application, the first acquisition module 41 is specifically used to identify the level to which the tags among the multiple tags belong; and combine tags at the same level to obtain at least one group of tag combinations.

[0221] Optionally, in an embodiment of the present application, the first acquisition module 41 is also used to query the hypernyms of the label in the multiple labels in the vocabulary database until the root node is reached, and determine the level to which the label belongs based on the path length from the label to the root node.

[0222] Optionally, in the embodiment of the present application, the first acquisition module 41 is further used to determine a synonym set of a tag among the multiple tags in the vocabulary database; and search for hypernyms of the synonym set in the vocabulary database step by step until a root node is reached.

[0223] Optionally, in the embodiment of the present application, the device further includes: a first processing module 43;

[0224] The first processing module 43 is used to filter out tags that meet a filter condition from the multiple tags to update the multiple tags.

[0225] Optionally, in the embodiment of the present application, the first acquisition module 41 is specifically used to acquire a tag list including multiple tags provided by a user;

[0226] The first extraction module 42 is specifically used to extract label features corresponding to the labels in the multiple groups of label combinations by using a multimodal model; the multimodal model is a pre-trained model.

[0227] Figure 4 The label processing device can perform Figure 1 The implementation principle and technical effect of the label processing method described in the illustrated embodiment will not be described in detail. The specific manner in which each module and unit performs operations in the label processing device in the above embodiment has been described in detail in the embodiment of the method, and will not be described in detail here.

[0228] Figure 5 A schematic diagram of the structure of an embodiment of an object processing device provided by the present application is shown in FIG. Figure 5 As shown, the device comprises:

[0229] A second acquisition module 51 is used to acquire an object to be identified;

[0230] A second extraction module 52, used to extract object features of the object to be identified;

[0231] The second determination module 53 is configured to determine the label features corresponding to the labels in the extracted multiple groups of label combinations; the multiple groups of label combinations are obtained by grouping the labels in the obtained multiple labels according to the hierarchical relationship.

[0232] The second matching module 54 is configured to match the multiple groups of label combinations with the object to be recognized respectively based on the object features and the label features, so as to obtain at least one target label corresponding to the object to be recognized.

[0233] Optionally, in the embodiment of the present application, the second matching module 54 is specifically configured to match the object to be recognized with the labels in the highest-level label combination based on the label features and the object features, so as to obtain the successfully matched labels; starting from the next-level label combination of the highest-level label combination, determine at least one label in the current-level label combination that has a superordinate-subordinate relationship with the successfully matched label in the previous-level label combination; match the at least one label with the object to be recognized, so as to obtain the successfully matched labels; use the at least one successfully matched label corresponding to the multiple groups of label combinations as at least one target label corresponding to the object to be recognized.

[0234] Optionally, in the embodiment of the present application, the second matching module 54 is specifically configured to calculate the feature similarity between the label features corresponding to the labels in the multiple groups of label combinations and the object features of the object to be recognized respectively; determine at least one target label that satisfies the similarity condition with similar features.

[0235] Optionally, in the embodiment of the present application, the second extraction module 52 is specifically configured to extract the object features of the object to be recognized by using a multimodal model.

[0236] The second determination module 53 is specifically configured to determine the label features corresponding to the labels in the multiple groups of label combinations extracted by using the multimodal model; the obtained multiple labels include a label list including multiple labels provided by a user.

[0237] Optionally, in the embodiment of the present application, the object to be recognized is an image to be recognized, and the device further includes: a second search module 55 and a second verification module 56.

[0238] The second determination module 53 is further configured to determine the image category of the image to be recognized based on the at least one target label; alternatively, the second search module 55 is configured to search for a target image that matches the image to be recognized based on the at least one target label; or; the second verification module 56 is configured to verify whether the image to be recognized meets the image requirements based on the at least one target label.

[0239] Figure 5 The object processing device described above can execute Figure 2The object processing method described in the illustrated embodiments, its implementation principle and technical effects will not be elaborated further. For the object processing device in the above embodiments, the specific manners in which each module and unit perform operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0240] An embodiment of the present application also provides a computing device, as Figure 6 shown, the computing device may include a storage component 61 and a processing component 62;

[0241] The storage component 61 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component to implement obtaining a plurality of labels; grouping the labels in the plurality of labels according to a hierarchical relationship to obtain a plurality of groups of label combinations; extracting label features respectively corresponding to the labels in the plurality of groups of label combinations; wherein the plurality of groups of label combinations are used to match with a to-be-identified object in a hierarchical order from high to low based on the label features and the object features of the to-be-identified object to obtain at least one target label of the to-be-identified object.

[0242] Of course, the computing device may also necessarily include other components, such as an input / output interface, a display component, a communication component, etc.

[0243] The input / output interface provides an interface between the processing component 62 and a peripheral interface module, and the above peripheral interface module may be an output device, an input device, etc. The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.

[0244] Among them, the processing component 62 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.

[0245] The storage component is configured to store various types of data to support operations on the terminal. The storage component may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disk.

[0246] The display component can be an electroluminescent (EL) element, a liquid crystal display, or a microdisplay with a similar structure, or a retina direct display or a similar laser scanning display.

[0247] It should be noted that when the above computing device implements Figure 1 the label processing method shown or Figure 2 the object processing method shown, it can be a physical device or an elastic computing host provided by a cloud computing platform, etc. It can be implemented as a distributed cluster composed of multiple servers or terminal devices, or can be implemented as a single server or a single terminal device.

[0248] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above Figure 1 label processing method shown or Figure 2 the object processing method shown. This computer-readable medium can be included in the computing device described in the above embodiments; or can exist separately without being assembled into the computing device.

[0249] The embodiment of the present application also provides a computer program product, which includes a computer program carried on a computer-readable storage medium, and when the computer program is executed by a computer, it can implement the label processing method as described above Figure 1 shown or Figure 2 the object processing method shown. In such an embodiment, the computer program can be downloaded and installed from the network and / or installed from a removable medium. When the computer program is executed by a processor, it executes various functions defined in the system of the present application.

[0250] It should be noted that the use of user data may be involved in the embodiments of the present application. In actual applications, it can be used in the solutions described herein within the scope permitted by applicable laws and regulations in compliance with the requirements of applicable laws and regulations in the country where it is located (for example, with the user's explicit consent, giving the user a practical notice, etc.).

[0251] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0252] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0253] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0254] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A label processing method, characterized in that, Comprising: Obtaining a plurality of tags provided by a user; Grouping the tags in the plurality of tags according to a hierarchical relationship to obtain multiple groups of tag combinations; wherein, the tags in the same group of tag combinations belong to the same level; Extracting the tag features respectively corresponding to the tags in the multiple groups of tag combinations; Grouping and storing the tag features; Wherein, the multiple groups of tag combinations are used to match with the object to be recognized in descending order of levels based on the tag features and the object features of the object to be recognized to obtain at least one target tag of the object to be recognized; Wherein, the multiple groups of tag combinations are specifically matched with the object to be recognized in the following manner: Based on the tag features and the object features, matching the object to be recognized with the tags in the highest-level tag combination to obtain a successfully matched tag; Starting from the next-level tag combination of the highest-level tag combination, determining at least one tag in the current-level tag combination that has a hyponymy relationship with the successfully matched tag in the previous-level tag combination; Matching the at least one tag with the object to be recognized to obtain a successfully matched tag; Taking the at least one successfully matched tag corresponding to the multiple groups of tag combinations as at least one target tag corresponding to the object to be recognized.

2. The method according to claim 1, wherein The grouping the tags in the plurality of tags according to a hierarchical relationship to obtain multiple groups of tag combinations includes: Identifying the levels to which the tags in the plurality of tags belong; Combining the tags at the same level to obtain at least one group of tag combinations.

3. The method according to claim 2, characterized in that, The identifying the levels to which the tags in the plurality of tags belong includes: Querying the hypernyms of the tags in the plurality of tags in a lexical database until reaching the root node, and determining the levels to which the tags belong according to the path lengths from the tags to the root node.

4. The method according to claim 3, characterized in that, The querying the hypernyms of the tags in the plurality of tags in a lexical database until reaching the root node includes: Determining the synsets of the tags in the plurality of tags in the lexical database; Gradually searching for the hypernyms of the synsets in the lexical database until reaching the root node.

5. The method according to claim 1, characterized in that After obtaining the plurality of tags, the method further includes: Screening out the tags that meet the screening conditions from the plurality of tags to update the plurality of tags.

6. The method according to claim 1, characterized in that, The obtaining a plurality of tags includes: Obtaining a tag list including a plurality of tags provided by a user; The extracting the tag features respectively corresponding to the tags in the multiple groups of tag combinations includes: Using a multi-modal model to extract the tag features respectively corresponding to the tags in the multiple groups of tag combinations; the multi-modal model is a pre-trained model.

7. A method for object processing, characterized in that, Comprising: Obtaining an object to be recognized; Extracting the object features of the object to be recognized; Determining the tag features corresponding to the tags in the extracted multiple groups of tag combinations; The multiple groups of tag combinations are obtained by grouping the tags in the obtained plurality of tags according to a hierarchical relationship; wherein, the tags in the same group of tag combinations belong to the same level; the plurality of tags are provided by a user; the tag features are grouped and stored; Based on the tag features and the object features, matching the object to be recognized with the tags in the highest-level tag combination to obtain a successfully matched tag; Starting from the next-level tag combination of the highest-level tag combination, determine at least one tag in the current-level tag combination that has a hierarchical relationship with the successfully matched tag in the previous-level tag combination; Match the at least one tag with the object to be recognized to obtain a successfully matched tag; Use the at least one successfully matched tag corresponding to the multiple groups of tag combinations as the at least one target tag corresponding to the object to be recognized.

8. The method according to claim 7, wherein The step of matching the multiple groups of tag combinations with the object to be recognized respectively based on the object features and the tag features to obtain at least one target tag corresponding to the object to be recognized includes: Calculate the feature similarity between the tag features corresponding to the tags in the multiple groups of tag combinations and the object features of the object to be recognized; Determine at least one target tag that satisfies the similarity condition with similar features.

9. The method according to claim 7, wherein The step of extracting the object features of the object to be recognized includes: Use a multimodal model to extract the object features of the object to be recognized; The step of determining the tag features corresponding to the tags in the multiple groups of tag combinations extracted includes: Determine the tag features corresponding to the tags in the multiple groups of tag combinations extracted by using the multimodal model; the multiple tags obtained include a tag list including multiple tags provided by the user.

10. The method according to claim 7, wherein When the object to be recognized is an image to be recognized, the method further includes: Based on the at least one target tag, determine the image category of the image to be recognized; Or; Based on the at least one target tag, search for a target image that matches the image to be recognized; Or; Based on the at least one target tag, verify whether the image to be recognized meets the image requirements.

11. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the tag processing method according to any one of claims 1 to 6, or the object processing method according to any one of claims 7 to 10.

12. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by the processing component, it implements the tag processing method according to any one of claims 1 to 6, or the object processing method according to any one of claims 7 to 10.

13. A computer program product, characterized in that, It includes a computer program / instructions, and when the computer program / instructions are executed by the processing component, it implements the tag processing method according to any one of claims 1 to 6, or the object processing method according to any one of claims 7 to 10.

Citation Information

Patent Citations

  • Text label generation method, model training method, text classification method and related equipment

    CN116127348A