Method for acquiring feature information of category description, image processing method and device
By automatically learning and iteratively updating the feature information of category descriptions, the bias problem caused by manually designed templates is solved, and the accuracy and adaptability of image recognition are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2022-04-29
- Publication Date
- 2026-04-24
AI Technical Summary
In existing technologies, manually designed category description templates are subject to human bias, resulting in low accuracy in image recognition tasks and requiring time-consuming attempts to obtain suitable descriptions from multiple templates.
By automatically learning the feature information of at least two category descriptions corresponding to each category and iteratively updating using the first loss function, more matching category descriptions are generated, thereby improving the accuracy and fit of image recognition tasks.
It improves the accuracy of image recognition tasks and the fit between different images in the same category, reduces human bias, and improves the accuracy of the recognition process.
Smart Images

Figure CN114998643B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a method for obtaining feature information describing categories, an image processing method, and an apparatus. Background Technology
[0002] Recent advances in vision-language models indicate that the features of images and the features of language descriptions within the same category are similar. Therefore, if it is necessary to determine the category of an object in an image from C categories, the category description corresponding to each of the C categories can be obtained. Since the features of the textual category description of the same category are similar to the features of the image, the features of the category description of each category can be used to assist in the recognition of objects in the image.
[0003] The current approach involves manually designing category description templates for each category, and then combining these templates with the category name to obtain the category description. For example, if the manually designed category description template is "This is a XXX" and the category name is "Cat," then the category description would be "This is a cat."
[0004] However, since manually designed category description templates can introduce human bias, they may not be optimal for this image recognition task. Furthermore, in order to obtain a suitable category description template, it is necessary to manually and repeatedly try multiple category description templates, which is time-consuming. Summary of the Invention
[0005] This application provides a method for obtaining feature information of category descriptions, an image processing method, and related devices. It automatically learns feature information of at least two category descriptions corresponding to each category, and the goal of iterative updating includes improving the accuracy of image recognition tasks, which is beneficial for obtaining category descriptions that are more suitable for recognition tasks; it is also beneficial for improving the adaptability with different images in the same category, so as to further improve the accuracy of the image recognition process.
[0006] To address the aforementioned technical problems, this application provides the following technical solutions:
[0007] In a first aspect, embodiments of this application provide a method for obtaining feature information of category descriptions, which can be used in the field of image processing within the field of artificial intelligence. The method includes: a first network device obtaining feature information of K category descriptions corresponding to each of C categories, where C and K are both integers greater than or equal to 2. The category description of the first category includes a category description template and the first category. For example, if the C categories are C different breeds of cats, and K is 3, then the three different category description templates can be "This is an XX", "This is a cat, the specific breed is XX", and "This is a cat of breed XX". If the category name of the first category is "American Shorthair", then the three different category descriptions corresponding to the first category are "This is an American Shorthair", "This is a cat, the specific breed is American Shorthair", and "This is a cat of breed American Shorthair".
[0008] The first network device generates predicted category information for an image based on the feature information of the K category descriptions corresponding to each first category and the feature information of the image. The predicted category of the image pointed to by the predicted category information includes C categories. Specifically, the first network device obtains the similarity between the features of the training image and the high-level features corresponding to each of the C categories in the target feature information set. According to the first loss function, the feature information of the K category descriptions corresponding to each first category is updated until the convergence condition is met. The goal of iterative updating using the first loss function is to improve the similarity between the predicted category information and the correct category information of the image.
[0009] In this implementation, feature information of K category descriptions corresponding to each of the C categories is obtained. Based on the feature information of the K category descriptions corresponding to each first category and the feature information of the image, predicted category information of the image is generated. Based on the correct category information of the image, the predicted category information, and a first loss function, the feature information of the K category descriptions is automatically updated until the convergence condition is met. The goal of iterative updating using the first loss function is to improve the similarity between the predicted category information and the correct category information of the image. Through the aforementioned scheme, the feature information of the K category descriptions corresponding to each category can be automatically learned, and the goal of iterative updating is to improve the accuracy of the image recognition task, which is beneficial to obtaining category descriptions that are more suitable for the recognition task. Since objects of the same category have various variations in different images, the most suitable category descriptions corresponding to different images of the same category may be different. In this scheme, obtaining K category descriptions corresponding to each category is beneficial to improving the fit with different images of the same category, thereby further improving the accuracy of the image recognition process.
[0010] In one possible implementation of the first aspect, the feature information described by the K categories corresponding to the first category includes first feature information and second feature information, and the position of the feature information of the first category in the first feature information is different from the position of the feature information of the first category in the second feature information.
[0011] In this implementation, due to the influence of factors such as pose, deformation and lighting conditions, there is diversity among different images of the same category. The feature information of the category description adapted to different images of the same category may be different. In this scheme, the position of the feature information of the first category in the first feature information and the position of the feature information of the first category in the second feature information are different. This is beneficial to improve the diversity of feature information of at least two categories obtained in the end, so as to improve the adaptability with different images of the same category, and thus improve the accuracy of image recognition results.
[0012] In one possible implementation of the first aspect, the feature information of the K category descriptions includes high-level features of the K category descriptions, where the high-level features of the category descriptions are features generated by the hidden layers / first neural network in the first neural network, and the first neural network is used to update the features of the category descriptions.
[0013] The first network device generates predicted category information for an image based on the feature information of the K category descriptions corresponding to each first category and the feature information of the image. This includes: the first network device can use a target model to model the high-level features of the category descriptions corresponding to the first category based on the high-level features of the K category descriptions corresponding to the first category, so as to determine the distribution information of the high-level features of the category descriptions corresponding to the first category. The target model can be a Gaussian distribution model, a Gaussian mixture distribution model, a von Mises distribution model, or other types of models.
[0014] The first network device performs a sampling operation based on the distribution information of the high-level features described by the category corresponding to each first category, obtaining a feature information set. The feature information set includes high-level features corresponding to each first category. For example, if a Gaussian distribution model is used to model the high-level features described by the category corresponding to the first category, the first network device can perform this sampling operation based on the mean and variance of the K high-level features described by the category corresponding to the first category. At least one high-level feature obtained through sampling follows the distribution of the high-level features described by the category corresponding to the first category. The first network device generates predicted category information for the image based on the image's feature information and the feature information set.
[0015] In this implementation, since the researchers found that the high-level features of multiple category descriptions corresponding to the same category are relatively concentrated, a sampling operation can be performed based on the distribution information of the high-level features of the category descriptions corresponding to each first category to obtain the high-level features corresponding to each first category. Based on the sampled high-level features and the feature information of the image, the predicted category information of the image is generated. Since the predicted features of the image are used to generate the function value of the first loss function, that is, the function value of the first loss function is obtained based on the sampled high-level features, the purpose of iterative update includes reducing the function value of the first loss function. That is, the purpose of iterative update includes obtaining more accurate predicted category information based on the sampled high-level features (that is, the high-level features around the high-level features of the K category descriptions). In other words, a higher update standard is set, which is conducive to obtaining better feature information of category descriptions, thereby improving the accuracy of image recognition.
[0016] In one possible implementation of the first aspect, the first network device acquires feature information of K category descriptions corresponding to each of the C categories, including: the first network device acquires the low-level features of the K category descriptions corresponding to each of the first categories, wherein the low-level features of the category descriptions are vectorized category descriptions; the low-level features of the category descriptions are input into a first neural network, and the low-level features of the category descriptions are updated by the first neural network to obtain the high-level features of the category descriptions.
[0017] The first network device updates the feature information of the K category descriptions corresponding to each first category, including: while keeping the parameters of the first neural network unchanged, the first network device performs gradient updates on the low-level features of the K category description templates corresponding to the C categories according to the function value of the first loss function, so as to obtain the updated low-level features of the K category descriptions corresponding to each first category.
[0018] In this implementation, keeping the parameters of the first neural network unchanged during the gradient update of the underlying features of the category description template helps reduce the number of parameters that need to be updated, thereby facilitating the faster acquisition of appropriate feature information and improving the efficiency of obtaining feature information of the category description.
[0019] In one possible implementation of the first aspect, the first network device updates the feature information of K category descriptions corresponding to each first category according to a first loss function, including: the first network device updates the feature information of K category descriptions corresponding to each first category according to a first loss function and a second loss function, wherein the objective of iterative updating using the second loss function includes reducing the similarity between the feature information of the K category description templates.
[0020] In this implementation, the function value of the second loss function is also used to update the feature information of the K category descriptions corresponding to each first category. The goal of iterative updating using the second loss function is to reduce the similarity between the feature information of at least two category description templates, that is, to increase the distance between the feature information of the K category descriptions corresponding to each first category, that is, to further improve the diversity of the feature information of the K category descriptions corresponding to each first category, so as to improve the adaptability with different images in the same category, thereby improving the accuracy of image recognition results.
[0021] In one possible implementation of the first aspect, the value of the first loss function is greater than or equal to the value of the objective function, where the objective function is the distance between the predicted category information and the correct category information of the image, and the objective of the iterative update includes reducing the value of the first loss function.
[0022] In this implementation, if directly calculating the sum of the function values of countless objective functions is difficult, the function value of the first loss function can be used instead. Since the function value of the first loss function is greater than or equal to the function value of the objective function, the goal of iterative update includes reducing the function value of the first loss function, which means that the goal of iterative update includes reducing the function value of the objective function. The objective function is the distance between the predicted category information and the correct category information of the image, which means reducing the distance between the predicted category information and the correct category information of the image. In other words, while maintaining the correct training objective, a simpler loss function can be used instead, which is beneficial to expanding the flexibility of this solution.
[0023] Secondly, embodiments of this application provide an image processing method applicable to the image processing field within the field of artificial intelligence. The method includes: a second network device extracting features from an image to obtain feature information; generating predicted category information for the image based on feature information corresponding to category descriptions of each of C categories and the image's feature information. The predicted category information points to an image whose predicted category includes C categories, where C is an integer greater than or equal to 2. The category description of the first category includes a category description template and the first category. The feature information corresponding to each of the C categories is obtained based on feature information from at least two category descriptions corresponding to each first category. The feature information from the at least two category descriptions corresponding to each first category is obtained through iterative updating using a first loss function. The goal of iterative updating using the first loss function is to improve the similarity between the predicted category information and the correct category information of the image.
[0024] In the second aspect of this application, the second network device can also be used to perform the steps performed by the first network device in the first aspect and various possible implementations of the first aspect. The feature information described by at least two categories corresponding to each of the first categories can be obtained by the methods in the first aspect and various possible implementations of the first aspect. The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in various possible implementations of the second aspect can all be found in the first aspect, and will not be repeated here.
[0025] Thirdly, embodiments of this application provide a device for acquiring feature information of category descriptions, which can be used in the field of image processing within the field of artificial intelligence. The device includes: an acquisition module, configured to acquire feature information of at least two category descriptions corresponding to each of C categories, where C is an integer greater than or equal to 2, and the category description of the first category includes a category description template and a first category; a generation module, configured to generate predicted category information of an image based on the feature information of the at least two category descriptions corresponding to each first category and the feature information of the image, wherein the predicted category of the image pointed to by the predicted category information includes C categories; and an update module, configured to update the feature information of the at least two category descriptions corresponding to each first category according to a first loss function until a convergence condition is met, wherein the objective of iterative updating using the first loss function includes improving the similarity between the predicted category information and the correct category information of the image.
[0026] In the third aspect of this application, the device for obtaining feature information of category description can also be used to perform the steps performed by the first network device in the first aspect and various possible implementations of the first aspect. The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the various possible implementations of the third aspect can all be found in the first aspect, and will not be repeated here.
[0027] Fourthly, embodiments of this application provide an image processing apparatus that can be used in the field of image processing within the field of artificial intelligence. The apparatus includes: an acquisition module for extracting features from an image to obtain feature information of the image; and a generation module for generating predicted category information of the image based on feature information of category descriptions corresponding to each of the C categories and the feature information of the image. The predicted category of the image pointed to by the predicted category information includes the C categories, where C is an integer greater than or equal to 2. The category description of the first category includes a category description template and the first category. The feature information corresponding to each of the C categories is obtained based on feature information of at least two category descriptions corresponding to each first category. The feature information of at least two category descriptions corresponding to each first category is obtained by iterative updating using a first loss function. The goal of iterative updating using the first loss function is to improve the similarity between the predicted category information and the correct category information of the image.
[0028] In the fourth aspect of this application, the image processing apparatus can also be used to perform the steps performed by the first network device in the second aspect and various possible implementations of the second aspect. The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the various possible implementations of the fourth aspect can be found in the second aspect, and will not be repeated here.
[0029] Fifthly, embodiments of this application provide a computer program product, which includes a program that, when run on a computer, causes the computer to perform the methods described in the first or second aspect above.
[0030] Sixthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods described in the first or second aspect above.
[0031] In a seventh aspect, embodiments of this application provide a network device including a processor and a memory, the processor being coupled to the memory, the memory being used to store a program; and the processor being used to execute the program in the memory, causing the network device to perform the method described in the first aspect above.
[0032] Eighthly, embodiments of this application provide a network device including a processor and a memory, the processor being coupled to the memory, the memory being used to store a program; the processor being used to execute the program in the memory, causing the network device to perform the method described in the second aspect above.
[0033] Ninthly, this application provides a chip system including a processor for supporting network devices in implementing the functions involved in the foregoing aspects, such as transmitting or processing data and / or information involved in the foregoing methods. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the second network device or communication device. This chip system may be composed of chips or may include chips and other discrete devices. Attached Figure Description
[0034] Figure 1 A schematic diagram of the main framework of artificial intelligence provided in the embodiments of this application;
[0035] Figure 2a A system architecture diagram of a system for obtaining feature information of category descriptions provided in embodiments of this application;
[0036] Figure 2b A flowchart illustrating a method for obtaining feature information describing a category provided in an embodiment of this application;
[0037] Figure 3 A flowchart illustrating a method for obtaining feature information describing a category provided in an embodiment of this application;
[0038] Figure 4 A schematic diagram of the feature information of the category description in the method for obtaining feature information of the category description provided in the embodiments of this application;
[0039] Figure 5 A schematic diagram illustrating the feature distribution of high-level features corresponding to multiple categories provided in the embodiments of this application;
[0040] Figure 6 A flowchart illustrating the distribution information of high-level features corresponding to multiple categories for determining the categories described in an embodiment of this application;
[0041] Figure 7 A flowchart illustrating a method for obtaining feature information describing a category provided in an embodiment of this application;
[0042] Figure 8 A schematic flowchart illustrating an image processing method provided in an embodiment of this application;
[0043] Figure 9a A schematic diagram illustrating the beneficial effects of the method for obtaining feature information of category description provided in the embodiments of this application;
[0044] Figure 9b A schematic diagram illustrating the beneficial effects of the method for obtaining feature information of category description provided in the embodiments of this application;
[0045] Figure 10 A schematic diagram of a device for acquiring feature information describing a category provided in an embodiment of this application;
[0046] Figure 11 A schematic diagram of an image processing apparatus provided in an embodiment of this application;
[0047] Figure 12 A schematic diagram of the structure of a network device provided in an embodiment of this application;
[0048] Figure 13 A schematic diagram of the structure of a network device provided in an embodiment of this application.
[0049] Figure 14 This is a schematic diagram of a chip structure provided in an embodiment of this application. Detailed Implementation
[0050] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0051] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0052] First, the overall workflow of the artificial intelligence system is described; please refer to [link / reference]. Figure 1 , Figure 1 The diagram illustrates a structural framework for artificial intelligence (AI). The framework is further elaborated below along two dimensions: the "Intelligent Information Chain" (horizontal axis) and the "IT Value Chain" (vertical axis). The "Intelligent Information Chain" reflects a series of processes from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom." The "IT Value Chain" reflects the value that AI brings to the information technology industry, from the underlying infrastructure of human intelligence and information (provided and processed through technological means) to the industrial ecosystem of the system.
[0053] (1) Infrastructure
[0054] The infrastructure provides computing power to support artificial intelligence systems, enabling communication with the external world and providing support through a basic platform. Communication with the outside world is achieved through sensors; computing power is provided by intelligent chips, which can specifically employ hardware acceleration chips such as central processing units (CPUs), embedded neural network processing units (NPUs), graphics processing units (GPUs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). The basic platform includes distributed computing frameworks and related platform guarantees and support, which may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to acquire data, and this data is provided to intelligent chips in the distributed computing system provided by the basic platform for computation.
[0055] (2) Data
[0056] The data at the next layer of infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, as well as IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0057] (3) Data processing
[0058] Data processing typically includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.
[0059] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training on data by symbolizing and formalizing it.
[0060] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.
[0061] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.
[0062] (4) General ability
[0063] After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0064] (5) Smart Products and Industry Applications
[0065] Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They are the encapsulation of overall artificial intelligence solutions, productizing intelligent information decision-making and realizing practical applications. Their application areas mainly include: intelligent terminals, intelligent manufacturing, intelligent transportation, smart homes, intelligent healthcare, intelligent security, autonomous driving, and smart cities.
[0066] The embodiments of this application can be applied to various application areas in the field of artificial intelligence, specifically for recognizing objects in images in various application areas. If the second network device determines the predicted category of an image from C categories, it can pre-acquire feature information corresponding to the category description of each of the C categories, and then determine the predicted category of the image based on the similarity between the feature information of the category description corresponding to each of the C categories and the feature information of the image.
[0067] In order to automatically learn the feature information of at least two category descriptions corresponding to each of the C categories, embodiments of this application provide a method for obtaining the feature information of category descriptions. Before introducing the aforementioned method, please refer to [the relevant documentation]. Figure 2a , Figure 2a A system architecture diagram for a system for acquiring feature information describing categories, as provided in the embodiments of this application, is shown. Figure 2a In the process, the feature information acquisition 200 for category description includes a first network device 210, a database 220, a second network device 230, and a data storage system 240. The second network device 230 includes a computing module 231.
[0068] The database 220 stores a training data set. The first network device 210 can obtain feature information of at least two category descriptions corresponding to each of the C categories, and use the training data set to iteratively update the feature information of at least two category descriptions corresponding to each first category until the convergence condition is met.
[0069] For details, please refer to Figure 2b , Figure 2bThis is a flowchart illustrating a method for obtaining feature information of category descriptions provided in an embodiment of this application. A1. A first network device obtains feature information of at least two category descriptions corresponding to each of C categories, where C is an integer greater than or equal to 2. The category description of a first category includes a category description template and the first category. A2. The first network device generates predicted category information for an image based on the feature information of the at least two category descriptions corresponding to each first category and the feature information of the image. The predicted category of the image pointed to by the predicted category information includes the C categories. A3. The first network device updates the feature information of the at least two category descriptions corresponding to each first category according to a first loss function until a convergence condition is met. The goal of iterative updating using the first loss function is to improve the similarity between the predicted category information and the correct category information of the image.
[0070] In this embodiment, the feature information of at least two category descriptions corresponding to each category can be automatically learned, and the goal of iterative update includes improving the accuracy of image recognition tasks. This is beneficial for obtaining category descriptions that are more suitable for the recognition task. Since objects of the same category have various changes in different images, the most suitable category descriptions corresponding to different images of the same category may be different. In this solution, obtaining at least two category descriptions corresponding to each category is beneficial for improving the adaptability to different images of the same category, thereby further improving the accuracy of the image recognition process.
[0071] The feature information obtained by the first network device 210, corresponding to the category descriptions of each of the C categories, can be deployed in various forms of second network devices 230. The second network device 230 can access data, code, etc., from the data storage system 240, and can also store data, instructions, etc., in the data storage system 240. The data storage system 240 can be located within the second network device 230, or it can be an external storage device relative to the second network device 230.
[0072] In some embodiments of this application, please refer to Figure 2a The second network device 230 can be configured in the client device, allowing the "user" to directly interact with it. For example, when the client device is a mobile phone or tablet, the second network device 230 can be a module in the host CPU of the phone or tablet used for image recognition. Alternatively, the second network device 230 can be a graphics processing unit (GPU) or neural network processor (NPU) in the phone or tablet, with the GPU or NPU acting as a coprocessor attached to the host processor, and tasks assigned by the host processor.
[0073] It is worth noting that Figure 2a This is merely a schematic diagram of the architecture of two image processing systems provided in this embodiment of the invention. The positional relationships between the devices, components, modules, etc. shown in the diagram do not constitute any limitation. For example, in some other embodiments of this application, the second network device 230 and the client device 250 can be separate independent devices. The second network device 230 interacts with the client device through a configured I / O interface. The "user" can input an image through the I / O interface 212 on the client device, and the second network device 230 returns the predicted category of the image to the client device through the I / O interface, providing it to the user.
[0074] Based on the above description, the specific implementation process of the iterative update stage and application stage of the feature information of the category description provided in the embodiments of this application will be described below.
[0075] I. Iterative Update Stage of Feature Information in Category Description
[0076] For specific details in the embodiments described in this application, please refer to [link / reference]. Figure 3 , Figure 3 This is a flowchart illustrating a method for obtaining feature information of category descriptions provided in an embodiment of this application. The method for obtaining feature information of category descriptions provided in an embodiment of this application may include:
[0077] 301. The first network device acquires feature information corresponding to K category descriptions for each of the C categories.
[0078] In this embodiment of the application, the first network device obtains feature information of K category descriptions corresponding to each of the C categories, where C is an integer greater than or equal to 2 and K is an integer greater than or equal to 2. The feature information of the K category descriptions corresponding to each of the C categories includes the high-level features of the aforementioned K category descriptions and may also include the low-level features of the aforementioned K category descriptions.
[0079] Specifically, for any one of the C categories (hereinafter referred to as "the first category" for ease of description), the first network device can first obtain the underlying features of the K category descriptions corresponding to the first category. The underlying features of the category description are obtained by vectorizing the category description. Each category description of the first category includes a category description template and the name of the first category.
[0080] As an example, if there are C categories representing C different breeds of cats, and K is 3, then the three different category description templates could be "This is an XX", "This is a cat, specifically breed XX", and "This is a cat of breed XX". If the first category is named "American Shorthair", then the three different category descriptions corresponding to the first category would be "This is an American Shorthair", "This is a cat, specifically breed American Shorthair", and "This is a cat of breed American Shorthair". This example is only for the convenience of understanding this scheme and is not intended to limit this scheme.
[0081] The first network device inputs the low-level features of each category description corresponding to the first category into the first neural network. The first neural network updates the low-level features of the category description to obtain the high-level features of each category description corresponding to the first category. The first neural network can be a text encoder. The high-level features of the category description are the hidden layers in the first neural network / the features generated by the first neural network. The hidden layers in the first neural network refer to the outputs of any intermediate layer of the first neural network.
[0082] The first network device performs the above method on each of the C categories to obtain feature information corresponding to K category descriptions for each of the C categories.
[0083] More specifically, regarding the process of "obtaining the underlying features of the K category descriptions corresponding to any first category among the C categories", in one implementation, the first network device can initialize the underlying features of the K category description templates, the first network device performs vectorization processing on the name of the first category to obtain the underlying features of the name of the first category, and the first network device combines the feature information of each category description template in the K category description templates with the feature information of the name of the first category to obtain the underlying features of the K category descriptions corresponding to the first category.
[0084] In another implementation, after obtaining K category description templates, the first network device can vectorize the K category description templates to obtain the underlying features of the K category description templates; the first network device can vectorize the name of the first category to obtain the underlying features of the name of the first category; the first network device can combine the feature information of each category description template in the K category description templates with the feature information of the name of the first category to obtain the underlying features of the K category descriptions corresponding to the first category.
[0085] In one implementation, the first network device can initialize K category description templates, which can also be called prompt templates. Each of the K category description templates can be combined with the class name of the first category to obtain K category descriptions corresponding to the first category. The first network device can then perform vectorization processing on the K category descriptions corresponding to the first category to obtain the underlying features of the K category descriptions corresponding to the first category.
[0086] The first network device can perform the above operation on each of the C categories, resulting in K category descriptions corresponding to each of the C categories. It should be noted that the K category descriptions corresponding to different categories within the C categories can be the same or different.
[0087] Optionally, the underlying features of the K categories corresponding to the same first category include a first underlying feature and a second underlying feature, and the position of the underlying feature of the name of the first category in the first underlying feature is different from the position of the underlying feature of the name of the first category in the second underlying feature; correspondingly, the high-level features of the K categories corresponding to the same first category include a first high-level feature and a second high-level feature, and the position of the high-level feature of the name of the first category in the first high-level feature is different from the position of the high-level feature of the name of the first category in the second high-level feature.
[0088] Furthermore, the K category descriptions may include a first category description and a second category description, and the position of the first category name in the first category description and the position of the first category name in the second category description may be different.
[0089] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic diagram illustrating the feature information of the category description in the method for obtaining feature information of the category description provided in the embodiments of this application. It should be noted that... Figure 4 The text describes how the feature information of category descriptions is displayed visually, as shown in the figure. It should be understood that the position of the category name feature information may differ across different category descriptions. Figure 4 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0090] In this embodiment, due to the influence of factors such as posture, deformation and lighting conditions, there is diversity among different images in the same category. The feature information of the category description adapted to different images in the same category may be different. In this solution, the position of the feature information of the first category in the first feature information and the position of the feature information of the first category in the second feature information are different. This is beneficial to improve the diversity of feature information of at least two categories obtained in the end, so as to improve the adaptability with different images in the same category, and thus improve the accuracy of image recognition results.
[0091] 302. The first network device determines the distribution information of the high-level features described by the category corresponding to the first category based on the high-level features described by at least two categories corresponding to the first category.
[0092] In some embodiments of this application, since the technicians have found in their research that the high-level features of multiple category descriptions corresponding to the same category are relatively concentrated, that is, the high-level features of multiple category descriptions corresponding to the same category are adjacent in the feature space, the first network device can model the distribution of the high-level features of the category descriptions corresponding to each of the first categories in the C categories based on the high-level features of the K category descriptions corresponding to each of the first categories in the C categories, so as to determine the distribution information of the high-level features of the category descriptions corresponding to each of the first categories in the C categories.
[0093] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 5 , Figure 5 This is a schematic diagram illustrating the feature distribution of high-level features corresponding to multiple categories provided in the embodiments of this application. Figure 5 It is obtained by visualizing the high-level features that describe the categories corresponding to multiple categories. Figure 5 Each point in the table represents a high-level feature describing a category, such as... Figure 5 As shown, the high-level features of multiple category descriptions corresponding to the same category are close together, while the high-level features of category descriptions corresponding to different categories are far apart. This should be understood. Figure 5 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0094] For any one of the C categories (hereinafter referred to as "Category 1" for ease of description), the first network device can use a target model to model the high-level features of the category description corresponding to Category 1, in order to determine the distribution information of the high-level features used to describe the category description corresponding to Category 1. The target model can be a Gaussian distribution model, a Gaussian mixture model, a von Mises distribution model, or other types of models, etc., which are not exhaustively listed here. As an example, if a Gaussian model is used to model the high-level features of the category description corresponding to Category 1, the first network device needs to obtain the mean and variance of the K high-level features of the category description corresponding to Category 1 (i.e., an example of the distribution information of the high-level features of the category description corresponding to Category 1). It should be noted that if other models are used for modeling, other types of distribution information need to be obtained, which are not limited here.
[0095] To further understand this scheme, the following is an example of the formulas for calculating the mean and variance of the high-level features described by the K categories corresponding to the first category:
[0096]
[0097]
[0098] Wherein, μ(P) K P represents the mean vector of the high-level features described by the K categories corresponding to the first category. k This represents the k-th category description template among the K category description templates corresponding to the first category, where K is the number of category description templates corresponding to the first category, w 1:C (P k This includes C high-level features corresponding to the k-th category description template, w 1:C (P k Each high-level feature in ) includes a high-level feature of the text description resulting from combining the k-th category description template with the category name of a category.
[0099] Σ(P K (w) represents the covariance matrix of the high-level features described by the K categories corresponding to the first category. 1:C (P k )-μ) T Represents w 1:C (P k The transpose of μ is given. The meanings of the other elements in equation (2) are the same as those of the elements in equation (1) above. They will not be repeated here. It should be understood that equations (1) and (2) are examples of C categories using the same category description template, and are not intended to limit this scheme.
[0100] The first network device performs the above operation on each of the C categories, thereby obtaining the distribution information of the high-level features of the category description corresponding to each of the C categories.
[0101] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 6 , Figure 6 This application provides a flowchart illustrating the distribution information of high-level features for determining category descriptions corresponding to multiple categories, as shown in the embodiments of this application. Figure 6 As shown, the first network device can obtain the low-level features of multiple category description templates, and combine the low-level features of each category description template with the low-level features of a category name to obtain the low-level features of a category description. Figure 6 The image shows three categories: dogs, birds, and cats.
[0102] The first network device inputs the low-level features corresponding to multiple category descriptions for each of the three categories into the first neural network, obtaining the high-level features output by the first neural network corresponding to multiple category descriptions for each of the three categories. Based on the high-level features corresponding to multiple category descriptions for each category, the first network device can obtain the distribution information of the high-level features of the category descriptions for each category. It should be understood that... Figure 6 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0103] 303. The first network device performs a sampling operation based on the distribution information of the high-level features described by the category corresponding to each first category to obtain a target feature information set, which includes the high-level features corresponding to each first category.
[0104] In some embodiments of this application, for any one of the C categories (hereinafter referred to as the "first category" for convenience), the first network device can sample at least one high-level feature based on the distribution information of the high-level features described by the K categories corresponding to the first category. For example, if a Gaussian distribution model is used to model the high-level features described by the category corresponding to the first category, the first network device can perform this sampling operation based on the mean and variance of the high-level features described by the K categories corresponding to the first category, and the sampled at least one high-level feature follows the distribution of the high-level features described by the category corresponding to the first category.
[0105] The first network device performs the aforementioned operation on each of the C categories to obtain at least one set of target feature information. Each set of target feature information includes a high-level feature corresponding to each of the C categories, that is, each set of target feature information includes C high-level features that correspond one-to-one with the C categories.
[0106] It should be noted that the high-level features corresponding to each first category in the target feature information set and the K high-level features corresponding to each first category obtained in step 301 may be different.
[0107] 304. The first network device generates predicted category information for the training image based on the high-level features of the category description corresponding to each first category and the feature information of the training image. The predicted category of the training image pointed to by the predicted category information includes C categories.
[0108] In this embodiment of the application, steps 302 and 303 are optional steps. If steps 302 and 303 are executed, the first network device can obtain the feature information of the training image and generate the predicted category information of the training image according to the target feature information set and the feature information of the training image. The predicted category of the training image pointed to by the predicted category information includes C categories.
[0109] Specifically, the first network device can be configured with multiple training data sets. In one implementation, each training data set may include a training image and the correct category information of that training image. The correct category of the object in the training image is included in C categories, meaning the correct category of the training image pointed to by the correct category information includes C categories. The first network device can input the training image into a second neural network to obtain the features of the training image generated by the second neural network. In another implementation, each training data set may directly include the feature information of a training image and the correct category information of that training image.
[0110] The first network device obtains the similarity between the features of the training image and the high-level features corresponding to each of the C categories in the target feature information set, and generates the predicted category information of the training image based on the aforementioned similarity; wherein, the higher the similarity between the features of the training image and the high-level features corresponding to a certain category (hereinafter referred to as "second category" for convenience), the greater the probability that the predicted category of the training image is the second category.
[0111] In this embodiment, since the technicians found that the high-level features of multiple category descriptions corresponding to the same category are relatively concentrated, a sampling operation can be performed based on the distribution information of the high-level features of the category descriptions corresponding to each first category to obtain the high-level features corresponding to each first category. Based on the high-level features obtained by sampling and the feature information of the image, the predicted category information of the image is generated. Since the predicted features of the image are used to generate the function value of the first loss function, that is, the function value of the first loss function is obtained based on the high-level features obtained by sampling, the purpose of iterative update includes reducing the function value of the first loss function. That is, the purpose of iterative update includes obtaining more accurate predicted category information based on the high-level features sampled (that is, the high-level features around the high-level features of the K category descriptions). That is, a higher update standard is set, which is conducive to obtaining better feature information of category descriptions, and thus conducive to improving the accuracy of image recognition.
[0112] If steps 303 and 304 are not executed, in one implementation, the first network device can also obtain a first feature information set. The first feature information set includes C high-level features corresponding one-to-one with the C categories. Each high-level feature in the first feature information set is obtained by averaging the high-level features describing the K categories corresponding to each first category. Then, based on the first feature information set and the feature information of the training image, predicted category information for the training image is generated. The specific implementation of the aforementioned steps can be found in the above description and will not be repeated here.
[0113] In another implementation, the first network device can also acquire a second feature information set, which includes C high-level features corresponding one-to-one with the C categories. Each high-level feature in the second feature information set is a high-level feature selected from the K high-level features describing each of the first categories. Then, based on the second feature information set and the feature information of the training images, predicted category information for the training images is generated. The specific implementation of the aforementioned steps can be found in the above description and will not be repeated here.
[0114] 305. The first network device generates the function value of the first loss function. The goal of iterative updating using the first loss function includes improving the similarity between the predicted category information of the training image and the correct category information of the training image.
[0115] In this embodiment of the application, the first network device can generate the function value of the first loss function based on the predicted category information and the correct category information of the training image. The goal of iterative updating using the first loss function includes improving the similarity between the predicted category information and the correct category information of the training image.
[0116] Furthermore, in one implementation, the first loss function can directly use the distance between the predicted class information and the correct class information of the training image. This distance can be cosine distance, Euclidean distance, L1 distance, L2 distance, or other types of distance, etc., which will not be exhaustively listed here. Alternatively, the first loss function can directly use the similarity between the predicted class information and the correct class information of the training image. This similarity can be cosine similarity, similarity based on Euclidean distance, or other types of similarity, etc., which will not be exhaustively listed here.
[0117] In another implementation, the objective function is the distance between the predicted category information and the correct category information of the image. The goal of iterative updates using the objective function can include reducing its value. In actual computation, step 304 can be repeated countless times to obtain countless predicted category information for the training image, thereby generating the sum of countless objective function values, which are then used for subsequent update operations. However, since the sum of countless objective function values cannot be calculated, the value of the first loss function can be used instead. The value of the first loss function is greater than or equal to the value of the objective function. The objective function is the distance between the predicted category information and the correct category information of the image. The goal of iterative updates using the first loss function includes reducing its value. Since the value of the first loss function is greater than or equal to the value of the objective function, the goal of iterative updates using the first loss function includes reducing the objective function value, i.e., reducing the distance between the predicted category information and the correct category information of the image, which can also be described as increasing the similarity between the predicted category information and the correct category information of the training image.
[0118] To further understand this scheme, the following example, using a Gaussian distribution model for modeling, reveals an example of the sum of the function values of numerous objective functions and the calculation formula for the first loss function:
[0119]
[0120] Among them, L(P K x represents the sum of the function values of an infinite number of objective functions. i Let y represent the i-th training image. i This represents the correct category information for the i-th training image. Represents the training data (x) i ,y i The expected value of all possibilities is calculated, and log(...) is a logarithmic function. Representative of w 1:C Calculate the mathematical expectation of all possibilities, w 1:C It follows a Gaussian distribution N(μ(P)K ),Σ(P K A random vector, μ(P) K ) and Σ(P K The meaning of ) can be found in the descriptions in equations (1) and (2) above, where e is the natural constant and z is the z-axis. i This represents the feature information of the i-th training image. It is z i transpose, Represents w 1:C In and y i This category corresponds to multiple high-level features, where τ is a predefined hyperparameter. Representing C categories Summing is performed on w c Represents w 1:C The multiple high-level features corresponding to the c-th category should be understood as follows: the examples in Equation (3) are only for the convenience of understanding this scheme and are not intended to limit this scheme.
[0121]
[0122] in, Represents the first loss function, ∑ c ... is the summation symbol, log(...) is the logarithmic function, and e is the natural constant. This represents when the i-th set of training data (x) is executed. i ,y i When ), the mean of the high-level features described by the K categories corresponding to the first category, μ c (P K () represents the mean of the high-level features described by the K categories corresponding to the c-th category out of C categories. Based on multiple A i,j A was obtained. i,j =Σ ii +Σ jj -Σ ij -Σ ji , Σ ii Σ represents the covariance matrix between the i-th category and the i-th category out of C categories. jj Σ represents the covariance matrix between the i-th category and the j-th category among C categories. ij Σ represents the covariance matrix between the i-th category and the j-th category. ji The covariance matrix between the j-th category and the ith category is represented by the following. The meanings of the other elements in Equation (4) can be understood by referring to the above descriptions of Equations (1) to (3). They will not be repeated here. It should be understood that the examples in Equation (4) are only for the convenience of understanding this scheme and are not intended to limit this scheme.
[0123] In this embodiment, if it is difficult to directly calculate the sum of the function values of numerous objective functions, the function value of the first loss function can be used instead. Since the function value of the first loss function is greater than or equal to the function value of the objective function, the goal of iterative update includes reducing the function value of the first loss function, that is, the goal of iterative update includes reducing the function value of the objective function. The objective function is the distance between the predicted category information and the correct category information of the image, that is, reducing the distance between the predicted category information and the correct category information of the image. In other words, while maintaining the correct training objective, a simpler loss function can be used instead, which is beneficial to expanding the flexibility of this solution.
[0124] 306. The first network device generates the function value of the second loss function. The goal of iterative updating using the second loss function is to reduce the similarity between the feature information of the K category description templates.
[0125] In some embodiments of this application, the first network device may also generate the function value of the second loss function. The function value of the second loss function may be obtained based on the similarity between the feature information of any two category description templates among the feature information of K category description templates, or the second loss function may be obtained based on the distance between the feature information of any two category description templates among the feature information of K category description templates. For the description of the concepts of "similarity" and "distance", please refer to the description in step 305, which will not be repeated here.
[0126] To further understand this scheme, the following is an example of the formula for the function value of the second loss function:
[0127]
[0128] Among them, L so This represents the function value of the second loss function, K represents the number of category description templates, and g(P) i ) represents the high-level feature of the i-th category description template among K category description templates, g(P j () represents the high-level feature of the j-th category description template among the K category description templates. <g(P i ),g(P j )> represents g(P i ) and g(P j The cosine distance between ) should be understood as follows: the example in Equation (5) is only for the convenience of understanding this scheme and is not intended to limit this scheme.
[0129] 307. The first network device updates the feature information of the K categories corresponding to each first category based on the function value of the first loss function.
[0130] In this embodiment of the application, step 306 is an optional step. If step 306 is executed, the first network device can perform a weighted summation of the function values of the first loss function and the second loss function to obtain the function value of the total loss function, and update the feature information of the K categories corresponding to each first category according to the function value of the total loss function.
[0131] In this embodiment, the function value of the second loss function is also used to update the feature information of the K category descriptions corresponding to each first category. The goal of iterative updating using the second loss function includes reducing the similarity between the feature information of at least two category description templates, that is, increasing the distance between the feature information of the K category descriptions corresponding to each first category, that is, further improving the diversity of the feature information of the K category descriptions corresponding to each first category, so as to improve the adaptability with different images in the same category, thereby improving the accuracy of image recognition results.
[0132] If step 306 is not executed, the first network device can update the feature information describing the K categories based on the function value of the first loss function. For a more intuitive understanding of this scheme, please refer to [link to relevant documentation]. Figure 7 , Figure 7 This is a flowchart illustrating a method for obtaining feature information describing a category, as provided in an embodiment of this application. Figure 7 The above can be combined with Figure 6 To understand the description, after obtaining the distribution information of the high-level features corresponding to each category, a sampling operation can be performed based on the distribution information of the high-level features corresponding to each category to obtain a high-level feature corresponding to each category. The first network device acquires the feature information of the training image, calculates the similarity between the feature information of the training image and a high-level feature corresponding to each category, and generates the predicted category information of the training image, such as... Figure 7 As shown, the feature information of the training image has the highest similarity to a high-level feature corresponding to the category of "cat". It should be understood that... Figure 7 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0133] Specifically, the first network device can perform gradient differentiation on the function value of the first loss function (or the function value of the total loss function) and backpropagate the differentiation result to update the feature information of the category descriptions corresponding to the C categories.
[0134] Furthermore, since the first network device is configured with a first neural network, a second neural network may also be configured optionally. In one implementation, during the backpropagation of the derivative result, the parameters of the first neural network and / or the second neural network can be updated. In another implementation, during the backpropagation of the derivative result, the first network device can keep the parameters of the first neural network and / or the second neural network unchanged.
[0135] In this embodiment of the application, keeping the parameters of the first neural network unchanged during the gradient update of the underlying features of the category description template is beneficial to reducing the number of parameters that need to be updated, thereby facilitating the faster acquisition of appropriate feature information and improving the efficiency of obtaining feature information of category description.
[0136] More specifically, in one implementation, the first network device can backpropagate the derivative result to directly update the underlying features of the K category description templates. Since the underlying features of each category description corresponding to each first category in the C categories include the underlying features of the category description template, the underlying features of the K category descriptions corresponding to each first category in the C categories are updated, and thus the high-level features of the K category descriptions corresponding to each first category in the C categories are also updated.
[0137] In another implementation, the first network device can backpropagate the derivative result to directly update the underlying features of the K category descriptions corresponding to each of the C categories, thereby also updating the high-level features of the K category descriptions corresponding to each of the C categories.
[0138] The first network device repeatedly performs the above steps to iteratively update the feature information of the K categories corresponding to each first category until the convergence condition is met. The convergence condition may include the convergence condition of the first loss function (optionally, also including the second loss function), or the number of iterations reaches a preset number.
[0139] In this embodiment, feature information of at least two category descriptions corresponding to each of the C categories is obtained. Based on the feature information of the at least two category descriptions corresponding to each first category and the feature information of the image, predicted category information of the image is generated. Based on the correct category information of the image, the predicted category information, and a first loss function, the feature information of the at least two category descriptions is automatically updated until the convergence condition is met. The goal of iterative updating using the first loss function is to improve the similarity between the predicted category information and the correct category information of the image. Through the aforementioned scheme, feature information of at least two category descriptions corresponding to each category can be automatically learned, and the goal of iterative updating is to improve the accuracy of the image recognition task, which is beneficial to obtaining a category description that is more suitable for the recognition task. Since objects of the same category have various changes in different images, the most suitable category description corresponding to different images of the same category may be different. In this scheme, obtaining at least two category descriptions corresponding to each category is beneficial to improving the fit with different images of the same category, thereby further improving the accuracy of the image recognition process.
[0140] II. Application Stage of Feature Information in Category Description
[0141] For specific details in the embodiments described in this application, please refer to [link / reference]. Figure 8 , Figure 8 This is a schematic flowchart illustrating an image processing method provided in an embodiment of this application. The image processing method provided in an embodiment of this application may include:
[0142] 801. The second network device extracts features from the image to obtain the image's feature information.
[0143] In this embodiment of the application, the second network device can input the image to be processed into the second neural network to obtain the feature information of the image generated by the second neural network.
[0144] 802. The second network device generates predicted category information for an image based on the feature information of the category description corresponding to each of the C categories and the feature information of the image. The predicted category of the image pointed to by the predicted category information includes the C categories. The category description of the first category includes a category description template and the first category. The feature information corresponding to each of the C categories is obtained based on the feature information of at least two category descriptions corresponding to each first category. The feature information of at least two category descriptions corresponding to each first category is obtained by iteratively updating using a first loss function. The goal of iteratively updating using the first loss function is to improve the similarity between the predicted category information and the correct category information of the image.
[0145] In this embodiment of the application, the specific implementation of step 802 can be found in [reference needed]. Figure 3The description of step 304 in the corresponding embodiment will not be repeated here. The "feature information corresponding to the category description of each of the C categories" can be the mean, median, or other values of "feature information of at least two category descriptions corresponding to each first category," and is not limited here. It should be noted that... Figure 8 In the corresponding embodiment, "feature information describing at least two categories corresponding to each first category" is based on Figure 3 The corresponding method is used to obtain, Figure 8 The meanings of each term can be found in the reference. Figure 3 The descriptions in the corresponding embodiments will not be repeated here.
[0146] The following section demonstrates the beneficial effects of the method for obtaining feature information in the category descriptions provided in this application, using experimental data. Please refer to [link / reference needed]. Figure 9a , Figure 9a This is a schematic diagram illustrating the beneficial effects of the method for obtaining feature information of category descriptions provided in the embodiments of this application. As shown in the figure, the asterisk group represents combining a manually designed category description template with C categories to obtain a category description; reference group one represents autonomously learning the feature information of a category description; reference group two represents directly performing image classification training, through... Figure 9a As can be seen from the examples, the feature information describing at least two categories obtained by using the scheme provided in the embodiments of this application can achieve the highest accuracy score.
[0147] In this application, the feature information of the same category name may be located in different positions in the feature information of different category descriptions. Please refer to [link / reference]. Figure 9b , Figure 9b This is a schematic diagram illustrating the beneficial effects of the method for obtaining feature information describing the category provided in the embodiments of this application. Figure 9b The results are shown on multiple datasets. Figure 9b Each vertical axis represents a different dataset, as shown in the figure. Compared with manually designed category description templates and reference two groups, the feature information of at least two category descriptions obtained by the scheme provided in this application can achieve the highest accuracy score.
[0148] exist Figures 1 to 8 Based on the corresponding embodiments, in order to better implement the above-described solutions of this application, related equipment for implementing the above solutions is also provided below. See details. Figure 10 , Figure 10This is a schematic diagram of a device for acquiring feature information of category descriptions provided in an embodiment of this application. The device 1000 for acquiring feature information of category descriptions includes: an acquisition module 1001, used to acquire feature information of at least two category descriptions corresponding to each of C categories, where C is an integer greater than or equal to 2, and the category description of the first category includes a category description template and a first category; a generation module 1002, used to generate predicted category information of an image based on the feature information of the at least two category descriptions corresponding to each first category and the feature information of the image, wherein the predicted category of the image pointed to by the predicted category information includes C categories; and an update module 1003, used to update the feature information of the at least two category descriptions corresponding to each first category according to a first loss function until a convergence condition is met, wherein the objective of iterative updating using the first loss function includes improving the similarity between the predicted category information and the correct category information of the image.
[0149] In one possible design, the feature information described by at least two categories corresponding to the first category includes first feature information and second feature information, and the position of the feature information of the first category in the first feature information and the position of the feature information of the first category in the second feature information are different.
[0150] In one possible design, the feature information of at least two category descriptions includes at least two high-level features of the category descriptions, where the high-level features of the category descriptions are features generated by the hidden layers / neural network in the neural network, and the neural network is used to update the features of the category descriptions;
[0151] The generation module 1002 is specifically used for: determining the distribution information of the high-level features of the category description corresponding to the first category based on the high-level features of at least two category descriptions corresponding to the first category; performing a sampling operation based on the distribution information of the high-level features of the category description corresponding to each first category to obtain a feature information set, the feature information set including the high-level features corresponding to each first category; and generating the predicted category information of the image based on the image feature information and the feature information set.
[0152] In one possible design, the acquisition module 1001 is specifically used to: acquire the low-level features of at least two category descriptions corresponding to each first category, wherein the low-level features of the category descriptions are category descriptions in vectorized form; input the low-level features of the category descriptions into a neural network, and update the low-level features of the category descriptions through the neural network to obtain the high-level features of the category descriptions;
[0153] The update module 1003 is specifically used to perform gradient updates on the low-level features of at least two category description templates corresponding to C categories based on the function value of the first loss function while keeping the parameters of the neural network unchanged, so as to obtain the updated low-level features of at least two category descriptions corresponding to each first category.
[0154] In one possible design, the update module 1003 is specifically used to update the feature information of at least two category descriptions corresponding to each first category according to a first loss function and a second loss function. The goal of iterative updating using the second loss function includes reducing the similarity between the feature information of at least two category description templates.
[0155] In one possible design, the value of the first loss function is greater than or equal to the value of the objective function, which is the distance between the predicted category information and the correct category information of the image. The goal of the iterative update includes reducing the value of the first loss function.
[0156] It should be noted that the information interaction and execution process between the modules / units in the feature information acquisition device 1000 for category description are different from those in this application. Figures 3 to 7 The various method embodiments are based on the same concept, and the details can be found in the descriptions of the method embodiments shown above in this application, which will not be repeated here.
[0157] See Figure 11 , Figure 11 This is a schematic diagram of an image processing apparatus provided in an embodiment of this application. The image processing apparatus 1100 includes: an acquisition module 1101, used to extract features from an image to obtain feature information of the image; and a generation module 1102, used to generate predicted category information of the image based on the feature information of the category description corresponding to each of the C categories and the feature information of the image. The predicted category of the image pointed to by the predicted category information includes the C categories, where C is an integer greater than or equal to 2.
[0158] The category description of the first category includes a category description template and the first category. The feature information corresponding to each first category in the C categories is obtained based on the feature information of at least two category descriptions corresponding to each first category. The feature information of at least two category descriptions corresponding to each first category is obtained by iteratively updating using a first loss function. The goal of iteratively updating using the first loss function is to improve the similarity between the predicted category information and the correct category information of the image.
[0159] In one possible design, the feature information described by at least two categories corresponding to the first category includes first feature information and second feature information, and the position of the feature information of the first category in the first feature information and the position of the feature information of the first category in the second feature information are different.
[0160] In one possible design, the feature information of the category descriptions corresponding to the C categories is obtained by iteratively updating using a first loss function and a second loss function, where the second loss function indicates the distance between the feature information of at least two category description templates.
[0161] It should be noted that the information interaction and execution process between the modules / units in the image processing apparatus 1100 are different from those in this application. Figure 8 The various method embodiments are based on the same concept, and the details can be found in the descriptions of the method embodiments shown above in this application, which will not be repeated here.
[0162] The following describes a network device provided in an embodiment of this application. Please refer to [link / reference]. Figure 12 , Figure 12 This is a schematic diagram of a network device provided in an embodiment of this application. Specifically, the network device 1200 includes: a receiver 1201, a transmitter 1202, a processor 1203, and a memory 1204 (wherein the number of processors 1203 in the second network device 1200 can be one or more). Figure 12 (Taking a processor as an example), the processor 1203 may include an application processor 12031 and a communication processor 12032. In some embodiments of this application, the receiver 1201, transmitter 1202, processor 1203, and memory 1204 may be connected via a bus or other means.
[0163] Memory 1204 may include read-only memory and random access memory, and provides instructions and data to processor 1203. A portion of memory 1204 may also include non-volatile random access memory (NVRAM). Memory 1204 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.
[0164] Processor 1203 controls the operation of network devices. In specific applications, the various components of network devices are coupled together through a bus system, which may include not only data buses but also power buses, control buses, and status signal buses. However, for clarity, all buses in the diagram are referred to as the bus system.
[0165] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 1203. The processor 1203 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 1203 or by instructions in software form. The processor 1203 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1203 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 1204. Processor 1203 reads the information in memory 1204 and, in conjunction with its hardware, completes the steps of the above method.
[0166] Receiver 1201 can be used to receive input digital or character information, and to generate signal inputs related to the settings and function control of the second network device. Transmitter 1202 can be used to output digital or character information through the first interface; transmitter 1202 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; transmitter 1202 may also include a display device such as a display screen.
[0167] In this embodiment of the application, the processor 1203 is used to execute... Figures 3 to 7 The image processing method executed by the second network device in the corresponding embodiment. It should be noted that the specific manner in which the application processor 12031 in processor 1203 executes the above steps is different from that in this application. Figures 3 to 7 The various method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 3 to 7 The corresponding method embodiments are the same, and for details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0168] This application also provides a network device; please refer to [link / reference]. Figure 13 , Figure 13 This is a schematic diagram of a network device provided in an embodiment of this application. Specifically, the network device 1300 is implemented by one or more servers. The first network device 1300 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1322 (e.g., one or more processors) and a memory 1332, and one or more storage media 1330 (e.g., one or more mass storage devices) for storing application programs 1342 or data 1344. The memory 1332 and storage media 1330 can be temporary or persistent storage. The program stored in the storage media 1330 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the first network device. Furthermore, the CPU 1322 may be configured to communicate with the storage media 1330 and execute the series of instruction operations in the storage media 1330 on the first network device 1300.
[0169] The first network device 1300 may also include one or more power supplies 1326, one or more wired or wireless network interfaces 1350, one or more input / output interfaces 1358, and / or one or more operating systems 1341, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0170] In this embodiment, the central processing unit 1322 is used to execute... Figure 8 The method for obtaining feature information describing the category executed by the first network device in the corresponding embodiment. It should be noted that the specific manner in which the central processing unit 1322 executes the above steps differs from that in this application. Figure 8 The various method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figure 8 The corresponding method embodiments are the same, and for details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0171] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned actions. Figures 3 to 7 The steps performed by the first network device in the method described in the illustrated embodiment, or causing the computer to perform the steps as described above. Figure 8 The steps performed by the second network device in the method described in the illustrated embodiment.
[0172] This application embodiment also provides a computer-readable storage medium storing a program for performing signal processing, which, when run on a computer, causes the computer to perform the aforementioned actions. Figures 3 to 7 The steps performed by the first network device in the method described in the illustrated embodiment, or causing the computer to perform the steps as described above. Figure 8 The steps performed by the second network device in the method described in the illustrated embodiment.
[0173] The second network device, first network device, second network device, or communication device provided in this application embodiment can specifically be a chip. The chip includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuitry. The processing unit can execute computer execution instructions stored in a storage unit to cause the chip to perform the aforementioned operations. Figures 3 to 7 The illustrated embodiment describes a method for obtaining feature information of a category, or, alternatively, for causing a chip to perform the above-described... Figure 8 The illustrated embodiment describes an image processing method. Optionally, the storage unit is a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).
[0174] For details, please refer to Figure 14 , Figure 14 This is a schematic diagram of a chip provided in an embodiment of this application. The chip can be represented as a neural network processor (NPU) 140. The NPU 140 is mounted as a coprocessor on the host CPU, and tasks are assigned by the host CPU. The core part of the NPU is the arithmetic circuit 140, which is controlled by the controller 1404 to extract matrix data from the memory and perform multiplication operations.
[0175] In some implementations, the arithmetic circuit 1403 internally includes multiple processing engines (PEs). In some implementations, the arithmetic circuit 1403 is a two-dimensional pulsating array. The arithmetic circuit 1403 can also be a one-dimensional pulsating array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1403 is a general-purpose matrix processor.
[0176] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from the weight memory 1402 and caches it in each PE of the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from the input memory 1401 and performs matrix operations with matrix B. The partial result or the final result of the obtained matrix is stored in the accumulator 1408.
[0177] Unified memory 1406 is used to store input and output data. Weight data is directly transferred to weight memory 1402 via Direct Memory Access Controller (DMAC) 1405. Input data is also transferred to unified memory 1406 via DMAC.
[0178] BIU stands for Bus Interface Unit, which is used for interaction between the AXI bus and the DMAC and the Instruction Fetch Buffer (IFB) 1409.
[0179] The Bus Interface Unit (BIU) 1410 is used by the instruction fetch memory 1409 to fetch instructions from external memory, and also by the memory access controller 1405 to fetch the original data of the input matrix A or the weight matrix B from external memory.
[0180] The DMAC is mainly used to move input data from external memory DDR to unified memory 1406, or to weight data to weight memory 1402, or to input data to input memory 1401.
[0181] The vector computation unit 1407 includes multiple arithmetic processing units that further process the output of the computation circuit as needed, such as vector multiplication, vector addition, exponential operations, logarithmic operations, size comparisons, etc. It is mainly used for computation in non-convolutional / fully connected layers of neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.
[0182] In some implementations, the vector computation unit 1407 can store the processed output vector in the unified memory 1406. For example, the vector computation unit 1407 can apply linear and / or nonlinear functions to the output of the computation circuit 1403, such as performing linear interpolation on feature planes extracted by convolutional layers, or accumulating a vector of values to generate activation values. In some implementations, the vector computation unit 1407 generates normalized values, pixel-level summed values, or both. In some implementations, the processed output vector can be used as activation input to the computation circuit 1403, for example, for use in subsequent layers of the neural network.
[0183] The instruction fetch buffer 1409 connected to the controller 1404 is used to store the instructions used by the controller 1404;
[0184] Unified memory 1406, input memory 1401, weighted memory 1402, and instruction fetch memory 1409 are all on-chip memories. External memory is proprietary to this NPU hardware architecture.
[0185] In the above-described method embodiments, the operations of each layer in the first neural network and the second neural network can be performed by the operation circuit 1403 or the vector calculation unit 1407.
[0186] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of a program in the first aspect of the method.
[0187] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0188] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, a first network device, or a network device, etc.) to execute the methods described in the various embodiments of this application.
[0189] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0190] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, first network device, or data center to another website, computer, first network device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a first network device or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
Claims
1. A method for obtaining feature information describing a category, characterized in that, The method includes: Obtain feature information of at least two category descriptions corresponding to each first category in C categories, where C is an integer greater than or equal to 2. The category description of the first category includes a category description template and the first category. The feature information of the category description includes high-level features of the category description. The high-level features of the category description are obtained by updating the low-level features of the category description through a neural network. The low-level features of the category description are obtained by vectorizing the category description. Based on the feature information of at least two categories corresponding to each first category and the feature information of the image, predictive category information of the image is generated, wherein the predicted category of the image pointed to by the predictive category information includes the C categories; According to the first loss function, the feature information of the at least two categories corresponding to each first category is updated until the convergence condition is met. The objective of iterative updating using the first loss function includes improving the similarity between the predicted category information and the correct category information of the image.
2. The method according to claim 1, characterized in that, The feature information described by the at least two categories corresponding to the first category includes first feature information and second feature information, wherein the position of the feature information of the first category in the first feature information is different from the position of the feature information of the first category in the second feature information.
3. The method according to claim 1 or 2, characterized in that, The feature information of the at least two category descriptions includes high-level features of the at least two category descriptions, wherein the high-level features of the category descriptions are hidden layers in the neural network / features generated by the neural network, and the neural network is used to update the features of the category descriptions; The step of generating predicted category information for the image based on feature information describing at least two categories corresponding to each first category and feature information of the image includes: Based on the high-level features of at least two category descriptions corresponding to the first category, determine the distribution information of the high-level features of the category descriptions corresponding to the first category; Based on the distribution information of the high-level features described by the category corresponding to each first category, a sampling operation is performed to obtain a feature information set, the feature information set including the high-level features corresponding to each first category; Based on the image's feature information and the set of feature information, predictive category information for the image is generated.
4. The method according to claim 3, characterized in that, The step of obtaining feature information corresponding to at least two category descriptions for each of the C categories includes: Obtain the underlying features of at least two category descriptions corresponding to each first category, wherein the underlying features of the category descriptions are vectorized forms of the category descriptions; The low-level features of the category description are input into the neural network, and the low-level features of the category description are updated by the neural network to obtain the high-level features of the category description; The updating of feature information describing the at least two categories corresponding to each first category includes: While keeping the parameters of the neural network unchanged, gradient updates are performed on the low-level features of at least two category description templates corresponding to the C categories based on the function value of the first loss function, so as to obtain the updated low-level features of the at least two category descriptions corresponding to each first category.
5. The method according to claim 1 or 2, characterized in that, The step of updating the feature information describing the at least two categories corresponding to each first category according to the first loss function includes: Based on the first loss function and the second loss function, the feature information of the at least two category descriptions corresponding to each first category is updated. The objective of iterative updating using the second loss function includes reducing the similarity between the feature information of the at least two category description templates.
6. The method according to claim 1 or 2, characterized in that, The value of the first loss function is greater than or equal to the value of the objective function, where the objective function is the distance between the predicted category information and the correct category information of the image, and the objective of the iterative update includes reducing the value of the first loss function.
7. An image processing method, characterized in that, The method includes: Feature extraction is performed on the image to obtain its feature information; Based on the feature information of the category description corresponding to each of the C categories and the feature information of the image, the predicted category information of the image is generated. The predicted category of the image pointed to by the predicted category information includes the C categories, where C is an integer greater than or equal to 2. The category description of the first category includes a category description template and the first category. The feature information corresponding to each of the C categories is obtained based on the feature information of at least two category descriptions corresponding to each first category. The feature information of at least two category descriptions corresponding to each first category is obtained by iteratively updating using a first loss function. The goal of iteratively updating using the first loss function is to improve the similarity between the predicted category information and the correct category information of the image. The feature information of the category description includes the high-level features of the category description. The high-level features of the category description are obtained by updating the low-level features of the category description through a neural network. The low-level features of the category description are obtained by vectorizing the category description.
8. The method according to claim 7, characterized in that, The feature information described by the at least two categories corresponding to the first category includes first feature information and second feature information, wherein the position of the feature information of the first category in the first feature information is different from the position of the feature information of the first category in the second feature information.
9. The method according to claim 7 or 8, characterized in that, The feature information of the category description corresponding to the C categories is obtained by iteratively updating using the first loss function and the second loss function, whereby the second loss function indicates the distance between the feature information of at least two category description templates.
10. A device for acquiring feature information describing a category, characterized in that, The device includes: The acquisition module is used to acquire feature information of at least two category descriptions corresponding to each first category in C categories, where C is an integer greater than or equal to 2. The category description of the first category includes a category description template and the first category. The feature information of the category description includes high-level features of the category description. The high-level features of the category description are obtained by updating the low-level features of the category description through a neural network. The low-level features of the category description are obtained by vectorizing the category description. The generation module is configured to generate predicted category information for the image based on feature information describing at least two categories corresponding to each first category and feature information of the image, wherein the predicted category information points to the predicted category of the image, which includes the C categories; The update module is used to update the feature information of at least two categories corresponding to each first category according to a first loss function until a convergence condition is met. The objective of iterative updating using the first loss function includes improving the similarity between the predicted category information and the correct category information of the image.
11. The apparatus according to claim 10, characterized in that, The feature information described by the at least two categories corresponding to the first category includes first feature information and second feature information, wherein the position of the feature information of the first category in the first feature information is different from the position of the feature information of the first category in the second feature information.
12. The apparatus according to claim 10 or 11, characterized in that, The feature information of the at least two category descriptions includes high-level features of the at least two category descriptions, wherein the high-level features of the category descriptions are hidden layers in the neural network / features generated by the neural network, and the neural network is used to update the features of the category descriptions; The generation module is specifically used for: Based on the high-level features of at least two category descriptions corresponding to the first category, determine the distribution information of the high-level features of the category descriptions corresponding to the first category; Based on the distribution information of the high-level features described by the category corresponding to each first category, a sampling operation is performed to obtain a feature information set, the feature information set including the high-level features corresponding to each first category; Based on the image's feature information and the set of feature information, predictive category information for the image is generated.
13. The apparatus according to claim 12, characterized in that, The acquisition module is specifically used for: Obtain the underlying features of at least two category descriptions corresponding to each first category, wherein the underlying features of the category descriptions are vectorized forms of the category descriptions; The low-level features of the category description are input into the neural network, and the low-level features of the category description are updated by the neural network to obtain the high-level features of the category description; The update module is specifically used to perform gradient updates on the low-level features of at least two category description templates corresponding to the C categories, based on the function value of the first loss function, while keeping the parameters of the neural network unchanged, so as to obtain the updated low-level features of the at least two category descriptions corresponding to each first category.
14. The apparatus according to claim 10 or 11, characterized in that, The update module is specifically used to update the feature information of the at least two category descriptions corresponding to each first category according to the first loss function and the second loss function. The goal of iterative updating using the second loss function is to reduce the similarity between the feature information of the at least two category description templates.
15. The apparatus according to claim 10 or 11, characterized in that, The value of the first loss function is greater than or equal to the value of the objective function, where the objective function is the distance between the predicted category information and the correct category information of the image, and the objective of the iterative update includes reducing the value of the first loss function.
16. An image processing apparatus, characterized in that, The device includes: The acquisition module is used to extract features from the image to obtain the feature information of the image; The generation module is used to generate predicted category information of the image based on the feature information of the category description corresponding to each first category in C categories and the feature information of the image. The predicted category of the image pointed to by the predicted category information includes the C categories, where C is an integer greater than or equal to 2. The category description of the first category includes a category description template and the first category. The feature information corresponding to each of the C categories is obtained based on the feature information of at least two category descriptions corresponding to each first category. The feature information of at least two category descriptions corresponding to each first category is obtained by iteratively updating using a first loss function. The goal of iteratively updating using the first loss function is to improve the similarity between the predicted category information and the correct category information of the image. The feature information of the category description includes the high-level features of the category description. The high-level features of the category description are obtained by updating the low-level features of the category description through a neural network. The low-level features of the category description are obtained by vectorizing the category description.
17. The apparatus according to claim 16, characterized in that, The feature information described by the at least two categories corresponding to the first category includes first feature information and second feature information, wherein the position of the feature information of the first category in the first feature information is different from the position of the feature information of the first category in the second feature information.
18. The apparatus according to claim 16 or 17, characterized in that, The feature information of the category description corresponding to the C categories is obtained by iteratively updating using the first loss function and the second loss function, whereby the second loss function indicates the distance between the feature information of at least two category description templates.
19. A computer program product, characterized in that, The computer program product includes a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 6, or causes the computer to perform the method as described in any one of claims 7 to 9.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 6, or causes the computer to perform the method as described in any one of claims 7 to 9.
21. A network device, characterized in that, It includes a processor and a memory, wherein the processor is coupled to the memory. The memory is used to store programs; The processor is configured to execute a program in the memory, causing the network device to perform the method as described in any one of claims 1 to 6.
22. A network device, characterized in that, It includes a processor and a memory, wherein the processor is coupled to the memory. The memory is used to store programs; The processor is configured to execute a program in the memory, causing the network to perform the method as described in any one of claims 7 to 9.
Citation Information
Patent Citations
User state mining and information recommending method, device and equipment
CN108984555A
Method and apparatus for generating image recognition model
CN110288049A