Salient Region Localization Model Training Method and Device
By training a model to identify and emphasize significant regions within product images, the method addresses the challenge of distinguishing fine-grained product categories, improving classification accuracy and reducing annotation requirements.
Patent Information
- Application Number
- CN202310356949.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-04-04
AI Technical Summary
Existing smart identification models face challenges in accurately distinguishing fine-grained product categories in retail settings due to similarities in color, size, pattern, and style among fast-moving consumer goods, and attention networks learned through hyperparameters are not universally applicable across different product series.
A method for training a model that focuses on identifying significant regions within product images, using convolutional kernels and loss functions to enhance feature extraction, allowing for better differentiation of distinct product features.
The proposed method improves the accuracy of product classification by emphasizing significant regions, reducing annotation effort and enhancing the model's ability to distinguish between fine-grained product categories.
Smart Images

Figure CN116363389B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, particularly to the fields of computer vision, deep learning, etc., and can be applied to scenarios such as smart cities. Specifically, it relates to a method and device for training a saliency region localization model, a method and device for identifying a saliency region, an electronic device, a storage medium, and a product. Background Art
[0002] Fast-moving consumer goods digitization and in-depth sales visits are emerging applications in recent years. With the maturity and development of deep learning, an important requirement in retail store digitization is the identification of fast-moving consumer goods. That is, to perform fine-grained categorization of products.
[0003] However, the colors, specifications, sizes, patterns, and styles of product appearances are very similar. Therefore, intelligent recognition models will encounter many difficulties when identifying fine-grained categories. Summary of the Invention
[0004] The present disclosure provides a method and device for training a saliency region localization model, a method and device for identifying a saliency region, an electronic device, a storage medium, and a product.
[0005] According to one aspect of the present disclosure, there is provided a method for training a saliency region localization model, the method including:
[0006] Obtain an image dataset of a product to be classified, and obtain a saliency image of the product to be classified; based on the saliency image, determine the saliency region of each image in the image dataset; extract the saliency features of the saliency region, and train an object detection model based on the saliency features to obtain a saliency region localization model.
[0007] According to a second aspect of the present disclosure, there is provided a method for identifying a saliency region, the method including:
[0008] Obtain an image of a product to be classified; input the image into the saliency region localization model to obtain the saliency region of the image, where the saliency region localization model is trained based on the training method described in the first aspect; identify the saliency region to determine the category of the product to be classified.
[0009] According to a third aspect of the present disclosure, there is provided a fine-grained saliency region localization model training device, the device including:
[0010] An acquisition module, configured to acquire an image dataset of product pairs to be classified, and acquire a saliency image of the product to be classified; A determination module, configured to determine a saliency region of each image in the image dataset based on the saliency image; A training module, configured to extract saliency features of the saliency region, and train an object detection model based on the saliency features to obtain a saliency region localization model.
[0011] According to a fourth aspect of the present disclosure, there is provided a saliency region recognition device, the device includes:
[0012] An acquisition module, configured to acquire an image of a product to be classified; An input module, configured to input the image into a saliency region localization model to obtain a saliency region of the image, where the saliency region localization model is trained based on the training device described in the third aspect; A recognition module, configured to recognize the saliency region and determine the category of the product to be classified.
[0013] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:
[0014] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the first aspect or the second aspect.
[0015] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the first aspect or the second aspect.
[0016] According to a seventh aspect of the present disclosure, there is provided a computer product, including a computer program, where the computer program, when executed by a processor, implements the method described in the first aspect or the second aspect.
[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0019] Figure 1 It shows a schematic flowchart of a method for training a saliency region localization model provided by an embodiment of the present disclosure;
[0020] Figure 2The flowchart shows a method for obtaining significant regions provided by an embodiment of the present disclosure;
[0021] Figure 3 The flowchart shows a method for determining significant regions provided by an embodiment of the present disclosure;
[0022] Figure 4 The flowchart shows a method for learning and training significant regions provided by an embodiment of the present disclosure;
[0023] Figure 5 The flowchart shows a method for learning and training significant regions provided by an embodiment of the present disclosure;
[0024] Figure 6 The flowchart shows a method for learning and training significant regions provided by an embodiment of the present disclosure;
[0025] Figure 7 The flowchart shows a method for identifying significant regions provided by an embodiment of the present disclosure;
[0026] Figure 8 The structural diagram shows a device for training a significant region localization model provided by an embodiment of the present disclosure;
[0027] Figure 9 The structural diagram shows a device for identifying significant regions provided by an embodiment of the present disclosure;
[0028] Figure 10 The schematic block diagram shows an example electronic device that can be used to implement the embodiments of the present disclosure. Detailed implementation manners
[0029] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0030] Fast-moving consumer goods digitization and in-depth sales visits are emerging applications in recent years. With the maturity and development of deep learning, the industry's prospect is to use AI to achieve data on the in-depth distribution of offline channels in retail stores, and obtain the progress of each salesperson's daily in-store inspections and verify the authenticity of the data. One of the more important requirements in retail store digitization is the identification of fast-moving consumer goods. Among a variety of products, there are many fine-grained categories, and the colors, specifications, sizes, patterns, and styles in the appearance of these categories are very similar, making it impossible to identify them quickly. Therefore, the model will encounter many difficulties when identifying fine-grained categories.
[0031] On this basis, related technologies propose adding an attention network to the model for identification. However, the attention network is obtained through hyperparameter learning and cannot be generalized to the application scenarios of different fine-grained products in different series.
[0032] Because the significant regions between different fine-grained products in different series are different, it is not possible to directly learn a general network to increase attention.
[0033] Therefore, the present disclosure proposes a method and device for training a significant region localization model. By learning and weighting according to the characteristics of the positions of significant regions between different fine-grained products, the obtained model can better distinguish the significant regions in different fine-grained products and improve the fine-grained recognition ability of the model.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0035] Figure 1 The flowchart shows a method for training a significant region localization model provided by an embodiment of the present disclosure. As Figure 1 shown, the method may include:
[0036] In step S110, an image dataset of the product to be classified is obtained, and a significant image of the product to be classified is obtained.
[0037] In the embodiments of the present disclosure, the product to be classified may be a fine-grained product. A fine-grained product is a refined division of a product and can be directly distinguished into a specific product. For example, a certain product in a certain series of XX brand. Among them, the product to be classified may be multiple products or a single product.
[0038] Embodiments of the present disclosure can obtain an image dataset of multiple fine-grained products to be distinguished, and further extract partial images from the obtained image dataset. The extracted partial images are used to determine the significant regions of the fine-grained products in the images. The images corresponding to the significant regions are obtained by means such as cropping.
[0039] It should be noted that the images corresponding to the significant regions can also be obtained based on other images, and the present disclosure does not make specific limitations here.
[0040] In step S120, based on the significant images, the significant regions of each image in the image dataset are determined.
[0041] In the embodiments of the present disclosure, the significant images of the products to be classified obtained can be compared and matched with the obtained image dataset, so as to determine the significant regions of the products to be classified in each image in the image dataset. This is convenient for subsequent recognition of the significant regions based on the target detection model to determine the category of the product to be classified in the image.
[0042] In step S130, the significant features of the significant regions are extracted, and a target detection model is trained based on the significant features to obtain a significant region localization model.
[0043] In the embodiments of the present disclosure, after determining the category of the product to be classified corresponding to the image, the significant features of the significant region of the image of the product to be classified can be further extracted, and thus the target detection model is trained based on the significant features, so that the target detection model has the ability to locate the significant region. Thus, the target detection model is realized to identify the category of the product to be classified in the located significant region, improving the accuracy of the recognition of the product to be classified.
[0044] The present disclosure annotates the corresponding significant regions for the obtained image dataset according to the significant images of the products to be classified, so as to reduce the annotation amount and improve the R & D efficiency. Further training the model according to the features of the significant regions can enable the model to have the ability to locate the significant regions, thereby improving the accuracy of identifying the categories of the products to be classified.
[0045] In the embodiments of the present disclosure, the implementation manners of obtaining the significant images of the products to be classified can refer to the following embodiments.
[0046] Figure 2The flowchart of a method for obtaining a salient region provided by an embodiment of the present disclosure is shown. As Figure 2 shown in
[0047] In step S210, a part of the image data is selected from the image dataset.
[0048] In an embodiment of the present disclosure, a part of the images can be selected from the obtained image dataset to label the salient regions of the selected part of the images. For example, if the obtained image dataset includes 10,000 images, 100 images can be selected from these 10,000 images. It should be noted that the 100 images include images of each product to be classified.
[0049] In step S220, the first salient image of the part of the image data is extracted.
[0050] In an embodiment of the present disclosure, the small images of the salient regions included in each image are obtained by cropping etc. according to the selected part of the images, so as to obtain the first salient image. In other words, the first salient image is the image of the salient region of each product to be classified.
[0051] In step S230, the image transformation processing is performed on the first salient image to obtain the salient image of the product to be classified.
[0052] In an embodiment of the present disclosure, the image transformation processing can be one or more of rotation, color jitter, and cropping. By performing one or more of rotation, color jitter (color jitter), and cropping on the first salient image, the sample data of the salient image is increased, and there is no need to re-obtain image data for cropping and other processing processes, saving resources. Among them, color jitter is to randomly transform the exposure, saturation, and hue, so that the model fits the scenes under different illuminations, thereby increasing the richness of the salient image.
[0053] The present disclosure increases the richness of the salient image by enhancing a part of the salient images, can reduce the annotation amount of the salient regions, and improve the annotation efficiency.
[0054] The present disclosure can also use image search, matching, etc. to label the salient regions for each image in the obtained image dataset. The implementation manner is as follows.
[0055] Figure 3 The flowchart of a method for determining a salient region provided by an embodiment of the present disclosure is shown. As Figure 3 shown in
[0056] In step S310, each image in the image dataset is divided according to a preset size to obtain multiple regions of the image.
[0057] In step S320, the saliency image is matched with the multiple regions to determine the regions that match the saliency image.
[0058] In step S330, the regions that match the saliency image are determined as the saliency regions of each image.
[0059] In the embodiments of the present disclosure, for each image of the product to be classified, the image can be divided into multiple regions of the same size. For example, each image is divided into multiple regions of 16*16 size.
[0060] In the present disclosure, the image regions can be searched by means of a sliding window, the searched regions are matched with the saliency image, the positions of the saliency regions in the image are determined, and the saliency regions are marked in the image.
[0061] In the embodiments of the present disclosure, after obtaining the data of the saliency region annotation of the image dataset, a relevant model can be trained based on the image dataset with the saliency regions marked, so as to determine a model that can locate the saliency regions of the images of the products to be classified.
[0062] In the embodiments of the present disclosure, the ppyoloe-plus-large model of the PaddleDetection model can be selected. Among them, the ppyoloe-plus-large model belongs to one of the object detection models (anchor free). The ppyoloe-plus-large model has a fast prediction speed and meets the speed requirements in the business of the present disclosure. Its implementation method is as follows:
[0063] Figure 4 The flowchart of a saliency region learning and training method provided by the embodiments of the present disclosure is shown, as Figure 4 shown in, the method may include:
[0064] In step S410, an object detection model is obtained.
[0065] In step S420, three layers of 1x1 convolutional kernels are added to the object detection model.
[0066] In step S430, saliency features are iteratively learned based on the convolutional kernels to obtain a first object detection model for predicting the saliency features of the products to be classified.
[0067] In step S440, the accuracy of the prediction results of the first object detection model is constrained based on a preset loss function to obtain a saliency region localization model.
[0068] In the embodiments of the present disclosure, a branch can be added to a target detection model (e.g., the ppyoloe - plus - large model). This branch can be a 3 - layer 1×1 convolutional kernel. In the present disclosure, significant features can be learned based on the added 3 - layer 1×1 convolutional kernel to obtain a first target detection model for predicting fine - grained product significant features.
[0069] Among them, the significant features learned by the convolutional kernel can be a mask of the significant region. The initial mask of the significant region is a randomly initialized feature.
[0070] In the present disclosure, the accuracy of the prediction results of the first target detection model can also be constrained based on a loss function, and through multiple iterative trainings, a significant region localization model can be obtained.
[0071] Figure 5 FIG. shows a schematic flowchart of a method for learning and training significant regions provided by the embodiments of the present disclosure. As Figure 5 shown therein, the method may include:
[0072] In step S510, the accuracy of the prediction results of the first target detection model is constrained based on a smooth mean absolute error loss function.
[0073] In step S520, in response to determining that the value of the smooth mean absolute error loss function is less than or equal to a first threshold, the significant features of the product to be classified predicted by the first target detection model are obtained.
[0074] In step S530, the predicted significant features of the product to be classified are weighted based on the significant features to obtain a significant region localization model.
[0075] In the present disclosure, a smooth mean absolute error loss function (smooth L1 loss) can also be selected in the loss function to constrain the accuracy of the prediction results of the first target detection model, so that the first target detection model learns the significant region to achieve the localization of the significant region in the image.
[0076] Further, in the process of training the target detection model based on an image dataset annotated with significant regions, when the value of the smooth L1 loss function is less than or equal to the first threshold, the significant features of the product to be classified predicted by the first target detection model can be obtained, and then the predicted significant features of the product to be classified are weighted based on the significant features. According to multiple iterations, the number of times of updating and iterating the weights increases until a significant region localization model is obtained.
[0077] Figure 6The flowchart shows a method for learning and training significant regions provided by an embodiment of the present disclosure. As shown in Figure 6 as follows, the method may include:
[0078] In step S610, obtain the number of iterations for learning significant features based on a convolutional kernel.
[0079] In step S620, in response to determining that the number of iterations is greater than a second threshold, obtain the value of the smoothed mean absolute error loss function corresponding to the number of iterations greater than the second threshold.
[0080] In step S630, compare the value of the smoothed mean absolute error loss function with the first threshold one by one until it is determined that the value of the smoothed mean absolute error loss function is less than or equal to the first threshold.
[0081] In an embodiment of the present disclosure, the value of the smoothed mean absolute error loss function can be determined by obtaining the number of iterations for learning significant features based on a convolutional kernel. For example, the first ten training processes (epochs) are not applicable. It should be understood that one epoch is equal to the process of training once using all samples in the training set (image dataset with significant regions marked). When a complete image dataset with significant regions marked passes through the neural network once and returns once, that is, a forward propagation and a backward propagation are performed, this process is called one epoch.
[0082] Enhance the significant region attention based on the weighted significant region features, so that the trained object detection model focuses on more discriminative significant feature positions, so that the trained object detection model can locate the significant regions of the product image to be classified, and obtain a fine-grained significant region localization model. The fine-grained significant region localization model can optimize the recognition of the product to be classified.
[0083] The optimization of the present disclosure for the recognition of products to be classified is based on the observation of business data. The corresponding significant region positions of each product to be classified are counted, and a method of bottom library retrieval and search is used to perform coarse-grained annotation of the significant region positions in the picture, reducing the annotation amount. And during the training process, significant region attention weighting is used to enable the trained object detection model to obtain features of higher-weight significant positions, thereby improving the fine-grained recognition ability of the model.
[0084] Based on a similar concept, the present disclosure also provides a method for using a fine-grained significant region localization model.
[0085] Figure 7 The flowchart shows a method for significant region recognition provided by an embodiment of the present disclosure. As shown in Figure 7 as follows, the method may include:
[0086] In step S710, an image of the product to be classified is acquired.
[0087] In the present disclosure, an image of the product to be classified that needs to be distinguished can be acquired by means such as photographing. The image should contain a salient feature for distinguishing the product to be classified.
[0088] In step S720, the image is input into the salient region localization model to obtain the salient region of the image.
[0089] Among them, the salient region localization model is trained based on the training method of the above-mentioned embodiment. In step S730, the salient region is recognized to determine the category of the product to be classified.
[0090] In the present disclosure, the acquired image can be input into a pre-trained fine-grained salient region localization model, and the fine-grained salient region localization model outputs the category to which the product to be classified in the image belongs.
[0091] The fine-grained salient region localization model provided by the present disclosure can optimize the recognition of the product to be classified, improve the recognition ability and accuracy of the product to be classified (fine-grained product), thereby reducing errors and waste of human resources.
[0092] Based on the same principle as the Figure 1 method shown in Figure 8 FIG. shows a schematic structural diagram of a salient region localization model training device provided by an embodiment of the present disclosure. As Figure 8 shown, the salient region localization model training device 800 may include:
[0093] An acquisition module 801, configured to acquire an image dataset of the product to be classified and acquire a salient image of the product to be classified; a determination module 802, configured to determine the salient region of each image in the image dataset based on the salient image; a training module 803, configured to extract the salient features of the salient region and train an object detection model based on the salient features to obtain a fine-grained salient region localization model.
[0094] In an embodiment of the present disclosure, the acquisition module 801 is configured to select a part of the image data in the image dataset; extract a first salient image of the part of the image data; and perform image transformation processing on the first salient image to obtain the salient image of the product to be classified.
[0095] In an embodiment of the present disclosure, the determining module 802 is configured to divide each image in the image dataset according to a preset size to obtain a plurality of regions of the image; match the saliency image with the plurality of regions, and determine the region that matches the saliency image; and determine the region that matches the saliency image as the saliency region of each image.
[0096] In an embodiment of the present disclosure, the training module 803 is configured to obtain a target detection model;
[0097] Add three layers of 1×1 convolutional kernels to the target detection model; iteratively learn the saliency features based on the convolutional kernels to obtain a first target detection model for predicting the saliency features of the product to be classified; and constrain the accuracy of the prediction result of the first target detection model based on a preset loss function to obtain a saliency region localization model.
[0098] In an embodiment of the present disclosure, the training module 803 is configured to constrain the accuracy of the prediction result of the first target detection model based on a smooth mean absolute error loss function; in response to determining that the value of the smooth mean absolute error loss function is less than or equal to a first threshold, obtain the saliency features of the product to be classified predicted by the first target detection model; and weight the predicted saliency features of the product to be classified based on the saliency features to obtain a fine-grained saliency region localization model.
[0099] In an embodiment of the present disclosure, the training module 803 is configured to obtain the number of iterations for learning the saliency features based on the convolutional kernels; in response to determining that the number of iterations is greater than a second threshold, obtain the value of the smooth mean absolute error loss function corresponding to the number of iterations greater than the second threshold; and compare the value of the smooth mean absolute error loss function with the first threshold one by one until it is determined that the value of the smooth mean absolute error loss function is less than or equal to the first threshold.
[0100] Based on the same principle as the Figure 7 method shown in Figure 9 FIG. shows a schematic structural diagram of a saliency region recognition device provided in an embodiment of the present disclosure, as Figure 9 shown, the saliency region recognition device 900 may include:
[0101] An acquisition module 901, configured to acquire an image of a product to be classified; an input module 902, configured to input the image into a saliency region localization model to obtain a saliency region of the image, where the saliency region localization model is trained based on a training device; and a recognition module 903, configured to recognize the saliency region and determine the category of the product to be classified.
[0102] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0103] In an exemplary embodiment, the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described in the above embodiments. The electronic device may be the above-mentioned computer or server.
[0104] In an exemplary embodiment, the readable storage medium may be a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method as described in the above embodiments.
[0105] In an exemplary embodiment, the computer program product includes a computer program which, when executed by a processor, implements the method as described in the above embodiments.
[0106] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0107] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0108] Figure 10 FIG. shows a schematic block diagram of an example electronic device 1000 that may be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0109] As Figure 10As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0110] Multiple components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disc, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0111] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the significant region localization model training method and the significant region recognition method. For example, in some embodiments, the fine-grained significant region localization model training method and the fine-grained significant region recognition method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the significant region localization model training method and the significant region recognition method described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the significant region localization model training method and the significant region recognition method by any other appropriate means (e.g., by means of firmware).
[0112] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0113] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0114] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0115] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0116] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0117] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0118] It should be understood that the various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0119] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for training a saliency region localization model, the method comprising: Obtaining an image dataset of a product to be classified, and obtaining a saliency image of the product to be classified; Based on the saliency image, determining the saliency region of each image in the image dataset; Extracting the saliency features of the saliency region, and training an object detection model based on the saliency features to obtain a saliency region localization model; The training the object detection model based on the saliency features to obtain a saliency region localization model includes: Obtaining an object detection model; Adding three layers of 1x1 convolutional kernels to the object detection model; Iteratively learning the saliency features based on the convolutional kernels to obtain a first object detection model for predicting the saliency features of the product to be classified; Constraining the accuracy of the prediction result of the first object detection model based on the smooth mean absolute error loss function; In response to determining that the value of the smooth mean absolute error loss function is less than or equal to a first threshold, obtaining the saliency features of the product to be classified predicted by the first object detection model; Weighting the predicted saliency features of the product to be classified based on the saliency features to obtain a saliency region localization model.
2. The method according to claim 1, wherein, The obtaining the saliency image of the product to be classified includes: Selecting a part of the image data in the image dataset; Extracting a first saliency image of the part of the image data; Performing image transformation processing on the first saliency image to obtain the saliency image of the product to be classified.
3. The method according to claim 1, wherein The determining the saliency region of each image in the image dataset based on the saliency image includes: Dividing each image in the image dataset according to a preset size to obtain multiple regions of the image; Matching the saliency image with the multiple regions to determine the regions that match the saliency image; Determining the regions that match the saliency image as the saliency regions of each image.
4. The method according to claim 1, wherein, The determining that the value of the smooth mean absolute error loss function is less than or equal to a first threshold includes: Obtaining the number of iterations of learning the saliency features based on the convolutional kernels; In response to determining that the number of iterations is greater than a second threshold, obtaining the value of the smooth mean absolute error loss function corresponding to the number of iterations greater than the second threshold; Comparing the value of the smooth mean absolute error loss function with the first threshold one by one until it is determined that the value of the smooth mean absolute error loss function is less than or equal to the first threshold.
5. A method for identifying a saliency region, the method comprising: Obtaining an image of a product to be classified; Inputting the image into a saliency region localization model to obtain the saliency region of the image, where the saliency region localization model is trained based on the training method according to any one of claims 1-4; Identifying the saliency region and determining the category of the product to be classified.
6. A device for training a saliency region localization model, the device comprising: An obtaining module, configured to obtain an image dataset of a product to be classified, and obtain a saliency image of the product to be classified; A determination module, configured to determine the salient region of each image in the image dataset based on the salient image; A training module, configured to extract the salient features of the salient region and train an object detection model based on the salient features to obtain a salient region localization model; The training module is configured to: Obtain an object detection model; Add three layers of 1x1 convolutional kernels to the object detection model; Iteratively learn the salient features based on the convolutional kernels to obtain a first object detection model for predicting the salient features of the product to be classified; Constrain the accuracy of the prediction result of the first object detection model based on the smooth mean absolute error loss function; In response to determining that the value of the smooth mean absolute error loss function is less than or equal to a first threshold, obtain the salient features of the product to be classified predicted by the first object detection model; Weight the predicted salient features of the product to be classified based on the salient features to obtain a fine-grained salient region localization model.
7. The apparatus according to claim 6, wherein, The obtaining module is configured to: Select part of the image data in the image dataset; Extract the first salient image of the part of the image data; Perform image transformation processing on the first salient image to obtain the salient image of the product to be classified.
8. The apparatus according to claim 6, wherein The determination module is configured to: Divide each image in the image dataset according to a preset size to obtain multiple regions of the image; Match the salient image with the multiple regions to determine the region that matches the salient image; Determine the region that matches the salient image as the salient region of each image.
9. The device according to claim 6, wherein, The determination module is configured to: Obtain the number of iterations for learning the salient features based on the convolutional kernels; In response to determining that the number of iterations is greater than a second threshold, obtain the value of the smooth mean absolute error loss function corresponding to the number of iterations greater than the second threshold; Compare the value of the smooth mean absolute error loss function with the first threshold one by one until it is determined that the value of the smooth mean absolute error loss function is less than or equal to the first threshold.
10. A salient region recognition device, the device includes: An obtaining module, configured to obtain an image of a product to be classified; An input module, configured to input the image into a salient region localization model to obtain the salient region of the image, and the salient region localization model is trained by the training device according to any one of claims 6-9; A recognition module, configured to recognize the salient region and determine the category of the product to be classified.
11. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.
13. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-5.
Citation Information
Patent Citations
Image classification method, device and equipment
CN111476310A
Image recognition model training method and device based on significance detection
CN112329810A