Feature extraction model training, sample retrieval method, device, and computer equipment
By using the initial feature extraction model that includes sample classification network, non-semantic feature extraction network and feature fusion network in feature extraction model training, the problem of time-consuming training of traditional feature extraction models is solved, and efficient feature extraction and sample retrieval is achieved.
Patent Information
- Application Number
- CN202111247520.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-10-26
AI Technical Summary
Traditional feature extraction models take time to train, and different models need to be trained separately to meet the needs of different features.
A feature extraction model training method is adopted to obtain a sample group containing the target sample, reference sample and category label, and input the samples into the initial feature extraction model, including sample classification network, non-semantic feature extraction network and feature fusion network. The initial features are output through these networks, and the feature loss and classification loss are calculated, the model parameters are adjusted until the convergence conditions are met, and the target feature extraction model is obtained.
It improves the efficiency of feature extraction model training, can learn semantic features, non-semantic features and fusion features at the same time, outputs diverse features, and is suitable for sample retrieval and other applications, improving the accuracy and efficiency of retrieval.
Smart Images

Figure CN114358109B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and particularly to a method, device, computer device, and storage medium for training a feature extraction model and retrieving samples. Background Art
[0002] With the development of computer technologies, retrieval technologies have emerged, such as image retrieval technology. Retrieval technology extracts features of a query sample and matches the features of the query sample with the features of samples stored in a sample retrieval library, so as to retrieve samples similar to the query sample in the sample retrieval library.
[0003] In traditional technologies, samples are usually input into a model for feature extraction, and one model is used to extract one type of feature. In this way, different models need to be trained separately for different features, which is time-consuming for training. Summary of the Invention
[0004] Based on this, to address the above technical problems, it is necessary to provide a method, device, computer device, and storage medium for training a feature extraction model and retrieving samples, which can improve the training efficiency.
[0005] A method for training a feature extraction model, the method comprising:
[0006] Obtain a first sample group, and input each sample in the first sample group into an initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and a class label corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network;
[0007] Output an initial classification feature and an initial semantic feature through the sample classification network, and output an initial non-semantic feature through the non-semantic feature extraction network;
[0008] Through the feature fusion network, fuse the initial semantic feature and the initial non-semantic feature of the same sample to obtain an initial fusion feature corresponding to each sample;
[0009] Calculate a loss based on the initial semantic feature, initial non-semantic feature, and initial fusion feature corresponding to the target sample and the reference sample to obtain a feature loss;
[0010] Calculate a loss based on the initial classification feature and the class label corresponding to the same sample to obtain a classification loss;
[0011] Based on the feature loss and the classification loss, adjust the model parameters of the initial feature extraction model until a convergence condition is met to obtain a target feature extraction model; the target feature extraction model is used to extract sample features of input samples, and the sample features are used for sample retrieval.
[0012] In one embodiment, the feature sizes corresponding to the initial semantic features, the initial non-semantic features, and the initial fusion features are the same.
[0013] A feature extraction model training device, the device includes:
[0014] A first sample group processing module, configured to obtain a first sample group and input each sample in the first sample group into an initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and a class label corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network;
[0015] A feature output module, configured to output an initial classification feature and an initial semantic feature through the sample classification network, and output an initial non-semantic feature through the non-semantic feature extraction network;
[0016] A feature fusion module, configured to fuse the initial semantic feature and the initial non-semantic feature of the same sample through the feature fusion network to obtain an initial fusion feature corresponding to each sample;
[0017] A feature loss determination module, configured to calculate a loss based on the initial semantic features, the initial non-semantic features, and the initial fusion features corresponding to the target sample and the reference sample to obtain a feature loss;
[0018] A classification loss determination module, configured to calculate a loss based on the initial classification feature and the class label corresponding to the same sample to obtain a classification loss;
[0019] A model parameter adjustment module, configured to adjust the model parameters of the initial feature extraction model based on the feature loss and the classification loss until a convergence condition is met to obtain a target feature extraction model; the target feature extraction model is used to extract the sample features of the input sample, and the sample features are used for sample retrieval.
[0020] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0021] Obtain a first sample group and input each sample in the first sample group into an initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and a class label corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network;
[0022] Output the initial classification features and initial semantic features through the sample classification network, and output the initial non-semantic features through the non-semantic feature extraction network;
[0023] Through the feature fusion network, fuse the initial semantic features and initial non-semantic features of the same sample to obtain the initial fusion features corresponding to each sample;
[0024] Calculate the loss based on the initial semantic features, initial non-semantic features and initial fusion features corresponding to the target sample and the reference sample to obtain the feature loss;
[0025] Calculate the loss based on the initial classification features and class labels corresponding to the same sample to obtain the classification loss;
[0026] Based on the feature loss and the classification loss, adjust the model parameters of the initial feature extraction model until the convergence condition is met to obtain the target feature extraction model; the target feature extraction model is used to extract the sample features of the input sample, and the sample features are used for sample retrieval.
[0027] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0028] Obtain a first sample group, and input each sample in the first sample group into the initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and class labels corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network and a feature fusion network;
[0029] Output the initial classification features and initial semantic features through the sample classification network, and output the initial non-semantic features through the non-semantic feature extraction network;
[0030] Through the feature fusion network, fuse the initial semantic features and initial non-semantic features of the same sample to obtain the initial fusion features corresponding to each sample;
[0031] Calculate the loss based on the initial semantic features, initial non-semantic features and initial fusion features corresponding to the target sample and the reference sample to obtain the feature loss;
[0032] Calculate the loss based on the initial classification features and class labels corresponding to the same sample to obtain the classification loss;
[0033] Based on the feature loss and the classification loss, adjust the model parameters of the initial feature extraction model until the convergence condition is met to obtain the target feature extraction model; the target feature extraction model is used to extract the sample features of the input sample, and the sample features are used for sample retrieval.
[0034] A sample retrieval method, the method comprising:
[0035] Obtaining a query sample and a set of candidate recall samples;
[0036] Inputting the candidate recall samples in the query sample and the set of candidate recall samples into a target feature extraction model to obtain a query sample feature corresponding to the query sample and a recall sample feature corresponding to the candidate recall samples;
[0037] Determining a retrieval result sample corresponding to the query sample from the set of candidate recall samples based on the query sample feature and the recall sample feature;
[0038] The training process of the target feature extraction model is as follows:
[0039] Obtaining a first sample group, and inputting each sample in the first sample group into an initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and a class label corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network;
[0040] Outputting an initial classification feature and an initial semantic feature through the sample classification network, outputting an initial non-semantic feature through the non-semantic feature extraction network, and fusing the initial semantic feature and the initial non-semantic feature of the same sample through the feature fusion network to obtain an initial fusion feature corresponding to each sample;
[0041] Calculating a loss based on the initial semantic feature, the initial non-semantic feature, and the initial fusion feature corresponding to the target sample and the reference sample to obtain a feature loss, and calculating a loss based on the initial classification feature and the class label corresponding to the same sample to obtain a classification loss;
[0042] Adjusting the model parameters of the initial feature extraction model based on the feature loss and the classification loss until a convergence condition is met to obtain a target feature extraction model.
[0043] A sample retrieval device, the device comprising:
[0044] A data acquisition module, configured to obtain a query sample and a set of candidate recall samples;
[0045] A data processing module, configured to input the candidate recall samples in the query sample and the set of candidate recall samples into a target feature extraction model to obtain a query sample feature corresponding to the query sample and a recall sample feature corresponding to the candidate recall samples;
[0046] A retrieval result determination module, configured to determine a retrieval result sample corresponding to the query sample from the candidate recall sample set based on the query sample feature and the recall sample feature;
[0047] The training process of the target feature extraction model is as follows:
[0048] Obtain a first sample group, and input each sample in the first sample group into an initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and a class label corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network;
[0049] Output an initial classification feature and an initial semantic feature through the sample classification network, output an initial non-semantic feature through the non-semantic feature extraction network, and fuse the initial semantic feature and the initial non-semantic feature of the same sample through the feature fusion network to obtain an initial fusion feature corresponding to each sample;
[0050] Calculate a loss based on the initial semantic feature, the initial non-semantic feature, and the initial fusion feature corresponding to the target sample and the reference sample to obtain a feature loss, and calculate a loss based on the initial classification feature and the class label corresponding to the same sample to obtain a classification loss;
[0051] Adjust the model parameters of the initial feature extraction model based on the feature loss and the classification loss until the convergence condition is met to obtain a target feature extraction model.
[0052] A computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0053] Obtain a query sample and a candidate recall sample set;
[0054] Input the query sample and the candidate recall samples in the candidate recall sample set into the target feature extraction model to obtain a query sample feature corresponding to the query sample and a recall sample feature corresponding to the candidate recall samples;
[0055] Determine a retrieval result sample corresponding to the query sample from the candidate recall sample set based on the query sample feature and the recall sample feature;
[0056] The training process of the target feature extraction model is as follows:
[0057] Obtain a first sample group, and input each sample in the first sample group into an initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and a class label corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network;
[0058] Output an initial classification feature and an initial semantic feature through the sample classification network, output an initial non-semantic feature through the non-semantic feature extraction network, and fuse the initial semantic feature and the initial non-semantic feature of the same sample through the feature fusion network to obtain an initial fusion feature corresponding to each sample;
[0059] Calculate a loss based on the initial semantic feature, initial non-semantic feature, and initial fusion feature corresponding to the target sample and the reference sample to obtain a feature loss, and calculate a loss based on the initial classification feature and the class label corresponding to the same sample to obtain a classification loss;
[0060] Based on the feature loss and the classification loss, adjust the model parameters of the initial feature extraction model until the convergence condition is satisfied to obtain a target feature extraction model.
[0061] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the following steps are implemented:
[0062] Obtain a query sample and a candidate recall sample set;
[0063] Input the query sample and the candidate recall samples in the candidate recall sample set into the target feature extraction model to obtain a query sample feature corresponding to the query sample and a recall sample feature corresponding to the candidate recall samples;
[0064] Based on the query sample feature and the recall sample feature, determine a retrieval result sample corresponding to the query sample from the candidate recall sample set;
[0065] The training process of the target feature extraction model is as follows:
[0066] Obtain a first sample group, and input each sample in the first sample group into an initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and a class label corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network;
[0067] Output the initial classification features and initial semantic features through the sample classification network, output the initial non-semantic features through the non-semantic feature extraction network, and fuse the initial semantic features and initial non-semantic features of the same sample through the feature fusion network to obtain the initial fusion features corresponding to each sample respectively;
[0068] Calculate the loss based on the initial semantic features, initial non-semantic features and initial fusion features corresponding to the target sample and the reference sample to obtain the feature loss, and calculate the loss based on the initial classification features and class labels corresponding to the same sample to obtain the classification loss;
[0069] Based on the feature loss and the classification loss, adjust the model parameters of the initial feature extraction model until the convergence condition is met to obtain the target feature extraction model.
[0070] The above feature extraction model training, sample retrieval method, device, computer device and storage medium obtain the first sample group and input each sample in the first sample group into the initial feature extraction model; the first sample group includes the target sample, the reference sample corresponding to the target sample and the class labels corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network and a feature fusion network; output the initial classification features and initial semantic features through the sample classification network, and output the initial non-semantic features through the non-semantic feature extraction network; fuse the initial semantic features and initial non-semantic features of the same sample through the feature fusion network to obtain the initial fusion features corresponding to each sample respectively; calculate the loss based on the initial semantic features, initial non-semantic features and initial fusion features corresponding to the target sample and the reference sample to obtain the feature loss; calculate the loss based on the initial classification features and class labels corresponding to the same sample to obtain the classification loss; based on the feature loss and the classification loss, adjust the model parameters of the initial feature extraction model until the convergence condition is met to obtain the target feature extraction model. In this way, a unified model is established to learn semantic features and non-semantic features, and at the same time learn the fusion features. The fusion features contain both semantic information and non-semantic information. Finally, the trained model can not only output semantic features and non-semantic features containing single-dimensional information, but also output fusion features containing two-dimensional information. Only one model needs to be trained to make the model output diverse features, improving the training efficiency. Subsequently, when performing sample retrieval, the candidate recall samples in the candidate recall sample set and the query sample can be input into the target feature extraction model to obtain the query sample features corresponding to the query sample and the recall sample features corresponding to the candidate recall samples. Based on the query sample features and the recall sample features, determine the retrieval result samples corresponding to the query sample from the candidate recall sample set. In this way, sample retrieval can be performed through the diverse features output by the model, improving the accuracy and efficiency of sample retrieval. Description of the Drawings
[0071] Figure 1 It is an application environment diagram of a feature extraction model training method and a sample retrieval method in an embodiment;
[0072] Figure 2 It is a schematic flowchart of a feature extraction model training method in an embodiment;
[0073] Figure 3 It is a schematic diagram of a model structure in an embodiment;
[0074] Figure 4 It is a schematic flowchart of training a feature extraction model in an embodiment;
[0075] Figure 5 It is a schematic flowchart of training a feature extraction model in another embodiment;
[0076] Figure 6 It is a schematic flowchart of a sample retrieval method in an embodiment;
[0077] Figure 7 It is a schematic flowchart of training an image feature extraction model in an embodiment;
[0078] Figure 8 It is a schematic flowchart of image retrieval in an embodiment;
[0079] Figure 9 It is a structural block diagram of a feature extraction model training device in an embodiment;
[0080] Figure 10 It is a structural block diagram of a sample retrieval device in an embodiment;
[0081] Figure 11 It is an internal structure diagram of a computer device in an embodiment;
[0082] Figure 12 It is an internal structure diagram of a computer device in an embodiment. Specific Embodiments
[0083] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0084] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0085] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0086] Computer Vision (CV) is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as target recognition and measurement in machine vision, and further performing image processing to make the computer-processed images more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0087] The key technologies of speech technology include automatic speech recognition technology, speech synthesis technology, and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.
[0088] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0089] The solution provided in the embodiments of this application involves technologies such as computer vision, speech, and machine learning in artificial intelligence, and will be specifically described through the following embodiments:
[0090] The feature extraction model training method and sample retrieval method provided in this application can be applied to an application environment as shown in Figure 1 In the figure. Among them, the terminal 102 communicates with the server 104 through the network. The terminal 102 can be, but is not limited to, a laptop computer, a smart phone, a tablet computer, a desktop computer, a smart TV, a vehicle-mounted terminal, and a portable wearable device. An application program is installed on the terminal. The application program can refer to a client installed in the terminal. A client (also known as an application client, APP client) refers to a program installed and running in the terminal; the application program can also refer to a non-installable application program, that is, an application program that can be used without downloading and installation. Such application programs are commonly known as applets and usually run as subroutines in the client; the application program can also refer to a web application program opened through a browser; and so on. The above various application programs are classified according to the application functions they provide. The types of application programs can include, but are not limited to: search application programs, instant messaging application programs, payment application programs, audio and video application programs, etc. The server 104 can be implemented by an independent server or a server cluster or cloud server composed of multiple servers.
[0091] Both the terminal 102 and the server 104 can be independently used to execute the feature extraction model training and sample retrieval methods provided in the embodiments of this application.
[0092] For example, a terminal obtains a first sample group, where the first sample group includes a target sample, a reference sample corresponding to the target sample, and class labels corresponding to each sample. The terminal inputs each sample in the first sample group into an initial feature extraction model, where the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network. An initial classification feature and an initial semantic feature are output through the sample classification network, an initial non-semantic feature is output through the non-semantic feature extraction network, and the initial semantic feature and the initial non-semantic feature of the same sample are fused through the feature fusion network to obtain an initial fusion feature corresponding to each sample. The terminal calculates a feature loss based on the initial semantic feature, the initial non-semantic feature, and the initial fusion feature corresponding to the target sample and the reference sample, and calculates a classification loss based on the initial classification feature and the class label corresponding to the same sample. The terminal adjusts the model parameters of the initial feature extraction model based on the feature loss and the classification loss until the convergence condition is satisfied, and obtains a target feature extraction model.
[0093] The terminal obtains a query sample and a candidate recall sample set, inputs the query sample and the candidate recall samples in the candidate recall sample set into the target feature extraction model, and obtains a query sample feature corresponding to the query sample and a recall sample feature corresponding to the candidate recall samples. The terminal determines a retrieval result sample corresponding to the query sample from the candidate recall sample set based on the query sample feature and the recall sample feature.
[0094] The terminal 102 and the server 104 can also cooperate to execute the feature extraction model training and sample retrieval methods provided in the embodiments of the present application.
[0095] For example, the server obtains a first sample group from the terminal, trains the initial feature extraction model based on the first sample group to obtain a target feature extraction model. The server obtains a query sample from the terminal, obtains a candidate recall sample set from the database, performs sample retrieval based on the target feature extraction model, and determines a retrieval result sample corresponding to the query sample from the candidate recall sample set. The server sends the retrieval result sample to the terminal.
[0096] The embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, sample retrieval (information search and information recommendation), etc.
[0097] In one embodiment, as Figure 2 shown, a feature extraction model training method is provided. Taking the example that this method is executed by a computer device, it can be understood that the computer device can be Figure 1 the terminal 102 shown, or the server 104. In this embodiment, the feature extraction model training method includes the following steps:
[0098] Step S202: Obtain a first sample group and input each sample in the first sample group into an initial feature extraction model. The first sample group includes a target sample, a reference sample corresponding to the target sample, and class labels corresponding to each sample. The initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network.
[0099] Among them, a sample refers to an object for transmitting information and presenting information, and can specifically be at least one of an image, speech, and text. There is a corresponding class label for the sample, and the class label is used to identify the class of the sample. For example, if the sample is an image, the class label corresponding to the sample can be animal breeds such as dogs, cats, and fish, plant breeds such as corals, pine trees, and osmanthus, and object types such as magnifying glasses, cabinets, and water bottles. The first sample group includes a target sample, a reference sample corresponding to the target sample, and class labels corresponding to each sample respectively. The first sample group is the training sample of the model and is used to train the initial feature extraction model. The reference sample corresponding to the target sample includes at least one of a positive sample and a negative sample corresponding to the target sample. The sample similarity between the target sample and the corresponding positive sample is greater than the sample similarity between the target sample and the corresponding negative sample. The first sample group can be at least one.
[0100] The feature extraction model is a machine learning or deep learning model for extracting sample features of the input sample. The input data of the feature extraction model is the sample, and the output data is the sample features corresponding to the sample. Different feature extraction models can be trained for different types of samples. For example, if the sample is an image-type sample, then an image feature extraction model can be trained; if the sample is a voice-type sample, then a voice feature extraction model can be trained.
[0101] The initial feature extraction model refers to the feature extraction model to be trained. The initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network. The sample classification network is used to classify the input data and can extract the semantic features and classification features of the input data. The input data of the sample classification network is the sample, and the output data is the classification features and semantic features of the sample. The non-semantic feature extraction network is used to extract the non-semantic features of the input data. The input data of the non-semantic feature extraction network is the sample, and the output data is the non-semantic features of the sample. The feature fusion network is used to fuse the semantic features and non-semantic features. The input data of the feature fusion network is the semantic features and non-semantic features, and the output data is the fused features. All kinds of features can be represented in the form of vectors or matrices.
[0102] Specifically, the computer device can obtain the first sample group locally, or from a terminal or a server, and train the initial feature extraction model based on the first sample group to obtain a target feature extraction model.
[0103] Step S204: Output the initial classification features and initial semantic features through the sample classification network, and output the initial non-semantic features through the non-semantic feature extraction network.
[0104] Among them, non-semantic features refer to features that do not have the ability of semantic measurement, and semantic features refer to features that have the ability of semantic measurement. Classification features are used to represent the features of sample categories. Based on the classification features, the predicted label of the sample, that is, the predicted category, can be determined. It can be understood that semantic features are output by the sample classification network. The sample classification network not only needs to extract features, but also needs to classify the features and predict the categories. That is, when the sample classification network extracts features, it not only considers the specific content of the sample but also considers the category to which the sample belongs. Therefore, semantic features have the ability of semantic measurement and are helpful for determining the category of the sample. The non-semantic feature extraction network only considers the content information of the sample, does not have the ability of semantic measurement, and cannot determine the category of the sample.
[0105] Specifically, the computer device inputs each sample in the first sample group into the initial feature extraction model respectively. The sample classification network of the initial feature extraction model processes the input samples to output the initial classification features and initial semantic features, and the non-semantic feature extraction network of the initial feature extraction model processes the input samples to output the initial non-semantic features.
[0106] In one embodiment, in order to improve the training efficiency of the model, the non-semantic feature extraction network and the sample classification network can share the underlying network parameters.
[0107] Step S206: Through the feature fusion network, fuse the initial semantic features and initial non-semantic features of the same sample to obtain the initial fusion features corresponding to each sample.
[0108] Specifically, after obtaining the initial semantic features and initial non-semantic features, further fuse the initial semantic features and initial non-semantic features of the same sample through the feature fusion network of the initial feature extraction model, so as to obtain the initial fusion features corresponding to the target sample and the reference sample respectively. Among them, fusing the features can specifically be concatenating the features, or compressing the features after concatenation to reduce the data volume of the features.
[0109] Step S208: Calculate the loss based on the initial semantic features, initial non-semantic features, and initial fusion features corresponding to the target sample and the reference sample to obtain the feature loss.
[0110] Specifically, the computer device may calculate a loss based on the initial semantic features, initial non-semantic features, and initial fusion features corresponding to the target sample and the reference sample, and obtain a feature loss. For example, the computer device may obtain a semantic feature loss based on the distance between the initial semantic features corresponding to the target sample and the reference sample, obtain a non-semantic feature loss based on the distance between the initial non-semantic features corresponding to the target sample and the reference sample, obtain a fusion feature loss based on the distance between the initial fusion features corresponding to the target sample and the reference sample, and obtain a feature loss by synthesizing the semantic feature loss, non-semantic feature loss, and fusion feature loss.
[0111] Step S210: Calculate a loss based on the initial classification features and class labels corresponding to the same sample, and obtain a classification loss.
[0112] Specifically, the computer device may obtain a classification loss based on the differences between the initial classification features and class labels corresponding to the target sample, and the differences between the initial classification features and class labels corresponding to the reference sample. For example, calculate the distance between the initial classification features and class labels corresponding to the same sample, and obtain a classification loss based on the calculation results corresponding to each sample. In one embodiment, the computer device may calculate the classification loss through a cross-entropy loss function.
[0113] Step S212: Based on the feature loss and the classification loss, adjust the model parameters of the initial feature extraction model until the convergence condition is met, and obtain a target feature extraction model; the target feature extraction model is used to extract the sample features of the input sample, and the sample features are used for sample retrieval.
[0114] Among them, the convergence condition may be at least one of the model iteration times reaching a preset number, the total loss being less than a preset loss, and the change rate of the total loss in a continuous preset number of iteration rounds being less than a preset threshold. The target feature extraction model refers to the feature extraction model that has been trained.
[0115] Specifically, after obtaining the feature loss and the target loss, the computer device may perform backpropagation based on the feature loss and the classification loss, update the model parameters of the initial feature extraction model, obtain an updated initial feature extraction model, and return to the step of inputting each sample in the first sample group into the initial feature extraction model, and continue training until the convergence condition is met, then the training is completed, and a target feature extraction model is obtained.
[0116] In one embodiment, the full set of training samples can be divided into different batches to obtain a first sample group for each batch. Each batch's first sample group is used for training, and multiple rounds of iteration of the full set of training samples are performed. When updating the model parameters, the gradient descent algorithm can be used for backpropagation. For example, the stochastic gradient descent method can be used to calculate the loss gradient, and the loss gradient is backpropagated to each network to update the model parameters.
[0117] After obtaining the target feature extraction model, the target feature extraction model can be used to extract the sample features of the input samples. For example, if the target feature extraction model is an image feature extraction model, then when an image is input into the image feature extraction model, the image feature extraction model can output the semantic features, non-semantic features, and fusion features of the image. The sample features extracted by the target feature extraction model can be used for sample retrieval. Sample retrieval refers to retrieving at least one sample from the sample library that is most similar to the target sample and the query sample. Based on the similarity between the sample features of two samples, it can be determined whether the two samples are similar. When performing sample retrieval, based on the similarity between the sample features of the query sample and each sample in the sample library, the samples in the sample library that are similar to the query sample can be determined. Then, the samples similar to the query sample can be used as the sample retrieval results of the query sample. When performing sample retrieval, all sample features can be used. Of course, to improve the retrieval efficiency, only one type of sample feature can also be used. Further, sample retrieval can be triggered passively. For example, in a search application, when the user enters search information in the search box for information retrieval, the search application will trigger sample retrieval, use the user's search information as the query sample, determine the sample retrieval results based on the query sample in the search library, and use the sample retrieval results as the search results to display the search results to the user. Sample retrieval can also be triggered automatically. For example, in a video and audio application, without the user entering search information, the video and audio application can automatically recommend information to the user. The video and audio application can automatically trigger sample retrieval, use the user's historical search results as the query sample, determine the sample retrieval results based on the query sample in the video and audio library, and use the sample retrieval results as the recommended results to actively display the recommended results to the user. Thus, it can be seen that the sample features extracted by the target feature extraction model can be applied to information search scenarios and information recommendation scenarios.
[0118] In one embodiment, the feature sizes corresponding to the initial semantic features, initial non-semantic features, and initial fusion features are the same.
[0119] Among them, the feature size refers to the data volume and scale of the feature. For example, if a semantic feature is represented by a 1*128 vector, then the feature size of this semantic feature can be 1*128.
[0120] Specifically, the feature sizes corresponding to the initial semantic features, the initial non-semantic features, and the initial fusion features can be the same. Then, when performing sample retrieval, on the premise of ensuring the retrieval accuracy, in order to improve the retrieval efficiency, only the fusion features that include both semantic information and non-semantic information can be used, effectively reducing the computational amount during retrieval. Calculate the similarity between the query sample and each sample in the sample library based on the fusion features corresponding to the query sample and each sample in the sample library, and determine the retrieved result sample corresponding to the query sample from the sample library based on the similarity.
[0121] In the above feature extraction model training method, by obtaining the first sample group and inputting each sample in the first sample group into the initial feature extraction model; the first sample group includes the target sample, the reference sample corresponding to the target sample, and the class label corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network; output the initial classification feature and the initial semantic feature through the sample classification network, and output the initial non-semantic feature through the non-semantic feature extraction network; through the feature fusion network, fuse the initial semantic feature and the initial non-semantic feature of the same sample to obtain the initial fusion feature corresponding to each sample; calculate the loss based on the initial semantic feature, the initial non-semantic feature, and the initial fusion feature corresponding to the target sample and the reference sample to obtain the feature loss; calculate the loss based on the initial classification feature and the class label corresponding to the same sample to obtain the classification loss; based on the feature loss and the classification loss, adjust the model parameters of the initial feature extraction model until the convergence condition is met to obtain the target feature extraction model. In this way, a unified model is established to learn semantic features and non-semantic features, and at the same time learn fusion features. The fusion features contain both semantic information and non-semantic information. The finally trained model can not only output semantic features and non-semantic features containing single-dimensional information, but also output fusion features containing two-dimensional information. Only one model needs to be trained to make the model output diverse features, improving the training efficiency.
[0122] In one embodiment, the sample classification network includes a semantic feature extraction sub-network and a semantic feature classification sub-network, and there is a shared network layer between the semantic feature extraction sub-network and the non-semantic feature extraction network. Outputting the initial classification feature and the initial semantic feature through the sample classification network, and outputting the initial non-semantic feature through the non-semantic feature extraction network includes:
[0123] Perform convolution processing on the input sample through the shared network layer to obtain the shared feature; perform feature processing on the shared feature through the feature processing layer of the semantic feature extraction sub-network to obtain the initial semantic feature; perform classification processing on the initial semantic feature through the semantic feature classification sub-network to obtain the initial classification feature; perform feature processing on the shared feature through the feature processing layer of the non-semantic feature extraction network to obtain the initial non-semantic feature.
[0124] Among them, the sample classification network includes a semantic feature extraction sub-network and a semantic feature classification sub-network. The semantic feature classification sub-network is connected after the semantic feature extraction sub-network, and the input data of the semantic feature classification sub-network is the output data of the semantic feature extraction sub-network. The input data of the semantic feature extraction sub-network is the sample, and the output data is the semantic feature. The input data of the semantic feature classification sub-network is the semantic feature, and the output data is the classification feature.
[0125] Furthermore, there is a shared network layer between the semantic feature extraction sub-network and the non-semantic feature extraction network, that is, the semantic feature extraction sub-network and the non-semantic feature extraction network share the underlying network structure. The shared network layer is the overlapping network structure of the semantic feature extraction sub-network and the non-semantic feature extraction network.
[0126] Specifically, after the sample is input into the initial feature extraction model, first, the input sample is subjected to convolutional processing through the shared network layer to extract the depth feature information of the input sample and obtain the shared feature. Then, the shared feature is respectively input into the subsequent network layers in the semantic feature extraction sub-network and the non-semantic feature extraction network. The shared feature is further processed by the subsequent feature processing layer of the semantic feature extraction sub-network to compress the feature and obtain the initial semantic feature. Finally, the initial semantic feature is classified by the semantic feature classification sub-network to obtain the initial classification feature. The shared feature is further processed by the subsequent feature processing layer of the non-semantic feature extraction network to compress the feature and obtain the initial non-semantic feature.
[0127] Reference Figure 3 , the feature extraction network includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network. The sample classification network includes a shared network layer, a first feature processing layer, and a semantic feature classification sub-network connected in sequence. The shared network layer and the first feature processing layer form the semantic feature extraction sub-network. The non-semantic feature extraction network includes a shared network layer and a second feature processing layer connected in sequence. The output data of the first feature processing layer is the semantic feature, the output data of the semantic feature classification sub-network is the classification feature, and the output data of the second feature processing layer is the non-semantic feature. The first feature processing layer and the second feature processing layer are respectively connected to the feature fusion network, and the output data of the first feature processing layer and the second feature processing layer serve as the input data of the feature fusion network. The output data of the feature fusion network is the fusion feature.
[0128] In one embodiment, the shared network layer includes the network structures shown in Table 1 and Table 2, that is, the shared network layer includes a convolutional layer and a pooling layer. The convolutional layer is used to extract features, and the pooling layer is used to compress features.
[0129] Table 1 Structure Table of ResNet-101 Feature Module
[0130]
[0131]
[0132] Table 2 Pooling layer structure for compressing the deep features output by ResNet-101 into a one-dimensional vector
[0133]
[0134] The feature processing layer of the semantic-free feature extraction network includes the network structure shown in Table 3. That is, the feature processing layer of the semantic-free feature extraction network includes a filtering layer and an embedding layer. The filtering layer is used to filter redundant features, and the embedding layer is used to compress features.
[0135] Table 3 Semantic-free module structure table, including semantic-free embedding extraction
[0136]
[0137] The feature processing layer of the semantic feature extraction sub-network and the semantic feature classification sub-network include the network structure shown in Table 4. That is, the feature processing layer of the semantic feature extraction sub-network includes a filtering layer and an embedding layer, and the semantic feature classification sub-network includes a classification layer, which is used to perform feature classification.
[0138] Table 4 Semantic module structure table, including semantic embedding extraction and semantic classification
[0139]
[0140]
[0141] The feature fusion network includes the network structure shown in Table 5. That is, the feature fusion network includes a merging layer and an embedding layer. The merging layer is used for feature splicing. The role of the merging layer is to concatenate the 1*128 feature vectors output by the embedding layer 1 in Table 3 and the embedding layer 2 in Table 4 end to end to form a 1*256 feature vector. The role of the embedding layer 3 in Table 5 is to perform information fusion on the concatenated 1*256 feature vector and compress it into a 1*128 feature vector. Training the initial feature extraction model established by the network layers from Table 1 to Table 5 can obtain the target feature extraction model. Among them, Conv1-Conv5 can be initialized with the parameters of ResNet101 pre-trained on the ImageNet dataset, and the newly added layers can be initialized with a Gaussian distribution with a variance of 0.01 and a mean of 0, such as each embedding layer.
[0142] Table 5 Fusion module structure table, including fusion embedding
[0143]
[0144] In the above embodiments, the sample classification network includes a semantic feature extraction sub-network and a semantic feature classification sub-network. There is a shared network layer between the semantic feature extraction sub-network and the non-semantic feature extraction network. The shared network layer can reduce the complexity of the model, reduce the data processing volume of the model, and reduce the parameters that the model needs to learn. Through the cooperation of the shared network layer, the semantic feature extraction sub-network and the semantic feature classification sub-network, the initial semantic feature and the initial classification feature are finally output.
[0145] In one embodiment, the reference sample includes a positive sample and a negative sample corresponding to the target sample. Based on the initial semantic feature, the initial non-semantic feature, and the initial fusion feature corresponding to the target sample and the reference sample, the loss is calculated to obtain the feature loss, including:
[0146] Based on the distance between the same type of features corresponding to the target sample and the positive sample, the positive semantic loss, the positive non-semantic loss, and the positive fusion loss are obtained; based on the distance between the same type of features corresponding to the target sample and the negative sample, the negative semantic loss, the negative non-semantic loss, and the negative fusion loss are obtained; based on the distance between the positive semantic loss and the negative semantic loss, the initial semantic loss is obtained, based on the distance between the positive non-semantic loss and the negative non-semantic loss, the initial non-semantic loss is obtained, and based on the distance between the positive fusion loss and the negative fusion loss, the initial fusion loss is obtained; the feature loss is obtained based on the initial semantic loss, the initial non-semantic loss, and the initial fusion loss.
[0147] Among them, the reference sample corresponding to the target sample includes the positive sample and the negative sample corresponding to the target sample, that is, the first sample group includes three types of samples. The similarity between the positive sample corresponding to the target sample and the target sample is greater than the similarity between the negative sample corresponding to the target sample and the target sample.
[0148] Specifically, if the reference sample includes the positive sample and the negative sample corresponding to the target sample, the training objective of the model can be to make the feature distance between the target sample and the positive sample less than the feature distance between the target sample and the negative sample. Then, when performing feature retrieval later, based on the sample features, a sample more similar to the query sample can be retrieved from a large number of samples.
[0149] When calculating the feature loss, the computer device can obtain the positive semantic loss, positive non-semantic loss, and positive fusion loss based on the distance between the same type of features corresponding to the target sample and the positive sample, and obtain the negative semantic loss, negative non-semantic loss, and negative fusion loss based on the distance between the same type of features corresponding to the target sample and the negative sample. For example, the positive fusion loss is obtained based on the distance between the initial fusion features corresponding to the target sample and the positive sample respectively. The computer device then further combines the positive and negative losses of the same type to obtain the feature loss. The computer device obtains the initial semantic loss based on the distance between the positive semantic loss and the negative semantic loss, obtains the initial non-semantic loss based on the distance between the positive non-semantic loss and the negative non-semantic loss, and obtains the initial fusion loss based on the distance between the positive fusion loss and the negative fusion loss. Finally, the computer device combines the initial semantic loss, initial non-semantic loss, and initial fusion loss to obtain the feature loss.
[0150] It can be understood that the computer device can adopt a custom algorithm or formula to calculate the distance between features, or adopt calculation methods such as Euclidean distance and Manhattan distance to calculate the distance between features.
[0151] In the above embodiments, the reference samples include the positive and negative samples corresponding to the target sample. Based on the distance between the same type of features corresponding to the target sample and the positive sample, the positive loss is obtained. Based on the distance between the same type of features corresponding to the target sample and the negative sample, the negative loss is obtained. Based on the positive loss and the negative loss, the feature loss is obtained. Based on the feature loss, the model parameters are adjusted, enabling the model to have the ability to distinguish positive and negative samples.
[0152] In one embodiment, obtaining the feature loss based on the initial semantic loss, initial non-semantic loss, and initial fusion loss includes:
[0153] Adjusting the parameters based on the semantic loss to update the initial semantic loss to obtain the intermediate semantic loss, and determining the target semantic loss based on the matching result between the intermediate semantic loss and the preset parameters; adjusting the parameters based on the non-semantic loss to update the initial non-semantic loss to obtain the intermediate non-semantic loss, and determining the target non-semantic loss based on the matching result between the intermediate non-semantic loss and the preset parameters; adjusting the parameters based on the fusion loss to update the initial fusion loss to obtain the intermediate fusion loss, and determining the target fusion loss based on the matching result between the intermediate fusion loss and the preset parameters; the parameter adjustment for the fusion loss is greater than the parameter adjustment for the semantic loss, and the parameter adjustment for the fusion loss is greater than the parameter adjustment for the non-semantic loss; generating the feature loss based on the target semantic loss, target non-semantic loss, and target fusion loss.
[0154] Among them, the loss adjustment parameter is used to adjust the loss and control the distance between the positive loss and the negative loss. The loss adjustment parameter can be set as needed. The loss adjustment parameters corresponding to different losses can be the same or different. The preset parameter can be set according to actual needs. For example, it can be set to 0.
[0155] Specifically, when fusing various losses, in order to further widen the feature distance between the target sample and the positive and negative samples, the computer device can adjust various losses and then fuse them. Taking the initial semantic loss as an example, the initial semantic loss can represent the distance between the positive semantic loss and the negative semantic loss. The computer device can obtain the semantic loss adjustment parameter, update the initial semantic loss based on the semantic loss adjustment parameter, amplify the initial semantic loss, and obtain the intermediate semantic loss. The computer device then matches the intermediate semantic loss with the preset parameter, compares the numerical values of the intermediate semantic loss and the preset parameter, and obtains the larger numerical value as the target semantic loss. The purpose of the target semantic loss is to widen the distance between the positive semantic loss and the negative semantic loss, so that similar samples and dissimilar samples can be distinguished based on the semantic features output by the model. Similarly, the computer device can update the initial non-semantic loss based on the non-semantic loss adjustment parameter to obtain the intermediate non-semantic loss, and determine the target non-semantic loss based on the matching result between the intermediate non-semantic loss and the preset parameter. The computer device can update the initial fusion loss based on the fusion loss adjustment parameter to obtain the intermediate fusion loss, and determine the target fusion loss based on the matching result between the intermediate fusion loss and the preset parameter.
[0156] Furthermore, the fusion loss adjustment parameter is greater than the semantic loss adjustment parameter, and the fusion loss adjustment parameter is greater than the non-semantic loss adjustment parameter. For example, the fusion loss adjustment parameter is 0.8, and the semantic loss adjustment parameter and the non-semantic loss adjustment parameter are 0.6. That is, after combining semantic information and non-semantic information, the fusion feature can further widen the feature distance between the target sample and the positive and negative samples, and further widen the distance between the positive fusion loss and the negative fusion loss. Subsequently, when performing sample retrieval, accurate retrieval results can also be obtained only based on the fusion feature.
[0157] After obtaining the target semantic loss, the target non-semantic loss, and the target fusion loss, the computer device can use the sum of the target semantic loss, the target non-semantic loss, and the target fusion loss as the feature loss, or use the weighted sum of the target semantic loss, the target non-semantic loss, and the target fusion loss as the feature loss. Among them, the weights corresponding to the target semantic loss, the target non-semantic loss, and the target fusion loss can be set as needed.
[0158] In one embodiment, the calculation formula of the target loss is as follows:
[0159] ltri = max(||x a - x p || - ||x a - x n || + α, 0)
[0160] where max(a, b) represents taking the maximum value of a and b. x a represents the sample feature of the target sample, and x p represents the sample feature of the positive sample corresponding to the target sample, and x n represents the sample feature of the negative sample corresponding to the target sample. ||x a - x p || represents calculating the L2 distance, i.e., the Euclidean distance, between x a and x p . α represents the loss adjustment parameter. The purpose of l tri is to make the distance between the target sample and the negative sample greater than the distance between the target sample and the positive sample by α. The target semantic loss, the target non - semantic loss, and the target fusion loss can all be calculated using the above formula, but the α corresponding to the target fusion loss is greater than the α corresponding to the target semantic loss and the target non - semantic loss.
[0161] In the above embodiments, generating a feature loss based on the target semantic loss, the target non - semantic loss, and the target fusion loss, and adjusting the model parameters based on this feature loss can improve the feature discrimination ability of the model.
[0162] In one embodiment, calculating a loss based on the initial classification feature and the class label corresponding to the same sample to obtain a classification loss, including:
[0163] Performing label encoding on the class labels of each sample to obtain corresponding label features; performing normalization processing on the classification features corresponding to each sample to obtain corresponding normalized features; performing logarithmic transformation on each normalized feature, and fusing the label features corresponding to the same sample and the logarithmically transformed normalized features to obtain classification sub - losses corresponding to each sample; obtaining the classification loss based on each classification sub - loss.
[0164] Among them, label encoding is used to convert the class label into data represented by binary. Each vector dimension in the label feature vector corresponds to a label category, and the specific vector value on the vector dimension represents whether the sample matches the corresponding label category. Normalization processing is used to map the vector values of the feature vector to a preset numerical range. The preset numerical range can be set as needed, for example, between 0 and 1.
[0165] Specifically, to facilitate the calculation of the classification loss, the computer device can process the classification features and class labels and convert them into data that is easy to compare and calculate. For the class labels, the computer device can perform label encoding on the class labels corresponding to the target samples and reference samples, and convert each class label into a label feature. For example, one-hot encoding can be used for label encoding. The computer device can perform normalization processing on the classification features corresponding to the target samples and reference samples respectively, map the vector values of the classification feature vectors to a preset numerical range, so as to obtain the normalized features corresponding to each sample respectively. For example, the softmax function can be used for normalization processing. The computer device performs a logarithmic transformation on each normalized feature, specifically, it can use a preset value as the base and the normalized feature as the true number for the logarithmic transformation. Finally, the computer device fuses the label features corresponding to the same sample and the normalized features after logarithmic transformation to obtain the classification sub-loss corresponding to each sample, and then combines the classification sub-losses to obtain the classification loss.
[0166] In one embodiment, the calculation formula of the classification loss is as follows:
[0167]
[0168] where p k represents the label feature corresponding to the k-th sample, q k represents the normalized feature obtained by performing normalization processing on the classification feature corresponding to the k-th sample, and N represents the number of samples.
[0169] In the above embodiment, by performing label encoding on the class labels, performing normalization processing and logarithmic transformation on the classification features, and then fusing the processing results of the two, an accurate classification loss can be obtained.
[0170] In one embodiment, the sample classification network includes a semantic feature extraction sub-network and a semantic feature classification sub-network. Based on the feature loss and the classification loss, adjusting the model parameters of the initial feature extraction model until the convergence condition is met to obtain the target feature extraction model, including:
[0171] Based on the feature loss and the classification loss, obtain the target loss, calculate the gradient of the target loss to obtain the loss gradient; update the loss gradient based on the first adjustment parameter to obtain the first loss, update the loss gradient based on the second adjustment parameter to obtain the second loss; the first adjustment parameter is less than the second adjustment parameter; adjust the network parameters of the semantic feature classification sub-network based on the first loss, and adjust the network parameters of other networks based on the second loss until the convergence condition is met to obtain the target feature extraction model.
[0172] Specifically, since the semantic feature classification sub-network involves classification processing and is prone to overfitting, in order to reduce the impact of the semantic feature classification sub-network on other networks, different losses can be adopted to update the semantic feature classification sub-network and other networks. After obtaining the updated loss and classification loss, the computer device can fuse the feature loss and classification loss to obtain the target loss. For example, the sum of the feature loss and classification loss can be used as the target loss, or the weighted sum of the feature loss and classification loss can be used as the target loss, and then the gradient of the target loss is calculated to obtain the loss gradient. The computer device can update the loss gradient based on the first adjustment parameter to obtain the first loss, adjust the network parameters of the semantic feature classification sub-network based on the first loss, update the loss gradient based on the second adjustment parameter to obtain the second loss, and adjust the network parameters of other networks based on the second loss until the convergence condition is met to obtain the target feature extraction model. Updating the loss gradient based on the adjustment parameter can specifically be multiplying the adjustment parameter and the loss gradient.
[0173] Among them, the first adjustment parameter is smaller than the second adjustment parameter, and the first adjustment parameter and the second adjustment parameter can be set as needed. For example, the first adjustment parameter is 10 times the second adjustment parameter, the first adjustment parameter is set to 0.005, and the second adjustment parameter is set to 0.0005. It can be understood that since the first adjustment parameter is smaller than the second adjustment parameter, the loss generated by classification will only be fully backpropagated to the semantic feature classification sub-network, and backpropagated to other networks in a multiple less than 1, thereby reducing the impact of the semantic feature classification sub-network on other networks and ensuring the overall training effect of the model. Further, if there is a shared network layer between the semantic feature extraction sub-network and the non-semantic feature extraction network, the fact that the first adjustment parameter is smaller than the second adjustment parameter can not only avoid overfitting of the classification information to the semantic embedding, but also avoid the non-semantic embedding being affected by classification due to excessive backpropagation of the classification gradient to the underlying shared features.
[0174] In the above embodiment, the target loss is obtained based on the feature loss and classification loss, the gradient of the target loss is calculated to obtain the loss gradient, the loss gradient is updated based on the first adjustment parameter to obtain the first loss, the loss gradient is updated based on the second adjustment parameter to obtain the second loss, the first adjustment parameter is smaller than the second adjustment parameter, the network parameters of the semantic feature classification sub-network are adjusted based on the first loss, and the network parameters of other networks are adjusted based on the second loss, which can improve the training effect of the model.
[0175] In one embodiment, as Figure 4 shown, before inputting each sample in the first sample group into the initial feature extraction model, the method further includes:
[0176] Step S402: Obtain a second sample group, input each sample in the second sample group into the candidate feature extraction model, and obtain a candidate non-semantic feature set, a candidate semantic feature set, and a candidate fusion feature set corresponding to the second sample group.
[0177] Step S404: Calculate a loss based on the candidate non-semantic feature set, the candidate semantic feature set, and the candidate fusion feature set to obtain a candidate loss.
[0178] Step S406: Based on the candidate loss, adjust the network parameters of the target network in the candidate feature extraction model until a first condition is satisfied to obtain an initial feature extraction model; the target network includes a sample classification network and a non-semantic feature extraction network.
[0179] Among them, the second sample group and the first sample group may include the same samples or different samples. The second sample group may also be at least one. The candidate feature extraction model is also a feature extraction model to be trained, and an initial feature extraction model is obtained after pre-training the candidate feature extraction model. Similar to the convergence condition, the first condition may also be that the number of model iterations reaches a preset number, the candidate loss is less than a preset loss, the change rate of the candidate loss in a continuous preset number of iteration rounds is less than a preset threshold, etc.
[0180] Specifically, since the fusion feature is obtained based on the semantic feature and the non-semantic feature, if the extraction effects of the semantic feature and the non-semantic feature are good, it is relatively easy to obtain an excellent fusion feature. To improve the model training efficiency, the networks for extracting semantic features and non-semantic features can be trained first. On the basis that the networks for extracting semantic features and non-semantic features perform well, comprehensive training can be carried out to finely adjust the overall model parameters, and then the target feature extraction model can be quickly obtained.
[0181] The computer device can obtain a second sample group, train a candidate feature extraction model based on the second sample group, and only adjust the network parameters of the sample classification network and the non-semantic feature extraction network in the candidate feature extraction model to obtain an initial feature extraction model. The computer device inputs each sample in the second sample group into the candidate feature extraction model, and the candidate feature extraction model outputs the candidate non-semantic features, candidate semantic features, and candidate fusion features corresponding to each sample in the second sample group, so as to obtain a candidate non-semantic feature set, a candidate semantic feature set, and a candidate fusion feature set corresponding to the second sample group. Referring to the calculation method of the feature loss, the computer device can obtain a candidate loss based on the candidate non-semantic feature set, the candidate semantic feature set, and the candidate fusion feature set. The computer device only adjusts the network parameters of the sample classification network and the non-semantic feature extraction network in the feature extraction model based on the candidate loss, and fixes the network parameters of the feature fusion network. The computer device can perform backpropagation based on the candidate loss, update the model parameters of the candidate feature extraction model, obtain an updated candidate feature extraction model, and return to the step of inputting each sample in the second sample group into the candidate feature extraction model, and continue training until the convergence condition is met, then the training is completed, and an initial feature extraction model is obtained.
[0182] In one embodiment, similar to adjusting the model parameters based on the feature loss and the classification loss, when adjusting the model parameters based on the candidate loss, the loss gradient of the candidate loss can also be adjusted based on a third adjustment parameter to obtain a third loss, the network parameters of the semantic feature classification sub-network are adjusted based on the third loss, the loss gradient of the candidate loss is adjusted based on a fourth adjustment parameter to obtain a fourth loss, and the network parameters of the non-semantic feature extraction network and the semantic feature extraction sub-network are adjusted based on the fourth loss until the first condition is met, and an initial feature extraction model is obtained. Among them, the third adjustment parameter is greater than the fourth adjustment parameter. For example, the third adjustment parameter is 0.005, and the fourth adjustment parameter is 0.0005.
[0183] In the above embodiment, the candidate feature extraction model is trained based on the second sample group, the sample classification network and the non-semantic feature extraction network in the model are adjusted to obtain an initial feature extraction model, and then the feature extraction model is fine-tuned to quickly obtain the target feature extraction model.
[0184] In one embodiment, as Figure 5 shown, before obtaining the second sample group and inputting each sample in the second sample group into the candidate feature extraction model, the method further includes:
[0185] Step S502, obtaining a third sample group, inputting each sample in the third sample group into the feature extraction model to be trained, and obtaining a non-semantic feature set corresponding to the third sample group.
[0186] Step S504: Calculate the loss based on the non-semantic feature set to obtain the initial loss.
[0187] Step S506: Based on the initial loss, adjust the model parameters of the non-semantic feature extraction network in the to-be-trained feature extraction model until the second condition is met, and obtain the candidate feature extraction model.
[0188] Among them, the third sample group and the first sample group may include the same samples or different samples. The third sample group may also be at least one. Similar to the convergence condition, the second condition may also be that the number of model iterations reaches a preset number, the initial loss is less than a preset loss, the change rate of the initial loss in a continuous preset number of iteration rounds is less than a preset threshold, etc.
[0189] Specifically, further considering that the convergence speeds of the networks for extracting semantic features and non-semantic features are different, the network for extracting non-semantic features has a slow convergence speed, and the network for extracting semantic features has a fast convergence speed. To improve the training effect of the model, the non-semantic feature extraction network can be trained first, the relevant model parameters of the non-semantic feature extraction network can be adjusted first, and then the next stage of training can be carried out.
[0190] The computer device can obtain the third sample group, train the to-be-trained feature extraction model based on the third sample group, only adjust the model parameters of the non-semantic feature extraction network in the model, and obtain the candidate feature extraction model. The computer device inputs each sample in the third sample group into the to-be-trained feature extraction model, and the to-be-trained feature extraction model outputs the non-semantic features corresponding to each sample in the third sample group, so as to obtain the non-semantic feature set corresponding to the third sample group. Referring to the calculation method of the feature loss, the computer device can obtain the initial loss based on the non-semantic feature set. The computer device only adjusts the network parameters of the non-semantic feature extraction network in the feature extraction model based on the initial loss. The computer device can perform backpropagation based on the initial loss, update the model parameters of the non-semantic feature extraction network, obtain the updated to-be-trained feature extraction model, and return to the step of inputting each sample in the third sample group into the to-be-trained feature extraction model and iterate to continue training until the convergence condition is met, then the training is completed, and the candidate feature extraction model is obtained.
[0191] In the above embodiments, training the to-be-trained feature extraction model based on the third sample group and adjusting the non-semantic feature extraction network in the model helps to maintain the convergence balance between the semantic branch and the non-semantic branch in the subsequent training stage and improve the training efficiency.
[0192] Of course, the computer device can also obtain a fourth sample group, input each sample in the fourth sample group into a feature extraction model with initialized parameters, obtain a non-semantic feature set corresponding to the fourth sample group, calculate loss information based on the non-semantic feature set corresponding to the fourth sample group, and adjust the network parameters of the non-semantic feature extraction network in the feature extraction model until the third condition is met, thereby obtaining an initial feature extraction model.
[0193] In one embodiment, the current sample group is any one of the first sample group, the second sample group, and the third sample group.
[0194] Obtaining the current sample group includes: obtaining a plurality of similar sample pairs; determining a current sample and a positive sample corresponding to the current sample from the current similar sample pairs, and determining a plurality of candidate samples from the remaining similar sample pairs; based on the sample similarity between the current sample and each candidate sample, determining at least one negative sample corresponding to the current sample from each candidate sample; using the positive sample and the negative sample corresponding to the current sample as the reference samples corresponding to the current sample, and obtaining at least one current sample group based on the current sample and the corresponding reference samples.
[0195] A similar sample pair refers to a sample pair in which two samples are labeled as the same or similar samples.
[0196] Specifically, when determining a sample triple, the computer device can obtain a plurality of similar sample pairs, randomly select a similar sample pair from the plurality of similar sample pairs as the current similar sample pair, use one sample in the current similar sample pair as the current sample, and the other sample as the positive sample corresponding to the current sample. The computer device then randomly selects a plurality of samples from the remaining similar sample pairs other than the current similar sample pair as candidate samples. For example, from each of the remaining similar sample pairs, one sample is randomly selected as a candidate sample. Then, the computer device calculates the sample similarity between the current sample and each candidate sample, and determines at least one negative sample corresponding to the current sample from each candidate sample based on the sample similarity. When calculating the sample similarity, the computer device can use a custom algorithm or a traditional text similarity calculation algorithm, image similarity calculation algorithm, voice similarity calculation algorithm, etc. When selecting negative samples, the candidate samples can be sorted in descending order of sample similarity, and several candidate samples with higher rankings are obtained as negative samples. After obtaining the positive sample and the negative sample corresponding to the current sample, the positive sample and the negative sample corresponding to the current sample are used as the reference samples corresponding to the current sample, and the current sample group is formed based on the current sample and the corresponding reference samples, and finally the same number of current sample groups as the number of negative samples is obtained.
[0197] It can be understood that a current sample group includes three samples, namely a current sample, a positive sample corresponding to the current sample, and a negative sample. If there are multiple negative samples corresponding to the current sample, then each negative sample and the current sample pair corresponding to the current sample form a current sample group, and finally multiple current sample groups are obtained. In addition, the class labels corresponding to each sample in the similar sample pairs are the same, that is, the sample and the positive sample corresponding to the sample have the same class label. However, the class labels corresponding to the positive sample and the negative sample corresponding to the same sample can be the same or different. The first sample group, the second sample group, and the third sample group can all be obtained by the above method. In one embodiment, the first sample group, the second sample group, and the third sample group can include the same similar sample pairs.
[0198] In the above embodiment, determining the negative sample corresponding to the current sample based on the sample similarity from the remaining similar sample pairs can quickly determine the negative sample on the basis of making full use of the existing data.
[0199] In one embodiment, determining at least one negative sample corresponding to the current sample from each candidate sample based on the sample similarity between the current sample and each candidate sample includes:
[0200] Inputting the current sample and each candidate sample into the matched current feature extraction model to obtain the sample feature sets corresponding to the current sample and each candidate sample respectively; the current feature extraction model matched with the first sample group is the initial feature extraction model, the current feature extraction model matched with the second sample group is the candidate feature extraction model, and the current feature extraction model matched with the third sample group is the to-be-trained feature extraction model; calculating the sample similarity between the current sample and each candidate sample based on the sample feature sets corresponding to the current sample and each candidate sample; classifying each candidate sample into a first type of sample and a second type of sample based on the sample similarity; the sample similarity corresponding to the first type of sample is greater than the sample similarity corresponding to the second type of sample; determining at least one negative sample corresponding to the current sample from the first type of samples.
[0201] Specifically, the similarity between samples can be calculated based on the sample features output by the model. Taking the first sample group as an example, when determining the first sample group, the target sample and each candidate sample can be input into the initial feature extraction model, and the sample feature set corresponding to the target sample and each candidate sample can be obtained according to the output data of the model. The sample feature set includes at least one of semantic features, non-semantic features and fusion features. Then, based on the sample feature set corresponding to the target sample and each candidate sample, the sample similarity between the target sample and each candidate sample is calculated. For example, the sample feature distance between the two samples is used as the sample similarity between the samples. The smaller the sample feature distance, the greater the sample similarity. The computer device can divide each candidate sample into a first class sample and a second class sample based on the sample similarity. The sample similarity corresponding to the first class sample is greater than the sample similarity corresponding to the second class sample. At least one sample is selected from the first class sample as at least one negative sample corresponding to the target sample. When performing sample classification, each candidate sample can be sorted in order from large to small in sample similarity, and a preset number of candidate samples with a higher sorting order are obtained as first class samples, and the remaining candidate samples are taken as second class samples. The preset number can be set as needed.
[0202] It can be understood that negative samples are determined from candidate samples with high sample similarity to the target samples, and the negative samples selected therefrom include a certain proportion of difficult samples. Subsequently, during model training, difficult samples help improve the model's feature discrimination capability, thereby improving model performance.
[0203] Similar to the method of determining the first sample group, when determining the second sample group, the current sample and each candidate sample can be input into the candidate feature extraction model, and the sample similarity can be calculated based on the output data of the model to determine the negative samples in the second sample group. When determining the third sample group, the current sample and each candidate sample can be input into the candidate feature extraction model, and the sample similarity can be calculated based on the output data of the model to determine the negative samples in the third sample group.
[0204] In the above embodiment, different training samples are used for model training at different training stages to improve the generalization ability of the model. Sample similarity is calculated based on the sample features output by the model trained in the previous stage, and the training samples for the next stage can be quickly determined. Negative samples are determined from the first type of samples with greater similarity, and difficult samples can be obtained, which help improve model performance during the training process.
[0205] In one embodiment, Figure 6 As shown, a sample retrieval method is provided, and the method is described by taking the method executed by a computer device as an example. It can be understood that the computer device can be Figure 1The terminal 102 shown may also be the server 104. In this embodiment, the sample retrieval method includes the following steps:
[0206] Step S602, obtain a query sample and a candidate recall sample set.
[0207] Among them, the query sample refers to the sample for which the most similar sample needs to be queried. The query sample can be at least one of an image, voice, and text. In an information search scenario, the query sample can be the search information submitted by the user through the terminal during a search. In an information recommendation scenario, the query sample can be the user reference information independently obtained by the computer device. The user reference information can be the user's attribute information. For example, the interests and hobbies in the user account registration information can be used as the user reference information. The user reference information can also be determined based on the user's historical behavior. For example, the user's historical search results and historical browsing samples can be used as the user reference information. The user reference information can reflect the user's interests and hobbies and the focus of attention. The candidate recall sample set refers to a sample library and a retrieval library, including a large number of candidate recall samples. At least one sample most similar to the query sample needs to be retrieved from the sample library as the retrieval result.
[0208] Specifically, when receiving a retrieval task or a query task, the computer device can obtain a query sample and a candidate recall sample set, and use the trained target feature extraction model to determine the retrieval result corresponding to the query sample from the candidate recall sample set.
[0209] Step S604, input the query sample and the candidate recall samples in the candidate recall sample set into the target feature extraction model to obtain the query sample feature corresponding to the query sample and the recall sample feature corresponding to the candidate recall sample.
[0210] Among them, the training process of the target feature extraction model is as follows: Obtain a first sample group, and input each sample in the first sample group into the initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and a class label corresponding to each sample. The initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network; output an initial classification feature and an initial semantic feature through the sample classification network, output an initial non-semantic feature through the non-semantic feature extraction network, and through the feature fusion network, fuse the initial semantic feature and the initial non-semantic feature of the same sample to obtain the initial fusion feature corresponding to each sample; calculate the loss based on the initial semantic feature, the initial non-semantic feature, and the initial fusion feature corresponding to the target sample and the reference sample to obtain a feature loss, and calculate the loss based on the initial classification feature and the class label corresponding to the same sample to obtain a classification loss; based on the feature loss and the classification loss, adjust the model parameters of the initial feature extraction model until the convergence condition is met to obtain the target feature extraction model.
[0211] It can be understood that the training process of the target feature extraction model can refer to the various embodiments of the foregoing feature extraction model training method, which will not be elaborated here.
[0212] Step S606: Based on the query sample feature and the recall sample feature, determine the retrieval result sample corresponding to the query sample from the candidate recall sample set.
[0213] Specifically, the computer device can input the query sample and the candidate recall samples in the candidate recall sample set into the target feature extraction model respectively, and obtain the query sample feature corresponding to the query sample and the recall sample features corresponding to each candidate recall sample according to the output data of the model. Both the query sample feature and the recall sample feature include at least one of semantic features, non-semantic features, and fusion features. The computer device can calculate the similarity between samples based on the query sample feature and the recall sample feature, and then obtain the candidate recall samples with higher similarity from the candidate recall sample set as the retrieval result samples corresponding to the query sample. Subsequently, the computer device can return the retrieval result samples to the sending end of the retrieval task and the query task.
[0214] In the above sample retrieval method, when training the model, a unified model is established to learn semantic features and non-semantic features, and at the same time learn fusion features. The fusion features contain both semantic information and non-semantic information. The finally trained model can not only output semantic features and non-semantic features containing single-dimensional information, but also output fusion features containing two-dimensional information. Only one model needs to be trained to make the model output diversified features, improving the training efficiency. Therefore, when performing sample retrieval, sample retrieval is performed through the diversified features output by the model, which can improve the accuracy and efficiency of sample retrieval.
[0215] In one embodiment, the recall sample feature includes at least one of semantic features, non-semantic features, and fusion features, and the query sample feature and the recall sample feature include the same type of sample features. Determining the retrieval result sample corresponding to the query sample from the candidate recall sample set based on the query sample feature and the recall sample feature includes:
[0216] Establish an index based on each recall sample feature corresponding to the same type to obtain a sample recall index corresponding to each type; perform sample retrieval from the corresponding sample recall index based on the query sample feature corresponding to the same type, and determine the retrieval result sample based on the sample retrieval result.
[0217] Specifically, when performing sample retrieval, it is necessary to use sample features of the same type and query for similar samples from the same perspective. Therefore, the query sample features and the retrieved sample features need to include sample features of the same type, and the retrieved sample features can include at least one of semantic features, non-semantic features, and fusion features.
[0218] Furthermore, to improve the retrieval efficiency, the computer device can build an index for retrieval based on the retrieved sample features. The index can represent the distribution of the retrieved sample features in space. Then, when performing retrieval, by determining the distribution area corresponding to the query sample features in the index, the retrieval result sample can be quickly determined from the candidate retrieved samples corresponding to the distribution area. Retrieval based on the index can avoid calculating the similarity between the query sample and the samples in the sample library pairwise each time, avoiding a large amount of complex calculations.
[0219] The computer device can build an index based on each retrieved sample feature corresponding to the same type to obtain a sample retrieval index corresponding to each type. For example, the non-semantic feature corresponds to the first sample retrieval index, the semantic feature corresponds to the second sample retrieval index, and the fusion feature corresponds to the third sample retrieval index. Correspondingly, it is necessary to perform sample retrieval from the corresponding sample retrieval index based on the query sample features corresponding to the same type. Finally, the computer device can determine the retrieval result sample from the candidate retrieved sample set based on the sample retrieval result.
[0220] In one embodiment, before obtaining the retrieval task, the computer device can pre-build the sample retrieval index. Then, after obtaining the retrieval task subsequently, the computer device can quickly perform sample retrieval based on the sample retrieval index and quickly determine the retrieval result sample.
[0221] In the above embodiment, performing sample retrieval through the sample retrieval index can improve the retrieval efficiency.
[0222] In one embodiment, building an index based on each retrieved sample feature corresponding to the same type to obtain a sample retrieval index corresponding to each type includes:
[0223] Performing feature clustering on each retrieved sample feature corresponding to the same type to obtain multiple clustering clusters corresponding to each type; there is a corresponding clustering center for each clustering cluster; based on each clustering cluster corresponding to the same type, obtaining a sample retrieval index corresponding to each type.
[0224] Specifically, when building an index, the computer device can perform feature clustering on the recall sample features corresponding to the same type, clustering the recall sample features that are adjacent in space into the same clustering cluster, and clustering the recall sample features that are at a certain distance in space into different clustering clusters, thereby obtaining multiple clustering clusters corresponding to each type respectively. Each clustering cluster is independent of each other, each clustering cluster corresponds to a local vector space respectively, and each clustering cluster has a corresponding clustering center. Similar to the recall sample features, the clustering center can also be represented by a vector or a matrix. The computer device can build a sample recall index based on the clustering clusters corresponding to the same type, using the clustering center as the main component of the index, and using the candidate recall samples whose sample features fall into the clustering cluster corresponding to the clustering center as the associated information of the index, and finally obtaining the sample recall index corresponding to each type respectively.
[0225] In one embodiment, if there are new samples in the sample library, the computer device can calculate the distances between the sample features of the new samples and each clustering center, and determine which clustering cluster the new samples belong to according to the distances, and associate the new samples with the clustering cluster to which the clustering center corresponding to the minimum distance belongs.
[0226] In the above embodiment, building an index based on clustering clusters can improve the subsequent retrieval efficiency.
[0227] In one embodiment, sample retrieval is performed from the corresponding sample recall index based on the query sample features corresponding to the same type, and the retrieval result samples are determined based on the sample retrieval results, including:
[0228] Based on the distances between the query sample features corresponding to the same type and each clustering center, determine the target clustering cluster from the corresponding clustering clusters respectively, and obtain the target clustering clusters corresponding to each type respectively; based on the distances between the query sample features corresponding to the same type and each recall sample feature within the target clustering cluster, determine the target sample feature from the recall sample features corresponding to the target clustering cluster; use the candidate recall samples corresponding to the target sample feature as the retrieval result samples.
[0229] Specifically, the computer device calculates the distances between the query sample features corresponding to the same feature type and the corresponding clustering centers, and uses the clustering cluster corresponding to the minimum distance as the target clustering cluster, thereby obtaining the target clustering clusters corresponding to each feature type respectively, so as to narrow down the retrieval range. Furthermore, the computer device performs retrieval among the candidate recall samples corresponding to the target clustering cluster, calculates the distances between the query sample features corresponding to the same feature type and each recall sample feature within the corresponding target clustering cluster, uses at least one recall sample feature with the closest distance as the target sample feature, and uses the candidate recall samples corresponding to the target sample feature as the final retrieval result samples. For different types of sample features, the retrieval result samples corresponding to each feature type can be obtained.
[0230] In the above embodiments, the retrieval range is first narrowed based on the clustering center, and then precise retrieval is performed within the target clustering cluster, which can improve the retrieval efficiency.
[0231] In one embodiment, when the query task corresponding to the query sample is a first type of task, the recalled sample features include fused features. When the query task corresponding to the query sample is a second type of task, the recalled sample features include semantic features and non-semantic features, and the query frequency of the first type of task is greater than that of the second type of task.
[0232] Specifically, since the fused features contain both semantic information and non-semantic information, in order to improve the retrieval efficiency, for tasks with high response and high-frequency requests, when retrieving the query samples of this type of task, fused features are used for retrieval. That is, when the query task corresponding to the query sample is the first type of task with a relatively high query frequency, the recalled sample features and the query sample features include fused features, and fused features are used for fast retrieval. When the query task corresponding to the query sample is the second type of task with a relatively low query frequency, the recalled sample features and the query sample features include semantic features and non-semantic features, and semantic features and non-semantic features are used for dual-feature accurate retrieval.
[0233] In the above embodiments, by using different retrieval methods for different tasks, the retrieval flexibility can be improved.
[0234] In a specific embodiment, the above feature extraction model training method and sample retrieval method can be applied to the image field.
[0235] 1. Data Preparation
[0236] Annotate the image similarity samples, and then mine the triple samples. Obtain the similar image pairs in each batch (batch represents the training batch, and there are bs similar image pairs in total), and mine the image triples in the image pairs of each batch. For example, for the image x in a certain similar image pair, randomly select one image from each of the remaining bs - 1 sample pairs, calculate its distance from x, sort them in ascending order of distance, and take the top 10 samples as negative samples, which respectively form triples with x and the positive sample of x. Then each x can generate 10 triples, and the entire batch can obtain 10 * bs triples. Annotate the categories corresponding to each image in the triples.
[0237] 2. Model Structure
[0238] The image feature extraction model is a dual-branch fusion model, and the model structure can refer to Figure 7。The underlying structure of the model is a shared network layer, and the shared network layer can be a shared CNN network. For example, the CNN network can be ResNet101 pre-trained based on the large-scale open-source data ImageNet. The specific structure of ResNet101 can be referred to in Table 1 and Table 2. The images in the triples are input into the model. The input images are processed by Table 1 and Table 2 in sequence, and then input into the dual-branch embedding extraction structure of Table 3 and Table 4 to generate the target non-semantic embedding (i.e., the output data of embedding layer 1) and semantic embedding (i.e., the output data of embedding layer 2). Then, the two embeddings are input into Table 5 to generate the fused embedding (i.e., the output data of embedding layer 3) and output.
[0239] 3. Model Training
[0240] Using the mined triple samples and semantic annotation information, supervised learning is performed on the model. The full amount of data is iterated for epoch (representing the number of training rounds), and the full amount of samples are processed in each round of iteration. Refer to Figure 7 , during the model training process, it is necessary to calculate the loss for semantic and non-semantic embeddings respectively, and it is also necessary to calculate the loss for the fused embedding. In addition, for the semantic module in Table 3, there is a classification layer above the embedding to provide semantic information, and this classification layer also needs to calculate the corresponding loss.
[0241] Considering that the convergence degrees of different branches are different, the non-semantic branch converges slowly and the semantic branch converges quickly, and the fused embedding is obtained based on the semantic embedding and non-semantic embedding, the training process is divided into three stages. First, the non-semantic embedding is trained with Loss1, and the non-semantic is learned for 5 epochs. In this stage, only the loss of the non-semantic branch is calculated and the relevant model parameters are updated. Then, the two basic embeddings are trained with Loss2, and the semantic + non-semantic branches are learned simultaneously for 10 epochs. In this stage, only the model parameters of the semantic + non-semantic branches are updated. Subsequently, all branches (semantic + non-semantic + fused branches) are trained with Loss3. Using the SGD stochastic gradient descent method, the calculated loss is backward calculated for gradients and the updated values of all model parameters are obtained, and the model is updated. For example, when the model loss does not decrease for 10 consecutive rounds, the model training is stopped.
[0242] loss1 = l triplet1
[0243] loss2 = l triplet1 +l triplet2 +l class
[0244] loss3 = l triplet1 + l triplet2 + l triplet3 + l class
[0245] Among them, l triplet1 represents the non-semantic loss, l triplet2 represents the semantic loss, l class represents the classification loss, l triplet3 represents the fusion loss.
[0246] In this way, the fused embedding in this training process is fused in a delayed manner. After training a good basic embedding, the fused embedding is fine-tuned, which can avoid the overall poor effect caused by directly fusing two basic embeddings with inconsistent convergence.
[0247] 4. Model Application
[0248] The trained model can be used to extract the non-semantic embedding, semantic embedding, and fused embedding of the input image. Referring to Figure 8 , first use the model to extract the embeddings of each image in the image library, and then determine the clustering centers of the embeddings. During retrieval, find the nearest n clustering centers (i.e., target clustering centers) according to the embedding of the query image, obtain the associated images of these clustering centers as candidate images, calculate the Euclidean distance between the embeddings of each candidate image and the embedding of the query image, and sort them from small to large. Obtain the top K images in the sorting as the final retrieval result.
[0249] Single retrieval system: This model combines two embeddings into a fused embedding, so that only one model inference is required during application to obtain an embedding containing semantic and non-semantic information. Under the fused embedding, only one retrieval system needs to be established, which can reduce the storage space of the retrieval system. It is very helpful in the case of limited inference or retrieval resources.
[0250] Dual retrieval system: A dual retrieval system can also be established based on the semantic and non-semantic embeddings obtained from model inference. Cluster the two features separately, establish an index system using the corresponding clustering centers respectively, retrieve separately, and then directly merge the retrieval results.
[0251] It can be understood that in the image search scenario and image recommendation scenario, a single retrieval system or a dual retrieval system can be established.
[0252] This embodiment has the following beneficial effects:
[0253] 1. Low inference time consumption and low storage resources: A unified model is established to learn two types of features, and the combined features are learned simultaneously. The advantage of the combined features is to borrow a feature vector that can express both semantic information and non-semantic information, and ultimately enable a model to output a high-quality and information-rich feature.
[0254] 2. Flexible single or dual-feature retrieval application system: Since both semantic and non-semantic information are reflected in a fused embedding, a retrieval system can be established only for this embedding. On the other hand, when the application service can support a larger retrieval memory, a dual-feature retrieval system can also be established by leveraging the dual embeddings of this model. In short, a retrieval system can be flexibly established using this model.
[0255] 3. The convergence performance of multi-task learning can be ensured through multi-stage training and fine-tuning of the combined features.
[0256] It should be understood that although Figure 2 , 4 , and the steps in the flowcharts of 5 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless specifically stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 , 4 , and at least a part of the steps in 5 may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in rotation with at least a part of the steps or stages in other steps or other steps.
[0257] In one embodiment, as Figure 9 shown, a feature extraction model training device is provided. This device can be a software module, a hardware module, or a combination of both to form a part of a computer device. Specifically, the device includes: a first sample group processing module 902, a feature output module 904, a feature fusion module 906, a feature loss determination module 908, a classification loss determination module 910, and a model parameter adjustment module 912, where:
[0258] The first sample group processing module 902 is configured to obtain a first sample group and input each sample in the first sample group into an initial feature extraction model; the first sample group includes target samples, reference samples corresponding to the target samples, and class labels corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network;
[0259] A feature output module 904, configured to output an initial classification feature and an initial semantic feature through a sample classification network, and output an initial non-semantic feature through a non-semantic feature extraction network;
[0260] A feature fusion module 906, configured to fuse the initial semantic feature and the initial non-semantic feature of the same sample through a feature fusion network to obtain an initial fusion feature corresponding to each sample;
[0261] A feature loss determination module 908, configured to calculate a loss based on the initial semantic feature, the initial non-semantic feature, and the initial fusion feature corresponding to a target sample and a reference sample to obtain a feature loss;
[0262] A classification loss determination module 910, configured to calculate a loss based on the initial classification feature and the class label corresponding to the same sample to obtain a classification loss;
[0263] A model parameter adjustment module 912, configured to adjust the model parameters of the initial feature extraction model based on the feature loss and the classification loss until a convergence condition is met to obtain a target feature extraction model; the target feature extraction model is used to extract the sample features of the input sample, and the sample features are used for sample retrieval.
[0264] In one embodiment, the sample classification network includes a semantic feature extraction sub-network and a semantic feature classification sub-network, and there is a shared network layer between the semantic feature extraction sub-network and the non-semantic feature extraction network. The feature output module is further configured to perform convolution processing on the input sample through the shared network layer to obtain a shared feature; perform feature processing on the shared feature through the feature processing layer of the semantic feature extraction sub-network to obtain an initial semantic feature; perform classification processing on the initial semantic feature through the semantic feature classification sub-network to obtain an initial classification feature; perform feature processing on the shared feature through the feature processing layer of the non-semantic feature extraction network to obtain an initial non-semantic feature.
[0265] In one embodiment, the reference sample includes a positive sample and a negative sample corresponding to the target sample. The feature loss determination module is further configured to obtain a positive semantic loss, a positive non-semantic loss, and a positive fusion loss based on the distance between the same type of features corresponding to the target sample and the positive sample; obtain a negative semantic loss, a negative non-semantic loss, and a negative fusion loss based on the distance between the same type of features corresponding to the target sample and the negative sample; obtain an initial semantic loss based on the distance between the positive semantic loss and the negative semantic loss, obtain an initial non-semantic loss based on the distance between the positive non-semantic loss and the negative non-semantic loss, obtain an initial fusion loss based on the distance between the positive fusion loss and the negative fusion loss; obtain a feature loss based on the initial semantic loss, the initial non-semantic loss, and the initial fusion loss.
[0266] In one embodiment, the feature loss determination module is further configured to update the initial semantic loss based on the semantic loss adjustment parameter to obtain an intermediate semantic loss, and determine the target semantic loss based on the matching result between the intermediate semantic loss and the preset parameter; update the initial non-semantic loss based on the non-semantic loss adjustment parameter to obtain an intermediate non-semantic loss, and determine the target non-semantic loss based on the matching result between the intermediate non-semantic loss and the preset parameter; update the initial fusion loss based on the fusion loss adjustment parameter to obtain an intermediate fusion loss, and determine the target fusion loss based on the matching result between the intermediate fusion loss and the preset parameter; the fusion loss adjustment parameter is greater than the semantic loss adjustment parameter, and the fusion loss adjustment parameter is greater than the non-semantic loss adjustment parameter; generate a feature loss based on the target semantic loss, the target non-semantic loss, and the target fusion loss.
[0267] In one embodiment, the classification loss determination module is further configured to perform label encoding on the class labels of each sample to obtain corresponding label features; perform normalization processing on the classification features corresponding to each sample to obtain corresponding normalized features; perform logarithmic transformation on each normalized feature, and fuse the label features corresponding to the same sample and the logarithmically transformed normalized features to obtain classification sub-losses corresponding to each sample; obtain a classification loss based on each classification sub-loss.
[0268] In one embodiment, the sample classification network includes a semantic feature extraction sub-network and a semantic feature classification sub-network. The model parameter adjustment module is further configured to obtain a target loss based on the feature loss and the classification loss, calculate a loss gradient for the target loss, obtain a first loss by updating the loss gradient based on a first adjustment parameter, and obtain a second loss by updating the loss gradient based on a second adjustment parameter; the first adjustment parameter is less than the second adjustment parameter; adjust the network parameters of the semantic feature classification sub-network based on the first loss, and adjust the network parameters of other networks based on the second loss until the convergence condition is met to obtain a target feature extraction model.
[0269] In one embodiment, the feature extraction model training device further includes:
[0270] A second sample group processing module, configured to obtain a second sample group, input each sample in the second sample group into a candidate feature extraction model to obtain a candidate non-semantic feature set, a candidate semantic feature set, and a candidate fusion feature set corresponding to the second sample group; calculate a loss based on the candidate non-semantic feature set, the candidate semantic feature set, and the candidate fusion feature set to obtain a candidate loss; adjust the network parameters of the target network in the candidate feature extraction model based on the candidate loss until the first condition is met to obtain an initial feature extraction model; the target network includes a sample classification network and a non-semantic feature extraction network.
[0271] In one embodiment, the feature extraction model training device further includes:
[0272] The third sample group processing module is used to obtain a third sample group, input each sample in the third sample group into the feature extraction model to be trained, and obtain a non-semantic feature set corresponding to the third sample group; calculate a loss based on the non-semantic feature set to obtain an initial loss; based on the initial loss, adjust the model parameters of the non-semantic feature extraction network in the feature extraction model to be trained until the second condition is satisfied, and obtain a candidate feature extraction model.
[0273] In one embodiment, the current sample group is any one of the first sample group, the second sample group, and the third sample group. The first sample group processing module, the second sample group processing module, and the third sample group processing module are also used to obtain multiple similar sample pairs; determine a current sample and a positive sample corresponding to the current sample from the current similar sample pairs, and determine multiple candidate samples from the remaining similar sample pairs; based on the sample similarity between the current sample and each candidate sample, determine at least one negative sample corresponding to the current sample from each candidate sample; use the positive sample and the negative sample corresponding to the current sample as the reference samples corresponding to the current sample, and obtain at least one current sample group based on the current sample and the corresponding reference samples.
[0274] In one embodiment, the first sample group processing module, the second sample group processing module, and the third sample group processing module are also used to input the current sample and each candidate sample into the corresponding current feature extraction model to obtain a sample feature set corresponding to the current sample and each candidate sample respectively; the current feature extraction model matched by the first sample group is the initial feature extraction model, the current feature extraction model matched by the second sample group is the candidate feature extraction model, and the current feature extraction model matched by the third sample group is the feature extraction model to be trained; calculate the sample similarity between the current sample and each candidate sample based on the sample feature sets corresponding to the current sample and each candidate sample; divide each candidate sample into a first type of sample and a second type of sample based on the sample similarity; the sample similarity corresponding to the first type of sample is greater than the sample similarity corresponding to the second type of sample; determine at least one negative sample corresponding to the current sample from the first type of samples.
[0275] In one embodiment, the feature sizes of the initial semantic features, the initial non-semantic features, and the initial fusion features are the same.
[0276] The above feature extraction model training device establishes a unified model to learn semantic features and non-semantic features, and simultaneously learns fusion features. The fusion features contain both semantic information and non-semantic information. The finally trained model can not only output semantic features and non-semantic features containing single-dimensional information, but also output fusion features containing two-dimensional information. Only by training one model can this model output diverse features, improving the training efficiency.
[0277] In one embodiment, asFigure 10 As shown, a sample retrieval device is provided. This device can be implemented as a software module, a hardware module, or a combination of both to be part of a computer device. Specifically, the device includes: a data acquisition module 1002, a data processing module 1004, and a retrieval result determination module 1006, where:
[0278] The data acquisition module 1002 is used to obtain a query sample and a candidate recall sample set.
[0279] The data processing module 1004 is used to input the candidate recall samples in the query sample and the candidate recall sample set into a target feature extraction model to obtain a query sample feature corresponding to the query sample and a recall sample feature corresponding to the candidate recall sample.
[0280] The retrieval result determination module 1006 is used to determine a retrieval result sample corresponding to the query sample from the candidate recall sample set based on the query sample feature and the recall sample feature.
[0281] The training process of the target feature extraction model is as follows:
[0282] Obtain a first sample group and input each sample in the first sample group into an initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and a class label corresponding to each sample. The initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network; output an initial classification feature and an initial semantic feature through the sample classification network, output an initial non-semantic feature through the non-semantic feature extraction network, and through the feature fusion network, fuse the initial semantic feature and the initial non-semantic feature of the same sample to obtain an initial fusion feature corresponding to each sample; calculate a loss based on the initial semantic feature, the initial non-semantic feature, and the initial fusion feature corresponding to the target sample and the reference sample to obtain a feature loss, and calculate a loss based on the initial classification feature and the class label corresponding to the same sample to obtain a classification loss; based on the feature loss and the classification loss, adjust the model parameters of the initial feature extraction model until the convergence condition is met to obtain the target feature extraction model.
[0283] In one embodiment, the recall sample feature includes at least one of a semantic feature, a non-semantic feature, and a fusion feature, and the query sample feature and the recall sample feature include the same type of sample feature. The retrieval result determination module includes:
[0284] An index building unit for building an index based on each recall sample feature corresponding to the same type to obtain a sample recall index corresponding to each type.
[0285] A sample retrieval unit is configured to retrieve samples from a corresponding sample recall index based on query sample features corresponding to the same type, and determine retrieved result samples based on the sample retrieval results.
[0286] In one embodiment, the index building unit is further configured to perform feature clustering on respective recall sample features corresponding to the same type to obtain multiple clustering clusters respectively corresponding to each type; there are corresponding clustering centers for each clustering cluster; based on the respective clustering clusters corresponding to the same type, obtain sample recall indexes respectively corresponding to each type.
[0287] In one embodiment, the sample retrieval unit is further configured to determine target clustering clusters from the corresponding respective clustering clusters based on the distances between the query sample features corresponding to the same type and each clustering center, to obtain target clustering clusters respectively corresponding to each type; determine target sample features from the respective recall sample features corresponding to the target clustering clusters based on the distances between the query sample features corresponding to the same type and each recall sample feature within the target clustering clusters; use the candidate recall samples corresponding to the target sample features as the retrieved result samples.
[0288] In one embodiment, when the query task corresponding to the query sample is a first type of task, the recall sample features include fusion features, and when the query task corresponding to the query sample is a second type of task, the recall sample features include semantic features and non-semantic features, and the query frequency of the first type of task is greater than that of the second type of task.
[0289] For the above sample retrieval device, when training the model, a unified model is established to learn semantic features and non-semantic features, and at the same time learn fusion features. The fusion features contain both semantic information and non-semantic information. The finally trained model can not only output semantic features and non-semantic features containing single-dimensional information, but also output fusion features containing two-dimensional information. Only by training one model can this model output diverse features, improving the training efficiency. Thus, when performing sample retrieval, sample retrieval is performed through the diverse features output by the model, which can improve the accuracy and efficiency of sample retrieval.
[0290] For the specific limitations of the feature extraction model training device and the sample retrieval device, reference can be made to the limitations of the feature extraction model training method and the sample retrieval method in the above text, which will not be elaborated here. Each module in the above feature extraction model training device and sample retrieval device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0291] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structural diagram may be as shown in Figure 11 . The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as similar sample pairs, candidate recall sample sets, and target feature extraction models. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a feature extraction model training method and a sample retrieval method.
[0292] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structural diagram may be as shown in Figure 12 . The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a feature extraction model training method and a sample retrieval method. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0293] Those skilled in the art can understand that Figure 11 , 12 The structures shown in are only block diagrams of some structures related to the solution of this application, and do not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0294] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0295] In one embodiment, a computer-readable storage medium is provided, storing a computer program which, when executed by a processor, implements the steps in the above method embodiments.
[0296] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions which are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.
[0297] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the various embodiments provided in the present application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0298] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0299] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for training a feature extraction model, characterized in that, The method includes: Obtaining a first sample group and inputting each sample in the first sample group into an initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and a class label corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network; the reference sample corresponding to the target sample includes at least one of a positive sample and a negative sample corresponding to the target sample; Outputting an initial classification feature and an initial semantic feature through the sample classification network, and outputting an initial non-semantic feature through the non-semantic feature extraction network; the initial classification feature is obtained by classifying the initial semantic feature, and the initial classification feature is a feature output by the sample classification network for characterizing the sample class; Fusing the initial semantic feature and the initial non-semantic feature of the same sample through the feature fusion network to obtain an initial fusion feature corresponding to each sample; Calculating a loss based on the initial semantic feature, the initial non-semantic feature, and the initial fusion feature corresponding to the target sample and the reference sample to obtain a feature loss, including: obtaining a semantic feature loss based on the distance between the initial semantic features corresponding to the target sample and the reference sample, obtaining a non-semantic feature loss based on the distance between the initial non-semantic features corresponding to the target sample and the reference sample, obtaining a fusion feature loss based on the distance between the initial fusion features corresponding to the target sample and the reference sample, and obtaining a feature loss based on the semantic feature loss, the non-semantic feature loss, and the fusion feature loss; Calculating a loss based on the initial classification feature and the class label corresponding to the same sample to obtain a classification loss, including: performing label encoding on the class labels of each sample to obtain corresponding label features, and obtaining a classification loss based on the difference between the initial classification feature and the label features corresponding to the same sample; Adjusting the model parameters of the initial feature extraction model based on the feature loss and the classification loss until a convergence condition is met to obtain a target feature extraction model; the target feature extraction model is used to extract sample features of input samples, and the sample features are used for sample retrieval.
2. The method according to claim 1, wherein The sample classification network includes a semantic feature extraction sub-network and a semantic feature classification sub-network, and there is a shared network layer between the semantic feature extraction sub-network and the non-semantic feature extraction network. Outputting an initial classification feature and an initial semantic feature through the sample classification network, and outputting an initial non-semantic feature through the non-semantic feature extraction network includes: Performing convolution processing on the input sample through the shared network layer to obtain a shared feature; Performing feature processing on the shared feature through the feature processing layer of the semantic feature extraction sub-network to obtain the initial semantic feature; Performing classification processing on the initial semantic feature through the semantic feature classification sub-network to obtain the initial classification feature; Performing feature processing on the shared feature through the feature processing layer of the non-semantic feature extraction network to obtain the initial non-semantic feature.
3. The method according to claim 1, wherein Calculating a loss based on the initial semantic features, initial non-semantic features, and initial fusion features corresponding to the target sample and the reference sample to obtain a feature loss, further comprising: Obtaining a positive semantic loss, a positive non-semantic loss, and a positive fusion loss based on the distances between the same type of features corresponding to the target sample and the positive sample; Obtaining a negative semantic loss, a negative non-semantic loss, and a negative fusion loss based on the distances between the same type of features corresponding to the target sample and the negative sample; Obtaining an initial semantic loss based on the distance between the positive semantic loss and the negative semantic loss, obtaining an initial non-semantic loss based on the distance between the positive non-semantic loss and the negative non-semantic loss, and obtaining an initial fusion loss based on the distance between the positive fusion loss and the negative fusion loss; Obtaining the feature loss based on the initial semantic loss, the initial non-semantic loss, and the initial fusion loss.
4. The method according to claim 3, wherein The obtaining the feature loss based on the initial semantic loss, the initial non-semantic loss, and the initial fusion loss includes: Updating the initial semantic loss based on a semantic loss adjustment parameter to obtain an intermediate semantic loss, and determining a target semantic loss based on the matching result between the intermediate semantic loss and a preset parameter; Updating the initial non-semantic loss based on a non-semantic loss adjustment parameter to obtain an intermediate non-semantic loss, and determining a target non-semantic loss based on the matching result between the intermediate non-semantic loss and a preset parameter; Updating the initial fusion loss based on a fusion loss adjustment parameter to obtain an intermediate fusion loss, and determining a target fusion loss based on the matching result between the intermediate fusion loss and a preset parameter; the fusion loss adjustment parameter is greater than the semantic loss adjustment parameter, and the fusion loss adjustment parameter is greater than the non-semantic loss adjustment parameter; Generating the feature loss based on the target semantic loss, the target non-semantic loss, and the target fusion loss.
5. The method according to claim 1, wherein Calculating a loss based on the initial classification features and class labels corresponding to the same sample to obtain a classification loss, further comprising: Performing label encoding on the class labels of each sample to obtain corresponding label features; Performing normalization processing on the classification features corresponding to each sample to obtain corresponding normalized features; Performing a logarithmic transformation on each normalized feature, fusing the label features corresponding to the same sample and the logarithmically transformed normalized features, to obtain classification sub-losses corresponding to each sample; Obtaining the classification loss based on each classification sub-loss.
6. The method according to claim 1, wherein The sample classification network includes a semantic feature extraction sub-network and a semantic feature classification sub-network. Adjusting the model parameters of the initial feature extraction model based on the feature loss and the classification loss until a convergence condition is met to obtain a target feature extraction model, including: Obtaining a target loss based on the feature loss and the classification loss, calculating a gradient of the target loss to obtain a loss gradient; Updating the loss gradient based on a first adjustment parameter to obtain a first loss, and updating the loss gradient based on a second adjustment parameter to obtain a second loss; the first adjustment parameter is less than the second adjustment parameter; Adjust the network parameters of the semantic feature classification sub-network based on the first loss, and adjust the network parameters of other networks based on the second loss until the convergence condition is met, thereby obtaining the target feature extraction model.
7. The method according to any one of claims 1 to 6, characterized in that Before obtaining the first sample group and inputting each sample in the first sample group into the initial feature extraction model, the method further includes: Obtain a second sample group, and input each sample in the second sample group into the candidate feature extraction model to obtain a candidate non-semantic feature set, a candidate semantic feature set, and a candidate fusion feature set corresponding to the second sample group; Calculate a loss based on the candidate non-semantic feature set, the candidate semantic feature set, and the candidate fusion feature set to obtain a candidate loss; the calculation method of the candidate loss is the same as that of the feature loss; Based on the candidate loss, adjust the network parameters of the target network in the candidate feature extraction model until the first condition is met, thereby obtaining the initial feature extraction model; the target network includes the sample classification network and the non-semantic feature extraction network.
8. The method according to claim 7, wherein Before obtaining the second sample group and inputting each sample in the second sample group into the candidate feature extraction model, the method further includes: Obtain a third sample group, and input each sample in the third sample group into the feature extraction model to be trained to obtain a non-semantic feature set corresponding to the third sample group; Calculate a loss based on the non-semantic feature set to obtain an initial loss; the calculation method of the initial loss is the same as that of the non-semantic feature loss; Based on the initial loss, adjust the model parameters of the non-semantic feature extraction network in the feature extraction model to be trained until the second condition is met, thereby obtaining the candidate feature extraction model.
9. The method according to claim 8, characterized in that, When the current sample group is any one of the first sample group, the second sample group, and the third sample group, obtaining the current sample group includes: Obtain a plurality of similar sample pairs; Determine the current sample and the positive sample corresponding to the current sample from the current similar sample pairs, and determine a plurality of candidate samples from the remaining similar sample pairs; Based on the sample similarity between the current sample and each candidate sample, determine at least one negative sample corresponding to the current sample from each candidate sample; Use the positive sample and the negative sample corresponding to the current sample as the reference samples corresponding to the current sample, and obtain at least one current sample group based on the current sample and the corresponding reference samples.
10. The method according to claim 9, wherein The determining, based on the sample similarity between the current sample and each candidate sample, of at least one negative sample corresponding to the current sample from each candidate sample includes: Input the current sample and each candidate sample into the matching current feature extraction model to obtain a sample feature set corresponding to the current sample and each candidate sample respectively; the current feature extraction model matching the first sample group is the initial feature extraction model, the current feature extraction model matching the second sample group is the candidate feature extraction model, and the current feature extraction model matching the third sample group is the feature extraction model to be trained. Based on the sample feature sets corresponding to the current sample and each candidate sample, calculate the sample similarity between the current sample and each candidate sample respectively; Divide each candidate sample into a first type of sample and a second type of sample based on the sample similarity; the sample similarity corresponding to the first type of sample is greater than the sample similarity corresponding to the second type of sample; Determine at least one negative sample corresponding to the current sample from the first type of samples.
11. The method according to any one of claims 1 to 6, characterized in that, The feature sizes of the initial semantic feature, the initial non-semantic feature, and the initial fusion feature are the same.
12. A sample retrieval method, characterized in that, The method includes: Obtain a query sample and a candidate recall sample set; Input the candidate recall samples in the query sample and the candidate recall sample set into the target feature extraction model to obtain the query sample feature corresponding to the query sample and the recall sample feature corresponding to the candidate recall sample; Based on the query sample feature and the recall sample feature, determine the retrieval result sample corresponding to the query sample from the candidate recall sample set; The training process of the target feature extraction model is as follows: Obtain a first sample group, and input each sample in the first sample group into the initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and class labels corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network; the reference sample corresponding to the target sample includes at least one of a positive sample and a negative sample corresponding to the target sample; Output an initial classification feature and an initial semantic feature through the sample classification network, output an initial non-semantic feature through the non-semantic feature extraction network, and fuse the initial semantic feature and the initial non-semantic feature of the same sample through the feature fusion network to obtain an initial fusion feature corresponding to each sample respectively; the initial classification feature is obtained by classifying the initial semantic feature, and the initial classification feature is a feature output by the sample classification network for characterizing the category of the sample; Calculate a loss based on the initial semantic feature, initial non-semantic feature, and initial fusion feature corresponding to the target sample and the reference sample to obtain a feature loss, including: obtaining a semantic feature loss based on the distance between the initial semantic features corresponding to the target sample and the reference sample, obtaining a non-semantic feature loss based on the distance between the initial non-semantic features corresponding to the target sample and the reference sample, obtaining a fusion feature loss based on the distance between the initial fusion features corresponding to the target sample and the reference sample, and obtaining a feature loss based on the semantic feature loss, the non-semantic feature loss, and the fusion feature loss; Calculate a loss based on the initial classification feature and the class label corresponding to the same sample to obtain a classification loss, including: performing label encoding on the class labels of each sample to obtain corresponding label features, and obtaining a classification loss based on the difference between the initial classification feature and the label features corresponding to the same sample; Based on the feature loss and the classification loss, adjust the model parameters of the initial feature extraction model until the convergence condition is satisfied to obtain the target feature extraction model.
13. The method according to claim 12, wherein The recall sample features include at least one of semantic features, non-semantic features, and fusion features, and the query sample features and the recall sample features include the same type of sample features; Determining the retrieval result sample corresponding to the query sample from the candidate recall sample set based on the query sample features and the recall sample features includes: Establishing an index based on each recall sample feature corresponding to the same type to obtain a sample recall index corresponding to each type; Performing sample retrieval from the corresponding sample recall index based on the query sample features corresponding to the same type, and determining the retrieval result sample based on the sample retrieval result.
14. The method according to claim 13, characterized in that, Establishing an index based on each recall sample feature corresponding to the same type to obtain a sample recall index corresponding to each type includes: Performing feature clustering on each recall sample feature corresponding to the same type to obtain multiple clustering clusters corresponding to each type; there is a corresponding clustering center for each clustering cluster; Based on each clustering cluster corresponding to the same type, obtaining the sample recall index corresponding to each type.
15. The method according to claim 14, characterized in that, Performing sample retrieval from the corresponding sample recall index based on the query sample features corresponding to the same type, and determining the retrieval result sample based on the sample retrieval result includes: Determining a target clustering cluster from each corresponding clustering cluster based on the distances between the query sample features corresponding to the same type and each clustering center, to obtain a target clustering cluster corresponding to each type; Determining target sample features from each recall sample feature corresponding to the target clustering cluster based on the distances between the query sample features corresponding to the same type and each recall sample feature within the target clustering cluster; Taking the candidate recall sample corresponding to the target sample feature as the retrieval result sample.
16. The method according to claim 13, wherein When the query task corresponding to the query sample is a first type of task, the recall sample features include fusion features, and when the query task corresponding to the query sample is a second type of task, the recall sample features include semantic features and non-semantic features, and the query frequency of the first type of task is greater than the query frequency of the second type of task.
17. A feature extraction model training device, characterized in that The device includes: A first sample group processing module, configured to obtain a first sample group and input each sample in the first sample group into an initial feature extraction model; the first sample group includes a target sample, a reference sample corresponding to the target sample, and a class label corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network; the reference sample corresponding to the target sample includes at least one of a positive sample and a negative sample corresponding to the target sample; A feature output module, configured to output an initial classification feature and an initial semantic feature through the sample classification network, and output an initial non-semantic feature through the non-semantic feature extraction network; the initial classification feature is obtained by classifying the initial semantic feature, and the initial classification feature is a feature output by the sample classification network for characterizing the sample category; A feature fusion module, which is used to fuse the initial semantic features and initial non-semantic features of the same sample through the feature fusion network to obtain the initial fusion features corresponding to each sample; A feature loss determination module, which is used to calculate the loss based on the initial semantic features, initial non-semantic features, and initial fusion features corresponding to the target sample and the reference sample to obtain the feature loss, including: obtaining the semantic feature loss based on the distance between the initial semantic features corresponding to the target sample and the reference sample, obtaining the non-semantic feature loss based on the distance between the initial non-semantic features corresponding to the target sample and the reference sample, obtaining the fusion feature loss based on the distance between the initial fusion features corresponding to the target sample and the reference sample, and obtaining the feature loss based on the semantic feature loss, the non-semantic feature loss, and the fusion feature loss; A classification loss determination module, which is used to calculate the loss based on the initial classification features and class labels corresponding to the same sample to obtain the classification loss, including: performing label encoding on the class labels of each sample to obtain the corresponding label features, and obtaining the classification loss based on the difference between the initial classification features and the label features corresponding to the same sample; A model parameter adjustment module, which is used to adjust the model parameters of the initial feature extraction model based on the feature loss and the classification loss until the convergence condition is met to obtain the target feature extraction model; the target feature extraction model is used to extract the sample features of the input sample, and the sample features are used for sample retrieval.
18. A sample retrieval device, characterized in that, The device includes: A data acquisition module, which is used to acquire the query sample and the candidate recall sample set; A data processing module, which is used to input the candidate recall samples in the query sample and the candidate recall sample set into the target feature extraction model to obtain the query sample features corresponding to the query sample and the recall sample features corresponding to the candidate recall samples; A retrieval result determination module, which is used to determine the retrieval result sample corresponding to the query sample from the candidate recall sample set based on the query sample features and the recall sample features; The training process of the target feature extraction model is as follows: Obtain the first sample group and input each sample in the first sample group into the initial feature extraction model; the first sample group includes the target sample, the reference sample corresponding to the target sample, and the class labels corresponding to each sample, and the initial feature extraction model includes a sample classification network, a non-semantic feature extraction network, and a feature fusion network; the reference sample corresponding to the target sample includes at least one of the positive sample and the negative sample corresponding to the target sample; Output the initial classification features and initial semantic features through the sample classification network, output the initial non-semantic features through the non-semantic feature extraction network, and fuse the initial semantic features and initial non-semantic features of the same sample through the feature fusion network to obtain the initial fusion features corresponding to each sample; the initial classification features are obtained by classifying the initial semantic features, and the initial classification features are the features output by the sample classification network for characterizing the sample category; Calculate a loss based on the initial semantic features, initial non-semantic features, and initial fusion features corresponding to the target sample and the reference sample to obtain a feature loss, including: obtaining a semantic feature loss based on the distance between the initial semantic features corresponding to the target sample and the reference sample, obtaining a non-semantic feature loss based on the distance between the initial non-semantic features corresponding to the target sample and the reference sample, obtaining a fusion feature loss based on the distance between the initial fusion features corresponding to the target sample and the reference sample, and obtaining a feature loss based on the semantic feature loss, the non-semantic feature loss, and the fusion feature loss; Calculate a loss based on the initial classification features and class labels corresponding to the same sample to obtain a classification loss, including: performing label encoding on the class labels of each sample to obtain corresponding label features, and obtaining a classification loss based on the difference between the initial classification features and the label features corresponding to the same sample; Adjust the model parameters of the initial feature extraction model based on the feature loss and the classification loss until the convergence condition is met to obtain a target feature extraction model.
19. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 16.
20. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 16.
21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Point cloud semantic segmentation method based on deep learning
CN111507982A
Street abnormal event detection method and device, equipment and medium
CN113515968A