Image processing method and device, computer equipment and computer readable storage medium
By using the target image recognition model in image processing for feature extraction and comparison, the problem of low identity recognition efficiency in the prior art is solved, and fast and accurate identity confirmation is achieved.
Patent Information
- Application Number
- CN202311504956.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to achieve efficient identity recognition, especially in image processing, and it is difficult to quickly and accurately determine the identity information in the image.
By acquiring the image information to be processed, performing feature extraction processing based on the target image recognition model, obtaining the recognition feature information, and comparing it with the preset image feature information, determining the image recognition result information, thereby confirming the identity information.
This method significantly improves the confirmation efficiency of identity information and can quickly and accurately identify and confirm the identity information in the image.
Smart Images

Figure CN119992614A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image processing method, apparatus, computer equipment and computer-readable storage medium. Background Art
[0002] With the development of society, effective identity recognition is required in more and more scenarios. Therefore, how to perform efficient identity recognition has become a problem that needs to be solved. Summary of the invention
[0003] The present application provides an image processing method that can compare faces and thereby determine corresponding identity information.
[0004] In a first aspect, the present application provides an image processing method, the method comprising:
[0005] Obtaining image information to be processed;
[0006] Performing feature extraction processing on the image information to be processed based on the target image recognition model to obtain recognition feature information;
[0007] Image recognition result information is determined based on the image feature information and the recognition feature information.
[0008] In a second aspect, the present application further provides an image processing device, the device comprising:
[0009] An acquisition module, used for acquiring image information to be processed;
[0010] An extraction module, used for performing feature extraction processing on the image information to be processed based on a target image recognition model to obtain recognition feature information;
[0011] The determination module is used to determine the image recognition result information based on the image feature information and the recognition feature information.
[0012] In a third aspect, the present application also provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in any one of the image processing methods.
[0013] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the steps in any one of the image processing methods described.
[0014] The image processing method provided by the present application compares the recognition feature information with the preset image feature information, so that the identity information of the image information to be processed can be determined based on the identity information of the preset image feature information matching the recognition feature information, thereby greatly improving the efficiency of confirming the identity information. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 is a schematic diagram of a scene of an image processing system provided in an embodiment of the present application;
[0017] Figure 2 is a schematic diagram of an embodiment of the image processing method in the embodiment of the present application;
[0018] Figure 3 is a schematic diagram of a functional module of an image processing device in an embodiment of the present application;
[0019] Figure 4 It is a schematic diagram of the structure of the terminal device in the embodiment of the present application. DETAILED DESCRIPTION
[0020] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0021] In the description of the present application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the feature. In the description of the present application, "plurality" means two or more, unless otherwise clearly and specifically defined.
[0022] In this application, the word "exemplary" is used to mean "used as an example, illustration or description". Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. At the same time, it is to be understood that in the specific implementation of this application, when user information, user data and other related data are involved, when the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data shall comply with relevant laws, regulations and standards of relevant countries and regions.
[0023] In order to enable any person skilled in the art to implement and use the present application, the following description is provided. In the following description, details are listed for the purpose of explanation. It should be understood that those of ordinary skill in the art will recognize that the present application can be implemented without using these specific details. In other examples, known structures and processes will not be elaborated in detail to avoid unnecessary details that make the description of the present application obscure. Therefore, the present application is not intended to be limited to the embodiments shown, but is consistent with the widest range of principles and features disclosed in the present application.
[0024] The present application provides an image processing method, apparatus, device and storage medium, which are described in detail below.
[0025] See also Figure 1 , Figure 1 Schematic diagram of a scene of an image processing system provided in an embodiment of the present application. The image processing system may include a terminal device 100 and a storage device 200. The storage device 200 may transmit data to the terminal device 100. Figure 1 The terminal device 100 in the embodiment can obtain the image data stored in the storage device 200 to execute the image processing method in the present application.
[0026] In the embodiment of the present application, the terminal device 100 includes but is not limited to a desktop computer, a portable computer, a network server, a PDA (Personal Digital Assistant), a tablet computer, a wireless terminal device, an embedded device, etc.
[0027] In an embodiment of the present application, communication between the terminal device 100 and the storage device 200 can be achieved through any communication method, including but not limited to mobile communications based on the 3rd Generation Partnership Project (3GPP), Long Term Evolution (LTE), and Worldwide Interoperability for Microwave Access (WiMAX), or computer network communications based on the TCP / IP Protocol Suite (TCP / IP) and User Datagram Protocol (UDP).
[0028] It should be noted that Figure 1 The scene diagram of the image processing system shown is only an example. The image processing system and scene described in the embodiment of the present application are intended to more clearly illustrate the technical solution of the embodiment of the present application, and do not constitute a limitation on the technical solution provided in the embodiment of the present application. Ordinary technicians in this field can know that with the evolution of the image processing system and the emergence of new business scenarios, the technical solution provided in the embodiment of the present application is also applicable to similar technical problems.
[0029] like Figure 2 As shown, Figure 2 This is a flow chart of an embodiment of an image processing method in an embodiment of the present application. The image processing method may include the following steps 201 to 203:
[0030] 201. Obtain image information to be processed.
[0031] In the embodiment of the present application, the image information to be processed can be acquired by a preset image sensor. When the infrared sensor of the image sensor detects that there is a living body approaching, the image sensor can start shooting, thereby acquiring the image information to be processed.
[0032] 202. Perform feature extraction processing on the image information to be processed based on the target image recognition model to obtain recognition feature information.
[0033] In the embodiment of the present application, after the image information to be processed is obtained, the image information to be processed can be input into the target image recognition model, and the target image recognition model is used to extract features of the image information to be processed to obtain recognition feature information. It should be noted that in the embodiment of the present application, any prior art can be referred to for extracting recognition feature information, and the specific details will not be repeated here.
[0034] 203. Determine image recognition result information based on the image feature information and the recognition feature information.
[0035] In an embodiment of the present application, the image feature information may be image feature information that is stored in advance. After obtaining the recognition feature information of the image information to be processed, the recognition feature information may be compared with the stored image feature information to obtain image recognition result information. When the recognition feature information is similar to the stored image feature information, the image recognition result information may indicate that the identity has been passed. If the recognition feature information is not similar to the stored image feature information, the image recognition result information may indicate that the identity has not been passed. Among them, the similarity calculation between the recognition feature information and the stored image feature information may be calculated by cosine similarity, which will not be described in detail here.
[0036] The image processing method provided by the present application compares the recognition feature information with the preset image feature information, so that the identity information of the image information to be processed can be determined based on the identity information of the preset image feature information matching the recognition feature information, thereby greatly improving the efficiency of confirming the identity information.
[0037] In order to better implement the embodiments of the present application, in one embodiment of the present application, the image feature information includes a number of basic feature information; based on the image feature information and the recognition feature information, the image recognition result information is determined, including:
[0038] Determine whether there is basic feature information matching the recognition feature information in the image feature information to obtain a first judgment result; when the first judgment result is yes, determine that the image recognition result information is the first recognition result; when the first judgment result is no, determine that the image recognition result information is the second recognition result.
[0039] In the embodiment of the present application, the pre-stored image feature information may include multiple basic feature information. After the identification feature information is obtained, each basic feature information in the image feature information can be compared with the identification feature information to obtain the similarity between each basic feature information and the identification feature information. If there is a basic feature information with the highest similarity to the identification feature information, and it is greater than the preset similarity threshold, the first judgment result is yes, and the first recognition result is obtained. If there is no basic feature information with the identification feature information whose similarity is greater than the preset similarity threshold, the first judgment result is no, and the second recognition result is obtained.
[0040] For example: if in a face recognition scenario, multiple facial features can be pre-stored; when a current face image is obtained, feature extraction is performed on the current face image to obtain face recognition feature information. At this time, the face recognition feature information can be respectively calculated for similarity with each pre-stored face feature. If there is a pre-set facial feature with the highest similarity to the face recognition feature information and it is greater than the similarity threshold, it can be determined that the face recognition has passed, that is, the first judgment result of the first recognition result can be obtained. If there is no pre-set facial feature with the face recognition feature information with a similarity greater than the similarity threshold, it can be determined that the face recognition has not passed, that is, the first judgment result of the second recognition result can be obtained.
[0041] Alternatively, if in an object recognition scene, multiple object features can be pre-stored, including water features, tree features, stone features, etc.; when a current object image is obtained, feature extraction is performed on the current object image to obtain the object's recognition feature information. At this time, the object's recognition feature information can be respectively calculated for similarity with the pre-stored object features. If there is a pre-existing object feature of a stone whose similarity with the object's recognition feature information is the highest and is greater than the similarity threshold, it can be determined that the object recognition has passed, that is, the first judgment result of the first recognition result can be obtained, and the first recognition result can characterize the currently recognized object as a stone. If there is no pre-existing object feature whose similarity with the object's recognition feature information is greater than the similarity threshold, it can be determined that the object recognition has not passed, that is, the first judgment result of the second recognition result can be obtained.
[0042] Alternatively, if in an animal recognition scenario, multiple animal features can be pre-stored, including monkey features, insect features, bird features, etc.; when a current animal image is obtained, feature extraction is performed on the current animal image to obtain the animal's recognition feature information. At this time, the animal's recognition feature information can be calculated for similarity with the pre-stored features of each object. If there is a pre-existing animal feature of a bird whose similarity to the animal's recognition feature information is the highest and is greater than the similarity threshold, it can be determined that the animal recognition has passed, that is, the first judgment result of the first recognition result can be obtained, and the first recognition result can characterize the currently recognized animal as a bird. If there is no pre-existing animal feature whose similarity to the animal's recognition feature information is greater than the similarity threshold, it can be determined that the object recognition has not passed, that is, the first judgment result of the second recognition result can be obtained.
[0043] In order to better implement the embodiment of the present application, in one embodiment of the present application, it is determined whether there is basic feature information matching the identification feature information in the image feature information to obtain a first determination result, including:
[0044] Determine one basic feature information as the feature information to be processed from the image feature information in sequence; calculate the feature distance between the feature information to be processed and the identification feature information to obtain the identification feature distance value; determine whether the identification feature distance value is less than the distance threshold to obtain a second judgment result; when the second judgment result is yes, determine the first judgment result to be yes, and end the judgment process corresponding to the first judgment result; when the second judgment result is no, determine whether all the basic feature information in the image feature information are calculated with the identification feature information for feature distance, and obtain a third judgment result; when the third judgment result is yes, determine the first judgment result to be no, and end the judgment process corresponding to the first judgment result; when the third judgment result is no, trigger the execution of determining one basic feature information as the feature information to be processed from the image feature information in sequence.
[0045] The above embodiment provides an implementation scheme for comparing each basic feature information in the image feature information with the identification feature information, and the above scheme can simultaneously perform the process of comparing each basic feature information in the image feature information with the identification feature information when performing feature comparison. In the embodiment of the present application, an implementation scheme is provided for sequentially determining a basic feature information from the image feature information as the feature information to be processed.
[0046] Specifically, compared to the identification process of each basic feature information in the image feature information and the identification feature information at the same time in the above-mentioned embodiment, the advantage of performing feature comparison in sequence is that if there is a basic feature information that matches the identification feature information, and the basic feature information that matches the identification feature information is not the last basic feature information in the comparison order, the subsequent comparison process can no longer be compared to save computing resources. For example: if in the sequential comparison process, there is a basic feature information and the identification feature information whose similarity is greater than a preset second similarity threshold, and this second similarity threshold is greater than the previous similarity threshold, then it can be determined that the current basic feature information and the identification feature information are completely matched, that is, there is no need to perform subsequent comparison calculations, and a second judgment result of yes is obtained, otherwise a second judgment result of no is obtained. In addition, in the embodiment of the present application, if all the basic feature information in the image feature information completes the feature distance calculation with the identification feature information, a third judgment result of yes can be obtained, otherwise a third judgment result of no is obtained. It should be noted that, in the embodiment of the present application, when sorting the various basic feature information, they can be randomly sorted, so as to obtain a comparison order, which is convenient for subsequent sequential comparison.
[0047] In addition, the above embodiment provides a similarity calculation method of cosine similarity, and the embodiment of the present application also provides a distance comparison method. For example: in the embodiment of the present application, after obtaining the recognition feature information, it can be compared with the preset image feature information, which can be specifically calculated by the formula: d(f(x), f(x i ))<ε, and this formula can be a related formula for calculating similarity.
[0048] The Euclidean distance between the recognition feature information and the preset image feature information can be calculated to obtain a comparison result. If there is at least one comparison result in which the Euclidean distance is less than a preset distance threshold, the basic feature information corresponding to the comparison result with the smallest Euclidean distance can be determined as the matching target feature information.
[0049] In order to better implement the embodiment of the present application, in one embodiment of the present application, the image feature information is updated based on the following steps:
[0050] When acquiring an image to be recorded, feature extraction is performed on the image to be recorded to obtain feature information of the image to be recorded corresponding to the image to be recorded; the feature information of the image to be recorded and the image to be recorded are stored to obtain image feature information; when it is determined to delete the feature information of the target image, the feature information of the target image in the current image feature information is deleted and the target image corresponding to the target image feature information is deleted to obtain image feature information.
[0051] The above-mentioned embodiment provides an implementation scheme for comparing identification feature information with pre-stored image feature information. The embodiment of the present application provides an implementation scheme for how to pre-store image feature information. Specifically, when the image to be recorded is obtained through an image sensor pre-installed on the terminal device. For example: when the user controls the terminal device to start collecting the image to be recorded on the terminal device, the terminal device can issue a voice prompt or text prompt to let the object or human body to be recorded align with the image sensor, at which time the image sensor starts working to collect the image to be recorded.
[0052] After the image to be recorded is collected, feature extraction can be performed on the image to be recorded to obtain image feature information of the image to be recorded. The feature extraction method can be performed according to any feature extraction method, and the specific embodiment of the present application is not limited.
[0053] Afterwards, when the acquisition of the image to be recorded and the feature extraction of the image to be recorded are completed, the image to be recorded and the corresponding feature information of the image to be recorded can be stored in association, for example, by setting the same identification number, etc. If the terminal device has previously stored historical images and image feature information corresponding to the historical images, the current image and the image feature information of the current image can be stored in the same storage location as the historical image and the feature information corresponding to the historical image in the storage device, so that the image and the image feature information can be quickly called when they are needed later.
[0054] In addition, in actual application scenarios, if it is necessary to delete the stored images and the corresponding image feature information, the user can perform the operation on the terminal device. When the user's deletion operation instruction is obtained, the corresponding user's deletion instruction is used to delete the target image specified by the user and the target image feature information corresponding to the target image. The deletion content not specified by the user is retained, thereby achieving the purpose of updating the image feature information.
[0055] However, with the development of artificial intelligence, neural network models are used in more and more scenarios. The effectiveness of neural network models usually depends on the model training process. For example, the more training samples the model receives during training, the better the model's recognition effect. Therefore, the better the model training scheme, the more it will affect the model's training effect.
[0056] In the prior art, in order to improve the performance of the model, it is usually chosen to increase the number of samples so that the model can improve its performance after receiving a large number of samples. However, it is usually impractical to increase the number of samples indefinitely. The specific reason is that the quality of samples obtained from open source means varies. When the quality of the obtained samples is poor, the effect of the training model will not be qualitatively improved. In addition, the cost of obtaining non-open source training samples with better quality cannot be controlled. Therefore, how to effectively train the model to improve the performance of the model has become a problem that needs to be solved in this field.
[0057] Therefore, in order to better implement the embodiments of the present application, in an embodiment of the present application, multiple training data are obtained; based on each training data, the target model is pre-trained to obtain target training data in each training data that does not meet the preset training requirements; and based on the target training data, the target model is then trained.
[0058] The training data in the embodiments of the present application can be any type of training data. For example: training text, training image, etc., specifically, the embodiments of the present application are not limited. At the same time, in the embodiments of the present application, in addition to obtaining training data from the corresponding storage medium, the training data can also be transmitted to the target model in real time to obtain the training data, specifically, the embodiments of the present application are also not limited. Among them, in the embodiments of the present application, the training data can carry labels.
[0059] In an embodiment of the present application, the method of pre-training the target model according to the training data can adopt supervised training or unsupervised training, and the specific need is to refer to whether the training data carries a label. For example: assuming that the training data carries a label, a supervised training method can be used. After the training data is input into the model, a predicted label can be obtained. At this time, it is sufficient to detect whether the predicted label is the real label carried. If it is not a real label, it can be determined that the training data currently input into the model does not meet the preset training requirements. At this time, the training data whose predicted label is not a real label can be determined as the target training data that does not meet the preset training requirements. Alternatively, when the training data used for model training does not carry a label, an unsupervised training method can be used. For example: each training data can be clustered, and after the clustering is completed, it can be detected whether the images in the cluster clusters are images of the same category, including: assuming that the training data is a face image, after the face image is clustered, it can be detected whether the face images in each cluster cluster are of the same person through other trained face recognition models. If there is a cluster in which a minority of face images and the majority of face images are not of the same person, it can be determined that the minority of face images are incorrectly clustered face images, and the incorrectly clustered face images can be regarded as target training data that does not meet the preset training requirements.
[0060] In the above description, different training methods such as supervised and unsupervised are provided to determine the target training data that does not meet the preset training requirements. It should be noted that the preset training requirements can be adjusted according to the actual training situation to better screen out the target training data that does not meet the preset training requirements. The specific embodiments of the present application are not limited.
[0061] In addition, it should be noted that in the embodiments of the present application, the target model can be any type of model, such as an image recognition model, a text recognition model, or any other model, and the specific embodiments of the present application are not limited thereto.
[0062] After the target training data that does not meet the preset training requirements is screened out according to the above description, it can be determined that the target training data is a relatively difficult data to identify for the target model. Therefore, if the target model can be successfully trained by the target training data, the performance of the target model can be greatly improved. Therefore, the target training data can be re-input into the target model for retraining. Among them, the retraining method can be any training method. For example: if the pre-training method is supervised training, the retraining method can still be supervised training. Specifically, the target training data and other training data except the target training data can be re-input into the target model after the parameter adjustment is completed, or the target training data can be re-input into the target model after the parameter adjustment is completed separately. If the target training data and other training data except the target training data are re-input into the target model, it is determined whether all the training data meet the preset training requirements. If not, the parameters of the target model are continuously adjusted until all the training data can meet the training requirements. Alternatively, when the training data re-input into the model is only the target training data, the above steps can still be repeated to adjust the model parameters until the model can obtain the training results corresponding to the target training data and meeting the training requirements.
[0063] The image processing method provided by the present application can first input each training data used for model training into the model for training, thereby obtaining corresponding pre-training results. After obtaining the pre-training results, it can be determined according to the pre-training results which target training data the model has not been effectively trained for. Therefore, on this basis, after the target training data is repeatedly input into the model for retraining, the model can be effectively retrained with the target training data, avoiding the waste of training samples caused by the target training data during the model training process, thereby improving the performance of the model.
[0064] In order to better implement the embodiment of the present application, in one embodiment of the present application, multiple training data are obtained, including:
[0065] Obtain target sample data; determine positive sample data and negative sample data of the target sample data; determine the target sample data, the positive sample data and the negative sample data as a first training data; regard the first training data as a training data until each training data is obtained.
[0066] The above embodiment provides a scheme for training a target model through training data. In order to further improve the training effect of the model, in the embodiment of the present application, the model can be trained in the form of a data group. Specifically, in the embodiment of the present application, the target sample data can be data in any form, including: image data, text data, etc., which is not limited in the specific present application. Assume that the target sample data is any face image A in the database. At this time, the positive sample data and negative sample data of the target sample data can be determined by a trained face recognition model. For example: a face image B with the same face as face image A can be determined from the database through the trained face recognition model, and the face image B is the positive sample data of the target sample data. At the same time, a face image C with a different face from the face image A can also be determined from the database or the network through the trained face recognition model, and the face image C is the negative sample data of the face image A. At this time, face image A, face image B and face image C together constitute a first training data, which can also be regarded as a training data group. With the training data constructed in this way, when training a model, the target model can learn the commonalities between the target sample data and the positive sample data. At the same time, the target model can also learn the differences between the target sample data and the negative sample data. Therefore, compared with directly training a model based on a single data, this training sample in the form of a data group can better help the model improve its performance.
[0067] In order to better implement the embodiment of the present application, in one embodiment of the present application, multiple training data are obtained, including:
[0068] Taking each first training data as the target first training data respectively, the target sample data corresponding to the target first training data and the negative sample data corresponding to the target first training data are fused to obtain the fused difficult negative sample data; determining the target sample data corresponding to the target first training data, the positive sample data corresponding to the target first training data and the difficult negative sample data corresponding to the target first training data as one second training data; treating the second training data as one training data until all training data are obtained.
[0069] The above embodiment provides a scheme for obtaining training data in the form of training data in the form of data groups. In order to further improve the performance of the model on this basis, the embodiment of the present application further optimizes the training data in the form of data groups. Specifically, it is necessary to fuse the negative sample data in the first training data with the target sample data, and use the fused data after fusion as the difficult negative sample data. The reason for doing this is that the negative sample data that is not similar to the target sample data is fused into the target sample data, which can make the difficult negative sample data have a certain similarity with the target sample data, increase the training difficulty of the model, and make it difficult for the model to learn the non-similarity features between the difficult negative sample and the target sample data. Therefore, if the model can learn the non-similarity between the difficult negative sample and the target sample data on this basis, the effect of the model will be better.
[0070] The method of fusing the target sample data with the negative sample data can be referred to formula (1). The specific formula (1) is as follows:
[0071] M(x, y)=(1-α)A(x, y)+αN(x, y)......(1)
[0072] In formula (1), A(x, y) is the target sample data, N(x, y) is the negative sample data, α is the set transparency parameter, and if the target sample data and the negative sample data are image data, (x, y) is the pixel coordinate. If the transparency parameter α is too small, the target sample data in the difficult negative sample data has too little information and cannot constitute a difficult negative sample; if α is too large, the target sample data in the difficult negative sample data has too much information, which can easily cause overfitting and cause the positive sample data to be recognized as different faces. Therefore, the transparency parameter α can be set between 0.4 and 0.6.
[0073] Therefore, assuming that there are target sample data A, positive sample data B and negative sample data C, the target sample data A and negative sample data C are fused using the above formula (1) to obtain fused difficult negative sample data D. Afterwards, the target sample data A, positive sample data B and difficult negative sample data D are taken as a second training data, and the second training data is regarded as one training data, so as to obtain multiple training data.
[0074] In order to better implement the embodiment of the present application, in one embodiment of the present application, multiple training data are obtained, including:
[0075] Acquire multiple sample data with labels, one sample data corresponds to one label; determine the number of clusters that is less than the number of label types; cluster each sample data according to the number of clusters to obtain each cluster; determine each training data according to the sample data in each cluster.
[0076] The above embodiment provides a scheme for obtaining training data in the form of data groups. The embodiment of the present application also provides a scheme for obtaining training data in the form of data groups by other means. Specifically, in the embodiment of the present application, if the sample data has its own identity identification label, and the same identity label represents that the sample data belongs to the same type of sample data, the sample data can be clustered to determine the positive sample data and the negative sample data of the sample data.
[0077] For example: suppose there are 100 face images, and the labels of these 100 face images are different in extreme cases, so these 100 face images are from different faces. At this time, it means that in theory, when 100 sample data are clustered without considering the labels, 100 clusters will be obtained. The reason is that if the performance of the model training is good enough, it can be determined that each face is not the same person, so 100 clusters will be formed, and each face image is a cluster. However, if these 100 face images are input into an untrained target model, the clustering results obtained by the model will most likely not form 100 clusters. In other words, there are face images that the model cannot recognize. Therefore, face images that the model cannot recognize can be used as the main training objects of the model. At the same time, the purpose of determining the number of clusters that is less than the number of label types is to ensure that at least one cluster contains sample data with different labels by determining the upper limit of the number of clusters. For example: Assume that there are N images of M people in total. In an extreme case, the number of clusters is set to K=M. Ideally, there are no negative sample pairs in each category. The smaller K is, the more negative sample pairs there are in a category, but the proportion of difficult negative sample pairs is small. You can set K=M / 2 to cluster the feature vectors, and divide the images into different clusters or categories according to the results of the algorithm. Based on this, after the clustering is completed, the label recognition program can be used to identify the label of each sample data in each cluster. If there are multiple types of labels in a cluster, it can be determined that the model is more likely to identify the sample data in the cluster as the same type, and the sample data in the cluster is more confusing to the model. Therefore, the data in the cluster can be used as training data. Therefore, the process can be repeated many times to obtain sufficient training data.
[0078] In order to better implement the embodiment of the present application, in one embodiment of the present application, each training data is determined according to the sample data in each cluster, including:
[0079] Take each cluster as the target cluster respectively, determine each sample data in the cluster as the target sample data; determine any sample data in the target cluster that has the same label as the target sample data as the positive sample data; determine any sample data in the target cluster that has a different label from the target sample data as the difficult negative sample data; determine the target sample data, positive sample data and difficult negative sample data as one training data, until all training data are obtained.
[0080] The above embodiment provides a method of using sample data with different labels in a cluster as training data. Similarly, in order to improve the training effect of the model, the sample data in each cluster can be used to construct training data in the form of a data group. In actual situations, the sample data with labels obtained are usually tens of thousands. Therefore, when clustering, there are more sample data in a cluster. At this time, a sample data can be randomly determined as the target sample data, and positive sample data and negative sample data can be found in the same cluster based on the label of the target sample data. For example: assuming that the label of the target sample data is A, sample data with the same label A is found in the same cluster as positive sample data, and a sample data with a label of non-A is randomly determined as difficult negative sample data. At this time, the target sample data labeled A, the positive sample data labeled A, and the negative sample data labeled non-A can be used as training data in the form of a data group. If there are pictures of different people in a cluster n i , i=1: M is a non-negative integer, n i = 0 means that there is no image of the i-th person in this cluster. Assume that n i =1, there is no image of the second person in this cluster, and it is not necessary to construct training data based on the current cluster. At the same time, based on the solution of this embodiment, the training data finally constructed may include repeated training data, and the repeated data can be removed.
[0081] In order to better implement the embodiment of the present application, in one embodiment of the present application, the target model is pre-trained according to each training data to obtain the target training data that does not meet the preset training requirements in each training data, including:
[0082] Each training data is input into the target model to obtain the pre-training results corresponding to each training data; according to each pre-training result, the training data with prediction errors is determined as the target training data.
[0083] In the embodiment of the present application, after each training data is obtained, the target model can be pre-trained. For example: if the training data is not in the form of a data group, after a training data is input into the model, if the training result obtained by the predicted label does not match the real label, it can be determined that the currently input training data does not meet the preset training requirements, and then the training data can be used as the target training data. If the training data is the training data in the form of a data group provided in the above embodiment, after a training data is input into the model, after the model learns the features of the positive sample data and the negative sample data, if the predicted label of the given positive sample data does not match the real label of the corresponding target sample data; or, if the predicted label of the given negative sample data matches the real label of the corresponding target sample data, it proves that the target model still cannot effectively identify the predicted result of the input data, and at this time, the training data can be regarded as the target training data that does not meet the preset training requirements.
[0084] It should be noted that if the input training data is training data in the form of a data set, the loss function of the target model can refer to formula (2). The specific formula (2) is as follows:
[0085] L(A,P,N)=max(||f(A)-f(P)|| 2 -||f(A)-f(N)|| 2 +α, 0)……(2)
[0086] Where A is the target sample data, P is the positive sample data, and N is the negative sample data or difficult negative sample data. The transparent parameter α can be determined by the formula (3): d(A, P) + α ≤ d(A, N) ... (3), d is the similarity distance. Specifically, during the training process, if pictures are randomly selected to form a data set (A, P, N), then the following condition is easily satisfied: d(A, P) + α ≤ d(A, N), so that the network cannot learn anything from it. Therefore, it is necessary to select a data set that is more difficult to train, that is, d(A, P) ≈ d(A, N), so that d(A, P) becomes smaller and d(A, N) becomes larger. Only by selecting a difficult data set can gradient descent play a role. Finally, the parameters learned through training make the encoding distance for the pictures of the same person very small; for the pictures of different people, the encoding distance is very large.
[0087] In order to better implement the embodiment of the present application, in one embodiment of the present application, retraining the target model according to the target training data includes:
[0088] The training weights of the target training data are adjusted; the target training data with the adjusted training weights are input into the target model to retrain the target model until the target model outputs the correct training results corresponding to the target training data.
[0089] The above embodiment provides a solution for retraining the model according to the target training data. In the embodiment of the present application, after the target training data is determined, the training weight of the target training data can be increased for all the training data, and the model can be retrained based on the adjusted weight. Specifically, after the weight of the target training data is adjusted, if the model receives training data with adjusted weights and training data without adjusted weights at this time, the model can focus on the target training data with adjusted weights; or, the training data with adjusted weights can be trained multiple times so that the model can effectively learn the features of the target training data, so as to help the model make effective and correct prediction results when it receives the target training data later.
[0090] Assuming that the number of training data in one epoch training is L, one epoch training can be understood as one training cycle or one training. Then the loss of each epoch sample is as shown in the following formula (4):
[0091]
[0092] Among them, A is the target sample data, P is the positive sample data, N is the negative sample data or difficult negative sample data; i is the ordinal number of the sample data in each epoch, and w is the weight of the corresponding sample data. First, the weight can be initialized. The specific initialization formula is shown in formula (5):
[0093]
[0094] As shown in formula (5), after the weights are initialized, the weight of each sample data is the same.
[0095] Assume that after one epoch of training, among the L group of samples, L1 group is correctly identified and L2 group is incorrectly identified. It is necessary to reduce the weight of the group that is correctly identified and increase the weight of the group that is incorrectly identified. And as the number of groups with incorrect identification decreases, the magnitude of the weight change should also decrease. Therefore, the following weight update formula (6) is designed to adjust the weight of the corresponding sample data:
[0096]
[0097] Among them, y i =1 means the i-th training sample is correctly identified, y i =0 indicates that the i-th training sample is recognized incorrectly. ω is the normalization coefficient, calculated as:
[0098] In order to better implement the embodiment of the present application, in one embodiment of the present application, the target model is pre-trained according to each training data to obtain the target training data that does not meet the preset training requirements in each training data, including:
[0099] Each training data is input into the target model, and the pre-training loss corresponding to each training data is determined; according to the pre-training loss corresponding to each training data, the training data that does not reach the preset training loss is determined as the target training data.
[0100] In the above embodiment, a method for determining target training data that does not meet the preset training requirements is provided by determining the prediction label output by the target model during pre-training. The embodiment of the present application also provides a method for determining target training data that does not meet the preset training requirements by determining the loss value of each training data.
[0101] Specifically, after inputting specific training data each time, the loss value corresponding to each training data can be obtained through the target model. If the obtained loss value is greater than a preset loss threshold, it can be determined that the currently input training data is the training data that is less than the preset training loss and is the target training data.
[0102] In order to better implement the embodiment of the present application, in one embodiment of the present application, retraining the target model according to the target training data includes:
[0103] Determine a retraining parameter according to the pre-training loss corresponding to the target training data and the pre-training loss corresponding to other training data excluding the target training data; if the retraining parameter is less than a preset value, repeatedly input each training data including the target training data into the target model to retrain the target model; if the retraining parameter is greater than or equal to the preset value, obtain additional training data; input the additional training data and each repetition including the target training data into the target model to retrain the target model.
[0104] In the above embodiment, a solution is provided for determining target training data that does not meet the preset training requirement based on the loss value. The embodiment of the present application also provides a better solution for determining target training data that does not meet the preset training requirement based on the loss value.
[0105] Specifically, an epoch can include multiple batches, which can be understood as the input of training data for training to the target model in batches during a training cycle, and a batch of training data can be regarded as a batch. In the previous batch construction, a batch can be regarded as a batch of fixed size generated according to the batch size given by the model, and the data in a batch is input into the model at one time. In the embodiment of the present application, assuming that the batch size is B, the loss J of B training data can be calculated. i , i=1:B. The sample loss value is divided into the following three intervals: 0, (0, α], (α, +∞), and the loss of B1 training data is 0, the loss of B2 training data is in the interval (0, α], and the loss of B3 training data is greater than α. When the loss is α, it means that d(A, P) ≈ d(A, N), so α is taken as the critical point. At this time, the retraining parameters can be determined by formula (7):
[0106]
[0107] Among them, B1, B2, and B3 are the pre-training losses corresponding to the target training data and the pre-training losses corresponding to other training data of the target training data; T is the retraining parameter.
[0108] At the same time, β1=1, β3<β2<1, these three loss value parameters represent the "correct degree" of the corresponding sample being identified. If the loss of a training data is 0, it means that this group of samples is completely correctly identified, so the correct degree β1=1; if the loss is between (0,α], it means that the distance between the target sample data and the positive sample is smaller than the distance between the target sample data and the negative sample, but the distance difference does not reach the boundary we set, so the correct degree β2=0.5; if the loss is greater than α, it means that the distance between the target sample data and the positive sample is greater than the distance between the target sample data and the negative sample, which means that the degree of correct identification is very small, β3=0.1. The value range of T is [0.1,1]. The larger T is, the higher the degree of correct identification of the training data in this batch is, which also indirectly reflects that there are fewer difficult negative samples in this batch, especially in the first few training epochs. Therefore, a certain number of difficult negative sample pairs can be added to this batch.
[0109] When T<0.7, it is considered that the batch has enough difficult negative samples and no more need to be added; when T>0.7, it is considered that it needs to be added. As the epoch increases, as the model training effect gets better and better, T in each batch will gradually increase, so the number of difficult negative samples added to the batch should decrease as the epoch increases. Based on the above analysis, we define the number of difficult negative samples to be added B' as follows, where E is the total number of training rounds.
[0110] It can be seen that in the embodiment of the present application, the preset value can be 0.7. Of course, the preset value can also be set according to the actual situation, and the specific embodiment of the present application is not limited. It should also be noted that if additional difficult negative sample data needs to be added, then the additional difficult negative sample data can be continued to be obtained according to the method in any of the above embodiments, and the specific details will not be repeated here. Whenever an additional difficult negative sample data is added, the retraining parameter can be continued to be calculated to determine whether the recalculated training parameter is still greater than 0.7 or less than 0.7, so as to determine whether it is necessary to continue to add difficult negative sample data, until the retraining parameter meets the conditions, no additional negative sample data will be added.
[0111] In addition, in an embodiment of the present application, the target model may include an input, a multi-layer CNN, a feature vector obtained by a fully connected layer, and a distance calculation. Two different inputs x1 and x2 are used, through two similar sub-networks with the same architecture, parameters, and weights. The two sub-networks are mirror images of each other, like conjoined twins. The input image is passed through CNN to obtain two vectors f(x1) and f(x2), which can be regarded as an encoding of the original image. The distance between the two feature vectors obtained is calculated. If the distance d(x1, x2) = ||f(x1)-f(x2)|| 2 If the value is less than or equal to the set threshold, the two face images belong to the same person.
[0112] In order to better implement the image processing method in the embodiment of the present application, in addition to the image processing method, an image processing device is also provided in the embodiment of the present application, such as Figure 3 As shown, the device 300 includes:
[0113] An acquisition module 301 is used to acquire image information to be processed;
[0114] An extraction module 302 is used to perform feature extraction processing on the image information to be processed based on the target image recognition model to obtain recognition feature information;
[0115] The determination module 303 is used to determine image recognition result information based on the image feature information and the recognition feature information.
[0116] The image processing device provided by the present application first acquires the image information to be processed through the acquisition module 301, then extracts the identification feature information of the image information to be processed through the extraction module 302, and finally compares the identification feature information with the preset image feature information through the determination module 303, so that the identity information of the image information to be processed can be determined based on the identity information of the preset image feature information that matches the identification feature information, thereby greatly improving the efficiency of confirming the identity information.
[0117] In some embodiments of the present application, the determination module 303 is specifically used to:
[0118] Determine whether there is basic feature information matching the identification feature information in the image feature information, and obtain a first determination result;
[0119] When the first judgment result is yes, determining the image recognition result information as the first recognition result;
[0120] When the first judgment result is no, the image recognition result information is determined to be a second recognition result.
[0121] In some embodiments of the present application, the determination module 303 is further configured to:
[0122] Determining a basic feature information as feature information to be processed from the image feature information in sequence;
[0123] Calculate the feature distance between the feature information to be processed and the identification feature information to obtain the identification feature distance value;
[0124] Determine whether the recognition feature distance value is less than a distance threshold, and obtain a second determination result;
[0125] When the second judgment result is yes, determine that the first judgment result is yes, and end the judgment process corresponding to the first judgment result;
[0126] When the second judgment result is no, it is judged whether all the basic feature information in the image feature information are subjected to feature distance calculation with the identification feature information to obtain a third judgment result;
[0127] When the third judgment result is yes, determine that the first judgment result is no, and end the judgment process corresponding to the first judgment result;
[0128] When the third judgment result is no, the triggering execution is to sequentially determine a basic feature information from the image feature information as the feature information to be processed.
[0129] In some embodiments of the present application, the device further includes a data updating module, which is specifically used to:
[0130] When the image to be recorded is obtained, feature extraction is performed on the image to be recorded to obtain feature information of the image to be recorded corresponding to the image to be recorded;
[0131] storing feature information of the image to be recorded and the image to be recorded, and obtaining the image feature information;
[0132] When it is determined to delete the target image feature information, the target image feature information in the current image feature information is deleted and the target image corresponding to the target image feature information is deleted to obtain the image feature information.
[0133] In some embodiments of the present application, the image processing device further includes a training data acquisition module, which is specifically used to:
[0134] Obtain target sample data;
[0135] Determine positive sample data and negative sample data of target sample data;
[0136] Determine the target sample data, the positive sample data, and the negative sample data as a first training data;
[0137] The first training data is regarded as one training data until each training data is acquired.
[0138] In some embodiments of the present application, the training data acquisition module is further used to:
[0139] Taking each first training data as the target first training data, respectively, performing data fusion on the target sample data corresponding to the target first training data and the negative sample data corresponding to the target first training data to obtain fused difficult negative sample data;
[0140] Determine the target sample data corresponding to the target first training data, the positive sample data corresponding to the target first training data, and the difficult negative sample data corresponding to the target first training data as a second training data;
[0141] The second training data is regarded as one training data until each training data is acquired.
[0142] In some embodiments of the present application, the training data acquisition module is further used to:
[0143] Obtain multiple sample data with labels, one sample data corresponds to one label;
[0144] Determine the number of clusters that is less than the number of label types;
[0145] Cluster each sample data according to the number of clusters to obtain each cluster;
[0146] The training data are determined according to the sample data in each cluster.
[0147] In some embodiments of the present application, the training data acquisition module is further used to:
[0148] Taking each cluster as the target cluster respectively, and determining each sample data in the cluster as the target sample data;
[0149] Determine any sample data in the target cluster that has the same label as the target sample data as positive sample data;
[0150] Determine any sample data in the target cluster that has a different label from the target sample data as difficult negative sample data;
[0151] The target sample data, the positive sample data, and the difficult negative sample data are determined as one training data until each training data is obtained.
[0152] In some embodiments of the present application, the image processing device further includes a pre-training module, which is specifically used to:
[0153] Input each training data into the target model to obtain the pre-training results corresponding to each training data;
[0154] According to each pre-training result, the training data with prediction errors is determined as the target training data.
[0155] In some embodiments of the present application, the image processing device further includes a retraining module, which is specifically used to:
[0156] Adjust the training weights of the target training data;
[0157] The target training data with adjusted training weights is input into the target model to retrain the target model until the target model outputs the correct training results corresponding to the target training data.
[0158] In some embodiments of the present application, the pre-training module is further used to:
[0159] Input each training data into the target model and determine the pre-training loss corresponding to each training data;
[0160] According to the pre-training loss corresponding to each training data, the training data that does not reach the preset training loss is determined as the target training data.
[0161] In some embodiments of the present application, the retraining module is further used to:
[0162] Determine a retraining parameter according to the pre-training loss corresponding to the target training data and the pre-training loss corresponding to other training data excluding the target training data;
[0163] If the retraining parameter is less than a preset value, each training data including the target training data is repeatedly input into the target model to retrain the target model;
[0164] If the retraining parameter is greater than or equal to a preset value, additional training data is obtained;
[0165] The additional training data and each repetition including the target training data are input into the target model to retrain the target model.
[0166] The present application also provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the steps of any image processing method in the present application. The terminal device integrates any image processing method provided in the present application. Figure 4 As shown, it shows a schematic diagram of the structure of the terminal device involved in the embodiment of the present application, specifically:
[0167] The terminal device may include one or more processing core processors 401, one or more computer-readable storage media memories 402, a power supply 403, an input unit 404 and other components. Those skilled in the art will appreciate that Figure 4 The terminal device structure shown in the figure does not constitute a limitation on the terminal device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them:
[0168] The processor 401 is the control center of the terminal device, and uses various interfaces and lines to connect various parts of the entire terminal device. By running or executing software programs and / or modules stored in the memory 402, and calling data stored in the memory 402, the processor 401 executes various functions of the terminal device and processes data, thereby monitoring the terminal device as a whole. Optionally, the processor 401 may include one or more processing cores; the processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, the user interface and the application program, etc., and the modem processor mainly processes wireless communication. It is understandable that the above-mentioned modem processor may not be integrated into the processor 401.
[0169] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0170] The terminal device also includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to manage charging, discharging, power consumption management and other functions through the power management system. The power supply 403 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.
[0171] The terminal device may further include an input unit 404, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0172] Although not shown, the terminal device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 401 in the terminal device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402, thereby realizing various functions, such as:
[0173] Obtaining image information to be processed;
[0174] Perform feature extraction processing on the image information to be processed based on the target image recognition model to obtain recognition feature information;
[0175] Based on the image feature information and the recognition feature information, image recognition result information is determined.
[0176] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0177] To this end, the embodiment of the present application provides a computer-readable storage medium, which may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any image processing method provided in the embodiment of the present application. For example, the computer program can be loaded by a processor to execute the following steps:
[0178] Obtaining image information to be processed;
[0179] Perform feature extraction processing on the image information to be processed based on the target image recognition model to obtain recognition feature information;
[0180] Based on the image feature information and the recognition feature information, image recognition result information is determined.
[0181] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the detailed description of other embodiments above, and will not be repeated here.
[0182] In specific implementation, the above units or structures can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units or structures can refer to the previous method embodiments, which will not be repeated here.
[0183] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0184] The above is a detailed introduction to an image processing method and device provided in an embodiment of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An image processing method, characterized in that: The method comprises: Obtaining image information to be processed; Performing feature extraction processing on the image information to be processed based on the target image recognition model to obtain recognition feature information; Image recognition result information is determined based on the image feature information and the recognition feature information.
2. The image processing method according to claim 1, characterized in that: The image feature information includes a number of basic feature information; The determining of image recognition result information based on the image feature information and the recognition feature information includes: Determine whether the basic feature information matching the identification feature information exists in the image feature information, and obtain a first determination result; When the first judgment result is yes, determining the image recognition result information as a first recognition result; When the first judgment result is no, the image recognition result information is determined to be a second recognition result.
3. The image processing method according to claim 2, characterized in that: The determining whether the basic feature information matching the identification feature information exists in the image feature information to obtain a first determination result includes: Determining one of the basic feature information as feature information to be processed from the image feature information in sequence; Calculating the feature distance between the feature information to be processed and the identification feature information to obtain an identification feature distance value; Determine whether the identification feature distance value is less than a distance threshold, and obtain a second determination result; When the second judgment result is yes, determining that the first judgment result is yes, and ending the judgment process corresponding to the first judgment result; When the second judgment result is no, judging whether all the basic feature information in the image feature information are subjected to feature distance calculation with the identification feature information, and obtaining a third judgment result; When the third judgment result is yes, determining that the first judgment result is no, and ending the judgment process corresponding to the first judgment result; When the third judgment result is no, the step of sequentially determining one of the basic feature information from the image feature information as the feature information to be processed is triggered.
4. The image processing method according to claim 1, characterized in that: The image feature information is updated based on the following steps: When acquiring an image to be recorded, performing feature extraction on the image to be recorded to obtain feature information of the image to be recorded corresponding to the image to be recorded; storing the image feature information to be recorded and the image to be recorded to obtain the image feature information; When it is determined to delete the target image feature information, the target image feature information in the current image feature information is deleted and the target image corresponding to the target image feature information is deleted to obtain the image feature information.
5. The image processing method according to claim 1, characterized in that: The target image recognition model is obtained based on the following steps: Get multiple training data; Performing initial training on the target image recognition model according to each of the training data to obtain initial training result information; The target image recognition model is trained according to the initial training result information.
6. The image processing method according to claim 5, characterized in that: The obtaining of a plurality of training data comprises: Obtain target sample data; Determine positive sample data and negative sample data of the target sample data; Determine the target sample data, the positive sample data, and the negative sample data as a first training data; Taking each of the first training data as the target first training data, respectively, performing data fusion on the target sample data corresponding to the target first training data and the negative sample data corresponding to the target first training data to obtain fused difficult negative sample data; Determine the target sample data corresponding to the target first training data, the positive sample data corresponding to the target first training data, and the difficult negative sample data corresponding to the target first training data as a second training data; Treat the second training data as one training data until all the training data are obtained; Or, obtain multiple sample data with labels, one sample data corresponds to one label; Determine the number of clusters that is less than the number of label types; Clustering each of the sample data according to the number of clusters to obtain each cluster; Taking each cluster as a target cluster, determining each sample data in the cluster as a target sample data; Determine any sample data in the target cluster that has the same label as the target sample data as positive sample data; Determine any sample data in the target cluster that has a different label from the target sample data as difficult negative sample data; The target sample data, the positive sample data, and the difficult negative sample data are determined as one training data until each of the training data is obtained.
7. The image processing method according to claim 5, characterized in that: The step of training the target image recognition model according to the initial training result information includes: According to the initial training result information, determining the training data with prediction errors as the target training data; adjusting the training weight of the target training data; Inputting the target training data for adjusting the training weights into the target model to train the target model until the target model outputs a correct training result corresponding to the target training data; Or, determining the pre-training loss corresponding to each of the training data according to the initial training result information; According to the pre-training loss corresponding to each of the training data, determine the training data that does not reach the preset training loss as the target training data; Determining training parameters according to the pre-training loss corresponding to the target training data and the pre-training loss corresponding to other training data excluding the target training data; If the retraining parameter is less than a preset value, repeatedly inputting each of the training data including the target training data into the target model to train the target model; If the training parameter is greater than or equal to the preset value, obtaining additional training data; The additional training data and each of the repetitions including the target training data are input into the target model to train the target model.
8. An image processing device, characterized in that: The device comprises: An acquisition module, used for acquiring image information to be processed; An extraction module, used for performing feature extraction processing on the image information to be processed based on a target image recognition model to obtain recognition feature information; The determination module is used to determine the image recognition result information based on the image feature information and the recognition feature information.
9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the steps of the image processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in the image processing method according to any one of claims 1 to 7.