Regression training method and device based on scale-inconsistent noise sparse data set

By employing a dual-path joint training method and loss function optimization, the problem of sparse regression label datasets with inconsistent scales and noise is solved, achieving stable convergence and efficient scoring of deep learning networks, which is suitable for large-scale automated aquaculture.

CN120974453APending Publication Date: 2025-11-18HEFEI LASSETER ROBOT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511110336.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies cannot effectively handle sparse regression label datasets with inconsistent scales and noise, causing deep learning networks to fail to converge or perform poorly during training.

Method used

A dual-path joint training method is adopted. The training dataset is divided by a pre-set network model to construct an image regression branch network and a label calibration branch network. The first and second loss functions are used for iterative training to generate calibration regression values ​​and predicted regression values, thereby reducing noise and unifying the scale.

Benefits of technology

It effectively handles the problem of inconsistent label scales, has robust learning capabilities for noisy labels, achieves stable convergence, and is suitable for automated animal condition scoring in large-scale farming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974453A_ABST
    Figure CN120974453A_ABST
Patent Text Reader

Abstract

The invention provides a regression training method and device based on a scale-inconsistent noisy sparse data set, relates to the field of animal feeding, and solves the problem that the prior art cannot work well on a scale-inconsistent noisy sparse regression label data set. And the technical problem that the deep learning network cannot converge or the final performance is poor in the training process is solved. The method comprises the following steps: acquiring a training data set D, and dividing the training data set D to obtain sub-data sets; performing dual-path joint training on the sub-data set through a preset network model to obtain a calibration regression value and a prediction regression value of each animal image sample; and performing iterative training on a currently trained preset network model based on the first loss function and the second loss function to obtain a trained image regression network model. The method is used in the animal body condition scoring process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of deep learning, in particular to a regression training method and device based on scale-inconsistent noisy sparse data sets. BACKGROUND

[0002] In large-scale animal production, in order to achieve the best feeding results with the minimum amount of feed, breeding experts evaluate the body condition of animals according to their own experience. In order to realize large-scale automation, a camera can be used to take pictures of animals, and a deep learning network can be applied to calculate the body condition of animals according to the animal images taken by the camera. The deep learning network needs to be fully trained first, and the training needs to collect a large number of expert scores of animal body condition as labels. However, the expert scores of animal body condition according to their own experience are subjective, and different experts may have different scores for the body condition of the same animal. Since the labels provided by the experts are subjective, they are usually noisy. When the labels in the data set come from multiple experts, the scale may not be consistent and sparse. Therefore, how to effectively design a regression training method for scale-inconsistent noisy sparse regression label data sets has become a technical problem to be solved. SUMMARY

[0003] The present application provides a regression training method and device based on scale-inconsistent noisy sparse data sets, which solves the technical problem that the prior art cannot work well on scale-inconsistent noisy sparse regression label data sets, resulting in that the deep learning network cannot converge during the training process or the final performance is poor.

[0004] To achieve the above purpose, the present application adopts the following technical solutions: In a first aspect, a regression training method based on scale-inconsistent noisy sparse data sets is provided, comprising: obtaining a training data set D and dividing the training data set D to obtain a sub-data set; wherein the training data set D includes animal image samples and corresponding body condition score regression value labels; performing double-path joint training on the sub-data set through a preset network model to obtain calibrated regression values and predicted regression values of each animal image sample; wherein the calibrated regression values are regression values after the preset network model is uniformly scaled and the noise is reduced, and the predicted regression values are animal body condition score regression values obtained by training and predicting the input animal image sample through the preset network model; iteratively training the preset network model based on a first loss function and a second loss function to obtain a trained image regression network model.

[0005] Based on the above technical scheme, in the regression training method based on the scale inconsistent noisy sparse data set provided in the application, the training data set D is obtained, the training data set D is divided to obtain a sub data set, the sub data set and the training data set D are trained through a preset network model, the calibration regression value and the prediction regression value of each image sample are obtained, the first loss value and the second loss value are obtained based on the loss calculation of the first loss function and the second loss function, the preset network model is fine-tuned and corrected based on the first loss value and the second loss value, and the trained image regression network model is obtained. It can effectively solve the problem of inconsistent label scale, has robust noise label learning ability, and shows stable convergence on different expert annotated data sets, and is suitable for automatic animal body condition scoring in large-scale breeding.

[0006] In combination with the above first aspect, in a possible implementation manner, the preset network model is constructed in the following manner: An image regression branch network S and a label calibration branch network T are constructed; wherein the image regression branch network S is used for inputting an image sample and outputting a prediction regression value corresponding to the image sample; and the label calibration branch network T is used for inputting an image sample and outputting a vector Q with a length equal to N, which is used by the image regression branch network S in a training process. A candidate regression value C is constructed based on a regression value label in the training data set D ; wherein the candidate regression value C is a sequence with a length of N, and satisfies the following formula: ; wherein, and are a minimum label value and a maximum label value in the training data set D respectively, is an nth element in the candidate regression value C , and the candidate regression value C is obtained.

[0007] In combination with the above first aspect, in a possible implementation manner, the training manner of the label calibration branch network T is as follows: A data set is randomly selected from the sub data set, including an image sample X and a label Y; The image sample X is input into the label calibration branch network T, and a vector Q is output, which is normalized to a probability distribution through a calculation formula Q=softmax(Q); wherein the vector Q is used as an occurrence probability of each element of the candidate regression value C.

[0008] In combination with the above first aspect, in a possible implementation manner, the training manner of the image regression branch network S is as follows: input the image sample X into the image regression branch network S, and extract basic features through convolution operation ; based on dynamic attention mask through the calculation formula obtain the attention feature ; wherein, is an element-wise multiplication operator in the channel dimension; the attention feature passes through a fully connected layer, and outputs a vector P, which is normalized to a probability distribution through the calculation formula P=softmax(P); wherein, the vector P is used as the candidate regression value the appearance probability of each element, the length of the vector P is equal to N, and N is a hyperparameter, which can be manually set according to experience; through the calculation formula calculate the expectation of the candidate regression value as the predicted regression value of the final image sample.

[0009] In combination with the above first aspect, in a possible implementation manner, the first loss function is obtained in the following manner: obtain a first loss function loss1 through the calculation formula loss1=KL(P||G) based on the vector P and the soft label G, and minimize the loss loss1 through an optimization algorithm; wherein, KL(P||G) is the KL divergence calculation of the vector P and the soft label G; the acquisition manner of the soft label G, comprising: calculate the calibration regression value U of the label calibration branch network T through the calculation formula ; calculate the absolute value of the distance difference between the calibration regression value U and the cluster center A through the calculation formula E=abs(U-A); construct a Gaussian function with the mean equal to U and the standard deviation equal to E, and obtain the soft label G through the calculation formula , and perform normalization operation ; wherein, is a Gaussian function, and the soft label G is obtained , which is used as the supervision label of the image regression branch network S and constructs a dynamic attention mask; the acquisition manner of the cluster center A, comprising: construct the cluster centers of k batches , based on the vector Q, take the index i corresponding to the maximum value in each row, and update the cluster center A through the calculation formula ; wherein, is the nth center value in the kth batch, is a weight coefficient, This represents the b-th element in U.

[0010] In conjunction with the first aspect mentioned above, in one possible implementation, the second loss function is obtained as follows: Based on the vector Q, a predetermined number of elements are deleted from each column, and the average entropy E1 of the vector Q is calculated using the formula E1=entropy(Q); where entropy() is used to calculate the average entropy. Based on the vector Q, select the maximum value in each row to form a vector QM= ; Through calculation formula Normalize the vector QM; The average entropy E2 of vector QM is calculated using the formula E2 = entropy(QM); The calibration regression value U and the label Y are divided into two equal parts according to their length to obtain vectors U1 and U2, and vectors Y1 and Y2; Through calculation formula The second loss function, loss2, is obtained, and the loss2 is minimized through an optimization algorithm. Here, eps is a constant term greater than zero, used to prevent anomalies such as division by zero or numerical overflow during calculation; r is a constant parameter used to balance numerical relationships in the loss calculation. These are the weight hyperparameters of the average entropy E1, average entropy E2, and loss L, respectively.

[0011] In conjunction with the first aspect above, in one possible implementation, the method of minimizing the loss through the optimization algorithm is as follows: Minimize the first loss function loss1 and the second loss function loss2 using the stochastic gradient descent algorithm, satisfying the following formula: ; in, and These are the learnable parameters in the image regression branch network S and the label calibration branch network T, respectively; t is the number of iterations, and η is the learning rate. and These are the gradients of loss1 and loss2 with respect to parameters param(S) and param(T), respectively.

[0012] In conjunction with the first aspect above, in one possible implementation, the dynamic attention mask is constructed in the following ways: Through calculation formula Obtain the mapping of learnable linear layers ;in, The number of channels for the intermediate features of the image regression branch network S; Through calculation formula Generating dynamic attention masks .

[0013] With the first aspect, in a possible implementation, the division manner of the sub-data set comprises: dividing the training data set D into K sub-data sets according to annotation sources , , ; wherein the label of the Kth sub-data set is annotated by the Kth expert.

[0014] The second aspect provides an electronic device, comprising a communication unit and a processing unit; the communication unit is configured to obtain a training data set D and divide the training data set D to obtain a sub-data set; wherein the training data set D comprises animal image samples and corresponding body condition score regression value labels; the processing unit is configured to perform double-path joint training on the sub-data set through a preset network model to obtain a calibrated regression value and a predicted regression value of each animal image sample; wherein the calibrated regression value is a regression value after the preset network model is uniformly scaled and noise is reduced, and the predicted regression value is an animal body condition score regression value obtained by training and prediction of the input animal image sample through the preset network model; and the preset network model is iteratively trained based on a first loss function and a second loss function to obtain a trained image regression network model.

[0015] The third aspect provides an electronic device, comprising a processor and a storage medium; the storage medium comprises instructions, and the processor is configured to execute the instructions to implement the method described in the first aspect and any possible implementation of the first aspect. The electronic device can be an electronic device or a chip in the electronic device.

[0016] The fourth aspect provides a regression training system based on scale-inconsistent noisy sparse data sets, comprising a data acquisition unit and an electronic device; wherein the data acquisition unit is configured to obtain a training data set D and divide the training data set D to obtain a sub-data set; wherein the training data set D comprises animal image samples and corresponding body condition score regression value labels; and the electronic device is configured to perform double-path joint training on the sub-data set through a preset network model to obtain a calibrated regression value and a predicted regression value of each animal image sample; wherein the calibrated regression value is a regression value after the preset network model is uniformly scaled and noise is reduced, and the predicted regression value is an animal body condition score regression value obtained by training and prediction of the input animal image sample through the preset network model; and the preset network model is iteratively trained based on a first loss function and a second loss function to obtain a trained image regression network model.

[0017] In a fifth aspect, the present application provides a computer readable storage medium, having stored therein instructions which, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect and any possible implementation manner of the first aspect.

[0018] In a sixth aspect, the present application provides a computer program product having instructions, which, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect and any possible implementation manner of the first aspect.

[0019] It should be understood that the description of technical features, technical solutions, advantages or the like in the present application does not imply that all features and advantages can be realized in any single embodiment. On the contrary, it can be understood that the description of a feature or advantage means that the specific technical feature, technical solution or advantage is included in at least one embodiment. Therefore, the description of technical features, technical solutions or advantages in the specification does not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions and advantages described in the embodiments can be combined in any appropriate manner. Those skilled in the art will understand that the embodiments can be implemented without one or more specific technical features, technical solutions or advantages of a specific embodiment. In other embodiments, additional technical features and advantages can be identified in specific embodiments without embodying all embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 A system architecture diagram of a regression training system provided by an embodiment of the present application; Figure 2 A flowchart of a regression training method based on scale inconsistent noisy sparse data sets provided by an embodiment of the present application; Figure 3 A structural schematic diagram of an electronic device provided by an embodiment of the present application; Figure 4 A hardware structural schematic diagram of an electronic device provided by an embodiment of the present application; DETAILED DESCRIPTION

[0021] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.

[0022] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0023] The regression training method based on scale-inconsistent noisy sparse datasets provided in this application can be applied to, for example... Figure 1 In the regression training system 100 shown, such as Figure 1 As shown, the communication system includes a data acquisition unit 10 and an electronic device 20.

[0024] The data acquisition device 10 is used to acquire a training dataset D and divide the dataset D into sub-datasets. The dataset D consists of noisy data samples with inconsistent label scales, including animal image samples and corresponding regression value labels for body condition scores. The electronic device 20 is used to perform dual-path joint training on the sub-datasets using a preset network model to obtain calibrated regression values ​​and predicted regression values ​​for each animal image sample. The calibrated regression value is the regression value after the preset network model has unified the scale and reduced noise, and the predicted regression value is the animal body condition score regression value predicted by the preset network model based on the input animal image sample. The preset network model is iteratively trained based on a first loss function and a second loss function to obtain a trained image regression network model.

[0025] To solve the technical problem that the prior art cannot work well on a scale-inconsistent and noisy sparse regression label data set, resulting in that a deep learning network cannot converge during training or the final performance is poor, an embodiment of the present application provides a regression training method based on a scale-inconsistent and noisy sparse data set, which comprises the following steps: obtaining a training data set D and dividing the training data set D to obtain a sub-data set; wherein the training data set D comprises animal image samples and corresponding body condition score regression value labels; performing double-path joint training on the sub-data set by using a preset network model to obtain calibration regression values and prediction regression values of each animal image sample; wherein the calibration regression value is a regression value after the preset network model unifies scales and reduces noise, and the prediction regression value is an animal body condition score regression value obtained by training and prediction of the input animal image sample by using the preset network model; iteratively training the preset network model based on a first loss function and a second loss function to obtain a trained image regression network model. Based on this, the label scale inconsistency problem can be effectively handled, and stable convergence is exhibited on different expert-annotated data sets, which is suitable for automatic animal body condition scoring in large-scale breeding.

[0026] As shown in Figure 2 , the regression training method based on a scale-inconsistent and noisy sparse data set provided by an embodiment of the present application comprises the following steps: S201, obtaining a training data set D and dividing the training data set D to obtain a sub-data set.

[0027] The training data set D comprises animal image samples and corresponding body condition score regression value labels.

[0028] In some implementations, the training data set D is divided into K sub-data sets according to the annotation sources , ,…, ; wherein the labels of the sub-data sets are annotated by the Kth expert.

[0029] It should be noted that different experts have different body condition scores for the same animal image sample, which is noisy and scale-inconsistent.

[0030] S202, performing double-path joint training on the sub-data set by using a preset network model to obtain calibration regression values and prediction regression values of each animal image sample.

[0031] The calibration regression value is a regression value after the preset network model unifies scales and reduces noise, and the prediction regression value is an animal body condition score regression value obtained by training and prediction of the input animal image sample by using the preset network model.

[0032] In some implementations, the preset network model is constructed in the following manner: an image regression branch network S and a label calibration branch network T are constructed; the image regression branch network S is configured to input an image sample and output a predicted regression value corresponding to the image sample; the label calibration branch network T is configured to input the image sample and output a vector Q with a length of N, which is used by the image regression branch network S in a training process; candidate regression values are constructed based on regression value labels in a training data set D ; the candidate regression values are a sequence with a length of N and satisfy the following formula: ; wherein, and are a minimum label value and a maximum label value in the training data set D, is an nth element in the candidate regression value , and the candidate regression value is obtained.

[0033] For example, the candidate regression value is {1.4, 1.8, 2.2, 2.6, 3.0, 3.4, 3.8, 4.2, 4.6, 5.0}.

[0034] S203, iteratively training the preset network model based on the first loss function and the second loss function to obtain a trained image regression network model.

[0035] For example, the first loss function is loss0=KL(V||G), which is minimized by an optimization algorithm to obtain a first loss function loss1. The second loss function loss is obtained by the calculation formula , and is minimized by an optimization algorithm to obtain a second loss function loss2.

[0036] Based on the above technical solutions, the regression training method based on scale inconsistent noisy sparse data sets provided in the application can effectively handle the label scale inconsistency problem, has robust noise label learning ability, and shows stable convergence on different expert annotated data sets, and is suitable for automatic animal body condition scoring in large-scale breeding.

[0037] In a possible implementation of the embodiments of the application, S202 can be implemented by S301 and S302 as follows, which are described in detail as follows: S301, the label calibration branch network T is used as a branch network to assist in training the image regression branch network S. In some implementations, the training manner of the label calibration branch network T is: randomly selecting one data set from the sub data sets , comprising: an image sample X and a label Y; inputting the image sample X into the label calibration branch network T to output a vector Q, and normalizing the vector Q into a probability distribution by a calculation formula Q = softmax(Q); wherein the vector Q is used as the appearance probability of each element of the candidate regression value C.

[0038] It should be noted that the single sampling is from the same sub data set, so as to avoid the scale inconsistency introduced by different labeling sources.

[0039] For example, the normalized probability distribution obtained by the calculation formula Q = softmax(Q) is Q = {0.09, 0.04, 0.19, 0.3, 0.19, 0.09, 0.03, 0.01, 0.05, 0.01}.

[0040] S302, constructing an image regression branch network S as a main network.

[0041] In some implementations, the training manner of the image regression branch network S is: inputting the image sample X into the image regression branch network S, extracting basic features by convolution operation , obtaining attention features based on the dynamic attention mask ; wherein, is an element-wise multiplication operator in the channel dimension; the attention features are input into a fully connected layer to output a vector P, and the vector P is normalized into a probability distribution by a calculation formula P = softmax(P); wherein the vector P is used as the appearance probability of each element of the candidate regression value ; and the expected value of the candidate regression value is calculated by a calculation formula as a prediction regression value of the final image sample.

[0042] For example, after dimension reduction by the fully connected layer, the normalized P = {0.07, 0.05, 0.21, 0.24, 0.19, 0.09, 0.05, 0.03, 0.01, 0.00} is obtained by the calculation formula P = softmax(P).

[0043] In a possible implementation of the embodiments of the present application, S203 can be implemented by S401 and S402 as follows, which are specifically described as follows: S401, calculating a first loss function.

[0044] In some implementations, the first loss function is obtained in the following manner: An initial loss function loss0 is obtained based on the predicted regression value V and the soft label G by a calculation formula loss0=KL(V||G), and is minimized by an optimization algorithm to obtain the first loss function loss1; wherein KL(V||G) is a KL divergence calculation of the predicted regression value V and the soft label G; The soft label G is obtained in the following manner: The calibration regression value U of the label calibration branch network T is calculated by a calculation formula The calibration regression value U of the label calibration branch network T is calculated by a calculation formula The absolute value of the distance difference between the calibration regression value U and the cluster center A is calculated by a calculation formula E=abs(U-A); A Gaussian function with a mean equal to U and a standard deviation equal to E is constructed, and the soft label G is obtained by a calculation formula and is normalized ; wherein, is a Gaussian function, and the soft label G is obtained , which is used as a supervision label of the image regression branch network S and constructs a dynamic attention mask; The cluster center A is obtained in the following manner: The cluster centers of k batches are constructed Based on the vector Q, the index i corresponding to the maximum value in each row is taken, and the cluster center A is updated by a calculation formula ; wherein, is the nth center value in the kth batch, is a weight coefficient, represents the bth element in U.

[0045] S402, a second loss function is calculated.

[0046] In some implementations, the second loss function is obtained in the following manner: Based on the vector Q, a preset number of elements in each column are deleted, the average entropy E1 of the vector Q is calculated by a calculation formula E1=entropy(Q); wherein entropy() is an average entropy calculation; Based on the vector Q, the maximum value in each row is selected to form a vector QM= ; The vector QM is normalized by a calculation formula ; The average entropy E2 of the vector QM is calculated by a calculation formula E2=entropy(QM); The calibration regression value U and the label Y are equally divided into two parts to obtain vectors U1 and U2, and vectors Y1 and Y2. The second loss function loss2 is obtained by calculating the formula and minimizing the second loss function loss2 by an optimization algorithm; wherein, eps is a constant term greater than zero, r is a constant term parameter, respectively, are weight hyperparameters of the average entropy E1, average entropy E2 and loss L.

[0047] S403, constructing a dynamic attention mask.

[0048] In some implementations, the dynamic attention mask is constructed in the following manner: by calculating the formula to obtain the mapping of the learnable linear layer ; wherein, is the channel number of the intermediate feature of the image regression branch network S; by calculating the formula to generate a dynamic attention mask .

[0049] For example, by calculating the formula to obtain the mapping of the learnable linear layer {0.2, 0.15, 0.3, 0.4, 0.25, 0.1, 0.05, 0.03,..., 0.01}, by calculating the formula to obtain {0.55, 0.54, 0.57, 0.6, 0.56, 0.52, 0.51, 0.50,..., 0.50}.

[0050] Based on the above technical solution, a regression training model is constructed, which takes the image regression branch network S as the main network and the label calibration branch network T as the auxiliary network. It has good effect on processing sparse regression label data sets with inconsistent scales and noise. In the demand for intelligent upgrading of breeding, this method focuses on the problem of accurate evaluation of animal body condition, provides reliable technical support for automatic body condition monitoring of large-scale breeding, and helps the intelligent and fine landing of breeding decision-making.

[0051] The above describes the solutions of the embodiments of the present application mainly from the perspective of device implementation. It can be understood that each device, for example, an electronic device, includes at least one of a corresponding hardware structure and a software module for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed herein, the present application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed by hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0052] The embodiments of the present application can divide the functional units of the electronic device according to the above method examples. For example, each functional unit can be divided according to each function, or two or more functions can be integrated in one processing unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. It should be noted that the division of the units in the embodiments of the present application is illustrative, and is only a logical functional division. When actually implemented, there can be another division manner.

[0053] In the case of using an integrated unit, Figure 3 A possible structural schematic diagram of the electronic device (denoted as electronic device 50) involved in the above embodiments is shown, which includes a processing unit 501 and a communication unit 502, and can also include a storage unit 503. Figure 3 The structural schematic diagram shown can be used to illustrate the structure of the electronic device involved in the above embodiments.

[0054] When Figure 3 When the structural schematic diagram shown is used to illustrate the structure of the electronic device involved in the above embodiments, the processing unit 501 is used to control and manage the actions of the electronic device, the communication unit 502 is used for communication between the electronic device and other devices, and the storage unit 503 is used to store the program code and data of the electronic device.

[0055] For example, the communication unit 502 is configured to obtain a training data set D and divide the training data set D to obtain a sub-data set; wherein the training data set D includes animal image samples and corresponding body condition score regression value labels. The processing unit 501 is configured to perform double-path joint training on the sub-data set through a preset network model to obtain a calibration regression value and a prediction regression value of each animal image sample; wherein the calibration regression value is a regression value after the preset network model is uniformly scaled and noise is reduced, and the prediction regression value is an animal body condition score regression value obtained by training and predicting the input animal image sample through the preset network model; and based on a first loss function and a second loss function, the preset network model is iteratively trained to obtain a trained image regression network model.

[0056] The processing unit 501 can be a processor or a controller, and the communication unit 502 can be a communication interface, a transceiver, a transceiver, a transceiver circuit, a transceiver device, etc. The communication interface is collectively referred to, and can include one or more interfaces. The storage unit 503 can be a memory. When the electronic device 50 is a chip, the processing unit 501 can be a processor or a controller, and the communication unit 502 can be an input interface and / or an output interface, a pin or a circuit, etc. The storage unit 503 can be a storage unit (for example, a register, a cache, etc.) within the chip, or a storage unit (for example, a read-only memory (ROM), a random access memory (RAM), etc.) located outside the chip.

[0057] The communication unit can also be referred to as a transceiver unit. The antenna and control circuit with transceiver function in the electronic device 50 can be regarded as the communication unit 502 of the electronic device 50, and the processor with processing function can be regarded as the processing unit 501 of the electronic device 50. Optionally, the device for realizing the receiving function in the communication unit 502 can be regarded as a communication unit, and the communication unit is used to execute the receiving steps in the embodiments of the application. The communication unit can be a receiver, a receiver, a receiving circuit, etc. The device for realizing the sending function in the communication unit 502 can be regarded as a sending unit, and the sending unit is used to execute the sending steps in the embodiments of the application. The sending unit can be a transmitter, a sender, a sending circuit, etc.

[0058] Figure 3 The integrated units in the above embodiments can be stored in a computer readable storage medium if they are realized in the form of software function modules and sold or used as independent products. Based on such understanding, the technical solutions of the embodiments of the application can be embodied in the form of software product, and the computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in the embodiments of the application. The storage medium for storing computer software product includes: U disk, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk, and various media that can store program codes.

[0059] Figure 3 The units in the above embodiments can also be referred to as modules, for example, the processing unit can be referred to as a processing module.

[0060] The embodiments of the application also provide a hardware structure diagram of an electronic device (denoted as electronic device 60), which is shown in Figure 4The electronic device 60 comprises a processor 601, and optionally further comprises a memory 602 connected with the processor 601.

[0061] In the first possible implementation, referring to Figure 4 The electronic device 60 further comprises a transceiver 603. The processor 601, the memory 602 and the transceiver 603 are connected through a bus. The transceiver 603 is used for communicating with other devices or communication networks. Optionally, the transceiver 603 can comprise a transmitter and a receiver. The device for realizing the receiving function in the transceiver 603 can be regarded as a receiver, and the receiver is used for executing the receiving steps in the embodiments of the present application. The device for realizing the sending function in the transceiver 603 can be regarded as a transmitter, and the transmitter is used for executing the sending steps in the embodiments of the present application.

[0062] Based on the first possible implementation, Figure 4 The structure diagram shown can be used for illustrating the structure of the electronic device involved in the above embodiments.

[0063] Among them, Figure 4 The system chip in the electronic device can also be illustrated. In this case, the actions performed by the electronic device can be realized by the system chip, and the specific actions performed can be referred to in the above, and will not be described here.

[0064] In the implementation process, each step in the method provided by the embodiment can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the software form. The steps of the method disclosed in the embodiments of the present application can be directly embodied as the execution completed by the hardware processor, or executed by the combination of the hardware and the software modules in the processor.

[0065] The processor in the present application can include, but is not limited to, at least one of the following: a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a microcontroller unit (MCU), or an artificial intelligence processor, and various computing devices running software, each of which can include one or more cores for executing software instructions to perform operations or processing. The processor can be a separate semiconductor chip, or can be integrated with other circuits as a semiconductor chip, for example, it can be integrated with other circuits (such as coding and decoding circuits, hardware acceleration circuits, or various bus and interface circuits) to form a SoC (system on chip), or it can be integrated as a built-in processor in an ASIC. The ASIC integrated with the processor can be packaged separately or packaged together with other circuits. In addition to including cores for executing software instructions to perform operations or processing, the processor can further include necessary hardware accelerators, such as field programmable gate arrays (FPGAs), PLDs (programmable logic devices), or logic circuits that implement special logic operations.

[0066] The memory in the embodiments of the present application can include at least one of the following types: read-only memory (ROM) or other types of static storage devices that can store static information and instructions, random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, and electrically erasable programmable read-only memory (EEPROM). In some scenarios, the memory can also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but not limited thereto.

[0067] The embodiments of the present application also provide a computer readable storage medium including instructions, which, when executed on a computer, cause the computer to perform any of the above methods.

[0068] The embodiments of the present application also provide a computer program product including instructions, which, when executed on a computer, cause the computer to perform any of the above methods.

[0069] The embodiment of the present application further provides a chip, comprising a processor and an interface circuit, the interface circuit being coupled with the processor, the processor being used to run computer programs or instructions to realize the method described above, and the interface circuit being used to communicate with other modules outside the chip.

[0070] In the above embodiments, the implementation can be achieved by software, hardware, firmware or any combination thereof, entirely or partially. When implemented by software, the implementation can be achieved in the form of a computer program product, entirely or partially. The computer program product comprises one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the flow or function described in the embodiments of the present application is generated, entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (for example, coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (for example, infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or data storage device including one or more servers, data centers, etc. integrated with the medium. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk (solid state disk, SSD)) and the like.

[0071] Although the present application is described herein in conjunction with various embodiments, other variations of the disclosed embodiments can be understood and implemented by those skilled in the art with reference to the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. Some measures described in mutually different dependent claims can be combined and produce a good result.

[0072] Although the present application has been described in connection with certain specific features and embodiments thereof, it is to be understood that it is provided as an example to the best of the applicant's knowledge and that various modifications and combinations of the described features and embodiments are possible and are within the spirit and scope of the application. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense, and all such modifications and variations are considered within the scope of the present application as defined by the following claims and their equivalents. Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the claims and their equivalents, the present application can be practiced otherwise than as specifically described.

Claims

1. A regression training method based on scale-inconsistent noisy sparse datasets, characterized in that, include: Obtain a training dataset D, and divide the training dataset D into sub-datasets; wherein, the training dataset D includes: animal image samples and corresponding regression value labels for body condition scores; The subset of data is jointly trained using a pre-defined network model via dual paths to obtain the calibration regression value and the prediction regression value for each animal image sample. The calibration regression value is the regression value after the pre-defined network model has been scaled and noise reduced, and the prediction regression value is the animal condition score regression value predicted by the input animal image sample after training with the pre-defined network model. The pre-set network model is iteratively trained based on the first loss function and the second loss function to obtain the trained image regression network model.

2. The method according to claim 1, characterized in that, The construction method of the preset network model includes: Construct an image regression branch network S and a label calibration branch network T; wherein, the image regression branch network S is used as input image samples and outputs the predicted regression value corresponding to the image samples; the label calibration branch network T is used as input image samples and outputs a vector Q of length N, which is used by the image regression branch network S during training. Candidate regression values ​​are constructed based on the regression value labels in the training dataset D. Wherein, the candidate regression value Given a sequence of length N, satisfy the following formula: ; in, and These are the minimum and maximum label values ​​in the training dataset D, respectively. Candidate regression values The nth element is used to obtain the candidate regression value. .

3. The method according to claim 2, characterized in that, The training method for the label calibration branch network T is as follows: Randomly select one dataset from the subset. This includes: image sample X and label Y; The image sample X is input into the label calibration branch network T, and the output vector Q is obtained. The vector Q is normalized into a probability distribution by calculating Q=softmax(Q). The vector Q is used as the probability of occurrence of each element of the candidate regression value C.

4. The method according to claim 3, characterized in that, The training method for the image regression branch network S is as follows: The image sample X is input into the image regression branch network S, and basic features are extracted through convolution operations. ; Based on dynamic attention mask Through calculation formula Obtain attention features ;in, This is the element-wise multiplication operator in the channel dimension; The attention features After passing through a fully connected layer, the output is a vector P, which is normalized to a probability distribution using the formula P = softmax(P); wherein, the vector P is used as the candidate regression value. The probability of each element appearing; Through calculation formula The expected value of the candidate regression value is calculated as the predicted regression value for the final image sample.

5. The method according to claim 4, characterized in that, The first loss function is obtained as follows: Based on the vector P and the soft label G, the first loss function loss1 is obtained by calculating loss1=KL(P||G), and the loss loss1 is minimized by an optimization algorithm; where KL(P||G) is the KL divergence between the vector P and the soft label G. The methods for obtaining the soft tag G include: Through calculation formula The calibration regression value U of the label calibration branch network T is calculated; The absolute value of the distance difference between the calibration regression value U and the cluster center A is calculated using the formula E=abs(UA). Construct a Gaussian function with mean U and standard deviation E, and calculate the formula. Obtain the soft label G and perform normalization. ;in, Given a Gaussian function, we obtain the soft label G= , used as supervision labels for the image regression branch network S and for constructing dynamic attention masks; The method for obtaining the cluster center A includes: Construct k batches of cluster centers Based on the vector Q, the index i corresponding to the maximum value in each row is taken, and the formula is used to calculate... Update the cluster center A; wherein, The nth center value in the kth batch. These are the weighting coefficients. This represents the b-th element in U.

6. The method according to claim 3, characterized in that, The second loss function is obtained as follows: Based on the vector Q, a predetermined number of elements are deleted from each column, and the average entropy E1 of the vector Q is calculated using the formula E1=entropy(Q); where entropy() is used to calculate the average entropy. Based on the vector Q, select the maximum value in each row to form a vector QM= ; Through calculation formula Normalize the vector QM; The average entropy E2 of vector QM is calculated using the formula E2 = entropy(QM); The calibration regression value U and the label Y are divided into two equal parts according to their length to obtain vectors U1 and U2, and vectors Y1 and Y2; Through calculation formula The second loss function, loss2, is obtained and minimized using an optimization algorithm; where eps is a constant term greater than zero, and r is a constant term parameter. These are the weight hyperparameters of the average entropy E1, average entropy E2, and loss L, respectively.

7. The method according to claim 5 or 6, characterized in that, The method for minimizing the loss using the optimization algorithm is as follows: Minimize the first loss function loss1 and the second loss function loss2 using the stochastic gradient descent algorithm, satisfying the following formula: ; in, and These are the learnable parameters in the image regression branch network S and the label calibration branch network T, respectively; t is the number of iterations, and η is the learning rate. and These are the gradients of loss1 and loss2 with respect to parameters param(S) and param(T), respectively.

8. The method according to claim 5, characterized in that, The dynamic attention mask is constructed in the following ways: Through calculation formula Obtain the mapping of learnable linear layers ;in, The number of channels for the intermediate features of the image regression branch network S; Through calculation formula Generate dynamic attention mask .

9. The method according to claim 1, characterized in that, The methods for dividing the dataset into sub-datasets include: The training dataset D is divided into K subsets according to the source of the annotation. , ,…, ; wherein, the subset dataset The label is assigned to the Kth expert.

10. An electronic device, characterized in that, The electronic device includes: a communication unit and a processing unit; The communication unit is used to acquire a training dataset D and divide the training dataset D into sub-datasets; wherein, the training dataset D includes: animal image samples and regression value labels of corresponding body condition scores; The processing unit is used to perform dual-path joint training on the subset of data using a preset network model to obtain the calibration regression value and the predicted regression value of each animal image sample; wherein, the calibration regression value is the regression value after the preset network model is scaled and noise is reduced, and the predicted regression value is the animal condition score regression value predicted by the input animal image sample after training with the preset network model; and the preset network model is iteratively trained based on the first loss function and the second loss function to obtain the trained image regression network model.