Image processing methods, apparatuses, electronic devices and software products

By generating time-dependent parameters of the loss sequence and dynamically adjusting the training strategy, the problem that manual parameter tuning cannot respond to the dynamics of model training in real time is solved, thereby improving the model's generalization ability and training efficiency.

CN122090149APending Publication Date: 2026-05-26CHINA UNITED NETWORK COMM GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNITED NETWORK COMM GRP CO LTD
Filing Date
2026-02-10
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, manual parameter tuning cannot respond to the dynamics of model training in real time, resulting in poor model generalization ability.

Method used

By acquiring images, determining a second image based on a preset training strategy, training the model, generating a first loss sequence, calculating a first parameter of time dependence, and dynamically adjusting the training strategy to adapt to the model state.

Benefits of technology

It achieves real-time response and improved generalization ability in model training, avoids poor training conditions, and reduces computational complexity and data resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090149A_ABST
    Figure CN122090149A_ABST
Patent Text Reader

Abstract

This application provides an image processing method, apparatus, electronic device, and program product. It relates to the field of image processing technology. The method includes: acquiring a first image; determining a second image from the first image based on a preset training strategy; training a first model for a first round based on the preset training strategy and the second image, and determining a first loss sequence, the first loss sequence indicating the stability of the model during training based on the preset training strategy; determining a first parameter corresponding to the first loss sequence, the first parameter indicating the time dependence of the first loss sequence; determining a target training strategy based on the first parameter; and training the first model for the first round based on the target training strategy and the first image. This can improve the generalization ability of the first model in image classification tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image processing method, apparatus, electronic device, and program product. Background Technology

[0002] In the field of deep learning, image classification is a core application scenario that has penetrated into key areas such as security monitoring, medical image diagnosis, and autonomous driving. Its model generalization ability directly determines the reliability and security of the model in practical applications.

[0003] Currently, the generalization ability of a model can be improved by manually adjusting hyperparameters to optimize training. However, in the above methods, manual parameter tuning cannot respond to the dynamics of model training in real time, which may result in poor generalization ability of the model. Summary of the Invention

[0004] This application provides an image processing method, apparatus, electronic device, and program product to improve the generalization ability of a model.

[0005] In a first aspect, embodiments of this application provide an image processing method, including:

[0006] Get the first image;

[0007] Based on a preset training strategy, the second image is determined from the first image;

[0008] Based on a preset training strategy and a second image, the first model is trained in the first round, and a first loss sequence is determined. The first loss sequence is used to indicate the stability of the model when the first model is trained based on the preset training strategy.

[0009] Determine the first parameter corresponding to the first loss sequence. The first parameter is used to indicate the time dependence of the first loss sequence.

[0010] Based on the first parameter, determine the target training strategy;

[0011] The first model is trained in the first round based on the target training strategy and the first image.

[0012] In one possible implementation, determining the first loss sequence includes:

[0013] During the first round of training of the first model, the loss values ​​for each round are obtained;

[0014] Based on the training time sequence, the loss values ​​of each round are sorted to obtain the first loss sequence.

[0015] In one possible implementation, determining the first parameter corresponding to the first loss sequence includes:

[0016] Divide the first loss sequence into a first number of first loss subsequences;

[0017] Based on each first loss subsequence, determine the first parameter corresponding to the first loss sequence.

[0018] In one possible implementation, the first parameter corresponding to each first loss subsequence is determined, including:

[0019] For each first loss subsequence, determine the second and third parameters corresponding to the first loss subsequence. The second parameter is used to indicate the stability of the first model based on the temporal dependence of the first loss subsequence, and the third parameter is used to indicate the stability of the first model based on the volatility of the first loss subsequence.

[0020] The first relationship is determined based on the second and third parameters corresponding to each first loss subsequence;

[0021] Based on the first relationship, determine the first parameter.

[0022] In one possible implementation, determining the target training strategy based on the first parameter includes:

[0023] Based on the first parameter, determine the time dependence of the first loss sequence;

[0024] Based on the degree of time dependence, the target training strategy is determined.

[0025] In one possible implementation, determining the time dependence of the first loss sequence based on the first parameter includes:

[0026] If the first parameter is less than or equal to the first threshold, the time dependence of the first loss sequence is determined as the first dependence.

[0027] If the first parameter is greater than the first threshold and less than the second threshold, the time dependence of the first loss sequence is determined as the second dependence.

[0028] If the first parameter is greater than or equal to the second threshold, the time dependence of the first loss sequence is determined as the third dependence.

[0029] In one possible implementation, the target training strategy is determined based on the degree of time dependence, including:

[0030] If the time dependence is the highest, increase the learning rate or noise in the preset training strategy to obtain the target training strategy.

[0031] If the time dependence is the second degree of dependence, the preset training strategy will be determined as the target training strategy.

[0032] If the time dependence is third-degree dependence, reduce the learning rate in the preset training strategy to obtain the target training strategy.

[0033] Secondly, embodiments of this application provide an image processing apparatus, comprising: an acquisition module, a first determining module, a first processing module, a second determining module, a third determining module, and a second processing module, wherein...

[0034] The acquisition module is used to acquire the first image;

[0035] The first determining module is used to determine the second image from the first image based on a preset training strategy;

[0036] The first processing module is used to train the first model in the first round based on a preset training strategy and a second image, and to determine a first loss sequence. The first loss sequence is used to indicate the stability of the model when training the first model based on the preset training strategy.

[0037] The second determining module is used to determine the first parameter corresponding to the first loss sequence. The first parameter is used to indicate the time dependence of the first loss sequence.

[0038] The third determining module is used to determine the target training strategy based on the first parameter;

[0039] The second processing module is used to train the first model in the first round based on the target training strategy and the first image.

[0040] In one possible implementation, the first processing module is specifically used for:

[0041] During the first round of training of the first model, the loss values ​​for each round are obtained;

[0042] Based on the training time sequence, the loss values ​​of each round are sorted to obtain the first loss sequence.

[0043] In one possible implementation, the second determining module is specifically used for:

[0044] Divide the first loss sequence into a first number of first loss subsequences;

[0045] Based on each first loss subsequence, determine the first parameter corresponding to the first loss sequence.

[0046] In one possible implementation, the second determining module is specifically used for:

[0047] For each first loss subsequence, determine the second and third parameters corresponding to the first loss subsequence. The second parameter is used to indicate the stability of the first model based on the temporal dependence of the first loss subsequence, and the third parameter is used to indicate the stability of the first model based on the volatility of the first loss subsequence.

[0048] The first relationship is determined based on the second and third parameters corresponding to each first loss subsequence;

[0049] Based on the first relationship, determine the first parameter.

[0050] In one possible implementation, the third determining module is specifically used for:

[0051] Based on the first parameter, determine the time dependence of the first loss sequence;

[0052] Based on the degree of time dependence, the target training strategy is determined.

[0053] In one possible implementation, the third determining module is specifically used for:

[0054] If the first parameter is less than or equal to the first threshold, the time dependence of the first loss sequence is determined as the first dependence.

[0055] If the first parameter is greater than the first threshold and less than the second threshold, the time dependence of the first loss sequence is determined as the second dependence.

[0056] If the first parameter is greater than or equal to the second threshold, the time dependence of the first loss sequence is determined as the third dependence.

[0057] In one possible implementation, the third determining module is specifically used for:

[0058] If the time dependence is the highest, increase the learning rate or noise in the preset training strategy to obtain the target training strategy.

[0059] If the time dependence is the second degree of dependence, the preset training strategy will be determined as the target training strategy.

[0060] If the time dependence is third-degree dependence, reduce the learning rate in the preset training strategy to obtain the target training strategy.

[0061] Thirdly, embodiments of this application provide an electronic device, including: at least one processor and a memory; the memory stores computer execution instructions; the at least one processor executes the computer execution instructions stored in the memory, causing the at least one processor to perform the image processing method as described in the first aspect and various possible designs of the first aspect.

[0062] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the image processing method described in the first aspect and various possible designs of the first aspect.

[0063] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the image processing method described in the first aspect and various possible designs of the first aspect.

[0064] In a sixth aspect, embodiments of this application provide a chip, the chip including at least one processor, the processor being configured to execute program instructions to implement the image processing method as described in the first aspect and various possible designs of the first aspect.

[0065] This application provides an image processing method, apparatus, electronic device, and program product. By calculating a first parameter characterizing the time dependence of a first loss sequence generated in the first round of training of a first model, and dynamically determining a target training strategy that adapts to the training state of the model based on the first parameter, this method replaces manual parameter tuning. It can respond to the dynamics of model training in real time and can determine the convergence of the first model during training based on the time dependence indicated by the first parameter, thus avoiding the first model from falling into an unfavorable training state and improving the generalization ability of the first model in image classification tasks. Attached Figure Description

[0066] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0067] Figure 1 A schematic flowchart of an image processing method provided in an embodiment of this application;

[0068] Figure 2 A flowchart illustrating a method for determining a first parameter provided in an embodiment of this application;

[0069] Figure 3 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this application;

[0070] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0071] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0072] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0073] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.

[0074] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.

[0075] It should be noted that the image processing method, apparatus, electronic device and program product provided in this application can be used in the field of image processing technology, or in any field other than the field of image processing technology. The application field of the image processing method, apparatus, electronic device and program product in this application is not limited.

[0076] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, they do not mean that the applicant has used or necessarily used the solution.

[0077] In the field of deep learning, image classification is a core application scenario that has penetrated into key areas such as security monitoring, medical image diagnosis, and autonomous driving. Its model generalization ability directly determines the reliability and security of the model in practical applications.

[0078] Currently, the generalization ability of a model can be improved by manually adjusting hyperparameters to optimize training. However, in the above methods, manual parameter tuning cannot respond to the dynamics of model training in real time, which may result in poor generalization ability of the model.

[0079] To address the aforementioned issues, this application provides an image processing method comprising: acquiring a first image; determining a second image from the first image based on a preset training strategy; training a first model for a first round based on the preset training strategy and the second image, and determining a first loss sequence, wherein the first loss sequence indicates the stability of the model during training based on the preset training strategy; determining a first parameter corresponding to the first loss sequence, wherein the first parameter indicates the time dependence of the first loss sequence; determining a target training strategy based on the first parameter; and training the first model for a first round based on the target training strategy and the first image.

[0080] In the above method, a first parameter representing the time dependence of the first loss sequence generated in the first round of training of the first model is calculated. Based on the first parameter, a target training strategy that adapts to the training state of the model is dynamically determined, replacing the manual parameter tuning method. This method can respond to the dynamics of model training in real time. Based on the time dependence indicated by the first parameter, the convergence of the first model during the training process can be determined, avoiding the first model from falling into an unfavorable training state, thereby improving the generalization ability of the first model in image classification tasks.

[0081] Furthermore, by determining the first loss sequence through the second image, the generalization state of the first model can be evaluated in real time without additional data, which can solve the problem of data resource consumption of the validation set in related technologies. The convergence of the first model can be determined through the first parameter, which can solve the problem of high computational complexity caused by calculating the eigenvalues ​​of the Hessian matrix or conducting parameter perturbation experiments to evaluate the convergence of the first model in related technologies.

[0082] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0083] Figure 1 This is a schematic flowchart illustrating an image processing method provided in an embodiment of this application. Please refer to [link / reference]. Figure 1 As shown, the method may include the following steps:

[0084] S101, Obtain the first image.

[0085] The execution subject of this application embodiment can be an electronic device or an image processing device disposed in an electronic device. The image processing device can be implemented by software or by a combination of software and hardware.

[0086] In one possible implementation, the first image refers to the initial set of original images acquired for training the first model. It is the basic data source for training the first model, and the type, quantity, and quality of the images it covers directly affect the model training effect.

[0087] S102. Based on the preset training strategy, determine the second image from the first image.

[0088] In one possible implementation, the preset training strategy refers to a set of training parameters and execution rules set based on experience or benchmarks of similar tasks in the initial stage of training the first model. This is used to train the first batch of the first model and provide an initial reference for subsequent strategy adjustments.

[0089] In one possible implementation, the preset training strategy may include, but is not limited to: the learning rate for training the first model, the noise for training the first model, and the size of the index set for training the first model.

[0090] In one possible implementation, the second image is a subset of images obtained by filtering, dividing, or preprocessing the first image according to a preset training strategy, and is used for the first round of training of the first model.

[0091] In one possible implementation, the second image can be determined from the first image based on the size of the index set in a preset training strategy.

[0092] For example, if the index set size of the first model being trained is 6 and the first round is 50, then during each round of training, 6 second images can be determined from the first image. During the first round of training of the first model, 300 second images can be determined from the first image. The explanation of the first round can be found in subsequent embodiments.

[0093] S103. Based on the preset training strategy and the second image, the first model is trained in the first round, and a first loss sequence is determined. The first loss sequence is used to indicate the stability of the model when the first model is trained based on the preset training strategy.

[0094] In one possible implementation, the first model refers to a deep learning model to be trained for a specific image classification task. Its network structure is designed according to the task requirements, and its parameters are continuously optimized through training to improve classification performance.

[0095] In one possible implementation, the first round can be preset, for example, to 50 times.

[0096] In one possible implementation, the first round of training refers to the first complete training process of the first model based on a preset training strategy and a second image.

[0097] In one possible implementation, the first loss sequence is a set of loss values ​​recorded in chronological order during the first round of training. It is used to reflect the stability of the first model in the first round of training. For example, the smaller the fluctuation of each loss value in the first loss sequence and the lower the overall trend, the more stable the training of the first model is.

[0098] In one possible implementation, the specific implementation method for training the first model in the first round based on a preset training strategy and a second image is as follows: Based on the preset training strategy, the parameters of the first model are initialized. For each round of the first training, the second image is loaded (the number of second images is equal to the size of the index set). The first model performs feature extraction and classification calculation on the second image. Then, using a loss function (e.g., cross-entropy loss function), the prediction result of the first model for the second image is compared with the true label of the second image, and the loss value for that round is calculated. Based on the loss value, the gradient of each parameter in the first model is calculated by the optimizer, and the weights of the first model are updated based on the learning rate in the preset training strategy.

[0099] In some implementations, determining the first loss sequence may include:

[0100] During the first round of training of the first model, the loss values ​​of each round are obtained; based on the training time order, the loss values ​​of each round are sorted to obtain the first loss sequence.

[0101] In one possible implementation, the training time sequence refers to the order in which the training rounds of the first model are advanced.

[0102] In this way, by recording the loss values ​​of each round in real time during the first round of training of the first model, and sorting the loss values ​​according to the order of rounds (training time order) to construct the first loss sequence, the temporal characteristics of the loss value changes during the training of the first model can be preserved. This indicates the evolution of the first model from initial adaptation to gradual convergence, providing basic data support for the determination of the first parameters. This allows the determination of the target training strategy to be more in line with the actual training dynamics of the first model, thereby avoiding the blindness of traditional manual parameter tuning and improving the training efficiency and generalization performance of the first model.

[0103] S104. Determine the first parameter corresponding to the first loss sequence. The first parameter is used to indicate the time dependence of the first loss sequence.

[0104] In one possible implementation, the first parameter is a quantitative indicator calculated based on the first loss sequence to indicate the degree of time dependence of the first loss sequence. The degree of time dependence can indicate the correlation characteristics of the loss value in the time dimension (such as the continuity and fluctuation correlation of the loss values ​​before and after). The larger the first parameter is, the stronger the degree of time dependence is indicated.

[0105] S105. Based on the first parameter, determine the target training strategy.

[0106] In one possible implementation, the target training strategy can be an optimized training strategy obtained by dynamically adjusting a preset training strategy, or it can be a preset training strategy. The target training strategy is used to adapt to the current generalization ability of the first model and prevent the first model from falling into a poor training state.

[0107] In some implementations, the specific methods for determining the target training strategy based on the first parameter are as follows:

[0108] Based on the first parameter, determine the time dependence of the first loss sequence; based on the time dependence, determine the target training strategy.

[0109] In some implementations, determining the time dependence of the first loss sequence based on the first parameter can include:

[0110] If the first parameter is less than or equal to the first threshold, the time dependence of the first loss sequence is determined as the first dependence; if the first parameter is greater than the first threshold and less than the second threshold, the time dependence of the first loss sequence is determined as the second dependence; if the first parameter is greater than or equal to the second threshold, the time dependence of the first loss sequence is determined as the third dependence.

[0111] In one possible implementation, the first threshold and the second threshold can be preset, with the second threshold being larger than the first value. For example, the first threshold is 0.5 and the second threshold is 0.7.

[0112] In one possible implementation, the degree of time dependence can indicate the correlation characteristics of each loss value in the first loss sequence, thereby indicating the continuity and stability of the change of loss values ​​during the training process of the first model. The degree of time dependence can include: first degree of dependence, second degree of dependence and third degree of dependence.

[0113] The first level of dependence is the lowest level of time dependence, indicating that the correlation between the loss values ​​in the first loss sequence is weak and the randomness is strong. The training state of the first model is unstable and prone to getting trapped in sharp minima. The second level of dependence is the medium level of time dependence, indicating that the correlation between the loss values ​​in the first loss sequence is moderate, the fluctuation is stable and shows a reasonable downward trend. The training state of the first model is in a stable convergence process. The third level of dependence is the highest level of time dependence, indicating that the correlation between the loss values ​​in the first loss sequence is strong and the persistence is good. The first model has converged to a flat region and the training state is stable.

[0114] In one possible implementation, the loss values ​​in the first loss sequence decrease steadily and the fluctuations before and after are closely related, indicating a high degree of time dependence; or the loss values ​​in the first loss sequence fluctuate randomly without obvious patterns, indicating a low degree of time dependence.

[0115] In this way, by comparing the first parameter with the preset first threshold and second threshold, three levels of time dependence are determined, and the correlation characteristics of the first loss sequence and the model training state are determined based on the level. This not only quantitatively represents the stability of the first model training process, but also provides a basis for judgment on the differentiated adjustment of subsequent training strategies, thereby matching the optimization needs of the current state of the first model and improving training efficiency and the generalization ability of the first model.

[0116] In some implementations, the target training strategy is determined based on the degree of time dependence, which may include:

[0117] If the time dependency is at the first level, increase the learning rate or noise in the preset training strategy to obtain the target training strategy; if the time dependency is at the second level, determine the preset training strategy as the target training strategy; if the time dependency is at the third level, decrease the learning rate in the preset training strategy to obtain the target training strategy.

[0118] In one possible implementation, the time dependency is the first dependency, indicating that the generalization ability of the first model needs to be improved. Therefore, the learning rate or noise in the preset training strategy can be increased to make the first model converge, and the loss sequence has a high degree of flatness to avoid getting trapped in sharp minima, thereby improving the generalization ability of the model.

[0119] It should be noted that if the time dependency is the highest level, the size of the index set in the preset training strategy can be reduced to obtain the target training strategy.

[0120] In one possible implementation, the time dependency level is the second dependency level, indicating that the generalization ability of the first model is moderate, so the training of the first model can continue based on the preset training strategy.

[0121] It should be noted that if the time dependence is second-degree dependence, the learning rate in the preset training strategy can be slightly increased, and / or the index set size in the preset training strategy can be slightly decreased to obtain the target training strategy.

[0122] In one possible implementation, the time dependency is the third dependency, indicating that the first model has good generalization ability. Therefore, the learning rate in the preset training strategy can be reduced to obtain the target training strategy.

[0123] In this way, by adopting differentiated training strategies according to different levels of time dependence, the training parameters can be accurately adapted to different states of the generalization ability of the first model. Furthermore, by adjusting the learning rate, noise, index set size, etc., the first model can converge to a flatter minimum point and avoid getting stuck in a sharp minimum. Thus, the generalization performance and training stability of the first model can be dynamically improved without consuming additional data resources.

[0124] S106. Based on the target training strategy and the first image, the first model is trained in the first round.

[0125] In one possible implementation, after reducing the learning rate in the preset training strategy to obtain the target training strategy, the following can also be done:

[0126] Obtain the third image; validate the first model based on the target training strategy and the third image, and calculate the error and feature value of the first model in classifying the third image. Based on the error and feature value, determine the generalization ability of the first model. If the generalization ability is strong, perform the first round of training on the first model based on the target training strategy and the first image; if the generalization ability is not strong, perform the first round of training on the first model based on the preset training strategy and the first image.

[0127] In one possible implementation, the third image is the original image set used for validation of the first model. The error is the deviation ratio between the predicted classification result of the first model and the true classification label of the third image when the first model classifies the third image. The eigenvalue is the largest eigenvalue of the Hessian matrix corresponding to the current parameters of the first model at the last iteration of validation. The eigenvalue is used to determine the sharpness of the minimum point of the loss function. The smaller the eigenvalue, the flatter the minimum point of the loss value and the stronger the generalization ability; the larger the eigenvalue, the sharper the extreme point of the loss value and the weaker the generalization ability.

[0128] In one possible implementation, the specific method for determining the generalization ability of the first model based on error and eigenvalues ​​is as follows:

[0129] If the error is less than or equal to the error threshold and the eigenvalue is less than or equal to the eigenvalue threshold, the generalization ability is strong; if the error is greater than the error threshold or the eigenvalue is greater than the eigenvalue threshold, the generalization ability is weak. The error threshold and the eigenvalue threshold can be preset, for example, the error threshold is 5% and the eigenvalue threshold is 10.

[0130] In one possible implementation, after determining the target training strategy, a fourth image can be obtained from the first image. After the first round of training, a second loss sequence can be determined, followed by a fourth parameter corresponding to the second loss sequence. Based on the fourth parameter, a new target training strategy is determined, and then the first model is trained for the first round again based on the new target training strategy. Here, the fourth image is used for the subsequent first round of training of the first model, the second loss sequence indicates the stability of the model when training the first model based on the target training strategy, and the fourth parameter indicates the time dependence of the second loss sequence. It should be noted that the specific implementation method can refer to the above embodiments, and will not be elaborated here.

[0131] In this embodiment, a first image is acquired; a second image is determined from the first image based on a preset training strategy; a first model is trained for a first round based on the preset training strategy and the second image, and a first loss sequence is determined, the first loss sequence indicating the stability of the model when training the first model based on the preset training strategy; a first parameter corresponding to the first loss sequence is determined, the first parameter indicating the time dependence of the first loss sequence; a target training strategy is determined based on the first parameter; and the first model is trained for a first round based on the target training strategy and the first image.

[0132] In the above method, a first parameter representing the time dependence of the first loss sequence generated in the first round of training of the first model is calculated. Based on the first parameter, a target training strategy that adapts to the training state of the model is dynamically determined, replacing the manual parameter tuning method. This method can respond to the dynamics of model training in real time. Based on the time dependence indicated by the first parameter, the convergence of the first model during the training process can be determined, avoiding the first model from falling into an unfavorable training state, thereby improving the generalization ability of the first model in image classification tasks.

[0133] Furthermore, by determining the first loss sequence through the second image, the generalization state of the first model can be evaluated in real time without additional data, which can solve the problem of data resource consumption of the validation set in related technologies. The convergence of the first model can be determined through the first parameter, which can solve the problem of high computational complexity caused by calculating the eigenvalues ​​of the Hessian matrix or conducting parameter perturbation experiments to evaluate the convergence of the first model in related technologies.

[0134] Based on the above embodiments, the following will be combined with... Figure 2 The method for determining the first parameter is explained. For example, Figure 2 A flowchart illustrating a method for determining a first parameter provided in an embodiment of this application is shown below. Figure 2 As shown, the method includes:

[0135] S201. Divide the first loss sequence into a first number of first loss subsequences.

[0136] In one possible implementation, the first loss subsequence is a series of subsequences obtained by splitting the first loss sequence into multiple subsequences of fixed length, which are used to refine the analysis of the state of the loss sequence at different training stages.

[0137] In one possible implementation, the first quantity can be preset, for example, to be 20. For example, if the first loss sequence includes the loss values ​​of 100 rounds of training and the first quantity is 20, then 5 (100 / 20) first loss subsequences can be obtained.

[0138] S202. Based on each first loss subsequence, determine the first parameter corresponding to the first loss sequence.

[0139] In some implementations, determining the first parameter corresponding to each first loss subsequence based on the first loss subsequences may include:

[0140] For each first loss subsequence, determine the second and third parameters corresponding to the first loss subsequence. The second parameter is used to indicate the stability of the first model based on the temporal dependence of the first loss subsequence, and the third parameter is used to indicate the stability of the first model based on the volatility of the first loss subsequence. Based on the second and third parameters corresponding to each first loss subsequence, determine the first relationship. Based on the first relationship, determine the first parameter.

[0141] In one possible implementation, the second parameter can be the range of each loss value in the first loss subsequence, that is, the difference between the maximum and minimum values ​​of each loss value in the first loss subsequence. The smaller the second parameter, the narrower the extreme fluctuation range of the loss value, the stronger the temporal dependence of the first model in this training phase (the loss change is continuous and stable), and the better the stability. The larger the second parameter, the more violent the extreme fluctuation of the loss value, the weaker the temporal dependence, and the worse the stability.

[0142] In one possible implementation, the third parameter can be the standard deviation of each loss value in the first loss subsequence, which is a quantitative indicator of the degree of dispersion of each loss value in the first loss subsequence from its average value. The smaller the third parameter, the more concentrated the fluctuation of the loss value around the average value is, the smaller the fluctuation of the first model in this training phase, and the stronger the stability. The larger the third parameter, the more obvious the discrete distribution of the loss value is, the greater the fluctuation of the first model, and the weaker the stability.

[0143] In one possible implementation, a first logarithmic value can be determined, then the ratio of the second parameter to the third parameter can be determined, then the logarithmic value of the ratio can be determined, then a first relationship graph can be determined to indicate the relationship between the logarithmic value of the ratio and the first logarithmic value, then the curve in the first relationship graph can be fitted to a straight line, and finally the slope of the straight line can be used as the second parameter.

[0144] In this embodiment, the first loss sequence is divided into fixed-length first loss subsequences according to a preset number of fights. The model stability of each subsequence is quantified from two dimensions: time dependence and volatility, using the range (second parameter) and standard deviation (third parameter). The first parameter is then determined based on the second and third parameters. This allows for a detailed analysis of the stability of the first loss sequence at different training stages. Furthermore, logarithmic transformation and linear fitting can eliminate data fluctuation interference, enabling the first parameter to more accurately and comprehensively indicate the training stability of the first model. This provides a reliable and quantitative basis for subsequent determination of time dependence and adjustment of training strategies, thereby enhancing the generalization ability of the first model.

[0145] It should be noted that the image processing method provided in this application embodiment can also be extended to other deep learning task fields, such as object detection, text classification, speech recognition, semantic segmentation, recommendation system, etc. For text-based deep learning task fields, the images mentioned in this application embodiment can be replaced with text, and for speech-based deep learning task fields, the images mentioned in this application embodiment can be replaced with speech. Further details are not provided here.

[0146] For example, Figure 3 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this application. Please refer to... Figure 3 The image processing device 300 includes: an acquisition module 301, a first determination module 302, a first processing module 303, a second determination module 304, a third determination module 305, and a second processing module 306, wherein...

[0147] Acquisition module 301 is used to acquire the first image;

[0148] The first determining module 302 is used to determine the second image from the first image based on a preset training strategy;

[0149] The first processing module 303 is used to train the first model in the first round based on a preset training strategy and a second image, and to determine a first loss sequence. The first loss sequence is used to indicate the stability of the model when training the first model based on the preset training strategy.

[0150] The second determining module 304 is used to determine the first parameter corresponding to the first loss sequence, and the first parameter is used to indicate the time dependence of the first loss sequence.

[0151] The third determining module 305 is used to determine the target training strategy based on the first parameter;

[0152] The second processing module 306 is used to train the first model in the first round based on the target training strategy and the first image.

[0153] The image processing apparatus provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0154] In one possible implementation, the first processing module 303 is specifically used for:

[0155] During the first round of training of the first model, the loss values ​​for each round are obtained;

[0156] Based on the training time sequence, the loss values ​​of each round are sorted to obtain the first loss sequence.

[0157] In one possible implementation, the second determining module 304 is specifically used for:

[0158] Divide the first loss sequence into a first number of first loss subsequences;

[0159] Based on each first loss subsequence, determine the first parameter corresponding to the first loss sequence.

[0160] In one possible implementation, the second determining module 304 is specifically used for:

[0161] For each first loss subsequence, determine the second and third parameters corresponding to the first loss subsequence. The second parameter is used to indicate the stability of the first model based on the temporal dependence of the first loss subsequence, and the third parameter is used to indicate the stability of the first model based on the volatility of the first loss subsequence.

[0162] The first relationship is determined based on the second and third parameters corresponding to each first loss subsequence;

[0163] Based on the first relationship, determine the first parameter.

[0164] In one possible implementation, the third determining module 305 is specifically used for:

[0165] Based on the first parameter, determine the time dependence of the first loss sequence;

[0166] Based on the degree of time dependence, the target training strategy is determined.

[0167] In one possible implementation, the third determining module 305 is specifically used for:

[0168] If the first parameter is less than or equal to the first threshold, the time dependence of the first loss sequence is determined as the first dependence.

[0169] If the first parameter is greater than the first threshold and less than the second threshold, the time dependence of the first loss sequence is determined as the second dependence.

[0170] If the first parameter is greater than or equal to the second threshold, the time dependence of the first loss sequence is determined as the third dependence.

[0171] In one possible implementation, the third determining module 305 is specifically used for:

[0172] If the time dependence is the highest, increase the learning rate or noise in the preset training strategy to obtain the target training strategy.

[0173] If the time dependence is the second degree of dependence, the preset training strategy will be determined as the target training strategy.

[0174] If the time dependence is third-degree dependence, reduce the learning rate in the preset training strategy to obtain the target training strategy.

[0175] The image processing apparatus provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0176] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 400 may include: a transceiver 401, a processor 402, and a memory 403.

[0177] Processor 402 executes computer execution instructions stored in memory, causing processor 402 to perform the scheme in the above embodiments. Processor 402 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0178] The memory 403 is connected to the processor 402 via the system bus and completes communication between them. The memory 403 is used to store computer program instructions.

[0179] Transceiver 401 can be used to obtain the task to be run and its configuration information.

[0180] The system bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. Transceivers are used to enable communication between database access devices and other computers (e.g., clients, read-write libraries, and read-only libraries). Memory may include random access memory (RAM) and may also include non-volatile memory.

[0181] This application also provides a chip for executing instructions, which is used to execute the image processing method described in the above embodiments.

[0182] This application also provides a computer-readable storage medium storing computer instructions that, when executed on a computer, cause the computer to perform the technical solution of the image processing method described above.

[0183] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0184] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0185] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.

[0186] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit composed of the above modules can be implemented in hardware or in the form of hardware plus software functional units.

[0187] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.

[0188] It should be understood that the aforementioned processor can be a CPU, or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0189] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk or optical disc, etc.

[0190] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0191] The aforementioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0192] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. The processor and storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components in an electronic control unit or main control device.

[0193] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. An image processing method, characterized in that, include: Get the first image; Based on a preset training strategy, a second image is determined from the first image; Based on the preset training strategy and the second image, the first model is trained in the first round, and a first loss sequence is determined. The first loss sequence is used to indicate the stability of the model when the first model is trained based on the preset training strategy. Determine a first parameter corresponding to the first loss sequence, wherein the first parameter is used to indicate the time dependence of the first loss sequence; Based on the first parameter, determine the target training strategy; Based on the target training strategy and the first image, the first model is trained in the first round.

2. The method according to claim 1, characterized in that, Determining the first loss sequence includes: During the first round of training of the first model, the loss values ​​for each round are obtained; Based on the training time sequence, the loss values ​​of each round are sorted to obtain the first loss sequence.

3. The method according to claim 1, characterized in that, Determining the first parameter corresponding to the first loss sequence includes: The first loss sequence is divided into a first number of first loss subsequences; Based on each first loss subsequence, the first parameter corresponding to the first loss sequence is determined.

4. The method according to claim 3, characterized in that, The step of determining the first parameter corresponding to the first loss sequence based on each first loss subsequence includes: For each first loss subsequence, a second parameter and a third parameter corresponding to the first loss subsequence are determined. The second parameter is used to indicate the model stability of the first model based on the temporal dependence of the first loss subsequence, and the third parameter is used to indicate the model stability of the first model based on the volatility of the first loss subsequence. Based on the second and third parameters corresponding to each of the first loss subsequences, the first relationship is determined; Based on the first relationship, the first parameter is determined.

5. The method according to claim 1, characterized in that, The step of determining the target training strategy based on the first parameter includes: Based on the first parameter, determine the time dependence of the first loss sequence; Based on the degree of time dependence, a target training strategy is determined.

6. The method according to claim 5, characterized in that, Determining the time dependence of the first loss sequence based on the first parameter includes: If the first parameter is less than or equal to the first threshold, the time dependence of the first loss sequence is determined to be the first dependence. If the first parameter is greater than the first threshold and less than the second threshold, the time dependence of the first loss sequence is determined to be the second dependence. If the first parameter is greater than or equal to the second threshold, the time dependence of the first loss sequence is determined to be the third dependence.

7. The method according to claim 5, characterized in that, The determination of the target training strategy based on the degree of time dependence includes: If the time dependence is the first degree of dependence, increase the learning rate or noise in the preset training strategy to obtain the target training strategy; If the degree of time dependence is the second degree of dependence, the preset training strategy is determined as the target training strategy; If the time dependence is the third degree of dependence, the learning rate in the preset training strategy is reduced to obtain the target training strategy.

8. An image processing apparatus, characterized in that, include: The module comprises an acquisition module, a first determination module, a first processing module, a second determination module, a third determination module, and a second processing module, wherein... The acquisition module is used to acquire the first image; The first determining module is used to determine a second image from the first image based on a preset training strategy; The first processing module is used to train the first model for a first round based on the preset training strategy and the second image, and to determine a first loss sequence, wherein the first loss sequence is used to indicate the stability of the model when training the first model based on the preset training strategy. The second determining module is used to determine a first parameter corresponding to the first loss sequence, wherein the first parameter is used to indicate the degree of time dependence of the first loss sequence; The third determining module is used to determine the target training strategy based on the first parameter; The second processing module is used to train the first model in the first round based on the target training strategy and the first image.

9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.