Model training method, information generation method, device, equipment and medium

By extracting the features of the training image samples and using the encoding and decoding models to generate the multi-layer decoding layer output results, the problems of large workload and low accuracy of manual labels in the prior art are solved, and the effect of fast and efficient training and improving the robustness of the model is achieved.

CN113297974BActive Publication Date: 2025-05-23BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110572088.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-25
Publication Date
2025-05-23
Estimated Expiration
2041-05-25

AI Technical Summary

Technical Problem

In the training of existing network models, manual annotation training image sample labels require a lot of work, and there may be labeling errors, affecting the model accuracy.

Method used

By extracting the features of the training image samples, generating feature vectors, using the encoding and decoding models to generate the output results of the multi-layer decoding layer, determining the difference information between the decoding layers, generating a loss value set, and adjusting the model parameters to improve robustness.

Benefits of technology

Without labeling the training image sample label, the encoding and decoding model is trained quickly and efficiently, which improves the robustness and accuracy of the model, reduces the workload of manual labeling and reduces the risk of labeling errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113297974B_ABST
    Figure CN113297974B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose a model training method, an information generation method, an apparatus, an electronic device, and a medium. A specific implementation of the method includes: extracting the features of each training image sample in a pre-acquired training image sample set to generate a first feature vector, and obtaining a first feature vector set; generating the output results of at least two decoding layers corresponding to the decoding model in the encoding and decoding model; determining the difference information of the output results between the first target decoding layer and the second target decoding layer in at least two decoding layers, and obtaining at least one difference information; generating a first loss value set based on at least one difference information; in response to determining that the training of the encoding and decoding model is not completed, adjusting the parameters of the encoding and decoding model based on the first loss value set. This implementation can train the encoding and decoding model quickly and efficiently without labeling the training image sample labels, thereby improving the robustness of the encoding and decoding model after training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a model training method, an information generation method, an apparatus, an electronic device, and a computer-readable medium. Background Art

[0002] At present, for the training of network models, each training image sample in the training image sample set often includes: a sample image and a label of the sample image. For the existing network model training, first, the sample image in the training sample image is input into the network model to obtain a predicted label. Then, the loss value is obtained by the difference between the predicted label and the label of the sample image. Finally, according to the above loss value, the parameters of the network model are updated by back propagation.

[0003] However, when using the above method to train the network model, the following technical problems often occur:

[0004] Training sample image labels are often manually annotated, and a large number of training image sample sets leads to a large workload. In addition, manual labeling may have labeling errors, which indirectly affects the accuracy of the network model. Summary of the invention

[0005] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0006] Some embodiments of the present disclosure propose model training methods, devices, electronic devices, and computer-readable media to solve one or more of the technical problems mentioned in the above background technology section.

[0007] In a first aspect, some embodiments of the present disclosure provide a model training method, including: extracting features of each training image sample in a pre-acquired training image sample set to generate a first feature vector, and obtaining a first feature vector set; generating output results of at least two decoding layers corresponding to a decoding model in the encoding and decoding model according to a coding and decoding model and the first feature vector set; determining difference information of output results between a first target decoding layer and a second target decoding layer in the at least two decoding layers, and obtaining at least one difference information, wherein the second target decoding layer is a decoding layer in the at least two decoding layers without the first target decoding layer; generating a first loss value set according to the at least one difference information; in response to determining that the training of the coding and decoding model is not completed, adjusting parameters of the coding and decoding model according to the first loss value set.

[0008] Optionally, the above-mentioned generation of output results of at least two decoding layers corresponding to the decoding model in the above-mentioned encoding and decoding model according to the encoding and decoding model and the above-mentioned first feature vector set includes: generating an input vector of the encoding model in the above-mentioned encoding and decoding model based on the above-mentioned first feature vector set; inputting the above-mentioned input vector into the above-mentioned encoding model to obtain the output results of each layer of the encoding layer; and generating the output results of the above-mentioned at least two layers of decoding layers based on the output results of the above-mentioned each layer of the encoding layer and the above-mentioned at least two layers of decoding layers.

[0009] Optionally, the training sample set includes: a source domain image sample set and a target domain image sample set; and the step of generating the input vector of the encoding model in the encoding and decoding model based on the first feature vector set includes: performing a vector dimension transformation on each first feature vector in the first feature vector set to generate a second feature vector, thereby obtaining a second feature vector set; and generating the input vector of the encoding model based on the second feature vector set and the first initial vector, wherein the first initial vector continuously changes with the training of the encoding and decoding model, and after the training of the encoding and decoding model is completed, the transformed first initial vector represents the global features of each sample in the source domain image sample set and the target domain image sample set.

[0010] Optionally, the output results of the at least two decoding layers are generated based on the output results of the above-mentioned encoding layers and the above-mentioned at least two decoding layers, including: using the output results of the last encoding layer in the above-mentioned encoding layers, which are related to the above-mentioned second feature vector set, as the input of the above-mentioned each decoding layer, and using the above-mentioned first initial vector and at least one second initial vector as the input of the above-mentioned decoding model, to generate the output results of the above-mentioned at least two decoding layers.

[0011] Optionally, adjusting the parameters of the encoding and decoding model according to the first loss value set includes: inputting the encoding results associated with the first initial vector in the output results of the encoding layers into the first domain discrimination model to obtain the second loss values; adjusting the parameters of the encoding model in the encoding and decoding model based on the second loss values, and adjusting the parameters of the encoding and decoding model based on the first loss value set.

[0012] Optionally, adjusting the parameters of the encoding and decoding model according to the first loss value set includes: inputting the encoding result sets associated with the second feature vector set in the output results of the encoding layers into the second domain discrimination model to obtain third loss value sets; adjusting the parameters of the encoding model in the encoding and decoding model based on the second loss values, and adjusting the parameters of the encoding and decoding model based on the first loss value set.

[0013] Optionally, adjusting the parameters of the encoding and decoding model according to the first loss value set includes: inputting the encoding result sets associated with the second feature vector set in the output results of the encoding layers into the second domain discrimination model to obtain third loss value sets; adjusting the parameters of the encoding model in the encoding and decoding model based on the third loss value sets, and adjusting the parameters of the encoding and decoding model based on the first loss value set.

[0014] Optionally, adjusting the parameters of the encoding and decoding model according to the first loss value includes: inputting each decoding result associated with the first initial vector among the output results of the decoding layers into the fourth discriminant model to obtain each fifth loss value; adjusting the parameters of the encoding and decoding model based on the first loss value set and the fifth loss values.

[0015] In a second aspect, some embodiments of the present disclosure provide an information generation method, comprising: acquiring a target image; inputting the target image into a pre-trained encoding and decoding model to obtain image information of each sub-item image present in the target image, wherein the encoding and decoding model is generated by a method such as any one of the embodiments of the present disclosure.

[0016] In a third aspect, some embodiments of the present disclosure provide a model training device, comprising: an extraction unit, configured to extract features of each training image sample in a pre-acquired training image sample set to generate a first feature vector, and obtain a first feature vector set; a first generation unit, configured to generate output results of at least two decoding layers corresponding to a decoding model in the encoding and decoding model according to the encoding and decoding model and the above-mentioned first feature vector set; a determination unit, configured to determine difference information of output results between a first target decoding layer and a second target decoding layer in the above-mentioned at least two decoding layers, and obtain at least one difference information, wherein the above-mentioned second target decoding layer is a decoding layer in the above-mentioned at least two decoding layers without the above-mentioned first target decoding layer; a second generation unit, configured to generate a first loss value set according to the above-mentioned at least one difference information; an adjustment unit, configured to adjust the parameters of the above-mentioned encoding and decoding model according to the above-mentioned first loss value set in response to determining that the training of the above-mentioned encoding and decoding model is not completed.

[0017] Optionally, the first generation unit is configured to: generate an input vector of the encoding model in the encoding and decoding model based on the first feature vector set; input the input vector into the encoding model to obtain the output results of each encoding layer; and generate the output results of the at least two decoding layers based on the output results of each encoding layer and the at least two decoding layers.

[0018] Optionally, the training sample set includes: a source domain image sample set and a target domain image sample set; and the first generation unit is configured to: generate the input vector of the encoding model in the encoding and decoding model based on the first feature vector set, including: performing a vector dimension transformation on each first feature vector in the first feature vector set to generate a second feature vector, and obtain a second feature vector set; based on the second feature vector set and the first initial vector, generate the input vector of the encoding model, wherein the first initial vector continuously changes with the training of the encoding and decoding model, and after the training of the encoding and decoding model is completed, the transformed first initial vector represents the global features of each sample in the source domain image sample set and the target domain image sample set.

[0019] Optionally, the first generation unit is configured to: use the output results of the last encoding layer in the above-mentioned encoding layers that are related to the above-mentioned second feature vector set as the input of the above-mentioned each decoding layer, and use the above-mentioned first initial vector and at least one second initial vector as the input of the above-mentioned decoding model to generate the output results of the above-mentioned at least two decoding layers.

[0020] Optionally, the adjustment unit is configured to: input each encoding result associated with the above-mentioned first initial vector among the output results of the above-mentioned encoding layers into the first domain discrimination model respectively to obtain each second loss value; adjust the parameters of the above-mentioned encoding model in the above-mentioned encoding and decoding model based on the above-mentioned each second loss value, and adjust the parameters of the above-mentioned encoding and decoding model based on the above-mentioned first loss value set.

[0021] Optionally, the adjustment unit is configured to: input each encoding result set associated with the above-mentioned second feature vector set among the output results of the above-mentioned encoding layers into the second domain discrimination model respectively to obtain each third loss value set; based on the above-mentioned each third loss value set, adjust the parameters of the above-mentioned encoding model in the above-mentioned encoding and decoding model, and based on the above-mentioned first loss value set, adjust the parameters of the above-mentioned encoding and decoding model.

[0022] Optionally, the adjustment unit is configured to: input each decoding result set associated with at least one second initial vector among the output results of the above-mentioned decoding layers into the third domain discrimination model to obtain each fourth loss value set; based on the above-mentioned first loss value set and the above-mentioned each fourth loss value set, adjust the parameters of the above-mentioned encoding and decoding model.

[0023] Optionally, the adjustment unit is configured to: input each decoding result associated with the above-mentioned first initial vector among the output results of the above-mentioned decoding layers into the fourth domain discrimination model to obtain each fifth loss value; and adjust the parameters of the above-mentioned encoding and decoding model based on the above-mentioned first loss value set and the above-mentioned each fifth loss value.

[0024] In a fourth aspect, some embodiments of the present disclosure provide an information generating device, comprising: an acquisition unit, configured to acquire a target image; an input unit, configured to input the target image into a pre-trained encoding and decoding model to obtain image information of each sub-item image existing in the target image, wherein the encoding and decoding model is generated by a method as in any embodiment of the present disclosure.

[0025] In a fifth aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect or the second aspect.

[0026] In a sixth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner in the first aspect or the second aspect is implemented.

[0027] The above-mentioned embodiments of the present disclosure have the following beneficial effects: the model training method of some embodiments of the present disclosure can train the encoding and decoding model quickly and efficiently without labeling the training image sample labels, thereby improving the robustness of the encoding and decoding model after training. Specifically, the training sample image labels are often manually labeled, and a large number of training image sample sets result in a large workload. In addition, there may be labeling errors in manual labeling, which indirectly affects the accuracy of the network model. Based on this, the model training method of some embodiments of the present disclosure can first extract the features of each training image sample in the pre-acquired training image sample set to generate a first feature vector for data support of the subsequent encoding and decoding model, and obtain a first feature vector set. Then, according to the encoding and decoding model and the above-mentioned first feature vector set, the output results of at least two layers of decoding layers corresponding to the decoding model in the above-mentioned encoding and decoding model are generated. In some embodiments of the present disclosure, the first feature vector set can be input into the encoding and decoding model to obtain the output results of at least two layers of decoding layers corresponding to the decoding model in the above-mentioned encoding and decoding model. Here, decoding by multiple decoding layers can make the accuracy of the subsequent training encoding and decoding model higher. The robustness of the training encoding and decoding model can be improved. Furthermore, the difference information of the output results between the first target decoding layer and the second target decoding layer in the at least two decoding layers is determined to obtain at least one difference information, wherein the second target decoding layer is a decoding layer in the at least two decoding layers without the first target decoding layer. Here, by determining the difference information of the output results between the first target decoding layer and the second target decoding layer to further train the encoding and decoding model, the accuracy of the encoding and decoding model can be higher and the robustness can be stronger. Then, according to the at least one difference information, a first loss value set is generated. Finally, in response to determining that the training of the encoding and decoding model is not completed, according to the first loss value set, the parameters of the encoding and decoding model are adjusted to further improve the robustness of the encoding and decoding network. Based on this, the model training method can quickly and efficiently train the encoding and decoding model without labeling the training image sample labels, thereby improving the robustness of the encoding and decoding model after training. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.

[0029] Figure 1 is a schematic diagram of an application scenario of a model training method according to some embodiments of the present disclosure;

[0030] Figure 2 is a flowchart of some embodiments of the model training method according to the present disclosure;

[0031] Figure 3 is a flowchart of other embodiments of the model training method according to the present disclosure;

[0032] Figure 4 is a flow chart of some embodiments of the information generation method according to the present disclosure;

[0033] Figure 5 is a schematic diagram of the structure of some embodiments of the model training device according to the present disclosure;

[0034] Figure 6 is a schematic diagram of the structure of some embodiments of the information generating device according to the present disclosure;

[0035] Figure 7 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0036] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0037] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0038] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0039] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0040] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0041] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0042] Figure 1 It is a schematic diagram of an application scenario of the model training method according to some embodiments of the present disclosure.

[0043] exist Figure 1In the application scenario, the electronic device 101 may first extract the features of each training image sample in the pre-acquired training image sample set 102 to generate a first feature vector, and obtain a first feature vector set 103. In this application scenario, the training image sample set 102 may include: a first training image sample 1021, a second training image sample 1022, and a third training image sample 1023. The features of the first training image sample 1021 are extracted to generate a first feature vector 1031. The features of the second training image sample 1022 are extracted to generate a first feature vector 1032. The features of the third training image sample 1023 are extracted to generate a first feature vector 1033. Then, according to the encoding and decoding model 104 and the first feature vector set 103, the output results of at least two decoding layers corresponding to the decoding model 106 in the encoding and decoding model 104 are generated. The encoding and decoding model 104 may include: an encoding model 105 and a decoding model 106. The coding model 105 may include: a first coding layer 1051, a second coding layer 1052, and a third coding layer 1053. The decoding model 106 may include: a first decoding layer 1061, a second decoding layer 1062, and a third decoding layer 1063. The output result of the first decoding layer 1061 may be an output result 107. The output result of the second decoding layer 1062 may be an output result 108. The output result of the third decoding layer 1063 may be an output result 109. Determine the difference information of the output results between the first target decoding layer and the second target decoding layer in the at least two decoding layers, and obtain at least one difference information 110. The second target decoding layer is a decoding layer in the at least two decoding layers without the first target decoding layer. In this application scenario, the third target decoding layer may be the third decoding layer 1063. The second target decoding layer may be the second decoding layer 1062. The first target decoding layer may be the first decoding layer 1061. The difference information 1101 may be the difference information between the output result 107 and the output result 109. The difference information 1102 may be the difference information between the output result 108 and the output result 109. A first loss value set 111 is generated according to the at least one difference information 110. The first loss set 111 includes: a first loss value 1111 corresponding to the first feature vector 1031, a first loss value 1112 corresponding to the first feature vector 1032, and a first loss value 1113 corresponding to the first feature vector 1033. Finally, in response to determining that the training of the encoding and decoding model 104 is not completed, the parameters of the encoding and decoding model 104 are adjusted according to the first loss value set 111.

[0044] It should be noted that the electronic device 101 can be hardware or software. When the electronic device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or it can be implemented as a single server or a single terminal device. When the electronic device is embodied as software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules for providing distributed services, for example, or it can be implemented as a single software or software module. No specific limitation is made here.

[0045] It should be understood that Figure 1 The number of electronic devices in the embodiment is only for illustration. Any number of electronic devices may be provided according to implementation requirements.

[0046] Continue to refer Figure 2 , shows a process 200 of some embodiments of the model training method according to the present disclosure. The model training method comprises the following steps:

[0047] Step 201 : extracting features of each training image sample in a pre-acquired training image sample set to generate a first feature vector, thereby obtaining a first feature vector set.

[0048] In some embodiments, the execution entity of the above model training (for example Figure 1 The electronic device shown in the figure can extract the features of each training image sample in the pre-acquired training image sample set to generate a first feature vector, thereby obtaining a first feature vector set. As an example, the training image sample set may include: a source domain image sample set and / or a target domain training sample set.

[0049] As an example, the execution entity may input each training image sample in the training image sample set into a pre-trained convolutional neural network (CNN) to generate a first feature vector, thereby obtaining a first feature vector set.

[0050] Step 202: Generate output results of at least two decoding layers corresponding to the decoding model in the encoding and decoding model according to the encoding and decoding model and the first feature vector set.

[0051] In some embodiments, the execution subject may generate the output results of at least two layers of decoding layers corresponding to the decoding model in the encoding and decoding model according to the encoding and decoding model and the first feature vector set. As an example, the encoding and decoding model may be a seq2seq (sequence to sequence) model. The encoding model may include at least one encoding layer. Each encoding layer and each decoding layer may be one of the following: a long short-term memory network (LSTM, Long Short-Term Memory), a GRU (Gate Recurrent Unit) model, a recurrent neural network (RNN). It can be understood that the encoding and decoding model may be used as a model for various tasks. For example, the encoding and decoding model may be a model for target object detection. It may also be a model for face recognition. It may also be a model for image segmentation. It is not limited here.

[0052] It should be noted that, in response to the encoding and decoding model being a target detection model, the output result of the above decoding model may be the item category probability and the location information of the sub-item image corresponding to the sub-item image that may exist in the training image sample.

[0053] As an example, the execution subject may first perform a vector dimension transformation on each feature vector in the first feature vector set to generate a third feature vector, thereby obtaining a third feature vector set. Then, the third feature vector set is input into the encoding and decoding model to generate output results of at least two decoding layers corresponding to the decoding model in the encoding and decoding model.

[0054] Step 203: determine difference information of output results between a first target decoding layer and a second target decoding layer in the at least two decoding layers to obtain at least one difference information.

[0055] In some embodiments, the execution subject may determine difference information of output results between a first target decoding layer and a second target decoding layer in the at least two decoding layers to obtain at least one difference information. The second target decoding layer is a decoding layer in the at least two decoding layers without the first target decoding layer. As an example, the first target decoding layer may be the last decoding layer of the at least two decoding layers.

[0056] As an example, if the encoding and decoding model is a target detection model, the execution entity may determine the difference information of the output results between the first target decoding layer and the second target decoding layer in the at least two decoding layers by the following formula:

[0057]

[0058] in, It can be the output result of the lth decoding layer in the decoding model It can represent the difference information between the output results of two decoding layers. KL(·,·) can represent KL divergence. ||·|| 1 It can represent L1 regularization. L1 It can be the loss weight coefficient. It can represent the category probability of the i-th item in the training sample image in the output of the l-th decoding layer of the decoding model. Similarly, It can represent the detection box of the i-th object in the training sample image in the output of the l-th decoding layer of the decoding model.

[0059] Here, the above encoding and decoding models are target detection models. It can be expressed by the following formula:

[0060]

[0061] Among them, N obj It can be the number of sub-item images included in the training sample images.

[0062] Here, by constraining the outputs of different decoding layers to be consistent, the training of the model can be simplified without the need for labeling, and its robustness in the encoding and decoding model can be improved.

[0063] It should be noted that for the decoding model, within a certain number of layers, the more decoding layers there are, the higher the accuracy of the corresponding decoded output will be. This makes the output of the lower decoding layer closer to the output of the higher decoding layer. This can indirectly make the output accuracy of the subsequent higher decoding layer higher.

[0064] Step 204: Generate a first loss value set based on the at least one difference information.

[0065] In some embodiments, the execution entity may generate a first loss value set based on the at least one difference information.

[0066] It should be noted that each training image sample has a corresponding first loss value. A training image sample set has a corresponding first loss value set. The above difference information can be the difference information between the output of the last decoding layer of the decoding model corresponding to each training image sample and the output of any decoding layer in the multiple decoding layers except the last decoding layer in the decoding model corresponding to each training image sample.

[0067] As an example, the first loss value corresponding to each training image sample can be determined by the following formula:

[0068]

[0069] in, It can be the first loss value. dec It can be the number of decoding layers in the decoding model. It can be the output result of the last decoding layer in the decoding model. l can be the level at which the second target decoding layer is located in the decoding model. It can be the output result of the lth decoding layer in the decoding model.

[0070] Step 204, in response to determining that the encoding and decoding model training is not completed, adjusting the parameters of the encoding and decoding model according to the first loss value set.

[0071] In some embodiments, in response to determining that the encoding and decoding model training is not completed, the execution entity may adjust the parameters of the encoding and decoding model according to the first loss value set.

[0072] As an example, in response to determining that the encoding and decoding model training is not completed, according to the first loss value set, the execution entity may adjust the parameters of the encoding and decoding model by back propagation.

[0073] The above-mentioned embodiments of the present disclosure have the following beneficial effects: the model training method of some embodiments of the present disclosure can train the encoding and decoding model quickly and efficiently without labeling the training image sample labels, thereby improving the robustness of the encoding and decoding model after training. Specifically, the training sample image labels are often manually labeled, and a large number of training image sample sets result in a large workload. In addition, there may be labeling errors in manual labeling, which indirectly affects the accuracy of the network model. Based on this, the model training method of some embodiments of the present disclosure can first extract the features of each training image sample in the pre-acquired training image sample set to generate a first feature vector for data support of the subsequent encoding and decoding model, and obtain a first feature vector set. Then, according to the above-mentioned first feature vector set and the encoding and decoding model, the output results of at least two layers of decoding layers corresponding to the decoding model in the above-mentioned encoding and decoding model are generated. In some embodiments of the present disclosure, the first feature vector set can be input into the encoding and decoding model to obtain the output results of at least two layers of decoding layers corresponding to the decoding model in the above-mentioned encoding and decoding model. Here, decoding by multiple decoding layers can make the accuracy of the subsequent training encoding and decoding model higher. The robustness of the training encoding and decoding model can be improved. Furthermore, the difference information of the output results between the first target decoding layer and the second target decoding layer in the at least two decoding layers is determined to obtain at least one difference information, wherein the second target decoding layer is a decoding layer in the at least two decoding layers without the first target decoding layer. Here, by determining the difference information of the output results between the first target decoding layer and the second target decoding layer to further train the encoding and decoding model, the accuracy of the encoding and decoding model can be higher and the robustness can be stronger. Then, according to the at least one difference information, a first loss value set is generated. Finally, in response to determining that the training of the encoding and decoding model is not completed, according to the first loss value set, the parameters of the encoding and decoding model are adjusted to further improve the robustness of the encoding and decoding network. Based on this, the model training method can quickly and efficiently train the encoding and decoding model without labeling the training image sample labels, thereby improving the robustness of the encoding and decoding model after training.

[0074] Further references Figure 3 , shows a process 300 of other embodiments of the model training method according to the present disclosure. The model training method comprises the following steps:

[0075] Step 301 : extracting features of each training image sample in a pre-acquired training image sample set to generate a first feature vector, thereby obtaining a first feature vector set.

[0076] Step 302: Perform vector dimensionality transformation on each first feature vector in the above first feature vector set to generate a second feature vector, and obtain a second feature vector set.

[0077] In some embodiments, the execution subject of the model training method (e.g., Figure 1 the electronic device 101 in

[0078] can perform vector dimensionality transformation on each first feature vector in the above first feature vector set to generate a second feature vector, and obtain a second feature vector set.

[0079] As an example, the above execution subject can perform vector embedding processing on each first feature vector in the above first feature vector set to change the vector dimensionality transformation, and obtain a second feature vector set.

[0080] As an example, the above execution subject can flatten each first feature vector in the above first feature vector set into a one-dimensional vector and perform feature mapping to generate a second feature vector, and obtain a second feature vector set.

[0081] Step 303: Generate the input vector of the above encoding model based on the above second feature vector set and the first initial vector.

[0082] In some embodiments, the above execution subject can generate the input vector of the above encoding model based on the above second feature vector set and the first initial vector. Among them, the above first initial vector changes continuously with the training of the above encoding and decoding models. After the training of the above encoding and decoding models ends, the transformed first initial vector represents the global features of each sample in the above source domain image sample set and the above target domain image sample set. Among them, the above training sample set includes: a source domain image sample set and a target domain image sample set.

[0083] As another example, the following formula can be referred to generate the input vector:

[0084]

[0085] where, Z 0It can be an input vector, that is, a vector obtained by concatenating the second feature vector set, the first initial vector, the position code, and the feature level code. It can be the first initial vector. It can be the Nth second eigenvector. E pos Can be a positional code. level It can be a feature level encoding. N can be the number of second feature vectors in the second feature vector set. D can be the dimension of feature embedding.

[0086] Step 304: Based on the output results of the encoding layers and the at least two decoding layers, the output results of the at least two decoding layers are generated.

[0087] In some embodiments, the execution entity may generate output results of the at least two decoding layers based on output results of the encoding layers and the at least two decoding layers.

[0088] As an example, the execution subject may first obtain the output result of the last coding layer among the coding layers, and then use the output result of the last coding layer as the input of the at least two coding layers to generate the output results of the at least two decoding layers.

[0089] In some optional implementations of some embodiments, the execution subject may use the output results of the last coding layer in the coding layers as the input of each decoding layer, and use the first initial vector and at least one second initial vector as the input of the decoding model to generate the output results of the at least two decoding layers. Each of the at least one second initial vector may be an embedding vector for object prediction (classification and positioning) of a source domain image sample or a target domain image sample. As the parameters of the coding and decoding model are adjusted, the at least one second initial vector changes accordingly. The output results of the last coding layer in the coding layers as the output results related to the second feature vector set may be the output results of the last coding layer in the coding layers minus the output results related to the first initial vector.

[0090] It should be noted that the input of the first decoding layer of the above decoding model can be the concatenation result of the output result of the last encoding layer, the first initial vector and at least one second initial vector. The input of the remaining decoding layers of the above decoding model can be the concatenation result of the output result of the previous decoding layer and the output result related to the above second feature vector set.

[0091] Step 305 : determining difference information of output results between a first target decoding layer and a second target decoding layer in the at least two decoding layers, and obtaining at least one difference information.

[0092] Step 306: Generate a first loss value set based on the at least one difference information.

[0093] Step 307: Input the encoding results associated with the first initial vector in the output results of the encoding layers into the first domain discrimination model to obtain the second loss values.

[0094] In some embodiments, the execution entity may input the encoding results associated with the first initial vector in the output results of the encoding layers into the first domain discrimination model to obtain the second loss values.

[0095] Here, the training sample set includes a source domain image sample set and a target domain image sample set. Existing encoding and decoding networks often assume that the training data and test data come from the same data distribution. In actual application scenarios, due to changes in weather, scene changes, etc., the data used by the training model and the data to be processed in the model deployment phase do not have the same data distribution, that is, there is a domain gap. The existence of the domain gap will significantly reduce the performance of the trained model in the test phase, that is, the generalization is poor. Unsupervised encoding and decoding models can train the model on a labeled source domain, and can generalize well to another unlabeled target domain with a different data distribution, thereby reducing the burden of manual labeling on the target domain. Therefore, this problem has important practical application value.

[0096] There is a certain difference between the data distribution in the source domain image sample set and the data distribution in the target domain image sample set. By inputting each encoding result associated with the above-mentioned first initial vector into the first domain discriminant model, feature alignment is performed on the scene layout and other aspects of the training sample image. Here, since the second feature vector set is obtained by expanding the feature map of the image, the first initial vector can aggregate the global feature information about the source domain image samples and the target domain image samples. In this way, the first domain discriminant model can be used to well align the global feature information of the second feature vector set in the encoding model.

[0097] Step 308, in response to determining that the encoding and decoding model training is not completed, adjusting the parameters of the encoding model in the encoding and decoding model based on the above-mentioned second loss values, and adjusting the parameters of the encoding and decoding model based on the above-mentioned first loss value set.

[0098] In some embodiments, in response to determining that the encoding and decoding model training is not completed, the execution subject may adjust the parameters of the encoding model in the encoding and decoding model based on the second loss values, and adjust the parameters of the encoding and decoding model based on the first loss value set. Since each second loss value is a loss value related to the encoding model in the encoding and decoding model, each second loss value is only used to adjust the parameters of the encoding model.

[0099] As an example, in response to determining that the encoding and decoding model training is not completed, the execution subject may determine the target loss value corresponding to the encoding model in the encoding and decoding model according to each second loss value, wherein the target loss value is obtained by the loss value corresponding to each encoding layer. The loss value corresponding to each encoding layer may be obtained by the following formula:

[0100]

[0101] in, It can be the loss value corresponding to the lth coding layer in the coding model. d can be the domain label. For the source domain image sample, d can be the value 0. For the source domain image sample, d can be the value 1. It can be the encoding result corresponding to the first initial vector corresponding to the lth encoding layer. It can be a first domain discriminant model. It may be the second loss value obtained by inputting the encoding result corresponding to the first initial vector and corresponding to the lth encoding layer in the encoding model into the first domain discrimination model.

[0102] Among them, the target loss value corresponding to the encoding model can be obtained by summing the loss values ​​corresponding to each encoding layer.

[0103] In some optional implementations of some embodiments, adjusting the parameters of the encoding and decoding model according to the first loss value set may include the following steps:

[0104] In the first step, each encoding result set associated with the second feature vector set in the output results of each encoding layer is input into the second domain discrimination model to obtain each third loss value set.

[0105] There is a certain difference between the data distribution of the source domain image sample set and the data distribution of the target domain image sample set. By inputting the encoding result sets associated with the above-mentioned second feature vector set in the output results of each encoding layer into the second domain discriminant model, the local features of the training sample image (for example, local texture, appearance and other features in the image) are migrated. Here, since the second feature vector set is obtained by expanding the feature map of the image, the second domain discriminant model can be used to align the local feature information of the encoder related to the second feature vector set.

[0106] It should be noted that the encoding result sets associated with the second feature vector set may be the output results of each encoding layer, excluding the encoding results associated with the first initial vector.

[0107] In the second step, based on the above-mentioned third loss value sets, the parameters of the above-mentioned encoding model in the above-mentioned encoding and decoding model are adjusted, and based on the above-mentioned first loss value set, the parameters of the above-mentioned encoding and decoding model are adjusted. Among them, since each third loss value is a loss value related to the encoding model in the encoding and decoding model, each third loss value set only adjusts the parameters of the above-mentioned encoding model.

[0108] As an example, in response to determining that the encoding and decoding model training is not completed, the execution subject may determine the target loss value corresponding to the encoding model in the encoding and decoding model based on the third loss value sets, wherein the target loss value is obtained by the loss value corresponding to each encoding layer. The loss value corresponding to each encoding layer may be obtained by the following formula:

[0109]

[0110] in, It can be the encoding result corresponding to the i-th second feature vector corresponding to the l-th encoding layer. It can be a second domain discriminant model.

[0111] Among them, the target loss value corresponding to the encoding model can be obtained by summing the loss values ​​corresponding to each encoding layer.

[0112] In some optional implementations of some embodiments, adjusting the parameters of the encoding and decoding model according to the first loss value set may include the following steps:

[0113] In the first step, each encoding result set associated with at least one second initial vector in the output results of each decoding layer is input into the third domain discrimination model to obtain each fourth loss value set.

[0114] There is a certain difference between the data distribution of the source domain image sample set and the data distribution of the target domain image sample set. By inputting the encoding result sets associated with at least one second initial vector in the output results of the above-mentioned decoding layers into the third domain discrimination model, feature alignment is performed on the individual level of the foreground objects of the training sample images.

[0115] The second step is to adjust the parameters of the encoding and decoding model based on the first loss value set and the fourth loss value sets.

[0116] As an example, in response to determining that the encoding and decoding model training is not completed, the execution entity may determine the target loss value corresponding to the decoding model in the encoding and decoding model based on the fourth loss value sets, wherein the target loss value is obtained by the loss value corresponding to each decoding layer. The loss value corresponding to each decoding layer may be obtained by the following formula:

[0117]

[0118] in, It may be the encoding result corresponding to the i-th second initial vector corresponding to the l-th encoding layer. It can be a third domain discriminant model.

[0119] Among them, the target loss value corresponding to the decoding model can be obtained by summing the loss values ​​corresponding to each decoding layer.

[0120] In some optional implementations of some embodiments, adjusting the parameters of the encoding and decoding model according to the first loss value set includes:

[0121] In the first step, each encoding result associated with the first initial vector in the output results of each decoding layer is input into the fourth discriminant model to obtain each fifth loss value.

[0122] There are certain differences between the data distribution of the source domain image sample set and the data distribution of the target domain image sample set. By inputting the encoding results associated with the first initial vector in the output results of the above decoding layers into the fourth domain discrimination model, feature alignment is performed on the training sample images in terms of the relationship between objects, the relationship between foreground and background, etc.

[0123] The second step is to adjust the parameters of the encoding and decoding model based on the first loss value set and the fifth loss values.

[0124] As an example, in response to determining that the encoding and decoding model training is not completed, the execution entity may determine the target loss value corresponding to the decoding model in the encoding and decoding model based on the fifth loss values, wherein the target loss value is obtained by the loss value corresponding to each decoding layer. The loss value corresponding to each decoding layer may be obtained by the following formula:

[0125]

[0126] in, It can be the decoding result of the first initial vector of the lth decoding layer. The decoding result of the first initial vector of the lth decoding layer may be input into the fifth loss value of the fourth domain discrimination model. It can be a fourth domain discriminant model.

[0127] Among them, the target loss value corresponding to the decoding model can be obtained by summing the loss values ​​corresponding to each decoding layer.

[0128] In some embodiments, the specific implementation of steps 301, 305 and 306 and the technical effects thereof can be referred to in Figure 2 Steps 201, 203 and 204 in the corresponding embodiment are not described in detail here.

[0129] from Figure 3 It can be seen that Figure 2 Compared with the description of some corresponding embodiments, Figure 3 The process 300 of the model training method in some corresponding embodiments further highlights the specific steps of feature alignment of the encoding model and the decoding model in the encoding model stage. Therefore, the encoding and decoding models trained by the schemes described in these embodiments can be more efficiently and accurately applied to the target domain image, making the encoding and model more robust.

[0130] Further references Figure 4 , shows a process 400 of another embodiment of the information generation method according to the present disclosure. The information generation method comprises the following steps:

[0131] Step 401, acquiring a target image.

[0132] In some embodiments, the execution subject of the information generation method (eg Figure 1 The electronic device 101) can acquire the target image through a wired or wireless method.

[0133] Step 402: input the target image into a pre-trained encoding and decoding model to obtain image information of each sub-item image in the target image.

[0134] In some embodiments, the execution subject may input the target image into a pre-trained encoding and decoding model to obtain image information of each sub-item image in the target image, wherein the image information may include the item category probability and item detection frame of the item corresponding to the sub-item image.

[0135] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: the information generating method can efficiently generate the image information of each sub-item image existing in the target image.

[0136] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a model training device, and these device embodiments are Figure 2 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0137] like Figure 5 As shown, a model training device 500 includes: an extraction unit 501, a first generation unit 502, a determination unit 503, a second generation unit 504 and an adjustment unit 505. The extraction unit 501 is configured to extract the features of each training image sample in the pre-acquired training image sample set to generate a first feature vector and obtain a first feature vector set; the first generation unit 502 is configured to generate the output results of at least two decoding layers corresponding to the decoding model in the encoding and decoding model according to the encoding and decoding model and the first feature vector set; the determination unit 503 is configured to determine the difference information of the output results between the first target decoding layer and the second target decoding layer in the at least two decoding layers to obtain at least one difference information, wherein the second target decoding layer is the decoding layer of the at least two decoding layers without the first target decoding layer; the second generation unit 504 is configured to generate a first loss value set according to the at least one difference information; the adjustment unit 505 is configured to adjust the parameters of the encoding and decoding model according to the first loss value set in response to determining that the training of the encoding and decoding model is not completed.

[0138] In some optional implementations of some embodiments, the first generation unit 502 in the model training device 500 is configured to: generate an input vector of the encoding model in the encoding and decoding model based on the first feature vector set; input the input vector into the encoding model to obtain the output results of each encoding layer; and generate the output results of the at least two decoding layers based on the output results of each encoding layer and the at least two decoding layers.

[0139] In some optional implementations of some embodiments, the training sample set includes: a source domain image sample set and a target domain image sample set; and the first generation unit 502 in the model training device 500 is configured to: based on the first feature vector set, generate the input vector of the encoding model in the encoding and decoding model, including: performing a vector dimension transformation on each first feature vector in the first feature vector set to generate a second feature vector, and obtain a second feature vector set; based on the second feature vector set and the first initial vector, generate the input vector of the encoding model, wherein the first initial vector continuously changes with the training of the encoding and decoding model, and after the training of the encoding and decoding model is completed, the transformed first initial vector represents the global features of each sample in the source domain image sample set and the target domain image sample set.

[0140] In some optional implementations of some embodiments, the first generation unit 502 in the model training device 500 is configured to: use the output results of the last encoding layer in the above-mentioned encoding layers that are related to the above-mentioned second feature vector set as the input of the above-mentioned each decoding layer, and use the above-mentioned first initial vector and at least one second initial vector as the input of the above-mentioned decoding model to generate the output results of the above-mentioned at least two decoding layers.

[0141] In some optional implementations of some embodiments, the adjustment unit 505 in the model training device 500 is configured to: input each encoding result associated with the above-mentioned first initial vector among the output results of the above-mentioned encoding layers into the first domain discrimination model respectively to obtain each second loss value; based on the above-mentioned each second loss value, adjust the parameters of the above-mentioned encoding model in the above-mentioned encoding and decoding model, and based on the above-mentioned first loss value set, adjust the parameters of the above-mentioned encoding and decoding model.

[0142] In some optional implementations of some embodiments, the adjustment unit 505 in the model training device 500 is configured to: input each encoding result set associated with the above-mentioned second feature vector set in the output results of the above-mentioned encoding layers into the second domain discrimination model respectively to obtain each third loss value set; based on the above-mentioned each third loss value set and the above-mentioned first loss value set, adjust the parameters of the above-mentioned encoding model in the above-mentioned encoding and decoding model.

[0143] In some optional implementations of some embodiments, the adjustment unit 505 in the model training device 500 is configured to: input each encoding result set associated with at least one second initial vector in the output results of the above-mentioned decoding layers into the third domain discrimination model to obtain each fourth loss value set; based on the above-mentioned each third loss value set, adjust the parameters of the above-mentioned encoding model in the above-mentioned encoding and decoding model, and based on the above-mentioned first loss value set, adjust the parameters of the above-mentioned encoding and decoding model.

[0144] In some optional implementations of some embodiments, the adjustment unit 505 in the model training device 500 is configured to: input each encoding result associated with the above-mentioned first initial vector in the output results of the above-mentioned decoding layers into the fourth discriminant model to obtain each fifth loss value; based on the above-mentioned first loss value set and the above-mentioned each fifth loss value, adjust the parameters of the above-mentioned encoding and decoding model.

[0145] It is understood that the units described in the device 500 are similar to those described in the reference Figure 2 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 500 and the units included therein, and will not be described in detail here.

[0146] Further references Figure 6 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an information generating device, and these device embodiments are Figure 4 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0147] like Figure 6 As shown, an information generating device 600 includes: an acquisition unit 601 and an input unit 605. The acquisition unit 601 is configured to acquire a target image; the input unit 602 is configured to input the target image into a pre-trained encoding and decoding model to obtain image information of each sub-item image existing in the target image, wherein the encoding and decoding model is generated by a method as in any embodiment of the present disclosure.

[0148] It is understood that the units described in the device 600 are similar to those described in the reference Figure 4 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 600 and the units included therein, and will not be described in detail here.

[0149] Reference below Figure 7 , which shows an electronic device (eg, Figure 1 A schematic diagram of the structure of an electronic device in (a) 700. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0150] like Figure 7 As shown, the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0151] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 The electronic device 700 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 7 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0152] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0153] It should be noted that the computer-readable medium in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0154] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0155] The computer-readable medium may be included in the electronic device; or it may exist independently without being installed in the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: extracts the features of each training image sample in the pre-acquired training image sample set to generate a first feature vector, and obtains a first feature vector set; generates output results of at least two decoding layers corresponding to the decoding model in the encoding and decoding model according to the encoding and decoding model and the first feature vector set; determines the difference information of the output results between the first target decoding layer and the second target decoding layer in the at least two decoding layers, and obtains at least one difference information, wherein the second target decoding layer is the decoding layer of the at least two decoding layers without the first target decoding layer; generates a first loss value set according to the at least one difference information; in response to determining that the training of the encoding and decoding model is not completed, adjust the parameters of the encoding and decoding model according to the first loss value set. Acquire a target image; input the target image into a pre-trained encoding and decoding model to obtain image information of each sub-item image present in the target image, wherein the encoding and decoding model is generated by a method as in any embodiment of the present disclosure.

[0156] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0157] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0158] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be arranged in a processor, for example, may be described as: a processor includes an extraction unit, a first generation unit, a determination unit, a second generation unit and an adjustment unit. The names of these units do not constitute a limitation on the units themselves in certain cases, for example, the extraction unit may also be described as "a unit that extracts the features of each training image sample in a pre-acquired training image sample set to generate a first feature vector and obtain a first feature vector set".

[0159] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0160] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.

Claims

1. A model training method, include: Extracting features of each training image sample in the pre-acquired training image sample set to generate a first feature vector, thereby obtaining a first feature vector set; Generate output results of at least two decoding layers corresponding to a decoding model in the encoding and decoding model according to the encoding and decoding model and the first feature vector set; Determine difference information of output results between a first target decoding layer and a second target decoding layer in the at least two decoding layers to obtain at least one difference information, wherein the second target decoding layer is a decoding layer in the at least two decoding layers without the first target decoding layer; generating a first loss value set according to the at least one difference information; In response to determining that the encoding and decoding model training is not completed, adjusting the parameters of the encoding and decoding model according to the first loss value set.

2. The method according to claim 1, in, The step of generating output results of at least two decoding layers corresponding to a decoding model in the encoding and decoding model according to the encoding and decoding model and the first feature vector set includes: Based on the first feature vector set, generating an input vector of the encoding model in the encoding and decoding model; Inputting the input vector into the coding model to obtain output results of each coding layer; Based on the output results of the encoding layers and the at least two decoding layers, the output results of the at least two decoding layers are generated.

3. The method according to claim 2, in, The training image sample set includes: a source domain image sample set and a target domain image sample set; and The step of generating an input vector of an encoding model in the encoding and decoding model based on the first feature vector set includes: Performing vector dimension transformation on each first eigenvector in the first eigenvector set to generate a second eigenvector, thereby obtaining a second eigenvector set; Based on the second feature vector set and the first initial vector, an input vector of the encoding model is generated, wherein the first initial vector is continuously transformed as the encoding and decoding model is trained, and after the encoding and decoding model training is completed, the transformed first initial vector represents the global features of each sample in the source domain image sample set and the target domain image sample set.

4. The method according to claim 3, in, The generating the output results of the at least two decoding layers based on the output results of the encoding layers and the at least two decoding layers comprises: The output results of the last encoding layer in the encoding layers that are related to the second feature vector set are used as inputs to the respective decoding layers, and the first initial vector and at least one second initial vector are used as inputs to the decoding model to generate output results of the at least two decoding layers.

5. The method according to claim 3, in, The adjusting the parameters of the encoding and decoding model according to the first loss value set includes: Inputting the encoding results associated with the first initial vector in the output results of the encoding layers into the first domain discrimination model to obtain second loss values; Based on the respective second loss values, the parameters of the encoding model in the encoding and decoding model are adjusted, and based on the first loss value set, the parameters of the encoding and decoding model are adjusted.

6. The method according to claim 3, in, The adjusting the parameters of the encoding and decoding model according to the first loss value set includes: Inputting the encoding result sets associated with the second feature vector set in the output results of the encoding layers into the second domain discrimination model respectively to obtain third loss value sets; Based on the respective third loss value sets, the parameters of the encoding model in the encoding and decoding model are adjusted, and based on the first loss value set, the parameters of the encoding and decoding model are adjusted.

7. The method according to claim 4, in, The adjusting the parameters of the encoding and decoding model according to the first loss value set includes: Inputting each decoding result set associated with at least one second initial vector in the output results of each decoding layer into a third domain discrimination model to obtain each fourth loss value set; Based on the first loss value set and the respective fourth loss value sets, the parameters of the encoding and decoding model are adjusted.

8. The method according to claim 4, in, The adjusting the parameters of the encoding and decoding model according to the first loss value includes: Inputting each decoding result associated with the first initial vector among the output results of each decoding layer into a fourth domain discrimination model to obtain each fifth loss value; Based on the first loss value set and the respective fifth loss values, the parameters of the encoding and decoding model are adjusted.

9. A method for generating information, include: Get the target image; The target image is input into a pre-trained encoding and decoding model to obtain image information of each sub-item image present in the target image, wherein the encoding and decoding model is generated by the method described in any one of claims 1-8.

10. A model training device, include: An extraction unit is configured to extract a feature of each training image sample in a pre-acquired training image sample set to generate a first feature vector, thereby obtaining a first feature vector set; A first generating unit is configured to generate output results of at least two decoding layers corresponding to a decoding model in the encoding and decoding model according to the encoding and decoding model and the first feature vector set; a determining unit configured to determine difference information of output results between a first target decoding layer and a second target decoding layer in the at least two decoding layers, and obtain at least one difference information, wherein the second target decoding layer is a decoding layer in the at least two decoding layers without the first target decoding layer; A second generating unit is configured to generate a first loss value set according to the at least one difference information; An adjustment unit is configured to adjust the parameters of the encoding and decoding model according to the first loss value set in response to determining that the encoding and decoding model training is not completed.

11. A target detection device, include: an acquisition unit, configured to acquire a target image; An input unit is configured to input the target image into a pre-trained encoding and decoding model to obtain image information of each sub-item image existing in the target image, wherein the encoding and decoding model is generated by the method described in any one of claims 1-8.

12. An electronic device, include: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.

13. A computer readable medium having a computer program stored thereon, in, When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Data processing model training method and device

    CN111639684A