Model training method and apparatus, and electronic device and storage medium
By enabling batch scaling of image sample sets, keyword generation, and model training parameter configuration within the model training window, the problem of cumbersome tool switching during model training in existing technologies is solved, thereby improving training efficiency and reducing user workload.
Patent Information
- Application Number
- PCT/CN2025/106483
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-30
- Filing Date
- 2025-07-01
- Publication Date
- 2026-02-05
AI Technical Summary
In existing technologies, the model training process requires users to switch between multiple auxiliary tools, which is cumbersome and requires a lot of manpower, resulting in low training efficiency and failing to improve work efficiency.
By enabling batch scaling of image sample sets, keyword generation, and model training parameter configuration within the model training window, a one-stop model training method is provided, reducing the need for users to switch between multiple tools.
It enables one-stop model training, improves training efficiency, reduces the need for users to switch between multiple auxiliary tools, and reduces labor time.
Smart Images

Figure CN2025106483_05022026_PF_FP_ABST
Abstract
Description
Model training method and device, electronic equipment and storage medium
[0001] Cross-reference to related applications
[0002] This application claims priority to the Chinese patent application No. 202411033522.1, filed on July 30, 2024, and entitled "Model training method and device, electronic equipment and storage medium", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0003] The present disclosure relates to the technical field of model training, and in particular, to a model training method, device, electronic equipment and storage medium. BACKGROUND
[0004] This section is intended to provide background or context to the embodiments of the disclosure recited in the claims. The description herein does not constitute admission that the prior art is prior art nor does it constitute an admission of any description in the section as prior art.
[0005] The current popular AI technology has great development space and use scenarios in various industry fields. With the rapid development of the AIGC field, the demand for AI technology in the game art field is also increasing. Different art styles of different projects need to use different AI models to produce materials. However, in the related art, when training a model, a user often needs to use multiple auxiliary tools to complete the entire model training process. This constant switching between various auxiliary tools is not only cumbersome to operate, but also requires a lot of manpower. SUMMARY
[0006] In view of the above, the purpose of the present disclosure is to provide a model training method, device, electronic equipment and storage medium to solve or partially solve the problems in the background art.
[0007] To achieve the above purpose, the present disclosure provides a model training method, which provides a graphical user interface through a terminal device, the graphical user interface being used to display a model training window available for user interaction; the method comprises:
[0008] In response to a confirmation operation of determining a save address of a target image sample set in the model training window, the target image sample set is obtained based on the save address, and the sample images in the target image sample set are batch scaled to obtain a scaled target image sample set;
[0009] In response to a keyword generation operation for the scaled target image sample set in the model training window, a preset keyword recognition model is used to add keyword labels to each sample image in the scaled target image sample set to obtain the target image sample set with added keywords.
[0010] In response to completion of the configuration operation of the training parameters of the target model in the model training window, the target model is trained based on the training parameters and the target image sample set to which the keywords are added.
[0011] Based on the same inventive concept, the exemplary embodiments of the present disclosure also provide a model training device, which provides a graphical user interface for a terminal device, the graphical user interface being used to display a model training window available for user interaction; the device comprises:
[0012] An image scaling module, in response to a confirmation operation of determining a storage address of a target image sample set in the model training window, acquires the target image sample set based on the storage address and performs batch scaling on sample images in the target image sample set to obtain a scaled target image sample set;
[0013] A keyword adding module, in response to a keyword generation operation on the scaled target image sample set in the model training window, adds keyword labels to each sample image in the scaled target image sample set through a preset keyword recognition model to obtain the target image sample set to which the keywords are added;
[0014] A model training module, in response to completion of the configuration operation of the training parameters of the target model in the model training window, trains the target model based on the training parameters and the target image sample set to which the keywords are added.
[0015] Based on the same inventive concept, the exemplary embodiments of the present disclosure also provide an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable by the processor, and the processor implements the model training method as described above when executing the program.
[0016] Based on the same inventive concept, the exemplary embodiments of the present disclosure also provide a non-transitory computer-readable storage medium, which stores computer instructions for causing a computer to execute the model training method as described above.
[0017] Based on the same inventive concept, the exemplary embodiments of the present disclosure also provide a computer program product, which comprises a computer program executable by one or more processors to cause the processors to execute the model training method as described above.
[0018] It can be seen from the above that the model training method, device, electronic equipment and storage medium provided by the present disclosure, in response to the confirmation operation of determining the storage address of the target image sample set in the model training window, obtain the target image sample set based on the storage address, and perform batch scaling on the sample images in the target image sample set to obtain the target image sample set after scaling; in response to the keyword generation operation for the target image sample set after scaling in the model training window, add keyword labels to each sample image in the target image sample set after scaling through a preset keyword recognition model to obtain the target image sample set with added keywords; in response to the configuration operation of the training parameters of the target model being completed in the model training window, train the target model based on the training parameters and the target image sample set with added keywords. Through the completion of all processes of model training in the model training window, one-stop training of the model is realized, the efficiency of model training is improved, the switching between multiple auxiliary works of the user is reduced, and the labor time of the user is reduced. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the present disclosure or the related art, the drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.
[0020] FIG. 1 is a schematic diagram of an application scenario according to one embodiment of the present disclosure;
[0021] FIG. 2 is a flow diagram of a model training method according to one embodiment of the present disclosure;
[0022] FIG. 3 is a schematic diagram of a model training window according to one embodiment of the present disclosure;
[0023] FIG. 4 is a schematic diagram of sample image scaling according to one embodiment of the present disclosure;
[0024] FIG. 5 is a schematic diagram of another sample image scaling according to one embodiment of the present disclosure;
[0025] FIG. 6 is a partial enlarged schematic diagram of a first model training window according to one embodiment of the present disclosure;
[0026] FIG. 7 is a partial enlarged schematic diagram of a second model training window according to one embodiment of the present disclosure;
[0027] FIG. 8 is a partial enlarged schematic diagram of a third model training window according to one embodiment of the present disclosure;
[0028] FIG. 9 is a structural schematic diagram of a model training device according to an embodiment of the present disclosure;
[0029] FIG. 10 is a structural schematic diagram of a specific electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] The principles and spirits of the present disclosure will be described below with reference to a number of exemplary embodiments. It should be understood that the embodiments are given only so that those skilled in the art can better understand and implement the present disclosure, and do not limit the scope of the present disclosure in any way. On the contrary, the embodiments are provided so that the present disclosure is more thorough and complete, and the scope of the present disclosure is fully conveyed to those skilled in the art.
[0031] It should be noted that, unless otherwise defined, technical terms or scientific terms used in the embodiments of the present disclosure should be understood as their common meanings to those skilled in the art. The terms "first", "second", and similar terms used in the embodiments of the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms "include", "contain", and similar terms mean that the elements or objects before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those clearly listed steps or units, but can include other steps or units that are not clearly listed or inherent to the process, method, product, or device. The terms "connect" or "connected" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", and the like are only used to represent relative positional relationships, and when the absolute positions of the described objects change, the relative positional relationships can also change accordingly. In addition, in the description of the present disclosure, the term "a plurality of" means two or more, unless otherwise specified. The term "and / or" describes the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.
[0032] According to the embodiments of the present disclosure, a model training method and device, an electronic device, and a storage medium are provided.
[0033] In this document, it should be understood that any number of elements in the drawings is used for illustration only and is not limiting, and any naming is only for differentiation and does not have any limiting meaning.
[0034] The principles and spirits of the present disclosure will be explained in detail below with reference to several representative embodiments of the present disclosure.
[0035] SUMMARY
[0036] In the related art, when training a model, the user often needs to use multiple auxiliary tools to complete the entire model training process. This constant switching between various auxiliary tools is not only cumbersome to operate, but also requires a lot of manpower. For example, in some related art, when training a model, the inevitable atlas processing requires key picture information to be cut out and cropped to 512 sizes, etc. Here, Photoshop and other art tools are required, and a certain artistic foundation is required. Moreover, in the related art, when adding keywords, manual means are generally required for addition, even with some keyword addition tools, manual checking and modification are still required. Therefore, the model training efficiency in the related art is low. In addition, in the related art, due to the non-reusable limitation of WebUI and other AI training tools, even if the user can master the skills of training the model, but each time of training needs to be transferred between different auxiliary tools several times, which cannot improve the work efficiency, at the same time, the process of training the model in the related art must be strictly in accordance with the fixed flow sequence, which is repetitive and mechanical, and a large amount of time and manpower is wasted.
[0037] To solve the above problems, the present disclosure provides a model training method, specifically comprising:
[0038] In response to a confirmation operation of determining the save address of the target image sample set in the model training window, the target image sample set is obtained based on the save address, and the sample images in the target image sample set are batch scaled to obtain the scaled target image sample set. In response to a keyword generation operation for the scaled target image sample set in the model training window, a preset keyword recognition model is used to add keyword labels to each sample image in the scaled target image sample set to obtain the target image sample set with added keywords. In response to a configuration operation of the training parameters of the target model in the model training window, the target model is trained based on the training parameters and the target image sample set with added keywords. By completing all processes of model training in the model training window, one-stop training of the model is realized, the efficiency of model training is improved, the switching between multiple auxiliary tools is reduced, and the labor time of the user is reduced.
[0039] After introducing the basic principles of the present disclosure, the various non-limiting embodiments of the present disclosure will be specifically introduced below.
[0040] Application scenario overview
[0041] In some specific application scenarios, the model training method of the present disclosure can be applied to various model training systems. As an example, referring to FIG. 1, the application scenario includes at least one server 102 and at least one terminal 101. The terminal device includes, but is not limited to, a desktop computer, a mobile phone, a mobile computer, a tablet computer, a media player, a smart wearable device, a personal digital assistant (PDA), or other electronic devices capable of implementing the above functions, etc. The server can be a standalone physical server, a server cluster composed of multiple physical servers or a distributed system, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (content distribution network), and basic cloud computing services such as big data and artificial intelligence platforms. The server and the terminal can communicate through a network to realize data transmission. The network can be a wired network or a wireless network, which is not limited in the present disclosure.
[0042] The server can be a server providing various services. Specifically, the server can be used to provide background services for applications running on the terminal. Optionally, in some implementations, the model training method provided by the embodiments of the present disclosure can be executed by a terminal device or a server. When executed by the server, in response to a confirmation operation of determining the storage address of the target image sample set in the model training window, the target image sample set is obtained based on the storage address, and the sample images in the target image sample set are batch scaled to obtain the scaled target image sample set; in response to a keyword generation operation for the scaled target image sample set in the model training window, a preset keyword recognition model is used to add keyword labels to each sample image in the scaled target image sample set to obtain the target image sample set with added keywords; and in response to a configuration operation of the training parameters of the target model being completed in the model training window, the target model is trained based on the training parameters and the target image sample set with added keywords. Optionally, the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or as a single software or software module. The embodiments of the present disclosure do not make specific limitations.
[0043] Optionally, the wireless networks or wired networks described above use standard communications technologies and / or protocols. The networks typically carry Internet traffic, but can also include private networks, which are not necessarily limited to the Internet, and can include any combination of local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), mobile, wired or wireless networks, private networks or virtual private networks (VPNs), any combination thereof, etc. In some embodiments, data exchanged over the one or more networks are represented using technologies and / or formats including, but not limited to, hypertext markup language (HTML), extensible markup language (XML), etc. In addition, all or some links can be encrypted using conventional encryption technologies, such as, but not limited to, secure socket layer (SSL), transport layer security (TLS), virtual private networks (VPNs), internet protocol security (IPsec), etc. In other embodiments, technologies other than, or in addition to, those described above can be used.
[0044] The model training method according to the example embodiments of the present disclosure will be described below in connection with specific application scenarios. It should be noted that the above-mentioned application scenarios are only shown for the purpose of facilitating the understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0045] It should be understood that, although each step in the flowchart of FIG. 2 is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in FIG. 2 can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these sub-steps or stages is not necessarily sequential, but can be alternately or alternately executed with other steps or at least part of the sub-steps or stages of other steps.
[0046] Example method
[0047] Referring to FIG. 2, the present embodiment provides a model training method. The execution subject of the method can be, but is not limited to, a server or a terminal device; the method provides a graphical user interface through the terminal device, the graphical user interface being used to display a model training window available for user interaction; the method comprises the following steps:
[0048] S101, in response to a confirmation operation of determining a storage address of a target image sample set in the model training window, acquiring the target image sample set based on the storage address, and performing batch scaling on sample images in the target image sample set to obtain a scaled target image sample set.
[0049] In a specific implementation, in order to facilitate user model training, the model training in the present embodiment is completed in a model training window provided by a terminal device, wherein the model training window is a pre-set operation window of a model training system, and the user can issue various instructions to the model training system through the model training window, thereby realizing model training. Optionally, the model to be trained in the present embodiment can be an image generation model, such as a LoRA model. Referring to FIG. 3, it is a schematic diagram of a model training window in the present disclosure, wherein the model training window comprises an image set processing unit, a keyword label unit, a model training unit and a model output unit. By combining multiple model training modules in the same training window, one-stop training of the model can be realized, and the efficiency of the user in model training is improved. By batch scaling the sample images in the target image sample set through the model training window, the operation of manually processing the image set through auxiliary tools can be reduced. For example, in the related art, unavoidable image set processing requires manual extraction of key picture information and cutting into a pre-set size. Here, not only an art tool such as Photoshop is needed, but also a certain artistic foundation is needed.
[0050] It should be noted that in the present embodiment, the confirmation operation of determining the storage address of the target image sample set can be set as needed, for example, it can be an operation of directly pasting the storage address of the target image sample set at a corresponding position in the model training window, or it can be an operation of confirming the storage address of the target image sample set through file directory searching at a corresponding position in the model training window. Optionally, the user can save the target image sample set to a certain position when performing model training, and record the storage position, so that the subsequent model training system can directly find the corresponding target image sample set through the storage address. Optionally, the target image sample set generally comprises multiple sample images, and these sample images can be images collected by the user through various ways for current model training.
[0051] Referring to FIG. 6, FIG. 6 is a partial enlarged view of a model training window in the present disclosure, and FIG. 6 mainly corresponds to the specific functions of the image set processing unit. The functions that can be implemented by FIG. 6 include but are not limited to adding the address of the image sample set (folder) that needs to be processed, automatically numbering the sample images in the image sample set, setting the cropping size, and the like. For example, in FIG. 6, when the mode control corresponds to increment-image, the system will automatically read the numbers of the sample images in the image sequence, i.e., 001, 002, 003, 004, and the like. The number control is 0 by default, so that the image read each time is from 0+1 (from the first image). Optionally, when there is a need, the user can also set the value corresponding to the number control to other numbers, so as to realize reading from a certain image in the middle. The path control is mainly used to copy and paste the storage address of the target image sample set. The scaling method control is used to determine the scaling method of the image, and the default is neares exact, i.e., the performance and restoration priority of the adjacent mode. The width and height controls are used to set the target size of the sample image after scaling. In general, the target size can be set to the default value, i.e., the width*height is 512*512, and the size 512 is the preferred size of the PC computer for AI training model. The cropping control is used to determine whether the image needs to be cropped and the cropping method, for example, when the cropping control is disabled, it means that cropping is disabled, so that the key information of the picture can be preserved as much as possible. The image name prefix control is used to set the prefix of the image name, and if it is not set, the image name prefix is 0 by default.
[0052] In order to accurately scale the sample images, in some embodiments, the sample images in the target image sample set are batch scaled to obtain the target image sample set after scaling, specifically including:
[0053] Obtaining the target size of the sample image after scaling;
[0054] For each sample image in the target image sample set, obtaining the original size of the sample image, scaling the sample image based on the original size and the target size to obtain the sample image after scaling;
[0055] Based on each sample image after scaling, the target image sample set after scaling is obtained.
[0056] In specific implementation, the target size of the image after scaling can be set by the user in the model training window, and then the system directly obtains the target size from the model training window, for example, the target size in FIG. 6 is 512*512 (unit: pixel). After obtaining the target size of the sample image after scaling, the sample image is scaled according to the original size of each sample image and the above target size to obtain the image after scaling of each sample image.
[0057] In order to preserve as much as possible the original information of the sample image after scaling, in some embodiments, the sample image is scaled based on the original size and the target size to obtain the scaled sample image, specifically comprising:
[0058] determining a scaling ratio based on the original size and the target size;
[0059] in response to the original size being smaller than the target size, inserting target pixel points in the sample image based on the scaling ratio to obtain the scaled sample image; wherein the color value of a target pixel point is determined by the color value of a pixel point adjacent to the target pixel point in the sample image;
[0060] in response to the original size being larger than the target size, determining a plurality of minimum scaling units in the sample image based on the scaling ratio; and determining a reserved pixel point from the minimum scaling units, and generating the scaled sample image based on the reserved pixel point; wherein the minimum scaling unit comprises a plurality of pixel points.
[0061] In specific implementation, in order to avoid key information on the sample image being cropped, in the embodiments of the present disclosure, the scaling of the image is realized by inserting or reducing pixel points in the original sample image. It should be noted that for an AI model, compressing a picture will not affect its recognition of the picture. The AI model will convert the picture we see into pixel points and other information through the decoding of the program, and then the AI model will "draw" in the virtual space. Finally, the decoding function of the AI model will convert the picture "drawn" by the AI model into the form of the picture seen by humans. Referring to FIG. 4, when the original size of the sample image is smaller than the target size, target pixel points need to be inserted in the sample image. The color value of a target pixel point is determined by the color value of a pixel point adjacent to the target pixel point in the sample image, for example, the color value of the target pixel point can be directly equal to the color value of the pixel point adjacent to the target pixel point. Referring to FIG. 5, when the original size of the sample image is larger than the target size, a part of the pixel points in the sample image need to be removed, that is, some reserved pixel points can be selected from the sample image, and then the scaled sample image is generated based on the reserved pixel points. Optionally, when determining the reserved pixel points, the sample image can be divided into a plurality of parts, each part being a minimum scaling unit, each minimum scaling unit comprising a plurality of pixel points, and then the reserved pixel points to be reserved and the excluded pixel points to be deleted are determined from the plurality of pixel points belonging to the same minimum scaling unit. Optionally, the reserved pixel points are randomly determined from the plurality of pixel points belonging to the same minimum scaling unit.
[0062] In some embodiments, the reserved pixels are determined from the minimum scaling unit, specifically including: determining a plurality of the reserved pixels from a plurality of pixel points included in the minimum scaling unit in sequence; wherein the determined reserved pixel each time is a pixel point farthest from the center of the determined reserved pixel among the remaining pixel points.
[0063] In some embodiments, the sample image is scaled based on the original size and the target size to obtain a scaled sample image, specifically including:
[0064] determining a target aspect ratio corresponding to the target size;
[0065] determining a foreground region of the sample image, and cropping the sample image with the center of the foreground region as a cropping center, so that the aspect ratio of the cropped sample image is equal to the target aspect ratio;
[0066] scaling the cropped sample image at a constant ratio to obtain the scaled sample image.
[0067] In specific implementation, it is considered that not all sample images are center-symmetric images, for example, the key information of many images may not be in the center of the image, but may be biased to a certain region. At this time, if the sample image is directly cropped according to the image center, the key information may be lost. Therefore, in the embodiments of the present disclosure, before the sample image is cropped, the foreground region, i.e. the non-background region, of the sample image is determined, and then the sample image is cropped with the center of the foreground region as the cropping center, instead of cropping the sample image with the center of the original sample image as the cropping center.
[0068] In some embodiments, the foreground region of the sample image is determined, specifically including:
[0069] inputting the sample image into the trained foreground recognition model to obtain the foreground region of the sample image.
[0070] In specific implementation, the foreground region of the sample image can be recognized by the trained foreground recognition model. The training of the foreground recognition model can refer to the training process of a neural network model in related technologies, and is not limited. Optionally, in some embodiments, a neural network model with image foreground recognition function in related technologies can also be directly used to recognize the foreground region of the sample image.
[0071] S102, in response to the keyword generation operation for the scaled target image sample set in the model training window, adding keyword labels to each sample image in the scaled target image sample set by a preset keyword recognition model to obtain the target image sample set with added keywords.
[0072] In specific implementation, when the keyword generation operation for the scaled target image sample set is received in the model training window, keyword labels are added to each sample image in the scaled target image sample set by a preset keyword recognition model to obtain the target image sample set with added keywords, thereby reducing the workload of manually adding keywords. Optionally, the specific preset keyword recognition model can be selected from a model with keyword recognition function in related technologies, and no limitation is made on this. For example, in some embodiments, the WD14 and / or Gemini model can be used for keyword recognition, and the recognized keywords are added to the corresponding sample images as keyword labels.
[0073] To accurately recognize keywords, in some embodiments, the preset keyword recognition model is multiple; keyword labels are added to each sample image in the scaled target image sample set by a preset keyword recognition model to obtain the target image sample set with added keywords, and specifically includes:
[0074] Each sample image in the scaled target image sample set is input into each preset keyword recognition model to obtain keyword labels output by each preset keyword recognition model;
[0075] The keyword labels that repeatedly appear in the keyword labels output by each preset keyword recognition model are determined as target keyword labels;
[0076] Based on the target keyword labels, keyword labels are added to each sample image in the scaled target image sample set to obtain the target image sample set with added keywords.
[0077] In specific implementation, considering that the ability of a single model to recognize keywords is limited and is prone to recognition errors, in this embodiment, multiple preset keyword recognition models can be used to recognize the keywords of the same target sample image, that is, the intersection of the keyword sets output by different preset keyword recognition models is used as the target keyword label.
[0078] In some embodiments, the method further includes:
[0079] Obtaining a wake-up word input by a user for the target image sample set from the model training window;
[0080] add the wake-up word to the target image sample set after scaling.
[0081] In implementation, in order to ensure the uniqueness of the wake-up word, in the embodiment, the user can manually input the wake-up word in the model training window, and then add the wake-up word to the target image sample set. Referring to FIG. 7, the user can input the wake-up word in the input position of the wake-up word module in the lower right corner, for example, in FIG. 7, the current wake-up word is "jianzhi".
[0082] S103, in response to completing the configuration operation of the training parameters of the target model in the model training window, training the target model based on the training parameters and the target image sample set to which the keyword is added.
[0083] In implementation, after obtaining the target image sample set to which the keyword is added, the sample data that can be used for model training is obtained, and then in response to the user completing the configuration operation of the training parameters of the target model in the model training window, the target model is trained according to the training parameters configured above and the target image sample set to which the keyword is added.
[0084] In some embodiments, the training parameters include the number of images loaded at the same time, the number of learning times of each sample image, the number of training groups corresponding to the target image sample set, and the saving epoch of the intermediate training result. The number of images loaded at the same time is used to determine the number of images trained at the same time in the model training process, for example, when the number of images loaded at the same time is 2, that is, 2 images are trained at the same time. It should be noted that although learning multiple different images at the same time can reduce the adjustment accuracy of each image, it can comprehensively capture the features of multiple images, and therefore the final effect can be better. The number of learning times of each sample image and the number of training groups corresponding to the target image sample set are both used to control the number of times of using the sample image. For example, when the number of sample images in a certain target image sample set is 10, the number of learning times of the corresponding sample image is 5, and the number of training groups corresponding to the target image sample set is 1, the number of times of model training is: 10*5*1=50 times.
[0085] It should be noted that in a neural network, an epoch refers to all training samples in the model being forward propagated and backward propagated once, that is, the process of training all training samples once. In each round of update, the number of network updates can be arbitrary, but it is usually set to traverse the dataset. Epochs is an important hyperparameter in the training process of neural networks, defined as a single training iteration of all batches in forward and backward propagation. In simple terms, an epoch is to input all data into the network to complete a forward calculation and backpropagation. The main role of setting the epoch is to divide the entire training process of the model into several segments, which can better observe and adjust the training of the model.
[0086] Referring to FIG. 8, it is a model training parameter setting interface in an embodiment of the present disclosure, wherein the ckpt_name control is used to determine the large model that the user wants to train the style of, the data_path control is used to determine the upper address of the processed atlas, the batch_size control is used to determine the number of images loaded at the same time, the max_train_epoches control is used to determine the maximum number of training rounds, the save_every_n_epochs control is used to determine saving once every n epochs. The output_name control is used to determine the name of the saved target model. The clip_skip control is used to determine the skip step, wherein 1 represents realism and 2 represents cartoon. The output_dir control is used to determine the saving location of the target model.
[0087] In some embodiments, after training the target model based on the training parameters and the target image sample set with added keywords, the method further comprises:
[0088] in response to receiving a picture generation instruction in the model training window; wherein the model training window connects a plurality of training completed target models;
[0089] determining the target model to be used from the plurality of training completed target models based on the picture generation instruction;
[0090] executing the picture generation instruction based on the target model to be used.
[0091] In the embodiments of the present disclosure, the user can also directly generate pictures in the model training window, that is, when a picture generation instruction is received in the model training window, the target model to be used is determined based on the picture generation instruction, and the picture generation instruction can include the name of the target model to be used. After determining the target model to be used, the picture generation is performed according to the target model to be used, that is, the picture generation instruction is completed. It should be noted that the target model to be used can be one or more, and when the target model to be used is multiple, the picture generation instruction is completed by multiple target models.
[0092] The model training method provided by the present disclosure, in response to determining the confirmation operation of the save address of the target image sample set in the model training window, obtains the target image sample set based on the save address, and performs batch scaling on the sample images in the target image sample set to obtain the scaled target image sample set; in response to the keyword generation operation of the scaled target image sample set in the model training window, a preset keyword recognition model is used to add keyword labels to each sample image in the scaled target image sample set to obtain the target image sample set with added keywords; in response to the configuration operation of the training parameters of the target model being completed in the model training window, the target model is trained based on the training parameters and the target image sample set with added keywords. Through all the processes of model training in the model training window, one-stop training of the model is realized, the efficiency of model training is improved, the switching between multiple auxiliary works of the user is reduced, and the labor time of the user is reduced.
[0093] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of the present embodiment can also be applied to a distributed scenario, and can be completed by multiple devices cooperating with each other. In this distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiments of the present disclosure, and the multiple devices can interact with each other to complete the method.
[0094] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order described above and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.
[0095] Exemplary device
[0096] Corresponding to any of the above-mentioned embodiment methods based on the same inventive concept, the present disclosure also provides a model training apparatus. The apparatus provides a graphical user interface for displaying a model training window available for user interaction through a terminal device.
[0097] Referring to FIG. 9, the model training apparatus comprises:
[0098] an image scaling module 201 configured to, in response to a confirmation operation of determining a save address of a target image sample set in the model training window, acquire the target image sample set based on the save address and perform batch scaling on sample images in the target image sample set to obtain a scaled target image sample set;
[0099] a keyword adding module 202 configured to, in response to a keyword generation operation for the scaled target image sample set in the model training window, add keyword labels to each sample image in the scaled target image sample set through a preset keyword recognition model to obtain the target image sample set with added keywords;
[0100] a model training module 203 configured to, in response to a configuration operation of training parameters of a target model in the model training window, train the target model based on the training parameters and the target image sample set with added keywords.
[0101] In some embodiments, the image scaling module specifically comprises:
[0102] an acquisition unit configured to acquire a target size after scaling of a sample image;
[0103] a scaling unit configured to, for each sample image in the target image sample set, acquire an original size of the sample image, perform scaling on the sample image based on the original size and the target size to obtain a scaled sample image;
[0104] an integration unit configured to obtain the target image sample set after scaling based on each scaled sample image.
[0105] In some embodiments, the scaling unit is specifically configured to perform:
[0106] determine a scaling ratio based on the original size and the target size;
[0107] in response to the original size being smaller than the target size, inserting target pixel points in the sample image based on the scaling ratio, to obtain a scaled sample image; wherein a color value of a target pixel point is determined based on color values of pixel points adjacent to the target pixel point in the sample image.
[0108] in response to the original size being greater than the target size, determining a plurality of minimum scaling units in the sample image based on the scaling ratio; and determining a reserved pixel point from the minimum scaling units, and generating the scaled sample image based on the reserved pixel point; wherein the minimum scaling unit includes a plurality of pixel points.
[0109] In some embodiments, the scaling unit is specifically configured to perform:
[0110] determining a plurality of reserved pixel points from the plurality of pixel points included in the minimum scaling unit in sequence; wherein each determined reserved pixel point is a pixel point farthest from the center of the determined reserved pixel point among the remaining plurality of pixel points.
[0111] In some embodiments, the scaling unit is specifically configured to perform:
[0112] determining a target aspect ratio corresponding to the target size;
[0113] determining a foreground region of the sample image, and cropping the sample image with the center of the foreground region as a cropping center, so that the aspect ratio of the cropped sample image is equal to the target aspect ratio;
[0114] scaling the cropped sample image at a constant ratio to obtain the scaled sample image.
[0115] In some embodiments, the scaling unit is specifically configured to perform:
[0116] inputting the sample image into a trained foreground recognition model to obtain a foreground region of the sample image.
[0117] In some embodiments, the preset keyword recognition model is a plurality of; the keyword adding module is specifically configured to perform:
[0118] inputting each sample image in the scaled target image sample set into each preset keyword recognition model respectively, to obtain keyword labels output by each preset keyword recognition model;
[0119] determining a keyword label that repeatedly appears in the keyword labels output by each preset keyword recognition model as a target keyword label;
[0120] add a keyword label to each sample image in the target image sample set after zooming based on the target keyword label, to obtain the target image sample set with added keywords.
[0121] In some embodiments, the apparatus further comprises a wake word module configured to perform:
[0122] obtaining a wake word input by a user for the target image sample set from the model training window;
[0123] adding the wake word to the target image sample set after zooming.
[0124] In some embodiments, the training parameters include the number of images loaded at the same time, the number of learning times for each sample image, the number of training groups corresponding to the target image sample set, and the saving rounds of intermediate training results.
[0125] In some embodiments, the apparatus further comprises an image generation module configured to perform:
[0126] in response to receiving a picture generation instruction in the model training window; wherein the model training window is connected to a plurality of trained target models;
[0127] determining the target model to be used from the plurality of trained target models based on the picture generation instruction;
[0128] executing the picture generation instruction based on the target model to be used.
[0129] For the convenience of description, the above apparatus is described in various modules in terms of functions. Of course, the functions of each module can be implemented in one or more software and / or hardware when implementing the present disclosure.
[0130] The model training apparatus of the above embodiments is used to implement the corresponding model training method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.
[0131] Based on the same inventive concept, corresponding to any of the above embodiment methods, the present disclosure also provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the model training method of any of the above embodiments.
[0132] Fig. 10 shows a more specific schematic diagram of the hardware structure of an electronic device according to the embodiment, which can include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040 and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030 and the communication interface 1040 are connected to each other through the bus 1050 for internal communication.
[0133] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, etc., for executing related programs to implement the technical solutions provided by the embodiments of the present specification.
[0134] The memory 1020 can be implemented by a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, the related program codes are stored in the memory 1020 and called and executed by the processor 1010.
[0135] The input / output interface 1030 is configured to connect to an input / output module to realize information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0136] The communication interface 1040 is configured to connect to a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0137] The bus 1050 includes a channel to transmit information between various components (such as the processor 1010, the memory 1020, the input / output interface 1030 and the communication interface 1040) of the device.
[0138] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only contain the components necessary to implement the embodiments of the present application, and does not necessarily contain all the components shown in the figure.
[0139] The electronic device of the above embodiment is used to implement the corresponding model training method in any of the preceding embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.
[0140] The memory 1020 stores machine readable instructions executable by the processor 1010, and when the electronic device is running, the processor 1010 and the memory 1020 communicate through the bus 1050, so that the processor 1010 executes the following instructions when running: in response to the confirmation operation of determining the save address of the target image sample set in the model training window, obtaining the target image sample set based on the save address, and performing batch scaling on the sample images in the target image sample set to obtain the target image sample set after scaling; in response to the keyword generation operation for the target image sample set after scaling in the model training window, adding keyword labels to each sample image in the target image sample set after scaling through a preset keyword recognition model to obtain the target image sample set with added keywords; and in response to the configuration operation of the training parameters of the target model in the model training window, training the target model based on the training parameters and the target image sample set with added keywords.
[0141] In a possible implementation, the instructions executed by the processor 1010 include performing batch scaling on the sample images in the target image sample set to obtain the target image sample set after scaling, specifically including:
[0142] Obtaining the target size of the sample image after scaling;
[0143] For each sample image in the target image sample set, obtaining the original size of the sample image, scaling the sample image based on the original size and the target size to obtain the sample image after scaling;
[0144] Obtaining the target image sample set after scaling based on each sample image after scaling.
[0145] In a possible implementation, the instructions executed by the processor 1010 include scaling the sample image based on the original size and the target size to obtain the scaled sample image, and the scaling specifically includes:
[0146] determining a scaling ratio based on the original size and the target size;
[0147] in response to the original size being smaller than the target size, inserting a target pixel point in the sample image based on the scaling ratio to obtain the scaled sample image; wherein a color value of the target pixel point is determined based on color values of pixel points adjacent to the target pixel point in the sample image;
[0148] in response to the original size being larger than the target size, determining a plurality of minimum scaling units in the sample image based on the scaling ratio; and determining a reserved pixel point from the minimum scaling units, and generating the scaled sample image based on the reserved pixel point; wherein the minimum scaling unit includes a plurality of pixel points.
[0149] In a possible implementation, the instructions executed by the processor 1010 include determining the reserved pixel point from the minimum scaling unit, and the determination specifically includes:
[0150] determining the reserved pixel point from the plurality of pixel points included in the minimum scaling unit in sequence; wherein the reserved pixel point determined each time is a pixel point farthest from the center of the reserved pixel point already determined among the remaining plurality of pixel points.
[0151] In a possible implementation, the instructions executed by the processor 1010 include scaling the sample image based on the original size and the target size to obtain the scaled sample image, and the scaling specifically includes:
[0152] determining a target aspect ratio corresponding to the target size;
[0153] determining a foreground region of the sample image, and cropping the sample image with a center of the foreground region as a cropping center, so that an aspect ratio of the cropped sample image is equal to the target aspect ratio;
[0154] proportionally scaling the cropped sample image to obtain the scaled sample image.
[0155] In a possible implementation, the instructions executed by the processor 1010 include determining the foreground region of the sample image, and the determination specifically includes:
[0156] inputting the sample image into a foreground recognition model trained to obtain the foreground region of the sample image.
[0157] In a possible implementation, in the instructions executed by the processor 1010, the preset keyword recognition model is multiple; keyword labels are added to each sample image in the scaled target image sample set through the preset keyword recognition model, to obtain the target image sample set with added keywords, and the method specifically includes:
[0158] Each sample image in the scaled target image sample set is input into each preset keyword recognition model, to obtain keyword labels output by each preset keyword recognition model;
[0159] A keyword label that repeatedly appears in the keyword labels output by each preset keyword recognition model is determined as a target keyword label;
[0160] Based on the target keyword label, a keyword label is added to each sample image in the scaled target image sample set, to obtain the target image sample set with added keywords.
[0161] In a possible implementation, in the instructions executed by the processor 1010, the method further includes:
[0162] An awakening word input by a user for the target image sample set is obtained from the model training window;
[0163] The awakening word is added to the scaled target image sample set.
[0164] In a possible implementation, in the instructions executed by the processor 1010, the training parameters include a number of images loaded at the same time, a learning number of each sample image, a number of training groups corresponding to the target image sample set, and a saving round of an intermediate training result.
[0165] In a possible implementation, in the instructions executed by the processor 1010, after the target model is trained based on the training parameters and the target image sample set with added keywords, the method further includes:
[0166] In response to receiving a picture generation instruction in the model training window; wherein the model training window is connected to multiple trained target models;
[0167] Based on the picture generation instruction, a target model to be used is determined from the multiple trained target models;
[0168] The picture generation instruction is executed based on the target model to be used.
[0169] In the above manner, when the electronic device is running, in response to a confirmation operation of determining a storage address of a target image sample set in the model training window, the target image sample set is acquired based on the storage address, and batch scaling is performed on sample images in the target image sample set, to obtain the target image sample set after scaling; in response to a keyword generation operation for the target image sample set after scaling in the model training window, a preset keyword recognition model is used to add keyword labels to each sample image in the target image sample set after scaling, to obtain the target image sample set with added keywords; and in response to a configuration operation of training parameters of a target model being completed in the model training window, the target model is trained based on the training parameters and the target image sample set with added keywords. Through the above process, one-stop training of the model is achieved, the efficiency of model training is improved, the switching between multiple auxiliary operations of the user is reduced, and the labor time of the user is reduced.
[0170] Exemplary program product
[0171] Corresponding to any of the above-mentioned embodiment methods based on the same inventive concept, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the model training method of any of the above embodiments.
[0172] The computer-readable medium of the present embodiment includes permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0173] The computer instructions stored in the storage medium of the above-mentioned embodiments are used to cause the computer to perform the model training method of any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which are not repeated here.
[0174] Based on the same inventive concept, the disclosure also provides a computer program product corresponding to any of the above-mentioned embodiment methods, which comprises a computer program. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to perform the model training method described in the above embodiments. The processor performing the corresponding steps can belong to the corresponding execution subject corresponding to each step in each embodiment of the scene editing method.
[0175] The computer program product of the above-mentioned embodiments is used to make the computer and / or the processor perform the model training method as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not repeated here.
[0176] It can be understood that before using the technical solutions of various embodiments of the disclosure, the type, use range, use scenario, etc. of the personal information involved will be informed to the user in a proper manner, and the authorization of the user will be obtained.
[0177] For example, in response to receiving the user's active request, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic devices, application programs, servers or storage media that perform the technical solutions of the disclosure according to the prompt information.
[0178] As an optional but not limited implementation manner, in response to accepting the user's active request, the way of sending prompt information to the user may, for example, be a pop-up window manner, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0179] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation of the disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the disclosure.
[0180] Those skilled in the art should understand that the embodiments of the disclosure can be implemented as a system, a method or a computer program product. Therefore, the disclosure can be embodied in the form of entire hardware, entire software (including firmware, resident software, microcode, etc.), or hardware and software combined, which are generally referred to as "circuitry", "module" or "system" herein. In addition, in some embodiments, the disclosure can also be implemented as a computer program product in one or more computer readable media, which contains computer readable program code.
[0181] Any combination of one or more computer readable medium can be utilized. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0182] A computer readable signal medium can include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0183] Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0184] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In an embodiment, multiple data storage devices can be used to store data. The data storage devices can be located on the same computer or on different computers.
[0185] It should be understood that each block of the flowchart and / or block diagram illustrations, and combinations of blocks in the flowchart and / or block diagram illustrations, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0186] These computer program instructions can also be stored in a computer- readable medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions which implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0187] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0188] The flowchart and block diagram in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0189] It should be noted that, although the foregoing detailed description has set forth several modules or units for the device for action execution, such division into modules or units is not mandatory. Rather, two or more modules or units described above can be embodied in a single module or unit according to the embodiments of the present disclosure. Conversely, a single module or unit described above can be divided into multiple modules or units according to the embodiments of the present disclosure.
[0190] Those skilled in the art should understand that the above discussion of any embodiment is merely exemplary and is not intended to be limiting of the scope of the disclosure including the claims; combinations of the above embodiments or different embodiments, or steps, can be made in the light of the above teachings, and the steps can be performed in any order, and there are many other variations of the disclosed embodiments in light of the above teachings, which are intended to be encompassed by the appended claims.
[0191] In addition, to simplify the description and discussion, and so as not to make the embodiments of the disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components can or can not be shown in the provided drawings. Furthermore, devices can be shown in block diagram form in order to avoid making the embodiments of the disclosure difficult to understand, and this also takes into account the fact that details regarding implementation of these block diagram devices are highly dependent on the platform in which the embodiments of the disclosure are to be implemented (i.e., these details should be well within the understanding of one of skill in the art). Where specific details (e.g., circuitry) are set forth in order to describe an illustrative embodiment of the disclosure, it should be apparent to those skilled in the art that the embodiments of the disclosure can be practiced without or with variation of these specific details. Thus, the description should not be viewed as limiting, but rather as merely descriptive.
[0192] While the disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations will be apparent to those skilled in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.
[0193] Embodiments of the disclosure are intended to cover all such alternatives, modifications, and variations as falling within the broad scope of the appended claims. Accordingly, any one of the steps of the disclosed embodiments can be performed in any order, and the steps can be performed in any order, and there are many other variations of the disclosed embodiments in light of the above teachings, which are intended to be encompassed by the appended claims.
Claims
1. A model training method, a graphical user interface is provided by a terminal device, the graphical user interface is used to display a model training window available for user interaction; the method comprises: in response to a confirmation operation of determining a save address of a target image sample set in the model training window, obtaining the target image sample set based on the save address, and performing batch scaling on sample images in the target image sample set to obtain a scaled target image sample set; in response to a keyword generation operation for the scaled target image sample set in the model training window, adding keyword labels to each sample image in the scaled target image sample set by a preset keyword recognition model to obtain the target image sample set with added keywords; in response to a configuration operation of training parameters of a target model in the model training window, training the target model based on the training parameters and the target image sample set with added keywords.
2. The method of claim 1, wherein, performing batch scaling on sample images in the target image sample set to obtain a scaled target image sample set, specifically comprising: obtaining a target size of the scaled sample image; for each sample image in the target image sample set, obtaining the original size of the sample image, scaling the sample image based on the original size and the target size to obtain the scaled sample image; obtaining the scaled target image sample set based on each scaled sample image.
3. The method of claim 1, wherein, scaling the sample image based on the original size and the target size to obtain the scaled sample image, specifically comprising: determining a scaling ratio based on the original size and the target size; in response to the original size being smaller than the target size, inserting target pixel points in the sample image based on the scaling ratio to obtain the scaled sample image; wherein the color value of the target pixel point is determined by the color value of the pixel points adjacent to the target pixel point in the sample image; in response to the original size being greater than the target size, determining a plurality of minimum scaling units in the sample image based on the scaling ratio; and determining a reserved pixel point from the minimum scaling unit, generating the scaled sample image based on the reserved pixel point; wherein the minimum scaling unit comprises a plurality of pixel points.
4. The method of claim 1, wherein, determining a reserved pixel point from the minimum scaling unit, specifically comprising: determining a plurality of reserved pixel points from the plurality of pixel points included in the minimum scaling unit in turn; wherein each determined reserved pixel point is the pixel point farthest from the center of the determined reserved pixel point among the remaining plurality of pixel points.
5. The method of claim 1, wherein, scaling the sample image based on the original size and the target size to obtain the scaled sample image, specifically comprising: determining a target aspect ratio corresponding to the target size; determining a foreground region of the sample image, and cropping the sample image with the center of the foreground region as the cropping center to make the aspect ratio of the cropped sample image equal to the target aspect ratio; The cropped sample image is scaled at a same ratio to obtain a scaled sample image.
6. The method of claim 1, wherein, The foreground region of the sample image is determined, specifically including: The sample image is input into the trained foreground recognition model to obtain the foreground region of the sample image.
7. The method of claim 1, wherein, The preset keyword recognition model is multiple; each sample image in the scaled target image sample set is added with a keyword label through a preset keyword recognition model to obtain the target image sample set with added keywords, specifically including: Each sample image in the scaled target image sample set is input into each preset keyword recognition model to obtain the keyword label output by each preset keyword recognition model; The keyword label that repeatedly appears in the keyword label output by each preset keyword recognition model is determined as a target keyword label; Based on the target keyword label, each sample image in the scaled target image sample set is added with a keyword label to obtain the target image sample set with added keywords.
8. The method of claim 1, wherein, The method further includes: An awakening keyword input by a user for the target image sample set is obtained from the model training window; The awakening keyword is added to the scaled target image sample set.
9. The method of claim 1, wherein, The training parameters include the number of images loaded at the same time, the number of learning times of each sample image, the number of training groups corresponding to the target image sample set, and the saving rounds of intermediate training results.
10. The method of claim 1, wherein, After training the target model based on the training parameters and the target image sample set with added keywords, the method further includes: In response to receiving a picture generation instruction in the model training window; wherein the model training window is connected to multiple trained target models; Based on the picture generation instruction, the target model to be used is determined from the multiple trained target models; The picture generation instruction is executed based on the target model to be used.
11. A model training device, which provides a graphical user interface for a terminal device, the graphical user interface being used to display a model training window available for user interaction; the device includes: An image scaling module, in response to a confirmation operation of determining a saving address of a target image sample set in the model training window, the target image sample set is obtained based on the saving address, and the sample images in the target image sample set are scaled in batches to obtain a scaled target image sample set; A keyword adding module, in response to a keyword generation operation for the scaled target image sample set in the model training window, each sample image in the scaled target image sample set is added with a keyword label through a preset keyword recognition model to obtain the target image sample set with added keywords; A model training module, in response to a configuration operation of training parameters of a target model in the model training window, the target model is trained based on the training parameters and the target image sample set with added keywords.
12. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable by the processor, the processor implementing the method of any one of claims 1 to 10 when executing the program.
13. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method of any one of claims 1 to 10.
Citation Information
Patent Citations
High-speed railway contact network image recognition method combining YOLOv3 and SENet
CN111582334A
Deep learning guide device and method
CN111931944A
Training method of image generation model and image generation method and device
CN115631261A
Image generation model training method and related device
CN116975347A
Image generation method, device, equipment, medium and product
CN118365987A