Data generation method and apparatus, and electronic device and computer-readable medium
By deploying the target plug-in on the client to process data generation requests, the problems of network latency, security, stability and cost in the data generation process in the prior art are solved, and more efficient and reliable data generation results are achieved.
Patent Information
- Application Number
- PCT/CN2024/133837
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-05
AI Technical Summary
In the process of data generation such as images, videos, texts, etc., the prior art has problems with network delay, data security, stability and reliability, and the content generation cost is relatively high.
By deploying the target plug-in on the client, using the plug-in to process data generation requests, local data generation and processing are implemented, avoiding the defects caused by collaborative processing with the server. The target plug-in is built based on the data generation model and has the function of processing data generation requests.
Improve data generation effect, enhance real-time, security and reliability, reduce content generation costs, and avoid delays caused by network factors.
Smart Images

Figure CN2024133837_05062025_PF_FP_ABST
Abstract
Description
Data generation method, device, electronic device, and computer-readable medium
[0001] This application claims priority to Chinese patent application No. 202311598782.9 filed on November 27, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby cited in their entirety as a part of this application. Technical Field
[0002] The present disclosure relates to a data generation method, device, electronic device, and computer-readable medium. Background Art
[0003] In some application scenarios, for data such as images, videos, and text, this data can be generated through collaboration between the client and the server. For ease of understanding, the image generation process in the context of text-to-image is used as an example.
[0004] As an example, the image generation process in the text image scenario can be specifically as follows: after the client detects a text image request triggered by the user, the client forwards the text image request to the server, so that the server can use the model with text image function that has been deployed in the server to process the request, obtain the generated image, and the server feeds the generated image back to the client for display. Summary of the Invention
[0005] The present disclosure provides a data generation method, device, electronic device, and computer-readable medium, which are conducive to improving data generation effects.
[0006] In order to achieve the above objectives, the technical solutions provided by the present disclosure are as follows:
[0007] The present disclosure provides a data generation method, which is applied to a client and includes:
[0008] receiving a data generation request;
[0009] The data generation request is processed using a target plug-in to obtain generated data; the target plug-in is constructed based on a data generation model corresponding to the data generation request, and the data generation model is used to process the data generation request.
[0010] In one possible implementation, the data generation request is used to request content generation processing based on reference information; the reference information includes at least one of text and image; the data generation model is a content generation model; and the generated data includes at least one generated image.
[0011] In a possible implementation manner, the client is deployed on a terminal device;
[0012] The step of processing the data generation request using the target plug-in to obtain generated data includes:
[0013] Sending the data generation request to the target plug-in deployed on the terminal device through the plug-in service deployed on the client;
[0014] The generated data fed back by the target plug-in in response to the data generation request is received.
[0015] In a possible implementation manner, the target plug-in includes a target model determined according to the data generation model; the target plug-in is configured to process the data generation request using the target model to obtain the generated data.
[0016] In one possible implementation, the target plug-in includes at least two candidate models, different candidate models implement different data processing processes, and the at least two candidate models include the target model; the target plug-in is also used to determine the target model corresponding to the data generation request from the at least two candidate models.
[0017] In a possible implementation manner, before processing the data generation request using the target plug-in, the method further includes:
[0018] In response to a client startup request, the client is started and the target plug-in is loaded.
[0019] In a possible implementation manner, after loading the target plug-in, the method further includes:
[0020] Initialize the loaded target plug-in;
[0021] The processing of the data generation request by using the target plug-in includes:
[0022] The data generation request is processed using the initialized target plug-in.
[0023] In a possible implementation manner, loading the target plug-in includes: asynchronously preloading the target plug-in.
[0024] In one possible implementation, the target plug-in is constructed based on at least one data generation model, different data generation models implement different data generation processes, and the at least one data generation model includes a data generation model corresponding to the data generation request.
[0025] In one possible implementation, the target plug-in includes a base model and at least one fine-tuning model; the base model and the at least one fine-tuning model are determined based on the at least one data generation model;
[0026] Before processing the data generation request using the target plug-in, the method further includes:
[0027] In response to a client startup request, starting the client and loading the basic model in the target plug-in;
[0028] The process of obtaining the generated data includes:
[0029] Determining a fine-tuning model corresponding to the data generation request from the at least one fine-tuning model;
[0030] Using the fine-tuning model to perform parameter correction processing on the loaded basic model to obtain the target model;
[0031] The data generation request is processed using the target model.
[0032] In a possible implementation manner, after receiving the data generation request, the method further includes:
[0033] Determine whether the target plug-in is in an available state;
[0034] The processing of the data generation request by using the target plug-in includes:
[0035] If it is determined that the target plug-in is in an available state, the data generation request is processed using the target plug-in.
[0036] In a possible implementation manner, after determining whether the target plug-in is in an available state, the method further includes:
[0037] If it is determined that the target plug-in is in an unavailable state, sending the data generation request to the server, and the server is configured to process the data generation request using the data generation model deployed on the server to obtain the generated data;
[0038] Receive the generated data fed back by the server.
[0039] In a possible implementation manner, the client is deployed on a terminal device;
[0040] After determining whether the target plug-in is in an available state, the method further includes:
[0041] If it is determined that the target plug-in is in an unavailable state, determining whether a plug-in description resource of the target plug-in is stored in the terminal device;
[0042] If it is determined that the plug-in description resource of the target plug-in is not stored in the terminal device, a resource demand request is sent to the server;
[0043] The plug-in description resource fed back by the server in response to the resource demand request is received and stored; the plug-in description resource is determined by the server according to the data generation model deployed on the server.
[0044] In a possible implementation manner, the target plug-in includes at least two data processing modules arranged in sequence, and the at least two data processing modules are obtained by splitting the data generation model;
[0045] The process of obtaining the generated data includes:
[0046] Load the i-th data processing module; i is a positive integer, and its initial value is 1;
[0047] Performing data processing using the i-th data processing module;
[0048] After obtaining the output result of the i-th data processing module, releasing the memory occupied by the i-th data processing module;
[0049] Update the i, and continue to execute the step of loading the i-th data processing module until a preset stop condition is reached, and determine the generated data according to the output result of the i-th data processing module.
[0050] In one possible implementation, the splitting process of the data generation model includes:
[0051] Splitting the data generation model into at least two sub-models, wherein different sub-models implement different data processing functions;
[0052] For any of the sub-models, if the memory requirement characterization data of the sub-model does not exceed the preset memory threshold, the sub-model is determined as the data processing module; if the memory requirement characterization data of the sub-model exceeds the preset memory threshold, the sub-model is split into at least two model fragments, and each of the model fragments is determined as the data processing module, and the memory requirement characterization data of each of the model fragments does not exceed the preset memory threshold.
[0053] In one possible implementation, the process of obtaining the at least two model fragments includes:
[0054] Obtain at least one candidate split description information, where different candidate split description information describe different model splitting methods;
[0055] For any of the candidate split description information, split the sub-model according to the candidate split description information to obtain a model split result corresponding to the candidate split description information, and determine the resource usage characterization data corresponding to the candidate split description information based on the model split result corresponding to the candidate split description information, wherein the resource usage characterization data includes computing resource reuse characterization data and memory usage characterization data;
[0056] selecting target split description information from the at least one candidate split description information based on the resource usage characterization data corresponding to each of the candidate split description information, wherein the resource usage balance degree presented by the resource usage characterization data corresponding to the target split description information is higher than the resource usage balance degree presented by the resource usage characterization data corresponding to any other candidate split description information in the at least one candidate split description information except the target split description information;
[0057] The at least two model segments are determined according to the model splitting result corresponding to the target splitting description information.
[0058] The present disclosure provides a data generation device, comprising:
[0059] A receiving unit, configured to receive a data generation request;
[0060] A processing unit is used to use a target plug-in to process the data generation request to obtain generated data; the target plug-in is constructed based on a data generation model corresponding to the data generation request, and the data generation model is used to process the data generation request.
[0061] The present disclosure provides an electronic device, the device comprising: a processor and a memory;
[0062] The memory is used to store instructions or computer programs;
[0063] The processor is configured to execute the instructions or computer programs in the memory so that the electronic device executes the data generation method provided by the present disclosure.
[0064] The present disclosure provides a computer-readable medium, characterized in that instructions or computer programs are stored in the computer-readable medium, and when the instructions or computer programs are executed on a device, the device executes the data generation method provided by the present disclosure.
[0065] The present disclosure provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, wherein the computer program contains program code for executing the data generation method provided by the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0067] FIG1 is a flow chart of a data generation method provided by an embodiment of the present disclosure;
[0068] FIG2 is a schematic diagram of a content generation process provided by an embodiment of the present disclosure;
[0069] FIG3 is a schematic diagram of a content generation process provided by an embodiment of the present disclosure;
[0070] FIG4 is a schematic diagram of a model splitting method provided by an embodiment of the present disclosure;
[0071] FIG5 is a schematic diagram of another model splitting method provided by an embodiment of the present disclosure;
[0072] FIG6 is a schematic diagram of an implementation of a multiple data generation process provided by an embodiment of the present disclosure;
[0073] FIG7 is a schematic structural diagram of a data generating device provided by an embodiment of the present disclosure;
[0074] FIG8 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0075] Research has found that for some application scenarios, such as content generation scenarios similar to text-generated images, image-generated images, text-generated videos, and image-generated videos, the content generation process in these scenarios can be implemented with the help of a collaborative approach between the client and the server; and the implementation process can be specifically as follows: after the client detects a content generation request triggered by the user, such as a text-generated image request, the client forwards the content generation request to the server, so that the server can use the model with content generation function that has been deployed on the server to process the request and obtain generated content, and the server feeds back the generated content to the client so that the client can display the generated content.
[0076] Research has also found that the content generation process achieved through collaboration between the client and the server as shown in the previous paragraph has the defects shown in ①-④ below.
[0077] ① Defects caused by network delays, specifically: because the client needs to forward the above-mentioned content generation request to the server through the network, when network delays occur, it will affect the response time of the content generation request, thereby affecting the user experience.
[0078] ② The defects in data security are as follows: because the client needs to forward the above-mentioned content generation request to the server for processing, some problems affecting data security may occur in the forwarding process and processing process of the content generation request, thereby affecting the user experience.
[0079] ③ Stability and reliability defects: Because the client and server are connected via a network, the stability and reliability of that network may affect the response process of the above-mentioned content generation request. For example, if the network fails or the network bandwidth is limited, it may cause the response time of the content generation request to increase, or even make the response to the content generation request impossible to complete.
[0080] ④ Defects in content generation costs, specifically: the use of computing resources and storage resources on the server may require additional costs, resulting in a relatively high cost for the content generation process achieved through collaboration between the client and the server.
[0081] Based on the above research, it can be known that in order to better improve the data generation experience, such as the content generation experience, the present disclosure provides a data generation method, which includes: for a client, after the client receives a data generation request, the client uses a target plug-in to process the data generation request to obtain generated data, so that data generation processing can be achieved on the client based on the plug-in, thereby effectively overcoming the defects existing when the data generation processing is achieved by the client and the server in a collaborative manner, thereby helping to improve the data generation effect. Among them, because the target plug-in is constructed based on the data generation model corresponding to the data generation request, the data processing function of the target plug-in includes the data processing function of the data generation model; and because the data generation model can be used to process the data generation request, the target plug-in can also be used to process the data generation request, so that after the client receives the data generation request, the client can directly use the target plug-in to process the data generation request, thereby making all related processes of the data generation request occur at the same end, so that the data generation processing can be single-ended closed-loop, thereby effectively improving the data generation process in terms of real-time, security, reliability, etc.
[0082] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0083] To better understand the technical solutions provided by the present disclosure, the data generation method provided by the present disclosure is described below with reference to some figures. As shown in Figure 1, the data generation method provided by the embodiment of the present disclosure includes the following S1-S2. Figure 1 is a flow chart of a data generation method provided by the embodiment of the present disclosure.
[0084] S1: The client receives a data generation request.
[0085] The client is used to provide some services to users, such as data generation services similar to text maps, etc. For example, the client can be implemented using the client shown in Figure 2 or Figure 3.
[0086] It should be noted that the present disclosure does not limit the implementation methods of the data generation service in the above paragraph. For example, in some application scenarios, the data generation service may refer to a content generation service. Moreover, the present disclosure does not limit the content generation service. For example, the content generation service may include a text-to-image service, a picture-to-image service, a text-to-video service, a picture-to-video service, etc.
[0087] In addition, the present disclosure does not limit the implementation method of the above client. For example, the client can be implemented using an existing or future user terminal that can interact with the user, such as an application (app) or a web page (web).
[0088] In addition, the client described above can be deployed on a terminal device so that the user of the terminal device can use the client to meet some of the user's needs, such as data generation needs such as text-to-image, image-to-image, text-to-video, and image-to-video. It should be noted that the present disclosure does not limit the implementation of the terminal device. For example, the terminal device can be a smartphone, a computer, a personal digital assistant (PDA), a tablet computer, etc.
[0089] A data generation request refers to a request triggered by a user on a client to request a certain data generation process; and the present disclosure does not limit the data generation request. For example, in some application scenarios, the data generation request can be implemented using the content generation request shown in Figure 2 or Figure 3. Among them, the content generation request is used to request content generation processing, such as text-to-image, image-to-image, text-to-video, and image-to-video processing. It can be seen that in one possible implementation, the data generation request can be used to request a certain content generation process, such as text-to-image, image-to-image, text-to-video, and image-to-video processing.
[0090] Based on the above, it can be seen that in some application scenarios, such as content generation scenarios, the data generation request described above can be used to request content generation processing based on reference information, so that the data generation request carries the reference information. The reference information refers to information required for reference when performing content generation processing; and this disclosure does not limit the reference information. For example, in some application scenarios, the reference information can include at least one of text and images. For ease of understanding, the following description is provided in conjunction with some application scenarios.
[0091] Scenario 1: In some scenarios, such as those involving text-generated images and text-generated videos, the data generation request can be used to request text-generated image processing or text-generated video processing; and when the data generation request carries at least prompt text, the data generation request can be used to request text-generated image processing or text-generated video processing based on the prompt text. Based on this, it can be seen that in one possible implementation, the reference information carried by the data generation request can include at least the prompt text, so that the data generation request can be used to request content generation processing, such as text-generated image or text-generated video processing, based on the prompt text.
[0092] Scenario 2: In some scenarios, such as image-to-image and image-to-video, the above data generation request can be used to request image-to-image processing or image-to-video processing; and when the data generation request carries at least an original image (source image), the data generation request can be used to request image-to-image processing or image-to-video processing based on the original image. Based on this, it can be seen that in one possible implementation, the reference information carried by the data generation request can at least include the original image, so that the data generation request can be used to request content generation processing based on the original image, such as image-to-image or image-to-video processing.
[0093] Scenario 3. In some scenarios, such as image generation scenarios based on multiple control information or video generation scenarios based on multiple control information, the above data generation request may carry multiple control information, so that the data generation request can be used to request image generation processing or video generation processing based on these control information. Among them, the multiple control information is used to describe the constraints required in the image generation processing or video generation processing; and the constraints described by different control information are different. In addition, the present disclosure does not limit the implementation methods of the multiple control information. For example, the multiple control information may include at least one text and at least one image. Based on this, it can be seen that in one possible implementation method, the reference information carried by the data generation request may at least include the at least one text and the at least one image, so that the data generation request can be used to request content generation processing based on the at least one text and the at least one image.
[0094] Based on the relevant content of S1 above, it can be known that for the terminal device used by the user, if a client is deployed on the terminal device, then when the user triggers a data generation request on the client, the client can perform some response processing for the data generation request, such as the response processing shown in S2 below, etc., to meet the user's data generation needs.
[0095] S2: The client uses the target plug-in to process the data generation request and obtain the generated data; the target plug-in is constructed based on the data generation model corresponding to the data generation request, and the data generation model is used to process the data generation request.
[0096] Among them, the target plug-in refers to the plug-in required to be used when the client processes the above data generation request, so that the client can use the target plug-in to complete the relevant processing process for the data generation request, so that the relevant processing process for the data generation request all occurs on the same end, so that the data generation processing can be closed-loop on a single end, thereby effectively overcoming the defects that exist when data generation processing is implemented in a multi-end collaborative manner.
[0097] In addition, for the target plug-in mentioned above, the target plug-in is constructed based on the data generation model corresponding to the data generation request mentioned above, so that the data processing function of the target plug-in includes the data processing function of the data generation model, so that when the data generation model can be used to process the data generation request, the target plug-in can also be used to process the data generation request. Among them, the data generation model refers to a pre-constructed model with the data generation request processing function; and the present disclosure does not limit the implementation method of the above data generation model. For example, in some application scenarios, the data generation model can be implemented using a content generation (Artificial Intelligence Generated Content, AIGC) model, so that the data generation model can be used to process any content generation request, such as a text-to-image request, an image-to-image request, a text-to-video request, or an image-to-video request.
[0098] Furthermore, regarding the data generation model in the previous paragraph, the data generation model is deployed on a device independent of the client, such as the server shown in FIG2 , so that the device can provide the client with the relevant resources of the target plug-in based on the data generation model. Thus, in one possible implementation, the target plug-in is constructed based on the data generation model deployed on the server that is capable of processing the data generation request, so that the target plug-in can be used in place of the data generation model. Data communication can be performed between the server and the client.
[0099] In addition, the present disclosure does not limit the implementation of the target plug-in described above. For ease of understanding, some examples are provided below for illustration.
[0100] As an example, in some application scenarios, such as scenarios where the model size is relatively small, when the target plug-in is constructed based on a data generation model deployed on the server that is capable of processing the data generation request, the target plug-in may include a target model determined based on the data generation model, so that the target plug-in can be used to directly process the data generation request using the target model. The target model refers to the model required to process the data generation request using the target plug-in; and the target model is determined based on the data generation model corresponding to the data generation request, so that the data generation function of the target model is consistent with the data generation function of the data generation model, thereby enabling the target model to be used in place of the data generation model.
[0101] It should be noted that the present disclosure does not limit the process of determining the target model in the above paragraph. For example, the process of determining the target model can be specifically: directly determining the data generation model corresponding to the above data generation request as the target model. For another example, the process of determining the target model can be: first extracting the core logic from the data generation model corresponding to the data generation request; then implementing the core logic according to the preset expression to obtain the target model. Among them, the preset expression refers to a pre-set method that can be interpreted on the terminal device; and the present disclosure does not limit the preset expression. For example, it can be pure C++.
[0102] As an example, in some application scenarios, in order to better save resources, a plug-in can be used to replace multiple data processing models on the server, such as an AIGC model, a speech recognition model, etc. Based on this, the present disclosure also provides a possible implementation of the above target plug-in, in which the target plug-in is constructed based on at least two data processing models, and the at least two data processing models include the data generation model corresponding to the above data generation request, so that the target plug-in includes candidate models determined based on each data processing model, so that the functions of the target plug-in include the functions of these data processing models, and thus the target plug-in can be used instead of these data processing models. Based on this, it can be seen that in a possible implementation, the target plug-in can include at least two candidate models, and the at least two candidate models include the above target model. Among them, because different candidate models are constructed based on different data processing models deployed by the server, so that the data processing processes implemented by different candidate models are different, the target plug-in including these candidate models can be used instead of these data processing models deployed by the server.
[0103] Based on the content of the above paragraph, it can be seen that in one possible implementation, when the above target plug-in includes at least two candidate models, and the at least two candidate models include a target model determined based on the data generation model corresponding to the above data generation request, the target plug-in can be used to determine the target model corresponding to the above data generation request from the at least two candidate models, and use the target model to process the data generation request. It should be noted that the present disclosure does not limit the implementation method of the aforementioned step of "determining the target model corresponding to the above data generation request from the at least two candidate models". For example, when the data generation request carries model description information, the step can specifically be: matching the model description information carried by the data generation request with the model description information of each candidate model, and determining the candidate model that best matches the model description information carried by the data generation request as the target model corresponding to the data generation request. Among them, the model description information is used to describe the characteristics of a model; and the present disclosure does not limit the implementation method of the model description information. For example, it can be implemented using a model configuration file.
[0104] As an example, in some application scenarios, such as scenarios where the model volume is relatively large, in order to better overcome the memory overflow defect caused by the large model volume, the present disclosure also provides a possible implementation method of the above target plug-in. In this implementation method, the target plug-in includes at least two data processing modules arranged in sequence. The at least two data processing modules are obtained by splitting the data generation model corresponding to the above data generation request, so that the target plug-in can complete the processing process for the data generation request by using these data processing modules in sequence, such as the module usage method shown in Figure 5 or Figure 6. The two data processing modules refer to the modules required to be used when the target plug-in processes the data generation request.
[0105] In addition, the present disclosure does not limit the splitting process of the data generation model corresponding to the above data generation request. For example, it can be specifically: directly dividing the data generation model into multiple parts, and using each of the divided parts as a data processing module.
[0106] For example, in order to better avoid the occurrence of memory overflow, the present disclosure also provides a possible implementation method of the splitting process of the data generation model corresponding to the above data generation request. Under this implementation method, the splitting process may include the following steps 11 to 13.
[0107] Step 11: Split the data generation model corresponding to the above data generation request into at least two sub-models, where different sub-models implement different data processing functions.
[0108] A sub-model refers to a model within the data generation model corresponding to the data generation request described above that implements a certain data processing function. For example, the sub-model may be the text feature extraction model, image feature encoding model, other information extraction model, image feature decoding model, denoising model, and control information adaptation model shown in Figures 4 or 5.
[0109] In addition, the present disclosure does not limit the implementation method of the above step 11. For example, the specific implementation method of step 11 can be: according to the data processing function, the data generation model corresponding to the above data generation request is split into models to obtain at least two sub-models, so that each sub-model represents a data processing function.
[0110] For another example, step 11 may specifically include: splitting the data generation model corresponding to the data generation request into at least two sub-models according to a pre-set model splitting rule, so that each sub-model can be used for one or more data processing functions. The model splitting rule refers to a rule pre-determined for an actual application scenario for splitting a model into multiple sub-models; and this disclosure does not limit the model splitting rule.
[0111] For example, in order to better improve the flexibility of sub-model splitting, the above step 11 can be specifically as follows: based on the correlation between each pair of adjacent network layers in the data generation model corresponding to the above data generation request, split at least two sub-models from the data generation model, so that the correlation between any pair of network layers in each sub-model is higher than the preset correlation threshold, so that each sub-model can realize one or more data processing functions as completely as possible, and thus can automatically perform model splitting processing under the premise of ensuring the functional integrity of each sub-model. Among them, for any pair of adjacent network layers in the data generation model, the correlation between the pair of adjacent network layers is used to characterize the correlation between the pair of adjacent network layers, so that the correlation between the pair of adjacent network layers can indicate whether the pair of adjacent network layers belong to different steps under the same data processing process; and the present disclosure does not limit the implementation method of the correlation. For example, in some application scenarios, the correlation can be obtained by manual annotation. For example, in some application scenarios, the correlation can be automatically determined during the training process of the data generation model, so that the correlation can be updated accordingly with the training update process of the data generation model.
[0112] Step 12: If the memory requirement representation data of the j-th sub-model in the data generation model corresponding to the above data generation request does not exceed the preset memory threshold, then the j-th sub-model is determined as the data processing module, j is a positive integer, j≤J, J is a positive integer, and J represents the number of sub-models of at least two sub-models in the data generation model corresponding to the above data generation request.
[0113] Among them, the j-th sub-model is used to represent any sub-model in the data generation model corresponding to the above data generation request; and the j-th sub-model is used to implement one or more data processing functions. In addition, the present disclosure does not limit the implementation method of the j-th sub-model. For example, when the data generation model includes the text feature extraction model, the image feature encoding model, the other information extraction model, the picture feature decoding model, and the denoising model and the control information adaptation model shown in Figure 4 or Figure 5, the j-th sub-model can be the text feature extraction model, the image feature encoding model, the other information extraction model or the picture feature decoding model; or, the j-th sub-model can be the denoising model and the control information adaptation model.
[0114] The memory requirement representation data of the j-th sub-model above is used to represent the memory required to be consumed when using the j-th sub-model; and the present disclosure does not limit the method for determining the memory requirement representation data of the j-th sub-model. For example, it can be implemented using any existing or future method that can measure the memory requirements of a model.
[0115] The preset memory threshold is used to indicate the upper limit of the memory consumption of a model; and the present disclosure does not limit the method for obtaining the preset memory threshold, for example, it can be manually set by the user. For example, in some application scenarios, in order to better avoid the occurrence of memory overflow, the preset memory threshold can be determined based on the memory parameters of the terminal device to ensure that the preset memory threshold can more accurately represent the upper limit of memory required when using a model under the terminal device. Among them, the memory parameter is used to describe the memory resource status in the terminal device, such as the maximum available memory, etc.; and the present disclosure does not limit the memory parameter. In addition, the present disclosure does not limit the implementation method of determining the preset memory threshold based on the memory parameter.
[0116] Based on the relevant content of step 12 above, it can be known that for the j-th sub-model in the data generation model corresponding to the above data generation request, if the memory requirement representation data of the j-th sub-model does not exceed the preset memory threshold, it can be determined that no memory overflow will occur when using the j-th sub-model. Therefore, in order to better improve efficiency, there is no need to further split the j-th sub-model, and it is only necessary to directly determine the j-th sub-model as a data processing module.
[0117] Step 13: If the memory requirement representation data of the j-th sub-model in the data generation model corresponding to the above data generation request exceeds the preset memory threshold, the j-th sub-model is split into at least two model fragments, and each model fragment is determined as a data processing module, and the memory requirement representation data of each model fragment does not exceed the preset memory threshold.
[0118] In the present disclosure, for the j-th sub-model in the above data generation model, if the memory requirement characterization data of the j-th sub-model exceeds the preset memory threshold, it can be determined that the possibility of memory overflow when using the j-th sub-model is relatively high. Therefore, in order to better avoid the occurrence of memory overflow, the j-th sub-model is split into at least two model fragments, such as fragment 1-fragment 4 shown in Figure 4 or Figure 5, so that the memory requirement characterization data of each model fragment does not exceed the preset memory threshold, so that memory overflow will not occur when using each model fragment, and each model fragment is determined as a data processing module.
[0119] In addition, for the j-th sub-model above, the at least two model fragments obtained by splitting the j-th sub-model have the following characteristics: Feature 1, for any model fragment, the model fragment includes one or more network layers in the j-th sub-model, so that the model fragment can represent a certain part of the j-th sub-model, so that the model fragment can realize the function required to be realized by the corresponding part of the j-th sub-model. Feature 2, different model fragments are used to represent different parts of the j-th sub-model, so that the union of the at least two model fragments can realize the data processing function of the j-th sub-model. Feature 3, there is no intersection between any two model fragments.
[0120] Furthermore, the present disclosure does not limit the implementation of the step of "splitting the j-th sub-model into at least two model segments" in step 13 above. For example, the j-th sub-model can be equally divided into multiple model segments. In another example, the implementation can be based on a pre-set sub-model splitting rule. The sub-model splitting rule refers to a rule pre-set based on the application scenario and used to split a sub-model with relatively high memory consumption.
[0121] In addition, in order to better improve the flexibility of sub-model splitting, the present disclosure also provides a possible implementation method of the splitting process of the j-th sub-model above. Under this implementation method, the splitting process can specifically include the following steps 131-134.
[0122] Step 131: Obtain at least one candidate split description information, where different candidate split description information describe different model splitting methods.
[0123] The kth candidate split description information refers to the information required when performing splitting processing on a sub-model according to the kth splitting method, so that the kth candidate split description information can represent the characteristics presented by the kth splitting method. k is a positive integer, k≤the number of candidate split description information in the at least one candidate split description information.
[0124] In addition, the present disclosure does not limit the implementation of the k-th candidate split description information. For example, the k-th candidate split description information may describe the various split positions required when splitting a sub-model according to the k-th splitting method. For another example, the k-th candidate split description information may be used to describe the splitting rules of the k-th splitting method, such as splitting position determination rules.
[0125] In addition, the present disclosure does not limit the method for obtaining the k-th candidate split description information above. For example, it can be set in advance by relevant personnel.
[0126] Step 132: For any candidate split description information, split the j-th sub-model above according to the candidate split description information to obtain the model splitting result corresponding to the candidate split description information, and determine the resource usage characterization data corresponding to the candidate split description information based on the model splitting result corresponding to the candidate split description information, and the resource usage characterization data includes computing resource reuse characterization data and memory usage characterization data.
[0127] The model splitting result corresponding to the k-th candidate split description information refers to the result obtained by splitting the j-th sub-model according to the k-th candidate split description information, so that the "model splitting result corresponding to the k-th candidate split description information" can represent the multiple segments obtained by splitting the j-th sub-model using the k-th splitting method. k is a positive integer, k≤the number of candidate split description information in the at least one candidate split description information.
[0128] The resource usage characterization data corresponding to the kth candidate split description information refers to the resource usage presented when the model splitting result corresponding to the kth candidate split description information is used to implement the data processing function described by the jth sub-model above, such as computing resource reuse and memory occupancy.
[0129] In addition, the present disclosure does not limit the implementation method of the resource usage characterization data corresponding to the k-th candidate split description information above. For example, it may at least include the computing resource reuse characterization data corresponding to the k-th candidate split description information and the memory usage characterization data corresponding to the k-th candidate split description information. Among them, the computing resource reuse characterization data corresponding to the k-th candidate split description information refers to the computing reuse situation presented when the model splitting result corresponding to the k-th candidate split description information is used to realize the data processing function described by the j-th sub-model above. The memory usage characterization data corresponding to the k-th candidate split description information refers to the memory occupancy situation presented when the model splitting result corresponding to the k-th candidate split description information is used to realize the data processing function described by the j-th sub-model.
[0130] In addition, the present disclosure does not limit the method for obtaining the resource usage characterization data corresponding to the k-th candidate split description information above. For example, it can be implemented using any existing or future method that can analyze the computational reuse and memory occupancy of some models. As an example, the process for obtaining the resource usage characterization data corresponding to the k-th candidate split description information can be: processing a sample request using the model splitting result corresponding to the k-th candidate split description information, and recording the computational reuse information and memory occupancy information involved in the sample request processing, so that after the sample request processing is completed, the resource usage characterization data corresponding to the k-th candidate split description information can be determined based on these computational reuse information and memory occupancy information. The sample request refers to the request required to analyze the computational reuse and memory occupancy of the j-th sub-model presented in a certain candidate split description information; and the sample request is similar to the above data generation request.
[0131] Step 133: Based on the resource usage characterization data corresponding to each candidate split description information, select target split description information from at least one of the above candidate split description information, and the resource usage characterization data corresponding to the target split description information presents a higher resource usage balance than the resource usage characterization data corresponding to any other candidate split description information in the at least one candidate split description information except the target split description information.
[0132] In the present disclosure, for the j-th sub-model above, after determining the resource usage characterization data corresponding to each candidate split description information for the j-th sub-model, target split description information can be selected from at least one candidate split description information above based on these resource usage characterization data, so that the target split description information is used to represent the candidate split description information that matches the j-th sub-model, so that the target split description information can represent the splitting method finally selected for the j-th sub-model. The target split description information refers to the candidate split description information selected from at least one candidate split description information and matching the j-th sub-model, so that the target split description information can represent the information required for splitting the j-th sub-model, such as each split position.
[0133] In addition, for the j-th sub-model above, the conditions required when selecting the target split description information that matches the j-th sub-model from at least one candidate split description information above include: the resource usage balance degree presented by the resource usage characterization data corresponding to the target split description information is higher than the resource usage balance degree presented by the resource usage characterization data corresponding to any other candidate split description information in the at least one candidate split description information except the target split description information. Among them, the resource usage balance degree presented by the resource usage characterization data corresponding to the k-th candidate split description information is used to characterize the balance degree between computational reuse and memory occupancy when the model splitting result corresponding to the k-th candidate split description information is used to realize the data processing function described by the j-th sub-model above, so that the resource usage balance degree can represent the balance presented by the model splitting result corresponding to the k-th candidate split description information under multi-dimensional performance; and the present disclosure does not limit the acquisition process of the resource usage balance degree. For example, it can be implemented using any existing or future balance degree calculation method under multi-dimensional performance, such as a method based on a pre-set balance calculation rule or a method based on a pre-set balance calculation formula, etc. k is a positive integer, k≤the number of candidate split description information in the at least one candidate split description information.
[0134] Based on the relevant content of step 133 above, it can be known that for the j-th sub-model above, after determining the resource usage characterization data corresponding to each candidate split description information for the j-th sub-model, the target split description information can be selected from at least one candidate split description information above based on these resource usage characterization data, so that the result obtained by splitting the j-th sub-model according to the target split description information can show the maximum balance between computational reuse and memory occupancy, so that the number of computational reuses can be reduced as much as possible while ensuring that no memory overflow occurs, which is conducive to achieving the maximum response speed without causing memory overflow.
[0135] Step 134: Determine at least two model segments corresponding to the j-th sub-model according to the model splitting result corresponding to the target splitting description information.
[0136] In the present disclosure, for the jth sub-model above, after selecting the target split description information for the jth sub-model from at least one candidate split description information above, at least two model fragments corresponding to the jth sub-model can be determined based on the model splitting result corresponding to the target split description information, so that the at least two model fragments can represent the result obtained by splitting the jth sub-model according to the splitting method described in the target split description information.
[0137] Based on the relevant contents of steps 131 to 134 above, it can be known that in some application scenarios, for the j-th sub-model in the data generation model corresponding to the above data generation request, if the memory requirement representation data of the j-th sub-model exceeds the preset memory threshold, it can be determined that there is a high possibility of memory overflow when using the j-th sub-model. Therefore, in order to better avoid the occurrence of memory overflow, the target split description information matching the j-th sub-model can be selected from some pre-set candidate split description information; and then, based on the model splitting result corresponding to the j-th sub-model under the target split description information, at least two model fragments corresponding to the j-th sub-model are determined. This is conducive to reducing the number of calculation reuses as much as possible while ensuring that no memory overflow occurs, thereby helping to increase the response speed as much as possible without causing memory overflow.
[0138] Based on the relevant contents of steps 11 to 13 above, it can be seen that in some application scenarios, for any sub-model in the data generation model corresponding to the above data generation request, if the memory requirement representation data of the sub-model does not exceed the preset memory threshold, it can be determined that no memory overflow will occur when the sub-model is used, so the sub-model can be directly determined as a data processing module; however, if the memory requirement representation data of the sub-model exceeds the preset memory threshold, it can be determined that the possibility of memory overflow when the sub-model is used is relatively high, so in order to better avoid the occurrence of memory overflow, the sub-model can be first split and processed to obtain at least two model fragments, so that the memory requirement representation data of each model fragment does not exceed the preset memory threshold; and then each model fragment is determined as a data processing module, so that it can effectively ensure that the memory requirement representation data of all data processing modules split from the data generation model do not exceed the preset memory threshold, thereby effectively ensuring that no memory overflow will occur when these data processing modules are used instead of the data generation model to implement the corresponding data generation process, which is conducive to avoiding defects caused by memory overflow. In addition, the present disclosure automatically splits the data generation model into at least two data processing modules by executing steps 11 to 13, which is conducive to improving the flexibility of model splitting, thereby effectively avoiding defects caused by manual model splitting.
[0139] In addition, for the data generation model corresponding to the above data generation request, the at least two data processing modules obtained by splitting the data generation model have the following characteristics: the at least two data processing modules are arranged in sequence, so that the data generation process described by the data generation model can be implemented by using the at least two data processing modules in sequence according to the arrangement order. In this way, the functions described by the data generation model can be better implemented with the help of the at least two data processing modules. It should be noted that the present disclosure does not limit the method for determining the arrangement sequence of each data processing module. For example, the arrangement sequence of each data processing module can be determined in the determination process of each data processing module.
[0140] Based on the relevant content of at least two data processing modules above, it can be seen that in some application scenarios, for the target plug-in constructed according to the data generation model corresponding to the data generation request above, in order to better avoid memory overflow when the target plug-in is used to implement data generation processing, the target plug-in can include at least two data processing modules arranged in sequence, and the at least two data processing modules are obtained by splitting the data generation model, so that the target plug-in can implement the function described by the data generation model with the help of the at least two data processing modules, such as processing the data generation request to obtain generated data, etc., so that it can be ensured that when the target plug-in is used to process the data generation request, no memory overflow will occur, thereby effectively ensuring the stability of the data generation processing implemented by a single end. Wherein, the generated data refers to the data obtained by processing the data generation request; and the present disclosure does not limit the generated data. For example, in some application scenarios, when the data generation request is used to request content generation processing based on reference information, the generated data can include at least one generated image, which refers to the result obtained by performing content generation processing based on the reference information. For example, for a text-to-image scenario or an image-to-image scenario, the generated data is a generated image. For example, for a scene of text-generated video or image-generated video, the generated data is a generated image sequence, such as a generated video.
[0141] Furthermore, for the target plug-in shown in the previous paragraph, since it includes at least two data processing modules arranged in sequence, to minimize the occurrence of memory overflow, the target plug-in can implement the corresponding data generation process by sequentially loading and releasing each data processing module. This ensures that sufficient memory is available during the use of each data processing module, effectively preventing memory overflow when implementing data generation processing with the target plug-in. To facilitate understanding of the foregoing, the following examples are provided for illustration.
[0142] As an example, when the above target plug-in includes at least two data processing modules arranged in sequence, and the at least two data processing modules are obtained by splitting the data generation model corresponding to the above data generation request, in order to avoid memory overflow, the above client can use the target plug-in to process the data generation request to obtain generated data, and the determination process of the generated data can include the following steps 21-24.
[0143] Step 21: Load the i-th data processing module; i is a positive integer, and its initial value is 1.
[0144] Here, i refers to the serial number of the data processing module to be used in the current round; and the initial value of i is 1.
[0145] The i-th data processing module refers to the data processing module that exists in the i-th arrangement position among the at least two data processing modules above.
[0146] Based on the relevant content of step 21 above, it can be known that for the current round, the i-th data processing module above is loaded into the memory so that the corresponding data processing process can be implemented later using the i-th data processing module already loaded in the memory.
[0147] Step 22: Use the i-th data processing module to perform data processing.
[0148] In the present disclosure, for the current round, after determining that the i-th data processing module has been loaded into the memory, it can be determined that the i-th data processing module is in an available state, so the i-th data processing module can be directly used to process the input data of the i-th data processing module. The input data of the i-th data processing module refers to the data input by the i-th data processing module; and the present disclosure does not limit the implementation method of the input data of the i-th data processing module. For example, if i=1, the input data of the i-th data processing module is determined based on the data generation request above, so that the input data of the i-th data processing module can at least include the reference information carried by the data generation request; if i>1, the input data of the i-th data processing module is determined based on the output data of the i-1-th data processing module, so that the input data of the i-th data processing module at least includes all or part of the output data of the i-1-th data processing module.
[0149] In addition, in some application scenarios, in order to better improve the data processing effect, some data processing modules need to be initialized before use. Based on this, the present disclosure also provides a possible implementation of step 22 above. Under this implementation, step 22 can be specifically as follows: if the i-th data processing module above meets the pre-set initialization conditions, then after determining that the i-th data processing module has been loaded into the memory, the i-th data processing module is initialized so that the initialized i-th data processing module has better data processing function; and then the initialized i-th data processing module is used to process data, which is conducive to improving the data generation effect. Among them, the initialization condition refers to the condition reached by the model that needs to be initialized; and the present disclosure does not limit the implementation of the initialization condition. For example, in some application scenarios, the initialization condition can be specifically as follows: the initialization parameters of the i-th data processing module are recorded in the preset storage space of the terminal device.
[0150] Based on the content of the above paragraph, it can be seen that for the current round, after determining that the above i-th data processing module has been loaded into the memory, it is determined whether the initialization parameters corresponding to the i-th data processing module are recorded in the preset storage space of the terminal device. If the initialization parameters corresponding to the i-th data processing module are recorded, it can be determined that the i-th data processing module meets the initialization conditions, and thus it can be determined that initialization processing is still required for the i-th data processing module, and then it can be determined that the i-th data processing module is in an unavailable state. Therefore, the i-th data processing module is initialized first to obtain the initialized i-th data processing module, so that the initialized i-th data processing module is in an available state; then the initialized i-th data processing module is used for data processing, which is conducive to improving the data processing effect.
[0151] Step 23: After obtaining the output result of the i-th data processing module, release the memory occupied by the i-th data processing module.
[0152] The output result of the i-th data processing module refers to the result obtained by the i-th data processing module performing data processing on the input data of the i-th data processing module.
[0153] Based on the relevant content of step 23 above, it can be known that for the current round, after obtaining the output result of the i-th data processing module above, it can be determined that the corresponding data processing process has been completed using the i-th data processing module, so that it can be determined that the memory occupied by the i-th data processing module can be recycled. Therefore, in order to better avoid memory overflow, the memory occupied by the i-th data processing module can be directly released to reduce the memory usage, so that the memory requirements used by subsequent modules can be met as much as possible, thereby effectively avoiding the occurrence of memory overflow.
[0154] Step 24: Update i and continue to execute step 21 and subsequent steps until a preset stop condition is reached, and determine to generate data based on the output result of the i-th data processing module.
[0155] The present disclosure does not limit the implementation of the above “updating i”, for example, it can be implemented using the following formula (1). i'=i+1 (1)
[0156] Where i' represents the updated value; i represents the value before the update.
[0157] The preset stopping condition refers to the condition that must be met to stop executing the multiple rounds of iterations; and the present disclosure does not limit the preset stopping condition. For example, the preset stopping condition may be: all modules in the at least two data processing modules have been traversed. For another example, when the "update i" step above is implemented using formula (1) above, the preset stopping condition may be: i equals the number of modules in the at least two data processing modules above.
[0158] In addition, the present disclosure does not limit the timing of determining the preset stop condition. For example, in some application scenarios, in order to maximize response speed, the timing of determining the preset stop condition can be earlier than the timing of executing the "update i" step above. Based on this, it can be seen that in one possible implementation, the above step 24 can be specifically as follows: after releasing the memory occupied by the i-th data processing module, determine whether the preset stop condition has been met. If the preset stop condition has not been met, update i and continue to execute the above step 21 and its subsequent steps; however, if the preset stop condition has been met, the generated data can be determined based on the output result of the i-th data processing module.
[0159] In addition, the present disclosure does not limit the implementation method of the above step of "determining the generated data based on the output result of the i-th data processing module". For example, in some application scenarios, it can be specifically as follows: determining the generated data based on the output result of the i-th data processing module, so that the generated data at least includes the output result of the i-th data processing module. For another example, it can be specifically as follows: directly determining the output result of the i-th data processing module as the generated data. For another example, it can be specifically as follows: determining the generated data based on the output result of at least one data processing module, so that the generated data includes the output result of the at least one data processing module, and the at least one data processing module includes the i-th data processing module.
[0160] Based on the relevant content of steps 21 to 24 above, it can be seen that for the above target plug-in, if the target plug-in includes at least two data processing modules arranged in sequence, and the at least two data processing modules are obtained by splitting the data generation model corresponding to the above data generation request, then when the client uses the target plug-in to process the data generation request to obtain generated data, it can use these data processing modules in sequence in a certain manner, such as the manner shown in Figure 4 or Figure 5, to obtain the generated data. This can effectively ensure that each data processing module has sufficient memory during use, thereby effectively avoiding defects caused by memory overflow, such as the inability to respond to the data generation request.
[0161] Based on the relevant content of the target plug-in above, it can be seen that for the target plug-in, since the target plug-in is constructed based on the data generation model corresponding to the above data generation request, the data processing function of the target plug-in includes the data processing function of the data generation model, so that the target plug-in can replace the data generation model to implement the corresponding function; and because the data generation model can be used to process the data generation request, the target plug-in can also be used to process the data generation request, so that after the client receives the data generation request, the client can directly use the target plug-in to process the data generation request, thereby making all related processes of the data generation request occur at the same end, so that the data generation processing can be single-ended closed-loop, thereby effectively improving the real-time, security, reliability and other aspects of the data generation process.
[0162] In addition, the present disclosure does not limit the above client's use of the target plug-in, that is, the implementation of S2 above. For ease of understanding, the following description is combined with some situations.
[0163] Case 1. In some application scenarios, in order to better improve the response speed of the above data generation request, the present disclosure also provides a possible implementation method of the above S2. Under this implementation method, when the target plug-in has been deployed in the client, the S2 can be specifically: after the client receives the data generation request, the client directly uses the target plug-in deployed in the client to process the data generation request and obtain the generated data.
[0164] It can be seen that in one possible implementation method, the above target plug-in can be directly deployed into the client, so that the client can directly use the target plug-in to process the data generation request, so that the processing of the data generation request involves almost no data communication process, which can effectively reduce the time required for data communication, thereby helping to improve the response speed to the data generation request, and thus helping to improve the data generation experience.
[0165] Case 2: In some application scenarios, in order to minimize the impact of plug-in deployment on client performance, the present disclosure also provides a possible implementation of S2 above. In this implementation, when the client is deployed on the terminal device, the above target plug-in is also deployed on the terminal device, and the client can communicate data with the target plug-in through the plug-in service deployed on the client, S2 can specifically include the following steps 31-32.
[0166] Step 31: The client sends a data generation request to the target plug-in deployed on the terminal device through the plug-in service deployed on the client.
[0167] The plug-in service is used to assist the client in calling the target plug-in; and the plug-in service is deployed in the client.
[0168] In addition, the present disclosure does not limit the implementation method of the above plug-in service. For example, it can be implemented in any existing or future manner that can assist the client in calling the plug-in. As an example, in one possible implementation method, as shown in Figure 3, the plug-in service may include a plug-in service interface and a plug-in host. The plug-in service interface is used to obtain a data generation request triggered by a user and send the data generation request to the plug-in host. The plug-in host is used to send the data generation request to the target plug-in so that the target plug-in can process the data generation request. It should be noted that the present disclosure does not limit the implementation method of the plug-in service interface and the plug-in host.
[0169] In addition, in order to better save plug-in deployment costs, the present disclosure also provides a possible implementation of the above plug-in service. Under this implementation, one plug-in service corresponds to multiple plug-ins, so that the client can call different plug-ins through the same plug-in service.
[0170] Based on the content of the above paragraph, the present disclosure also provides a possible implementation method of the above target plug-in determination process. Under this implementation method, when the plug-in service deployed on the client corresponds to multiple candidate plug-ins, the target plug-in determination process can be specifically as follows: determine the target plug-in corresponding to the above data generation request from the multiple candidate plug-ins. Among them, the candidate plug-in refers to a plug-in that can be called with the help of the plug-in service; and the correspondence between the multiple candidate plug-ins and the plug-in service can be pre-set. In addition, the present disclosure does not limit the implementation method of the multiple candidate plug-ins. For example, the multiple candidate plug-ins may include content generation plug-ins, voice recognition plug-ins, three-dimensional model construction plug-ins, etc. The content generation plug-in is used to process any content generation request. The voice recognition plug-in is used to process any voice recognition request. The three-dimensional model construction plug-in is used to process any three-dimensional model construction request.
[0171] Step 32: The client receives the generated data fed back by the target plug-in deployed on the terminal device in response to the data generation request.
[0172] Based on the relevant content of steps 31 to 32 above, it can be seen that in one possible implementation method, the above target plug-in can be completely independent of the client, and the target plug-in and the client are deployed on the same terminal device, so that the client can call the target plug-in with the help of the plug-in service deployed in the client; and the calling process can be specifically as follows: after the user triggers a data generation request on the client, the client first sends a data generation request to the target plug-in deployed on the terminal device through the plug-in service deployed on the client, so that the target plug-in processes the data generation request to obtain generated data, and feeds the generated data back to the client for display. Among them, because the target plug-in and the client are deployed on the same terminal device, the relevant processes of the data generation request all occur on the terminal device, so that single-end processing of the data generation request can be achieved; and because the target plug-in is completely independent of the client, the relevant processes of the target plug-in are also completely independent of the relevant processes of the client, so that the relevant processes of the target plug-in will hardly affect the relevant processes of the client, so that single-end processing of the data generation request can be achieved without affecting the client performance, which is conducive to improving the data generation experience without affecting the user's client usage experience.
[0173] Based on the relevant contents of S1 to S2 above, it can be seen that for the data generation method provided by the embodiment of the present disclosure, after the client receives the data generation request, the client uses the target plug-in to process the data generation request to obtain generated data, so that data generation processing can be realized on the client based on the plug-in, thereby effectively overcoming the defects existing when data generation processing is realized by means of a collaborative method between the client and the server, thereby helping to improve the data generation effect. Among them, because the target plug-in is constructed based on the data generation model corresponding to the data generation request, the data processing function of the target plug-in includes the data processing function of the data generation model; and because the data generation model can be used to process the data generation request, the target plug-in can also be used to process the data generation request, so that after the client receives the data generation request, the client can directly use the target plug-in to process the data generation request, thereby making all related processes of the data generation request occur at the same end, so that the data generation processing can be closed in a single end, thereby effectively improving the effect of the data generation process in terms of real-time, security, reliability, etc.
[0174] In fact, in order to better improve the response speed, the present disclosure also provides a possible implementation of the above data generation method, in which the data generation method not only includes the above S1-S2, but also the following step 41. The execution time of step 41 is earlier than the execution time of S2.
[0175] Step 41: In response to the client startup request, start the client and load the target plug-in.
[0176] The client startup request is used to request startup of the client; and the present disclosure does not limit the implementation method of the client startup request. For example, it can be implemented using any existing or future client startup request.
[0177] Based on the relevant content of step 41 above, it can be seen that in some application scenarios, for a terminal device used by a user, after the terminal device detects that the user has triggered a client startup request, it starts the client and loads the target plug-in. In this way, the target plug-in can be loaded as early as possible, so that when the user subsequently triggers a data generation request on the client, the client can directly use the loaded target plug-in to process the data generation request and obtain the generated data. This can effectively improve the response speed to the data generation request, thereby helping to improve the user experience.
[0178] In fact, in some application scenarios, in order to minimize the impact of plug-in loading on client startup, the present disclosure also provides a possible implementation of the above data generation method. In this implementation, the data generation method not only includes the above S1-S2, but also the following step 42. The execution time of step 42 is earlier than the execution time of S2.
[0179] Step 42: In response to the client start request, start the client and asynchronously preload the target plug-in.
[0180] It should be noted that the present disclosure does not limit the implementation method of the asynchronous preloading in the above step 42. For example, it can be implemented using any existing or future asynchronous preloading method.
[0181] Based on the relevant content of step 42 above, it can be seen that in some application scenarios, for a terminal device used by a user, after the terminal device detects that the user has triggered a client startup request, it starts the client and asynchronously preloads the target plug-in so that the loading process of the target plug-in does not affect the startup process of the client, thereby being able to load the target plug-in as early as possible without affecting the startup speed of the client, which is conducive to improving user experience.
[0182] In practice, in some application scenarios, the target plug-in described above, such as one built based on a diffusion model, may require initialization to improve performance. Based on this, the present disclosure also provides a possible implementation of the data generation method described above. In this implementation, the data generation method may include at least step 43 below. Step 43 is executed later than step 41 or step 42.
[0183] Step 43: Initialize the loaded target plug-in so that the client can subsequently use the initialized target plug-in to process the above data generation request.
[0184] It should be noted that the present disclosure does not limit the implementation method of the initialization processing in step 43 above. For example, it can use the initialization parameters pre-stored for the target plug-in in the terminal device to initialize the loaded target plug-in so that the initialized target plug-in has better performance.
[0185] Based on the relevant content of step 43 above, it can be known that in some application scenarios, for a terminal device used by a user, after the terminal device detects that the user has triggered a client startup request, it starts the client and loads (or asynchronously preloads) the target plug-in; then the loaded target plug-in is initialized to obtain an initialized target plug-in, so that after the client receives a data generation request triggered by the user, the client can directly use the initialized target plug-in to process the data generation request and obtain generated data. Among them, because the initialized target plug-in has better performance, the generated data obtained by using the initialized target plug-in is more accurate, which is conducive to improving the data generation effect.
[0186] In fact, in order to better improve the diversity of data generation, the present disclosure also provides a possible implementation of the above-mentioned target plug-in. Under this implementation, the target plug-in is constructed based on at least one data generation model. Different data generation models implement different data generation processes. The at least one data generation model includes the data generation model corresponding to the above-mentioned data generation request. Among them, the at least one data generation model refers to the model required to be based on when constructing the target plug-in; and the present disclosure does not limit the at least one data generation model. For example, in some application scenarios, such as multi-style image generation scenarios or multi-style video generation scenarios, the at least one data generation model may include at least one style content generation model; and different style content generation models are used to implement content generation processing under different styles, such as image generation processing or video generation processing. Among them, the nth style content generation model is used to implement content generation processing under the nth style, n is a positive integer, n≤N, N is a positive integer, and N represents the number of models in the at least one style content generation model. In addition, the present disclosure does not limit the implementation of the at least one style content generation model. For example, the at least one style content generation model may include the image generation model of style 1, the image generation model of style 2, ..., the image generation model of style N shown in FIG6 .
[0187] In addition, in some application scenarios, in order to realize as diverse data generation processes as possible on a terminal device with limited resources, such as image generation processes under N styles, etc. The present disclosure also provides a possible implementation of the above target plug-in. Under this implementation, when the target plug-in is constructed based on at least one data generation model, the data generation processes realized by different data generation models are different, and the at least one data generation model includes the data generation model corresponding to the above data generation request, the target plug-in may include a base model and at least one fine-tuning model determined based on the at least one data generation model, so that the target plug-in can realize a variety of data generation processes with the help of the base model and the at least one fine-tuning model. Among them, the base model refers to a model determined based on the at least one data generation model and can be used as a base model, so that the base model can present some common features of the various data generation processes described by the at least one data generation model, so that the base model can represent the base model required to be used when using the target plug-in to realize any data generation process described by the at least one data generation model, such as the base model shown in Figure 6. The nth fine-tuning model refers to a model determined based on the nth data generation model and used to provide parameter fine-tuning information, such as a lora model, so that the nth fine-tuning model can exhibit certain characteristics unique to the nth data generation model, so that the nth data generation model can be subsequently obtained by using the nth fine-tuning model to perform parameter correction on the base model, so that the nth fine-tuning model and the base model can subsequently be used to replace the nth data generation model. Moreover, the present disclosure does not limit the implementation method of the nth fine-tuning model. For example, the nth fine-tuning model can be implemented using the fine-tuning model of style n shown in Figure 6, where n is a positive integer, n≤N, and N is a positive integer, and N represents the number of models in the at least one fine-tuning model. In particular, because the storage resources occupied by the base model and the at least one fine-tuning model are much smaller than the storage resources occupied by the at least one data generation model, the terminal device used to store the base model and the at least one fine-tuning model can implement as diverse image generation processing as possible within limited storage resources, which is conducive to increasing the diversity of data generation and thus better meeting the data generation needs of users.
[0188] In addition, for the target plug-in shown in the previous paragraph, when the target plug-in is used to implement a certain data generation process, in order to better avoid the occurrence of memory overflow, the base model in the target plug-in and the fine-tuning model corresponding to the data generation process can be loaded at different time points. This can effectively avoid the memory pressure caused by loading the base model and the fine-tuning model at the same time, thereby effectively overcoming the defects caused by the large amount of memory required when directly loading the data generation model used to implement the data generation process, such as the image generation model of style 1 shown in Figure 6, etc., and thus help to better avoid the occurrence of memory overflow, which is conducive to improving the stability of the data generation process. For a better understanding, a possible implementation of the above data generation method is described below as an example.
[0189] As an example, in one possible implementation, when the above target plug-in includes a base model and at least one fine-tuning model, the base model and the at least one fine-tuning model are determined based on at least one data generation model, the data generation processes implemented by different data generation models are different, and the at least one data generation model includes the data generation model corresponding to the above data generation request, the data generation method provided by the present disclosure may include the following steps 51-54.
[0190] Step 51: In response to the client start request, start the client and load (or asynchronously preload) the basic model in the target plug-in.
[0191] Step 52: After the client receives the data generation request, the client determines the fine-tuning model corresponding to the data generation request from at least one fine-tuning model of the target plug-in.
[0192] It should be noted that the present disclosure does not limit the implementation method of the above step 52. For example, when the above data generation request carries model description information, the step 52 can specifically be: matching the model description information carried by the data generation request with the model description information of each fine-tuning model in the target plug-in, and determining the fine-tuning model that best matches the model description information carried by the data generation request as the fine-tuning model corresponding to the data generation request.
[0193] It should also be noted that the present disclosure does not limit the execution entity of the step "determining the fine-tuning model corresponding to the data generation request from at least one fine-tuning model of the target plug-in" in the above step 52. For example, the execution entity of this step can be the client or the target plug-in.
[0194] Step 53: Utilize the fine-tuning model corresponding to the above data generation request to perform parameter correction processing on the loaded basic model to obtain the data generation model corresponding to the data generation request.
[0195] It should be noted that the present disclosure does not limit the implementation method of the above step 53. For example, it can be implemented by adopting any existing or future method that can perform parameter correction processing on a basic model based on a fine-tuning model, such as a lora model.
[0196] It should be noted that the present disclosure does not limit the execution subject of the above step 53. For example, the execution subject of step 53 may be a client or a target plug-in.
[0197] Step 54: The client processes the data generation request using the data generation model corresponding to the above data generation request.
[0198] Based on the relevant content of steps 51 to 54 above, it can be seen that in some application scenarios, the target plug-in can use a basic model and some fine-tuning models to implement multiple data generation processes, so as to achieve as diverse data generation processes as possible on terminal devices with limited resources, thereby effectively overcoming the defects caused by limited terminal device resources, and thus helping to improve data generation diversity, which is conducive to improving user experience.
[0199] In fact, in some application scenarios, in order to better improve user experience, the present disclosure also provides a possible implementation of the above data generation method. In this implementation, the data generation method at least includes the following steps 61 to 63.
[0200] Step 61: The client receives a data generation request.
[0201] It should be noted that for the relevant content of step 61, please refer to S1 above.
[0202] Step 62: The client determines whether the target plug-in is in an available state.
[0203] In the present disclosure, for the client, after the client receives the data generation request triggered by the user, the client can determine whether the target plug-in corresponding to the data generation request is in an available state. If the target plug-in is in an available state, it can be determined that the target plug-in can immediately process the data generation request, so the client can directly use the target plug-in to process the data generation request; however, if the target plug-in is in an unavailable state, it can be determined that the target plug-in cannot immediately (or even cannot) process the data generation request, so the client can use other methods, such as the method shown in step 64 below, to complete the processing process of the data generation request.
[0204] It should be noted that this disclosure does not limit the implementation of the above-mentioned step of "determining whether the target plug-in is in an available state." For example, it can be specifically: if the client determines that the target plug-in is in a deployment phase, such as a loading phase or an initialization phase, then the client can determine that the target plug-in is in an unavailable state; however, if the client determines that the target plug-in has completed the entire deployment phase, then the client can determine that the target plug-in is in an available state. The deployment phase refers to the phase in which the target plug-in is deployed to the client or terminal device; and this disclosure does not limit the deployment phase. For example, the deployment phase can include a loading phase and an initialization phase.
[0205] Step 63: If the client determines that the target plug-in is in an available state, the client uses the target plug-in to process the above data generation request to obtain generated data.
[0206] Based on the relevant content of steps 61 to 63 above, it can be seen that in some application scenarios, for the above client, after the client receives the data generation request triggered by the user, the client can determine whether the target plug-in corresponding to the data generation request is in an available state, so that when it is determined that the target plug-in is in an available state, the client can directly use the target plug-in to process the data generation request, which is conducive to improving the response speed.
[0207] In fact, in some application scenarios, in order to further improve the response speed, the present disclosure also provides a possible implementation of the above data generation method. In this implementation, the data generation method may include at least steps 61 to 63 above, and steps 64 and 65 below. In particular, the execution time of step 64 is later than the execution time of step 62.
[0208] Step 64: If the client determines that the target plug-in is in an unavailable state, the client sends the above data generation request to the server, and the server is used to process the data generation request using the data generation model deployed on the server to obtain generated data.
[0209] The data generation model deployed on the server refers to the data generation model corresponding to the above data generation request, so that the data generation model deployed on the server can be used to build the above target plug-in; and the relevant content of the data generation model can be found above.
[0210] Based on the relevant content of step 64 above, it can be seen that for the above client, after the client receives the data generation request triggered by the user, if the client determines that the target plug-in corresponding to the data generation request is in an unavailable state, it can be determined that the client cannot use the target plug-in to complete the processing process for the data generation request. Therefore, in order to improve the response speed as much as possible, the client can directly send the data generation request to the server, so that the server can use the data generation model deployed on the server to process the data generation request, obtain the generated data, and feed the generated data back to the client for display, so that the client can display the generated data to the user in the shortest possible time, which is conducive to improving the response speed and thus improving the user experience.
[0211] It should be noted that the present disclosure does not limit the association relationship between the model involved in the above-mentioned target plug-in and the data generation model deployed on the server. For example, the model involved in the target plug-in is determined based on the data generation model deployed on the server, so that the model involved in the target plug-in and the data generation model deployed on the server have the same function, so that the target plug-in can use the model involved in the target plug-in to implement the function described by the data generation model deployed on the server, and thus the target plug-in can replace the data generation model deployed on the server to implement the corresponding function.
[0212] It should also be noted that the present disclosure does not limit the implementation method of the above server. For example, it can be implemented using any existing or future server, such as an independent server, a cluster server or a cloud server.
[0213] Step 65: The client receives the generated data fed back by the server.
[0214] Based on the relevant content of steps 61 to 65 above, it can be seen that in some application scenarios, for the above client, after the client receives the data generation request triggered by the user, the client can determine whether the target plug-in corresponding to the data generation request is in an available state. If the target plug-in is in an available state, it can be determined that the target plug-in can immediately process the data generation request, so the client can directly use the target plug-in to process the data generation request; however, if the target plug-in is in an unavailable state, it can be determined that the target plug-in cannot immediately (or even cannot) process the data generation request, so the client can directly send the data generation request to the server, so that the client can use the corresponding model deployed on the server to complete the processing process for the data generation request, and obtain and display the generated data fed back by the server, which is conducive to improving the response speed and thus improving the user experience.
[0215] In fact, in some application scenarios, there may be multiple reasons why the target plug-in is in an unavailable state, such as the target plug-in is in the deployment stage or the relevant resources of the target plug-in do not exist in the terminal device. Therefore, in order to better improve the data generation effect, the present disclosure also provides a possible implementation method of the above data generation method. Under this implementation method, when the above client is deployed on the terminal device, the data generation method can at least include the following steps 71-73.
[0216] Step 71: If the client determines that the target plug-in is in an unavailable state, the client determines whether a plug-in description resource of the target plug-in is stored in the terminal device.
[0217] The target plug-in description resources refer to the resources required to deploy the target plug-in on a client or terminal device. Furthermore, the present disclosure does not limit the implementation of the target plug-in description resources. For example, the target plug-in description resources may include loading phase resources and initialization phase resources. The loading phase resources refer to the resources required to load the target plug-in. The initialization phase resources refer to the resources required to initialize the target plug-in.
[0218] Step 72: If the client determines that the terminal device does not store the plug-in description resource of the target plug-in, the client sends a resource request to the server.
[0219] The resource demand request is used to request the server to provide the plug-in description resource of the target plug-in. The present disclosure does not limit the implementation of the resource demand request. For example, the resource demand request can be implemented using the plug-in download request shown in FIG. 2 .
[0220] Based on the relevant content of step 72 above, it can be seen that for the client deployed on the terminal device, after the client receives the data generation request triggered by the user, if the client determines that the target plug-in corresponding to the data generation request is in an unavailable state and determines that the plug-in description resource of the target plug-in is not stored in the terminal device, it can be determined that the client has received the data generation requirement that is the same or similar to the data generation requirement described by the data generation request for the first time, and thus it can be determined that the client does not have the ability to process the data generation requirement alone. Therefore, the client can directly send a resource requirement request to the server, so that after the server receives the resource requirement request, the server can feedback the plug-in description resource of the target plug-in determined according to the data generation model in the server to the client, so that the client can complete the deployment process for the target plug-in by using the plug-in description resource, so that the client can use the deployed target plug-in to process the data generation requirement that is the same or similar to the data generation requirement described by the data generation request that is subsequently triggered. It should be noted that the present disclosure does not limit the implementation method of the deployment process of the target plug-in. For example, it can include the loading process of the target plug-in and the initialization process of the target plug-in.
[0221] Step 73: The client receives and stores the plug-in description resource fed back by the service client in response to the resource demand request; the plug-in description resource is determined by the server based on the data generation model deployed on the server.
[0222] Based on the relevant contents of steps 71 to 73 above, it can be known that in some application scenarios, for the client deployed on the terminal device above, when the client receives a data generation request, if the client has never processed a data generation requirement that is the same as or similar to the data generation requirement described by the data generation request, then the client not only needs to rely on the server to complete the processing process for the data generation request, but also needs to obtain the relevant resources of the target plug-in corresponding to the data generation request from the server, so that after the target plug-in is deployed on the client or terminal device based on these resources, the client can use the deployed target plug-in to process the subsequently received data generation requirements that are the same as or similar to the data generation requirement described by the data generation request, etc., so that plug-in deployment can be realized on demand, thereby effectively avoiding the waste of resources caused by deploying a large number of plug-ins that users do not need on the terminal device, and thus making more effective use of the limited resources in the terminal device.
[0223] Based on the above data generation method, it can be seen that the present disclosure provides a technical solution for single-end data generation processing based on plug-ins, such as content generation processing, and the technical solution has the advantages shown in (1) to (8) below.
[0224] (1) Regarding the target plug-in mentioned above, it implements data processing by utilizing some lightweight algorithms. This can, to a certain extent, overcome the drawbacks caused by the large data generation model required to build the target plug-in, such as the inability to deploy the target plug-in on the terminal device or the inability to run the model involved in the target plug-in on the terminal device.
[0225] (2) For the target plug-in mentioned above, the model involved in the target plug-in is determined by the server by extracting the core logic from the data generation model deployed by the server. This can effectively overcome the defect caused by the large data generation model required to build the target plug-in.
[0226] (3) For the target plug-in mentioned above, the model involved in the target plug-in is obtained by the server by converting the data generation model deployed by the server from one expression method to another expression method, so that the expression method of the model involved in the target plug-in is more compatible with the operating environment of the terminal device, which can effectively solve the defects caused by the difference between the operating environment of the server and the operating environment of the terminal device. Among them, the former expression method refers to the expression method required to be used by the server when running the data generation model deployed by the server, such as Python. The latter expression method refers to the expression method required to be used by the terminal device when running the model involved in the target plug-in, such as pure C++.
[0227] (IV) For the target plug-in mentioned above, the computationally intensive parts involved in the use of the target plug-in can be implemented using heterogeneous computing methods. This is conducive to fully utilizing the computing resources on the terminal device, such as graphics processing unit (GPU), neural processing unit (NPU), digital signal processing (DSP) and other chip computing resources.
[0228] (5) The present disclosure processes models, such as the AIGC model, into plug-ins so that the model does not occupy the client's installation package space, thereby allowing users to download and dynamically load the model on demand on the client. This can effectively overcome the model size problem. For example, because the model deployed on the server is relatively large, when the model is deployed along with the client's installation package, the installation package size will be relatively large, thereby affecting the user's installation package download experience and the client's update experience.
[0229] (6) The present disclosure avoids memory overflow by splitting a large model into multiple small models connected in series, loading the small model currently needed according to the configuration and destroying it immediately after the small model is run, thereby freeing up sufficient memory for subsequent small model runs, which is conducive to improving memory stability.
[0230] (7) The present disclosure not only adopts a technical solution in which the client uses a plug-in to realize data generation and processing, but also adopts a technical solution in which the client and the server realize data generation and processing through collaboration when the data is processed for the first time or the plug-in is not fully deployed. This can effectively improve the response speed, thereby improving the user experience.
[0231] (8) The present disclosure adopts a basic model and multiple fine-tuning models, such as multiple Lora small models, to realize multiple data generation processes. This can effectively solve the problem that the richness of the data processing process is difficult to expand due to the limited resources of the terminal device, thereby helping to improve the diversity of data generation and processing, and further helping to improve the user experience.
[0232] Based on the data generation method provided in the embodiments of the present disclosure, the embodiments of the present disclosure also provide a data generation device, which will be explained and illustrated below in conjunction with Figure 7. Figure 7 is a schematic diagram of the structure of a data generation device provided in the embodiments of the present disclosure. It should be noted that for the technical details of the data generation device provided in the embodiments of the present disclosure, please refer to the relevant content of the data generation method above.
[0233] As shown in FIG7 , the data generation device 700 provided in an embodiment of the present disclosure includes:
[0234] Receiving unit 701, configured to receive a data generation request;
[0235] The processing unit 702 is used to process the data generation request using a target plug-in to obtain generated data; the target plug-in is constructed based on a data generation model corresponding to the data generation request, and the data generation model is used to process the data generation request.
[0236] In one possible implementation, the data generation request is used to request content generation processing based on reference information; the reference information includes at least one of text and image; the data generation model is a content generation model; and the generated data includes at least one generated image.
[0237] In a possible implementation manner, the client is deployed on a terminal device;
[0238] The processing unit 702 is specifically configured to: send the data generation request to the target plug-in deployed on the terminal device through the plug-in service deployed on the client; and receive generated data fed back by the target plug-in in response to the data generation request.
[0239] In a possible implementation manner, the target plug-in includes a target model determined according to the data generation model; the target plug-in is configured to process the data generation request using the target model to obtain the generated data.
[0240] In one possible implementation, the target plug-in includes at least two candidate models, different candidate models implement different data processing processes, and the at least two candidate models include the target model; the target plug-in is also used to determine the target model corresponding to the data generation request from the at least two candidate models.
[0241] In a possible implementation manner, the data generating device 700 further includes:
[0242] The first responding unit is configured to respond to a client start request, start the client, and load the target plug-in.
[0243] In a possible implementation manner, the data generating device 700 further includes:
[0244] Initialization unit, used to initialize the loaded target plug-in;
[0245] The processing unit 702 is specifically configured to process the data generation request using the initialized target plug-in.
[0246] In a possible implementation manner, the first response unit is specifically configured to: start the client in response to a client start request, and asynchronously preload the target plug-in.
[0247] In one possible implementation, the target plug-in is constructed based on at least one data generation model, different data generation models implement different data generation processes, and the at least one data generation model includes a data generation model corresponding to the data generation request.
[0248] In one possible implementation, the target plug-in includes a base model and at least one fine-tuning model; the base model and the at least one fine-tuning model are determined based on the at least one data generation model;
[0249] The data generating device 700 further includes:
[0250] A second responding unit, configured to respond to a client start request, start the client, and load a basic model in the target plug-in;
[0251] The processing unit 702 is specifically used to: determine the fine-tuning model corresponding to the data generation request from the at least one fine-tuning model; use the fine-tuning model to perform parameter correction processing on the loaded basic model to obtain the target model; and use the target model to process the data generation request.
[0252] In a possible implementation manner, the data generating device 700 further includes:
[0253] A first determining unit, configured to determine whether the target plug-in is in an available state;
[0254] The processing unit 702 is specifically configured to: if it is determined that the target plug-in is in an available state, use the target plug-in to process the data generation request.
[0255] In a possible implementation manner, the data generating device 700 further includes:
[0256] A first sending unit is configured to send the data generation request to a server if it is determined that the target plug-in is in an unavailable state, and the server is configured to process the data generation request using a data generation model deployed on the server to obtain the generated data;
[0257] The first receiving unit is configured to receive the generated data fed back by the server.
[0258] In a possible implementation manner, the client is deployed on a terminal device;
[0259] The data generating device 700 further includes:
[0260] a second determining unit, configured to determine whether a plug-in description resource of the target plug-in is stored in the terminal device if it is determined that the target plug-in is in an unavailable state;
[0261] A second sending unit is configured to send a resource request to the server if it is determined that the plug-in description resource of the target plug-in is not stored in the terminal device;
[0262] The second receiving unit is configured to receive and store the plug-in description resource fed back by the server in response to the resource demand request; the plug-in description resource is determined by the server according to a data generation model deployed on the server.
[0263] In a possible implementation manner, the target plug-in includes at least two data processing modules arranged in sequence, and the at least two data processing modules are obtained by splitting the data generation model;
[0264] The processing unit 702 is specifically configured to:
[0265] Load the i-th data processing module; i is a positive integer, and its initial value is 1;
[0266] Performing data processing using the i-th data processing module;
[0267] After obtaining the output result of the i-th data processing module, releasing the memory occupied by the i-th data processing module;
[0268] Update the i, and continue to execute the step of loading the i-th data processing module until a preset stop condition is reached, and determine the generated data according to the output result of the i-th data processing module.
[0269] In one possible implementation, the splitting process of the data generation model includes: splitting the data generation model into at least two sub-models, and different sub-models implement different data processing functions; for any of the sub-models, if the memory requirement representation data of the sub-model does not exceed the preset memory threshold, then the sub-model is determined as the data processing module; if the memory requirement representation data of the sub-model exceeds the preset memory threshold, then the sub-model is split into at least two model fragments, and each of the model fragments is determined as the data processing module, and the memory requirement representation data of each of the model fragments does not exceed the preset memory threshold.
[0270] In one possible implementation, the process of obtaining the at least two model fragments includes: obtaining at least one candidate split description information, where different candidate split description information describe different model splitting methods; for any of the candidate split description information, splitting the sub-model according to the candidate split description information to obtain the model splitting result corresponding to the candidate split description information, and determining the resource usage characterization data corresponding to the candidate split description information based on the model splitting result corresponding to the candidate split description information, the resource usage characterization data including computing resource reuse characterization data and memory usage characterization data; selecting target split description information from the at least one candidate split description information based on the resource usage characterization data corresponding to each of the candidate split description information, the resource usage balance degree presented by the resource usage characterization data corresponding to the target split description information being higher than the resource usage balance degree presented by the resource usage characterization data corresponding to any other candidate split description information in the at least one candidate split description information except the target split description information; determining the at least two model fragments based on the model splitting result corresponding to the target split description information.
[0271] In a possible implementation, the client includes the data generating device 700 .
[0272] Based on the relevant content of the above-mentioned data generating device 700, it can be known that for the data generating device 700 provided by the embodiment of the present disclosure, after the data generating device 700 receives the data generation request, the data generating device 700 uses the target plug-in to process the data generation request to obtain generated data. In this way, data generation processing can be performed based on the plug-in on the data generating device 700, thereby effectively overcoming the defects that exist when data generation processing is implemented with the help of multi-terminal collaboration, and thus helping to improve the data generation effect. Among them, because the target plug-in is constructed based on the data generation model corresponding to the data generation request, the data processing function of the target plug-in includes the data processing function of the data generation model; and because the data generation model can be used to process the data generation request, the target plug-in can also be used to process the data generation request, so that after the data generation device 700 receives the data generation request, the data generation device 700 can directly use the target plug-in to process the data generation request, thereby making all related processes of the data generation request occur at the same end, so that the data generation processing can be single-ended closed-loop, thereby effectively improving the real-time, security, reliability and other aspects of the data generation process.
[0273] In addition, an embodiment of the present disclosure also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the data generation method provided in the embodiment of the present disclosure.
[0274] Referring to FIG8 , a schematic diagram of the structure of an electronic device 800 suitable for implementing embodiments of the present disclosure is shown. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG8 is merely an example and should not limit the functionality or scope of use of embodiments of the present disclosure.
[0275] As shown in Figure 8, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0276] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although FIG8 shows the electronic device 800 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0277] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0278] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0279] The embodiments of the present disclosure further provide a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes any implementation of the data generation method provided in the embodiments of the present disclosure.
[0280] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0281] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0282] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0283] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device can perform the method.
[0284] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0285] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0286] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit / module does not, in some cases, limit the unit itself.
[0287] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0288] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0289] It should be noted that the various embodiments of this disclosure are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the descriptions of the systems or devices disclosed in the embodiments for similarities and differences between them. Since the systems or devices disclosed in the embodiments correspond to the methods disclosed in the embodiments, their descriptions are relatively simple, and reference can be made to the descriptions of the methods for any related details.
[0290] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0291] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0292] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0293] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present disclosure. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments shown herein, but is intended to be construed in the widest manner consistent with the principles and novel features disclosed herein.
Claims
1. A data generation method, applied to a client, comprising: receiving a data generation request; Processing the data generation request using the target plug-in to obtain generated data; The target plug-in is constructed according to a data generation model corresponding to the data generation request, and the data generation model is used to process the data generation request.
2. The method according to claim 1, wherein: The data generation request is used to request content generation processing based on reference information; the reference information includes at least one of text and image; The data generation model is a content generation model; The generated data includes at least one generated image.
3. The method according to claim 1 or 2, wherein: The client is deployed on a terminal device; The step of using the target plug-in to process the data generation request to obtain generated data includes: Sending the data generation request to a target plug-in deployed on the terminal device through the plug-in service deployed on the client; The generated data fed back by the target plug-in in response to the data generation request is received.
4. The method according to claim 1, wherein: The target plug-in includes a target model determined according to the data generation model; The target plug-in is used to process the data generation request using the target model to obtain the generated data.
5. The method according to claim 4, wherein: The target plug-in includes at least two candidate models, different candidate models implement different data processing processes, and the at least two candidate models include the target model; The target plug-in is further used to determine a target model corresponding to the data generation request from the at least two candidate models.
6. The method according to any one of claims 1 to 5, wherein: Before processing the data generation request by using the target plug-in, the method further includes: In response to a client startup request, the client is started and the target plug-in is loaded.
7. The method according to claim 6, wherein: After loading the target plug-in, the method further includes: Initializing the loaded target plug-in; The using the target plug-in to process the data generation request includes: The data generation request is processed using the initialized target plug-in.
8. The method according to claim 6 or 7, wherein: The loading of the target plug-in comprises: Asynchronously preloads the target plugin.
9. The method according to claim 1, wherein: The target plug-in is constructed based on at least one data generation model, and different data generation models implement different data generation processes. The at least one data generation model includes a data generation model corresponding to the data generation request.
10. The method according to claim 9, wherein: The target plug-in includes a base model and at least one fine-tuning model; the base model and the at least one fine-tuning model are determined according to the at least one data generation model; Before processing the data generation request by using the target plug-in, the method further includes: In response to a client startup request, starting the client and loading a basic model in the target plug-in; The process of obtaining the generated data includes: Determining a fine-tuning model corresponding to the data generation request from the at least one fine-tuning model; Using the fine-tuning model to perform parameter correction processing on the loaded basic model to obtain the target model; The data generation request is processed using the target model.
11. The method according to claim 1, wherein: After receiving the data generation request, the method further includes: Determine whether the target plug-in is in an available state; The using the target plug-in to process the data generation request includes: If it is determined that the target plug-in is in an available state, the data generation request is processed using the target plug-in.
12. The method according to claim 11, wherein: After determining whether the target plug-in is in an available state, the method further includes: If it is determined that the target plug-in is in an unavailable state, the data generation request is sent to a server, and the server is used to process the data generation request using the data generation model deployed on the server to obtain the generated data; Receive the generated data fed back by the server.
13. The method according to claim 11, wherein: The client is deployed on a terminal device; After determining whether the target plug-in is in an available state, the method further includes: If it is determined that the target plug-in is in an unavailable state, determining whether a plug-in description resource of the target plug-in is stored in the terminal device; If it is determined that the plug-in description resource of the target plug-in is not stored in the terminal device, a resource demand request is sent to the server; The plug-in description resource fed back by the server in response to the resource demand request is received and stored; the plug-in description resource is determined by the server according to the data generation model deployed on the server.
14. The method according to claim 1, wherein: The target plug-in includes at least two data processing modules arranged in sequence, and the at least two data processing modules are obtained by splitting the data generation model; The process of obtaining the generated data includes: Load the i-th data processing module; i is a positive integer, and the initial value of i is 1; Using the i-th data processing module to perform data processing; After obtaining the output result of the i-th data processing module, releasing the memory occupied by the i-th data processing module; Update the i, and continue to execute the step of loading the i-th data processing module until a preset stop condition is reached, and determine the generated data according to the output result of the i-th data processing module.
15. The method according to claim 14, wherein: The splitting process of the data generation model includes: Splitting the data generation model into at least two sub-models, wherein different sub-models implement different data processing functions; For any of the sub-models, if the memory requirement characterization data of the sub-model does not exceed the preset memory threshold, the sub-model is determined as the data processing module; if the memory requirement characterization data of the sub-model exceeds the preset memory threshold, the sub-model is split into at least two model fragments, and each of the model fragments is determined as the data processing module, and the memory requirement characterization data of each of the model fragments does not exceed the preset memory threshold.
16. The method according to claim 15, wherein: The process of obtaining the at least two model fragments includes: Obtain at least one candidate split description information, where different candidate split description information describe different model splitting methods; For any of the candidate split description information, the sub-model is split according to the candidate split description information to obtain a model split result corresponding to the candidate split description information, and based on the model split result corresponding to the candidate split description information, resource usage characterization data corresponding to the candidate split description information is determined, wherein the resource usage characterization data includes computing resource reuse characterization data and memory usage characterization data; Selecting target split description information from the at least one candidate split description information according to the resource usage characterization data corresponding to each of the candidate split description information, wherein the resource usage balance degree presented by the resource usage characterization data corresponding to the target split description information is higher than the resource usage balance degree presented by the resource usage characterization data corresponding to any other candidate split description information in the at least one candidate split description information except the target split description information; The at least two model segments are determined according to the model splitting result corresponding to the target splitting description information.
17. A data generating device, comprising: A receiving unit configured to receive a data generation request; The processing unit is configured to use a target plug-in to process the data generation request to obtain generated data, wherein the target plug-in is constructed based on a data generation model corresponding to the data generation request, and the data generation model is used to process the data generation request.
18. An electronic device comprising a processor and a memory, in, The memory is configured to store instructions or computer programs; The processor is configured to execute the instructions or computer programs in the memory so that the electronic device performs the method according to any one of claims 1 to 16.
19. A computer readable medium storing instructions or a computer program, wherein: When the instructions or computer programs are executed on a device, the device is caused to execute the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Service-based SDK calling method and device, electronic equipment and storage medium
CN113377465A
Data entry method and device, intelligent terminal and storage medium
CN115629788A
File processing method and device
CN116126804A
Image generation method and device
CN116883528A
Image processing method, processing apparatus and processing device
WO2019091181A1