Generating method and device of text graph model and hypernetwork

By using the hypernetwork to generate network parameters suitable for the target subject to update the literary and graphic model and fine-tune it based on the target image, the problem that generative artificial intelligence in the prior art is difficult to generate personalized images quickly and efficiently, and a more efficient generation of personalized literary and graphic model is achieved.

CN117475032BActive Publication Date: 2025-06-06BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311388651.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-24
Publication Date
2025-06-06
Estimated Expiration
2043-10-24

AI Technical Summary

Technical Problem

In the prior art, in generative artificial intelligence, it is difficult to quickly and efficiently generate personalized customized images for reference images, and cannot effectively meet the diverse needs of users.

Method used

By acquiring the target image and text data, the first network parameters are obtained using the pre-generated hypernetwork, the initial literary graphic model is updated, and the updated literary graphic model is then fine-tuned based on the target image to generate a personalized literary graphic model associated with the target subject.

Benefits of technology

The parameter scale of processing in the generation process of personalized cultural graph models is reduced, the generation efficiency is improved, and more efficient and reliable conditions are provided for hypernetwork to process images of various subject types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117475032B_ABST
    Figure CN117475032B_ABST
Patent Text Reader

Abstract

The present disclosure proposes a method and device for generating a text graph model and a hypernetwork, which relates to the field of computer technology, especially to the field of artificial intelligence technology such as computer vision, natural language processing, large models, and deep learning, and can be applied to content generation scenarios based on artificial intelligence. The method includes: obtaining a target image and text data containing a target subject; inputting the target image into a pre-generated hypernetwork to obtain a first network parameter output by the hypernetwork; based on the first network parameter, updating the initial text graph model to obtain an updated text graph model; inputting text data into the updated text graph model to obtain a predicted image output by the updated text graph model; based on the difference between the predicted image and the target image, fine-tuning the updated text graph model to obtain a personalized text graph model associated with the target subject.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, in particular to the field of artificial intelligence technology such as computer vision, natural language processing, large models, and deep learning, and specifically to a method and device for generating a text graph model and a hypernetwork. Background Art

[0002] In the application of generative artificial intelligence (AIGC), the key to meeting the diverse needs of users is to quickly customize reference images through text-based graph models to generate results of the reference image subject in different semantic scenarios. Summary of the invention

[0003] The present disclosure aims to solve one of the technical problems in the related art at least to some extent.

[0004] According to a first aspect of the present disclosure, a method for generating a cultural graph model is provided, comprising:

[0005] Acquire target image and text data containing target subject;

[0006] Inputting the target image into a pre-generated hypernetwork to obtain a first network parameter output by the hypernetwork;

[0007] Based on the first network parameters, the initial Wensheng graph model is updated to obtain an updated Wensheng graph model;

[0008] Inputting the text data into the updated text graph model to obtain a predicted image output by the updated text graph model;

[0009] Based on the difference between the predicted image and the target image, the updated text graph model is fine-tuned to obtain a personalized text graph model associated with the target subject.

[0010] According to a second aspect of the present disclosure, a device for generating a cultural graph model is provided, comprising:

[0011] A first acquisition module is used to acquire a target image and text data containing a target subject;

[0012] A second acquisition module is used to input the target image into a pre-generated hypernetwork to obtain a first network parameter output by the hypernetwork;

[0013] A third acquisition module, configured to update the initial Wensheng graph model based on the first network parameter to acquire an updated Wensheng graph model;

[0014] A fourth acquisition module, configured to input the text data into the updated text graph model, and acquire a predicted image output by the updated text graph model;

[0015] A fifth acquisition module is used to fine-tune the updated Wensheng graph model based on the difference between the predicted image and the target image, so as to acquire a personalized Wensheng graph model associated with the target subject.

[0016] According to a third aspect of the present disclosure, a method for generating a hypernetwork is provided, comprising:

[0017] Acquire a training data set, wherein the training data set includes a plurality of sample images and target network parameters to be generated, wherein the plurality of sample images include subject types;

[0018] Inputting the sample image into an initial hypernetwork to obtain a third network parameter output by the initial hypernetwork;

[0019] Based on the difference between the third network parameter and the target network parameter, the initial super network is modified until a super network associated with the subject type is obtained.

[0020] According to a fourth aspect of the present disclosure, a device for generating a hypernetwork is provided, comprising:

[0021] A sixth acquisition module is used to acquire a training data set, wherein the training data set includes a plurality of sample images and target network parameters to be generated, wherein the plurality of sample images include subject types;

[0022] A seventh acquisition module, used for inputting the sample image into an initial super network to obtain a third network parameter output by the initial super network;

[0023] an eighth acquisition module, configured to modify the initial super network based on the difference between the third network parameter and the target network parameter until a super network associated with the subject type is acquired,

[0024] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0025] at least one processor; and

[0026] a memory communicatively connected to the at least one processor; wherein,

[0027] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for generating a cultural graph model as described in the first aspect, or execute the method for generating a hypernetwork as described in the third aspect.

[0028] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method for generating a text graph model as described in the first aspect, or to execute the method for generating a hypernetwork as described in the third aspect.

[0029] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement the method for generating a text graph model as described in the first aspect, or perform the steps of the method for generating a hypernetwork as described in the third aspect.

[0030] The text graph model, the method and the device for generating a hypernetwork provided by the present disclosure have the following beneficial effects:

[0031] In the present disclosure, a first network parameter suitable for a target subject is first generated by using a hypernetwork to update an initial Wensheng graph model, and then the updated Wensheng graph model is fine-tuned based on a target image to generate a personalized Wensheng graph model associated with the target subject, thereby reducing the scale of parameters processed during the generation of the personalized Wensheng graph model and improving the generation efficiency of the personalized Wensheng graph model. In addition, by using a training data set associated with a subject type, the network parameters output by the initial hypernetwork for a sample image and the difference between the network parameters and the target network parameters are obtained to iteratively correct the initial hypernetwork to obtain a hypernetwork associated with the subject type, thereby improving the efficiency and reliability of the hypernetwork in processing images of various subject types and providing conditions for generating a personalized Wensheng graph model.

[0032] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The above and / or additional aspects and advantages of the present disclosure will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, which are used to better understand the present solution and do not constitute a limitation of the present disclosure, wherein:

[0034] Figure 1 is a flow chart of a method for generating a text graph model according to an embodiment of the present disclosure;

[0035] Figure 2 is a flow chart of a method for generating a cultural graph model according to another embodiment of the present disclosure;

[0036] Figure 3 is a flow chart of a method for generating a hypernetwork according to an embodiment of the present disclosure;

[0037] Figure 4 is a schematic diagram of the structure of a device for generating a text graph model according to an embodiment of the present disclosure;

[0038] Figure 5 is a schematic diagram of the structure of a device for generating a super network according to an embodiment of the present disclosure;

[0039] Figure 6 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0040] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0041] The embodiments of the present disclosure relate to the fields of artificial intelligence technology such as computer vision, natural language processing, large models, and deep learning.

[0042] Artificial Intelligence (AI) is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence.

[0043] Computer vision refers to the use of cameras and computers to replace the human eye to identify, track and measure targets, and further perform graphic processing so that the computer processing becomes an image that is more suitable for human eye observation or transmission to instrument detection.

[0044] Natural Language Processing (NLP) is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language.

[0045] Large Language Model (LLM), or big model, refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. Large language models can handle a variety of natural language tasks, such as text classification, question answering, dialogue, etc., and are an important path to artificial intelligence.

[0046] Deep learning is the process of learning the inherent laws and representation levels of sample data. The information obtained in the learning process is of great help in interpreting data such as text, images, and sounds. The ultimate goal of deep learning is to enable machines to have analytical learning capabilities like humans and to be able to recognize data such as text, images, and sounds.

[0047] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0048] The following describes a text graph model, a method and apparatus for generating a hypernetwork, an electronic device, and a storage medium according to an embodiment of the present disclosure with reference to the accompanying drawings.

[0049] It should be noted that the execution subject of the method for generating the text graph model and the hypernetwork of this embodiment is a device for generating the text graph model and the hypernetwork, which can be implemented by software and / or hardware, and can be configured in an electronic device, which can include but is not limited to a terminal, a server, etc. In the embodiment of the present disclosure, the example of configuring the device for generating the text graph model and the hypernetwork in the system for generating the text graph model and the hypernetwork is described.

[0050] Figure 1 It is a flowchart of a method for generating a text graph model proposed according to an embodiment of the present disclosure.

[0051] like Figure 1 As shown, the method for generating the text graph model includes:

[0052] S101: Acquire a target image and text data including a target subject.

[0053] The target subject is a subject that needs to be included in the image generated according to the text graph model to be generated. For example, it can be a specific face, a car, a building, a plant, etc., and the present disclosure does not limit this.

[0054] In the disclosed embodiment, a target image containing a target subject can be obtained through a network, and text data corresponding to the target subject can be determined. For example, when the target subject is a face, the text data can be "a face"; or, when the target subject is a vehicle, the corresponding text data can be "a car", etc.

[0055] S102: Input the target image into a pre-generated hypernetwork to obtain a first network parameter output by the hypernetwork.

[0056] The pre-generated hypernetwork refers to a hypernetwork that has been trained in advance using a data set. The first network parameter is a network parameter of a part of the network layer of the Wensheng graph model predicted by the hypernetwork when generating an image of the target subject.

[0057] Optionally, the target image can be first input into the encoder in the hypernetwork to obtain the first image feature corresponding to the target image. Then, the first image feature is spliced ​​with the preset first tensor and input into the decoder of the hypernetwork to obtain the first network parameter output by the decoder. Thus, the output first network parameter can better reflect the characteristics of the target subject, providing conditions for improving the personalization of the Wensheng graph model.

[0058] Among them, the encoder can be a series of stacked convolutional layers, which is used to extract the graphic features of the subject in the image.

[0059] The first image feature may be a feature tensor including image structure, color, etc. The preset first tensor may be a zero-initialized weight layer, and the dimension of the first tensor is the same as the dimension of the first network parameter.

[0060] It should be noted that, when the decoder outputs the first network parameters, it can obtain more accurate and reliable first network parameters by performing multiple feedback iterations on the network parameters output each time until a preset number of iterations or other hyperparameters are met.

[0061] Optionally, when the dimension of the first image feature is different from the dimension of the first tensor, the first image feature may be linearly transformed to obtain a second image feature having the same dimension as the first tensor. The second image feature may then be concatenated with the first tensor and input into a decoder in the hypernetwork to obtain the first network parameter output by the decoder.

[0062] In the disclosed embodiment, a first image feature with different dimensions is mapped to a second image feature with the same dimension through a linear transformation layer, and then the second image feature is concatenated with the first tensor and input into a decoder to obtain the first network parameters, thereby solving the problem that the image feature and the first tensor cannot be concatenated, improving the generalization ability of the decoder of the hypernetwork to obtain the first network parameters, and improving the performance of the hypernetwork.

[0063] Optionally, based on the type of the target subject, multiple hypernetworks for different types of target subjects can be pre-trained, so that the personalized features of different types of subjects can be better extracted, and the network parameters suitable for different types of subjects can be more accurately determined, thereby providing conditions for improving the generation efficiency of personalized text graph models. Therefore, a pre-generated target hypernetwork associated with the type of the target subject can be obtained first, and then the target image can be input into the target hypernetwork to obtain the first network parameter output by the target hypernetwork.

[0064] S103: Based on the first network parameter, the initial text graph model is updated to obtain an updated text graph model.

[0065] In the disclosed embodiment, the system for generating the Wensheng graph model can directly use the first network parameters to replace the corresponding initial network parameters in the initial Wensheng graph model to obtain an updated Wensheng graph model. Thus, the first network parameters corresponding to the target image are directly used to replace the corresponding initial network parameters in the Wensheng graph model, so that the parameters of the updated Wensheng graph model are more suitable for the target subject, reducing the difficulty and complexity of updating the network parameters for generating a personalized Wensheng graph model suitable for the target subject.

[0066] It should be noted that the corresponding network parameters in the initial Wensheng graph model may be parameters in some key layer networks in the Wensheng graph model, such as the cross attention layer and the self attention layer, etc., and the present disclosure does not limit this.

[0067] Alternatively, the first network parameter may be added to the corresponding initial network parameter in the initial text graph model to update the initial text graph model to obtain an updated text graph model. For example, the first network parameter may be weighted fused with the initial network parameter, and the fused network parameter may be used to update the initial text graph model.

[0068] S104: Input text data into the updated text-graph model to obtain a predicted image output by the updated text-graph model.

[0069] In the disclosed embodiment, after the text graph model is updated using the first network parameter, the text graph model can generate personalized results related to the target subject according to the text data, but there may still be a situation where the similarity between the generated image and the target image does not meet the requirements.

[0070] Therefore, it is also necessary to input the text data corresponding to the target image into the updated text graph model, obtain the predicted image output by the updated text graph model, determine the difference between the predicted image and the target image, and further optimize the updated text graph model to further improve the accuracy and robustness of the text graph model.

[0071] S105: Based on the difference between the predicted image and the target image, fine-tune the updated text graph model to obtain a personalized text graph model associated with the target subject.

[0072] In the disclosed embodiment, the system for generating the Wensheng graph model can calculate the loss between the predicted image and the target image, and then fine-tune the updated Wensheng graph model based on the loss value, so as to obtain a personalized Wensheng graph model associated with the target subject. Since some parameters in the updated Wensheng graph model are generated by the hypernetwork based on the target image containing the target subject, the complexity and time consumption of fine-tuning the updated Wensheng graph model are greatly reduced.

[0073] In this embodiment, the system for generating a text graph model first obtains a target image and text data containing a target subject, then inputs the target image into a pre-generated hypernetwork to obtain a first network parameter output by the hypernetwork, then updates the initial text graph model based on the first network parameter to obtain an updated text graph model, then inputs the text data into the updated text graph model to obtain a predicted image output by the updated text graph model, and fine-tunes the updated text graph model based on the difference between the predicted image and the target image to obtain a personalized text graph model associated with the target subject. Thus, by first generating a first network parameter suitable for the target subject using the hypernetwork, the initial text graph model is updated, and then fine-tuning the updated text graph model based on the target image to generate a personalized text graph model associated with the target subject, thereby reducing the scale of parameters processed in the generation process of the personalized text graph model and improving the generation efficiency of the personalized text graph model.

[0074] Figure 2 It is a flowchart of a method for generating a text graph model proposed in another embodiment of the present disclosure.

[0075] like Figure 2 As shown, the method for generating the text graph model includes:

[0076] S201: Acquire a target image and text data including a target subject.

[0077] S202: Input the target image into a pre-generated hypernetwork to obtain a first network parameter output by the hypernetwork.

[0078] The description of S201 and S202 can be specifically referred to the above embodiment, which will not be repeated here.

[0079] S203: Perform weighted fusion on the first network parameter and the corresponding initial network parameter in the initial text graph model to obtain fused second network parameters.

[0080] In the disclosed embodiment, a weight value pair corresponding to the first network parameter and the initial network parameter in the initial text graph model may be obtained first, and then weighted fusion may be performed based on the weight value pair to obtain a fused second network parameter.

[0081] Optionally, the first weight value corresponding to the first network parameter and the second weight value corresponding to the initial network parameter may be determined according to the type of the target subject, and then the first network parameter and the corresponding initial network parameter in the initial text graph model are weightedly fused based on the first weight value and the second weight value.

[0082] The type of the target subject may include a face image, a vehicle image, etc. The first weight value is used to control the influence of the image feature corresponding to the target subject in the process of updating the Wensheng graph model, and the second weight value is used to control the update degree of the Wensheng graph model.

[0083] It should be noted that the distribution of the first weight value and the second weight value corresponding to different types of subjects can be determined according to actual needs, and may be the same or different. For example, for a subject type with more details (such as a face image, etc.), the first weight value corresponding to the first network parameter may be higher, and the second weight value corresponding to the initial network parameter may be lower; or, for a subject type with fewer details (such as cups, cars and other objects), the first weight value corresponding to the first network parameter may be lower, and the second weight value corresponding to the initial network parameter may be higher. This disclosure does not limit this.

[0084] In the disclosed embodiment, the system for generating the Wensheng graph model can respectively calculate the product of the first network parameter and the first weight value, and the product of the initial network parameter and the second weight value, and then add the two products to obtain the fused second network parameter. Thus, the relative importance of the two network parameters can be adjusted according to the type of the target subject, making the fused second network parameter more reliable, and providing conditions for improving the personalization of the updated Wensheng graph model.

[0085] S204: Using the fused second network parameters, the corresponding initial network parameters in the initial Wensheng graph model are replaced to obtain the Wensheng graph model after the network parameters are replaced.

[0086] In the disclosed embodiment, the second network parameters obtained by weighted fusion of the first network parameters and the corresponding initial network parameters in the initial text graph model are used to replace the corresponding initial network parameters in the initial text graph model, so as to complete the update of the text graph model, thereby further improving the applicability and personalization of the text graph model.

[0087] S205: Input the text data into the updated text-graph model, and obtain the predicted image output by the updated text-graph model.

[0088] S206: Based on the difference between the predicted image and the target image, fine-tune the updated text graph model to obtain a personalized text graph model associated with the target subject.

[0089] The description of S205 and S206 can be specifically referred to the above embodiment, which will not be repeated here.

[0090] In this embodiment, after obtaining the first network parameters output by the hypernetwork, the system for generating the Wensheng graph model first performs weighted fusion on the first network parameters with the corresponding initial network parameters in the initial Wensheng graph model to obtain fused second network parameters, and then uses the fused second network parameters to replace the corresponding initial network parameters in the initial Wensheng graph model to obtain the Wensheng graph model after network parameter replacement, so that the Wensheng graph model can be updated by combining the first network parameters with the corresponding initial network parameters in the initial Wensheng graph model, thereby further improving the personalization of the generated Wensheng graph model.

[0091] Figure 3 It is a flowchart of a method for generating a hypernetwork proposed in an embodiment of the present disclosure.

[0092] like Figure 3 As shown, the method for generating the hypernetwork includes:

[0093] S301: Obtain a training data set.

[0094] The training data set includes multiple sample images and target network parameters to be generated. In addition, the multiple sample images include subject types.

[0095] In the disclosed embodiment, the hypernetwork generation system can obtain relevant public training data sets according to the target subject type of the text graph model to be generated, so as to train a hypernetwork that can extract features from a target image containing the target subject.

[0096] S302: Input the sample image into the initial hypernetwork to obtain a third network parameter output by the initial hypernetwork.

[0097] In the disclosed embodiment, the initial super network includes an initial encoder and a decoder. The initial encoder can be used to extract image features from a sample image, and then the third network parameters can be obtained based on the image features and the initial decoder.

[0098] It should be noted that in order to avoid overfitting of the sample graph when the hypernetwork obtains the third network parameters, the target network parameters to be generated in the training data set can be used to perform regularization supervision on the third network parameters output by the hypernetwork.

[0099] Optionally, the sample image can be input into the initial encoder in the initial hypernetwork to obtain a third image feature output by the initial encoder, and then, when the dimension of the third image feature is the same as that of the target network parameter, the third image feature is concatenated with the preset second tensor and input into the initial decoder in the initial hypernetwork to obtain a third network parameter output by the initial decoder.

[0100] The dimension of the second tensor is the same as the dimension of the target network parameter, and the second tensor can be the same as the first tensor.

[0101] In the disclosed embodiment, by judging whether the third image feature output by the initial encoder is the same as the dimension of the target network parameter, if the third image feature is the same, the third network parameter is obtained by concatenation with the tensor and decoding. This can avoid overfitting of the sample image by the hypernetwork, making the third network parameter output by the initial hypernetwork more reliable and accurate, and providing conditions for improving the correction efficiency of the initial hypernetwork.

[0102] Optionally, when the dimensions of the third image feature and the target network parameter are different, the third image feature can be transformed first to obtain the fourth image feature, and then the fourth image feature is concatenated with the second tensor and input into the initial decoder in the initial hypernetwork to obtain the third network parameter output by the initial decoder.

[0103] Among them, the dimension of the fourth image feature is the same as the dimension of the target network parameter.

[0104] In the disclosed embodiment, by judging whether the third image feature output by the initial encoder is the same as the dimension of the target network parameter, if it is not the same, the third image feature is mapped to the fourth image feature of the same dimension as the target network parameter through linear mapping, and then concatenated with the tensor, and decoded to obtain the third network parameter. This can solve the problem that the extracted image feature is different in dimension from the target network parameter, improve the generalization ability of obtaining network parameters, and provide conditions for improving the correction efficiency of the initial super network.

[0105] S303: Based on the difference between the third network parameters and the target network parameters, the initial super network is modified until a super network associated with the subject type is obtained.

[0106] In the disclosed embodiment, the hypernetwork generation system can calculate the loss between the third network parameters and the target network parameters, and then use the loss to correct the initial hypernetwork. After that, the corrected hypernetwork can be used to recalculate the third network parameters, and then the loss with the target network parameters can be calculated... Through multiple feedback iterations, until the difference between the output third network parameters and the target network parameters is less than the empirical value, or greater than the number of iterations, etc., a hypernetwork associated with the subject type can be obtained.

[0107] In this embodiment, the hypernetwork generation system first obtains a training data set, and then inputs a sample image in the training data set into the initial hypernetwork to obtain a third network parameter output by the initial hypernetwork, and then modifies the initial hypernetwork based on the difference between the third network parameter and the target network parameter until a hypernetwork associated with the subject type is obtained. Thus, by using the training data set associated with the subject type, obtaining the network parameters output by the initial hypernetwork for the sample image, and the difference between the network parameters and the target network parameters, the initial hypernetwork is iteratively modified to obtain a hypernetwork associated with the subject type, which can improve the efficiency and reliability of the hypernetwork in processing images of each subject type, and provide conditions for generating a personalized Wensheng graph model.

[0108] Figure 4 It is a structural schematic diagram of a device for generating a text graph model proposed in an embodiment of the present disclosure.

[0109] like Figure 4 As shown, the device 400 for generating the text graph model includes:

[0110] The first acquisition module 401 is used to acquire a target image and text data containing a target subject;

[0111] A second acquisition module 402 is used to input the target image into a pre-generated hypernetwork to obtain a first network parameter output by the hypernetwork;

[0112] A third acquisition module 403 is used to update the initial Wensheng graph model based on the first network parameter to obtain an updated Wensheng graph model;

[0113] A fourth acquisition module 404 is used to input text data into the updated text graph model to obtain a predicted image output by the updated text graph model;

[0114] The fifth acquisition module 405 is used to fine-tune the updated text graph model based on the difference between the predicted image and the target image, and acquire a personalized text graph model associated with the target subject.

[0115] In some embodiments, the second acquisition module 402 is specifically configured to:

[0116] Inputting the target image into an encoder in the hypernetwork to obtain a first image feature corresponding to the target image;

[0117] The first image feature is concatenated with a preset first tensor and input into a decoder in the hypernetwork to obtain a first network parameter output by the decoder, wherein the dimension of the first tensor is the same as the dimension of the first network parameter.

[0118] In some embodiments, the second acquisition module 402 is specifically configured to:

[0119] When the dimension of the first image feature is different from the dimension of the first tensor, linearly transform the image feature to obtain a second image feature with the same dimension as the first tensor;

[0120] The second image feature is concatenated with the first tensor and input into a decoder in the hypernetwork to obtain the first network parameter output by the decoder.

[0121] In some embodiments, the third acquisition module 403 is specifically used to:

[0122] The first network parameters are used to replace the corresponding initial network parameters in the initial Wensheng graph model to obtain the Wensheng graph model after the network parameters are replaced.

[0123] In some embodiments, the third acquisition module 403 is specifically used to:

[0124] Performing weighted fusion of the first network parameter and the corresponding initial network parameter in the initial Wensheng graph model to obtain a fused second network parameter;

[0125] The fused second network parameters are used to replace the corresponding initial network parameters in the initial Wensheng graph model to obtain the Wensheng graph model after the network parameters are replaced.

[0126] In some embodiments, the third acquisition module 403 is specifically used to:

[0127] Determine the first weight corresponding to the first network parameter according to the type of the target subject.

[0128] The second weight value corresponding to the weight value and the initial network parameter;

[0129] Based on the first weight value and the second weight value, the first network parameter is weightedly fused with the corresponding initial network parameter in the initial text graph model.

[0130] In some embodiments, the second acquisition module 402 is specifically configured to:

[0131] Obtaining a pre-generated target hypernetwork associated with a type of a target subject;

[0132] The target image is input into the target hypernetwork to obtain the first network parameter output by the target hypernetwork.

[0133] It should be noted that the above explanation of the method for generating a text graph model is also applicable to the device for generating a text graph model of this embodiment, and will not be repeated here.

[0134] In this embodiment, the system for generating a text graph model first obtains a target image and text data containing a target subject, then inputs the target image into a pre-generated hypernetwork to obtain a first network parameter output by the hypernetwork, then updates the initial text graph model based on the first network parameter to obtain an updated text graph model, then inputs the text data into the updated text graph model to obtain a predicted image output by the updated text graph model, and fine-tunes the updated text graph model based on the difference between the predicted image and the target image to obtain a personalized text graph model associated with the target subject. Thus, by first generating a first network parameter suitable for the target subject using the hypernetwork, the initial text graph model is updated, and then fine-tuning the updated text graph model based on the target image to generate a personalized text graph model associated with the target subject, thereby reducing the scale of parameters processed in the generation process of the personalized text graph model and improving the generation efficiency of the personalized text graph model.

[0135] Figure 5 It is a structural diagram of a super network generation device proposed in an embodiment of the present disclosure.

[0136] like Figure 5 As shown, the super network generation device 500 includes:

[0137] The sixth acquisition module 501 is used to acquire a training data set, wherein the training data set includes a plurality of sample images and target network parameters to be generated, wherein the plurality of sample images contain subjects of the same type;

[0138] A seventh acquisition module 502 is used to input the sample image into the initial super network to obtain a third network parameter output by the initial super network;

[0139] The eighth acquisition module 503 is used to modify the initial super network based on the difference between the third network parameter and the target network parameter until a super network associated with the subject type is acquired.

[0140] In some embodiments, the seventh acquisition module 502 is specifically configured to:

[0141] Input the sample image into the initial encoder in the initial hypernetwork to obtain the third image feature output by the initial encoder;

[0142] When the dimension of the third image feature is the same as that of the target network parameter, the third image feature is concatenated with the preset second tensor and input into the initial decoder in the initial super network to obtain the third network parameter output by the initial decoder, wherein the dimension of the second tensor is the same as that of the target network parameter.

[0143] In some embodiments, the seventh acquisition module is further used to:

[0144] When the dimensions of the third image feature and the target network parameter are different, transforming the third image feature to obtain a fourth image feature, wherein the dimension of the fourth image feature is the same as the dimension of the target network parameter;

[0145] The fourth image feature is concatenated with the second tensor and input into an initial decoder in the initial hypernetwork to obtain a third network parameter output by the initial decoder.

[0146] It should be noted that the above explanation of the method for generating a hypernetwork is also applicable to the apparatus for generating a hypernetwork in this embodiment, and will not be repeated here.

[0147] In this embodiment, the hypernetwork generation system first obtains a training data set, and then inputs a sample image in the training data set into the initial hypernetwork to obtain a third network parameter output by the initial hypernetwork, and then modifies the initial hypernetwork based on the difference between the third network parameter and the target network parameter until a hypernetwork associated with the subject type is obtained. Thus, by using the training data set associated with the subject type, obtaining the network parameters output by the initial hypernetwork for the sample image, and the difference between the network parameters and the target network parameters, the initial hypernetwork is iteratively modified to obtain a hypernetwork associated with the subject type, which can improve the efficiency and reliability of the hypernetwork in processing images of each subject type, and provide conditions for generating a personalized Wensheng graph model.

[0148] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0149] Figure 6 A schematic block diagram of an example electronic device 500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0150] like Figure 6As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0151] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0152] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as the method for generating a Wensheng graph model and a hypernetwork. For example, in some embodiments, the Wensheng graph model and the method for generating a hypernetwork may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the Wensheng graph model and the method for generating a hypernetwork described above may be executed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the method for generating the text graph model and the hypernetwork in any other appropriate manner (for example, by means of firmware).

[0153] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0154] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0155] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0156] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0157] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0158] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.

[0159] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0160] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In the description of the present disclosure, the words "if" and "if" used can be interpreted as "at the time of" or "when" or "in response to determination" or "under the circumstances of".

[0161] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for generating a text graph model, include: Acquire target image and text data containing target subject; Inputting the target image into a pre-generated hypernetwork to obtain a first network parameter output by the hypernetwork; Based on the first network parameters, the initial Wensheng graph model is updated to obtain an updated Wensheng graph model; Inputting the text data into the updated text graph model to obtain a predicted image output by the updated text graph model; Based on the difference between the predicted image and the target image, fine-tuning the updated Wensheng graph model to obtain a personalized Wensheng graph model associated with the target subject; The super network is generated in the following way: The sample image is input into the initial hypernetwork to obtain a third network parameter, and the initial hypernetwork is modified according to the difference between the third network parameter and the target network parameter to obtain a hypernetwork associated with the subject type of the sample image.

2. The method according to claim 1, in, Inputting the target image into a pre-generated hypernetwork and obtaining a first network parameter output by the hypernetwork comprises: Inputting the target image into an encoder in the hypernetwork to obtain a first image feature corresponding to the target image; The first image feature is concatenated with a reference tensor and input into a decoder in the super network to obtain a first network parameter output by the decoder, wherein the dimension of the reference tensor is the same as the dimension of the first network parameter.

3. The method according to claim 2, in, The step of splicing the first image feature with the preset first tensor and inputting the resultant information into a decoder in the hypernetwork to obtain a first network parameter output by the decoder includes: When the dimension of the first image feature is different from the dimension of the first tensor, linearly transform the image feature to obtain a second image feature with the same dimension as the first tensor; The second image feature is concatenated with the reference tensor and then input into a decoder in the hypernetwork to obtain a first network parameter output by the decoder.

4. The method according to claim 1, in, The updating of the initial text graph model based on the network parameters to obtain an updated text graph model includes: The first network parameters are used to replace the corresponding initial network parameters in the initial Wensheng graph model to obtain the Wensheng graph model after the network parameters are replaced.

5. The method according to claim 1, in, The updating of the initial text graph model based on the network parameters to obtain an updated text graph model includes: Performing weighted fusion of the first network parameter and the corresponding initial network parameter in the initial text graph model to obtain fused second network parameters; The fused second network parameters are used to replace the corresponding initial network parameters in the initial Wensheng graph model to obtain the Wensheng graph model after the network parameters are replaced.

6. The method according to claim 5, in, The first network parameter is weightedly fused with the corresponding initial network parameter in the initial text graph model, including: Determining, according to the type of the target subject, a first weight value corresponding to the first network parameter and a second weight value corresponding to the initial network parameter; Based on the first weight value and the second weight value, the first network parameter is weightedly fused with the corresponding initial network parameter in the initial text graph model.

7. The method according to claim 1, in, The step of inputting the target image into a pre-generated hypernetwork to obtain a first network parameter output by the hypernetwork includes: obtaining a pre-generated target hypernetwork associated with the type of the target subject; The target image is input into the target hypernetwork to obtain a first network parameter output by the target hypernetwork.

8. The method according to any one of claims 1 to 7, in, The method for generating the hypernetwork specifically includes: Acquire a training data set, wherein the training data set includes a plurality of sample images and target network parameters to be generated, wherein the subject types included in the plurality of sample images are the same; Inputting the sample image into an initial hypernetwork to obtain a third network parameter output by the initial hypernetwork; Based on the difference between the third network parameter and the target network parameter, the initial super network is modified until a super network associated with the subject type is obtained.

9. The method according to claim 8, in, The step of inputting the sample image into an initial hypernetwork to obtain a third network parameter output by the initial hypernetwork includes: Inputting the sample image into an initial encoder in an initial hypernetwork to obtain a third image feature output by the initial encoder; When the dimension of the third image feature is the same as that of the target network parameter, the third image feature is concatenated with a preset second tensor and input into an initial decoder in the initial super network to obtain the third network parameter output by the initial decoder, wherein the dimension of the second tensor is the same as that of the target network parameter.

10. The method according to claim 9, in, After obtaining the third image feature output by the initial encoder, the method further includes: In a case where the dimensions of the third image feature and the target network parameter are different, transforming the third image feature to obtain a fourth image feature, wherein the dimension of the fourth image feature is the same as the dimension of the target network parameter; The fourth image feature is concatenated with the second tensor and input into an initial decoder in the initial hypernetwork to obtain a third network parameter output by the initial decoder.

11. A device for generating a text graph model, include: A first acquisition module is used to acquire a target image and text data containing a target subject; A second acquisition module is used to input the target image into a pre-generated hypernetwork to obtain a first network parameter output by the hypernetwork; A third acquisition module, configured to update the initial Wensheng graph model based on the first network parameter to acquire an updated Wensheng graph model; A fourth acquisition module, configured to input the text data into the updated text graph model, and acquire a predicted image output by the updated text graph model; a fifth acquisition module, configured to fine-tune the updated Wensheng graph model based on a difference between the predicted image and the target image, and acquire a personalized Wensheng graph model associated with the target subject; The super network is generated in the following way: The sample image is input into the initial hypernetwork to obtain a third network parameter, and the initial hypernetwork is modified according to the difference between the third network parameter and the target network parameter to obtain a hypernetwork associated with the subject type of the sample image.

12. The device according to claim 11, in, The second acquisition module is specifically used for: Inputting the target image into an encoder in the hypernetwork to obtain a first image feature corresponding to the target image; The first image feature is concatenated with a preset first tensor and input into a decoder in the hypernetwork to obtain a first network parameter output by the decoder, wherein the dimension of the first tensor is the same as the dimension of the first network parameter.

13. The device according to claim 12, in, The second acquisition module is specifically used for: When the dimension of the first image feature is different from the dimension of the first tensor, linearly transform the image feature to obtain a second image feature with the same dimension as the first tensor; The second image feature is concatenated with the first tensor and input into a decoder in the hypernetwork to obtain a first network parameter output by the decoder.

14. The device according to claim 11, in, The third acquisition module is specifically used for: The first network parameters are used to replace the corresponding initial network parameters in the initial Wensheng graph model to obtain the Wensheng graph model after the network parameters are replaced.

15. The device according to claim 11, in, The third acquisition module is specifically used for: Performing weighted fusion of the first network parameter and the corresponding initial network parameter in the initial text graph model to obtain fused second network parameters; The fused second network parameters are used to replace the corresponding initial network parameters in the initial Wensheng graph model to obtain the Wensheng graph model after the network parameters are replaced.

16. The device according to claim 15, in, The third acquisition module is specifically used for: Determining, according to the type of the target subject, a first weight value corresponding to the first network parameter and a second weight value corresponding to the initial network parameter; Based on the first weight value and the second weight value, the first network parameter is weightedly fused with the corresponding initial network parameter in the initial text graph model.

17. The device according to claim 11, in, The second acquisition module is specifically used for: obtaining a pre-generated target hypernetwork associated with the type of the target subject; The target image is input into the target hypernetwork to obtain a first network parameter output by the target hypernetwork.

18. The device according to any one of claims 11 to 17, further comprising: include: A sixth acquisition module, used to acquire a training data set, wherein the training data set includes a plurality of sample images and target network parameters to be generated, wherein the subject types included in the plurality of sample images are the same; A seventh acquisition module, used for inputting the sample image into an initial super network to obtain a third network parameter output by the initial super network; An eighth acquisition module is used to modify the initial super network based on the difference between the third network parameter and the target network parameter until a super network associated with the subject type is acquired.

19. The device according to claim 18, in, The seventh acquisition module is specifically used for: Inputting the sample image into an initial encoder in an initial hypernetwork to obtain a third image feature output by the initial encoder; When the dimension of the third image feature is the same as that of the target network parameter, the third image feature is concatenated with a preset second tensor and input into an initial decoder in the initial super network to obtain the third network parameter output by the initial decoder, wherein the dimension of the second tensor is the same as that of the target network parameter.

20. The device according to claim 19, in, The seventh acquisition module is further used for: In a case where the dimensions of the third image feature and the target network parameter are different, transforming the third image feature to obtain a fourth image feature, wherein the dimension of the fourth image feature is the same as the dimension of the target network parameter; The fourth image feature is concatenated with the second tensor and input into an initial decoder in the initial hypernetwork to obtain a third network parameter output by the initial decoder.

21. An electronic device, include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for generating a culture graph model according to any one of claims 1 to 10.

22. A non-transitory computer-readable storage medium storing computer instructions, It is characterized in that in, The computer instructions are used to enable the computer to execute the method for generating a Wensheng graph model according to any one of claims 1 to 10.

23. A computer program product, It is characterized in that The invention comprises a computer program, which, when executed by a processor, implements the steps of the method for generating a text graph model according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Knowledge distillation-based vehicle part segmentation method and related equipment

    CN115063589A