Instruction information storage method, device and medium based on artificial intelligence speech model
By acquiring object noise reduction intensity information and generating a mask image set, and using a large language model to generate and store image redrawing instruction information, the problem of low efficiency in manual customization is solved, and efficient and high-quality image redrawing instruction generation and storage are achieved.
Patent Information
- Application Number
- CN202510456353.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Customizing image redrawing instructions manually on intelligent platforms is inefficient, leading to wasted computing resources and manpower. Existing technologies struggle to efficiently generate high-quality image redrawing instructions.
By acquiring the object denoising intensity information set of the target object information set, a mask image set is generated, and image redrawing instruction information is generated using a pre-trained large language model and stored on the server side of the intelligent platform.
It enables efficient and accurate generation of image redrawing instruction information, improving the efficiency and quality of image redrawing while reducing computing resource consumption and labor costs.
Smart Images

Figure CN120407834B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the field of large language models supporting voice processing, in particular to an instruction information storage method, device and medium based on an artificial intelligence voice model. BACKGROUND
[0002] At present, images have become an effective form of information display. With the continuous upgrading of hardware devices, the quality requirements for images are also getting higher and higher. For the generation of image redrawing instruction information of a target image, the commonly used way is: the image uploading object manually edits according to the image requirements of the required redrawing image to obtain the image redrawing instruction information.
[0003] However, when the above-mentioned way is used to generate image redrawing instruction information, the following technical problems often exist:
[0004] The efficiency of manually customizing image redrawing instruction information on a school intelligent platform is low. It needs to constantly repeat trial and error of image redrawing instruction information, which leads to occupation of large language models, so that the computing resources are always occupied and a lot of human cost is wasted.
[0005] The above information disclosed in the background section is only intended to enhance the understanding of the background of the present inventive concept, and therefore, it can include information that does not form the prior art known to those of ordinary skill in the art in the country. SUMMARY
[0006] The summary section of the present disclosure is used to introduce the concepts in a brief form, which will be described in detail in the specific embodiments section. The summary section of the present disclosure is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0007] Some embodiments of the present disclosure propose an instruction information storage method, device and medium based on an artificial intelligence voice model to solve one or more of the technical problems mentioned in the background section.
[0008] In a first aspect, some embodiments of the present disclosure provide an instruction information storage method based on an artificial intelligence speech model, comprising: obtaining an object noise reduction intensity information set corresponding to a target object information set set on a target school intelligent platform; in response to receiving image redrawing instruction generation information for a target image, obtaining image description information and an image object information set corresponding to the target image; generating a mask image corresponding to each image object information in the image object information set to obtain a mask image set; for each image object information, the following first generation step is performed: in response to determining that the target object information set includes the image object information, obtaining target object noise reduction intensity information corresponding to the image object information from the object noise reduction intensity information set, and determining a target mask image corresponding to the image object information; packing the target object noise reduction intensity information, the image description information, the target mask image, and the target image to obtain packed information; using a pre-trained large language model supporting speech processing to generate image redrawing instruction information corresponding to each packed information in the obtained packed information set, to obtain an image redrawing instruction information set, wherein the large language model is a language model trained based on the target object information set and the object noise reduction intensity information set; storing the image redrawing instruction information set on a server corresponding to the target school intelligent platform.
[0009] In a second aspect, some embodiments of the present disclosure provide an instruction information storage device based on an artificial intelligence speech model, comprising: a first acquisition unit configured to acquire a set of object noise reduction intensity information corresponding to a set of target object information set on a target school intelligent platform; a second acquisition unit configured to acquire image description information and a set of image object information corresponding to a target image in response to receiving image redrawing instruction generation information for the target image; a first generation unit configured to generate a mask image corresponding to each image object information in the set of image object information to obtain a set of mask images; an execution unit configured to perform the following first generation step for each image object information: acquiring target object noise reduction intensity information corresponding to the image object information from the set of object noise reduction intensity information in response to determining that the set of target object information includes the image object information, and determining a target mask image corresponding to the image object information; packing the target object noise reduction intensity information, the image description information, the target mask image, and the target image to obtain packed information; a second generation unit configured to generate image redrawing instruction information corresponding to each packed information in the obtained set of packed information using a pre-trained large language model supporting speech processing to obtain a set of image redrawing instruction information, wherein the large language model is a language model trained based on the set of target object information and the set of object noise reduction intensity information; and an instruction information storage unit configured to store the set of image redrawing instruction information on a server corresponding to the target school intelligent platform.
[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; and a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.
[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer readable medium having a computer program stored thereon, wherein the program is executed by a processor to implement the method described in any implementation manner of the first aspect.
[0012] The above various embodiments of the present disclosure have the following beneficial effects: the instruction information storage method based on the artificial intelligence speech model of some embodiments of the present disclosure can efficiently and high-quality generate image redrawing instructions for target images to meet the diversified image needs of target objects. Specifically, the reason why the generation of related image redrawing instructions is not efficient is that the efficiency of manually customizing image redrawing instruction information on the school intelligent platform is low. It needs to constantly repeat trial and error image redrawing instruction information, which occupies large language models, so that computing resources are always occupied and a lot of human cost is wasted. Based on this, the instruction information storage method based on the artificial intelligence speech model of some embodiments of the present disclosure, first, obtains the object noise reduction intensity information set corresponding to the target object information set set on the target school intelligent platform. Here, by obtaining the object noise reduction intensity information set corresponding to the object information set, the determination efficiency of the image noise reduction intensity information of the subsequent image object information can be greatly improved. Then, in response to receiving the image redrawing instruction generation information for the target image, the image description information and the image object information set corresponding to the target image are obtained for subsequent generation of image redrawing instructions for each image object information. Next, the mask image corresponding to each image object information in the image object information set is generated to obtain a mask image set, so as to replace the image object with high quality for the mask image, for subsequent generation of high-quality redrawing images. Then, for each image object information in the image object information set, the following first generation step is performed: first, in response to determining that the target object information set includes the image object information, the target object noise reduction intensity information corresponding to the image object information is obtained from the object noise reduction intensity information set, and the target mask image corresponding to the image object information is determined to provide data basis for subsequent generation of image redrawing instruction information. Third, the target object noise reduction intensity information, the image description information, the target mask image and the target image are packaged to obtain packaging information, so as to realize the integration of data. Using a pre-trained large language model supporting speech processing, the image redrawing instruction information corresponding to each packaging information in the obtained packaging information set can be accurately and efficiently generated to obtain an image redrawing instruction information set, wherein the large language model is a language model trained based on the target object information set and the object noise reduction intensity information set. Finally, the image redrawing instruction information set is stored in the server corresponding to the target school intelligent platform, so as to facilitate the subsequent generation of redrawing images. In summary, by determining whether the image object information is in the target object information set, the corresponding object noise reduction intensity information is quickly determined. In addition, through the large language model, the image redrawing instruction information for the packaged data can be efficiently and accurately generated to obtain high-quality redrawing images. BRIEF DESCRIPTION OF DRAWINGS
[0013] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the annexed drawings in which: like reference numerals refer to like elements throughout. The annexed drawings are intended for purposes of illustration only and shall not limit the scope of the disclosure. The drawings are not necessarily drawn to scale, and certain components, elements, and / or structures that aid in the understanding of the present disclosure are shown by exaggerated size or in diagrammatic form.
[0014] Figure 1 is a flowchart of some embodiments of an instruction information storage method based on an artificial intelligence voice model according to the present disclosure;
[0015] Figure 2 is a structural schematic diagram of some embodiments of an instruction information storage apparatus based on an artificial intelligence voice model according to the present disclosure;
[0016] Figure 3 is a structural schematic diagram of an electronic device suitable for use to implement some embodiments of the present disclosure. DETAILED DESCRIPTION
[0017] Embodiments of the present disclosure will be described in detail with reference to the drawings, wherein the same or similar components are denoted by the same reference numerals, and thus, their duplicate description will be omitted. While the present disclosure is susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and will herein be described in detail. It should be understood, however, that there is no intent to limit the disclosure to the particular embodiments disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the disclosure. Like numbers refer to like elements throughout the description of the figures.
[0018] It is also to be noted that, for the sake of brevity, the figures herein may not show all components of the devices, only those components necessary in understanding the embodiments. While the embodiments will be described in the general context of a method, it should be understood that the embodiments can be implemented in combination with or separately from various types of devices and can be implemented in various types of devices. Further, while the embodiments will be described in the general context of a method, some embodiments can be implemented in one or more systems or apparatuses.
[0019] It should be noted that the terms "first", "second", and the like, used in the description and in the claims of the present disclosure are used only for distinguishing between similar elements and do not necessarily have an ordinal or chronological significance.
[0020] It should be noted that the terms "one", "multiple", and the like, used in the description and in the claims of the present disclosure are used only for illustrative purposes and should be understood as "one or more" unless otherwise indicated in the context.
[0021] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are used only for illustrative purposes and are not intended to limit the scope of the messages or information.
[0022] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, in which some embodiments of the present disclosure are shown. Like numbers refer to like elements throughout the several views.
[0023] Reference Figure 1, shows the flow 100 of some embodiments of the instruction information storage method based on the artificial intelligence voice model according to the present disclosure. The instruction information storage method based on the artificial intelligence voice model comprises the following steps:
[0024] Step 101, obtaining the object noise reduction intensity information set corresponding to the target object information set set on the target school intelligent platform.
[0025] In some embodiments, the execution subject (for example, an electronic device) of the above-mentioned instruction information storage method based on the artificial intelligence voice model can obtain the object noise reduction intensity information set corresponding to the target object information set set on the target school intelligent platform through wired connection or wireless connection. Wherein, the target school intelligent platform can be a processing platform for school scene supporting various intelligent processing. For example, the target school intelligent platform supports the following intelligent processing operations: image cropping operation, image synthesis operation, video editing operation, text generation operation. The target object information can be the object information corresponding to the target object. For example, the object information can be the object identifier of the target object. For the campus scene, the corresponding target object can be one of the following: campus personnel, campus scene. That is, the target object information set can include object information for school personnel and object information for each campus scene. The target object information in the target object information set and the object noise reduction intensity information in the above-mentioned object noise reduction intensity information set have a one-to-one correspondence. The object noise reduction intensity information can represent the noise reduction intensity of the element content of the target object in the image. The higher the noise reduction intensity, the clearer the target object in the image. For example, the object noise reduction intensity information can be but not limited to at least one of the following: first-level noise reduction intensity information, second-level noise reduction intensity information, third-level noise reduction intensity information. The noise reduction intensity corresponding to the first-level noise reduction intensity information is higher than the noise reduction intensity corresponding to the second-level noise reduction intensity information. The noise reduction intensity corresponding to the second-level noise reduction intensity information is higher than the noise reduction intensity corresponding to the third-level noise reduction intensity information. The object noise reduction intensity information can also be numerical information. The higher the corresponding numerical value, the higher the object content noise reduction intensity. In practice, for the campus personnel in the campus, multiple object noise reduction intensities can be set, that is, each object noise reduction intensity information can include multiple object noise reduction intensities. The campus scene is a scene represented by the campus landmark. For example, the campus landmark can be the campus library, and can also be the campus cafeteria, and can also be the campus gate. Each campus scene information can also have multiple object noise reduction intensities set.
[0026] It should be noted that the object denoising intensity information set and the target object information set can be obtained by specific training of the large language model. In practice, the specific training can be: obtaining a training data set corresponding to the target object information set. The training data can include: an image corresponding to the target object information and an image after denoising at a denoising intensity. The image corresponding to the target object information is used as input data, and the image after denoising is used as training purpose. The initial large language model is trained to obtain a large language model that can process the image corresponding to the target object information at a denoising intensity. Thus, the large language model can summarize the image denoising situation to obtain the object denoising intensity information set and the target object information set. Then, the object denoising intensity information set and the target object information set are recorded on the school intelligent platform for image redrawing processing of the object.
[0027] In step 102, in response to receiving the image redrawing instruction generation information for the target image, the image description information and the image object information set corresponding to the target image are obtained.
[0028] In some embodiments, in response to receiving the image redrawing instruction generation information for the target image, the execution subject can obtain the image description information and the image object information set corresponding to the target image.
[0029] The target image can be an image with improved quality of object content in the image. Specifically, the quality improvement can be the quality improvement of the object content corresponding to the object in the image. The image redrawing instruction generation information can be request information indicating that the image redrawing instruction is to be generated. The image redrawing instruction can be a prompt (Prompt) for redrawing the target image. By inputting the image redrawing instruction and the target image into the large language model, a corresponding redrawing image can be efficiently generated. The image description information can be a comprehensive description of the description feature information set corresponding to the target image. That is, the image description information describes the feature content under each feature corresponding to the target image. In practice, the description features in the description feature set and the description feature information in the description feature information set have a one-to-one correspondence. The description feature information can be the feature content corresponding to the description feature. The description feature set can include but is not limited to at least one of the following: image style feature, image color tone feature, image texture feature, image object category feature. The image object information in the image object information set can be the object information of the image object. In practice, the image object can be an element object in the image. That is, the image object information set can be the element information corresponding to each image element in the target image.
[0030] As an example, the execution subject can obtain the image description information through an AI tagging tool.
[0031] Step 103: Generate a mask image corresponding to each image object in the above image object information set to obtain a mask image set.
[0032] In some embodiments, the execution entity may generate a mask image corresponding to each image object in the image object information set to obtain a mask image set.
[0033] As an example, the aforementioned execution entity can use the natural language mask generation tool DINO to generate a mask image corresponding to each image object in the aforementioned image object information set, thereby obtaining a mask image set.
[0034] In some optional implementations of certain embodiments, generating the mask image corresponding to each image object in the aforementioned image object information set may include the following steps:
[0035] The first step is to determine the object region information in the target image that corresponds to the content of the image object information. The object region information can be the location information of the image object within the target image. In practice, the object region information can be in coordinate form. The content corresponding to the image object information can be the image content corresponding to the image object information.
[0036] As an example, the aforementioned execution entity can input the target image and the aforementioned image object information into the object region information generation model to obtain object region information. The object region information generation model can be a neural network model that generates the location of the object region in the image. In practice, the object region information generation model can be a YOLO model. Specifically, it can be any version of the YOLO model. The object region information generation model can be trained using conventional model training methods. In practice, the object region information generation model can be trained using conventional object recognition model training methods. That is, it can be trained using conventional YOLO model training methods.
[0037] The second step is to perform binarization processing on the target image based on the object region information to generate a binarized image, which serves as the mask image corresponding to the image object information.
[0038] As an example, the aforementioned execution entity can set the pixel values of the object region information in the target image to a first value, and the pixel values of the remaining regions to a second value, to generate a binarized image as a mask image. The first and second values can be preset. For example, the first value could be 12, and the corresponding second value could be 255.
[0039] Step 104: For each image object information, perform the following first generation step:
[0040] In response to determining that the target object information set includes the image object information, the execution subject can obtain target object denoising intensity information corresponding to the image object information from the object denoising intensity information set, and determine a target mask image corresponding to the image object information.
[0041] In response to determining that the target object information set includes the image object information, the execution subject can obtain target object denoising intensity information corresponding to the image object information from the object denoising intensity information set, and determine a target mask image corresponding to the image object information.
[0042] As an example, the execution subject can query the target object denoising intensity information corresponding to the image object information from the object denoising intensity information set by querying the target object denoising intensity information. The implementation of the target mask image refers to the implementation of the target object denoising intensity information.
[0043] In some optional implementations of some embodiments, after step 1042, the steps further include:
[0044] First, in response to determining that the target object information set does not include the image object information, at least one target object information corresponding to an object semantic similarity greater than a target degree between the object semantic corresponding to the image object information and the object semantic corresponding to the image object information is filtered out from the target object information set. Wherein, the object semantic can represent the object category corresponding to the image object information. The object category can be what kind of object the object in the image is. The semantic similarity can be the object category similarity. Specifically, the object category corresponding to each image object information can be determined by an object category table. The object category similarity between two object categories can be determined by a preset semantic similarity table. The semantic similarity can be a value between 0 and 1. The higher the value, the more similar the semantic content.
[0045] Second, at least one object denoising intensity information corresponding to the at least one target object information is determined. Wherein, there is a one-to-one correspondence between the target object information in the at least one target object information and the object denoising intensity information in the at least one object denoising intensity information.
[0046] Third, at least one semantic similarity corresponding to the at least one target object information is determined. Wherein, there is a one-to-one correspondence between the target object information in the at least one target object information and the semantic similarity in the at least one semantic similarity.
[0047] In the fourth step, at least one target intensity weight information corresponding to the at least one semantic similarity is generated. There is a one-to-one correspondence between a semantic similarity in the at least one semantic similarity and a target intensity weight information in the at least one target intensity weight information. The target intensity weight information can represent the importance of the object noise reduction intensity information. The higher the corresponding target intensity weight information, the more important the corresponding object noise reduction intensity information.
[0048] As an example, the at least one semantic similarity is input into a target intensity weight information conversion model to generate the at least one target intensity weight information. The target intensity weight information conversion model can be a neural network model that converts semantic similarity into corresponding target intensity weight information. That is, the target intensity weight information conversion model can represent the mapping relationship between semantic similarity and target intensity weight information. In practice, the target intensity weight information conversion model can be a fully connected layer with regression output. The regression content can be the content output by the model in the regression type. The target intensity weight information conversion model can be obtained by continuously updating the parameters of the fully connected layer through target training data sets and conventional model training. For another example, the target intensity weight information conversion model can also be a conventional regression model, such as a linear regression model. Through the target intensity weight information conversion model, the mapping relationship between semantic similarity and target intensity weight information is established.
[0049] In practice, a training data set between semantic similarity and target intensity weight information is obtained. Through the training data set, the model parameters of the initial intensity weight information conversion model are updated to obtain the target intensity weight information conversion model.
[0050] In the fifth step, a target intensity weight information in the at least one target intensity weight information is multiplied by a corresponding object noise reduction intensity information in the at least one object noise reduction intensity information to generate a multiplication result, thereby obtaining at least one multiplication result.
[0051] In the sixth step, an average value corresponding to the at least one multiplication result is determined as the target image noise reduction intensity information.
[0052] In step 1042, the target object noise reduction intensity information, the image description information, the target mask image, and the target image are packaged to obtain packaging information.
[0053] In some embodiments, the execution subject can package the target object noise reduction intensity information, the image description information, the target mask image, and the target image to obtain the packaging information.
[0054] In step 105, a pre-trained large language model supporting voice processing is used to generate image redrawing instruction information corresponding to each piece of packaged information in the obtained set of packaged information, to obtain a set of image redrawing instruction information.
[0055] In some embodiments, the execution subject described above can use a pre-trained large language model supporting voice processing to generate image redrawing instruction information corresponding to each piece of packaged information in the obtained set of packaged information, to obtain a set of image redrawing instruction information. The large language model can support voice output and text input. That is, the large language model can be a multi-modal large language model supporting multi-modal input. When the input is voice, the large language model converts the voice into text and then processes based on the text. The large language model can be a large language model that focuses on model training of target object information and corresponding object noise reduction intensity information. So that the large language model can realize noise reduction processing under the object noise reduction intensity for the object content under each target object information. The large language model for image noise reduction refers to a model that can process image noise reduction tasks through natural language processing technology. These models usually combine natural language processing (NLP), computer vision (CV), and audio processing technologies, and use the powerful language understanding and generation capabilities of large language models to understand and process image content and audio content. For example, the large language model can be a model based on Stable Diffusion, and can also be a time series denoising model based on the Transformer architecture. Stable Diffusion is composed of multiple parts, including a text understanding component, an image information creator, and an image generator. The text understanding component converts text information into a digital representation and then inputs it into the image generator to generate high-quality images. Its efficient internal structure and multi-step generation process make the generated images have higher quality, faster running speed, and less resource consumption. The Transformer architecture captures long-range dependencies in sequence data through self-attention mechanisms. This mechanism can identify and reduce noise by learning pixel relationships in images when processing images. Combined with deep learning techniques, the Transformer model can exhibit superior performance in complex image noise reduction tasks. The image redrawing instruction information can be instruction information (i.e., prompt information) representing image redrawing of a target image.
[0056] As an example, first, the execution subject described above can generate prompt information representing generation of image redrawing instruction information from packaged data. Then, the prompt information and the packaged data are input into the large language model to obtain the image redrawing instruction information.
[0057] In some optional implementations of some embodiments, the execution subject described above can use a pre-trained large language model to generate image redrawing instruction information corresponding to each piece of packaged information in the obtained set of packaged information, including the following steps:
[0058] In a first step, the target object fills in the redrawing requirement information on the redrawing demand processing page corresponding to the target school intelligent platform. The redrawing requirement information can be information input by the target object to require the content corresponding to the redrawing picture. For example, the redrawing requirement information can be: "the resolution corresponding to the redrawing picture is the target resolution".
[0059] In a second step, the generation instruction information is generated to generate the image redrawing instruction information according to the redrawing requirement information and the packaging information.
[0060] In a third step, the generation instruction information is input into the large language model to obtain the image redrawing instruction information.
[0061] As an example, the execution subject can input the generation instruction information and the packaging information into the large language model to obtain the image redrawing instruction information.
[0062] In step 106, the image redrawing instruction information set is stored on the server corresponding to the target school intelligent platform.
[0063] In some embodiments, the execution subject can store the image redrawing instruction information set and the image identifier corresponding to the target image in the form of a key-value pair on the server corresponding to the target school intelligent platform.
[0064] In some optional implementations of some embodiments, after step 106, the steps further include:
[0065] In a first step, in response to receiving the image redrawing request initiated by the target object on the target school intelligent platform, each image redrawing instruction information in the image redrawing instruction information set and the target image are displayed on the instruction information display page in the target school intelligent platform. The instruction information display page includes an instruction information editing control. The instruction information display page can be a page for displaying the image redrawing instruction information set. The instruction information editing control can be a control for editing the image redrawing instruction information. In practice, on the instruction information display page, there is a corresponding instruction information editing control at the target position of each image redrawing instruction information.
[0066] Second step, in response to detecting that the target object clicks the selection information for at least one image redraw instruction information on the instruction information display page, input the at least one image redraw instruction information into the large language model to generate at least one redrawn image required by the target object. Wherein, at least one image redraw instruction information can be selected by the target object. For example, the selection information can be clicking the image redraw instruction information on the instruction information display page. The redrawn image in the at least one redrawn image has a one-to-one correspondence with the image redraw instruction information in the at least one image redraw instruction information. In practice, the at least one image redraw instruction information and the corresponding at least one package information can be input into the large language model to generate at least one redrawn image required by the target object.
[0067] Third step, display the at least one redrawn image on the redrawn image display page in the target school intelligent platform. Wherein, the redrawn image display page has a data transmission control and at least one redrawn image regeneration control corresponding to the at least one redrawn image. The data transmission control can be a control that supports the target object to transmit the at least one redrawn image.
[0068] Fourth step, in response to determining that the target object clicks the target redrawn image regeneration control, pop up the instruction information display pop-up window of the image redraw instruction information corresponding to the target redrawn image regeneration control. Wherein, the instruction information display pop-up window includes: instruction information editing control. The instruction information editing control can be a control for editing the instruction information.
[0069] Fifth step, in response to determining that the target object clicks the instruction information editing control and edits the image redraw instruction information corresponding to the target redrawn image regeneration control, obtains the edited image redraw instruction information.
[0070] Sixth step, store the edited image redraw instruction information and display it on the instruction information display page.
[0071] Seventh step, in response to detecting that the target object clicks the selection information for the edited image redraw instruction information on the instruction information display page, input the edited image redraw instruction information into the large language model to generate the target redrawn image required by the target object.
[0072] Eighth step, display the target redrawn image on the redrawn image display page.
[0073] In the process of adopting technical solutions to solve the above technical problems mentioned in the background, the following problems often accompany: "generating a redrawn image using a large language model is often accurate and efficient, but for specific scenarios (e.g., the large language model has not been trained and learned relevant semantic feature information), resulting in inaccurate redrawn images." The inventors considered the lack of accurate checking of the output of the large language model, and decided to use the following scheme to solve it:
[0074] In some optional implementations of some embodiments, after the above at least one image redrawing instruction information is input into the above large language model to generate at least one redrawn image required by the target object, the method further comprises:
[0075] First, for each of the at least one redrawn image, perform a second generation step:
[0076] Substep 1, obtain the description feature set of the image description information corresponding to the redrawn image. Each description feature in the description feature set can be a pre-determined image feature. The description feature is a feature used to describe the content of the image. In practice, the description feature set can include but is not limited to at least one of the following: image style feature, image texture feature, image color tone feature, image object category feature.
[0077] Substep 2, obtain the image description information extraction model corresponding to the description feature set, wherein the image description information extraction model is a neural network model with the description feature set as the target output. That is, the output content corresponding to the image description information extraction model is the feature information set corresponding to the description feature set. The image description information extraction model can be a neural network model that extracts the feature content of the description feature set in the image. That is, the model input of the image description information extraction model is the image, and the output is the description feature information corresponding to the image. The image description information extraction model can extract all-around feature information through a plurality of convolution layers in series (for example, 11 convolution layers in series), and then output the feature content of each description feature in the description feature set through a plurality of parallel fully connected layers. The image description information extraction model can be trained in conjunction with the corresponding image application scenario. For example, the image description information extraction model can be trained in conjunction with the image recognition model. That is, the image description information extraction model can be used as a feature extraction module of the image recognition model.
[0078] Sub-step 3, input the redrawing image into the image description information extraction model to generate redrawing image description information, wherein the redrawing image description information includes a set of redrawing description feature information corresponding to the set of description features. The description features in the set of description features and the redrawing description feature information in the set of redrawing description feature information have a one-to-one correspondence. The redrawing description feature information can be image description feature information corresponding to the redrawing image. The redrawing description feature information can be feature content of the description feature for the redrawing image.
[0079] Sub-step 4, determine the set of description feature information corresponding to the set of description features included in the image description information. The description features in the set of description features and the description feature information in the set of description feature information have a one-to-one correspondence. The description feature information can be feature content corresponding to the description feature.
[0080] Sub-step 5, generate a feature information difference between the set of redrawing description feature information and the set of description feature information.
[0081] As an example, first, the execution subject can perform information difference between each redrawing description feature information in the set of redrawing description feature information and the corresponding description feature information in the set of description feature information to generate information difference values, and obtain a set of information difference values. Then, the average value corresponding to the set of information difference values is determined as the feature information difference.
[0082] Sub-step 6, in response to determining that the feature information difference is less than a predetermined information difference, generate an initial mask image corresponding to the redrawing image. The predetermined information difference can be a predetermined information difference. For example, the preset information difference can be a preset information difference value. When the feature information difference is higher than or equal to the preset information difference, it indicates that the feature difference is large. When the feature information difference is less than the preset information difference, it indicates that the feature difference is small. Specifically, the generation method of the initial mask image can refer to the generation method of the mask image.
[0083] In response to determining that the mask image difference between the initial mask image and the corresponding target mask image is less than the predetermined difference information, the target image and the redrawn image are input into an object noise reduction intensity information generation model to generate noise reduction intensity information. The mask image difference can be the image pixel difference between the initial mask image and the target mask image. The object noise reduction intensity information generation model can be a neural network model that generates object noise reduction intensity information. The object noise reduction intensity information generation model can be a neural network model that generates the noise reduction difference between the target image and the redrawn image. In practice, the object noise reduction intensity information generation model can be a multi-layer convolutional neural network model connected in series. Specifically, the object noise reduction intensity information generation model can include a first feature extraction layer for extracting image feature semantic content corresponding to the target image, a second feature extraction layer for extracting image feature semantic content corresponding to the redrawn image, a feature fusion layer for fusing the output of the first feature extraction layer and the output of the second feature extraction layer, and a series of fully connected layers for regression output. The first feature extraction layer and the second feature extraction layer can be convolutional layers connected in series with different numbers of layers. The feature fusion layer can be a feature information splicing layer.
[0084] In response to determining that the difference between the noise reduction intensity information and the corresponding object noise reduction intensity information is less than the preset difference value, information indicating that the redrawn image passes the image verification is generated.
[0085] Alternatively, in response to determining that the difference between the noise reduction intensity information and the corresponding object noise reduction intensity information is greater than or equal to the preset difference value, information indicating that the redrawn image fails the image verification is generated.
[0086] Optionally, the image description information extraction model comprises: a first image feature extraction model, a first description feature information extraction model set corresponding to the description feature set, a first description feature information output layer set corresponding to the description feature set, a second image feature extraction model, a second description feature information output layer set corresponding to the description feature set, and a script output layer. The second image feature extraction model comprises a plurality of convolutional neural networks connected in series. The plurality of convolutional neural networks connected in series comprises at least one convolutional neural network connected in series, and the at least one convolutional neural network connected in series is at a target position in the plurality of convolutional neural networks connected in series. The target position can be a predetermined number of positions. The convolutional neural networks in the at least one convolutional neural network connected in series have a one-to-one correspondence with the second description feature information output layers in the second description feature information output layer set. The first image feature extraction model can be a neural network model for extracting image feature information. In practice, the first image feature extraction model can be an 8-layer convolutional neural network connected in series. The first description feature information extraction model in the first description feature information extraction model set has a one-to-one correspondence with the description features in the description feature set. That is, the first description feature information extraction model can be a neural network model for extracting features for the description features. In practice, the first description feature information extraction model can be a multi-layer convolutional neural network connected in series. The description features in the description feature set have a one-to-one correspondence with the first description feature information output layers in the first description feature information output layer set. The first description feature information output layer can be a fully connected layer. The second image feature extraction model can be a neural network model for extracting image feature information. The description features in the description feature set have a one-to-one correspondence with the second description feature information output layers in the second description feature information output layer set. In practice, the second description feature information output layer can be a multi-layer convolutional neural network connected in series. The script output layer can be a network layer for generating a script. In practice, the script output layer can be a recurrent neural network model. The plurality of convolutional neural networks connected in series can be a 10-layer convolutional neural network connected in series. The number of network layers corresponding to the at least one convolutional neural network connected in series is the same as the number of output layers corresponding to the second description feature information output layer set.
[0087] Optionally, the inputting the redrawing image into the image description information extraction model to generate redrawing image description information can comprise the following steps:
[0088] Firstly, the redrawing image is input into the first image feature extraction model to generate first image feature information. The first image feature information can represent the vector form of the image feature semantic content corresponding to the redrawing image.
[0089] Secondly, input each of the first image feature information into each of the first description feature information extraction model in the first description feature information extraction model set to generate first description feature extraction information, and obtain a first description feature extraction information set.
[0090] Thirdly, input each of the first description feature extraction information in the first description feature extraction information set into the corresponding first description feature information output layer in the first description feature information output layer set to output first initial description feature information, and obtain a first initial description feature information set.
[0091] Fourthly, input the first description feature extraction information set into the second image feature extraction model to generate at least one convolution output result for the at least one serially connected convolutional neural network and an overall convolution output result for the plurality of serially connected convolutional neural networks.
[0092] Fifthly, input each of the at least one convolution output result into the corresponding second description feature information output layer in the second description feature information output layer set to generate second description feature information, and obtain a second initial description feature information set.
[0093] Sixthly, determine the information difference between the first initial description feature information set and the second initial description feature information set. The specific implementation is not described here.
[0094] Seventhly, in response to determining that the information difference is less than a predetermined degree difference value, input the overall convolution output result into the script output layer to generate script information. The script information can be a description script corresponding to the redrawn image.
[0095] Eighthly, generate the redrawn image description information according to the script information, the first initial description feature information set and the second initial description feature information set.
[0096] As an example, firstly, perform feature information extraction on the script information to generate a fourth initial description feature information set. Then, according to a voting mechanism, filter out a first actual description feature information set corresponding to the description feature set from the first initial description feature information set, the second initial description feature information set and the fourth initial description feature information set. Finally, determine the first actual description feature information set as the redrawn image description information.
[0097] Optionally, the image description information extraction model further comprises: at least one second description feature information extraction model corresponding to the at least one serially connected convolutional neural network, and a third description feature information output layer set. There is a one-to-one correspondence between a convolutional neural network in the at least one serially connected convolutional neural network and a second description feature information extraction model in the at least one second description feature information extraction model. There is a one-to-one correspondence between a third description feature information output layer in the third description feature information output layer set and a description feature in the description feature set.
[0098] Optionally, the step further comprises:
[0099] First, in response to determining that the information difference is greater than or equal to the predetermined degree difference value, for each convolutional output result in the at least one convolutional output result, the following second generation step is performed:
[0100] First sub-step, input the convolutional output result into the corresponding second description feature information extraction model in the at least one second description feature information extraction model to generate second description feature extraction information.
[0101] Second sub-step, input the second description feature extraction information into the corresponding third description feature information output layer in the third description feature information output layer set to output third initial description feature information.
[0102] Second, generate the redrawn image description information according to the first initial description feature information set, the obtained third initial description feature information set, and the second initial description feature information set.
[0103] As an example, first, the execution subject can use a voting mechanism to generate a second actual description feature information set corresponding to the description feature set according to the first initial description feature information set, the obtained third initial description feature information set, and the second initial description feature information set. Finally, the second actual description feature information set is determined as the redrawn image description information.
[0104] Optionally, after the above at least one image redrawing instruction information is input into the above large language model to generate at least one redrawing image required by the above target object, the redrawing image is subjected to image checking as another application point of the present disclosure, solving the problem that the large language model output is not accurate enough, resulting in a large content deviation in the redrawing image. Based on this, for each redrawing image, the feature information difference between the description feature information corresponding to the redrawing image and the description feature information in the original image description information is extracted to preliminarily determine whether there is a problem of large feature deviation. On this basis, by generating an initial mask image corresponding to the redrawing image and comparing it with the original corresponding target mask image, the noise reduction strength is further determined, and in the case where the noise reduction strength is less than a predetermined difference value, it is determined that the generation of the redrawing image has no problem. In addition, in the process of preliminarily determining the feature deviation, it is crucial to extract the redrawing image feature information in the redrawing image. Based on this, by the first image feature extraction model included in the above image description information extraction model, the first description feature information extraction model set corresponding to the above description feature set, the first description feature information output layer set corresponding to the above description feature set, the second image feature extraction model, the second description feature information output layer set corresponding to the above description feature set and the script output layer, and the specific design of the corresponding model structure, the accurate extraction of the description feature information can be realized, and the accuracy of the determined feature information difference is guaranteed.
[0105] The above various embodiments of the present disclosure have the following beneficial effects: the instruction information storage method based on the artificial intelligence voice model of some embodiments of the present disclosure can efficiently and high-quality generate image redrawing instructions for target images to meet the diversified image needs of target objects. Specifically, the reason why the generation of related image redrawing instructions is not efficient is that it is inefficient to manually customize image redrawing instruction information on the school intelligent platform. It needs to constantly repeat trial and error image redrawing instruction information, which occupies large language models, so that computing resources are always occupied and a lot of human cost is wasted. Based on this, the instruction information storage method based on the artificial intelligence voice model of some embodiments of the present disclosure, first, obtains the object noise reduction intensity information set corresponding to the target object information set set on the target school intelligent platform. Here, by obtaining the object noise reduction intensity information set corresponding to the object information set, the determination efficiency of the image noise reduction intensity information of the image object information can be greatly improved. Then, in response to receiving the image redrawing instruction generation information for the target image, the image description information and the image object information set corresponding to the target image are obtained for subsequent generation of image redrawing instructions for each image object information. Next, the mask image corresponding to each image object information in the image object information set is generated to obtain a mask image set, so as to replace the image object with high quality for the mask image, for subsequent generation of high-quality redrawing images. Then, for each image object information in the image object information set, the following first generation step is performed: first, in response to determining that the target object information set includes the image object information, the target object noise reduction intensity information corresponding to the image object information is obtained from the object noise reduction intensity information set, and the target mask image corresponding to the image object information is determined to provide data basis for subsequent generation of image redrawing instruction information. Third, the target object noise reduction intensity information, the image description information, the target mask image and the target image are packaged to obtain packaging information, so as to realize the integration of data. Using a pre-trained large language model, the image redrawing instruction information corresponding to each packaging information in the obtained packaging information set can be accurately and efficiently generated to obtain an image redrawing instruction information set, wherein the large language model is a language model trained based on the target object information set and the object noise reduction intensity information set. Finally, the image redrawing instruction information set is stored in the server corresponding to the target school intelligent platform, so as to facilitate the subsequent generation of redrawing images. In summary, by determining whether the image object information is in the target object information set, the corresponding object noise reduction intensity information is quickly determined. In addition, through the large language model, the image redrawing instruction information for the packaged data can be efficiently and accurately generated to obtain high-quality redrawing images.
[0106] Further reference Figure 2As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of instruction information storage devices based on artificial intelligence speech models, which device embodiments correspond to those method embodiments shown in Figure 1 The instruction information storage devices based on artificial intelligence speech models can be applied in various electronic devices.
[0107] As shown in Figure 2 An instruction information storage device 200 based on an artificial intelligence speech model includes a first acquisition unit 201, a second acquisition unit 202, a first generation unit 203, an execution unit 204, a second generation unit 205, and an instruction information storage unit 206. The first acquisition unit 201 is configured to acquire a set of object noise reduction intensity information corresponding to a set of target object information set on a target school intelligent platform; the second acquisition unit 202 is configured to acquire a set of image description information and image object information corresponding to a target image in response to receiving image redrawing instruction generation information for the target image; the first generation unit 203 is configured to generate a mask image corresponding to each image object information in the set of image object information, obtaining a set of mask images; the execution unit 204 is configured to perform the following first generation steps for each image object information: in response to determining that the set of target object information includes the image object information, acquiring target object noise reduction intensity information corresponding to the image object information from the set of object noise reduction intensity information, and determining a target mask image corresponding to the image object information; packing the target object noise reduction intensity information, the image description information, the target mask image, and the target image to obtain packed information; the second generation unit 205 is configured to generate image redrawing instruction information corresponding to each packed information in the set of obtained packed information using a pre-trained large language model supporting speech processing, obtaining a set of image redrawing instruction information, wherein the large language model is a language model trained based on the set of target object information and the set of object noise reduction intensity information; and the instruction information storage unit 206 is configured to store the set of image redrawing instruction information on a server corresponding to the target school intelligent platform.
[0108] It can be understood that the units described in the instruction information storage device 200 based on artificial intelligence speech models correspond to the respective steps in the methods described with reference to Figure 1 Therefore, the operations, features, and beneficial effects described above for the methods also apply to the instruction information storage device 200 based on artificial intelligence speech models and the units contained therein, which will not be described again.
[0109] The following will be described with reference to Figure 3It shows a schematic diagram of the structure of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0110] like Figure 3 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0111] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0112] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0113] Note that the computer-readable medium or media used to provide the computer program sequence to the computer system can be embedded in a computer program product, which comprises all the respective features, which are provided with the computer program sequence, and which are enumerated above. It is understood that the computer-readable medium or media described herein are included in the computer program product, or are a component of the computer program product. In some embodiments of the disclosure, the computer-readable storage medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In some embodiments of the disclosure, a computer-readable storage medium can be any tangible medium that contains, or stores a program for use by or in connection with an instruction execution system, apparatus, or device. In some embodiments of the disclosure, a computer-readable signal medium can include a computer-readable storage medium in baseband or propagated as a carrier wave in a propagated data signal, which contains a computer-readable program code. Such a propagated signal can take a wide variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0114] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0115] The computer readable medium can be included in the electronic device, or can exist separately from the electronic device. The computer readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: obtain a set of object denoising strength information corresponding to a set of target object information set on a target school intelligent platform; in response to receiving image redrawing instruction generation information for a target image, obtain image description information and a set of image object information corresponding to the target image; generate a mask image corresponding to each image object information in the set of image object information to obtain a set of mask images; for each image object information, perform the following first generation step: in response to determining that the set of target object information includes the image object information, obtaining target object denoising strength information corresponding to the image object information from the set of object denoising strength information, and determining a target mask image corresponding to the image object information; packing the target object denoising strength information, the image description information, the target mask image and the target image to obtain packed information; using a pre-trained large language model supporting voice processing, generating image redrawing instruction information corresponding to each packed information in the obtained set of packed information to obtain a set of image redrawing instruction information, wherein the large language model is a language model trained based on the set of target object information and the set of object denoising strength information; and storing the set of image redrawing instruction information on a server corresponding to the target school intelligent platform.
[0116] Computer program code for carrying out operations of some embodiments of the disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0117] The flow and block diagrams in the drawings represent possible architectural, functional, and operational architectures of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block can represent a module, a segment, or a portion of code that comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and software.
[0118] The units described in some embodiments of the present disclosure can be implemented by means of software, or can be implemented by means of hardware. The described units can also be arranged in a processor, for example, it can be described that: a processor includes a first acquisition unit, a second acquisition unit, a first generation unit, an execution unit, a second generation unit and an instruction information storage unit. Among them, the name of these units does not constitute a limitation to the unit itself in some cases, for example, the instruction information storage unit can also be described as "a unit for storing the image redrawing instruction information set in the server corresponding to the target school intelligent platform".
[0119] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), etc.
[0120] The above description is merely some of the preferred embodiments of the present disclosure and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the application involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above inventive concept. For example, the above features are replaced with the technical features disclosed in the embodiments of the present disclosure (but not limited to) having similar functions to form technical solutions.
Claims
1. An instruction information storage method based on an artificial intelligence voice model, comprising: obtaining a set of object noise reduction intensity information corresponding to a set of target object information set on a target school intelligent platform; in response to receiving an image redrawing instruction generation information for a target image, obtaining image description information and a set of image object information corresponding to the target image; generating a mask image corresponding to each image object information in the set of image object information to obtain a set of mask images; for each image object information, performing the following first generation step: in response to determining that the set of target object information includes the image object information, obtaining target object noise reduction intensity information corresponding to the image object information from the set of object noise reduction intensity information, and determining a target mask image corresponding to the image object information; packaging the target object noise reduction intensity information, the image description information, the target mask image and the target image to obtain packaged information; using a pre-trained large language model supporting voice processing, generating image redrawing instruction information corresponding to each packaged information in the obtained set of packaged information to obtain a set of image redrawing instruction information, wherein the large language model is a language model trained based on the set of target object information and the set of object noise reduction intensity information; storing the set of image redrawing instruction information on a corresponding server of the target school intelligent platform.
2. The method of claim 1, wherein, After the step of obtaining target object noise reduction intensity information corresponding to the image object information from the set of object noise reduction intensity information in response to determining that the set of target object information includes the image object information, and determining a target mask image corresponding to the image object information, the method further comprises: in response to determining that the image object information is not included, filtering at least one target object information from the set of target object information, wherein the semantic similarity between the corresponding object semantics of the at least one target object information and the corresponding object semantics of the image object information is greater than a target degree; determining at least one object noise reduction intensity information corresponding to the at least one target object information; determining at least one semantic similarity corresponding to the at least one target object information; generating at least one target intensity weight information for the at least one semantic similarity; multiplying target intensity weight information in the at least one target intensity weight information and corresponding object noise reduction intensity information in the at least one object noise reduction intensity information to generate a multiplication result, obtaining at least one multiplication result; determining an average value corresponding to the at least one multiplication result as target object noise reduction intensity information.
3. The method of claim 1, wherein, The step of generating a mask image corresponding to each image object information in the set of image object information comprises: determining object region information in the target image corresponding to the content of the image object information; performing binaryzation processing on the target image according to the object region information to generate a binaryzation image as a mask image corresponding to the image object information.
4. The method of claim 1, wherein, The step of using a pre-trained large language model supporting voice processing to generate image redrawing instruction information corresponding to each packaged information in the obtained set of packaged information comprises: obtaining redrawing requirement information filled in a redrawing requirement processing page corresponding to the target school intelligent platform by the target object; generating generation instruction information representing generation of image redrawing instruction information according to the redrawing requirement information and the packaging information; inputting the generation instruction information into the large language model to obtain the image redrawing instruction information.
5. The method of claim 1, wherein, The method further comprises: in response to receiving an image redrawing request initiated by the target object on the target school intelligent platform, displaying each image redrawing instruction information in the image redrawing instruction information set and instruction information display page of the target image on the target school intelligent platform, wherein the instruction information display page comprises an instruction information editing control; in response to detecting that the target object clicks selection information for at least one image redrawing instruction information on the instruction information display page, inputting the at least one image redrawing instruction information into the large language model to generate at least one redrawing image required by the target object; displaying the at least one redrawing image on a redrawing image display page in the target school intelligent platform, wherein the redrawing image display page comprises a data transmission control and at least one redrawing image regeneration control corresponding to the at least one redrawing image; in response to determining that the target object clicks a target redrawing image regeneration control, popping up an instruction information display pop-up window of image redrawing instruction information corresponding to the target redrawing image regeneration control, wherein the instruction information display pop-up window comprises an instruction information editing control; in response to determining that the target object clicks the instruction information editing control and edits the image redrawing instruction information corresponding to the target redrawing image regeneration control, obtaining edited image redrawing instruction information; storing the edited image redrawing instruction information and displaying it on the instruction information display page; in response to detecting that the target object clicks selection information for the edited image redrawing instruction information on the instruction information display page, inputting the edited image redrawing instruction information into the large language model to generate a target redrawing image required by the target object; displaying the target redrawing image on the redrawing image display page.
6. An instruction information storage device based on an artificial intelligence voice model, comprising: a first obtaining unit configured to obtain a set of object noise reduction intensity information corresponding to a set of target object information set on a target school intelligent platform; a second obtaining unit configured to obtain image description information and a set of image object information corresponding to a target image in response to receiving image redrawing instruction generation information for the target image; a first generating unit configured to generate a mask image corresponding to each image object information in the set of image object information to obtain a set of mask images; The execution unit is configured to, for each image object information, perform the following first generation step: in response to determining that the target object information set includes the image object information, obtaining target object denoising intensity information corresponding to the image object information from the object denoising intensity information set, and determining a target mask image corresponding to the image object information; packaging the target object denoising intensity information, the image description information, the target mask image and the target image to obtain packaged information; The second generation unit is configured to generate image redrawing instruction information corresponding to each packaged information in the obtained packaged information set by using a pre-trained large language model supporting voice processing, to obtain a set of image redrawing instruction information, wherein the large language model is a language model trained based on the target object information set and the object denoising intensity information set; The instruction information storage unit is configured to store the set of image redrawing instruction information on the server corresponding to the target school intelligent platform.
7. An electronic device, comprising: one or more processors; a memory device having stored thereon one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1-5.
8. A computer readable medium having stored thereon a computer program, wherein, The program is executed by the processor to implement the method of any one of claims 1-5. The program is executed by the processor to implement the method of any one of claims 1-5.
Citation Information
Patent Citations
Image redrawing method and device, computer equipment and storage medium
CN117808917A
Image processing method and apparatus, electronic device and storage medium
US20250086806A1