Instruction information storage method and device based on artificial intelligence voice model, and medium
By obtaining object noise reduction intensity information and generating mask images, and using pre-trained large language models to generate image redraw instruction information, the problem of inefficiency of artificial customization is solved, efficient and high-quality image redraw instruction generation and storage is achieved, and the computing resource utilization rate of the intelligent platform is improved.
Patent Information
- Application Number
- CN202510456353.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-11
AI Technical Summary
In the prior art, the inefficiency of artificially customizing image redrawing instruction information on intelligent platforms leads to waste of computing resource occupation and labor costs, and the redrawing of images in scenarios where large language models are not trained is not accurate enough.
By obtaining the object noise reduction intensity information set of the target object information set, a mask image is generated and the image redraw instruction information is generated using the pre-trained large language model, and stored on the intelligent platform, combining the image description information and the object information set for efficient and accurate image redraw instruction generation.
It improves the efficiency and quality of image redrawing instructions generation, reduces computing resource occupation and labor costs, and ensures high-quality redrawing effect under diversified image requirements.
Smart Images

Figure CN120407834A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of large language models that support speech processing, and more particularly to a method, device, and medium for storing instruction information based on an artificial intelligence speech model. Background Art
[0002] Currently, images have become an effective form of presenting information. With the continuous upgrade of hardware devices, the quality requirements for images are also getting higher and higher. For the generation of image redrawing instruction information for a target image, the commonly adopted method is that the image upload object manually edits according to the image requirements of the desired redrawn image to obtain the image redrawing instruction information.
[0003] However, when using the above method to generate image redrawing instruction information, the following technical problems often exist:
[0004] The efficiency of manually customizing and inputting image redrawing instruction information on the school intelligent platform is low. It is necessary to continuously repeat and trial-and-error the image redrawing instruction information, which will occupy the large language model, resulting in the continuous occupation of computing resources and wasting a lot of human costs.
[0005] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention
[0006] This content part of the present disclosure is used to introduce the concepts in a brief form, and these concepts will be described in detail in the subsequent detailed implementation part. This content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0007] Some embodiments of the present disclosure propose a method, device, and medium for storing instruction information based on an artificial intelligence speech model to solve one or more of the technical problems mentioned in the above background art section.
[0008] In a first aspect, some embodiments of the present disclosure provide a method for storing instruction information based on an artificial intelligence voice model, including: obtaining an object noise reduction intensity information set corresponding to a target object information set set on a target school intelligent platform; in response to receiving information generated by an image redrawing instruction for a target image, obtaining image description information and an image object information set corresponding to the target image; generating a mask image corresponding to each image object information in the image object information set to obtain a mask image set; for each image object information, performing the following first generation step: in response to determining that the target object information set includes the image object information, obtaining target object noise reduction intensity information corresponding to the image object information from the object noise reduction intensity information set, and determining a target mask image corresponding to the image object information; packing the target object noise reduction intensity information, the image description information, the target mask image, and the target image to obtain packed information; using a pre-trained large language model supporting voice processing to generate image redrawing instruction information corresponding to each packed information in the obtained packed information set to obtain an image redrawing instruction information set, where the large language model is a language model specifically trained based on the target object information set and the object noise reduction intensity information set; storing the image redrawing instruction information set on a server corresponding to the target school intelligent platform.
[0009] In a second aspect, some embodiments of the present disclosure provide an instruction information storage device based on an artificial intelligence voice model, comprising: a first acquisition unit, configured to acquire an object noise reduction intensity information set corresponding to a target object information set set on a target school intelligent platform; a second acquisition unit, configured to generate information in response to receiving an image redrawing instruction for a target image, and acquire image description information and an image object information set corresponding to the above target image; a first generation unit, configured to generate a mask image corresponding to each image object information in the above image object information set, and obtain a mask image set; an execution unit, configured to execute the following first generation step for each image object information: in response to determining that the above target object information set includes the above image object information, generating a mask image from the above object noise reduction intensity information set; The target object noise reduction intensity information corresponding to the above-mentioned image object information is obtained from the information set, and the target mask image corresponding to the above-mentioned image object information is determined; the target object noise reduction intensity information, the above-mentioned image description information, the above-mentioned target mask image and the above-mentioned target image are packaged to obtain packaged information; the second generation unit is configured to use a pre-trained large language model that supports speech processing to generate image redrawing instruction information corresponding to each packaged information in the obtained packaged information set, and obtain an image redrawing instruction information set, wherein the above-mentioned large language model is a language model that is specifically trained based on the above-mentioned target object information set and the object noise reduction intensity information set; the instruction information storage unit is configured to store the above-mentioned image redrawing instruction information set on the server corresponding to the above-mentioned target school intelligent platform.
[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.
[0012] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: The instruction information storage method based on the artificial intelligence voice model in some embodiments of the present disclosure can efficiently and high-quality generate image redrawing instructions for a target image to meet the diverse image needs of the target object. Specifically, the reason for the inefficient generation of relevant image redrawing instructions is that the efficiency of manually customizing and inputting image redrawing instruction information on the school intelligent platform is low. It is necessary to continuously repeat and trial-and-error the image redrawing instruction information, which will occupy the large language model, resulting in the continuous occupation of computing resources and wasting a lot of labor costs. Based on this, in some embodiments of the present disclosure, the instruction information storage method based on the artificial intelligence voice model, first, obtains the object noise reduction intensity information set corresponding to the target object information set set on the target school intelligent platform. Here, by obtaining the object noise reduction intensity information set corresponding to the target object information set, the determination efficiency of the image noise reduction intensity information of the subsequent image object information can be greatly improved. Then, in response to receiving the image redrawing instruction generation information for the target image, obtains the image description information and the image object information set corresponding to the above target image for subsequent generation of corresponding image redrawing instructions for each image object information. Next, generates a mask image corresponding to each image object information in the above image object information set to obtain a mask image set, so as to facilitate high-quality replacement of the image object for the mask image and be used for subsequent generation of high-quality redrawn images. Then, for each image object information in the above image object information set, the following first generation step is performed: The first step, in response to determining that the above target object information set includes the above image object information, obtains the target object noise reduction intensity information corresponding to the above image object information from the above object noise reduction intensity information set, and determines the target mask image corresponding to the above image object information to provide a data basis for subsequent generation of image redrawing instruction information. The third step, packages the above target object noise reduction intensity information, the above image description information, the above target mask image, and the above target image to obtain a packaged information to achieve data integration. Using the pre-trained large language model supporting speech processing, can accurately and efficiently generate the image redrawing instruction information corresponding to each packaged information in the obtained packaged information set to obtain an image redrawing instruction information set, where the above large language model is a language model specifically trained based on the above target object information set and the object noise reduction intensity information set. Finally, stores the above image redrawing instruction information set on the server corresponding to the above target school intelligent platform for subsequent generation of redrawn images. In summary, by determining whether the image object information is in the target object information set, the corresponding object noise reduction intensity information can be quickly determined. In addition, through the large language model, the image redrawing instruction information for the packaged data can be efficiently and accurately generated, so as to obtain high-quality redrawn images subsequently. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0014] Figure 1 is a flowchart of some embodiments of a method for storing instruction information based on an artificial intelligence speech model according to the present disclosure;
[0015] Figure 2 is a schematic structural diagram of some embodiments of a device for storing instruction information based on an artificial intelligence speech model according to the present disclosure;
[0016] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Specific Embodiments
[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0018] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0019] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.
[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".
[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0022] The present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.
[0023] Reference Figure 1, which shows the flow 100 of some embodiments of the instruction information storage method based on an artificial intelligence voice model according to the present disclosure. The instruction information storage method based on an artificial intelligence voice model includes the following steps:
[0024] Step 101, obtain an object noise reduction intensity information set corresponding to a target object information set set on a target school intelligent platform.
[0025] In some embodiments, the execution subject (for example, an electronic device) of the above instruction information storage method based on an artificial intelligence voice model can obtain an object noise reduction intensity information set corresponding to a target object information set set on a target school intelligent platform through a wired connection method or a wireless connection method. Among them, the target school intelligent platform can be a processing platform for school scenarios that supports various intelligent processing. For example, the target school intelligent platform supports the following intelligent processing operations: image cropping operation, image synthesis operation, video editing operation, copywriting generation operation. The target object information can be the object information corresponding to the target object. For example, the object information can be the object identifier of the target object. For the campus scenario, the corresponding target object can be one of the following: campus personnel, campus scenarios. That is, the target object information set can include: object information for school personnel and object information for each campus scenario. There is a one-to-one correspondence between the target object information in the target object information set and the object noise reduction intensity information in the above object noise reduction intensity information set. The object noise reduction intensity information can represent the noise reduction intensity for denoising the element content of the target object in the image. The higher the noise reduction intensity, the clearer the target object in the image. For example, the object noise reduction intensity information can be, but is not limited to, at least one of the following: first-level noise reduction intensity information, second-level noise reduction intensity information, third-level noise reduction intensity information. The noise reduction intensity corresponding to the first-level noise reduction intensity information is higher than the noise reduction intensity corresponding to the second-level noise reduction intensity information. The noise reduction intensity corresponding to the second-level noise reduction intensity information is higher than the noise reduction intensity corresponding to the third-level noise reduction intensity information. The object noise reduction intensity information can also be information in numerical form. The higher the corresponding value, the higher the noise reduction intensity of the object corresponding content. In practice, for campus personnel in the campus, multiple object noise reduction intensities can be set, that is, each object noise reduction intensity information can include: multiple object noise reduction intensities. The campus scenario is a scenario represented by campus landmarks. For example, the campus landmark can be the campus library, or the campus cafeteria, or the campus gate. Multiple object noise reduction intensities can also be set for each campus scenario information.
[0026] It should be noted that the object noise reduction intensity information set and the target object information set can be obtained through specific training of a large language model. In practice, the specific training can be as follows: Obtain the training data set corresponding to the target object information set. Among them, the training data can include: the image corresponding to the target object information and the image after noise reduction under the noise reduction intensity of the image. Use the image corresponding to the target object information as the input data and the image after noise reduction as the training objective to perform model training on the initial large language model to obtain a large language model that can perform processing under the noise reduction intensity on the image corresponding to the target object information. Thus, the large language model can summarize the image noise reduction situation that can be executed to obtain the object noise reduction intensity information set and the target object information set. Then, record the object noise reduction intensity information set and the target object information set on the school intelligent platform for the object to perform image redrawing processing.
[0027] Step 102, in response to receiving the information for generating an image redrawing instruction for the target image, obtain the image description information and the image object information set corresponding to the above target image.
[0028] In some embodiments, in response to receiving the information for generating an image redrawing instruction for the target image, the above execution subject can obtain the image description information and the image object information set corresponding to the above target image.
[0029] Among them, the target image can be an image for improving the quality of the object content in the image. Specifically, the quality improvement can be the quality improvement of the object content corresponding to the object in the image. The information for generating an image redrawing instruction can be the request information indicating the generation of an image redrawing instruction. The image redrawing instruction can be the prompt content (Prompt) for redrawing the target image. By inputting the image redrawing instruction and the target image into the large language model, a corresponding redrawn image can be efficiently generated. The image description information can be the comprehensive description information of the description feature information set corresponding to the target image. That is, the image description information describes the feature content under each feature corresponding to the target image. In practice, there is a one-to-one correspondence between the description features in the description feature set and the description feature information in the description feature information set. The description feature information can be the feature content corresponding to the description feature. The description feature set can include at least one of the following: image style feature, image tone feature, image texture feature, image object category feature. The image object information in the image object information set can be the object information of the image object. In practice, the image object can be an element object in the image. That is, the image object information set can be the element information corresponding to each image element in the target image.
[0030] As an example, the above execution subject can obtain the image description information through an AI tagging tool.
[0031] Step 103: Generate a mask image corresponding to each image object information in the above image object information set to obtain a mask image set.
[0032] In some embodiments, the above execution entity may generate a mask image corresponding to each image object information in the above image object information set to obtain a mask image set.
[0033] As an example, the above execution entity may use the natural language mask generation tool DINO to generate a mask image corresponding to each image object information in the above image object information set to obtain a mask image set.
[0034] In some optional implementation manners of some embodiments, generating a mask image corresponding to each image object information in the above image object information set may include the following steps:
[0035] First step: Determine the object region information related to the content corresponding to the above image object information in the above target image. The object region information may be the position information of the image object in the target image. In practice, the object region information may be information in coordinate form. The content corresponding to the image object information may be the image content corresponding to the image object information.
[0036] As an example, the above execution entity may input the target image and the above image object information into an object region information generation model to obtain the object region information. Among them, the object region information generation model may be a neural network model that generates the object region position of the image object information in the image. In practice, the object region information generation model may be a YOLO model. Specifically, it may be any version of the model in the YOLO model. The object region information generation model may be trained through a conventional model training method. In practice, the object region information generation model may be trained through the model training method of a conventional object recognition model. That is, the model training method of a conventional YOLO model may be used for training.
[0037] Second step: Perform binarization processing on the above target image according to the above object region information to generate a binarized image as the mask image corresponding to the above image object information.
[0038] As an example, the above execution entity may set the pixel values in the object region information of the target image to a first value, and the pixel values in the remaining regions to a second value to generate a binarized image as the mask image. The first value and the second value may be preset. For example, the first value may be 12, and the corresponding second value may be 255.
[0039] Step 104: For each image object information, perform the following first generation step:
[0040] Step 1041, in response to determining that the above target object information set includes the above image object information, obtain the target object noise reduction intensity information corresponding to the above image object information from the above object noise reduction intensity information set, and determine the target mask image corresponding to the above image object information.
[0041] In some embodiments, in response to determining that the above target object information set includes the above image object information, the above execution entity may obtain the target object noise reduction intensity information corresponding to the above image object information from the above object noise reduction intensity information set, and determine the target mask image corresponding to the above image object information.
[0042] As an example, the above execution entity may query the target object noise reduction intensity information corresponding to the above image object information from the above object noise reduction intensity information set by means of querying the target object noise reduction intensity information. For the implementation manner corresponding to the target mask image, refer to the implementation of the target object noise reduction intensity information.
[0043] In some optional implementation manners of some embodiments, after step 1042, the steps further include:
[0044] First step, in response to determining that the above image object information is not included, screen out at least one target object information from the above target object information set, where the semantic similarity between the corresponding object semantics of the target object information and the corresponding object semantics of the above image object information is greater than a target degree. Among them, the object semantics may represent the object category corresponding to the image object information. The object category may be what type of object the object in the image is. The semantic similarity may be the object category similarity. Specifically, the object category corresponding to each image object information may be determined through an object category table. The object category similarity between two object categories may be determined through a preset semantic similarity table. The semantic similarity may be a value between 0 and 1, and the higher the value, the more similar the semantic content is.
[0045] Second step, determine at least one object noise reduction intensity information corresponding to the above at least one target object information. Among them, there is a one-to-one correspondence between the target object information in the at least one target object information and the object noise reduction intensity information in the above at least one object noise reduction intensity information.
[0046] Third step, determine at least one semantic similarity corresponding to the above at least one target object information. Among them, there is a one-to-one correspondence between the target object information in the at least one target object information and the semantic similarity in the above at least one semantic similarity.
[0047] Step 4: Generate at least one target intensity weight information for the above-mentioned at least one semantic similarity. There is a one-to-one correspondence between the semantic similarity in the at least one semantic similarity and the target intensity weight information in the at least one target intensity weight information. The target intensity weight information can characterize the importance degree corresponding to the object noise reduction intensity information. The higher the corresponding target intensity weight information, the more important the corresponding object noise reduction intensity information.
[0048] As an example, input the above-mentioned at least one semantic similarity into the target intensity weight information conversion model to generate at least one target intensity weight information. Among them, the target intensity weight information conversion model can be a neural network model that converts semantic similarity into corresponding target intensity weight information. That is, the target intensity weight information conversion model can represent the mapping relationship between semantic similarity and target intensity weight information. In practice, the target intensity weight information conversion model can be a fully connected layer with the output being regression content. The regression content can be the content whose model output is of the regression type. The target intensity weight information conversion model can be obtained by continuously updating the parameters of the fully connected layer through the target training dataset and the conventional model training. For another example, the target intensity weight information conversion model can also be a conventional regression model, such as a linear regression model. Through the target intensity weight information conversion model, a mapping relationship between semantic similarity and target intensity weight information is established.
[0049] In practice, obtain the training dataset between semantic similarity and target intensity weight information. Through the training dataset, update the model parameters of the initial intensity weight information conversion model to obtain the target intensity weight information conversion model.
[0050] Step 5: Multiply the target intensity weight information in the above-mentioned at least one target intensity weight information by the corresponding object noise reduction intensity information in the above-mentioned at least one object noise reduction intensity information to generate a multiplication result, and obtain at least one multiplication result.
[0051] Step 6: Determine the average value corresponding to the above-mentioned at least one multiplication result as the target image noise reduction intensity information.
[0052] Step 1042: Package the above-mentioned target object noise reduction intensity information, the above-mentioned image description information, the above-mentioned target mask image, and the above-mentioned target image to obtain the package information.
[0053] In some embodiments, the above-mentioned execution subject can package the above-mentioned target object noise reduction intensity information, the above-mentioned image description information, the above-mentioned target mask image, and the above-mentioned target image to obtain the package information.
[0054] Step 105: Use a pre-trained large language model that supports speech processing to generate image redrawing instruction information corresponding to each packaging information in the obtained packaging information set, thereby obtaining an image redrawing instruction information set.
[0055] In some embodiments, the above-mentioned execution entity can use a pre-trained large language model that supports speech processing to generate image redrawing instruction information corresponding to each packaging information in the obtained packaging information set, thereby obtaining an image redrawing instruction information set. Among them, the large language model can support speech output and text input. That is, the large language model can be a multi-modal large language model that supports multi-modal input. When the input is speech, the large language model will convert the speech into text and then process it based on the text. The large language model can be a large language model that is trained with a focus on target object information and the corresponding object noise reduction intensity information. This enables the large language model to perform noise reduction processing at the object noise reduction intensity for the object content under each target object information. The large language model for image noise reduction refers to a model that can process image noise reduction tasks through natural language processing technology. These models usually combine technologies of natural language processing (NLP), computer vision (CV), and audio processing, and utilize the powerful language understanding and generation capabilities of the large language model to achieve the understanding and processing of image content and audio content. For example, the large language model can be a model based on Stable Diffusion, or a time series denoising model based on the Transformer architecture. Stable Diffusion consists of multiple parts, including a text understanding component, an image information creator, and an image generator. The text understanding component converts text information into a digital representation and then inputs it into the image generator to generate high-quality images. Its efficient internal structure and multi-step generation process result in higher-quality images, faster running speed, and less resource consumption. The Transformer architecture captures long-range dependencies in sequential data through the self-attention mechanism. This mechanism can identify and reduce noise by learning the pixel relationships in the image when processing images. Combining with deep learning technology, the Transformer model can demonstrate superior performance in complex image noise reduction tasks. The image redrawing instruction information can be instruction information (i.e., prompt information) representing image redrawing of the target image.
[0056] As an example, first, the above-mentioned execution entity can generate prompt information representing the generation of image redrawing instruction information based on the packaging data. Then, the prompt information and the packaging data are input into the large language model to obtain the image redrawing instruction information.
[0057] In some optional implementation manners of some embodiments, the above-mentioned execution entity can use a pre-trained large language model to generate image redrawing instruction information corresponding to each packaging information in the obtained packaging information set, including the following steps:
[0058] Step 1: Obtain the redrawing requirement information filled in by the target object on the redrawing requirement processing page corresponding to the intelligent platform of the target school. Among them, the redrawing requirement information may be the information entered by the target object for the content corresponding to the redrawing map. For example, the redrawing requirement information may be: "The resolution corresponding to the redrawing map is the target resolution."
[0059] Step 2: Generate generation instruction information representing generating image redrawing instruction information according to the above redrawing requirement information and the above packaging information.
[0060] Step 3: Input the above generation instruction information into the above large language model to obtain the above image redrawing instruction information.
[0061] As an example, the above execution entity may input the generation instruction information and the above packaging information into the above large language model to obtain the above image redrawing instruction information.
[0062] Step 106: Store the above image redrawing instruction information set on the server corresponding to the intelligent platform of the target school.
[0063] In some embodiments, the above execution entity may store the above image redrawing instruction information set and the image identifier corresponding to the target image in the form of key-value pairs on the server corresponding to the intelligent platform of the target school.
[0064] In some optional implementation manners of some embodiments, after step 106, the steps further include:
[0065] Step 1: In response to receiving the image redrawing request initiated by the target object on the intelligent platform of the target school, display each image redrawing instruction information in the above image redrawing instruction information set and the instruction information display page of the above target image in the intelligent platform of the target school. Among them, the above instruction information display page includes: an instruction information editing control. The instruction information display page may be a page for displaying the image redrawing instruction information set. The instruction information editing control may be a control for editing the image redrawing instruction information. In practice, on the instruction information display page, at the target position of each image redrawing instruction information, there is a corresponding instruction information editing control.
[0066] Step 2: In response to detecting that the above-mentioned target object clicks on the selection information for at least one image redrawing instruction information on the above-mentioned instruction information display page, input the above-mentioned at least one image redrawing instruction information into the above-mentioned large language model to generate at least one redrawn image required by the above-mentioned target object. Among them, at least one image redrawing instruction information can be selected by the target object. For example, the selection information can be clicking on the image redrawing instruction information on the instruction information display page. There is a one-to-one correspondence between the redrawn images in the at least one redrawn image and the image redrawing instruction information in the at least one image redrawing instruction information. In practice, the above-mentioned at least one image redrawing instruction information and the corresponding at least one packaging information can be input into the above-mentioned large language model to generate at least one redrawn image required by the above-mentioned target object.
[0067] Step 3: Display the above-mentioned at least one redrawn image on the redrawing diagram display page in the above-mentioned target school intelligent platform. Among them, there are data transmission controls and at least one redrawing diagram regeneration control corresponding to the above-mentioned at least one redrawn image on the above-mentioned redrawing diagram display page. The data transmission control can be a control that supports the target object to transmit at least one redrawn image.
[0068] Step 4: In response to determining that the above-mentioned target object clicks on the target redrawing diagram regeneration control, pop up an instruction information display pop-up window for the image redrawing instruction information corresponding to the above-mentioned target redrawing diagram regeneration control. Among them, the above-mentioned instruction information display pop-up window includes: an instruction information editing control. The instruction information editing control can be a control for editing and processing instruction information.
[0069] Step 5: In response to determining that the above-mentioned target object clicks on the above-mentioned instruction information editing control and edits the image redrawing instruction information corresponding to the above-mentioned target redrawing diagram regeneration control, obtain the edited image redrawing instruction information.
[0070] Step 6: Store the above-mentioned edited image redrawing instruction information and display it on the above-mentioned instruction information display page.
[0071] Step 7: In response to detecting that the above-mentioned target object clicks on the selection information for the above-mentioned edited image redrawing instruction information on the above-mentioned instruction information display page, input the above-mentioned edited image redrawing instruction information into the above-mentioned large language model to generate the target redrawn image required by the above-mentioned target object.
[0072] Step 8: Display the above-mentioned target redrawn image on the above-mentioned redrawing diagram display page.
[0073] In the process of adopting technical solutions to solve the above technical problems mentioned in the background art, the following problems often arise: "Generating redrawn images using large language models is often accurate and efficient. However, for specific scenarios (e.g., scenarios that the large language model has not been trained on and lacks relevant semantic feature information), the output redrawn images are not precise enough." Considering the drawback of the lack of precise verification for the output of large language models, the inventors decided to adopt the following solution to address this issue:
[0074] In some optional implementation manners of some embodiments, after inputting the at least one image redrawing instruction information into the large language model to generate at least one redrawn image required by the target object, the method further includes:
[0075] First step, for each of the at least one redrawn image, perform a second generation step:
[0076] Sub-step 1, obtain the description feature set of the image description information corresponding to the redrawn image. Among them, each description feature in the description feature set can be a pre-determined image feature. The description feature is a feature used to describe the content of the image. In practice, the description feature set can include but is not limited to at least one of the following: image style feature, image texture feature, image tone feature, image object category feature.
[0077] Sub-step 2, obtain the image description information extraction model corresponding to the description feature set, where the image description information extraction model is a neural network model with the description feature set as the target output. That is, the output content corresponding to the image description information extraction model is the feature information set corresponding to the description feature set. The image description information extraction model can be a neural network model that extracts the feature content under the description feature set in the image. That is, the model input of the image description information extraction model is an image, and the output is the respective description feature information corresponding to the image. The image description information extraction model can extract all-round feature information through a multi-layer cascaded convolutional layer (e.g., an 11-layer cascaded convolutional layer), and then output the feature content under each description feature in the description feature set through multiple parallel fully connected layers. The image description information extraction model can be trained in cooperation with the corresponding image application scenario. For example, the image description information extraction model can be trained in cooperation with an image recognition model. That is, the image description information extraction model can serve as the feature extraction module of the image recognition model.
[0078] Sub-step 3: Input the above redrawn image into the above image description information extraction model to generate redrawn image description information. The redrawn image description information includes a redrawn description feature information set corresponding to the above description feature set. There is a one-to-one correspondence between the description features in the description feature set and the redrawn description feature information in the redrawn description feature information set. The redrawn description feature information may be the image description feature information corresponding to the redrawn image. The redrawn description feature information may be the feature content under the description feature for the redrawn image.
[0079] Sub-step 4: Determine the description feature information set corresponding to the above description feature set included in the above image description information. There is a one-to-one correspondence between the description features in the description feature set and the description feature information in the description feature information set. The description feature information may be the feature content corresponding to the description feature.
[0080] Sub-step 5: Generate the feature information difference between the above redrawn description feature information set and the above description feature information set.
[0081] As an example, first, the above execution entity can perform information subtraction on each redrawn description feature information in the redrawn description feature information set and the corresponding description feature information in the description feature information set to generate information differences, obtaining an information difference set. Then, determine the average value corresponding to the information difference set as the feature information difference.
[0082] Sub-step 6: In response to determining that the above feature information difference is less than a predetermined information difference, generate an initial mask image corresponding to the above redrawn image. The predetermined information difference may be a pre-determined information difference. For example, the preset information difference may be a pre-set information difference value. When the feature information difference is higher than or equal to the preset information difference, it indicates a relatively large feature difference. When the feature information difference is less than the preset information difference, it indicates a relatively small feature difference. Specifically, the generation method of the initial mask image can refer to the generation method of the mask image.
[0083] Sub-step 7: In response to determining that the mask image difference between the above initial mask image and the corresponding target mask image is less than the predetermined difference information, input the above target image and the above redrawn image into an object noise reduction intensity information generation model to generate noise reduction intensity information. Among them, the mask image difference can be the image pixel difference between the above initial mask image and the above target mask image. The object noise reduction intensity information generation model can be a neural network model for generating object noise reduction intensity information. The object noise reduction intensity information generation model can be a neural network model for generating the noise reduction difference between the target image and the redrawn image. In practice, the object noise reduction intensity information generation model can be a convolutional neural network model connected in series in multiple layers. Specifically, the object noise reduction intensity information generation model can include: a first feature extraction layer for extracting the semantic content of the image features corresponding to the target image, a second feature extraction layer for extracting the semantic content of the image features corresponding to the redrawn image, a feature fusion layer for fusing the corresponding outputs of the first feature extraction layer and the second feature extraction layer, and a fully connected layer connected in series for regression output. The first feature extraction layer and the second feature extraction layer can be convolutional layers connected in series with different numbers of layers. The feature fusion layer can be a feature information splicing layer.
[0084] Step 2: In response to determining that the difference between the above noise reduction intensity information and the corresponding object noise reduction intensity information is less than the preset difference, generate information indicating that the above redrawn image passes the image verification.
[0085] Optionally, in response to determining that the difference between the above noise reduction intensity information and the corresponding object noise reduction intensity information is greater than or equal to the preset difference, generate information indicating that the above redrawn image fails the image verification.
[0086] Optionally, the above-mentioned image description information extraction model includes: a first image feature extraction model, a first set of description feature information extraction models corresponding to the above-mentioned description feature set, a first set of description feature information output layers corresponding to the above-mentioned description feature set, a second image feature extraction model, a second set of description feature information output layers corresponding to the above-mentioned description feature set, and a copywriting output layer. The second image feature extraction model includes: a plurality of serially-connected convolutional neural networks. The plurality of serially-connected convolutional neural networks includes: at least one serially-connected convolutional neural network, and the target position of the at least one serially-connected convolutional neural network in the plurality of serially-connected convolutional neural networks. The target position can be the first predetermined number of positions. There is a one-to-one correspondence between the convolutional neural network in the at least one serially-connected convolutional neural network and the second description feature information output layer in the second set of description feature information output layers. Among them, the first image feature extraction model can be a neural network model for extracting image feature information. In practice, the first image feature extraction model can be an 8-layer serially-connected convolutional neural network. There is a one-to-one correspondence between the first description feature information extraction models in the first set of description feature information extraction models and the description features in the description feature set. That is, the first description feature information extraction model can be a neural network model for specifically extracting features for the description features. In practice, the first description feature information extraction model can be a multi-layer serially-connected convolutional neural network. There is a one-to-one correspondence between the description features in the description feature set and the first description feature information output layers in the first set of description feature information output layers. The first description feature information output layer can be a fully-connected layer. The second image feature extraction model can be a neural network model for extracting image feature information. There is a one-to-one correspondence between the description features in the above-mentioned description feature set and the second description feature information output layers in the second set of description feature information output layers. In practice, the second description feature information output layer can be a multi-layer serially-connected convolutional neural network. The copywriting output layer can be a network layer for generating copywriting. In practice, the copywriting output layer can be a recurrent neural network model. The plurality of serially-connected convolutional neural networks can be a 10-layer serially-connected convolutional neural network. The number of network layers corresponding to the at least one serially-connected convolutional neural network is the same as the number of output layers corresponding to the second set of description feature information output layers.
[0087] Optionally, inputting the above-mentioned redrawn image into the above-mentioned image description information extraction model to generate redrawn image description information may include the following steps:
[0088] First step, input the above-mentioned redrawn image into the above-mentioned first image feature extraction model to generate first image feature information. The first image feature information can be information in the form of a vector representing the semantic content of the image features corresponding to the redrawn image.
[0089] In the second step, input the above first image feature information into each first description feature information extraction model in the above first description feature information extraction model set to generate first description feature extraction information, and obtain a first description feature extraction information set.
[0090] In the third step, input each first description feature extraction information in the above first description feature extraction information set into the corresponding first description feature information output layer in the above first description feature information output layer set to output first initial description feature information, and obtain a first initial description feature information set.
[0091] In the fourth step, input the above first description feature extraction information set into the above second image feature extraction model to generate at least one convolution output result for the above at least one serially connected convolutional neural network and an overall convolution output result for the above multiple serially connected convolutional neural networks.
[0092] In the fifth step, input each convolution output result in the above at least one convolution output result into the corresponding second description feature information output layer in the above second description feature information output layer set to generate second description feature information, and obtain a second initial description feature information set.
[0093] In the sixth step, determine the information difference between the above first initial description feature information set and the above second initial description feature information set. The specific implementation method will not be elaborated here.
[0094] In the seventh step, in response to determining that the above information difference is less than a predetermined degree difference, input the above overall convolution output result into the above copywriting output layer to generate copywriting information. Among them, the copywriting information can be the description copywriting corresponding to the redrawn image.
[0095] In the eighth step, generate the above redrawn image description information according to the above copywriting information, the above first initial description feature information set, and the above second initial description feature information set.
[0096] As an example, first, perform feature information extraction on the above copywriting information to generate a fourth initial description feature information set. Then, according to the voting mechanism, screen out a first actual description feature information set corresponding to the description feature set from the above first initial description feature information set, the above second initial description feature information set, and the above fourth initial description feature information set. Finally, determine the above first actual description feature information set as the redrawn image description information.
[0097] Optionally, the above image description information extraction model further includes: at least one second description feature information extraction model corresponding to the at least one serially connected convolutional neural network, and a third description feature information output layer set. Among them, there is a one-to-one correspondence between the convolutional neural network in the at least one serially connected convolutional neural network and the second description feature information extraction model in the at least one second description feature information extraction model. There is a one-to-one correspondence between the third description feature information output layer in the third description feature information output layer set and the description feature in the description feature set.
[0098] Optionally, the steps further include:
[0099] First step, in response to determining that the above information difference is greater than or equal to the above predetermined degree difference, for each of the above at least one convolutional output results, perform the following second generation steps:
[0100] First sub-step, input the above convolutional output result into the corresponding second description feature information extraction model in the above at least one second description feature information extraction model to generate second description feature extraction information.
[0101] Second sub-step, input the above second description feature extraction information into the corresponding third description feature information output layer in the above third description feature information output layer set to output third initial description feature information.
[0102] Second step, generate the above redrawn image description information according to the above first initial description feature information set, the obtained third initial description feature information set, and the above second initial description feature information set.
[0103] As an example, first, the above execution entity can use a voting mechanism to generate a second actual description feature information set corresponding to the description feature set according to the first initial description feature information set, the obtained third initial description feature information set, and the above second initial description feature information set. Finally, determine the above second actual description feature information set as the redrawn image description information.
[0104] Optionally, after inputting the at least one image redrawing instruction information into the large language model to generate at least one redrawn image required by the target object as described above, image verification of the redrawn image is another inventive point of the present disclosure, which solves the problem that "the output of the large language model is not accurate enough, resulting in a large content deviation in the redrawn image". Based on this, in the present disclosure, for each redrawn image, by extracting the feature information difference between the description feature information corresponding to the redrawn image and the description feature information in the original image description information, it is preliminarily determined whether there is a problem of a large feature deviation. On this basis, by generating an initial mask image corresponding to the redrawn image and comparing the difference with the corresponding target mask image of the original, the noise reduction intensity is further determined. When the noise reduction intensity is less than a predetermined difference, it is determined that there is no problem with the generation of the redrawn image. In addition, in the process of preliminarily determining the feature deviation, it is crucial to extract the redrawing feature information in the redrawing diagram. Based on this, through the first image feature extraction model included in the above image description information extraction model, the first description feature information extraction model set corresponding to the description feature set, the first description feature information output layer set corresponding to the description feature set, the second image feature extraction model, the second description feature information output layer set corresponding to the description feature set, and the copywriting output layer, as well as the specific design of the corresponding model structure, accurate extraction of the description feature information can be achieved, ensuring the accuracy of determining the feature information difference.
[0105] The above embodiments of the present disclosure have the following beneficial effects: The method for storing instruction information based on an artificial intelligence voice model according to some embodiments of the present disclosure can efficiently and high-quality generate image redrawing instructions for a target image to meet the diverse image needs of the target object. Specifically, the reason for the inefficient generation of relevant image redrawing instructions is that the efficiency of manually customizing input image redrawing instruction information on the school intelligent platform is low. It is necessary to continuously repeat and trial-and-error the image redrawing instruction information, which will occupy the large language model, resulting in the continuous occupation of computing resources and the waste of a large amount of labor costs. Based on this, the method for storing instruction information based on an artificial intelligence voice model according to some embodiments of the present disclosure, first, obtains the object noise reduction intensity information set corresponding to the target object information set set on the target school intelligent platform. Here, by obtaining the object noise reduction intensity information set corresponding to the target object information set, the determination efficiency of the image noise reduction intensity information of the subsequent image object information can be greatly improved. Then, in response to receiving the image redrawing instruction generation information for the target image, obtains the image description information and the image object information set corresponding to the above target image for subsequent generation of corresponding image redrawing instructions for each image object information. Next, generates a mask image corresponding to each image object information in the above image object information set to obtain a mask image set, so as to facilitate the high-quality replacement of the image object for the mask image and be used for subsequent generation of high-quality redrawn images. Then, for each image object information in the above image object information set, the following first generation step is performed: The first step, in response to determining that the above target object information set includes the above image object information, obtains the target object noise reduction intensity information corresponding to the above image object information from the above object noise reduction intensity information set, and determines the target mask image corresponding to the above image object information to provide a data basis for subsequent generation of image redrawing instruction information. The third step, packs the above target object noise reduction intensity information, the above image description information, the above target mask image and the above target image to obtain a packed information to achieve data integration. Using the pre-trained large language model, the image redrawing instruction information corresponding to each packed information in the obtained packed information set can be accurately and efficiently generated to obtain an image redrawing instruction information set, wherein the above large language model is a language model specifically trained based on the above target object information set and the object noise reduction intensity information set. Finally, stores the above image redrawing instruction information set on the server corresponding to the above target school intelligent platform for subsequent generation of redrawn images. In summary, by determining whether the image object information is in the target object information set, the corresponding object noise reduction intensity information can be quickly determined. In addition, through the large language model, the image redrawing instruction information for the packed data can be efficiently and accurately generated, so that high-quality redrawn images can be obtained subsequently.
[0106] Further reference Figure 2, as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an instruction information storage device based on an artificial intelligence speech model, and these device embodiments correspond to Figure 1 the method embodiments shown. The instruction information storage device based on the artificial intelligence speech model can be specifically applied to various electronic devices.
[0107] As Figure 2 shown, an instruction information storage device 200 based on an artificial intelligence speech model includes: a first acquisition unit 201, a second acquisition unit 202, a first generation unit 203, an execution unit 204, a second generation unit 205, and an instruction information storage unit 206. Among them, the first acquisition unit 201 is configured to acquire an object noise reduction intensity information set corresponding to a target object information set set on a target school intelligent platform; the second acquisition unit 202 is configured to acquire image description information and an image object information set corresponding to the target image in response to generating information for an image redrawing instruction for the target image; the first generation unit 203 is configured to generate a mask image corresponding to each image object information in the image object information set to obtain a mask image set; the execution unit 204 is configured to, for each image object information, perform the following first generation step: in response to determining that the target object information set includes the image object information, acquire target object noise reduction intensity information corresponding to the image object information from the object noise reduction intensity information set, and determine a target mask image corresponding to the image object information; package the target object noise reduction intensity information, the image description information, the target mask image, and the target image to obtain packaged information; the second generation unit 205 is configured to use a pre-trained large language model supporting speech processing to generate image redrawing instruction information corresponding to each packaged information in the obtained packaged information set to obtain an image redrawing instruction information set, where the large language model is a language model specifically trained based on the target object information set and the object noise reduction intensity information set; the instruction information storage unit 206 is configured to store the image redrawing instruction information set on a server corresponding to the target school intelligent platform.
[0108] It can be understood that the units described in the instruction information storage device 200 based on the artificial intelligence speech model correspond to each step in the method described with reference to Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the instruction information storage device 200 based on the artificial intelligence speech model and the units included therein, and will not be elaborated here.
[0109] Next, with reference to Figure 3, which shows a schematic structural diagram of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0110] As Figure 3 shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0111] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wireline to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included. Figure 3 Each block shown in
[0112] Specifically, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the methods of some embodiments of the present disclosure are performed.
[0113] It should be noted that, in some embodiments of the present disclosure, the above-mentioned computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0114] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0115] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: obtain an object noise reduction intensity information set corresponding to a target object information set set on a target school intelligent platform; in response to generating information in response to an image redrawing instruction for a target image, obtain image description information and an image object information set corresponding to the target image; generate a mask image corresponding to each image object information in the image object information set to obtain a mask image set; for each image object information, perform the following first generation step: in response to determining that the target object information set includes the image object information, obtain target object noise reduction intensity information corresponding to the image object information from the object noise reduction intensity information set, and determine a target mask image corresponding to the image object information; package the target object noise reduction intensity information, the image description information, the target mask image, and the target image to obtain packaged information; use a pre-trained large language model supporting speech processing to generate image redrawing instruction information corresponding to each packaged information in the obtained packaged information set to obtain an image redrawing instruction information set, where the large language model is a language model specifically trained based on the target object information set and the object noise reduction intensity information set; store the image redrawing instruction information set in a server corresponding to the target school intelligent platform.
[0116] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0118] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a first acquisition unit, a second acquisition unit, a first generation unit, an execution unit, a second generation unit, and an instruction information storage unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the instruction information storage unit can also be described as "the unit that stores the above-mentioned image redrawing instruction information set in the server corresponding to the above-mentioned target school intelligent platform".
[0119] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0120] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.
Claims
1. An instruction information storage method based on an artificial intelligence voice model, comprising: Obtaining an object noise reduction intensity information set corresponding to a target object information set set on a target school intelligent platform; In response to receiving information generated by an image redrawing instruction for a target image, obtaining image description information and an image object information set corresponding to the target image; Generating a mask image corresponding to each image object information in the image object information set to obtain a mask image set; For each image object information, perform the following first generation step: In response to determining that the target object information set includes the image object information, obtaining target object noise reduction intensity information corresponding to the image object information from the object noise reduction intensity information set, and determining a target mask image corresponding to the image object information; Packing the target object noise reduction intensity information, the image description information, the target mask image, and the target image to obtain packed information; Using a pre-trained large language model supporting speech processing to generate image redrawing instruction information corresponding to each piece of packed information in the obtained packed information set to obtain an image redrawing instruction information set, wherein the large language model is a language model specifically trained based on the target object information set and the object noise reduction intensity information set; Storing the image redrawing instruction information set on the server corresponding to the target school intelligent platform.
2. The method according to claim 1, wherein, After the step of, in response to determining that the target object information set includes the image object information, obtaining target object noise reduction intensity information corresponding to the image object information from the object noise reduction intensity information set, and determining a target mask image corresponding to the image object information, the method further includes: In response to determining that the image object information is not included, screening out at least one target object information from the target object information set whose semantic similarity between the corresponding object semantics and the corresponding object semantics of the image object information is greater than a target degree; Determining at least one object noise reduction intensity information corresponding to the at least one target object information; Determining at least one semantic similarity corresponding to the at least one target object information; Generating at least one target intensity weight information for the at least one semantic similarity; Multiplying the target intensity weight information in the at least one target intensity weight information by the corresponding object noise reduction intensity information in the at least one object noise reduction intensity information to generate a multiplication result to obtain at least one multiplication result; Determining an average value corresponding to the at least one multiplication result as the target object noise reduction intensity information.
3. The method according to claim 1, wherein The step of generating a mask image corresponding to each image object information in the image object information set includes: Determining object region information in the target image related to the content corresponding to the image object information; Performing binarization processing on the target image according to the object region information to generate a binarized image as the mask image corresponding to the image object information.
4. The method according to claim 1, wherein The step of using a pre-trained large language model supporting speech processing to generate image redrawing instruction information corresponding to each piece of packed information in the obtained packed information set includes: Obtain the redrawing requirement information filled in by the target object on the redrawing requirement processing page corresponding to the intelligent platform of the target school; Generate generation instruction information indicating that image redrawing instruction information is generated according to the redrawing requirement information and the packaging information; Input the generation instruction information into the large language model to obtain the image redrawing instruction information.
5. The method according to claim 1, wherein, The method further includes: In response to receiving an image redrawing request initiated by the target object on the intelligent platform of the target school, display each image redrawing instruction information in the image redrawing instruction information set and the instruction information display page of the target image on the intelligent platform of the target school, where the instruction information display page includes: an instruction information editing control; In response to detecting that the target object clicks on the selection information for at least one image redrawing instruction information on the instruction information display page, input the at least one image redrawing instruction information into the large language model to generate at least one redrawn image required by the target object; Display the at least one redrawn image on the redrawing diagram display page in the intelligent platform of the target school, where the redrawing diagram display page has a data transmission control and at least one redrawing diagram regeneration control corresponding to the at least one redrawn image; In response to determining that the target object clicks on the target redrawing diagram regeneration control, pop up an instruction information display pop-up window for the image redrawing instruction information corresponding to the target redrawing diagram regeneration control, where the instruction information display pop-up window includes: an instruction information editing control; In response to determining that the target object clicks on the instruction information editing control and edits the image redrawing instruction information corresponding to the target redrawing diagram regeneration control, obtain the edited image redrawing instruction information; Store the edited image redrawing instruction information and display it on the instruction information display page; In response to detecting that the target object clicks on the selection information for the edited image redrawing instruction information on the instruction information display page, input the edited image redrawing instruction information into the large language model to generate the target redrawn image required by the target object; Display the target redrawn image on the redrawing diagram display page.
6. An instruction information storage device based on an artificial intelligence voice model, including: A first acquisition unit configured to acquire an object noise reduction intensity information set corresponding to a target object information set set on an intelligent platform of a target school; A second acquisition unit configured to, in response to receiving image redrawing instruction generation information for a target image, acquire image description information and an image object information set corresponding to the target image; A first generation unit configured to generate a mask image corresponding to each image object information in the image object information set to obtain a mask image set; An execution unit, configured to, for each piece of image object information, perform the following first generation step: in response to determining that the target object information set includes the image object information, obtain target object noise reduction intensity information corresponding to the image object information from the object noise reduction intensity information set, and determine a target mask image corresponding to the image object information; Pack the target object noise reduction intensity information, the image description information, the target mask image, and the target image to obtain packed information; A second generation unit, configured to use a pre-trained large language model supporting speech processing to generate image redrawing instruction information corresponding to each piece of packed information in the obtained packed information set, to obtain an image redrawing instruction information set, where the large language model is a language model specifically trained based on the target object information set and the object noise reduction intensity information set; An instruction information storage unit, configured to store the image redrawing instruction information set on a server corresponding to the target school intelligent platform.
7. An electronic device, comprising: One or more processors; A storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-5.
8. A computer-readable medium having a computer program stored thereon, wherein, The program, when executed by the processor, implements the method according to any one of claims 1-5.
Citation Information
Patent Citations
Image redrawing method and device, computer equipment and storage medium
CN117808917A
Image processing method and apparatus, electronic device and storage medium
US20250086806A1