Interesting medicine science popularization production system and method based on artificial intelligence
By comparing and adjusting the input keyword vectors and image vectors in the popular image generation model of interesting medicine, the problem of insufficient fun in the generation process is solved, and more efficient image generation is achieved, and the generated images are more in line with user needs.
Patent Information
- Application Number
- CN202510748046.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing popular image generation model of interesting medicine science lacks targeted adjustments during the generation process, resulting in the generated images being less interesting and unable to effectively attract the target audience.
By comparing the differences between the input keyword vectors, output keyword vectors and image vectors when the adjacent image is generated, the effective indicators of the description of the input text and the personified eigenvalues are quantified, and the input keyword vectors are adjusted to balance the authenticity and fun of the image.
It improves the accuracy of popularization of interesting medical science images, reduces the number of generations, and the generated images are more in line with user needs and expectations.
Smart Images

Figure CN120298533A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of popular science image generation, and in particular to an interesting medical popular science production system and method based on artificial intelligence. Background Art
[0002] Interesting popular science medical knowledge posters and comics are an effective way to spread medical and health knowledge through vivid and interesting visual representations. Their content usually uses bright colors, illustrated anthropomorphic images, and humorous character designs, aiming to convey health information to the public in a simple and clear manner, covering topics such as nutrition, exercise, mental health, disease prevention, etc.
[0003] The existing generation of interesting popular science medical knowledge posters mainly relies on artificial intelligence models that can process popular science texts. Their generation process depends on the popular science text and descriptive text input by users, and multiple user inputs and adjustments are required to meet the needs of users. However, there is often a lack of targeted adjustment in terms of interestingness, which easily leads to lower interestingness of the generated images, unable to effectively attract the target audience, and thus multiple adjustments are needed. Summary of the Invention
[0004] In order to solve the technical problem that the existing generation of interesting popular science medical knowledge posters mainly relies on artificial intelligence models that can process popular science texts, and their generation process depends on the popular science text and descriptive text input by users, and multiple user inputs and adjustments are required to meet the needs of users. However, there is often a lack of targeted adjustment in terms of interestingness, which easily leads to lower interestingness of the generated images, the purpose of the present invention is to provide an interesting medical popular science production system and method based on artificial intelligence, and the specific technical solutions adopted are as follows:
[0005] In the production process of interesting medical popular science images, obtain the input keyword vector corresponding to the input text, the output keyword vector corresponding to the output image, and the image vector each time an image is generated;
[0006] Compare the differences between the input keyword vectors, the differences between the output keyword vectors, and the differences between the image vectors during two adjacent image generations, and determine the description effectiveness index of the input text each time an image is generated;
[0007] Obtain the anthropomorphic vector of each image generation and compare it with the anthropomorphic vector of the comparison image, so as to quantify the anthropomorphic eigenvalue vector of the input text each time an image is generated; based on the description effectiveness index and the anthropomorphic eigenvalue vector of the input text each time an image is generated, determine the interestingness feature vector of the input text;
[0008] Adjust the input keyword vector for the input text in the next image generation based on the interesting feature vector of the input text during the current image generation to obtain an adjusted input keyword vector, which is used in the production process of interesting medical science popularization pictures to obtain the next output image.
[0009] Furthermore, the method for obtaining the described effective index includes:
[0010] Compare the differences between the input keyword vectors during two adjacent image generations to determine the text change index;
[0011] Compare the differences between the output keyword vectors during two adjacent image generations to determine the content change index;
[0012] Compare the differences between the image vectors during two adjacent image generations to determine the image change index;
[0013] During two adjacent image generations, take the value obtained by performing negative correlation mapping and normalization on the product of the text change index and the content change index as the effective factor;
[0014] Take the value obtained by normalizing the product of the image change index and the effective factor as the description effective index of the input text during the latter image generation in two adjacent image generations.
[0015] Furthermore, the method for obtaining the text change index includes:
[0016] Compare two input keyword vectors to determine the low-dimensional vector and the high-dimensional vector, and perform zero-padding on the low-dimensional vector to obtain a padded vector;
[0017] In the two input keyword vectors, take the value obtained by performing negative correlation mapping and normalization on the cosine similarity between the high-dimensional vector and the padded vector as the text change index.
[0018] Furthermore, the method for obtaining the content change index includes:
[0019] Compare two output keyword vectors to determine the low-dimensional vector and the high-dimensional vector, and perform zero-padding on the low-dimensional vector to obtain a feature vector;
[0020] In the two output keyword vectors, take the value obtained by performing negative correlation mapping and normalization on the cosine similarity between the high-dimensional vector and the feature vector as the content change index.
[0021] Furthermore, the method for obtaining the image change index includes:
[0022] Calculate the Euclidean distance between two image vectors as the image change index.
[0023] Further, the method for obtaining the anthropomorphic eigenvalue vector includes:
[0024] In each image generation, the input keyword vector is used as the input of a pre-trained neural network to obtain the anthropomorphic vector corresponding to the generated image and the corresponding relationship between the element values in the anthropomorphic vector and the element values in the output keyword vector;
[0025] Obtain the anthropomorphic vector of the comparison image and use it as the comparison vector, where the comparison image includes the original electron microscope image and the standard portrait;
[0026] In the anthropomorphic vector in each image generation and each comparison vector, the cosine similarity between the element values at the same position is used as the feature factor;
[0027] Multiply the value obtained by performing negative correlation mapping on the feature factor between the element value at each position in the anthropomorphic vector of the generated image and the original electron microscope image by the feature factor between the element value at each position in the anthropomorphic vector of the generated image and the standard portrait, and normalize the obtained product. The value is used as the element value at the same position in the anthropomorphic eigenvalue vector, so as to obtain the anthropomorphic eigenvalue vector corresponding to the input text in each image generation.
[0028] Further, the method for obtaining the interestingness feature vector includes:
[0029] In each image generation, the number of element values in the input keyword vector corresponding to each element value in the anthropomorphic eigenvalue vector is used as the quantity factor of each element value in the anthropomorphic eigenvalue vector;
[0030] In each image generation, the product of the quantity factor of each element value in the anthropomorphic eigenvalue vector and the description effectiveness index of the input text is used as the adjustment parameter corresponding to each element value in the anthropomorphic eigenvalue vector;
[0031] Normalize the product of each element value in the anthropomorphic eigenvalue vector and the corresponding adjustment parameter, and use the value as the element value at the corresponding position in the interestingness feature vector, so as to obtain the interestingness feature vector of the input text in each image generation.
[0032] Further, the acquisition of the input keyword adjustment vector includes:
[0033] After aligning the interestingness feature vector of the input text in the current image generation and the input keyword vector of the input text in the next image generation based on the zero-padding method, a vector group is obtained;
[0034] In the vector group, the vector corresponding to the input keyword vector is used as the first vector, and the vector corresponding to the interestingness feature vector is used as the second vector;
[0035] In the first vector and the second vector, for the element values at the same position, the product of the first balance coefficient and the element value in the first vector is used as the first adjustment factor, and the product of the second balance coefficient and the element value in the second vector is used as the second adjustment factor. The sum value of the first adjustment factor and the second adjustment factor is used as the element value at the same position in the input keyword adjustment vector, so as to obtain the input keyword adjustment vector of the input text for the next image generation.
[0036] Further, in the production process of the interesting medical science popularization map, obtaining the input keyword vector corresponding to the input text, the output keyword vector corresponding to the output image, and the image vector for each image generation includes:
[0037] In the generation process of the interesting medical science popularization map, the image vector of the output image is extracted by a pre-trained CLIP model, the input keyword vector is obtained by processing the input text based on the BioBERT model, and the output keyword vector corresponding to the output image is generated by an image description model.
[0038] An interesting medical science popularization production system based on artificial intelligence includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory. When at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor, the steps of the interesting medical science popularization production method based on artificial intelligence are implemented.
[0039] The present invention has the following beneficial effects:
[0040] The generation process of interesting medical popular science pictures is a process that requires continuous adjustment. The model automatically generates images based on the input text. Subsequently, users can continue to input text to adjust the images in order to obtain interesting medical popular science images that meet the requirements and can effectively attract the target audience. During the production process of interesting medical popular science pictures, for each image generation, the present invention can obtain the input keyword vector, output keyword vector, and image vector corresponding to the image generation, which can be used to analyze the image generation situation in subsequent processes. Each time text is input, it is actually another processing of the image. In order to quantify the effectiveness of the input text for image processing, the input keyword vectors, output keywords, and image vectors during adjacent image generations can be compared, so as to obtain the description effectiveness index of the input text for each image generation. Further, due to the influence of the interestingness and authenticity of the image, the main issue lies in the trade-off between retaining the original morphological characteristics and the degree of anthropomorphism. In order to avoid the lack of interestingness due to excessive authenticity of the generated image, it is necessary to further analyze the expression methods of interestingness. The anthropomorphic vectors of each image generation and the anthropomorphic vectors of the comparison images can be obtained and compared to obtain the anthropomorphic eigenvalue vector of the input text for each image generation, which is used to reflect the anthropomorphic characteristics of the output image; furthermore, by combining the description effectiveness index and the anthropomorphic eigenvalue vector of the input text for each image generation, the interestingness feature vector of the input text can be determined, which can characterize the interestingness feature. Finally, the model is improved based on the interestingness feature vector, that is, the input keyword vector of the input text for the next image generation is adjusted to obtain the input keyword adjustment vector, which balances authenticity and interestingness, thereby improving the accuracy of the generation result, effectively reducing the number of generations, and generating more interesting medical popular science images that meet the user's needs and expectations. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0042] Figure 1 It is a flowchart of a method for an interesting medical popular science production method based on artificial intelligence provided by an embodiment of the present invention;
[0043] Figure 2 It is a flowchart of a method for obtaining a description effectiveness index provided by an embodiment of the present invention;
[0044] Figure 3 It is a system block diagram of an interesting medical popular science production system based on artificial intelligence provided by an embodiment of the present invention;
[0045] Figure 4 The figure shows a schematic structural diagram of a fun medical science popularization production system based on artificial intelligence provided by an embodiment of the present invention. Detailed implementation manners
[0046] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following combines the accompanying drawings and preferred embodiments to specifically describe a fun medical science popularization production system and method based on artificial intelligence proposed by the present invention, including its specific implementation manners, structures, features and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0048] The following specifically describes the specific solutions of a fun medical science popularization production system and method based on artificial intelligence provided by the present invention with reference to the accompanying drawings.
[0049] Please refer to Figure 1 , which shows a method flow chart of a fun medical science popularization production method based on artificial intelligence provided by an embodiment of the present invention. The method includes the following steps:
[0050] Step S1: During the production of fun medical science popularization pictures, obtain the input keyword vector corresponding to the input text, the output keyword vector corresponding to the output image, and the image vector each time an image is generated.
[0051] Fun popular science medical knowledge posters and comics are an effective way to spread medical and health knowledge through vivid and interesting visual presentations. Through story-telling and humorous elements, these forms not only enhance the audience's memory but also promote information sharing, reduce learning barriers, and are applicable to school education, public health promotion, and social media marketing.
[0052] The actual generation process of fun popular science pictures is realized by the input text and anthropomorphic feature descriptions. The input text is used to provide actual science popularization information keywords to confirm the element type information of the final finished fun popular science pictures, while the fun attributes more depend on the anthropomorphic features of different element types, thereby realizing the fun evaluation of the generated pictures.
[0053] DALL-E 2 is an advanced image generation model developed by OpenAI, capable of generating high-quality images based on text descriptions provided by users. This model can understand complex text prompts and generate vivid images that match them. In addition, it also supports image editing functions, allowing users to make subtle adjustments to the generated images or generate different variants of a certain image, thus achieving more flexible creation. Therefore, in this embodiment of the present invention, interesting medical science popularization pictures are mainly obtained based on this model.
[0054] First, in the production process of interesting medical science popularization pictures, it is necessary to go through input text - output image - input text (for adjustment purposes) - output image… until an interesting medical science popularization picture that meets the user's needs is obtained. During this process, it is necessary to obtain the input keyword vector corresponding to the input text, the output keyword vector corresponding to the output image, and the image vector of the output image each time an image is generated. Preferably, in an embodiment of the present invention, the image vector of the output image can be extracted through a pre-trained CLIP model, the input keyword vector can be obtained by processing the input text based on the BioBERT model, and the output keyword vector corresponding to the output image can be generated through an image description model (such as ViT-GPT2).
[0055] It should be noted that the DALL-E 2 model, CLIP model, BioBERT model, and ViT-GPT2 model are all well-known technologies, and the specific processes will not be elaborated here.
[0056] Step S2: Compare the differences between the input keyword vectors, the differences between the output keyword vectors, and the differences between the image vectors during two adjacent image generations to determine the description effectiveness index of the input text for each image generation.
[0057] The main factor affecting the accuracy of the DALL-E 2 model's parsing of the relationship between text and image is the accuracy of the description of the morphological characteristics of the popular science object required by the descriptive text. For the keywords in the user's input text, they have a text anchor effect. That is, when the keyword exists in the input text, there will be a corresponding element description in the image generated by the DALL-E 2 model. When the keyword does not exist, it will achieve fuzzy matching description, and its anchor effect is not obvious enough. For users, the change of the input text on the image obtained from the previous input text is actually a reprocessing of the generated image. By comparing the differences between the input keyword vectors, the differences between the output keyword vectors, and the differences between the image vectors during two adjacent image generations, the description effectiveness index of the input text for each image generation can be determined, which is used to reflect the corresponding relationship between the input text and the output image and characterize the effective degree of the change in the style information of the input text.
[0058] Preferably, in an embodiment of the present invention, the method for obtaining the effective index includes:
[0059] Please refer to Figure 2 , which shows the flowchart of the method for obtaining the effective index in an embodiment of the present invention. The method includes the following steps:
[0060] For the input keyword vector corresponding to the input text during a certain image generation, the smaller the change compared to the previous time, the smaller the change in the output keyword vector of the output image, and the greater the change in the corresponding output image content. Then it indicates that the change in the style of the text information of the input text does not damage the accuracy of science popularization and can directly act on the style dimension, so the degree of usefulness is higher.
[0061] Step S201: Compare the difference between the input keyword vectors during two adjacent image generations to determine the text change index.
[0062] In view of the fact that the dimensions of the input keyword vectors during two adjacent image generations may be inconsistent, so first compare the two input keyword vectors to determine the low-dimensional vector and the high-dimensional vector, and perform zero-padding on the low-dimensional vector to obtain the padded vector, so that the two vectors have the same dimension.
[0063] Then, calculate the cosine similarity between the high-dimensional vector and the padded vector among the two input keyword vectors. The greater the cosine similarity, the more similar the input texts are, and the smaller the adjustment degree. Therefore, the value obtained by performing negative correlation mapping and normalization on the cosine similarity is used as the text change index. The larger the text change index, the greater the difference between the input texts of two adjacent times, and the more obvious the anchoring effect of the input keyword vector will be. The negative correlation mapping and normalization operations here can be performed using the formula , where is the exponential function with the natural constant e as the base, and x represents the independent variable.
[0064] Step S202: Compare the difference between the output keyword vectors during two adjacent image generations to determine the content change index.
[0065] Similarly, in view of the fact that the dimensions of the output keyword vectors corresponding to the output images during two adjacent image generations may be inconsistent, so compare the two output keyword vectors to determine the low-dimensional vector and the high-dimensional vector, and perform zero-padding on the low-dimensional vector to obtain the feature vector.
[0066] Then, among the two output keyword vectors, the value obtained by performing negative correlation mapping and normalization on the cosine similarity between the high-dimensional vector and the feature vector is used as the content change index. The larger the cosine similarity here, the less information is lost, indicating that the styles of the input keyword vectors are relatively similar. Therefore, after using negative correlation mapping and normalization to correct the logical relationship, the larger the content change index obtained, the greater the style change of the picture, and the lower the accuracy of the science popularization. The negative correlation mapping and normalization operation here can be performed using the formula , where represents the exponential function with the natural constant e as the base, and x represents the independent variable.
[0067] Step S203: Compare the differences between the image vectors during two adjacent image generations to determine the image change index.
[0068] Since the dimensions of the image vectors are the same, when measuring the differences between the image vectors, the Euclidean distance between the two image vectors can be directly calculated as the image change index. At this time, the larger the image change index, the more obvious the style change of the image.
[0069] Step S204: During two adjacent image generations, comprehensively consider the text change index, the content change index, and the image change index to determine the description effectiveness index of the input text during the latter image generation in the two adjacent image generations.
[0070] Based on the foregoing analysis, it can be seen that when the image style changes significantly, and the text change index and the content change index are smaller, it indicates that the style change of the text information in the input text does not damage the accuracy of the science popularization and can directly act on the style dimension, so the usefulness is higher.
[0071] The larger the text change index, the greater the difference between the input texts in two adjacent times. The larger the content change index, the greater the style change of the picture, and the lower the accuracy of the science popularization. Therefore, during two adjacent image generations, the value obtained by performing negative correlation mapping and normalization on the product of the text change index and the content change index is used as the effectiveness factor. The larger the effectiveness factor, the smaller the degree of damage to the science popularization accuracy caused by the style change of the text information in the input text. The negative correlation mapping and normalization operation here can be performed using the formula , where represents the exponential function with the natural constant e as the base, and x represents the independent variable.
[0072] Finally, the product of the image change index and the effective factor is normalized to obtain the value as the description effective index of the input text in the latter image generation of two adjacent image generation. The larger the description effective index, the greater the change in the input text can lead to a greater change in the image content without destroying the accuracy of popular science. Therefore, the higher the effectiveness, the more useful the change in the input text. Normalization is a technical means well known to those skilled in the art, and the normalization function can be linear normalization or standard normalization, etc. The specific normalization method is not limited here.
[0073] Step S3: Obtain the anthropomorphic vector each time an image is generated and compare it with the anthropomorphic vector of the comparison image, so as to quantify the anthropomorphic feature value vector of the input text each time an image is generated; based on the description effectiveness index of the input text each time an image is generated and the anthropomorphic feature value vector, determine the interesting feature vector of the input text.
[0074] Interest is very important in interesting medical popular science pictures, which can enhance the audience's memory and promote information sharing through humorous elements or story-telling narration; the relationship between the interest and authenticity of the image mainly lies in the trade-off between retaining the morphological characteristics and the degree of anthropomorphism, that is, the more obvious the medical knowledge is reflected in the output image, the more the image is inclined to the real form (such as the morphological distribution of the cell wall and flagella of Vibrio cholerae), and the anthropomorphic eyes, mouth, arms and other elements on the output image are more inclined to the result of interest, but the image that becomes a human being will obviously lack medical popular science information. Therefore, in the embodiment of the present invention, it is necessary to further analyze the interesting expression method, by comparing the difference between the anthropomorphic vector at each image generation and the anthropomorphic vector of the comparison image, evaluate whether the style adjustment leads to content distortion, obtain the anthropomorphic feature value vector (reflecting the style strength), and combine the description of the effective index (reflecting the accuracy of the content) to generate the interesting feature vector, so as to provide data support for the subsequent adjustment process.
[0075] First, the anthropomorphic vector at each image generation and the anthropomorphic vector of the comparison image can be obtained and compared, thereby quantifying the anthropomorphic feature value vector of the input text at each image generation.
[0076] Preferably, in one embodiment of the present invention, the method for obtaining the anthropomorphic feature value vector includes:
[0077] When judging the anthropomorphic features, the anthropomorphic vector can be first used to quantify the degree of anthropomorphism of the image in multiple dimensions each time the image is generated. The dimensions of the anthropomorphic vector include appearance features (such as facial features, body structure and clothing), behavioral features (such as movement flexibility and emotional expression), and language ability (such as conversation ability and language style). By scoring these dimensions, an anthropomorphic vector can be preliminarily constructed. For example, a neural network can be trained, and the input keyword vector corresponding to the input text of the image is used as the input of the neural network. Scoring can be performed from the aforementioned multiple dimensions to obtain the scoring values on each dimension to form an anthropomorphic vector. In addition, since a dimension in the anthropomorphic vector may correspond to multiple keywords, such as appearance features, including facial features, body structure and clothing, when the anthropomorphic vector is obtained, the corresponding relationship between each dimension in the anthropomorphic vector and the element value in the input keyword vector can be determined.
[0078] Each time an image is generated, the input keyword vector can be used as the input of the pre-trained neural network to obtain the anthropomorphic vector corresponding to the generated image and the corresponding relationship between the element values in the anthropomorphic vector and the element values in the output keyword vector.
[0079] Similarly, an anthropomorphic vector of a comparison image is obtained and used as a comparison vector, wherein the comparison image includes an original electron microscope image and a standard portrait. The original electron microscope image is used as a comparison image with an anthropomorphic degree of 0, and can be obtained through an image acquisition device, such as using a microscope and an imaging device to obtain an image of a real Vibrio cholerae; the standard portrait is used as a comparison image with an anthropomorphic degree of 10.
[0080] Then, in the anthropomorphic vector and each comparison vector at each time of image generation, the cosine similarity between the element values at the same position is used as the characteristic factor. The characteristic factor is used to characterize the similarity of the degree of anthropomorphism between the output image and the comparison image at each time of image generation. For example, if the characteristic factor between the output image and the standard portrait is larger, the degree of anthropomorphism of the output image is higher, and the interest is higher; conversely, if the characteristic factor between the output image and the original electron microscope image is larger, the authenticity of the output image is higher, and there may be a risk of lower interest.
[0081] Finally, perform a negative correlation mapping process on the characteristic factors between the element values at each position in the anthropomorphic vector of the generated image and the original electron microscope image, so as to obtain an interestingness factor, which is used to reflect the interestingness of the element values at each position in the anthropomorphic vector of the generated image. The larger this value is, the lower the similarity with the original electron microscope image, and the higher the interestingness. Then multiply the interestingness factor by the characteristic factors between the element values at each position in the anthropomorphic vector of the generated image and the standard portrait, and use the normalized value of the obtained product as the element value at the same position in the anthropomorphic eigenvalue vector, so as to obtain the anthropomorphic eigenvalue vector corresponding to the input text for each image generation. The negative correlation mapping here can adopt the formula , where represents the exponential function with the natural constant e as the base, and x represents the independent variable.
[0082] Thus, the anthropomorphic eigenvalue vector corresponding to the input text for each image generation can be obtained.
[0083] The anthropomorphic eigenvalue vector describes the interestingness of the image in terms of the anthropomorphic style. However, in the actual comparison process, the manifestation of interestingness is related to the description effectiveness of the keywords in the input text, and each element value in the anthropomorphic eigenvalue vector corresponds to one or more keywords in the input text. Therefore, in this embodiment of the present invention, the description effectiveness index of the input text for each image generation and the anthropomorphic eigenvalue vector are combined to determine the interestingness feature vector of the input text.
[0084] Preferably, in an embodiment of the present invention, the method for obtaining the interestingness feature vector includes:
[0085] In this embodiment of the present invention, a linkage mechanism combining feature emphasis and content effectiveness is established. The purpose is to improve the anthropomorphic features with higher effectiveness. Therefore, for each image generation, the number of element values in the input keyword vector corresponding to each element value in the anthropomorphic eigenvalue vector is used as the quantity factor of each element value in the anthropomorphic eigenvalue vector.
[0086] Then, for each image generation, the product of the quantity factor of each element value in the anthropomorphic eigenvalue vector and the description effectiveness index of the input text is used as the adjustment parameter corresponding to each element value in the anthropomorphic eigenvalue vector. The larger the adjustment parameter is, the more keywords the element value corresponds to and the higher the effectiveness. Then the anthropomorphic feature of this element value should be amplified.
[0087] Therefore, finally, the value obtained by normalizing the product of each element value in the anthropomorphic feature value vector and the corresponding adjustment parameter is used as the element value at the corresponding position in the interestingness feature vector, so as to obtain the interestingness feature vector of the input text for each image generation. The normalization is a technical means well-known to those skilled in the art, and the choice of the normalization function can be linear normalization or standard normalization, etc. The specific normalization method is not limited herein.
[0088] Step S4: Adjust the input keyword vector of the input text for the next image generation based on the interestingness feature vector of the input text for the current image generation, so as to obtain an input keyword adjustment vector, which is used in the production process of the interesting medical science popularization map to obtain the next output image.
[0089] Based on the foregoing steps, the interestingness feature vector of the input text for each image generation can be determined. Then, the keyword vector of the input text for the next image generation can be adjusted based on the interestingness feature vector of the current image generation, so that the interestingness feature is more prominent, and an input keyword adjustment vector is obtained. Finally, it is used in the production process of the interesting science popularization map to obtain the next output image.
[0090] Preferably, in an embodiment of the present invention, the method for obtaining the input keyword adjustment vector includes:
[0091] In the interestingness feature vector of the input text for the current image generation and the input keyword vector of the input text for the next image generation, zero-padding is performed on the low-dimensional vector to obtain a vector group. At this time, in the vector group, the dimensions of the original interestingness feature vector and the input keyword vector can be kept consistent. In this vector group, the vector corresponding to the input keyword vector is used as the first vector, and at the same time, the vector corresponding to the interestingness feature vector is used as the second vector.
[0092] Then, a balance coefficient is introduced to fuse the element values in the first vector and the second vector: in the first vector and the second vector, for the element values at the same position, the product of the first balance coefficient and the element value in the first vector is used as the first adjustment factor, the product of the second balance coefficient and the element value in the second vector is used as the second adjustment factor, and the sum value of the first adjustment factor and the second adjustment factor is used as the element value at the same position in the input keyword adjustment vector, so as to obtain the input keyword adjustment vector of the input text for the next image generation.
[0093] It should be noted that since the purpose of the interestingness feature vector is to adjust the interestingness, when obtaining the input keyword adjustment vector, the proportion of the element values in the input keyword vector should be larger. Therefore, the value range of the first balance coefficient is (0.5, 1), and the value range of the second balance coefficient is (0, 0.5), and the sum of the two should be 1. In this embodiment of the present invention, the first balance coefficient can be set to 0.7, and then the second balance coefficient is 0.3.
[0094] Thus, the input keyword adjustment vector of the input text for the next image generation can be obtained, and then this vector is applied to the DALL-E2 model to obtain the next output image. Moreover, the next output image takes into account the influence of the interestingness feature, so it will better meet the requirements of the interesting popular science picture. Therefore, the number of adjustments by the user can be reduced as much as possible, enabling the model to obtain the interesting popular science image required by the user faster and more accurately.
[0095] In summary, the generation process of the interesting medical popular science picture is a process that requires continuous adjustment. The model automatically generates an image based on the input text, and then the user can continue to input text to adjust the image in order to obtain an interesting medical popular science picture that meets the requirements and can effectively attract the target audience. In the production process of the interesting medical popular science picture, in this embodiment of the present invention, for each image generation, the input keyword vector, output keyword vector, and image vector corresponding to the image generation are obtained, which can be used to analyze the image generation situation in the subsequent process. Each time text is input, it is actually another processing of the image. To quantify the effectiveness of the input text for image processing, the input keyword vectors, output keywords, and image vectors during two adjacent image generations can be compared to obtain the description effectiveness index of the input text for each image generation. Further, due to the influence of the interestingness and authenticity of the image, the main issue lies in the trade-off between retaining the original morphological features and the degree of anthropomorphism. To avoid the generated image being too realistic and lacking interestingness, it is necessary to further analyze the expression method of interestingness. The anthropomorphic vector of each image generation and the anthropomorphic vector of the comparison image can be obtained and compared to obtain the anthropomorphic eigenvalue vector of the input text for each image generation, which is used to reflect the anthropomorphic characteristics of the output image. Furthermore, by combining the description effectiveness index of the input text for each image generation and the anthropomorphic eigenvalue vector, the interestingness feature vector of the input text can be determined, which can characterize the interestingness feature. Finally, based on the interestingness feature vector, the model is improved, that is, the input keyword vector of the input text for the next image generation is adjusted to obtain the input keyword adjustment vector, which balances the authenticity and interestingness, thereby improving the accuracy of the generation result, effectively reducing the number of generations, and generating more interesting medical popular science pictures that meet the user's needs and expectations.
[0096] An embodiment of the present invention also provides an AI-based interesting medical science popularization production system, including a processor and a memory. The memory stores at least one instruction, at least one program, a code set or an instruction set. When the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor, the steps of an AI-based interesting medical science popularization production method are implemented.
[0097] Please refer to Figure 3 , which shows a system block diagram of an AI-based interesting medical science popularization production system, including a data acquisition module 301 for implementing step S1 in the above method embodiment; a description validity analysis module 302 for implementing step S2 in the above method embodiment; an interestingness analysis module 303 for implementing step S3 in the above method embodiment; and an output image adjustment module 304 for implementing step S4 in the above method embodiment.
[0098] It should be noted that for the system provided in the above embodiment, only the division of the above function modules is used for illustration. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the computer device is divided into different function modules to complete all or part of the functions described above. In addition, an AI-based interesting medical science popularization production system and an embodiment of an AI-based interesting medical science popularization production method provided in the above embodiment belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.
[0099] Please refer to Figure 4 , which shows a schematic system structure diagram of an AI-based interesting medical science popularization production system provided by an embodiment of the present invention, including a processor 400, a memory 401, a bus 402 and a communication interface 403. The processor 400, the communication interface 403 and the memory 401 are connected through the bus 402. Among them, the memory 401 may include high-speed random access memory. The bus 402 may be an ISA bus, a PCI bus or an EISA bus, etc. The processor 400 may be an integrated circuit chip with signal processing capabilities. The memory 401 stores at least one instruction, at least one program, a code set or an instruction set. When the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor, the steps in an AI-based interesting medical science popularization production method are implemented.
[0100] It should be noted that the above sequence of embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0101] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized.
Claims
1. A method for producing interesting medical popular science based on artificial intelligence, characterized in that The method includes: In the production process of the interesting medical science popularization map, obtaining the input keyword vector corresponding to the input text, the output keyword vector corresponding to the output image, and the image vector each time an image is generated; Comparing the differences between the input keyword vectors, the differences between the output keyword vectors, and the differences between the image vectors during two adjacent image generations, and determining the description effectiveness index of the input text each time an image is generated; Obtaining the anthropomorphic vector of each image generation and comparing it with the anthropomorphic vector of the comparison image, so as to quantify the anthropomorphic eigenvalue vector of the input text each time an image is generated; Based on the description effectiveness index and the anthropomorphic eigenvalue vector of the input text each time an image is generated, determining the interestingness feature vector of the input text; Adjusting the input keyword vector of the input text for the next image generation based on the interestingness feature vector of the input text for the current image generation to obtain an input keyword adjustment vector, which is used in the production process of the interesting medical science popularization map to obtain the next output image.
2. The method for producing interesting medical popular science based on artificial intelligence according to claim 1, wherein The method for obtaining the description effectiveness index includes: Comparing the differences between the input keyword vectors during two adjacent image generations to determine the text change index; Comparing the differences between the output keyword vectors during two adjacent image generations to determine the content change index; Comparing the differences between the image vectors during two adjacent image generations to determine the image change index; During two adjacent image generations, taking the value obtained by performing negative correlation mapping and normalization on the product of the text change index and the content change index as the effective factor; Taking the value obtained by normalizing the product of the image change index and the effective factor as the description effectiveness index of the input text for the latter image generation during two adjacent image generations.
3. A method for producing interesting medical popular science based on artificial intelligence according to claim 2, characterized in that, The method for obtaining the text change index includes: Comparing two input keyword vectors to determine the low-dimensional vector and the high-dimensional vector, and performing zero-padding on the low-dimensional vector to obtain a padded vector; During two input keyword vectors, taking the value obtained by performing negative correlation mapping and normalization on the cosine similarity between the high-dimensional vector and the padded vector as the text change index.
4. A method for producing interesting medical popular science based on artificial intelligence according to claim 2, characterized in that, The method for obtaining the content change index includes: Comparing two output keyword vectors to determine the low-dimensional vector and the high-dimensional vector, and performing zero-padding on the low-dimensional vector to obtain a feature vector; During two output keyword vectors, taking the value obtained by performing negative correlation mapping and normalization on the cosine similarity between the high-dimensional vector and the feature vector as the content change index.
5. A method for producing interesting medical popular science based on artificial intelligence according to claim 2, characterized in that, The method for obtaining the image change index includes: Calculating the Euclidean distance between two image vectors as the image change index.
6. The method for producing interesting medical popular science based on artificial intelligence according to claim 1, characterized in that The method for obtaining the anthropomorphic eigenvalue vector includes: During each image generation, taking the input keyword vector as the input of a pre-trained neural network to obtain the anthropomorphic vector corresponding to the generated image and the corresponding relationship between the element values in the anthropomorphic vector and the element values in the output keyword vector; Obtaining the anthropomorphic vector of the comparison image and using it as the comparison vector, where the comparison image includes the original electron microscope image and the standard portrait; For the anthropomorphic vector and each comparison vector during each image generation, the cosine similarity between the element values at the same position is used as the feature factor; The value obtained by performing negative correlation mapping on the feature factor between the element value at each position in the anthropomorphic vector of the generated image and the original electron microscope image is multiplied by the feature factor between the element value at each position in the anthropomorphic vector of the generated image and the standard portrait. The value obtained after normalizing the resulting product is used as the element value at the same position in the anthropomorphic feature value vector, thereby obtaining the anthropomorphic feature value vector corresponding to the input text during each image generation.
7. A method for producing interesting medical popular science based on artificial intelligence according to claim 6, characterized in that, The method for obtaining the interestingness feature vector includes: During each image generation, the number of element values corresponding to each element value in the anthropomorphic feature value vector in the input keyword vector is used as the quantity factor of each element value in the anthropomorphic feature value vector; During each image generation, the product of the quantity factor of each element value in the anthropomorphic feature value vector and the description effectiveness index of the input text is used as the adjustment parameter corresponding to each element value in the anthropomorphic feature value vector; The value obtained by normalizing the product of each element value in the anthropomorphic feature value vector and the corresponding adjustment parameter is used as the element value at the corresponding position in the interestingness feature vector, thereby obtaining the interestingness feature vector of the input text during each image generation.
8. A method for producing interesting medical popular science based on artificial intelligence according to claim 1, characterized in that The obtaining of the input keyword adjustment vector includes: After aligning the interestingness feature vector of the input text during the current image generation and the input keyword vector of the input text during the next image generation based on the zero-padding method, a vector group is obtained; In the vector group, the vector corresponding to the input keyword vector is used as the first vector, and the vector corresponding to the interestingness feature vector is used as the second vector; In the first vector and the second vector, for the element values at the same position, the product of the first balance coefficient and the element value in the first vector is used as the first adjustment factor, and the product of the second balance coefficient and the element value in the second vector is used as the second adjustment factor. The sum value of the first adjustment factor and the second adjustment factor is used as the element value at the same position in the input keyword adjustment vector, thereby obtaining the input keyword adjustment vector of the input text during the next image generation.
9. A method for producing interesting medical popular science based on artificial intelligence according to claim 1, characterized in that During the production process of the interesting medical science popularization map, obtaining the input keyword vector, output keyword vector, and image vector corresponding to the input text during each image generation includes: During the generation process of the interesting medical science popularization map, the image vector of the output image is extracted through a pre-trained CLIP model, the input keyword vector is obtained by processing the input text based on the BioBERT model, and the output keyword vector corresponding to the output image is generated through an image description model.
10. An interesting medical science popularization production system based on artificial intelligence, characterized in that, It includes a processor and a memory. At least one instruction, at least one program, code set, or instruction set is stored in the memory. When at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor, the steps of an interesting medical science popularization production method based on artificial intelligence as described in any one of claims 1-9 are implemented.
Citation Information
Patent Citations
Retrieval recommendation method and device, computer readable storage medium and electronic equipment
CN109783727A
Focus image processing method and related device
CN113724188A
Medical science popularization video production method and system based on UGC mode
CN114554246A
Instruction data safety early warning method based on intelligent robot automatic control system
CN115958609A
Building structure column image thumbnail intelligent generation method for intelligent construction
CN116542859A
Cited By
Large model-based medical science popularization image-text generation method and system
CN121564138A