An interesting medical science popularization production system and method based on artificial intelligence

By analyzing the differences in input keywords and image vectors in the process of popular science images in the interesting medicine, quantifying the description effectiveness and anthropomorphic features, and adjusting the input keyword vectors, the problem of insufficient fun in the existing technology is solved, and more efficient popular science image generation in the interesting medicine is achieved.

CN120298533BActive Publication Date: 2025-08-29BEIJING ZONGHENG WUSHUANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510748046.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-29
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing popular science images for interesting medicine rely on popular science text and description text input by users. There is a lack of targeted adjustment during the generation process, resulting in low interest in the image and difficulty in effectively attracting the target audience.

Method used

By obtaining the input keyword vector, output keyword vector and image vector during each image generation, comparing the differences between adjacent generations, quantifying the effective indicators of the description of the input text and the personified eigenvalues, adjusting the input keyword vectors to balance the fun and authenticity, using the pre-trained CLIP and BioBERT models to extract the vectors, and combining the DALL-E2 model to generate images.

Benefits of technology

It improves the accuracy and efficiency of popularization of interesting medical science images, reduces the number of user adjustments, and the generated images are more in line with user needs and expectations, and can effectively attract the target audience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298533B_ABST
    Figure CN120298533B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of popular science image generation, and specifically to an interesting medical popular science production system and method based on artificial intelligence. Each time an image is generated, the input keyword vector, output keyword vector and image vector are synchronously collected, and the description effectiveness index of the text adjustment is quantified by comparing the vector difference between two adjacent generated images; in response to the need for optimizing image fun, anthropomorphic feature analysis is introduced, and the difference in the anthropomorphic vector between the current generated image and the comparison image is calculated to obtain the characteristic value, and a multi-dimensional fun feature vector is formed in combination with the description effectiveness index; a dynamic adjustment mechanism for model parameters is established based on the characteristic vector, and the medical authenticity and the degree of artistic anthropomorphism are balanced by adaptively correcting the input keyword vector, so as to achieve precise control of the generation effect. In summary, the present invention effectively solves the problem that traditional generation models are difficult to strike a balance between fun and accuracy in the field of medical popular science, and can significantly reduce the number of generation times and improve user satisfaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of popular science image generation, and in particular to an interesting medical popular science production system and method based on artificial intelligence. Background Art

[0002] Interesting popular science medical knowledge posters and cartoons are an effective way to spread medical and health knowledge through vivid and interesting visual expressions. Their content usually uses bright colors, anthropomorphic illustrations and humorous character designs. They aim to convey health information to the public in a concise and clear manner, covering topics such as nutrition, exercise, mental health, disease prevention, etc.

[0003] Existing methods for generating interesting popular science medical knowledge posters mainly rely on artificial intelligence models that can process popular science texts. The generation process relies on popular science texts and descriptive texts input by users, and meets user needs through multiple user inputs and adjustments. However, there is often a lack of targeted adjustments in terms of fun, which easily leads to the generated images being less interesting and unable to effectively attract the target audience, thus requiring multiple adjustments. Summary of the Invention

[0004] In order to solve the technical problem that the existing generation of interesting popular science medical knowledge posters mainly relies on an artificial intelligence model that can process popular science texts, and its generation process relies on popular science texts and description texts input by users, and meets user needs through multiple user inputs and adjustments, but often lacks targeted adjustments in terms of fun, which easily leads to low fun of generated images, the purpose of the present invention is to provide an interesting medical popular science production system and method based on artificial intelligence. The technical solutions adopted are as follows:

[0005] In the production process of interesting medical popular science pictures, the input keyword vector corresponding to the input text, the output keyword vector corresponding to the output image, and the image vector are obtained each time an image is generated;

[0006] Compare the differences between the input keyword vectors, the output keyword vectors, and the image vectors during two consecutive image generation times to determine the description effectiveness index of the input text during each image generation;

[0007] Obtaining and comparing the anthropomorphic vector of each image generated with the anthropomorphic vector of the comparison image, thereby quantifying the anthropomorphic feature value vector of the input text at each image generation; determining the interestingness feature vector of the input text based on the description effectiveness index of the input text at each image generation and the anthropomorphic feature value vector;

[0008] Based on the interesting feature vector of the input text when the current image is generated, the input keyword vector of the input text when the next image is generated is adjusted to obtain an input keyword adjustment vector, which is used in the production process of interesting medical popular science pictures to obtain the next output image.

[0009] Furthermore, the method for obtaining the description effective indicator includes:

[0010] Compare the differences between the input keyword vectors when two adjacent images are generated to determine the text change index;

[0011] Compare the differences between the output keyword vectors of two adjacent image generation times to determine the content change index;

[0012] Compare the difference between the image vectors when two adjacent images are generated to determine the image change index;

[0013] When two adjacent images are generated, the product of the text change index and the content change index is negatively correlated and normalized to obtain the value as the effective factor;

[0014] The value obtained by normalizing the product of the image change index and the effectiveness factor is used as the description effectiveness index of the input text in the latter image generation between two adjacent image generation.

[0015] Furthermore, the method for obtaining the text change index includes:

[0016] Compare the two input keyword vectors, determine the low-dimensional vector and the high-dimensional vector, and pad the low-dimensional vector with zeros to obtain the padding vector;

[0017] In the two input keyword vectors, the cosine similarity between the high-dimensional vector and the filling vector is negatively correlated and normalized to obtain a value which is used as the text change indicator.

[0018] Furthermore, the method for obtaining the content change indicator includes:

[0019] Compare the two output keyword vectors, determine the low-dimensional vector and the high-dimensional vector, and fill the low-dimensional vector with zeros to obtain the feature vector;

[0020] In the two output keyword vectors, the cosine similarity between the high-dimensional vector and the feature vector is negatively correlated and normalized to a value that serves as the content change indicator.

[0021] Furthermore, the method for obtaining the image change index includes:

[0022] Calculate the Euclidean distance between two image vectors as an image change indicator.

[0023] Furthermore, the method for obtaining the anthropomorphic eigenvalue vector includes:

[0024] Each time an image is generated, the input keyword vector is used as the input of the pre-trained neural network to obtain the anthropomorphic vector corresponding to the generated image and the correspondence between the element values ​​in the anthropomorphic vector and the element values ​​in the output keyword vector;

[0025] Obtaining an anthropomorphic vector of a comparison image and using it as a comparison vector, wherein the comparison image includes an original electron microscope image and a standard human portrait;

[0026] In the anthropomorphic vector and each comparison vector at each image generation, the cosine similarity between the element values ​​at the same position is used as the feature factor;

[0027] The value obtained by negatively correlating the characteristic factor between the element value at each position in the anthropomorphic vector of the generated image and the original electron microscope image is multiplied by the characteristic factor between the element value at each position in the anthropomorphic vector of the generated image and the standard portrait. The obtained product is normalized and used as the element value at the same position in the anthropomorphic characteristic value vector, thereby obtaining the anthropomorphic characteristic value vector corresponding to the input text each time the image is generated.

[0028] Furthermore, the method for obtaining the interesting feature vector includes:

[0029] Each time an image is generated, the number of element values ​​corresponding to each element value in the input keyword vector in the anthropomorphic feature value vector is used as the quantity factor of each element value in the anthropomorphic feature value vector;

[0030] When each image is generated, the product of the quantity factor of each element value in the anthropomorphic feature value vector and the description effectiveness index of the input text is used as the adjustment parameter corresponding to each element value in the anthropomorphic feature value vector;

[0031] The product of each element value in the anthropomorphic feature value vector and the corresponding adjustment parameter is normalized and used as the element value at the corresponding position in the interesting feature vector, thereby obtaining the interesting feature vector of the input text each time an image is generated.

[0032] Furthermore, the acquisition of the input keyword adjustment vector includes:

[0033] After aligning the interesting feature vector of the input text during the current image generation with the input keyword vector of the input text during the next image generation using the zero-padding method, a vector group is obtained;

[0034] In the vector group, the vector corresponding to the input keyword vector is used as the first vector, and the vector corresponding to the interesting feature vector is used as the second vector;

[0035] In the first vector and the second vector, for the element value at the same position, the product of the first balance coefficient and the element value in the first vector is used as the first adjustment factor, the product of the second balance coefficient and the element value in the second vector is used as the second adjustment factor, and the sum of the first adjustment factor and the second adjustment factor is used as the element value at the same position in the input keyword adjustment vector, so as to obtain the input keyword adjustment vector of the input text when the next image is generated.

[0036] Furthermore, in the process of producing the interesting medical popular science illustration, obtaining the input keyword vector corresponding to the input text, the output keyword vector corresponding to the output image, and the image vector each time an image is generated includes:

[0037] In the process of generating interesting medical popular science pictures, the image vector of the output image is extracted through the pre-trained CLIP model, the input text is processed based on the BioBERT model to obtain the input keyword vector, and the output keyword vector corresponding to the output image is generated through the image description model.

[0038] A system for producing interesting medical popular science based on artificial intelligence includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set. When the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor, the steps of the method for producing interesting medical popular science based on artificial intelligence are implemented.

[0039] The present invention has the following beneficial effects:

[0040] The generation process of interesting medical popular science pictures is a process that requires constant adjustment. The model will automatically generate images based on the input text, and then the user can continue to input text to adjust the image in order to obtain interesting medical popular science images that meet the requirements and can effectively attract the target audience. In the production process of interesting medical popular science pictures, the present invention obtains the input keyword vector, output keyword vector and image vector corresponding to each image generation, which can be used to analyze the image generation situation in the subsequent process. Each time text is input, it is actually another processing of the image. In order to quantify the effectiveness of the input text for image processing, the input keyword vector, output keyword and image vector of two adjacent image generations can be compared to obtain the description effectiveness index of the input text each time the image is generated. Furthermore, since the interest and authenticity of the image are mainly affected by the trade-off between retaining its own morphological characteristics and the degree of anthropomorphism, in order to avoid generating an image with too high authenticity but lack of interest, it is necessary to further analyze the interesting expression method. The anthropomorphic vector of each image generated can be obtained and compared with the anthropomorphic vector of the comparison image to obtain the anthropomorphic feature value vector of the input text at each image generation, which is used to reflect the anthropomorphic characteristics of the output image. Furthermore, the description effectiveness index of the input text at each image generation and the anthropomorphic feature value vector are combined to determine the interestingness feature vector of the input text, which can characterize the interestingness. Finally, based on the interestingness feature vector, the model is improved. Specifically, the input keyword vector of the input text for the next image generation is adjusted to obtain the input keyword adjustment vector. This balances authenticity and interestingness, thereby improving the accuracy of the generated results, effectively reducing the number of generation times, and producing interesting medical popular science images that better meet user needs and expectations. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0042] Figure 1 A flowchart of a method for producing interesting medical popular science content based on artificial intelligence provided by one embodiment of the present invention;

[0043] Figure 2 A flowchart of a method for obtaining effective indicators provided by one embodiment of the present invention;

[0044] Figure 3 This is a system block diagram of an interesting medical science popularization production system based on artificial intelligence provided by one embodiment of the present invention;

[0045] Figure 4 A schematic diagram of the system structure of an interesting medical popular science production system based on artificial intelligence provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0046] To further illustrate the technical means and effectiveness of the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effectiveness of an AI-based fun medical science production system and method proposed by the present invention. In the following description, different references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0047] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0048] The following describes in detail a specific scheme of an interesting medical science popularization production system and method based on artificial intelligence provided by the present invention in conjunction with the accompanying drawings.

[0049] See also Figure 1 , which shows a method flow chart of a method for producing interesting medical popular science based on artificial intelligence provided by one embodiment of the present invention, the method comprising the following steps:

[0050] Step S1: During the production process of the interesting medical popular science illustrations, the input keyword vector corresponding to the input text, the output keyword vector corresponding to the output image, and the image vector are obtained each time an image is generated.

[0051] Interesting popular science medical knowledge posters and comics are an effective way to spread medical and health knowledge through vivid and interesting visual expressions. Through storytelling and humorous elements, these forms not only enhance the audience's memory, but also promote information sharing and reduce learning barriers. They are suitable for school education, public health promotion and social media marketing.

[0052] The actual generation process of interesting science popularization pictures is realized by inputting text and anthropomorphic feature description. The input text is used to provide actual popular science information keywords to confirm the element type information of the final interesting science popularization picture, while the interesting attributes depend more on the anthropomorphic features of different element types, thereby realizing the interestingness judgment of the generated picture.

[0053] DALL-E2 is an advanced image generation model developed by OpenAI that generates high-quality images based on user-provided text descriptions. The model can understand complex text prompts and generate vivid images to match them. Furthermore, it supports image editing, allowing users to make subtle adjustments to generated images or generate different variations of a given image, enabling more flexible creative work. Therefore, in this embodiment of the present invention, this model is primarily used to generate interesting medical science illustrations.

[0054] First, the production process of fun medical popular science illustrations requires a process of input text, output image, input text (for adjustment purposes), output image, and so on, until a fun medical popular science illustration that meets user needs is obtained. During this process, it is necessary to obtain the input keyword vector corresponding to the input text, the output keyword vector corresponding to the output image, and the image vector of the output image for each image generation. Preferably, in one embodiment of the present invention, the image vector of the output image can be extracted using a pre-trained CLIP model, the input text is processed based on the BioBERT model to obtain the input keyword vector, and the output keyword vector corresponding to the output image is generated using an image description model (such as ViT-GPT2).

[0055] It should be noted that the DALL-E2 model, CLIP model, BioBERT model, and ViT-GPT2 model are all well-known technologies, and the specific processes will not be described here.

[0056] Step S2: Compare the differences between the input keyword vectors, the output keyword vectors, and the image vectors during two adjacent image generation processes to determine the description validity index of the input text during each image generation process.

[0057] The main factor affecting the accuracy of the DALL-E2 model's analysis of the relationship between text and image is the accuracy of the descriptive text in describing the morphological features of the required popular science object. Keywords in the user's input text act as text anchors. That is, when the keyword exists in the input text, the corresponding element will be depicted in the image generated by the DALL-E2 model. When the keyword does not exist, a fuzzy matching description will be implemented, and its anchoring effect is not obvious. For the user, the changes made to the input text on the image obtained by the previous input text are actually reprocessing of the generated image. By comparing the differences between the input keyword vectors, the differences between the output keyword vectors, and the differences between the image vectors when two adjacent images are generated, the description effectiveness index of the input text can be determined for each image generation. This index is used to reflect the correspondence between the input text and the output image, and to characterize the effectiveness of the change in the style information of the input text.

[0058] Preferably, in one embodiment of the present invention, the method for obtaining the description effective indicator includes:

[0059] See also Figure 2 , which shows a flow chart of a method for obtaining effective indicators in one embodiment of the present invention, the method comprising the following steps:

[0060] For the input keyword vector corresponding to the input text during a certain image generation, the smaller the change relative to the previous one, and the smaller the change in the output keyword vector of the output image, while the corresponding output image content changes more, it means that the style change of the text information of the input text does not undermine the accuracy of popular science, and can directly affect the style dimension, and therefore the more useful it is.

[0061] Step S201: comparing the difference between the input keyword vectors generated during two adjacent image generation operations to determine a text change index.

[0062] Since the dimensions of the input keyword vectors when two adjacent images are generated may be inconsistent, we first compare the two input keyword vectors to determine the low-dimensional vector and the high-dimensional vector, and then pad the low-dimensional vector with zeros to obtain the padding vector so that the dimensions of the two vectors remain consistent.

[0063] Then, in the two input keyword vectors, the cosine similarity between the high-dimensional vector and the filling vector is calculated. The larger the cosine similarity, the more similar the input texts are, and the smaller the adjustment degree is. Therefore, the cosine similarity is negatively correlated and normalized, and the value after the text change index is used as the text change index. The larger the text change index, the greater the difference between the two adjacent input texts, and the more obvious the anchor point role of the input keyword vector will be. The negative correlation mapping and normalization operation here can be used as the formula ,in, To express the exponential function with the natural constant e as the base, x represents the independent variable.

[0064] Step S202: comparing the differences between the output keyword vectors of two adjacent image generation times to determine a content change index.

[0065] Similarly, since the dimensions of the output keyword vectors corresponding to the output images when two adjacent images are generated may be inconsistent, the two output keyword vectors are compared to determine the low-dimensional vector and the high-dimensional vector, and the low-dimensional vector is zero-filled to obtain the feature vector.

[0066] Then, in the two output keyword vectors, the cosine similarity between the high-dimensional vector and the feature vector is negatively correlated and normalized to obtain the value as the content change indicator. The larger the cosine similarity here, the less information is lost, and the style of the input keyword vector will be closer. Therefore, after using negative correlation mapping and normalization to correct the logical relationship, the larger the content change index obtained, the greater the style change of the image, and the accuracy of the popular science will be reduced. The negative correlation mapping and normalization operation here can be performed using the formula ,in, To express the exponential function with the natural constant e as the base, x represents the independent variable.

[0067] Step S203: comparing the difference between the image vectors generated during two adjacent image generations to determine an image change index.

[0068] Since the dimensions of image vectors are consistent, when measuring the differences between image vectors, the Euclidean distance between the two image vectors can be directly calculated as the image change index. At this time, the larger the image change index, the more obvious the change in image style.

[0069] Step S204: When two adjacent images are generated, the text change index, the content change index, and the image change index are comprehensively considered to determine the description validity index of the input text when the latter of the two adjacent images is generated.

[0070] Based on the above analysis, when the image style changes significantly, and the text change index and content change index are smaller, it means that the style change of the text information of the input text does not destroy the accuracy of popular science, and can directly affect the style dimension, so the degree of usefulness is higher.

[0071] The larger the text change index, the greater the difference between the two adjacent input texts. The larger the content change index, the greater the change in the style of the image, and the lower the accuracy of the popular science. Therefore, when two adjacent images are generated, the product of the text change index and the content change index is negatively correlated and normalized to the value obtained as the effective factor. The larger the effective factor, the less damage the change in the style of the text information of the input text has on the accuracy of the popular science. The negative correlation mapping and normalization operation here can be performed using the formula ,in, To express the exponential function with the natural constant e as the base, x represents the independent variable.

[0072] Finally, the product of the image change index and the effectiveness factor is normalized to obtain the value of the descriptive effectiveness index for the input text in the second image generation between two consecutive image generation attempts. A larger descriptive effectiveness index indicates that a small change in the input text can lead to a significant change in the image content without compromising the accuracy of the popular science. Therefore, a higher effectiveness index indicates that the change in the input text is more useful. Normalization is a well-known technique for those skilled in the art, and the normalization function can be linear normalization or standard normalization, among others. The specific normalization method is not limited here.

[0073] Step S3: Obtain the anthropomorphic vector at each image generation and the anthropomorphic vector of the comparison image and compare them, so as to quantify the anthropomorphic feature value vector of the input text at each image generation; determine the interesting feature vector of the input text based on the description effectiveness index and the anthropomorphic feature value vector of the input text at each image generation.

[0074] Interest is crucial in entertaining medical popular science images. Through humorous elements or storytelling, they can enhance audience memory and promote information sharing. The relationship between an image's interest and authenticity lies primarily in the trade-off between preserving the image's morphological characteristics and the degree of anthropomorphism. Specifically, the more pronounced the medical knowledge is in the output image, the more realistic the image is (e.g., the morphological distribution of the cell wall and flagella of Vibrio cholerae). Anthropomorphic eyes, mouths, arms, and other elements in the output image contribute to a more entertaining outcome. However, a completely human-like image would clearly lack medical popular science information. Therefore, in this embodiment of the present invention, further analysis of the entertaining presentation is required. By comparing the anthropomorphic vectors of each generated image with those of the comparison image, an assessment is made as to whether style adjustments have resulted in content distortion. This results in an anthropomorphic feature value vector (reflecting style strength) and, combined with a descriptive effectiveness index (reflecting content accuracy), an interesting feature vector is generated to provide data support for subsequent adjustments.

[0075] First, the anthropomorphic vector at each image generation and the anthropomorphic vector of the comparison image can be obtained and compared, thereby quantifying the anthropomorphic feature value vector of the input text at each image generation.

[0076] Preferably, in one embodiment of the present invention, the method for obtaining an anthropomorphic eigenvalue vector includes:

[0077] When judging anthropomorphic features, the anthropomorphic vector can be first used to quantify the degree of anthropomorphism of the image in multiple dimensions each time the image is generated. The dimensions of the anthropomorphic vector include appearance features (such as facial features, body structure and clothing), behavioral features (such as movement flexibility and emotional expression), and language ability (such as conversation ability and language style). By scoring these dimensions, an anthropomorphic vector can be preliminarily constructed. For example, a neural network can be trained, and the input keyword vector corresponding to the input text of the image is used as the input of the neural network. Scoring can be performed from the aforementioned multiple dimensions to obtain the score values ​​on each dimension to form an anthropomorphic vector. In addition, since a dimension in the anthropomorphic vector may correspond to multiple keywords, such as appearance features, including facial features, body structure and clothing, when obtaining the anthropomorphic vector, the correspondence between each dimension in the anthropomorphic vector and the element value in the input keyword vector can be determined.

[0078] Each time an image is generated, the input keyword vector can be used as the input of a pre-trained neural network to obtain the anthropomorphic vector corresponding to the generated image and the correspondence between the element values ​​in the anthropomorphic vector and the element values ​​in the output keyword vector.

[0079] Similarly, an anthropomorphic vector of a comparison image is obtained and used as a comparison vector, wherein the comparison image includes an original electron microscope image and a standard human portrait. The original electron microscope image serves as a comparison image with an anthropomorphic degree of 0, which can be obtained through an image acquisition device, such as using a microscope and an imaging device to obtain an image of a real Vibrio cholerae; the standard human portrait serves as a comparison image with an anthropomorphic degree of 10.

[0080] Then, in the anthropomorphic vector and each comparison vector at each time of image generation, the cosine similarity between the element values ​​at the same position is used as the characteristic factor. The characteristic factor is used to characterize the similarity of the degree of anthropomorphism between the output image and the comparison image at each time of image generation. For example, if the characteristic factor between the output image and the standard portrait is larger, the degree of anthropomorphism of the output image is higher, and the more interesting it is; conversely, if the characteristic factor between the output image and the original electron microscope image is larger, the authenticity of the output image is higher, and there may be a risk of lower interest.

[0081] Finally, the element value at each position in the anthropomorphic vector of the generated image and the characteristic factor between the original electron microscope image are negatively correlated to obtain the interest factor, which is used to reflect the interest of the element value at each position in the anthropomorphic vector of the generated image. The larger the value, the lower the similarity with the original electron microscope image, and the higher the interest. The interest factor is multiplied by the characteristic factor between the element value at each position in the anthropomorphic vector of the generated image and the standard portrait, and the obtained product is normalized as the element value at the same position in the anthropomorphic characteristic value vector, thereby obtaining the anthropomorphic characteristic value vector corresponding to the input text each time the image is generated. The negative correlation mapping here can be used as the formula ,in, To express the exponential function with the natural constant e as the base, x represents the independent variable.

[0082] At this point, the anthropomorphic feature value vector corresponding to the input text each time an image is generated can be obtained.

[0083] The anthropomorphic eigenvalue vector describes the interestingness of the image from the anthropomorphic style, but in the actual comparison process, the interestingness performance will be related to the descriptive effectiveness of the keywords of the input text, and each element value in the anthropomorphic eigenvalue vector will correspond to one or more keywords of the input text. Therefore, in this embodiment of the present invention, the descriptive effectiveness index of the input text and the anthropomorphic eigenvalue vector are combined each time an image is generated to determine the interestingness feature vector of the input text.

[0084] Preferably, in one embodiment of the present invention, the method for obtaining the interesting feature vector includes:

[0085] In this embodiment of the present invention, a linkage mechanism combining feature emphasis and content effectiveness is established, with the aim of improving the anthropomorphic features with a higher degree of effectiveness. Therefore, each time an image is generated, the number of element values ​​corresponding to each element value in the anthropomorphic feature value vector in the input keyword vector is used as the quantity factor of each element value in the anthropomorphic feature value vector.

[0086] Then, each time an image is generated, the product of the quantity factor of each element value in the anthropomorphic feature value vector and the description effectiveness index of the input text is used as the adjustment parameter corresponding to each element value in the anthropomorphic feature value vector. The larger the adjustment parameter, the more keywords corresponding to the element value and the higher the effectiveness, so the anthropomorphic feature of the element value should be amplified.

[0087] Therefore, the product of each element value in the anthropomorphic feature value vector and the corresponding adjustment parameter is normalized and used as the element value at the corresponding position in the interestingness feature vector, thereby obtaining the interestingness feature vector of the input text for each image generation. Normalization is a technical method well known to those skilled in the art, and the normalization function can be linear normalization or standard normalization, etc. The specific normalization method is not limited here.

[0088] Step S4: Based on the interesting feature vector of the input text when the current image is generated, the input keyword vector of the input text when the next image is generated is adjusted to obtain an input keyword adjustment vector, which is used in the production process of interesting medical popular science pictures to obtain the next output image.

[0089] Based on the above steps, the interesting feature vector of the input text can be determined each time an image is generated. Then, based on the interesting feature vector of the current image generation, the keyword vector of the input text during the next image generation can be adjusted to make the interesting feature more prominent, thereby obtaining the input keyword adjustment vector. Finally, it is used in the production process of interesting popular science pictures to obtain the next output image.

[0090] Preferably, in one embodiment of the present invention, the method for obtaining the input keyword adjustment vector includes:

[0091] In the interesting feature vector of the input text during the current image generation and the input keyword vector of the input text during the next image generation, the low-dimensional vectors are zero-filled to obtain a vector group. At this time, in the vector group, the dimensions of the original interesting feature vector and the input keyword vector can be kept consistent. In this vector group, the vector corresponding to the input keyword vector can be used as the first vector, and the vector corresponding to the interesting feature vector can be used as the second vector.

[0092] Then, a balance coefficient is introduced to fuse the element values ​​in the first vector and the second vector: for the element value at the same position in the first vector and the second vector, the product of the first balance coefficient and the element value in the first vector is used as the first adjustment factor, the product of the second balance coefficient and the element value in the second vector is used as the second adjustment factor, and the sum of the first adjustment factor and the second adjustment factor is used as the element value at the same position in the input keyword adjustment vector, thereby obtaining the input keyword adjustment vector of the input text during the next image generation.

[0093] It should be noted that, given that the purpose of the interesting feature vector is to adjust the interestingness, when obtaining the input keyword adjustment vector, the proportion of element values ​​in the input keyword vector should be larger. Therefore, the value range of the first balance coefficient is (0.5, 1), and the value range of the second balance coefficient is (0, 0.5), and the sum of the two should be 1. In this embodiment of the present invention, the first balance coefficient can be set to 0.7, then the second balance coefficient is 0.3.

[0094] At this point, the input keyword adjustment vector of the input text when generating the next image can be obtained, and then the vector can be applied to the DALL-E2 model to obtain the next output image. The next output image takes into account the influence of interesting features and will therefore be more in line with the requirements of interesting science images. Therefore, the number of user adjustments can be reduced as much as possible, allowing the model to obtain the interesting science images required by users faster and more accurately.

[0095] In summary, the generation process of interesting medical popular science pictures is a process that requires constant adjustment. The model will automatically generate images based on the input text, and then the user can continue to input text to adjust the image in order to obtain interesting medical popular science images that meet the requirements and can effectively attract the target audience. In the production process of interesting medical popular science pictures, the embodiment of the present invention obtains the input keyword vector, output keyword vector and image vector corresponding to each image generation, which can be used to analyze the image generation situation in the subsequent process. Each time text is input, it is actually another processing of the image. In order to quantify the effectiveness of the input text for image processing, the input keyword vector, output keyword and image vector of two adjacent image generations can be compared to obtain the description effectiveness index of the input text each time the image is generated. Furthermore, since the influence of the fun and authenticity of the image mainly lies in the trade-off between retaining its own morphological characteristics and the degree of anthropomorphism, in order to avoid generating an image with too high authenticity but lack of fun, it is necessary to further analyze the interesting expression method. The anthropomorphic vector of each image generated can be obtained and compared with the anthropomorphic vector of the comparison image to obtain the anthropomorphic feature value vector of the input text at each image generation, which is used to reflect the anthropomorphic characteristics of the output image. Furthermore, the description effectiveness index of the input text at each image generation and the anthropomorphic feature value vector are combined to determine the interestingness feature vector of the input text, which can characterize the interestingness. Finally, based on the interestingness feature vector, the model is improved. Specifically, the input keyword vector of the input text for the next image generation is adjusted to obtain the input keyword adjustment vector. This balances authenticity and interestingness, thereby improving the accuracy of the generated results, effectively reducing the number of generation times, and producing interesting medical popular science images that better meet user needs and expectations.

[0096] An embodiment of the present invention also provides an interesting medical popular science production system based on artificial intelligence, including a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and when the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor, the steps of an interesting medical popular science production method based on artificial intelligence are implemented.

[0097] See also Figure 3 , which shows a system block diagram of an interesting medical popular science production system based on artificial intelligence, including a data acquisition module 301, used to implement step S1 in the above method embodiment; a description validity analysis module 302, used to implement step S2 in the above method embodiment; a fun analysis module 303, used to implement step S3 in the above method embodiment; and an output image adjustment module 304, used to implement step S4 in the above method embodiment.

[0098] It should be noted that the system provided in the above embodiment is merely an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the above embodiment provides an interesting medical science popularization production system based on artificial intelligence and an interesting medical science popularization production method based on artificial intelligence, which are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0099] See also Figure 4 , which shows a system structure diagram of an interesting medical popular science production system based on artificial intelligence provided by an embodiment of the present invention, including a processor 400, a memory 401, a bus 402 and a communication interface 403, wherein the processor 400, the communication interface 403 and the memory 401 are connected via the bus 402; wherein the memory 401 may include a high-speed random access memory, the bus 402 may be an ISA bus, a PCI bus or an EISA bus, etc., and the processor 400 may be an integrated circuit chip with signal processing capabilities; the memory 401 stores at least one instruction, at least one program, a code set or an instruction set, and when the at least one instruction, at least one program, a code set or an instruction set is loaded and executed by the processor, the steps in an interesting medical popular science production method based on artificial intelligence are implemented.

[0100] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0101] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

Claims

1. A method for producing interesting medical popular science based on artificial intelligence, characterized in that: The method comprises: In the production process of interesting medical popular science pictures, the input keyword vector corresponding to the input text, the output keyword vector corresponding to the output image, and the image vector are obtained each time an image is generated; Compare the differences between the input keyword vectors, the output keyword vectors, and the image vectors during two consecutive image generation times to determine the description effectiveness index of the input text during each image generation; Obtaining and comparing the anthropomorphic vector of each image generated with the anthropomorphic vector of the comparison image, thereby quantifying the anthropomorphic feature value vector of the input text at each image generation; determining the interestingness feature vector of the input text based on the description effectiveness index of the input text at each image generation and the anthropomorphic feature value vector; Based on the interesting feature vector of the input text during the current image generation, the input keyword vector of the input text during the next image generation is adjusted to obtain an input keyword adjustment vector, which is used in the production process of the interesting medical popular science illustration to obtain the next output image; The method for obtaining the anthropomorphic eigenvalue vector includes: Each time an image is generated, the input keyword vector is used as the input of the pre-trained neural network to obtain the anthropomorphic vector corresponding to the generated image and the correspondence between the element values ​​in the anthropomorphic vector and the element values ​​in the output keyword vector; Obtaining an anthropomorphic vector of a comparison image and using it as a comparison vector, wherein the comparison image includes an original electron microscope image and a standard human portrait; In the anthropomorphic vector and each comparison vector at each image generation, the cosine similarity between the element values ​​at the same position is used as the feature factor; The value obtained by negatively correlating the characteristic factor between the element value at each position in the anthropomorphic vector of the generated image and the original electron microscope image is multiplied by the characteristic factor between the element value at each position in the anthropomorphic vector of the generated image and the standard portrait. The obtained product is normalized and used as the element value at the same position in the anthropomorphic characteristic value vector, thereby obtaining the anthropomorphic characteristic value vector corresponding to the input text during each image generation; Describe the effectiveness index, which is used to reflect the correspondence between the input text and the output image, and characterize the effectiveness of the change in the style information of the input text; The anthropomorphism vector is used to quantify the degree of anthropomorphism of the image in multiple dimensions each time the image is generated.

2. The method for producing interesting medical popular science based on artificial intelligence according to claim 1, characterized in that: The method for obtaining the description effective indicator includes: Compare the differences between the input keyword vectors when two adjacent images are generated to determine the text change index; Compare the differences between the output keyword vectors of two adjacent image generation times to determine the content change index; Compare the difference between the image vectors when two adjacent images are generated to determine the image change index; When two adjacent images are generated, the product of the text change index and the content change index is negatively correlated and normalized to obtain the value as the effective factor; A value obtained by normalizing the product of the image change index and the effectiveness factor is used as a description effectiveness index of the input text in the latter image generation between two adjacent image generation.

3. The method for producing interesting medical popular science based on artificial intelligence according to claim 2, characterized in that: The method for obtaining the text change indicator includes: Compare the two input keyword vectors, determine the low-dimensional vector and the high-dimensional vector, and pad the low-dimensional vector with zeros to obtain the padding vector; In the two input keyword vectors, the cosine similarity between the high-dimensional vector and the filling vector is negatively correlated and normalized to obtain a value which is used as the text change indicator.

4. The method for producing interesting medical popular science based on artificial intelligence according to claim 2, characterized in that: The method for obtaining the content change indicator includes: Compare the two output keyword vectors, determine the low-dimensional vector and the high-dimensional vector, and fill the low-dimensional vector with zeros to obtain the feature vector; In the two output keyword vectors, the cosine similarity between the high-dimensional vector and the feature vector is negatively correlated and normalized to a value that serves as the content change indicator.

5. The method for producing interesting medical popular science based on artificial intelligence according to claim 2, characterized in that: The method for obtaining the image change index includes: Calculate the Euclidean distance between two image vectors as an image change indicator.

6. The method for producing interesting medical popular science based on artificial intelligence according to claim 1, characterized in that: The method for obtaining the interesting feature vector includes: Each time an image is generated, the number of element values ​​corresponding to each element value in the input keyword vector in the anthropomorphic feature value vector is used as the quantity factor of each element value in the anthropomorphic feature value vector; When each image is generated, the product of the quantity factor of each element value in the anthropomorphic feature value vector and the description effectiveness index of the input text is used as the adjustment parameter corresponding to each element value in the anthropomorphic feature value vector; The product of each element value in the anthropomorphic feature value vector and the corresponding adjustment parameter is normalized and used as the element value at the corresponding position in the interesting feature vector, thereby obtaining the interesting feature vector of the input text each time an image is generated.

7. The method for producing interesting medical popular science based on artificial intelligence according to claim 1, characterized in that: The acquisition of the input keyword adjustment vector includes: After aligning the interesting feature vector of the input text during the current image generation with the input keyword vector of the input text during the next image generation using the zero-padding method, a vector group is obtained; In the vector group, the vector corresponding to the input keyword vector is used as the first vector, and the vector corresponding to the interesting feature vector is used as the second vector; In the first vector and the second vector, for the element value at the same position, the product of the first balance coefficient and the element value in the first vector is used as the first adjustment factor, the product of the second balance coefficient and the element value in the second vector is used as the second adjustment factor, and the sum of the first adjustment factor and the second adjustment factor is used as the element value at the same position in the input keyword adjustment vector, so as to obtain the input keyword adjustment vector of the input text when the next image is generated.

8. The method for producing interesting medical popular science based on artificial intelligence according to claim 1, characterized in that: In the process of producing the interesting medical popular science illustrations, obtaining the input keyword vector corresponding to the input text, the output keyword vector corresponding to the output image, and the image vector each time an image is generated includes: In the process of generating interesting medical popular science pictures, the image vector of the output image is extracted through the pre-trained CLIP model, the input text is processed based on the BioBERT model to obtain the input keyword vector, and the output keyword vector corresponding to the output image is generated through the image description model.

9. An interesting medical science popularization production system based on artificial intelligence, characterized by: The invention comprises a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and when the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor, the steps of the method for producing interesting medical popular science based on artificial intelligence as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Building structure column image thumbnail intelligent generation method for intelligent construction

    CN116542859A

  • Medical image report automatic quality control error correction system and method

    CN119724466A