Content generation method, apparatus, device, storage medium, and program product

CN122822201APending Publication Date: 2026-09-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510350110.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]但是,上述医学检查报告的解读服务的灵活性和适配性较差,无法较好地解读各种医学检查报告,致使生成的解读结果较为片面,准确性和可读性均不足

Benefits of technology

[0017]第四方面,本公开实施例还提供了一种计算机可读存储介质,该存储介质存储有计算机程序,当计算机程序被处理器执行时,使得处理器实现本公开任意实施例所说明的内容生成方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822201A_ABST
    Figure CN122822201A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a content generation method, apparatus, device, storage medium and program product. The method comprises: in response to an image uploading operation, determining a received image as a target image; if it is determined that the target image belongs to a medical examination report, determining each field content contained in the target image; generating a model prompt word based on each field content, and calling a report interpretation model based on the model prompt word to generate a report interpretation result of the target image; wherein the report interpretation model is obtained by fine-tuning a pre-trained generative model using a medical examination report sample. In this way, the generalization ability of the report interpretation model can be improved, and the flexibility of medical report interpretation and the readability and accuracy of the medical report interpretation result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a content generation method, apparatus, device, storage medium, and program product. Background Technology

[0002] With the development of internet technology, more and more users are accustomed to searching for medical-related information online. However, due to the highly specialized nature of medicine, the information available online is limited. This is especially true for understanding and searching medical examination reports, where the internet struggles to provide satisfactory results. Therefore, online services for interpreting medical examination reports have emerged. These services can perform character recognition and information extraction on images of medical reports uploaded by users, and then use certain medical rules (such as whether indicators exceed normal ranges) or machine learning algorithms to understand and organize the extracted information, generating interpretations of the medical examination reports.

[0003] However, the above-mentioned medical examination report interpretation services are not flexible and adaptable, and cannot interpret various medical examination reports well, resulting in interpretation results that are rather one-sided, with insufficient accuracy and readability. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure provides a content generation method, apparatus, device, storage medium, and program product.

[0005] In a first aspect, embodiments of this disclosure provide a content generation method, the method comprising:

[0006] In response to an image upload operation, the received image is identified as the target image;

[0007] If it is determined that the target image belongs to a medical examination report, then the content of each field contained in the target image is determined;

[0008] Model prompts are generated based on the content of each field, and the report interpretation model is invoked based on the model prompts to generate the report interpretation result of the target image; wherein, the report interpretation model is obtained by fine-tuning a pre-trained generative model using samples of medical examination reports.

[0009] Secondly, embodiments of this disclosure also provide a content generation apparatus, the apparatus comprising:

[0010] The target image determination module is used to determine the received image as the target image in response to the image upload operation;

[0011] The field content determination module is used to determine the content of each field contained in the target image if it is determined that the target image belongs to a medical examination report;

[0012] The report interpretation result generation module is used to generate model prompt words based on the content of each field, and to call the report interpretation model based on the model prompt words to generate the report interpretation result of the target image; wherein, the report interpretation model is obtained by fine-tuning a pre-trained generative model using medical examination report samples.

[0013] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:

[0014] processor;

[0015] Memory, used to store executable instructions;

[0016] The processor is used to read executable instructions from memory and execute the executable instructions to implement the content generation method described in any embodiment of this disclosure.

[0017] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the content generation method described in any embodiment of this disclosure.

[0018] Fifthly, embodiments of this disclosure also provide a computer program product for executing the content generation method described in any embodiment of this disclosure.

[0019] The content generation method, apparatus, device, storage medium, and program product of this disclosure are capable of responding to an image upload operation by identifying a received image as a target image; if the target image is determined to belong to a medical examination report, the content of each field contained in the target image is determined; model prompts are generated based on the content of each field, and a report interpretation model obtained by fine-tuning a pre-trained generative model using medical examination report samples is invoked based on the model prompts to generate a report interpretation result for the target image; this realizes the interpretation of images from medical examination reports using a report interpretation model adapted to the medical field. On the one hand, it can leverage a widely trained generative model of a certain scale to improve the generalization ability of the report interpretation model, without the need for predefined rules or templates, thus improving the flexibility of medical report interpretation and its adaptability to various medical examination reports; on the other hand, it can utilize the learning ability and contextual analysis ability of the report interpretation model to automatically capture complex nonlinear relationships between multiple medical indicators, and combine medical history, symptoms, and examination data to generate more comprehensive, fluent, and user-friendly medical report interpretation results, improving the readability and accuracy of the interpretation results, thereby enhancing the user experience.

[0020] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. Attached Figure Description

[0021] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0022] Figure 1 A flowchart illustrating a content generation method provided in an embodiment of this disclosure;

[0023] Figure 2 A schematic diagram of the startup interface for a medical examination report function provided in an embodiment of this disclosure;

[0024] Figure 3 A schematic diagram of a session interface provided in an embodiment of this disclosure;

[0025] Figure 4 A schematic diagram of an image upload interface provided in an embodiment of this disclosure;

[0026] Figure 5 A schematic diagram illustrating the display of report interpretation results provided in an embodiment of this disclosure;

[0027] Figure 6 A flowchart illustrating another content generation method provided in this embodiment of the disclosure;

[0028] Figure 7 This is a schematic diagram of the structure of a content generation apparatus provided in an embodiment of the present disclosure;

[0029] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0032] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0034] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0035] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0036] Medical examination reports encompass information about the body's condition obtained through various medical means. These can include test reports such as blood and urine routine tests, as well as reports from electrocardiograms, endoscopic examinations of internal cavities of organs, pathology reports, and functional examination reports of organs. There are two main types of solutions for automatically interpreting medical examination reports. One type is based on medical rule systems, which uses preset medical report interpretation templates to automatically generate fixed summaries based on abnormal values ​​of specific indicators (e.g., values ​​exceeding the normal range). However, the fixed templates used in this type of method result in poor flexibility and readability in interpreting medical reports. Furthermore, it can only analyze single indicators and lacks the ability to perform multi-indicator correlation analysis, leading to lower accuracy in the interpretation results. The other type is based on machine learning algorithms, which uses traditional machine learning algorithms to extract and analyze important keywords or medical entities (such as disease names, drugs, symptoms, etc.) in medical examination reports and generate concise report interpretation results, such as "elevated white blood cell count" or "anemia." However, this type of method suffers from problems such as insufficient contextual understanding, limited generalization ability, and poor readability due to unfriendly interpretation results.

[0037] Based on the above, this disclosure provides a content generation scheme that utilizes a pre-trained generative model (such as a language model) of a certain scale based on medical examination report samples to obtain a report interpretation model adapted to the medical examination reports. This report interpretation model is then used to interpret the uploaded medical examination report images. This eliminates the need to set medical rules and templates, improving the flexibility of report interpretation. Furthermore, the learning ability of the report interpretation model, combined with relevant medical knowledge (such as nonlinear relationships between multiple indicators) and user-related historical information (such as historical medical records, historical diagnoses, and symptoms), improves the accuracy of the report interpretation results. Additionally, the natural language generation capability of the report interpretation model enhances the fluency and user-friendliness of the report interpretation results, thereby improving readability and user experience.

[0038] The content generation method provided in this disclosure is applicable to scenarios involving the automatic interpretation of medical examination reports, such as search scenarios in the medical field and human-computer dialogue scenarios in the medical field. This method can be executed by a content generation device, which can be implemented in software and / or hardware. This device can be integrated into an electronic device with client-side functionality (such as an application, app, or webpage) that enables the aforementioned scenarios. This electronic device may include, but is not limited to, smartphones, personal digital assistants (PDAs), tablet computers (Tablet PCs), laptops, desktop computers, or medical self-service terminals.

[0039] Figure 1 A flowchart illustrating a content generation method provided in an embodiment of this disclosure is shown. Figure 1 As shown, the content generation method may include the following steps:

[0040] S110. In response to the image upload operation, the received image is identified as the target image.

[0041] The image upload operation is an interactive operation used to upload images, which can be implemented through pre-set interactive controls with image upload functionality. The target image is the image to be interpreted and processed, such as content generation.

[0042] Specifically, electronic devices can provide users with various forms of interactive interfaces, such as interactive controls or prompts for uploading images. Users can perform image upload operations through these interfaces. In response to the image upload operation, the electronic device can identify the received image as the target image for this processing.

[0043] In some embodiments, S110 includes: displaying an image upload interface and displaying an image upload control in the image upload interface in response to a trigger operation of the report interpretation control; and determining the received image as the target image in response to a trigger operation of the image upload control.

[0044] The report interpretation control is the interactive control corresponding to the interactive entry point used to trigger the report interpretation function; for example, it could be an interactive button with the aforementioned functions. The image upload interface is a visual interface displaying relevant prompts and interactive controls for uploading images; it can be a page, a floating window, or a floating layer. The image upload control is an interactive control with image upload functionality; for example, it could be a "Take a Photo" button or a "Select Image from Album" interactive button.

[0045] Specifically, see Figure 2The client for interpreting medical examination reports can be implemented in any form, such as a mini-program, application, or webpage. When a user launches the client, the electronic device can display the client's startup interface 200, which shows brief information about the client (such as an icon and a brief introduction like "xxx Doctor: Your Personalized Health Consultation Assistant"), a report interpretation control 210, and a session entry control 220 (such as a "View / Enter" control) to access the client's session page. The user can trigger the medical examination report interpretation function by performing a trigger operation on the report interpretation control 210, which will cause the electronic device to display the image upload interface.

[0046] Alternatively, users can also trigger operations on the aforementioned session entry control 220 to enter, for example... Figure 3 The human-computer interaction interface 300 is shown. This interface 300 is used to display the real-time conversation content between the user and the client. A report interpretation control 320 can be displayed in the area surrounding the input box 310 of the conversation content on the interface 300, or in the top area of ​​the interface 300. When the user performs a trigger operation on the report interpretation control 320, the electronic device can display an image upload interface.

[0047] The above image upload interface can be viewed as follows: Figure 4 As shown in the floating window 400, it can display prompts for the report interpretation function (such as function introduction, types of interpretable medical examination reports, requirements for uploaded images, etc.), examples of uploaded images, and image upload controls. The image upload controls can be a "Take a Photo" button 410 that triggers the image acquisition function, and / or a "Select from Album" button 420 that triggers the image selection function, etc.

[0048] When the user triggers the "Take Photo" button 410 or the "Select from Album" button 420, the electronic device can execute the corresponding image acquisition or image selection process to obtain the target image determined by the user. This allows for faster acquisition of the target image through explicit interactive controls.

[0049] S120. If it is determined that the target image belongs to a medical examination report, then determine the content of each field contained in the target image.

[0050] Specifically, the target image uploaded by the user may not be a medical examination report, in which case the electronic device cannot correctly perform the report interpretation function. Therefore, the electronic device first uses a machine learning model or a matching rule based on the characteristics of medical reports to identify whether the target image belongs to a medical examination report. If the target image is not a medical examination report, the process ends and the user is given a relevant error message. If the target image is a medical examination report, information can be extracted from the target image to extract fields and their values ​​that have report analysis value, which will be used as the content of each field. For example, for a routine blood test report, fields such as age, gender, specimen type, test item name, test result, reference range of test value, and unit of test result can be extracted from the target image.

[0051] S130. Generate model prompt words based on the content of each field, and call the report interpretation model based on the model prompt words to generate the report interpretation result of the target image.

[0052] The model prompt is a text input by the user to guide the model in generating specific outputs. The report interpretation model is obtained by fine-tuning a pre-trained generative model using samples of medical examination reports. This generative model can be a language model of a certain scale; it can be a general-purpose base model pre-trained with a large number of diverse training samples, or a medical-specific base model pre-trained with medical-related training samples. The medical examination report samples can include field content from various types of medical examination reports and their corresponding expert interpretations.

[0053] Specifically, to improve the adaptability of the generative model to the interpretation of medical examination reports, this embodiment of the disclosure can pre-obtain a certain number of medical examination report samples and use these samples to fine-tune and train the generative model, obtaining a report interpretation model more suitable for the medical examination report interpretation scenario. This allows the generative model to undergo larger-scale training, improving its generalization ability. Furthermore, it enables the model to learn the nonlinear relationships between multiple medical examination indicators based on its existing knowledge, thereby possessing the ability to comprehensively analyze the relationships between multiple medical examination indicators and their corresponding symptoms, improving the accuracy of the medical report interpretation results. Moreover, the natural language processing capabilities of the generative model can be leveraged to improve the fluency and human-likeness of the output medical report interpretation results, thus enhancing the readability of the report interpretation results.

[0054] After obtaining the content of each field in the target image, the electronic device can use this content to fill in the placeholders in a pre-built prompt template, forming model prompts. Then, the electronic device inputs the model prompts into the report interpretation model to trigger its execution. Following the output format of the model prompts, it outputs a summary of the model's interpretation of the target image, i.e., the report interpretation result. Afterward, the electronic device can display the report interpretation result in the session interface. Figure 5 As shown, after the user uploads the target image, the electronic device can display the report interpretation result 510 in the session interface 500.

[0055] The content generation method provided in the above embodiments of this disclosure can respond to an image upload operation and determine the received image as the target image. If the target image is determined to belong to a medical examination report, the content of each field contained in the target image is determined. Model prompt words are generated based on the content of each field, and based on the model prompt words, a report interpretation model obtained by fine-tuning a pre-trained generative model using samples of medical examination reports is called to generate the report interpretation result of the target image. This realizes the interpretation of images of medical examination reports using a report interpretation model adapted to the medical field. On the one hand, it can improve the generalization ability of the report interpretation model by leveraging a widely trained generative model of a certain scale, without the need for predefined rules or templates, thus improving the flexibility of medical report interpretation and its adaptability to various medical examination reports. On the other hand, it can automatically capture the complex nonlinear relationships between multiple medical indicators by leveraging the learning ability and contextual analysis ability of the report interpretation model, and combine medical history, symptoms, and examination data to generate more comprehensive, fluent, and humanized medical report interpretation results, improving the readability and accuracy of the interpretation results, thereby enhancing the user experience.

[0056] Figure 6 This is a flowchart illustrating another content generation method provided in this disclosure. In some embodiments, the content generation method can further refine the step of "determining the received image as the target image in response to an image upload operation." Based on this, the step of "generating model prompts based on the content of each field" can be further refined. In other embodiments, the content generation method can further refine the step of "determining the content of each field contained in the target image if it is determined that the target image belongs to a medical examination report." See also... Figure 6 The content generation method specifically includes the following steps:

[0057] S601. In response to the start operation of the conversation function, the conversation interface is displayed, and the conversation content is received in real time through the conversation interface.

[0058] Among them, real-time conversation content refers to the specific content of conversation messages obtained / received in real time during human-computer interaction.

[0059] Specifically, see [link to relevant documentation] Figure 2 When a user performs a trigger operation on the session entry control 220, it indicates that the user wants to initiate the human-computer dialogue function for health consultation. The electronic device can then respond to this trigger operation by displaying, as shown below. Figure 3 The example shown is the conversation interface 300. Users can input their medical questions through this interface, and the client can provide feedback based on a language model specific to the medical field in its backend. These medical questions and feedback can all serve as the content of the real-time conversation.

[0060] S602. Analyze the content of the instant conversation to determine whether there is any intention to interpret the report during the conversation.

[0061] Specifically, in addition to the explicit interactive controls mentioned in the aforementioned embodiments, the interactive entry point for the medical examination report interpretation function can also be determined automatically through the content of the human-computer dialogue to identify whether the user has a need / intent to interpret the medical examination report (i.e., report interpretation intent). This allows for better identification of user needs, enabling the timely provision of an interactive entry point for the image upload function, thus better guiding the user to use the medical examination report interpretation function and further enhancing the user experience. Based on this, the electronic device can perform intent recognition on the real-time acquired dialogue content to determine whether a report interpretation intent exists during the dialogue.

[0062] In some embodiments, S602 can be implemented as follows: using target matching rules and / or a session analysis model to analyze the real-time session content and determine whether there is a report interpretation intent during the session.

[0063] The target matching rule is a pre-set keyword matching rule that can be configured based on relevant keywords in the medical examination report and the matching business scenario. For example, if the matching business scenario is a coarse matching scenario, the target matching rule can be implemented as a keyword matching rule related to the medical examination report with relatively weak matching constraints (i.e., the first matching rule), such as a rule to match specific keywords like "medical report," "physical examination," "laboratory test," "blood test data," "complete blood count," and "urinalysis," as well as similar keywords. Conversely, if the matching business scenario is a precise matching scenario, the target matching rule can be implemented as a keyword matching rule related to the medical examination report with relatively strong matching constraints (i.e., the second matching rule), such as a rule to match specific keywords like "medical report," "physical examination," "laboratory test," "blood test data," "complete blood count," and "urinalysis." The conversation analysis model is a pre-trained language model whose function is to identify the intent / needs of the medical examination report through text analysis.

[0064] Specifically, electronic devices can use methods such as target matching rules, session analysis models, and combinations of target matching rules and session analysis models to analyze the content of each real-time session in order to identify whether the user intends to interpret reports during the session.

[0065] In some examples, analyzing real-time conversation content using target matching rules to determine whether a report interpretation intent exists during the conversation can be implemented as follows: Electronic devices can use predefined target matching rules to perform keyword matching on each piece of real-time conversation content. If a match is successful, it can be determined that a report interpretation intent exists during the conversation; if a match fails, it can be determined that no report interpretation intent exists during the conversation. This allows for intent determination using relatively simple keyword matching rules, reducing the implementation cost of report interpretation intent recognition to some extent and improving the efficiency of report interpretation intent recognition.

[0066] In other examples, a conversation analysis model is used to analyze real-time conversation content to determine whether there is an intent to interpret reports during the conversation. This can be implemented as follows: the electronic device can generate model prompts using the content of each real-time conversation, and input these prompts into the conversation analysis model to trigger its execution. The model outputs a result indicating whether there is an intent to interpret reports, such as whether a medical report is mentioned and whether report interpretation is required. If the model output indicates that no medical report is mentioned, or that a medical report is mentioned but interpretation is not required, it can be determined that there is no intent to interpret reports during the conversation. Conversely, if the model output indicates that a medical report is mentioned and interpretation is required, it can be determined that there is an intent to interpret reports during the conversation. This can improve the accuracy of intent recognition to some extent.

[0067] In other examples, the instant conversation content is analyzed using target matching rules and conversation analysis models to determine whether there is an intention to interpret reports during the conversation. This can be achieved by: using the first matching rule to perform keyword matching on each instant conversation content; if the match is successful, the conversation analysis model is called based on the target content to generate the model output results.

[0068] The model output indicates whether the conversation process contains an intention to interpret a report. The target content includes at least one of the following: each instant conversation content, successfully matched keywords, and the instant conversation content corresponding to the keywords.

[0069] Specifically, to further improve the accuracy of intent recognition while controlling implementation costs and efficiency, electronic devices can first use a first matching rule with relatively low matching accuracy to perform keyword matching on each instant conversation content. If the match fails, it indicates that the possibility of each instant conversation content having a report interpretation intent is very small, and it can be directly determined that there is no report interpretation intent, thus ending the current report interpretation process. This eliminates the need to execute the conversation analysis model, saving operating costs. If the match succeeds, it indicates that the possibility of each instant conversation content having a report interpretation intent is relatively high. Then, model prompts can be constructed using the target content and input into the conversation analysis model for intent recognition, obtaining the model output results. If the model output indicates that a medical report is mentioned and report interpretation is required, then it can be determined that the conversation process has a report interpretation intent. If the model output indicates that a medical report is not mentioned, or a medical report is mentioned but report interpretation is not required, then it can be determined that the conversation process does not have a report interpretation intent. This further improves the accuracy of intent recognition.

[0070] In other examples, the analysis of real-time conversation content using target matching rules and conversation analysis models can determine whether there is a report interpretation intent during the conversation. This can be achieved as follows: based on each real-time conversation content, the conversation analysis model is invoked to generate model output results; if the model output results indicate the existence of a report interpretation intent, then the second matching rule is used to perform keyword matching on each real-time conversation content; if the matching is successful, it is determined that there is a report interpretation intent during the conversation; if the matching fails, it is determined that there is no report interpretation intent during the conversation.

[0071] Specifically, considering that the conversation analysis model is more accurate in identifying intent than the first matching rule, but it may still have some recognition errors due to insufficient training, model illusions, and other reasons, this embodiment first constructs model prompt words using the content of each instant conversation, and then uses these to call the conversation analysis model to generate model output results. If the model output results indicate that there is no intention to interpret the report, the current report interpretation process ends. If the model output results indicate that there is an intention to interpret the report, it means that there is a significant need for report interpretation in the conversation process. At this time, the second matching rule can be used to perform more rigorous keyword matching on the content of each instant conversation. If the matching fails, it means that the model identification is incorrect, and there is no intention to interpret the report in the conversation process. If the matching is successful, it can be determined that there is an intention to interpret the report in the conversation process. This can further improve the accuracy of identifying the intention to interpret the report, thereby further improving the timeliness of the subsequent display of the image upload entry, avoiding interference with the user, and further improving the user experience.

[0072] S603. If so, the prompt conversation content for the medical report interpretation function will be displayed in the conversation interface; the prompt conversation content includes the function entry control.

[0073] The prompt session content is used to prompt users to upload images of their medical examination reports. This content may include relevant prompt text and function entry controls. These function entry controls are the interactive controls that lead to the medical report interpretation function; for example, they could be an "Upload Report" button.

[0074] Specifically, if the conversation does not involve a report interpretation intent, the electronic device can continue to answer and provide feedback on the medical questions entered by the user. If the conversation does involve a report interpretation intent, the electronic device can display the prompt conversation content as immediate feedback to the user in the conversation interface.

[0075] See also Figure 3 After the electronic device recognizes the intent of the instant conversation content "high cholesterol in blood routine test", it determines that there is a need for report interpretation. Then, it can display a prompt conversation content 330 after the instant conversation content in the conversation interface 300, and display the "upload report" button 331 in it to guide the user to upload the image of the medical examination report.

[0076] S604. In response to the trigger operation of the function entry control, display the image upload interface and display the image upload control in the image upload interface.

[0077] Specifically, if a user performs a trigger operation on the function entry control of the "Upload Report" button (example 331), the electronic device can display the following in response to the trigger operation: Figure 4 The image upload interface is shown, and the image upload control is displayed in it.

[0078] S605. In response to the trigger operation of the image upload control, the received image is identified as the target image.

[0079] S606. Perform image recognition and information extraction on the target image to determine the content of each field contained in the target image.

[0080] Specifically, after acquiring the target image, the electronic device can use image recognition and information extraction algorithms, such as optical character recognition, to identify and extract the fields contained in the target image and obtain the content of each field.

[0081] It should be noted that during the image recognition and information extraction process, personal user information may be ignored to protect user privacy.

[0082] S607. Use the third matching rule to perform keyword matching on the content of each field.

[0083] The third matching rule is another pre-set keyword matching rule with relatively weak matching constraints. It can be set according to the content characteristics of the medical examination report. For example, keyword matching rules can be set using keywords unique to medical examination reports such as "hospital", "indicator unit", "physician / doctor", "indicator value", "reference range / reference value" (such as whether the above keywords exist, the number of times the keywords are successfully matched, etc.).

[0084] Specifically, electronic devices can first use a third matching rule to perform keyword matching on the content of each field to determine whether the target image has specific content of a medical examination report, thereby determining whether it belongs to a medical examination report.

[0085] S608. If the match is successful, the target image is determined to belong to a medical examination report.

[0086] Specifically, if a matching keyword is found, it indicates that the target image belongs to a medical examination report, and the subsequent steps continue. If no matching keyword is found, i.e., the match fails, it indicates that the target image does not belong to a medical examination report, and the process ends, reducing resource waste and improving operational efficiency.

[0087] S606′: Input the target image into the preset classification model to determine the classification probability that the target image belongs to the image category corresponding to the medical examination report.

[0088] Among them, the preset classification model is a pre-trained model with classification function, which can predict the probability that the target image belongs to a medical examination report.

[0089] Specifically, after obtaining the target image, the target image can be input into a preset classification model, and the model can be run to output the classification probability of the target image belonging to the image category corresponding to the medical examination report.

[0090] S607′ If the classification probability exceeds the preset threshold, the target image is determined to belong to a medical examination report.

[0091] The preset threshold is a pre-set probability critical value.

[0092] Specifically, if the classification probability does not exceed a preset threshold, it indicates that the target image does not belong to a medical examination report, and the process ends. Conversely, if the classification probability exceeds the preset threshold, it can be determined that the target image belongs to a medical examination report.

[0093] S608′ Perform image recognition and information extraction on the target image to determine the content of each field contained in the target image.

[0094] Specifically, electronic devices can utilize image recognition and information extraction algorithms, such as optical character recognition, to identify and extract fields contained in a target image, thereby obtaining the content of each field.

[0095] S609. Use the fourth matching rule to perform keyword matching on the content of each field.

[0096] The fourth matching rule is another pre-set keyword matching rule with stronger matching constraints than the third matching rule. It can be set according to the content characteristics of the medical examination report.

[0097] Specifically, based on the preliminary determination that the target image belongs to a medical examination report through a preset classification model, the fourth matching rule can be used to perform keyword matching on the content of each field in the target image to further determine whether the target image belongs to a medical examination report.

[0098] S610. If the match is successful, the target image is determined to belong to a medical examination report.

[0099] S611. If the matching fails, the target image is determined not to belong to the medical examination report.

[0100] S612. Based on the content of each field and the content of each instant conversation, generate model prompt words, and call the report interpretation model based on the model prompt words to generate the report interpretation result of the target image.

[0101] Specifically, based on determining that the user needs report interpretation through the content of each instant conversation, model prompts can be constructed using the content of each field in the target image and the content of each instant conversation. These prompts are then used to invoke the report interpretation model. This provides the report interpretation model with contextual information from the current conversation (such as the user's medical history, symptoms, and past medical records) as supplementary information, increasing the amount of information the model can capture and thus further improving the accuracy of report interpretation.

[0102] The content generation method provided in the above embodiments can respond to the startup operation of the conversation function, display the conversation interface, and receive real-time conversation content through the conversation interface; analyze the real-time conversation content to determine whether there is an intention to interpret the report during the conversation; if so, display the prompt conversation content of the medical report interpretation function in the conversation interface; the prompt conversation content includes a function entry control; respond to the trigger operation of the function entry control, display the image upload interface, and display the image upload control in the image upload interface; respond to the trigger operation of the image upload control, determine the received image as the target image; thereby, it can better identify the user's needs, provide the user with the interactive entry point of the image upload function in a timely manner, so as to better guide the user to use the medical examination report interpretation function, and further improve the user experience. Furthermore, by generating model prompt words based on the content of each field and each real-time conversation content, the report interpretation model can be further provided with the conversation context content of the current conversation as supplementary information, thereby increasing the amount of information that the model can capture, and further improving the accuracy of report interpretation. In addition, by combining image recognition and information extraction with a third matching rule, resource waste can be reduced to a certain extent and image judgment efficiency can be improved based on identifying whether the target image is a medical examination report. By combining a pre-defined classification model with image recognition and information extraction, the accuracy of image judgment can be improved to some extent, beyond simply identifying whether a target image is a medical examination report. Furthermore, by combining a pre-defined classification model with image recognition, information extraction, and a fourth matching rule, the accuracy of image judgment can be further enhanced, thereby reducing resource waste in report interpretation and improving operational efficiency.

[0103] The following are embodiments of the content generation apparatus provided in this invention. This apparatus and the content generation methods in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the content generation apparatus, please refer to the embodiments of the above content generation methods.

[0104] Figure 7 A schematic diagram of the structure of a content generation apparatus provided in an embodiment of this disclosure is shown. For example... Figure 7 As shown, the content generation apparatus 700 may include:

[0105] The target image determination module 710 is used to determine the received image as the target image in response to the image upload operation;

[0106] The field content determination module 720 is used to determine the content of each field contained in the target image if it is determined that the target image belongs to a medical examination report;

[0107] The report interpretation result generation module 730 is used to generate model prompt words based on the content of each field, and call the report interpretation model based on the model prompt words to generate the report interpretation result of the target image; wherein, the report interpretation model is obtained by fine-tuning the pre-trained generative model using medical examination report samples.

[0108] The content generation apparatus provided in the above embodiments of this disclosure can respond to an image upload operation and determine the received image as the target image; if the target image is determined to be a medical examination report, the content of each field contained in the target image is determined; model prompt words are generated based on the content of each field, and a report interpretation model obtained by fine-tuning a pre-trained generative model using medical examination report samples is called based on the model prompt words to generate the report interpretation result of the target image; this realizes the interpretation of images of medical examination reports using a report interpretation model adapted to the medical field. On the one hand, it can improve the generalization ability of the report interpretation model by using a widely trained generative model of a certain scale, and does not require predefined rules or templates, thus improving the flexibility of medical report interpretation and its adaptability to various medical examination reports; on the other hand, it can automatically capture the complex nonlinear relationships between multiple medical indicators by using the learning ability and context analysis ability of the report interpretation model, and combine medical history, symptoms and examination data to generate more comprehensive, fluent and humanized medical report interpretation results, improving the readability and accuracy of the interpretation results, thereby improving the user experience.

[0109] In some embodiments, the target image determination module 710 includes:

[0110] The real-time conversation content receiving submodule is used to respond to the start operation of the conversation function, display the conversation interface, and receive real-time conversation content through the conversation interface.

[0111] The Report Interpretation Intent Determination submodule is used to analyze the content of real-time conversations to determine whether a report interpretation intent exists during the conversation.

[0112] The prompt conversation content display submodule is used to display the prompt conversation content for the medical report interpretation function in the conversation interface if the condition is met; the prompt conversation content includes the function entry control;

[0113] The image upload control display submodule is used to display the image upload interface in response to the trigger operation of the function entry control, and to display the image upload control in the image upload interface;

[0114] The target image determination submodule is used to determine the received image as the target image in response to the trigger operation of the image upload control.

[0115] In some embodiments, the report interpretation intent determination submodule is specifically used for:

[0116] By using target matching rules and / or conversation analysis models, the content of real-time conversations is analyzed to determine whether there is an intention to interpret reports during the conversation; the conversation analysis model is a pre-trained language model.

[0117] Furthermore, the report interpretation intent determination submodule is specifically used for:

[0118] The first matching rule is used to perform keyword matching on the content of each instant conversation;

[0119] If a match is successful, the conversation analysis model is invoked based on the target content to generate the model output results; the model output results indicate whether there is a report interpretation intent in the conversation process; the target content includes at least one of the following: each instant conversation content, successfully matched keywords, and the instant conversation content corresponding to the keywords.

[0120] Optionally, the report interpretation intent determination submodule is specifically used for:

[0121] The session analysis model is invoked based on the content of each real-time session to generate the model output results;

[0122] If the model output indicates an intention to interpret a report, then the second matching rule is used to perform keyword matching on the content of each instant conversation.

[0123] If a match is found, it is determined that the conversation process contains an intent to interpret a report;

[0124] If a match fails, it is determined that the session does not contain a report interpretation intent.

[0125] In some embodiments, the report interpretation result generation module 730 is specifically used for:

[0126] Based on the content of each field and each real-time conversation, model prompt words are generated.

[0127] In some embodiments, the field content determination module 720 is specifically used for:

[0128] Perform image recognition and information extraction on the target image to determine the content of each field contained in the target image;

[0129] Use the third matching rule to perform keyword matching on the content of each field;

[0130] If a match is successful, the target image is determined to belong to a medical examination report.

[0131] In other embodiments, the field content determination module 720 is specifically used for:

[0132] Input the target image into a preset classification model to determine the probability that the target image belongs to the image category corresponding to the medical examination report;

[0133] If the classification probability exceeds a preset threshold, the target image is determined to belong to a medical examination report;

[0134] Image recognition and information extraction are performed on the target image to determine the content of each field contained in the target image.

[0135] Furthermore, in some embodiments, the field content determination module 720 is also specifically used for:

[0136] After performing image recognition and information extraction on the target image and determining the content of each field contained in the target image, the fourth matching rule is used to perform keyword matching on the content of each field.

[0137] If the match is successful, the target image is determined to belong to a medical examination report;

[0138] If the match fails, it is determined that the target image does not belong to the medical examination report.

[0139] In other embodiments, the target image determination module 710 is specifically used for:

[0140] In response to the triggering operation of the report interpretation control, the image upload interface is displayed, and the image upload control is displayed in the image upload interface;

[0141] In response to a trigger operation on the image upload control, the received image is identified as the target image.

[0142] The content generation apparatus provided in the embodiments of the present invention can execute the content generation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0143] It is worth noting that in the embodiments of the above content generation device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of this disclosure.

[0144] This disclosure also provides an electronic device that may include a processor and a memory, the memory being used to store executable instructions. The processor may be used to read the executable instructions from the memory and execute the executable instructions to implement the content generation method described in the above embodiments.

[0145] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure is shown.

[0146] like Figure 8As shown, the electronic device 800 may include a processing unit 801 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output interface (I / O interface) 805 is also connected to the bus 804.

[0147] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touch screens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data.

[0148] It should be noted that, Figure 8 The illustrated electronic device 800 is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein. That is, although... Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0149] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the network isolation policy management method of any embodiment of this disclosure.

[0150] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the processor to implement the network isolation policy management method in any embodiment of this disclosure.

[0151] It should be noted that the computer-readable medium described above in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. Computer-readable storage media can be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, radio frequency (RF), etc., or any suitable combination thereof.

[0152] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as Hypertext Transfer Protocol (HTTP), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0153] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0154] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the network isolation policy management method described in any embodiment of this disclosure.

[0155] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0157] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0158] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0159] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0160] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A content generation method, characterized in that, include: In response to an image upload operation, the received image is identified as the target image; If it is determined that the target image belongs to a medical examination report, then the content of each field contained in the target image is determined; Model prompts are generated based on the content of each field, and the report interpretation model is invoked based on the model prompts to generate the report interpretation result of the target image; wherein, the report interpretation model is obtained by fine-tuning a pre-trained generative model using samples of medical examination reports.

2. The method according to claim 1, characterized in that, The step of determining the received image as the target image in response to the image upload operation includes: In response to the activation of the conversation function, the conversation interface is displayed, and real-time conversation content is received through the conversation interface; The content of the instant conversation is analyzed to determine whether there is any intention to interpret a report during the conversation; If so, the prompt session content for the medical report interpretation function will be displayed in the session interface; the prompt session content includes function entry controls; In response to the triggering operation of the function entry control, the image upload interface is displayed, and the image upload control is displayed in the image upload interface; In response to a trigger operation on the image upload control, the received image is identified as the target image.

3. The method according to claim 2, characterized in that, The analysis of the real-time conversation content to determine whether there is an intent to interpret a report during the conversation includes: The instantaneous conversation content is analyzed using target matching rules and / or a conversation analysis model to determine whether the intention to interpret the report exists during the conversation; the conversation analysis model is a pre-trained language model.

4. The method according to claim 3, characterized in that, By utilizing target matching rules and a session analysis model, the content of the real-time session is analyzed to determine whether the intent to interpret the report exists during the session, including: Using the first matching rule, keyword matching is performed on the content of each instant conversation; If a match is successful, the conversation analysis model is invoked based on the target content to generate model output results; the model output results indicate whether the conversation process contains the report interpretation intent; the target content includes at least one of the instant conversation content, the successfully matched keywords, and the instant conversation content corresponding to the keywords.

5. The method according to claim 3, characterized in that, By using matching rules and a session analysis model, the content of the instant session is analyzed to determine whether the intent to interpret the report exists during the session, including: Based on the real-time conversation content, the conversation analysis model is invoked to generate model output results; If the model output indicates the intention to interpret the report, then the second matching rule is used to perform keyword matching on each of the instant conversation contents; If the match is successful, it is determined that the session process contains the intent to interpret the report; If the match fails, it is determined that the session process does not contain the intent to interpret the report.

6. The method according to claim 2, characterized in that, The generation of model prompt words based on the content of each of the fields includes: The model prompt words are generated based on the content of each field and the content of each real-time conversation.

7. The method according to claim 1, characterized in that, If it is determined that the target image belongs to a medical examination report, then the content of each field contained in the target image is determined, including: The target image is subjected to image recognition and information extraction to determine the content of each field contained in the target image; The third matching rule is used to perform keyword matching on the content of each field. If a match is found, the target image is determined to belong to the medical examination report.

8. The method according to claim 1, characterized in that, If it is determined that the target image belongs to a medical examination report, then the content of each field contained in the target image is determined, including: The target image is input into a preset classification model to determine the probability that the target image belongs to the image category corresponding to the medical examination report. If the classification probability exceeds a preset threshold, then the target image is determined to belong to the medical examination report; The target image is subjected to image recognition and information extraction to determine the content of each field contained in the target image.

9. The method according to claim 8, characterized in that, After performing image recognition and information extraction on the target image to determine the content of each field contained in the target image, the method further includes: The fourth matching rule is used to perform keyword matching on the content of each of the fields. If a match is successful, the target image is determined to belong to the medical examination report; If the match fails, it is determined that the target image does not belong to the medical examination report.

10. The method according to claim 1, characterized in that, The step of determining the received image as the target image in response to the image upload operation includes: In response to a trigger operation on the report interpretation control, an image upload interface is displayed, and an image upload control is displayed in the image upload interface; In response to a trigger operation on the image upload control, the received image is identified as the target image.

11. A content generation apparatus, characterized in that, include: The target image determination module is used to determine the received image as the target image in response to the image upload operation; The field content determination module is used to determine the content of each field contained in the target image if it is determined that the target image belongs to a medical examination report; The report interpretation result generation module is used to generate model prompt words based on the content of each field, and to call the report interpretation model based on the model prompt words to generate the report interpretation result of the target image; wherein, the report interpretation model is obtained by fine-tuning a pre-trained generative model using medical examination report samples.

12. An electronic device, characterized in that, include: processor; Memory, used to store executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the content generation method according to any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, causes the processor to implement the content generation method according to any one of claims 1-10.

14. A computer program product, characterized in that, The computer program product is used to implement the content generation method according to any one of claims 1-10.