Intelligent writing assistant based on Mendel randomization system

Through the Mendel randomization system's intelligent writing assistant, the author's unique writing genes are identified, diverse writing variants are generated, and personalized style reinforcement suggestions are provided, which solves the problem of insufficient causal reasoning in existing writing tools, and achieves the diversity and scientific rigor of writing content.

CN120578744APending Publication Date: 2025-09-02CHENGDU KNOWLEDGE VISION SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510745303.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Existing writing auxiliary tools rely too much on word co-occurrence statistics, lack causal reasoning, difficult to capture and strengthen the author's unique style, easy to fall into patterning within generation, lack of diversity, and lack of scientific basis for revision suggestions.

Method used

The Mendel randomization system is adopted to identify the author's unique writing genes by analyzing text phenotype characteristics, generate diverse writing variants, use the random allocation principle of genetic variation, provide personalized style enhancement suggestions, detect potential confounding variables, generate rigorous causal expressions, eliminate interfering factors, and predict writing effects.

Benefits of technology

The diversity and scientific rigor of writing content are achieved, the patterning of generated content is avoided, and the suggestions for improvement of evidence-based writing are provided, the author's personalized style is respected, and the content innovation is ensured randomness and naturalness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578744A_ABST
    Figure CN120578744A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of intelligent assistants, and provides an intelligent writing assistant based on a Mendel randomization system, which comprises a database, a multi-modal MR NLP integration module, a causal engine reasoning module, a content processing module and a causal writing guidance module. The NLP integration module of the multi-mode MR is used for combining genetic data with other modes through the multi-mode MR; the causal engine inference module is used for establishing a'genotype-phenotype 'relation network among text elements by simulating Mendel randomized causal inference logic, and analyzing and identifying real effective factors in writing through tool variables; the causal writing guidance module is used for identifying writing elements really influencing the reaction of the reader; the assistant identifies unique writing genes of authors by analyzing phenotypic characteristics of texts, provides personalized style strengthening suggestions, and outputs Mendel randomization to research causal relationships on the basis of a database, so that confounding factors are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent assistants, and more particularly, to an intelligent writing assistant based on a Mendelian randomization system. Background Art

[0002] Traditional natural language processing technology is based on statistical language models, rule engines and templated writing systems, and the application of machine learning methods in text generation; modern deep learning writing assistance is based on large language models of Transformer architecture, neural sequence-to-sequence generation models, and the application of attention mechanisms in the coherence of long texts.

[0003] However, existing writing assistance tools have the following limitations: over-reliance on word co-occurrence statistics, lack of causal reasoning, unexplainable decision-making process, easy interference from confounding factors, lack of scientific basis for modification suggestions, difficulty in capturing and strengthening the author's unique style "gene", and the generated content easily falls into patterning and lacks true diversity.

[0004] Mendelian randomization (MR) is a statistical method that uses genetic variation as an instrumental variable to infer causal relationships between exposure factors and outcomes. Therefore, applying this scientific methodology to the field of intelligent writing can build a writing assistance system with unique advantages. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the purpose of the present invention is to provide an intelligent writing assistant based on the Mendelian randomization system. By analyzing the "phenotypic" characteristics of the text, the author's unique "writing gene" is identified, and personalized style enhancement suggestions are provided. Based on the database, the Mendelian randomization study causal relationship is output to avoid confounding factors. At the same time, it can detect potential confounding variables in the text, recommend instrumental variables, and generate more rigorous causal statements; through the anti-confusion suggestion module, interfering factors are eliminated and the core expression is focused; through the phenotypic prediction module, the possible effects of different writing styles are predicted, avoiding the generated content from falling into stereotypes and increasing the diversity of the generated content.

[0006] To achieve the above object, the present invention provides the following technical solutions: An intelligent writing assistant based on a Mendelian randomization system includes a database, a multimodal MR NLP integration module, a causal engine reasoning module, a content processing module, and a causal writing guidance module. The multimodal MR NLP integration module combines genetic data with other modalities through multimodal MR to enhance causal inference. The causal engine reasoning module establishes a "genotype-phenotype" relationship network between text elements by imitating the causal inference logic of Mendelian randomization, and identifies truly effective factors in writing through instrumental variable analysis. The causal writing guidance module is used to identify writing elements that truly influence reader responses. The causal writing guidance module includes a causal-driven text generation unit, a counterfactual content correction unit, and an academic writing assistance unit.

[0007] The present invention is further configured as follows: the database is used to integrate the GWAS database, knowledge graph, and academic corpus.

[0008] The present invention is further configured as follows: the content processing module includes a content generation unit and a style optimization unit; the content generation unit automatically generates diversified writing variants based on the random allocation principle of genetic variation, ensuring the randomness and naturalness of content innovation.

[0009] The present invention is further configured as follows: the style optimization unit is used to analyze the "phenotypic" characteristics of the text, identify the author's unique "writing gene", and provide personalized style enhancement suggestions.

[0010] The present invention is further configured as follows: the causal-driven text generation unit is based on a database, outputs Mendelian randomization to study causal relationships, and avoids confounding factors; the counterfactual content correction unit is used to detect potential confounding variables in the text, recommend instrumental variables, and generate a more rigorous causal statement.

[0011] The present invention is further configured as follows: the academic writing assistance unit includes an MR literature mining subunit and a chart generation subunit; the MR literature mining subunit is used to automatically scan the database and extract MR research conclusions as arguments; the chart generation subunit is used to visualize the ternary relationship of "gene-exposure-outcome" as a causal diagram and embed it into the paper.

[0012] The present invention is further configured to include an anti-confusion suggestion module, a genetic diversity generation module, and a phenotype prediction module; the anti-confusion suggestion module is used to eliminate interference factors and focus on core expression; the genetic diversity generation module is based on an algorithm to generate naturally varying content options; and the phenotype prediction module is used to predict the possible effects of different writing methods.

[0013] The present invention is further configured as follows: the NLP integration module of the multimodal MR includes a cross-modal instrumental variable construction unit, a multimodal causal graph generation unit and a dynamic hybrid control unit.

[0014] The present invention is further configured as follows: the cross-modal instrumental variable construction unit models correlation through a gene-phenotype database + user text embedding vector → multimodal variational autoencoder; the multimodal causal graph generation unit automatically generates an interactive graph to display the causal chain of "gene → brain image → symptom description"; the dynamic confounding control unit extracts environmental data as an additional covariate, inputs it into a multimodal MR model, and generates a corrected text.

[0015] The advantages of the present invention are: 1. This invention establishes a "genotype-phenotype" relationship network between text elements by imitating the causal inference logic of Mendelian randomization, and identifies the truly effective factors in writing through instrumental variable analysis. Based on the random allocation principle of genetic variation, it automatically generates diverse writing variants to ensure the randomness and naturalness of content innovation. It integrates scientific research methods into the creative process, providing writers with intelligent assistance that combines scientific rigor and artistic creativity.

[0016] 2. This invention analyzes the "phenotypic" characteristics of the text, identifies the author's unique "writing gene", and provides personalized style enhancement suggestions. Based on the database, it outputs Mendelian randomization to study causal relationships and avoid confounding factors. At the same time, it can detect potential confounding variables in the text, recommend instrumental variables, and generate more rigorous causal statements.

[0017] 3. The present invention eliminates interference factors and focuses on core expression through the anti-confusion suggestion module; through the phenotype prediction module, it predicts the possible effects of different writing methods, avoids the generated content from falling into stereotypes, and increases the diversity of the generated content. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 The figure shows a framework diagram of an intelligent writing assistant based on a Mendelian randomization system according to the present invention.

[0019] Figure 2 This is a framework diagram of the cause-and-effect writing guidance module of the present invention.

[0020] Figure 3 This is a framework diagram of the academic writing assistance unit of the present invention.

[0021] Figure 4 This is a framework diagram of the NLP integration module of the multimodal MR of the present invention. DETAILED DESCRIPTION

[0022] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0023] It should be noted that, unless otherwise specified, all technical and scientific terms used in this application have the same meaning as commonly understood by ordinary technicians in the technical field to which this application belongs.

[0024] In the present invention, unless otherwise specified, directions such as "up" and "down" are generally used with respect to the directions shown in the drawings, or with respect to the vertical, perpendicular or gravity directions; similarly, for ease of understanding and description, "left" and "right" are generally used with respect to the left and right shown in the drawings; "inside" and "outside" refer to the inside and outside relative to the outline of each component itself, but the above-mentioned directions are not used to limit the present invention. Example

[0025] See also Figure 1-4 , the present invention provides the following technical solutions: An intelligent writing assistant based on a Mendelian randomization system specifically includes a database, a multimodal MR NLP integration module, a causal engine reasoning module, a content processing module, and a causal writing guidance module; the multimodal MR NLP integration module combines genetic data with other modalities (images, text, and environment) through multimodal MR to enhance causal inference; the causal engine reasoning module establishes a "genotype-phenotype" relationship network between text elements by imitating the causal inference logic of Mendelian randomization, and identifies truly effective factors in writing through instrumental variable analysis; the causal writing guidance module is used to identify writing elements that truly influence reader responses; the causal writing guidance module includes a causal-driven text generation unit, a counterfactual content correction unit, and an academic writing assistance unit.

[0026] Working principle of this embodiment 1: By imitating the causal inference logic of Mendelian randomization, a "genotype-phenotype" relationship network is established between text elements, and the truly effective factors in writing are identified through instrumental variable analysis. Based on the random allocation principle of genetic variation, diverse writing variants are automatically generated to ensure the randomness and naturalness of content innovation.

[0027] By analyzing the "phenotypic" characteristics of the text, identifying the author's unique "writing gene", and providing personalized style enhancement suggestions, it outputs Mendelian randomization to study causal relationships based on the database to avoid confounding factors. At the same time, it can detect potential confounding variables in the text, recommend instrumental variables, and generate more rigorous causal statements. Example

[0028] See also Figure 1-4,This second embodiment makes the following improvements based on the first embodiment. Specifically, the database is used to integrate GWAS databases (such as UK Biobank), knowledge graphs (such as Wikidata), and academic corpora.

[0029] The content processing module includes a content generation unit and a style optimization unit; the content generation unit is based on the random allocation principle of genetic variation, automatically generating diverse writing variants to ensure the randomness and naturalness of content innovation.

[0030] The style optimization unit is used to analyze the "phenotypic" characteristics of the text, identify the author's unique "writing genes", and provide personalized style enhancement suggestions.

[0031] The causal-driven text generation unit is based on a database and outputs Mendelian randomization to study causal relationships and avoid confounding factors; the counterfactual content correction unit is used to detect potential confounding variables in the text, recommend instrumental variables, and generate more rigorous causal statements.

[0032] The academic writing assistance unit includes an MR literature mining subunit and a chart generation subunit; the MR literature mining subunit is used to automatically scan the database and extract MR research conclusions as arguments; the chart generation subunit is used to visualize the ternary relationship of "gene-exposure-outcome" as a causal graph (DAG) and embed it into the paper.

[0033] It also includes an anti-confusion suggestion module, a genetic diversity generation module, and a phenotype prediction module; the anti-confusion suggestion module is used to eliminate interference factors and focus on core expression; the genetic diversity generation module is based on an algorithm to generate natural variation content options; the phenotype prediction module is used to predict the possible effects of different writing methods.

[0034] The NLP integration module of multimodal MR includes a cross-modal instrumental variable construction unit, a multimodal causal graph generation unit, and a dynamic confounding control unit.

[0035] The cross-modal instrumental variable construction unit models correlation through the gene-phenotype database (such as GTEx) + user text embedding vector → multimodal variational autoencoder (MM-VAE); the multimodal causal graph generation unit automatically generates an interactive graph showing the causal chain of "gene → brain imaging → symptom description"; the dynamic confounding control unit extracts environmental data (such as satellite remote sensing) as additional covariates, inputs it into the multimodal MR model (such as MVMR), and generates corrected text.

[0036] Working principle of the second embodiment: Mendelian randomization uses genetic variants (such as SNPs) as instrumental variables to infer the causal relationship between exposure factors and outcomes. In smart writing, this logic can be abstracted as: Instrumental variables: Keywords, grammatical structures or knowledge graph nodes entered by users serve as “genetic tools” to exclude confounding factors (such as fuzzy expressions).

[0037] Causal inference: Generate highly credible content (such as popular science explanations and academic discussions) through the correlation between instrumental variables and text objectives (such as thematic coherence and logical chains).

[0038] Core function design 1) Causation-driven text generation Hypothetical scenario: User inputs "Does smoking cause lung cancer?" MR logic: Instrumental variable: Automatically extract "CHRNA5 gene (known to be associated with smoking addiction)" as an instrumental variable.

[0039] Causal generation: Based on the medical literature database, the output is "Mendelian randomization studies have shown that people with CHRNA5 gene variants have an increased tendency to smoke and a significantly increased risk of lung cancer, supporting the causal effect of smoking." Advantages: Avoids confounding factors (such as "smokers may also drink alcohol") and improves scientificity.

[0040] (2) Counterfactual content correction Detect potential confounding variables in the text (such as "economic level affects educational attainment") and recommend instrumental variables (such as "implementation time of the Compulsory Education Law") to generate more rigorous causal statements.

[0041] (3) Academic writing assistance MR literature mining: Automatically scan databases such as PubMed and extract MR research conclusions as evidence.

[0042] Graph generation: Visualize the ternary relationship of "gene-exposure-outcome" as a causal graph (DAG) and embed it in the paper.

[0043] MR itself needs to meet three major assumptions (correlation, independence, and exclusivity), and an NLP equivalent verification mechanism needs to be designed.

[0044] This assistant combines multimodal MR (such as gene-image-text association) to generate richer content. It combines the rigor of causal science with the generation capabilities of AI. It is particularly suitable for writing scenarios that require evidence support, promoting the paradigm upgrade from "correlated description" to "causal expression". It combines the statistical assumptions of Mendelian randomization (MR) with the verification mechanism of natural language processing (NLP) and extends it to multimodal MR (such as gene-image-text association). This is the core challenge and innovation of building an intelligent writing assistant.

[0045] NLP Equivalence Verification Mechanism: Algorithmic Mapping of the Three Major MR Assumptions 1. Relevance Genetic definition: The instrumental variable (IV) must be strongly associated with the exposure factor (e.g., a SNP is significantly associated with smoking behavior in a GWAS).

[0046] NLP equivalent: Input: User-provided "instrumental variables" (such as the keyword "CHRNA5 gene") and "exposure factors" (such as "smoking").

[0047] Verification algorithm: The co-occurrence frequency and semantic similarity (such as BERT embedding cosine distance) between the two are calculated through knowledge graphs (such as Biomedical Wikidata).

[0048] If the similarity is lower than the threshold (e.g., F statistic < 10), it indicates “weak instrumental variable risk” and it is recommended to replace or supplement the evidence.

[0049] 2. Exchangeability Genetically defined: IVs should not affect the outcome through confounding factors (e.g., SNPs should not be associated with income level).

[0050] NLP equivalent: Promiscuous detection: Use causal discovery models (such as the PC algorithm) to scan text for potential confounding variables (e.g., “smokers are likely to drink alcohol”).

[0051] Example: When the sentence "People with CHRNA5 gene mutations have a high risk of lung cancer" is input, the system detects a missing variable "smoking frequency" and prompts you to verify the independence of the supplementary instrumental variables.

[0052] Adversarial training: Introducing confounding factor discriminators into generative models to reduce implicit biases (such as socioeconomic bias).

[0053] 3. Exclusion Restriction Genetic definition: IV can only affect the outcome through exposure factors and there is no other path. NLP equivalent: Path blocking: Construct a text causal diagram (e.g., using a DAG generation model) to detect illegal paths from instrumental variables to outcomes (e.g., “CHRNA5 gene → immunosuppression → lung cancer”).

[0054] If an illegal pathway is detected, counterfactual generation is triggered: "If smoking behavior does not exist, is the CHRNA5 gene still associated with lung cancer?"

[0055] NLP integration of multimodal MR refers to combining genetic data with other modalities (images, text, and context) to enhance causal inference. The intelligent writing assistant must support the following features: 1. Construction of cross-modal instrumental variables Gene-text instrumental variables: For example, by associating "language-related genes (such as FOXP2)" with user writing styles (such as sentence length and vocabulary complexity), we can infer the influence of genetics on expression.

[0056] Technical implementation: Genotype database (such as GTEx) + user text embedding vector → Multimodal Variational Autoencoder (MM-VAE) to model correlation.

[0057] 2. Multimodal Causal Graph Generation Input: User query "Genetic and environmental factors of depression".

[0058] Output: Text layer: MR conclusion (such as "SLC6A4 gene is associated with depression risk, p=1e-5").

[0059] Imaging layer: embedding fMRI data (such as default mode network activity) as a mediating variable.

[0060] Visualization: Automatically generate interactive maps showing the causal chain from "gene → brain imaging → symptom description".

[0061] 3. Dynamic hybrid control Environmental modality integration: For example, analyzing whether air pollution (PM2.5 data) confuses gene-asthma text descriptions.

[0062] method: Environmental data (e.g., satellite remote sensing) can be extracted as additional covariates and input into multimodal MR models (e.g., MVMR).

[0063] Generate the corrected text: "After adjusting for PM2.5, the association between the IL33 gene and asthma decreased by 30%."

[0064] In summary, by transforming the statistical rigor of MR into NLP computable verification rules and integrating multimodal data, this assistant can achieve: 1. Causally credible text generation (beyond correlation description); 2. Cross-modal evidence synthesis (e.g., joint inference of genes, images, and text); 3. Automatic bias control (simulating the MR hypothesis testing process).

[0065] Ultimately, it will promote the AI ​​writing revolution from "statistical correlation" to "causal expression", avoid the correlation trap of traditional writing AI, focus on causality, provide evidence-based writing improvement suggestions, respect the personalized development of the author's original "genotype", and generate results with scientific interpretability.

[0066] This system integrates scientific research methods into the creative process, providing writers with intelligent assistance that combines scientific rigor and artistic creativity.

[0067] Obviously, the embodiments described above are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0068] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, tasks, devices, components and / or combinations thereof.

[0069] It should be noted that the terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0070] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

[0071] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. An intelligent writing assistant based on the Mendelian randomization system, characterized by: It includes database, multimodal MR NLP integration module, causal engine reasoning module, content processing module, and causal writing guidance module; The NLP integration module of the multimodal MR combines genetic data with other modalities through multimodal MR to enhance causal inference; The causal engine reasoning module mimics the causal inference logic of Mendelian randomization to establish a "genotype-phenotype" relationship network between text elements and identifies the truly effective factors in writing through instrumental variable analysis; The causal writing guidance module is used to identify writing elements that truly affect readers' responses; the causal writing guidance module includes a causal-driven text generation unit, a counterfactual content correction unit, and an academic writing assistance unit.

2. The intelligent writing assistant based on the Mendelian randomization system according to claim 1, characterized in that: The database is used to integrate GWAS database, knowledge graph, and academic corpus.

3. The intelligent writing assistant based on the Mendelian randomization system according to claim 2, characterized in that: The content processing module includes a content generation unit and a style optimization unit; the content generation unit automatically generates diverse writing variants based on the random allocation principle of genetic variation, ensuring the randomness and naturalness of content innovation.

4. The intelligent writing assistant based on the Mendelian randomization system according to claim 3, characterized in that: The style optimization unit is used to analyze the "phenotypic" characteristics of the text, identify the author's unique "writing genes", and provide personalized style enhancement suggestions.

5. The intelligent writing assistant based on the Mendelian randomization system according to claim 4, characterized in that: The causal-driven text generation unit is based on a database and outputs Mendelian randomization to study causal relationships and avoid confounding factors; the counterfactual content correction unit is used to detect potential confounding variables in the text, recommend instrumental variables, and generate more rigorous causal statements.

6. The intelligent writing assistant based on the Mendelian randomization system according to claim 5, characterized in that: The academic writing assistance unit includes an MR literature mining subunit and a chart generation subunit; the MR literature mining subunit is used to automatically scan the database and extract MR research conclusions as arguments; the chart generation subunit is used to visualize the ternary relationship of "gene-exposure-outcome" as a causal diagram and embed it into the paper.

7. The intelligent writing assistant based on the Mendelian randomization system according to claim 6, characterized in that: It also includes an anti-confusion suggestion module, a genetic diversity generation module and a phenotype prediction module; the anti-confusion suggestion module is used to eliminate interference factors and focus on core expression; the genetic diversity generation module is based on an algorithm to generate natural variation content options; the phenotype prediction module is used to predict the possible effects of different writing methods.

8. The intelligent writing assistant based on the Mendelian randomization system according to claim 7, characterized in that: The NLP integration module of the multimodal MR includes a cross-modal instrumental variable construction unit, a multimodal causal graph generation unit, and a dynamic hybrid control unit.

9. The intelligent writing assistant based on the Mendelian randomization system according to claim 8, characterized in that: The cross-modal instrumental variable construction unit models correlation through a genotype database + user text embedding vector → multimodal variational autoencoder; the multimodal causal graph generation unit automatically generates an interactive graph showing the causal chain of "gene → brain image → symptom description"; the dynamic confounding control unit extracts environmental data as an additional covariate, inputs it into the multimodal MR model, and generates corrected text.