Ring genomics visualization workflow generation method based on large model driving

By using large-model-driven semantic parsing and multi-objective constraint solving, combined with positive and negative example retrieval and executable validation, a circular genomics visualization workflow is generated, which solves the problems of high threshold and complex configuration of existing tools and achieves fast and accurate visualization result generation.

CN121884964BActive Publication Date: 2026-07-21COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
Filing Date
2026-01-20
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing bioinformatics tools, visualization tools for gene variation, expression levels, and chromosome structure require users to manually configure parameters, resulting in a high learning threshold for beginners, cumbersome workflows, and difficulty in quickly generating accurate visualization results based on users' natural language commands.

Method used

A large model-driven approach is adopted, which parses natural language input through a semantic parsing module to generate an intent structure model. Combined with multi-objective constraint solving and positive and negative example retrieval of candidate circular layout schemes, the constraint objectives are optimized and verified, visual instructions are constructed, and verification and local repair are performed through an executable verification model to generate an interactive visual analysis model of the circular genome.

Benefits of technology

It lowers the operational threshold, improves research efficiency, automatically optimizes and generates scientific data mapping and aesthetically pleasing circular layouts, reduces runtime error rates, possesses trial-and-error capabilities similar to human experts, and solves the problem of difficulty in quickly generating accurate visualization results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884964B_ABST
    Figure CN121884964B_ABST
Patent Text Reader

Abstract

The application provides a large model driving-based annular genomics visualization workflow generation method, relates to the technical field of biological information, and comprises the following steps: a semantic analysis module of a natural language input large language model is analyzed to generate an intention structure model; a validation constraint target is constructed for the intention structure model and the characteristics of each genomic data, a candidate annular layout scheme is generated by solving the constraint retrieval target, positive and negative examples of the candidate annular layout scheme are retrieved, the constraint patch is updated based on the evidence set of the retrieval, and the validation constraint target is updated; the visualization instruction is obtained by constructing based on the optimized validation constraint target and the constraint patch; the visualization instruction is verified according to the executable verification model, local repair is performed and verification is recycled when the verification fails, and the annular genomic interactive visual analysis model is output until the verification is passed. The application solves the problem that it is difficult to quickly generate accurate visualization results according to user natural language instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics, and more specifically, to a method for generating circular genomics visualization workflows based on large model-driven approaches. Background Technology

[0002] In existing bioinformatics technologies, the visualization of gene variation, expression levels, and chromosome structure typically relies on pie charts. However, most of these tools require users to manually configure parameters and possess certain programming skills. This not only presents a high learning curve for beginners and results in a cumbersome and complex workflow, but also presents the challenge of quickly generating accurate visualization results based on user natural language commands.

[0003] Therefore, there is an urgent need for a method to generate circular genomics visualization workflows based on large models, which solves the problem of difficulty in quickly generating accurate visualization results based on user natural language commands. Summary of the Invention

[0004] The purpose of this invention is to provide a method for generating a circular genomics visualization workflow based on a large model, thereby improving the aforementioned problems. To achieve this objective, the technical solution adopted by this invention is as follows:

[0005] Firstly, this application provides a method for generating a large-model-driven circular genomics visualization workflow, including:

[0006] Acquire multiple genome data;

[0007] Natural language is input into the semantic parsing module of the large language model for parsing, generating an intent structure model;

[0008] The intent structure model and the features of each of the genome data are used to construct and verify constraint objectives. By solving the constraint retrieval objectives through multi-objective constraint solutions, candidate circular layout schemes are generated.

[0009] The candidate ring layout scheme is searched for positive and negative examples. Constraints are patched and the verification constraint target is updated based on the retrieved evidence set to obtain the optimized verification constraint target.

[0010] Based on the optimized verification constraint objective and the constraint patch, a visualization command is obtained;

[0011] The visualization instructions are validated based on the executable validation model. If the validation fails, local repair is performed and the validation is repeated until the validation passes, and the circular genome interactive visual analysis model is output.

[0012] Secondly, this application also provides a large-model-driven circular genomics visualization workflow generation device, including:

[0013] The acquisition module is used to acquire multiple genome data.

[0014] The parsing module is used to parse natural language input into the semantic parsing module of the large language model and generate an intent structure model;

[0015] The constraint module is used to construct and verify constraint targets for the intent structure model and the features of each of the genome data, and to generate candidate circular layout schemes by solving multi-objective constraints on the constraint retrieval targets.

[0016] The retrieval module is used to perform positive and negative example retrieval on the candidate ring layout scheme, perform constraint patching and update the verification constraint target through the retrieved evidence set, and obtain the optimized verification constraint target.

[0017] A construction module is used to build based on the optimization verification constraint target and the constraint patch to obtain visualization instructions;

[0018] The verification module is used to verify the visualization instructions according to the executable verification model. When the verification fails, it performs local repair and loops the verification until the verification passes and outputs the circular genome interactive visual analysis model.

[0019] The beneficial effects of this invention are as follows:

[0020] This invention improves research efficiency by introducing semantic understanding from a large language model, allowing researchers to focus on the biological problem itself without needing to concern themselves with the underlying drawing code implementation. It effectively solves the technical problems of traditional bioinformatics visualization tools, such as high barriers to entry, complex configurations, and difficulty in verifying layout logic. By combining multi-objective constraint solving with positive and negative example retrieval techniques for candidate circular layout schemes, it automatically optimizes and generates scientific data mappings and aesthetically pleasing circular layout schemes. Subsequently, by introducing an executable verification model and a local repair mechanism, it acquires trial-and-error capabilities similar to human experts. It can self-correct when faced with non-standard genomic data or logically conflicting drawing commands, reducing runtime error rates. In summary, this invention solves the problem of the difficulty in quickly generating accurate visualization results based on user natural language commands.

[0021] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of the method for generating a circular genomics visualization workflow based on a large model, as described in this embodiment of the invention.

[0024] Figure 2 This is a schematic diagram of the structure of the circular genomics visualization workflow generation device based on a large model driven by the present invention.

[0025] The diagram is labeled as follows: 800, a device for generating circular genomics visualization workflows based on large models; 801, processor; 802, memory; 803, multimedia component; 804, I / O interface; 805, communication component. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0027] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0028] Example 1:

[0029] This embodiment provides a method for generating a circular genomics visualization workflow based on a large model.

[0030] See Figure 1 The figure shows that the method includes steps S1 to S5, including:

[0031] S1: Acquire multiple genome data;

[0032] In this step, the genomic data includes genomic sequence information, annotation information, functional information, and other genome-related biological data.

[0033] Preferably, the genomic data includes a chromosome translocation dataset and a gene expression profile dataset;

[0034] S2: The semantic parsing module of the large language model parses the natural language input and generates an intent structure model;

[0035] In this step, the user's requests in natural language are transformed into instructions that the computer can understand. The computer quickly understands the user's natural language, and the user does not need to have in-depth knowledge of complex visualization tools. The computer can perform visualization operations of circular genomics simply by using natural language, which greatly reduces the operation threshold and time cost.

[0036] S3: Construct and verify constraint targets for the intent structure model and the features of each of the genome data, and generate candidate circular layout schemes by solving multi-objective constraints on the constraint retrieval targets;

[0037] To clarify the specific method for obtaining candidate ring layout schemes, step S3 includes S31 to S34, specifically:

[0038] S31: Jointly compile and verify the constraint target by combining the intention structure model and the data feature set in the genomic data, and generate a constraint intermediate representation;

[0039] In this step, the intent structure model and the set of data features in the genomic data are jointly compiled into a machine-readable and verifiable constrained intermediate representation, which includes hard constraints that must be met and soft constraints that can be optimized. Hard constraints include non-overlapping tracks and labels, minimum readability threshold, and visual channel consistency, while soft constraints include maximizing information density, minimizing visual clutter, prioritizing the display of user-focused areas, and minimizing interaction burden.

[0040] S32: Based on the hard constraints in the intermediate constraint representation, solve the feasibility of the candidate layout space to obtain a feasible candidate set;

[0041] In this step, in the predefined candidate layout space, the feasibility of each potential layout is determined according to the hard constraints, layouts that violate any hard constraints are eliminated, and a feasible candidate set that satisfies at least all hard constraints is selected to obtain the feasible candidate set.

[0042] S33: Perform candidate layout judgment on the feasible candidate set. If the constraint intermediate representation is not satisfied, perform pruning and backoff processing until a feasible solution is output.

[0043] In this step, the hard constraint satisfaction of each candidate layout in the feasible candidate set is checked. If it is found that it is not satisfied, pruning or rollback operations are performed until all the retained layouts fully meet the hard constraint requirements, and the final set of feasible solutions is output.

[0044] S34: The feasible solution is scored and iteratively optimized according to the preset multi-objective evaluation index. The optimized feasible solutions are sorted according to the scoring results, and the solution with the best score is selected as a candidate to generate a candidate ring layout scheme.

[0045] In this step, the soft constraints in the intermediate constraint representation are transformed into a multi-objective scoring function for scoring, and the scoring results are obtained. The feasible solution is quantitatively evaluated and iteratively optimized by combining a weight adaptive strategy. The optimized feasible solutions are sorted by the scoring results, and the layout with the highest comprehensive score is selected as the optimal candidate, generating a high-quality and verifiable candidate ring layout scheme.

[0046] S4: Perform positive and negative example retrieval on the candidate ring layout scheme, patch the constraints and update the verification constraint target through the retrieved evidence set, and obtain the optimized verification constraint target;

[0047] To clarify the specific method for obtaining the optimization verification constraint objective, step S4 includes S41 to S46, specifically:

[0048] S41: Construct the candidate ring layout scheme and the preset intent features to obtain the constructed retrieval signature;

[0049] In this step, the construction of the retrieval signature comprehensively covers the core feature dimensions, including intent features (user's task type, focus, and interaction needs), data features (annotation density level, sample size level, and combination of mutation types), layout structure (track set, track order, and channel mapping category), and constraint summary (key thresholds and ranges of hard constraints in the initial constraints). Through the structured integration of multi-dimensional features, a retrieval identifier that can accurately match relevant cases in the knowledge base is formed.

[0050] S42: Based on the constructed retrieval signature, perform positive example evidence retrieval to obtain positive example retrieval results;

[0051] In this step, the constructed retrieval signature is used as the query condition to filter out successful practical experiences that are compatible with the current candidate layout scheme. The positive example retrieval results specifically include recommended track combination methods, optimal parameter ranges, efficient interaction combination schemes, and the applicable scenarios and key success conditions of these successful cases.

[0052] S43: Based on the constructed retrieval signature, perform negative example evidence retrieval to obtain negative example retrieval results;

[0053] In this step, for the failure case library and the disabled pattern library in the constructed retrieval signature, failure experience with similar features to the current candidate scheme is extracted. The negative example retrieval results include track or label strategies that lead to unreadable behavior in high-density annotation and multi-sample comparison scenarios in typical failure layout patterns, specific conditions that trigger failure, types of violated constraints, disabled parameter combinations (parameter combinations that will lead to a surge in overlap rate), and repair suggestions.

[0054] S44: Generate constraint patches based on the negative example retrieval results, and transform the failure conditions in the negative example retrieval results through constraint patches to obtain new hard constraints;

[0055] In this step, failure conditions in the negative example retrieval results are transformed into new hard constraints with mandatory binding force. For failure conditions where color matching is unreadable under a specific background, they are transformed into mandatory standards for color contrast. This is to ensure that the new hard constraints can accurately avoid the identified failure risks and are semantically consistent with the initial constraints, and are machine-readable and verifiable.

[0056] S45: Convert the layout space pruning rules for the disabled combinations in the negative example retrieval results, and merge the layout space pruning rules with the newly added hard constraints to obtain the updated verification constraint target;

[0057] In this step, the prohibited parameter combinations and conflicting pairings of track type and channel mapping in the negative example retrieval results are transformed into layout space pruning rules, explicitly prohibiting such invalid combinations from appearing in subsequent layout generation. Then, according to the priority rules of security-readability-aesthetics-preference, the layout space pruning rules and the newly added hard constraints are integrated into the initial constraints to obtain the updated verification constraint target, which resolves the potential conflicts between constraints, forms the updated verification constraint target, and achieves full coverage of failure modes by the constraint system.

[0058] S46: Optimize the updated verification constraint target based on the positive example retrieval results. Adjust and interactively combine the threshold of the newly added hard constraint using the recommended parameters in the positive example retrieval results to obtain the optimized verification constraint target.

[0059] In this step, based on the successful parameter range and interaction combination experience in the positive example retrieval results, the threshold of the newly added hard constraint in the updated verification constraint objective is finely adjusted to avoid the threshold setting being too strict, resulting in too few feasible solutions. At the same time, the efficient interaction combination and track configuration in the positive examples are transformed into the basis for adjusting the weight of the soft constraint optimization objective. This allows the optimization verification constraint objective to retain the ability to avoid failure modes while also drawing on successful experience to improve the usability and optimization effect of the layout, ultimately forming a constraint system that takes into account both constraint rigidity and optimization flexibility.

[0060] S5: Based on the optimized verification constraint target and the constraint patch, construct and obtain visualization instructions;

[0061] To clarify the specific method for obtaining visualization instructions, step S5 includes S51 to S53, specifically:

[0062] S51: Based on the optimization verification constraint objective, the candidate ring layout schemes are scored and ranked in multiple dimensions to obtain the preferred layout set;

[0063] In this step, based on the optimization verification constraint objective, the candidate circular layout schemes are first judged according to the satisfaction of hard constraints, and schemes that do not meet the core constraints are eliminated. Then, for the remaining schemes that have passed the hard constraint screening, the soft constraints are weighted according to the weights adaptively allocated by the user intent structure model to obtain a comprehensive score and sort them in descending order. Finally, the top-N schemes with the highest scores are selected to form the preferred layout set.

[0064] S52: Perform differentiable and nondifferentiable index decomposition on the preferred layout set, and obtain the Pareto front by performing spatial parallel search through a differentiable rendering proxy model;

[0065] To clarify the specific method for obtaining the Pareto frontier, step S52 includes S521 to S524, which specifically include:

[0066] S521: Perform differentiable and nondifferentiable index decomposition on the preferred layout set to obtain differentiable and nondifferentiable indices;

[0067] In this step, the evaluation metrics of the preferred layout set are precisely mapped to specific constraints such as tracks, channels, and interactions, and a binary decomposition of differentiable and non-differentiable metrics is performed accordingly. Differentiable metrics include quantitative parameters that support gradient calculation, such as geometric conflict overlap rate, spacing compliance rate, and visual channel mapping adaptation, while non-differentiable metrics include subjective parameters that rely on experience-based judgment, such as visual clutter perception, interaction burden assessment, and the highlighting effect of attention areas.

[0068] S522: Calculate the gradient of the differentiable index with respect to the layout latent code in the latent code space according to the differentiable rendering proxy model, and obtain the gradient update direction;

[0069] In this step, the differentiable rendering proxy model is pre-trained using a large number of layout samples and rendering effect data to simulate the real rendering process and output differentiable index values. The latent code space maps the key parameters of the layout (radius, orbital order, channel mapping rules). The gradient direction of the differentiable index as the latent code changes is calculated using the gradient descent method, allowing the latent code adjustment path optimized by the differentiable index to obtain the gradient update direction. This provides directional guidance for precise optimization.

[0070] S523: Based on zero-order Markov Monte Carlo and same-policy reinforcement learning to estimate the expected reward, perform a multi-dimensional continuous layout latent code space parallel search on the non-differentiable index to obtain the non-differentiable latent code distribution.

[0071] In this step, the zero-order Markov Monte Carlo method is used to address the problem that gradients of non-differentiable indicators cannot be directly calculated. By randomly sampling latent code space samples of the non-differentiable indicators and evaluating the performance of the corresponding non-differentiable indicators, the same-policy reinforcement learning method, based on the current search policy, transforms the optimization effect of non-differentiable indicators into expected rewards, guiding the search process to focus on high-reward regions. The parallel search of the multi-dimensional continuous layout of the latent code space simultaneously covers multiple dimensions such as orbital parameters and interaction configurations, efficiently mining latent code combinations that meet the optimization requirements of non-differentiable indicators, and finally forming a statistically significant distribution of non-differentiable latent codes.

[0072] S524: Based on the Pareto dominance criterion, the gradient update direction and the distribution of the non-differentiable latent code are integrated and the non-dominated solution set is selected to obtain the Pareto front.

[0073] In this step, the Pareto dominance criterion is that the differentiable and non-differentiable indices corresponding to a certain latent code combination are not inferior to other combinations, and at least one index is better. During the integration process, the high-quality samples in the distribution of differentiable and non-differentiable latent codes guided by the gradient update direction need to be evaluated in a unified manner. Dominated latent code combinations are eliminated, and the non-dominated solution set is retained as the Pareto front. This front contains multiple optimal layout latent codes that weigh the optimization effects of different indices, providing diverse candidates for subsequent evolutionary optimization.

[0074] S53: The Pareto front is iteratively optimized according to the self-adversarial evolution mechanism of ring topology preservation to generate an evolutionary layout population;

[0075] In this step, the self-adversarial evolution mechanism for maintaining the circular topology is used to ensure that the layout always meets the core structural requirements of circular genome visualization during the evolution process, without problems such as track breakage or topological disorder. The self-adversarial evolution mechanism introduces adversarial logic between the generator and the discriminator. The generator performs mutation, crossover and other operations based on the Pareto front latent code to generate new layouts. The discriminator is responsible for evaluating the constraint satisfaction, index optimization effect and topological rationality of the new layouts. Through multiple rounds of iteration, layouts that combine constraint compliance, index superiority and topological stability are selected to form a stable and diverse population of evolutionary layouts.

[0076] S54: Compile the optimal layout in the evolutionary layout population into consistent instructions, and use the constraint patch to perform conflict detection and automatic correction on the compilation instructions to obtain visual instructions.

[0077] In this step, the optimal layout in the evolutionary layout population is comprehensively evaluated based on constraint satisfaction reports, index scores, and topological integrity to select the optimal layout scheme that meets the requirements. Subsequently, the optimal layout is used to perform conflict detection using the constraint patch. Based on the disabled modes and newly added hard constraints in the patch, layout defects that violate parameter combinations or fail to meet constraints are automatically corrected to avoid failure modes. Finally, the corrected scheme is compiled with consistency instructions to transform its orbit definition, channel mapping, and interaction rules into traceable, structured, and visual instructions.

[0078] S6: Validate the visualization instructions according to the executable validation model. If the validation fails, perform local repair and iterate the validation until the validation passes, and output the circular genome interactive visual analysis model.

[0079] To clarify the specific method for obtaining the circular genome interactive visual analysis model, step S6 includes S61 to S64, specifically:

[0080] S61: Verify the visualization instructions in the sandbox environment according to the executable verification model, drive the visualization instructions to execute and collect rendering verification indicators to obtain rendering indicator values;

[0081] In this step, the executable verification model is an executable verifier, which includes a rendering sandbox and an indicator calculation module. Within the sandbox environment, it drives the execution and rendering of the visualization instructions, simultaneously collecting multi-dimensional verification indicators covering geometric conflict, readability, data consistency, and interactive accessibility, ultimately generating quantified rendering indicator values. S62: The rendering indicator values ​​are input into a preset genomics visualization constraint objective function for back-substitution calculation to obtain the violation set;

[0082] In this step, the rendered index value is input into a preset genomics visualization constraint objective function for back-substitution calculation, the deviation between the index and the constraint threshold is identified and quantified, thereby generating a violation set that includes specific violation constraint clauses, precise location of violation tracks or component positions, assessment of violation severity and root cause analysis results, and further determining whether the violation has triggered the boundary conditions of negative example evidence.

[0083] S63: Select the minimum cost repair operator for the category of the violation set according to the hierarchical local repair strategy, and generate a new visualization instruction;

[0084] In this step, based on the severity of the violation set, the least-cost repair operator is selected first, and parameter-level (adjusting basic parameters), component-level (replacing type or encoding), and layout-level (rearranging or splitting tracks) repair operations are executed sequentially. The root cause of the repair, operator type, and parameter change differences are recorded simultaneously to generate a new visualization instruction after repair. For stubborn errors that cannot be repaired, the evidence-level repair mechanism is triggered to update the retrieved evidence set and generate a new constraint patch.

[0085] S64: Validate the new visualization instructions by fine-tuning the preset circular layout, and output the circular genome interactive visual analysis model when the validation result meets the standard.

[0086] In this step, the convergence of the new visualization instructions is verified by fine-tuning the preset ring layout. When all hard constraints are met and the soft constraint scores are met, a ring-shaped interactive visual analysis model of the genome containing executable configuration and component code is directly output, and a constraint satisfaction report and repair path record are generated simultaneously for updating the design knowledge base. If the preset iteration limit or parameter change threshold is reached, a rollback mechanism is triggered, returning to the most recently verified version to perform conservative modifications to avoid the efficiency loss of full regeneration, until the convergence condition is met.

[0087] Example 2:

[0088] This embodiment provides a large-model-driven circular genomics visualization workflow generation device, the device including:

[0089] The acquisition module is used to acquire multiple genome data.

[0090] The parsing module is used to parse natural language input into the semantic parsing module of the large language model and generate an intent structure model;

[0091] The constraint module is used to construct and verify constraint targets for the intent structure model and the features of each of the genome data, and to generate candidate circular layout schemes by solving multi-objective constraints on the constraint retrieval targets.

[0092] To clarify the specific methods for obtaining the constraint module, the following are included:

[0093] The compilation unit is used to jointly compile and verify the constraint target of the intent structure model and the set of data features in the genomic data, and generate an intermediate constraint representation.

[0094] The solution unit is used to solve the feasibility of the candidate layout space based on the hard constraints in the intermediate constraint representation, and obtain a set of feasible candidates;

[0095] The pruning unit is used to perform candidate layout judgment on the feasible candidate set. When the constraint intermediate representation is not satisfied, pruning and back-off processing is performed until a feasible solution is output.

[0096] The iteration unit is used to score and iteratively optimize the feasible solution according to the preset multi-objective evaluation index, sort the optimized feasible solutions according to the scoring results, select the solution with the best score as a candidate, and generate candidate ring layout schemes.

[0097] The retrieval module is used to perform positive and negative example retrieval on the candidate ring layout scheme, perform constraint patching and update the verification constraint target through the retrieved evidence set, and obtain the optimized verification constraint target.

[0098] To clarify the specific methods for obtaining information from the retrieval module, the following are included:

[0099] A construction unit is used to construct the candidate ring layout scheme and the preset intent features to obtain the constructed retrieval signature;

[0100] The positive example retrieval unit is used to perform positive example evidence retrieval based on the constructed retrieval signature to obtain positive example retrieval results.

[0101] The negative example retrieval unit is used to perform negative example evidence retrieval based on the constructed retrieval signature and obtain negative example retrieval results.

[0102] The patching unit is used to generate constraint patches based on the negative example retrieval results, and to transform the failure conditions in the negative example retrieval results through the constraint patches to obtain new hard constraints;

[0103] The conversion unit is used to convert the layout space pruning rules into the disabled combinations in the negative example retrieval results, and merge the layout space pruning rules with the newly added hard constraints to obtain the updated verification constraint target.

[0104] The optimization unit is used to optimize the updated verification constraint target based on the positive example retrieval results. It adjusts and interactively combines the threshold of the newly added hard constraint with the recommended parameters in the positive example retrieval results to obtain the optimized verification constraint target.

[0105] A construction module is used to build based on the optimization verification constraint target and the constraint patch to obtain visualization instructions;

[0106] To clarify the specific methods for obtaining the building modules, the following are included:

[0107] The scoring unit is used to score and rank the candidate ring layout schemes in multiple dimensions based on the optimization verification constraint objective, so as to obtain the preferred layout set.

[0108] The decomposition unit is used to perform differentiable and nondifferentiable index decomposition on the preferred layout set, and to obtain the Pareto front by performing spatial parallel search through a differentiable rendering proxy model.

[0109] To clarify the specific methods for obtaining the decomposition units, the following are included:

[0110] Decomposition subunits are used to decompose the preferred layout set into differentiable and nondifferentiable indices to obtain differentiable and nondifferentiable indices.

[0111] The computational subunit is used to calculate the layout latent code gradient of the differentiable index in the latent code space according to the differentiable rendering proxy model, and obtain the gradient update direction.

[0112] The search subunit is used to perform a multi-dimensional continuous layout latent code space parallel search on the non-differentiable index based on the expected reward estimated by zero-order Markov Monte Carlo and same-policy reinforcement learning to obtain the non-differentiable latent code distribution.

[0113] The filtering subunit is used to integrate the gradient update direction and the non-differentiable latent code distribution based on the Pareto dominance criterion and filter out the non-dominated solution set to obtain the Pareto front.

[0114] An iterative optimization unit is used to iteratively optimize the Pareto front according to the self-adversarial evolutionary mechanism of ring topology preservation, and generate an evolutionary layout population.

[0115] The instruction compilation unit is used to compile the optimal layout in the evolutionary layout population into consistent instructions, and to perform conflict detection and automatic correction on the compiled instructions through the constraint patch to obtain visual instructions.

[0116] The verification module is used to verify the visualization instructions according to the executable verification model. When the verification fails, it performs local repair and loops the verification until the verification passes and outputs the circular genome interactive visual analysis model.

[0117] It should be noted that the specific manner in which each module performs its operation in the apparatus described in the above embodiments has been described in detail in the embodiments of the method, and will not be elaborated here.

[0118] Example 3:

[0119] Corresponding to the above method embodiments, this embodiment also provides a large model-driven circular genomics visualization workflow generation device. The large model-driven circular genomics visualization workflow generation device described below and the large model-driven circular genomics visualization workflow generation method described above can be referred to in correspondence.

[0120] Figure 2 This is a block diagram illustrating a large model-driven circular genomics visualization workflow generation device 800 according to an exemplary embodiment. Figure 2 As shown, the large model-driven circular genomics visualization workflow generation device 800 may include: a processor 801 and a memory 802. The large model-driven circular genomics visualization workflow generation device 800 may also include one or more of the following: a multimedia component 803, an I / O interface 804, and a communication component 805.

[0121] The processor 801 controls the overall operation of the large model-driven circular genomics visualization workflow generation device 800 to complete all or part of the steps in the aforementioned large model-driven circular genomics visualization workflow generation method. The memory 802 stores various types of data to support the operation of the large model-driven circular genomics visualization workflow generation device 800. This data may include, for example, instructions for any application or method operating on the large model-driven circular genomics visualization workflow generation device 800, as well as application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as keyboards, mice, and buttons. These buttons can be virtual or physical. Communication component 805 is used for wired or wireless communication between the large model-driven circular genomics visualization workflow generation device 800 and other devices. Wireless communication includes Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof. Therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, and an NFC module.

[0122] In an exemplary embodiment, the large model-driven circular genomics visualization workflow generation device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the above-described large model-driven circular genomics visualization workflow generation method.

[0123] Example 4:

[0124] Corresponding to the above method embodiments, this embodiment also provides a medium. The medium described below can be referred to in relation to the above-described method for generating a circular genomics visualization workflow based on a large model.

[0125] A medium storing a computer program, which, when executed by a processor, implements the steps of the large model-driven circular genomics visualization workflow generation method described in the above method embodiments.

[0126] The medium can specifically be any medium capable of storing program code, such as a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0127] The above are merely preferred embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0128] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for generating a circular genomics visualization workflow driven by a large model, characterized in that, include: Acquire multiple genome data; Natural language is input into the semantic parsing module of the large language model for parsing, generating an intent structure model; The intent structure model and the features of each of the genome data are used to construct and verify constraint objectives. By solving the constraint retrieval objectives through multi-objective constraint solutions, candidate circular layout schemes are generated. The candidate ring layout scheme is searched for positive and negative examples. Constraints are patched and the verification constraint target is updated based on the retrieved evidence set to obtain the optimized verification constraint target. Based on the optimized verification constraint objective and the constraint patch, a visualization command is obtained; The specific methods for obtaining the visualization instructions include: Based on the optimization verification constraint objective, the candidate ring layout schemes are scored and ranked in multiple dimensions to obtain the preferred layout set; The preferred layout set is decomposed into differentiable and non-differentiable indices to obtain differentiable and non-differentiable indices. The evaluation metrics of the preferred layout set are precisely mapped to specific constraints of tracks, channels, and interactions, and a binary decomposition of differentiable and non-differentiable metrics is performed accordingly. Differentiable metrics include geometric conflict overlap rate, spacing compliance rate, visual channel mapping adaptation, and other quantitative parameters that support gradient calculation, while non-differentiable metrics include visual clutter perception, interaction burden assessment, attention area highlighting effect, and other subjective parameters that rely on experience judgment. The gradient of the differentiable index with respect to the layout latent code is calculated in the latent code space based on the differentiable rendering proxy model to obtain the gradient update direction; The differentiable rendering proxy model is pre-trained and generated using a large number of layout samples and rendering effect data to simulate the real rendering process and output differentiable index values. The latent code space maps the radius of the layout, the orbit order, the channel mapping rules, and other key parameters. The gradient direction of the differentiable index as the latent code changes is calculated using the gradient descent method, so that the latent code adjustment path optimized by the differentiable index is obtained, and the gradient update direction is obtained. Based on zero-order Markov Monte Carlo and same-policy reinforcement learning to estimate the expected reward, a multi-dimensional continuous layout latent code space parallel search is performed on the non-differentiable index to obtain the non-differentiable latent code distribution. The zero-order Markov Monte Carlo method is used to solve the problem that the gradient of non-differentiable indicators cannot be directly calculated. By randomly sampling latent code space samples of the non-differentiable indicators and evaluating the performance of the corresponding non-differentiable indicators, the same-policy reinforcement learning is based on the current search policy and transforms the optimization effect of non-differentiable indicators into expected rewards, guiding the search process to focus on high reward regions. The parallel search of the multi-dimensional continuous layout of the latent code space covers orbital parameters, interaction configurations and other multiple dimensions, and mines latent code combinations that meet the optimization requirements of non-differentiable indicators, ultimately forming a non-differentiable latent code distribution. Based on the Pareto dominance criterion, the gradient update direction and the non-differentiable latent code distribution are integrated and the non-dominated solution set is selected to obtain the Pareto front. The Pareto front is iteratively optimized based on the self-adversarial evolutionary mechanism of ring topology preservation to generate an evolutionary layout population. The optimal layout in the evolutionary layout population is compiled with consistent instructions, and the compilation instructions are conflict-detected and automatically corrected by the constraint patch to obtain visual instructions. The visualization instructions are validated based on the executable validation model. If the validation fails, local repair is performed and the validation is repeated until the validation passes, and the circular genome interactive visual analysis model is output.

2. The method for generating a circular genomics visualization workflow based on a large model as described in claim 1, comprising constructing and validating constraint objectives for the intent structure model and the features of each of the genomic data, and generating candidate circular layout schemes by solving multi-objective constraints on the constraint retrieval objectives, including: The intention structure model and the set of data features in the genomic data are jointly compiled and verified to validate the constraint target, generating an intermediate constraint representation. Based on the hard constraints in the intermediate constraint representation, the feasibility of the candidate layout space is solved to obtain a feasible candidate set. The feasible candidate set is judged by candidate layout. If the intermediate representation of the constraint is not satisfied, pruning and backtracking are performed until a feasible solution is output. The feasible solutions are scored and iteratively optimized according to the preset multi-objective evaluation index. The optimized feasible solutions are sorted according to the scoring results, and the solution with the best score is selected as a candidate to generate a candidate ring layout scheme.

3. The method for generating a circular genomics visualization workflow based on a large model as described in claim 1, wherein positive and negative examples are retrieved for the candidate circular layout schemes, and constraint patches are applied and the validation constraint target is updated using the retrieved evidence set to obtain an optimized validation constraint target, including: The candidate ring layout scheme and the preset intent features are constructed to obtain the constructed retrieval signature; Based on the constructed retrieval signature, positive example evidence retrieval is performed to obtain positive example retrieval results; Based on the constructed retrieval signature, negative example evidence retrieval is performed to obtain negative example retrieval results; Constraint patches are generated based on the negative example retrieval results. The failure conditions in the negative example retrieval results are transformed through the constraint patches to obtain new hard constraints. The disabled combinations in the negative example search results are transformed into layout space pruning rules. The layout space pruning rules and the newly added hard constraints are then merged to obtain the updated verification constraint target. The updated verification constraint objective is optimized based on the positive example retrieval results. The threshold of the newly added hard constraint is adjusted and interactively combined using the recommended parameters in the positive example retrieval results to obtain the optimized verification constraint objective.

4. A large-model-driven circular genomics visualization workflow generation device, characterized in that, include: The acquisition module is used to acquire multiple genome data. The parsing module is used to parse natural language input into the semantic parsing module of the large language model and generate an intent structure model; The constraint module is used to construct and verify constraint targets for the intent structure model and the features of each of the genome data, and to generate candidate circular layout schemes by solving multi-objective constraints on the constraint retrieval targets. The retrieval module is used to perform positive and negative example retrieval on the candidate ring layout scheme, perform constraint patching and update the verification constraint target through the retrieved evidence set, and obtain the optimized verification constraint target. A construction module is used to build based on the optimization verification constraint target and the constraint patch to obtain visualization instructions; The building module includes: The scoring unit is used to score and rank the candidate ring layout schemes in multiple dimensions based on the optimization verification constraint objective, so as to obtain the preferred layout set. Decomposition subunits are used to decompose the preferred layout set into differentiable and nondifferentiable indices to obtain differentiable and nondifferentiable indices. The evaluation metrics of the preferred layout set are precisely mapped to specific constraints of tracks, channels, and interactions, and a binary decomposition of differentiable and non-differentiable metrics is performed accordingly. Differentiable metrics include geometric conflict overlap rate, spacing compliance rate, visual channel mapping adaptation, and other quantitative parameters that support gradient calculation, while non-differentiable metrics include visual clutter perception, interaction burden assessment, attention area highlighting effect, and other subjective parameters that rely on experience judgment. The computational subunit is used to calculate the layout latent code gradient of the differentiable index in the latent code space according to the differentiable rendering proxy model, and obtain the gradient update direction. The differentiable rendering proxy model is pre-trained and generated using a large number of layout samples and rendering effect data to simulate the real rendering process and output differentiable index values. The latent code space maps the radius of the layout, the orbit order, the channel mapping rules, and other key parameters. The gradient direction of the differentiable index as the latent code changes is calculated using the gradient descent method, so that the latent code adjustment path optimized by the differentiable index is obtained, and the gradient update direction is obtained. The search subunit is used to perform a multi-dimensional continuous layout latent code space parallel search on the non-differentiable index based on the expected reward estimated by zero-order Markov Monte Carlo and same-policy reinforcement learning to obtain the non-differentiable latent code distribution. The zero-order Markov Monte Carlo method is used to solve the problem that the gradient of non-differentiable indicators cannot be directly calculated. By randomly sampling latent code space samples of the non-differentiable indicators and evaluating the performance of the corresponding non-differentiable indicators, the same-policy reinforcement learning is based on the current search policy and transforms the optimization effect of non-differentiable indicators into expected rewards, guiding the search process to focus on high reward regions. The parallel search of the multi-dimensional continuous layout of the latent code space covers orbital parameters, interaction configurations and other multiple dimensions, and mines latent code combinations that meet the optimization requirements of non-differentiable indicators, ultimately forming a non-differentiable latent code distribution. The filtering subunit is used to integrate the gradient update direction and the non-differentiable latent code distribution based on the Pareto dominance criterion and filter out the non-dominated solution set to obtain the Pareto front. An iterative optimization unit is used to iteratively optimize the Pareto front according to the self-adversarial evolutionary mechanism of ring topology preservation, and generate an evolutionary layout population. The instruction compilation unit is used to compile the optimal layout in the evolutionary layout population into consistent instructions, and to perform conflict detection and automatic correction on the compiled instructions through the constraint patch to obtain visual instructions. The verification module is used to verify the visualization instructions according to the executable verification model. When the verification fails, it performs local repair and loops the verification until the verification passes and outputs the circular genome interactive visual analysis model.

5. The large-model-driven circular genomics visualization workflow generation device according to claim 4, wherein the constraint module comprises: The compilation unit is used to jointly compile and verify the constraint target of the intent structure model and the set of data features in the genomic data, and generate an intermediate constraint representation. The solution unit is used to solve the feasibility of the candidate layout space based on the hard constraints in the intermediate constraint representation, and obtain a set of feasible candidates; The pruning unit is used to perform candidate layout judgment on the feasible candidate set. When the constraint intermediate representation is not satisfied, pruning and back-off processing is performed until a feasible solution is output. The iteration unit is used to score and iteratively optimize the feasible solution according to the preset multi-objective evaluation index, sort the optimized feasible solutions according to the scoring results, select the solution with the best score as a candidate, and generate candidate ring layout schemes.

6. The large-model-driven circular genomics visualization workflow generation device according to claim 4, wherein the retrieval module comprises: A construction unit is used to construct the candidate ring layout scheme and the preset intent features to obtain the constructed retrieval signature; The positive example retrieval unit is used to perform positive example evidence retrieval based on the constructed retrieval signature to obtain positive example retrieval results. The negative example retrieval unit is used to perform negative example evidence retrieval based on the constructed retrieval signature and obtain negative example retrieval results. The patching unit is used to generate constraint patches based on the negative example retrieval results, and to transform the failure conditions in the negative example retrieval results through the constraint patches to obtain new hard constraints; The conversion unit is used to convert the layout space pruning rules into the disabled combinations in the negative example retrieval results, and merge the layout space pruning rules with the newly added hard constraints to obtain the updated verification constraint target. The optimization unit is used to optimize the updated verification constraint target based on the positive example retrieval results. It adjusts and interactively combines the threshold of the newly added hard constraint with the recommended parameters in the positive example retrieval results to obtain the optimized verification constraint target.