Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

132 results about "Information density" patented technology

Information density is the amount of human-readable information in a unit of screen real estate such as a square inch. Minimalism. As in other design areas, there is a significant and pervasive tendency in user interface design toward minimalism.

A feature editing method for large model content security

The application discloses a feature editing method for large model content security, which compares and analyzes the sparse coding features of a chat assistant constructed based on a large language model under positive user input and negative user input, extracts the internal response differences of the model to different semantic directions, and the mechanism can automatically and accurately identify the key feature dimensions highly related to the semantic direction of the target attribute. The model activation is mapped to a sparse feature space by using a sparse autoencoder, and each dimension of the feature has independent and interpretable semantic meaning. By injecting a feature guide vector in the space, the interference of the control process on the text grammar, fluency and information density is significantly reduced. The sparse representation mechanism is introduced to structure the intermediate activation features in the reasoning process of the large language model and to intervene in a targeted manner, so that the reply of the chat assistant to the user input conforms to the preset safety specification, and the safety and controllability of the chat assistant in the interaction with the user are improved.
Owner:ZHEJIANG UNIV +1

Public opinion event multi-mode semantic fusion modeling and abstract generation method and system

The invention discloses a public opinion event multi-mode semantic fusion modeling and abstract generation method and system, and relates to the field of natural language processing and social network analysis. Through the multi-mode semantic fusion technology, the short text understanding ability is improved, and the problems of semantic fuzziness and network language diversification are solved. Meanwhile, through a cross-window event cluster matching technology, an event evolution path with time continuity is constructed, and comprehensive capture of event dynamic characteristics is realized. Besides, the structured event abstract is automatically generated by utilizing the generative model, so that the consistency and the information density of the abstract are improved, and the actual application requirements are met. Through the innovations, the defects in the aspects of semantic comprehension, dynamic modeling and abstract generation in the prior art can be effectively overcome, a more efficient and accurate solution is provided for monitoring and analysis of public opinion events, and the method has wide application prospects in the fields of public opinion monitoring, emergency early warning, social media data analysis and the like.
Owner:NORTHEASTERN UNIV CHINA

Voice segmentation intelligent editing system based on deep learning

PendingCN121260170ASpeech recognitionSpeech segmentationInformation density
The invention relates to the technical field of voice signal processing, and discloses a voice segmentation intelligent editing system based on deep learning. The system comprises a voice feature extraction module, a segmentation boundary detection module, a semantic content analysis module, an editing strategy generation module and a real-time quality evaluation module. The voice feature extraction module collects multi-dimensional voice features and timestamp information, and verifies feature integrity and timeliness; the segmentation boundary detection module identifies voice pause intervals and semantic turning nodes and divides segmentation units and boundary types; a semantic content analysis module extracts text content and emotion features of each segment, and analyzes semantic topic relevance and information density; an editing strategy generation module formulates a segmentation retention rule and a sequence adjustment scheme, and matches user preferences and scene demands; the real-time quality evaluation module monitors voice fluency and information integrity in the editing process and analyzes splicing errors and user feedback. According to the system, intelligent processing of the whole voice editing process is realized.
Owner:SHENZHEN JYEOO NETWORK TECH CO LTD

Efficient annotation-driven hierarchical fault positioning method

The invention provides an efficient annotation-driven hierarchical fault positioning method, which comprises the following steps of: guiding a large language model to intelligently analyze and generate semantic annotations of a text based on a positioning algorithm of large model annotations, and performing fine-grained software fault positioning from a file level to a function level and then to a position level by using the annotation-based hierarchical positioning method. And an efficient hierarchical progressive sorting and screening algorithm is used for ensuring the software fault positioning efficiency. According to the method, the defect positioning efficiency can be ensured while the positioning accuracy is ensured. The positioning algorithm based on large model annotation strategically balances the information density, provides enough context clues, and improves the model understanding ability; according to the annotation-based hierarchical positioning method, positioning is divided into file / function / position hierarchies, and annotations are added by using a large model cue word project in sequence, so that the positioning accuracy is improved; and an efficient hierarchical progressive sorting and screening algorithm is provided, so that the optimal balance between the performance and the cost is realized.
Owner:NANJING UNIV

Feature editing method for large model content security

The invention discloses a large model content security-oriented feature editing method, which comprises the following steps of: comparing and analyzing sparse coding features of a chat assistant constructed on the basis of a large language model under positive user input and negative user input, and extracting internal response differences of the model in different semantic directions; the mechanism can automatically and accurately identify key feature dimensions highly related to the semantic direction of the target attribute. A sparse auto-encoder is utilized to activate and map the model to a sparse feature space, and each dimension of feature has an independent and interpretable semantic meaning. By injecting a feature guide vector into the space, the interference of a control process on text grammar, fluency and information density is remarkably reduced. A sparse representation mechanism is introduced, structural modeling and targeted intervention are carried out on intermediate activation features in a big language model reasoning process, replies input by a chat assistant to a user are guided to conform to a preset safety specification, and the safety and controllability of the chat assistant in interaction with the user are improved.
Owner:ZHEJIANG UNIV +1

Processing and transmitting active regions of display for improved performance

A system is disclosed, including a display, a processor and a memory. The memory stores instructions that, when executed by the processor, configure the system to perform operations. Active region data is generated that includes, for each of one or more active regions, active region location data and active region content. The active region data is transmitted to a display having a display area. For each active region, the active region content is displayed at an active region location of the display area based on the active region location data, the active region content being displayed at a higher spatiotemporal information density than content displayed in the display area outside of the active regions.
Owner:SNAP INC

Forage grass yield prediction model construction method based on deep learning

The invention relates to the technical field of deep learning, in particular to a forage grass yield prediction model construction method based on deep learning, which comprises the steps of constructing a multi-source data fusion module, establishing a cold start mechanism, designing a multi-modal deep learning prediction network, integrating a physical constraint mechanism and constructing a management decision support system. A seasonal attribution analysis function is realized; in the prior art, a simple data superposition or static weighted fusion scheme is generally adopted, and inherent defects of deficiency, different scales and heterogeneity of multi-source data are difficult to process, so that the fusion feature quality is poor; according to the method, firstly, a data blank is accurately filled through an intelligent algorithm based on space-time continuity, then heterogeneous data is unified to a standard grid by using a multi-scale pyramid engine, and finally, deep fusion is performed through an attention mechanism for dynamically calculating importance of each data source; the integrity, the consistency and the information density of the input data are remarkably improved, and a solid and reliable data foundation is laid for subsequent accurate prediction.
Owner:Garze Tibetan Autonomous Prefecture Animal Husbandry Science Research Institute (Garze Tibetan Autonomous Prefecture Yak Industry Development Center)

RAG content generation method and system based on dynamic slicing and adaptive fusion

The invention discloses an RAG content generation method and system based on dynamic slicing and self-adaptive fusion, and relates to the field of natural language process.The method comprises the steps that in response to a user query instruction, semantic fusion processing is conducted on a plurality of related text blocks, and a preliminary text sequence is traversed; detecting a logic breakpoint between adjacent text blocks through a context association identifier, recalling a middle connection text block corresponding to the logic breakpoint, and inserting the middle connection text block into a sequence to generate a fusion text stream; performing adaptive information density adjustment on the fused text stream to obtain optimized context content; generating structured prompt information based on the optimized context content and the user query instruction; and calling the target large language model to generate a user query result based on the structured prompt information. The answer accuracy and logic continuity of the target large language model for user query are effectively improved, and the utilization efficiency of model computing resources is optimized.
Owner:HANGZHOU WEIMING XINKE TECH CO LTD +1

Intelligent operation and maintenance method and system based on large language model

The embodiment of the invention provides an intelligent operation and maintenance method and system based on a large language model. According to the method, firstly, at least M multi-modal original logs generated by a micro-service cluster are collected, compression processing is conducted on the multi-modal original logs by means of a large language model, at least N multi-modal target logs are obtained, the log data volume can be greatly reduced, and the effective information density is improved; synchronously collecting system operation indexes of timestamps corresponding to the multi-modal target logs and determining an association relationship between the system operation indexes to reduce invalid diagnosis workload; a multi-type heterogeneous operation and maintenance topological graph is constructed based on the incidence relation, a current fault is accurately positioned, the problem that root causes are difficult to quickly position through traditional manual analysis of mesh topology is solved, and the root cause recognition time is shortened; and finally, a preset historical fault knowledge base is called, a target repairing script matched with the current fault is generated in combination with the large language model, repairing of the current fault is achieved, time consumed by manual intervention is shortened, and therefore the fault diagnosis efficiency is remarkably improved.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

File data content accurate and deep analysis and interpretation method based on AI

The invention belongs to the technical field of artificial intelligence, and particularly relates to an AI-based file data content accurate and deep analysis and interpretation method, which comprises the following steps: acquiring multi-format file data and file meta-information, constructing an AI analysis network, extracting text semantic vectors, image visual features, table structure information and document layout features, and constructing a multi-dimensional semantic map. According to a user query intention, semantic extension is performed in combination with a domain knowledge base, an enhanced semantic description vector is generated through a graph attention mechanism, semantic reasoning and relation mining are performed by adopting an improved knowledge distillation Transform model, and a deep analysis conclusion is generated through a multi-hop reasoning model in combination with file data complexity, information density and user requirements. And generating a personalized interpretation report in combination with a user role and a task scene, and outputting an analysis result through a visual interface. Therefore, the problems of poor understanding ability, poor file adaptability and the like in the prior art are solved.
Owner:WUHAN CHANGYUAN HONGTIAN DATA INFORMATION TECHNOLOGY CO LTD

Multi-modal large model question and answer method, device and equipment based on attention entropy and medium

The invention discloses a multi-modal large model question and answer method, device and equipment based on attention entropy and a medium, and relates to the field of artificial intelligence and machine learning. Comprising the following steps: inputting a target image and a problem text into a multi-modal large model to obtain an image marking sequence and a text lexical element sequence; connecting an rth pruning layer in front of an rth decoding layer of the decoder, calculating an attention matrix from rth vision to text and an attention matrix from rth text to vision according to an rth image mark sequence and an rth text lexical element sequence, and determining an ith information density weight corresponding to an ith image mark in the rth image mark sequence; based on the information density weight corresponding to each image mark in the rth image mark sequence, screening out an rth reserved image mark sequence; and inputting the rth reserved image mark sequence and the rth text lexical element sequence into the rth decoding layer to obtain an answer text output by the large language model, so as to effectively balance the calculation efficiency and the semantic integrity under the condition of not needing additional training.
Owner:TSINGHUA UNIVERSITY

Recursive deep text query method, system and equipment and storage medium

The invention discloses a recursive deep text query method, system and device and a storage medium, and the method comprises the steps: firstly creating an initial query and initializing a hierarchical depth, executing a first-layer query, carrying out evolution processing, and then carrying out retrieval; if the information density is insufficient, sub-query is generated to execute second-layer query, and retrieval is performed after second evolution processing; if the second-layer retrieval is contradictory, third-layer query is executed, retrieval is performed after timeline query and information source tracing verification are passed, and a third-layer retrieval result is finally output; through multi-level query and targeted evolution processing, query can be optimized step by step, more information dimensions can be covered, information gaps can be filled up, and retrieval comprehensiveness and accuracy can be improved; and a third-layer verification mechanism can effectively check contradictions by means of timelines and source tracing, so that the accuracy and reliability of results are guaranteed, and better retrieval results are provided for users.
Owner:ZHILU CLOUD (SHENZHEN) ARTIFICIAL INTELLIGENCE CO LTD

A method for serialization extraction of highly variable exons

PendingCN122290698AInformation densityExon
This invention discloses an efficient RNA data preprocessing method to address the problems of low processing efficiency and low information density in high-throughput sequencing data. Its core steps include: (1) introducing a parallel processing scheme for high-throughput sequence data, rapidly mapping RNA-seq data to a reference genome to generate a BAM file; (2) extracting base sequences and expression levels and storing them as compact PKL format files; (3) extracting all exon position information by parsing the genome annotation file; (4) combining multi-sample expression level data to screen for highly variable exons and constructing a high-information-density feature list based on the sample set; and (5) accurately extracting target sequences from the preprocessed file based on this list. Compared to traditional methods, this innovative approach achieves triple optimization: full-process parallel processing for accelerated computation, high-compression data storage, and adaptive feature selection. Processing speed is increased by 3-5 times, and data volume is reduced by more than 90%, making it suitable for high-throughput RNA-seq data analysis with large sample sizes.
Owner:TIANJIN UNIV

A code library UML diagram automatic generation method and system based on abstract syntax tree and large language model cooperation

The application discloses a code library UML diagram automatic generation method and system based on cooperation of abstract syntax tree and large language model, comprising an AST analysis and structure extraction layer, which is used for being responsible for abstract syntax tree analysis on source files in the code library and extracting skeleton structure information of the code library, and a large language model semantic analysis layer, which is used for receiving structured code skeleton data output by the AST analysis layer and submitting the structured data to the large language model for semantic analysis through a carefully designed prompt word (Prompt), wherein the application extracts structured skeleton information of the code library through AST analysis instead of directly processing original source code, compresses the information volume to 10% to 20% of the original code amount, greatly improves the effective information density of the input large language model, enables a larger scale code library to be covered in a limited context window, and solves the problem that the existing large language model scheme cannot process a large scale code library due to Token waste.
Owner:BEIJING SIMPLE POINT TECH CO LTD

Industrial multi-modal data extraction method and system based on large language model

The invention discloses an industrial multi-modal data extraction method and system based on a large language model, and relates to the technical field of large language models.The industrial multi-modal data extraction method comprises the steps that industrial multi-modal data are collected, high-dimensional features are extracted through a convolutional neural network, and unified feature vectors are formed through splicing after normalization; determining a first retention ratio search range by using the PCA accumulated variance and the reconstruction error elbow point, and generating an initial retention ratio in combination with the semantic information density; taking the proportion as a center, generating a plurality of candidate retention proportions by adopting Gaussian perturbation, performing semantic compression on each proportion, extracting multi-dimensional features, and calculating a compression score; and mapping the proportion and the score into a screening vector, forming a continuous scoring surface through two-dimensional mapping and interpolation, analyzing a scoring extreme value, screening out an optimal retention proportion, and finally applying the optimal retention proportion to industrial semantic compression. The problems that in existing industrial multi-modal data semantic compression, the retention ratio is too high, so that real-time performance is reduced, and the retention ratio is too low, so that semantics are lost are solved.
Owner:SUZHOU ZHIYOU QIUSUO INTELLIGENT TECHNOLOGY CO LTD

Power grid analysis program code migration method and system based on large model

The invention relates to a power grid analysis program code migration method and system based on a large model, and belongs to the technical field of software development and intellectualization. The method comprises the steps that preprocessing of structuralization, semantization and information density optimization is conducted on MATLAB source codes, and standardized codes are generated; disambiguation anchor point cue words are constructed based on a power grid calculation scene, wherein the disambiguation anchor point cue words comprise index, operation and return value mapping rules and examples; the standardized codes and the cue words are input into a large model, four-layer semantic alignment migration is conducted by the large model according to the sequence of an algorithm layer, a data layer, a control layer and a calculation layer, verification logic is executed after migration of each layer to ensure semantic alignment, and Python codes are obtained; and finally, verifying the migration code and iteratively optimizing the cue word until the core index is consistent with the MATLAB source code. According to the method, the problems of semantic ambiguity and nesting overload existing when a large model migrates a power grid program are solved through the integrated design of preprocessing, migration and verification, and high-precision and automatic code migration is achieved.
Owner:FUJIAN AGRI & FORESTRY UNIV

A surface defect detection method based on feature pre-fusion and mask guidance

The application discloses a surface defect detection method based on feature pre-fusion and mask guidance, and mainly solves the problem of low detection accuracy of weak distinguishability and small scale defects in the prior art. The implementation scheme is as follows: 1) obtaining a data set and a detection label; 2) constructing a defect detection model; 3) constructing a loss function; 4) training the defect detection model; and 5) reasoning and obtaining a detection result. The surface defect detection model constructed by the application realizes the expansion of the receptive field and the information diffusion through the feature pre-fusion, enhances the context within the feature map, and effectively improves the detection accuracy of weak distinguishability defects. The mask label of the defect boundary box is introduced through multi-stage feature fusion, so that the information density of the defect area is increased, and the detection accuracy of small scale defects is improved.
Owner:CENT SOUTH UNIV

Video platform intelligent content recommendation method based on deep learning

The invention discloses a video platform intelligent content recommendation method based on deep learning, and relates to the technical field of intelligent recommendation, and the method comprises the following specific steps: data collection, collecting user historical watching data including video attribute data, user behavior data and scene feature data to form a basic data set; video information density calculation: based on the collected video attribute data, extracting visual, audio and text multi-dimensional features by using a multi-modal deep learning model, constructing a video information density evaluation system, and converting video contents into quantifiable information density values; according to the method, the user cognitive load tolerance model is constructed by analyzing the video information density and the user behavior data, the acceptance capability of the user to contents with different information densities can be accurately grasped, the current cognitive load state of the user is dynamically judged in combination with the real-time scene information of the user, and the recommended content is accurately matched with the cognitive ability of the user.
Owner:SHENZHEN JUWANGSHI TECH CO LTD

Enhanced training method for improving analysis capability of middle text of large language model

The invention discloses a reinforced training method for improving the middle-section text analysis capability of a large language model, and relates to the field of natural language processing and generative large language models.The method comprises the steps that middle-section double-window construction is conducted according to a dynamic scale factor, extraction and splicing are conducted from original long text corpora based on window positions, and the middle-section text corpora are obtained; obtaining a preliminary training sample; performing middle text information density evaluation on the preliminary training sample, and verifying and optimizing the preliminary training sample to obtain an optimized sample; and inputting the optimized sample and the updated position identification sequence into a large language model, and realizing iterative training of the large language model through a position-aware middle loss weighting mechanism to obtain an optimized large language model. According to the method, the retrieval and understanding accuracy of the model in the process of processing the middle position information of the sequence can be remarkably improved in a balanced manner, the U-shaped performance curve which is sunken originally is successfully leveled, and the training performance bottleneck is overcome.
Owner:UNIV OF SCI & TECH OF CHINA

Redundancy elimination method and device for voice transfer text, and electronic equipment

The invention discloses a voice transcription text redundancy elimination method and device and electronic equipment, and belongs to the technical field of text processing. The method comprises the following steps: performing redundancy elimination processing on a to-be-processed voice transcription text based on a preset text redundancy processing rule to obtain a first redundancy elimination processing text; performing redundancy elimination processing on the first redundancy elimination processing text by adopting a part-of-speech and position double-weighted statistical language model method to obtain candidate redundancy words in the first redundancy elimination processing text; by calculating fluency gains of sentences in the first redundancy removal processing text before and after deleting each candidate redundant word, screening to obtain redundant words; and performing redundancy elimination processing on the first redundancy elimination processing text according to the screened redundant words. According to the method, the language information density is effectively improved, so that the quality of the voice transliteration text after redundancy elimination processing is improved.
Owner:HANVON CORP

Recommended medical content methods, procedures, products, devices, equipment, and storage media

This disclosure provides a method, program product, apparatus, device, and storage medium for recommending medical content, comprising: obtaining a first set of medical content, including at least one piece of medical content describing medical viewpoints, wherein the at least one piece of medical content has a content information density score and a description depth rating; obtaining historical data of users, including a rating of the user's level of understanding of the corresponding medical viewpoints; determining a rating score for at least one piece of medical content based on the level of understanding rating and the description depth rating; determining a score for the user's associated interactive behavior towards at least one piece of medical content; determining a first weight for the rating score and a second weight for the associated interactive behavior score based on the information density score; obtaining a ranking score for at least one piece of medical content; and presenting at least one piece of medical content to the user.
Owner:ASTRAZENECA PHARMACEUTICAL (CHINA) CO LTD

Methods for displaying information on a human-machine interface of a motor vehicle, computer program products, human-machine interfaces, and motor vehicles

A method is described for a human-machine interface (4) for controlling a motor vehicle (2), wherein the human-machine interface (4) has at least one display device (8, 10, 12) for displaying information (34, 36, 38, 40, 42, 44), wherein at least one driving condition parameter is obtained, the driving condition parameter characterizing the driving condition of the motor vehicle (2), and wherein a load parameter characterizing the cognitive load of the driver (6) of the motor vehicle (2) is determined by means of an algorithm (20) based on the obtained driving condition parameter, thereby adapting the information density of the information to be displayed by means of the at least one display device (8, 10, 12) according to the load parameter, wherein the algorithm (20) includes a dataset having driving condition parameters, wherein a load parameter is assigned to each driving condition parameter, wherein the load parameter is based on at least one measured measurement parameter characterizing the driver's behavior. A computer program product, a human-machine interface, and a motor vehicle are also described.
Owner:PEUGEOT CITROEN AUTOMOBILES SA

Natural language query analysis and database field matching method based on large model

The embodiment of the specification provides a natural language query analysis and database field matching method based on a large model, which comprises the following steps: performing intent analysis and entity recognition on a user natural language query based on a large language model (LLM) to extract a potential keyword set; for each extracted keyword, binding an entity type through an STAM mechanism and constructing an entity-type mapping relationship; constructing a column description vector library, combining an approximate nearest neighbor (ANN) algorithm to realize efficient retrieval, and combining a multi-strategy fusion matching mechanism to combine the results of semantic vector matching, edit distance matching and traceability enhanced matching to form a final candidate column set; connecting a database metadata interface, analyzing information and constructing a basic structure to generate an enhanced semantic field description, and outputting in a four-tuple structure; reversing the corresponding table through column matching, judging the table structure value in combination with the information density in the table and the connectivity with other tables, eliminating redundant tables, and optimizing the database representation.
Owner:数字郑州科技有限公司

A live room information display method, system, device and medium

The application relates to the technical field of live content display and interaction optimization, and discloses a live room information display method, system, device and medium, which comprises the following steps: collecting audience data and analyzing audience viewing behavior in real time; automatically matching viewing modes according to the audience viewing behavior and displaying different information levels for different viewing modes; analyzing the audience's barrage and automatically screening the barrage priority; and dynamically displaying live content according to the adjusted information level and barrage priority. Through intelligent dynamic adjustment of the display level of live content and the barrage priority, more personalized and efficient information display is realized. According to the viewing behavior, interaction frequency and device network status of the audience, the system automatically adjusts the information density and display level, avoids information overload or loss, and improves the user experience. Meanwhile, based on the emotional analysis of the barrage, the interaction frequency and the content relevance, high-value barrages are preferentially displayed, and the audience's interaction sense and participation are enhanced.
Owner:NANTONG CHENGXIN INFORMATION TECHNOLOGY CO LTD

An eye array camera-based target detection neural network accelerator

The application belongs to the technical field of neural network acceleration control, and discloses a target detection neural network accelerator based on an array camera, which comprises a multi-view image acquisition module, a data preprocessing module, a heterogeneous computing module, a target fusion module and a trajectory generation module. The computing unit of the neural network accelerator is divided into two types of high-precision and low-power consumption. The multi-view images collected by the array camera are evaluated for target value, and high-value images are dynamically allocated to high-precision units, and low-value images are dynamically allocated to low-power consumption units, so that the adaptive scheduling of computing resources is realized, and the excessive calculation of low information density areas can be effectively avoided. At the same time, after the multi-view images are processed heterogeneously, the same target classification is realized, the low-power consumption unit is used to output the detection result as the spatial constraint of the subsequent output of the high-precision unit, cross-path target fusion is realized, and the pipeline blockage and the increase of end-to-end delay are minimized.
Owner:SHANGHAI UNIV

Vehicle-mounted multimedia information fusion display method

PendingCN121375822ADriver/operatorIn vehicle
The invention discloses a vehicle-mounted multimedia information fusion display method, particularly relates to the technical field of multimedia information fusion display, and effectively inhibits interference of single-mode noise and perception delay on cognitive load evaluation through a fusion analysis method of multi-mode time reference alignment and credibility weighting. The real-time performance and accuracy of driver state recognition are remarkably improved, a layout reconstruction control function is dynamically triggered based on the cognitive load level, interface information density and priority can be adjusted in a self-adaptive mode under different driving situations, balance control over information simplification and supplementation is achieved, therefore, the driving cognitive burden is relieved, operation smoothness is maintained, and meanwhile, the driving experience is improved. Through continuous monitoring of interface adjustment frequency and fixation point change rate, layout reconstruction oscillation intensity is quantified, a cognitive feedback loop instability risk model is constructed in combination with cognitive load change, and a potential cyclic oscillation trend is identified and intervened in advance.
Owner:深圳市鼎微科技有限公司

Intelligent agent for sending and receiving official documents based on artificial intelligence and method for generating dataset

The application discloses a document receiving and sending intelligent agent based on artificial intelligence and a data set generation method, and particularly relates to the field of artificial intelligence, and relates to intelligent decision-making and trainable data construction of document receiving, distribution, handling, batch handling, circulation, archiving and other businesses. The method models the evidence pointing structure, establishes a stable association between the handling conclusion and the evidence fragments in the text and the attachments, forms a traceable instruction, evidence and action unified sample structure, filters high information density samples through event graph construction and multi-view consistency anomaly analysis, and synthesizes controllable difficult cases in the form of counterfactual disturbance to improve the coverage of complex situations. In the reasoning stage, the process constraint execution mechanism is fused, the candidate actions are filtered and the rejection reason record is output, the model output is ensured to be consistent with the executable action set of the business, and the audit playback and closed-loop incremental learning are supported.
Owner:JINAN ZHONGGUAN INCUBATOR TECH CO LTD

Text branch-based ai conversation structure graph and context dynamic assembly method

The application provides a text branch-based AI conversation structure graph and a context dynamic assembly method. The method allows the user to open a new conversation unit as a text branch by designating the selected content as a context anchor from any conversation unit, thereby dynamically organizing the conversation into a visual structure graph with each unit as a node. The context dynamic assembly algorithm is based on the context anchor and the graph, intelligently evaluating the relevance of each historical conversation and the current input from the dimensions of semantic association, structural proximity, logical context, etc., and calculating the comprehensive weight and sorting for the most relevant content. Finally, the algorithm intelligently balances the fidelity and abstractness of the information at each level according to the sorting and the budget of the smallest semantic unit (Token), constructs a context load with optimal information density and dynamic fidelity, and provides it to the AI, significantly improving its understanding and response ability to the user's real intention in complex multi-threaded conversations.
Owner:李仲毅

Warehouse-level long code-oriented high-density completion input construction and completion method

The invention discloses a high-density completion input construction and completion method for warehouse-level long codes. The method comprises the following steps: firstly, acquiring a current code snippet near a to-be-complemented position and a background long code context, performing structured analysis on the context, dividing candidate code units, and determining a code unit containing the to-be-complemented position as a query code unit; performing multi-channel correlation retrieval on the candidate code units, screening to obtain candidate subsets, performing function-level or class-level importance evaluation and reordering on the candidate subsets, performing self-adaptive truncation according to an attenuation relationship of adjacent candidate importance gains, further generating a fine-grained semantic unit sequence, dividing functional blocks, and obtaining a fine-grained semantic unit sequence; according to the method, importance evaluation is carried out on the function blocks, the function blocks are selected or cut under budget constraint, and the function blocks and current code snippets are jointly used as completion input, so that information density and task correlation are improved, reasoning cost and response delay are reduced, and stability and availability of warehouse-level long code completion are improved.
Owner:HANGZHOU DIANZI UNIV

Knowledge blind area perception distillation method and system for target language model and readable storage medium

PendingCN122509308AReduce deployment complexityInjection is precise and efficientInformation densityData mining
This application relates to a knowledge blind spot perception distillation method, system, and readable storage medium for target language models. The method includes: segmenting domain documents into semantic fragments; generating a candidate question-answer pair set containing knowledge category labels using a first language model; obtaining the target language model's answers to each question without context; calculating a knowledge mastery score through fact-level comparison; filtering out an unknown question-answer pair set that the model does not yet possess; performing hash bucketing clustering on the unknown question-answer pair set according to category labels to obtain multiple topic clusters; and allocating character sub-budgets to each topic cluster based on the proportion of question-answer pairs; generating structured text based on the sub-budgets and unknown question-answer pairs; and finally assembling the text into distilled prompt words for use by the target model. This application effectively filters out knowledge already possessed by the model, increases the information density of the prompt words, and reduces the computational and storage overhead of edge deployment.
Owner:HANGZHOU TANYUAN INNOVATION CULTURE TECHNOLOGY CO LTD