Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

358 results about "Data synthesis" patented technology

Data synthesis. meta analysis A method that uses statistical techniques to combine results from different studies and obtain a quantitative estimate of the overall effect of a particular intervention or variable on a defined outcome—i.e., it is a statistical process for pooling data from many clinical trials to glean a clear answer.

Training data synthesis method and device based on error extrapolation and inference chain analysis, medium and program product

The invention provides a training data synthesis method and device based on error extrapolation and inference chain analysis, a medium and a program product. The method comprises the steps of obtaining an initial sample set; performing multiple sampling reasoning on the problem of each task sample by using a small language model to generate a plurality of reasoning chains; calculating an overall error score of each reasoning chain based on a preset error evaluation rule, and determining a to-be-corrected reasoning chain; the inference chain to be corrected and the corresponding question are input into the large language model together, and a corrected answer is generated; forming a new task sample by the question and the corrected answer, and finely adjusting partial parameters of the small language model; repeatedly executing the process until the performance index change rate of the model on the task evaluation set is lower than a preset threshold value, and outputting a final task sample; and forming a training sample set by a plurality of final task samples, and performing all-parameter fine tuning on the small language model. According to the method, the training data self-optimization path is constructed by taking the model error as guidance, so that the semantic consistency and the data validity are improved.
Owner:SHANGHAI COOPERS TECHNOLOGY CO LTD

Training of multi-modality object detectors

Techniques for determining a presence of an object, especially an object such as animal or debris, in a path of a vehicle, are discussed herein. For example, sensors of various modalities, which may include multispectral sensors, may capture data representing an environment the vehicle is traversing. In examples, one or more trained machine learned (ML) models, operating on a vehicle computing system, may detect and / or classify objects in the environment, based on input data of one or more modalities or spectral bands. The ML models may be pre-trained using training data including real sensor data, synthetic data, and / or augmented data, along with auto-generated annotations. In some examples, hyperspectral data may be used to identify materials associated with detected objects. A confidence score associated with the detection of the object may also be computed. The vehicle may be controlled based on detection of the object and its classification.
Owner:ZOOX INC

Three-dimensional data synthesis method, electronic equipment, storage medium and program product

The invention provides a three-dimensional data synthesis method, electronic equipment, a storage medium and a program product, and the method comprises the steps: obtaining a plurality of texture-free three-dimensional grid models, carrying out the rendering of each three-dimensional grid model, obtaining a plurality of view images, and constructing a training sample pair in combination with a semantic description text; based on the training sample pair, utilizing a low-rank adaptation technology to carry out fine tuning on the first text graph model to obtain a second text graph model with multi-view consistency understanding ability; modeling a latent space semantic difference between the rendered image of the to-be-deformed three-dimensional grid model and the target semantic text based on the second text graph model, constructing an optimization constraint and updating parameters, and obtaining a deformed three-dimensional grid model conforming to target semantics; and generating a texture image by using the texture generation model, and attaching the texture image to the surface of the deformed three-dimensional grid model to obtain synthetic three-dimensional data with textures. The reconstruction precision, the diversity expression ability and the cross-category generalization ability of the synthesized three-dimensional data are obviously enhanced.
Owner:SHANG HAI JIE YUE XING CHEN ZHI NENG KE JI YOU XIAN GONG SI

Radio interference identification method based on electromagnetic spectrum monitoring

The invention discloses a radio interference identification method based on electromagnetic spectrum monitoring, and the method comprises the following steps: S1, collecting spectrum data in an electromagnetic environment through broadband spectrum monitoring equipment, and carrying out the digital sampling processing; s2, adopting a diffusion probability model to automatically remove environmental noise and non-interference signals; s3, automatically generating an expansion data set of edge and small sample interference signals by adopting an attention-enhanced data synthesis expansion technology; s4, extracting multi-dimensional time-frequency domain dynamic correlation characteristics of the spectrum interference signal based on a Transform structure; s5, adopting an attention enhancement self-supervision algorithm to generate a self-supervision soft label of an unknown interference type; s6, dynamic back diffusion iteration is carried out through the diffusion probability model, and an interference classification result is obtained; and S7, constructing a closed-loop dynamic feedback mechanism by using an interference classification result. According to the invention, accurate identification of radio interference signals is realized, and the accuracy, dynamic adaptability and automation degree of interference signal classification are improved.
Owner:WUHAN HAIHUA XINTONG TECH CO LTD

Anti-fact fair synthesis data generation method and device based on causal reasoning

The invention provides an anti-fact fair data synthesis method and device based on causal reasoning, and aims to generate high-quality synthesis data meeting the fairness requirement by mining the causal relationship between observable features. The synthesis method comprises the following steps: extracting observable features, sensitive features and labels from original data, extracting potential features through a variational automatic codec, and constructing a causal relationship graph; designing a generator according to a topological sequence of the causal relationship graph, connecting a causal path, inputting the potential features and the related features into the generator in sequence, and constructing a data generation process conforming to a causal structure; introducing a discriminator to carry out adversarial training on a generation result and original data, and optimizing generator parameter distribution; finally, synthetic data meeting fairness requirements are generated. According to the method, effective regulation and control on the influence of sensitive characteristics and strict constraint on a causal structure are realized, the generated data has higher fairness and interpretability, and the method can be applied to the fields with higher fairness requirements, such as finance, medical treatment and education.
Owner:JINAN UNIVERSITY

Liquid flash TDCR multi-nuclide beta spectrum analysis method and system based on artificial intelligence

The invention belongs to the technical field of nuclear radiation measurement, and relates to a liquid flash TDCR multi-nuclide beta spectrum analysis method and system based on artificial intelligence. The method comprises the following steps: constructing a numerical model of a liquid flash detector spectrometer by using a Monte Carlo technology to simulate the energy spectrum response of a single nuclide under different quenching conditions; constructing a training database by adopting a parameterized data synthesis algorithm; constructing a multi-task mixed spectrum analysis neural network model; carrying out model training by utilizing the constructed database; and inputting an actual measurement spectrogram, and outputting the multi-nuclide absolute activity, the detector efficiency and the decomposition energy spectrum contribution curve in real time by using the trained neural network model. According to the method, differential distribution characteristics and quenching response curve characteristics of the beta continuous energy spectrum are creatively fused, a neural network architecture with physical mechanism constraints is constructed, and liquid flash TDCR multi-nuclide energy spectrum characteristic decoupling and accurate activity solving are achieved under the unknown quenching condition.
Owner:SHANDONG UNIV

Object detection using multispectral data

Techniques for determining a presence of an object, especially an object such as animal or debris, in a path of a vehicle, are discussed herein. For example, sensors of various modalities, which may include multispectral sensors, may capture data representing an environment the vehicle is traversing. In examples, one or more trained machine learned (ML) models, operating on a vehicle computing system, may detect and / or classify objects in the environment, based on input data of one or more modalities or spectral bands. The ML models may be pre-trained using training data including real sensor data, synthetic data, and / or augmented data, along with auto-generated annotations. In some examples, hyperspectral data may be used to identify materials associated with detected objects. A confidence score associated with the detection of the object may also be computed. The vehicle may be controlled based on detection of the object and its classification.
Owner:ZOOX INC

Single crystal turbine blade crystal orientation defect detection system based on adaptive optics

The invention discloses a single crystal turbine blade crystal orientation defect detection system based on adaptive optics, particularly relates to the field of nondestructive testing, and is used for solving the problem that thermal damage and signal consistency are difficult to consider in acoustic imaging of complex structural parts. Through partition energy threshold setting, dynamic pulse planning, photoacoustic processing and unified coordinate cross-subarea data synthesis, local thermal damage caused by excessive concentration of energy is effectively avoided in the material detection process, and the requirement for collecting high-bandwidth signals can be met in a large range. The signal-to-noise ratio and the detection efficiency are improved by regulating and controlling pulse energy and a scanning path, the amplitude comparability and the image coherence among regions are ensured through signal normalization correction and subregion image set splicing, high-resolution and comprehensive-coverage acoustic mapping is output, the crystal orientation defects of the single crystal turbine blade are positioned and evaluated, the detection precision is improved, the measurement period is shortened, and the method is suitable for large-scale popularization and application. And the method has an applicable value for aero-engine parts requiring rapid and high-precision nondestructive detection.
Owner:ANHUI GONGYUAN ELECTRONIC TECHNOLOGY CO LTD

Method for synthesizing electronic medical record data based on semantic processing

The invention discloses an electronic medical record data synthesis method based on semantic processing, and relates to the technical field of medical informatization, and the method comprises the following steps: S1, constructing a probabilistic medical knowledge graph; s2, generating a semantic representation vector; s3, constructing a multi-dimensional dynamic space-time atlas; s4, generating a discrete personalized disease course event sequence with space-time coordinates; s5, taking the discrete personalized disease course event sequence and the corresponding medical entity semantic representation vector as condition input, guiding the improved TSDiff model to execute an iterative denoising process, and outputting a multi-dimensional random disease course trajectory; s6, forming multi-modal electronic medical record data; and S7, performing multi-dimensional quality evaluation on the multi-modal electronic medical record data. According to the method, the limitations of logic inconsistency, modal splitting and model capability solidification in a traditional synthesis method are overcome, and an efficient and accurate solution is provided.
Owner:BEIJING INTELLIGENT DECISION MEDICAL TECH CO LTD

Training data generation method and device of network attack recognition model, and electronic equipment

The invention discloses a training data generation method and device for a network attack recognition model and electronic equipment, and relates to the field of network security, and the method comprises the steps: collecting to-be-recognized emails from a plurality of information sources, carrying out the classification processing and labeling processing of the preprocessed to-be-recognized emails, obtaining a labeled email sample set, and storing the labeled email sample set in a database; extracting content features, sender features, structural features and attachment features to obtain a sample feature set, performing data synthesis by using a generative adversarial network model, generating a synthesized mail sample, converting the labeled mail sample and the synthesized mail sample into a preset model data format, generating a model training data set, training a preset large language model, and obtaining a new mail sample; and identifying whether the contact emails are phishing emails or not through the model. According to the method and the device, the technical problems that omissions of mail filtering are easily caused and the user satisfaction is influenced because a model for identifying phishing mails cannot update training data for identifying an attack mode in time in related technologies are solved.
Owner:TRAVELSKY TECHNOLOGY LIMITED

Risk zoning method, device and equipment for helicopter flight and storage medium

The embodiment of the invention provides a risk zoning method and device for helicopter flight, equipment and a storage medium. The method is applied to the technical field of data processing, and comprises the following steps: carrying out classification and threshold division on meteorological elements according to a fusion data set and different flight scenes to obtain a meteorological element threshold breakpoint set; based on the meteorological element threshold breakpoint set, performing frequency statistics and data synthesis on the fused data set to obtain a multi-layer meteorological element frequency data set; obtaining an initial clustering result of each meteorological element through clustering analysis according to the meteorological element frequency data set; according to the initial clustering result, carrying out risk grade division on each meteorological element to obtain a standardized single-element risk zoning result; and based on a single-element risk zoning result, adopting a multi-element fusion method to obtain a comprehensive risk zoning map. In this way, the technical problem that in the prior art, due to the fact that multiple elements jointly participate in clustering, the risk level is difficult to judge can be solved.
Owner:CHINESE PEOPLES LIBERATION ARMY AVIATION COLLEGE

Asset management coding model modeling method based on federated learning

The invention relates to the field, and particularly discloses an asset management coding model modeling method based on federated learning, and the method comprises the steps of heterogeneous asset coding alignment, a dynamic federated aggregation mechanism, small sample data source optimization and dynamic weight adjustment. Through a dynamic federation aggregation mechanism, resources are efficiently utilized, the real-time response capability is improved, through comparison between a historical state and a current state, short-term noise fluctuation is filtered, a model is prevented from being misled by an abnormal value, and rapid modeling is achieved for a newly-accessed data source through data synthesis and model pre-training. Quality-driven aggregation is improved, a high-weight data source is dominant in parameter updating, and the accuracy of the model is improved.
Owner:CHINA NAT INST OF STANDARDIZATION

Target trajectory tracking method and product based on azimuth statistical analysis and prediction

The invention discloses a target trajectory tracking method and product based on azimuth statistical analysis and prediction, and the method comprises the steps: converting a horizontal array manifold vector according to a set interval, carrying out the array signal processing of the array manifold vector at each angle according to a set scanning interval, and obtaining the beam width at the angle; performing spatial spectrum synthesis on the horizontal array receiving data according to a plane wave array signal processing method to obtain a horizontal array data synthesis azimuth spectrum; presetting an initial azimuth angle of a to-be-tracked target, and solving a target azimuth angle at the current moment in combination with a target operation attribute corresponding to the azimuth angle, an array azimuth estimation error and a beam width; within the statistical time, adopting a least square method to predict a target azimuth angle, and adopting a target azimuth angle track compensation value to compensate a target azimuth angle prediction value; the target azimuth angle at the current moment is extracted according to the target azimuth angle predicted value, and then continuous and stable tracking of the target azimuth trajectory is completed.
Owner:INST OF ACOUSTICS CHINESE ACAD OF SCI

Systems and methods for generating synthetic data, and training and testing conversational artificial intelligence platforms

Methods and systems for generating and employing synthetic data are disclosed. The synthetic data is generated by defining roles for a plurality of speakers and inputting the roles to at least one Large Language Model (LLM), which in turn successively generates statements of each speaker which are responsive to generated statements for the other speaker based on the defined roles. Each successive set of statements are input to the LLM to generate additional statements of the speakers to obtain synthetic dialog data. The synthetic dialog data can be used to test and / or train neural networks as well as various platforms, including conversation analytics platforms.
Owner:SESTEK SES & ILETISIM BILGISAYAR TEKNOLOJILERI TIC & SAN AS

Mineral spectral characteristic unmixing device using generative adversarial network

The invention discloses a mineral spectral feature unmixing device using a generative adversarial network, which comprises a multi-modal data preprocessing module used for correcting and enhancing original mineral spectral data, solving the problem of small sample training and providing high-quality input for subsequent unmixing, and a generative network module used for fusing noise and geological text description, and providing high-quality input for subsequent unmixing. The dynamic adversarial training module is used for initializing end member features based on comparative learning of a mineral symbiosis sequence and improving priori cognition of the model on a mineral combination rule, and the end member reconstruction verification module is used for cyclically reconstructing a single albedo matrix by using a generator network unit and a discriminator network unit; nonlinear scattering and atmospheric noise interference are effectively eliminated through the model conversion unit and the 3D convolution kernel, the data synthesis capability of the conditional generative adversarial network is combined, the small sample training bottleneck is relieved, and the prior constraint of an end member feature combination rule is enhanced based on mineral symbiosis sequence comparative learning through a dynamic adversarial training mechanism.
Owner:CHINA AERO GEOPHYSICAL SURVEY & REMOTE SENSING CENT FOR LAND & RESOURCES

Nursing scene data enhancement and generation method for intention recognition model training

The invention discloses a nursing scene data enhancement and generation method for intention recognition model training, and belongs to the technical field of artificial intelligence and smart old-age care. The method aims at solving the problems that in the prior art, high-quality nursing scene training data is deficient, and the obtaining cost is high. According to the core technical scheme, the method comprises the steps that firstly, a staring sequence and other context information (such as time, place and physiological signals) of a user are processed through a multi-modal information fusion model, and a structured initial situation vector is generated; secondly, inputting the vector into a scene generation model combined with a nursing knowledge base, and automatically generating a batch of basic nursing scene data with intention labels; key points are that a data enhancement module is introduced, and a plurality of innovative strategies such as situation element disturbance, virtual physiological data synthesis, gaze path variation and virtual scene deduction are adopted to deeply process basic data, so that the diversity and complexity of a data set are greatly enriched; and finally, combining the basic data with the enhanced data to construct a final comprehensive training data set. According to the method, large-scale and high-fidelity training data can be generated in a low-cost and high-efficiency manner, and the accuracy and robustness of the intention recognition model in a real nursing environment are remarkably improved.
Owner:PEKING UNIV SHENZHEN GRADUATE SCHOOL

Photovoltaic power generation prediction system based on multi-source analysis model

The invention relates to the technical field of electric power prediction, discloses a photovoltaic power generation prediction system based on a multi-source analysis model, and aims to solve the problems that an existing photovoltaic power generation prediction technology is insufficient in individual and isomerized node prediction capability, and prediction is limited due to dependence on historical data in a newly-accessed and data-missing node scene. The system comprises a multi-dimensional node feature quantification module, a virtual historical data synthesis module, a power prediction module based on an enhanced data set, a closed-loop deviation traceability and correction module and a compensation strategy execution module oriented to a specific scene. Through adoption of the technical scheme, high-confidence prediction can be provided for blank or data missing nodes on the premise of not depending on historical data of the target node, and the precision, the coverage rate and the dynamic adaptability of distributed photovoltaic prediction are remarkably improved.
Owner:STATE GRID INFO TELECOM GREAT POWER SCI & TECH +2

Method for generating audio deep learning training data based on adversarial neural network

The invention relates to a method for generating audio deep learning training data based on an adversarial neural network (GAN), which is a data synthesis technology based on a generative adversarial network (GAN) and is used for generating audio training samples of scarce categories so as to effectively expand the diversity of a data set. According to the method, the generation capability of the GAN model is utilized to simulate the audio samples of the rare category, and the problem that the samples of the rare category are insufficient in a traditional data set can be solved. Through the technology, researchers can reduce the cost of manually collecting samples, the consumption of human resources is greatly reduced, and the universality of the data set in category coverage is ensured at the same time. In addition, the generated diversified samples can enhance the generalization ability of the model, so that the model is more stable when processing changeful data in the real world. According to the method, the training efficiency of the model can be effectively improved, the robustness of the model in different scenes can be improved, and the precision and practicability of an audio classification task in specific field application are further promoted.
Owner:CCTV INT NETWORK WUXI CO LTD

Multi-modal large model training method and device, equipment and storage medium

One or more embodiments of the invention provide a multi-modal large model training method, apparatus and device, and a storage medium, and the multi-modal large model comprises a visual coding layer used for generating image features corresponding to an image, and a large language model used for generating a reply text based on the image features and a query text; the method comprises the following steps: aiming at aligning an image feature space of a visual coding layer and a text representation space of a large language model, adjusting a multi-modal large model to obtain an aligned multi-modal large model; based on pre-training samples which are constructed in a data synthesis mode and correspond to various tasks in the at least one task in the medical scene, pre-training the aligned multi-modal large model to obtain a pre-trained multi-modal large model; and performing fine tuning on the pre-trained multi-modal large model based on a fine tuning sample which is obtained through a manual labeling mode and corresponds to the target task in the medical scene to obtain the multi-modal large model used for executing the target task.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Interactive scene data synthesis and key point visibility updating algorithm based on monocular vision

The invention relates to the technical field of multi-person posture estimation, in particular to an interactive scene data synthesis and key point visibility updating algorithm based on monocular vision, which comprises the following steps of: 1, arranging single-person picture data under monocular shooting vision, processing an input single-person picture under monocular shooting vision by using a GrondingDINO visual language open type target detection model, and obtaining a target detection result; and the Person is used as a retrieval keyword to identify a character individual in the picture. Through fine processing of the steps, especially introduction of target matching, target position adjustment and key point visibility updating methods, the quality of the synthesized image is greatly improved, the fusion degree of the target and the background is optimized through accurate matching and natural target pasting, and the image quality is improved. The space consistency and the visual naturalness of the synthetic image are enhanced, the problem of overfitting is effectively avoided through the improvement, the diversity of training data is enhanced, and therefore the generalization ability of the model is improved.
Owner:GUANGZHOU VIRTUAL POWER NETWORK TECH CO LTD

Data Synthesis Using Generative Models

In disclosed techniques a system generates, using a generative model, current synthetic communications, including inputting conditions for the synthetic communications into the trained generative model. The system generates the trained generative model by iteratively performing multiple operations until a discriminator of the generative model determines that synthetic communications output by the generative model satisfy a difference threshold. The operations include: generating, by a generator of the generative model, based on existing communications, a training synthetic communications, determining, by the discriminator of the generative model, differences between the existing communications and the training synthetic communications, and updating the generator based on the differences. Using the current synthetic communications and the existing communications, the system trains another model to evaluate newly initiated communications. The disclosed data synthesis techniques may advantageously enable discovery of concealed patterns, which in turn improves detection of processing systems that execute models trained on the synthetic data.
Owner:PAYPAL INC

Embedded data synthesis method and device integrating retrieval and large model distillation and medium

The invention provides an embedded data synthesis method and device fusing retrieval and large model distillation and a medium. The method comprises the following steps of: preprocessing an unstructured document in a vertical field, and dividing the unstructured document into multi-granularity text blocks with a hierarchical association relationship; forming a context based on the combination of the multi-granularity text blocks, injecting disturbance information corresponding to the priori knowledge in the vertical field into the context, calling a generative model to generate a retrieval query according to the context, and determining a target text block corresponding to the retrieval query as an initial positive sample; false negative sample text blocks are filtered according to the incidence relation between the text blocks, and a positive sample set and a negative sample set are formed; and constructing a comparative learning training sample, and training the semantic representation model by using the comparative learning training sample to generate an embedded vector for the retrieval task. According to the method, the retrieval task construction efficiency and authenticity can be improved, the positive sample coverage integrity is improved, and the contrast learning training stability and retrieval precision are enhanced.
Owner:北京衔远有限公司

Mathematical application question solving method and system based on data synthesis

The invention belongs to the technical field of mathematical solving, and provides a mathematical application question solving method and system based on data synthesis, and the method comprises the steps: obtaining a to-be-solved mathematical application question; constructing a reasoning chain of the mathematical application questions based on a large language model; judging whether the constructed reasoning chain is converted into a program auxiliary chain or not according to the prompt judgment word; if conversion is needed, the large language model generates an inference chain and a program auxiliary chain, the program auxiliary chain is executed by using a code interpreter, and solving of the mathematical application question is completed; otherwise, solving the mathematical application problem directly according to the reasoning chain.
Owner:UNIV OF JINAN

Equipment data-free federation incremental learning method and system under resource limitation

The invention discloses a resource-limited equipment data-free federation incremental learning method and system, and relates to the technical field of transfer learning, and the method comprises the steps: training a model based on an incremental data flow through a deviation correction mechanism, obtaining a trained local parameter, uploading the trained local parameter to a cloud, and carrying out the data-free federation incremental learning of the equipment; calculating a difference value of the parameters to obtain a gradient of each edge computing device, calculating a federal average updating direction, calculating an aggregation weight according to a deviation degree between the updating gradient and the federal average updating direction, and performing weighted fusion on the updating gradient to obtain an updated global model; a synthesis sample of the old task is generated with the purpose of minimizing diversity loss; and carrying out knowledge distillation based on the synthetic sample, migrating old task knowledge to the updated global model, obtaining a final global model, and issuing the final global model to an edge end. Through deviation correction training, deviation degree and information entropy dual-perception aggregation and attention-guided data-free synthesis and distillation, federal incremental learning of resource-constrained edge equipment is realized.
Owner:HUAQIAO UNIVERSITY

Remote sensing image semantic segmentation method based on multi-stage subtitle driven diffusion model

The invention discloses a remote sensing image semantic segmentation method based on a multi-stage subtitle-driven diffusion model, and the method comprises the steps: firstly obtaining an original remote sensing image semantic segmentation data set, designing an instance segmentation strategy, and obtaining a remote sensing single-target instance sub-image data set; secondly, providing a two-stage remote sensing semantic subtitle generation algorithm, and migrating a pre-training diffusion model based on Stable Diffusion to a remote sensing scene by combining a cutting instance and a conditional fine tuning strategy to obtain a diffusion model adaptive to the remote sensing scene; thirdly, constructing a multi-layer weighted attention image-semantic mask joint generation framework based on the adaptive model, and generating a high-quality remote sensing target image and a semantic mask; then, providing a cross-scale semantic constraint data synthesis method based on a ground sampling distance to obtain enhanced remote sensing image data; and finally, training a divider by using the enhanced data to realize accurate semantic segmentation of the remote sensing image. The method can effectively alleviate the dependence of annotation data, and improves the segmentation precision and scene adaptability.
Owner:HOHAI UNIV

Evaluation data synthesis system integrating multi-model collaborative question setting and multi-strategy filtering

The invention belongs to the technical field of artificial intelligence, and particularly relates to an evaluation data synthesis system integrating multi-model collaborative question setting and multi-strategy filtering. The system comprises an evaluation category definition and seed question bank construction module which is used for systematically defining an evaluation target and constructing a multi-dimensional seed question bank; the multi-model collaborative question setting module is used for generating an initial question pool by calling a large language model and taking an output result of the evaluation category definition and seed question bank construction module as input; the multi-model automatic quality inspection module is used for filtering low-quality questions in the initial question pool in two rounds by introducing a large language model as a quality inspection judgment model; the multi-strategy duplicate removal module is used for performing duplicate removal processing on the question pool by adopting a semantic vector rapid matching and large language model semantic judgment mode; and the seed question bank iteration module is used for taking the residual evaluation questions in the question pool as high-quality evaluation questions and completely supplementing the high-quality evaluation questions to the initial multi-dimensional seed question bank to form an updated seed question bank.
Owner:HANGZHOU YUANYU INTELLIGENT TECHNOLOGY CO LTD

Three-dimensional radar echo reflectivity variational auto-encoder pre-training method

The invention discloses a three-dimensional radar echo reflectivity variational auto-encoder pre-training method, which belongs to the technical field of meteorological radar data analysis, and comprises the following steps: extracting advanced features from three-dimensional high-resolution radar data through a space encoder based on an attention mechanism; mapping the advanced features into determined probability distribution parameters using a probabilistic encoder; sampling according to the distribution parameters by using a re-parameterization technique to obtain continuous potential variables; mapping the potential variables back to a pixel space by using a probability decoder based on an attention mechanism to complete data reconstruction; and finally, constructing a mixed loss function consisting of a mean square error and KL divergence by utilizing a reconstruction result and a distribution parameter to train the model. According to the invention, by optimizing the network architecture and introducing the attention mechanism, high-quality reconstruction of high-resolution three-dimensional radar data is realized while extremely low video memory requirements and low parameter quantity are ensured, and the practical value of data synthesis and enhancement is remarkably improved.
Owner:CHENGDU UNIV OF INFORMATION TECH +2

Multi-modal sample data synthesis and labeling integration method and device, equipment and storage medium

The invention discloses a multi-modal sample data synthesis and labeling integration method and device, equipment and a storage medium, and relates to the technical field of information extraction. The method comprises the following steps: firstly, grouping a collected sample data set, calculating a reasoning confidence average value, a standard deviation and a multi-modal large model false detection rate of each group of first sample data, and determining a target group according to the reasoning confidence average value and the standard deviation; and generating a target sample feature part in combination with the grouping condition of the target group and the grouping noise, and combining the feature part and the first actual label part into target sample data after model prediction and rechecking. Meanwhile, according to the target grouping demand number, the target sample generation number and the conversion coefficient, the supplementary collection number is determined, a supplementary collection data set is obtained, and after merging, image and context features of incremental samples are extracted and coded and fused. And finally, inputting the fusion code into the model, and optimizing the model by directly using the loss value or the corrected loss value according to whether the fusion code is a collection type, thereby effectively filling the weak region of the sample and improving the performance of the model.
Owner:BEIJING DIGITAL CHINA CLOUD COMPUTING CO LTD

Diversified text sensitive data synthesis method for desensitization effect evaluation

The invention provides a diversified text sensitive data synthesis method for desensitization effect evaluation, which comprises the following steps: constructing a sensitive entity system meeting desensitization evaluation requirements, the sensitive entity system comprises a general field, a medical field and a financial field, and each field comprises a plurality of entity types; obtaining an original data set, counting entity distribution on the original data set, constructing a target distribution model, and designing a diversified strategy based on entity types and sentence patterns; guiding the large language model to generate candidate corpora according to the target distribution model and the diversification strategy, performing character-level alignment labeling and consistency verification on the candidate corpora, and generating a synthetic data set based on the candidate corpora; and performing multi-dimensional quality verification on the synthetic data set from the data layer, the entity layer and the semantic layer to obtain an evaluation result, and feeding back the evaluation result to a closed-loop controller to adjust a quota and a generation parameter so as to generate a final synthetic data set. The method can be applied to validity evaluation of a data desensitization tool, and the problems of single evaluation dimension, data sparsity and the like in text desensitization evaluation are solved.
Owner:SUN YAT SEN UNIV

CARS-LA-ICP-MS combined multimode analysis system and analysis method

ActiveCN121410094AMaterial analysis by electric/magnetic meansOptical parametric amplifierElement analysis
The invention discloses a CARS-LA-ICP-MS combined multimode analysis system. The CARS-LA-ICP-MS combined multimode analysis system comprises a fusion light path module, a mass spectrometry gas path module, a three-dimensional mobile station, a CARS signal acquisition system and a computer control system, the fusion light path module comprises a laser source, an optical parameter amplifier, a light pulse time delay system, three dichroscopes, a light beam adjusting system, a galvanometer, a movable objective lens and a camera; the mass spectrometry gas circuit module comprises a carrier gas supply device, a multimode analysis sample cell and a mass spectrometry device; the multimode analysis sample cell comprises an upper cover, a lower cover and a side wall shell. On the other hand, the invention discloses an analysis method which comprises initialization setting, CARS analysis, mass spectrum analysis task parameter setting, laser ablation mass spectrum analysis and data synthesis. Through coaxial optical path design and an intelligent algorithm, deep fusion of chemical imaging and elemental analysis is realized, and a brand new research tool is provided for the fields of life science, material science and the like.
Owner:SHANGHAICHEMLABINSTRUMENTCO LTD