Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

270 results about "Data synthesis" patented technology

Data synthesis. meta analysis A method that uses statistical techniques to combine results from different studies and obtain a quantitative estimate of the overall effect of a particular intervention or variable on a defined outcome—i.e., it is a statistical process for pooling data from many clinical trials to glean a clear answer.

Training of multi-modality object detectors

Techniques for determining a presence of an object, especially an object such as animal or debris, in a path of a vehicle, are discussed herein. For example, sensors of various modalities, which may include multispectral sensors, may capture data representing an environment the vehicle is traversing. In examples, one or more trained machine learned (ML) models, operating on a vehicle computing system, may detect and / or classify objects in the environment, based on input data of one or more modalities or spectral bands. The ML models may be pre-trained using training data including real sensor data, synthetic data, and / or augmented data, along with auto-generated annotations. In some examples, hyperspectral data may be used to identify materials associated with detected objects. A confidence score associated with the detection of the object may also be computed. The vehicle may be controlled based on detection of the object and its classification.
Owner:ZOOX INC

Single crystal turbine blade crystal orientation defect detection system based on adaptive optics

The invention discloses a single crystal turbine blade crystal orientation defect detection system based on adaptive optics, particularly relates to the field of nondestructive testing, and is used for solving the problem that thermal damage and signal consistency are difficult to consider in acoustic imaging of complex structural parts. Through partition energy threshold setting, dynamic pulse planning, photoacoustic processing and unified coordinate cross-subarea data synthesis, local thermal damage caused by excessive concentration of energy is effectively avoided in the material detection process, and the requirement for collecting high-bandwidth signals can be met in a large range. The signal-to-noise ratio and the detection efficiency are improved by regulating and controlling pulse energy and a scanning path, the amplitude comparability and the image coherence among regions are ensured through signal normalization correction and subregion image set splicing, high-resolution and comprehensive-coverage acoustic mapping is output, the crystal orientation defects of the single crystal turbine blade are positioned and evaluated, the detection precision is improved, the measurement period is shortened, and the method is suitable for large-scale popularization and application. And the method has an applicable value for aero-engine parts requiring rapid and high-precision nondestructive detection.
Owner:ANHUI GONGYUAN ELECTRONIC TECHNOLOGY CO LTD

Method for synthesizing electronic medical record data based on semantic processing

The invention discloses an electronic medical record data synthesis method based on semantic processing, and relates to the technical field of medical informatization, and the method comprises the following steps: S1, constructing a probabilistic medical knowledge graph; s2, generating a semantic representation vector; s3, constructing a multi-dimensional dynamic space-time atlas; s4, generating a discrete personalized disease course event sequence with space-time coordinates; s5, taking the discrete personalized disease course event sequence and the corresponding medical entity semantic representation vector as condition input, guiding the improved TSDiff model to execute an iterative denoising process, and outputting a multi-dimensional random disease course trajectory; s6, forming multi-modal electronic medical record data; and S7, performing multi-dimensional quality evaluation on the multi-modal electronic medical record data. According to the method, the limitations of logic inconsistency, modal splitting and model capability solidification in a traditional synthesis method are overcome, and an efficient and accurate solution is provided.
Owner:BEIJING INTELLIGENT DECISION MEDICAL TECH CO LTD

Systems and methods for generating synthetic data, and training and testing conversational artificial intelligence platforms

Methods and systems for generating and employing synthetic data are disclosed. The synthetic data is generated by defining roles for a plurality of speakers and inputting the roles to at least one Large Language Model (LLM), which in turn successively generates statements of each speaker which are responsive to generated statements for the other speaker based on the defined roles. Each successive set of statements are input to the LLM to generate additional statements of the speakers to obtain synthetic dialog data. The synthetic dialog data can be used to test and / or train neural networks as well as various platforms, including conversation analytics platforms.
Owner:SESTEK SES & ILETISIM BILGISAYAR TEKNOLOJILERI TIC & SAN AS

Nursing scene data enhancement and generation method for intention recognition model training

The invention discloses a nursing scene data enhancement and generation method for intention recognition model training, and belongs to the technical field of artificial intelligence and smart old-age care. The method aims at solving the problems that in the prior art, high-quality nursing scene training data is deficient, and the obtaining cost is high. According to the core technical scheme, the method comprises the steps that firstly, a staring sequence and other context information (such as time, place and physiological signals) of a user are processed through a multi-modal information fusion model, and a structured initial situation vector is generated; secondly, inputting the vector into a scene generation model combined with a nursing knowledge base, and automatically generating a batch of basic nursing scene data with intention labels; key points are that a data enhancement module is introduced, and a plurality of innovative strategies such as situation element disturbance, virtual physiological data synthesis, gaze path variation and virtual scene deduction are adopted to deeply process basic data, so that the diversity and complexity of a data set are greatly enriched; and finally, combining the basic data with the enhanced data to construct a final comprehensive training data set. According to the method, large-scale and high-fidelity training data can be generated in a low-cost and high-efficiency manner, and the accuracy and robustness of the intention recognition model in a real nursing environment are remarkably improved.
Owner:PEKING UNIV SHENZHEN GRADUATE SCHOOL

Photovoltaic power generation prediction system based on multi-source analysis model

The invention relates to the technical field of electric power prediction, discloses a photovoltaic power generation prediction system based on a multi-source analysis model, and aims to solve the problems that an existing photovoltaic power generation prediction technology is insufficient in individual and isomerized node prediction capability, and prediction is limited due to dependence on historical data in a newly-accessed and data-missing node scene. The system comprises a multi-dimensional node feature quantification module, a virtual historical data synthesis module, a power prediction module based on an enhanced data set, a closed-loop deviation traceability and correction module and a compensation strategy execution module oriented to a specific scene. Through adoption of the technical scheme, high-confidence prediction can be provided for blank or data missing nodes on the premise of not depending on historical data of the target node, and the precision, the coverage rate and the dynamic adaptability of distributed photovoltaic prediction are remarkably improved.
Owner:STATE GRID INFO TELECOM GREAT POWER SCI & TECH +2

Method for generating audio deep learning training data based on adversarial neural network

The invention relates to a method for generating audio deep learning training data based on an adversarial neural network (GAN), which is a data synthesis technology based on a generative adversarial network (GAN) and is used for generating audio training samples of scarce categories so as to effectively expand the diversity of a data set. According to the method, the generation capability of the GAN model is utilized to simulate the audio samples of the rare category, and the problem that the samples of the rare category are insufficient in a traditional data set can be solved. Through the technology, researchers can reduce the cost of manually collecting samples, the consumption of human resources is greatly reduced, and the universality of the data set in category coverage is ensured at the same time. In addition, the generated diversified samples can enhance the generalization ability of the model, so that the model is more stable when processing changeful data in the real world. According to the method, the training efficiency of the model can be effectively improved, the robustness of the model in different scenes can be improved, and the precision and practicability of an audio classification task in specific field application are further promoted.
Owner:CCTV INT NETWORK WUXI CO LTD

Embedded data synthesis method and device integrating retrieval and large model distillation and medium

The invention provides an embedded data synthesis method and device fusing retrieval and large model distillation and a medium. The method comprises the following steps of: preprocessing an unstructured document in a vertical field, and dividing the unstructured document into multi-granularity text blocks with a hierarchical association relationship; forming a context based on the combination of the multi-granularity text blocks, injecting disturbance information corresponding to the priori knowledge in the vertical field into the context, calling a generative model to generate a retrieval query according to the context, and determining a target text block corresponding to the retrieval query as an initial positive sample; false negative sample text blocks are filtered according to the incidence relation between the text blocks, and a positive sample set and a negative sample set are formed; and constructing a comparative learning training sample, and training the semantic representation model by using the comparative learning training sample to generate an embedded vector for the retrieval task. According to the method, the retrieval task construction efficiency and authenticity can be improved, the positive sample coverage integrity is improved, and the contrast learning training stability and retrieval precision are enhanced.
Owner:北京衔远有限公司

Mathematical application question solving method and system based on data synthesis

The invention belongs to the technical field of mathematical solving, and provides a mathematical application question solving method and system based on data synthesis, and the method comprises the steps: obtaining a to-be-solved mathematical application question; constructing a reasoning chain of the mathematical application questions based on a large language model; judging whether the constructed reasoning chain is converted into a program auxiliary chain or not according to the prompt judgment word; if conversion is needed, the large language model generates an inference chain and a program auxiliary chain, the program auxiliary chain is executed by using a code interpreter, and solving of the mathematical application question is completed; otherwise, solving the mathematical application problem directly according to the reasoning chain.
Owner:UNIV OF JINAN

Equipment data-free federation incremental learning method and system under resource limitation

The invention discloses a resource-limited equipment data-free federation incremental learning method and system, and relates to the technical field of transfer learning, and the method comprises the steps: training a model based on an incremental data flow through a deviation correction mechanism, obtaining a trained local parameter, uploading the trained local parameter to a cloud, and carrying out the data-free federation incremental learning of the equipment; calculating a difference value of the parameters to obtain a gradient of each edge computing device, calculating a federal average updating direction, calculating an aggregation weight according to a deviation degree between the updating gradient and the federal average updating direction, and performing weighted fusion on the updating gradient to obtain an updated global model; a synthesis sample of the old task is generated with the purpose of minimizing diversity loss; and carrying out knowledge distillation based on the synthetic sample, migrating old task knowledge to the updated global model, obtaining a final global model, and issuing the final global model to an edge end. Through deviation correction training, deviation degree and information entropy dual-perception aggregation and attention-guided data-free synthesis and distillation, federal incremental learning of resource-constrained edge equipment is realized.
Owner:HUAQIAO UNIVERSITY

Remote sensing image semantic segmentation method based on multi-stage subtitle driven diffusion model

The invention discloses a remote sensing image semantic segmentation method based on a multi-stage subtitle-driven diffusion model, and the method comprises the steps: firstly obtaining an original remote sensing image semantic segmentation data set, designing an instance segmentation strategy, and obtaining a remote sensing single-target instance sub-image data set; secondly, providing a two-stage remote sensing semantic subtitle generation algorithm, and migrating a pre-training diffusion model based on Stable Diffusion to a remote sensing scene by combining a cutting instance and a conditional fine tuning strategy to obtain a diffusion model adaptive to the remote sensing scene; thirdly, constructing a multi-layer weighted attention image-semantic mask joint generation framework based on the adaptive model, and generating a high-quality remote sensing target image and a semantic mask; then, providing a cross-scale semantic constraint data synthesis method based on a ground sampling distance to obtain enhanced remote sensing image data; and finally, training a divider by using the enhanced data to realize accurate semantic segmentation of the remote sensing image. The method can effectively alleviate the dependence of annotation data, and improves the segmentation precision and scene adaptability.
Owner:HOHAI UNIV

Evaluation data synthesis system integrating multi-model collaborative question setting and multi-strategy filtering

The invention belongs to the technical field of artificial intelligence, and particularly relates to an evaluation data synthesis system integrating multi-model collaborative question setting and multi-strategy filtering. The system comprises an evaluation category definition and seed question bank construction module which is used for systematically defining an evaluation target and constructing a multi-dimensional seed question bank; the multi-model collaborative question setting module is used for generating an initial question pool by calling a large language model and taking an output result of the evaluation category definition and seed question bank construction module as input; the multi-model automatic quality inspection module is used for filtering low-quality questions in the initial question pool in two rounds by introducing a large language model as a quality inspection judgment model; the multi-strategy duplicate removal module is used for performing duplicate removal processing on the question pool by adopting a semantic vector rapid matching and large language model semantic judgment mode; and the seed question bank iteration module is used for taking the residual evaluation questions in the question pool as high-quality evaluation questions and completely supplementing the high-quality evaluation questions to the initial multi-dimensional seed question bank to form an updated seed question bank.
Owner:HANGZHOU YUANYU INTELLIGENT TECHNOLOGY CO LTD

Three-dimensional radar echo reflectivity variational auto-encoder pre-training method

The invention discloses a three-dimensional radar echo reflectivity variational auto-encoder pre-training method, which belongs to the technical field of meteorological radar data analysis, and comprises the following steps: extracting advanced features from three-dimensional high-resolution radar data through a space encoder based on an attention mechanism; mapping the advanced features into determined probability distribution parameters using a probabilistic encoder; sampling according to the distribution parameters by using a re-parameterization technique to obtain continuous potential variables; mapping the potential variables back to a pixel space by using a probability decoder based on an attention mechanism to complete data reconstruction; and finally, constructing a mixed loss function consisting of a mean square error and KL divergence by utilizing a reconstruction result and a distribution parameter to train the model. According to the invention, by optimizing the network architecture and introducing the attention mechanism, high-quality reconstruction of high-resolution three-dimensional radar data is realized while extremely low video memory requirements and low parameter quantity are ensured, and the practical value of data synthesis and enhancement is remarkably improved.
Owner:CHENGDU UNIV OF INFORMATION TECH +2

Multi-modal sample data synthesis and labeling integration method and device, equipment and storage medium

The invention discloses a multi-modal sample data synthesis and labeling integration method and device, equipment and a storage medium, and relates to the technical field of information extraction. The method comprises the following steps: firstly, grouping a collected sample data set, calculating a reasoning confidence average value, a standard deviation and a multi-modal large model false detection rate of each group of first sample data, and determining a target group according to the reasoning confidence average value and the standard deviation; and generating a target sample feature part in combination with the grouping condition of the target group and the grouping noise, and combining the feature part and the first actual label part into target sample data after model prediction and rechecking. Meanwhile, according to the target grouping demand number, the target sample generation number and the conversion coefficient, the supplementary collection number is determined, a supplementary collection data set is obtained, and after merging, image and context features of incremental samples are extracted and coded and fused. And finally, inputting the fusion code into the model, and optimizing the model by directly using the loss value or the corrected loss value according to whether the fusion code is a collection type, thereby effectively filling the weak region of the sample and improving the performance of the model.
Owner:BEIJING DIGITAL CHINA CLOUD COMPUTING CO LTD

Diversified text sensitive data synthesis method for desensitization effect evaluation

The invention provides a diversified text sensitive data synthesis method for desensitization effect evaluation, which comprises the following steps: constructing a sensitive entity system meeting desensitization evaluation requirements, the sensitive entity system comprises a general field, a medical field and a financial field, and each field comprises a plurality of entity types; obtaining an original data set, counting entity distribution on the original data set, constructing a target distribution model, and designing a diversified strategy based on entity types and sentence patterns; guiding the large language model to generate candidate corpora according to the target distribution model and the diversification strategy, performing character-level alignment labeling and consistency verification on the candidate corpora, and generating a synthetic data set based on the candidate corpora; and performing multi-dimensional quality verification on the synthetic data set from the data layer, the entity layer and the semantic layer to obtain an evaluation result, and feeding back the evaluation result to a closed-loop controller to adjust a quota and a generation parameter so as to generate a final synthetic data set. The method can be applied to validity evaluation of a data desensitization tool, and the problems of single evaluation dimension, data sparsity and the like in text desensitization evaluation are solved.
Owner:SUN YAT SEN UNIV

CARS-LA-ICP-MS combined multimode analysis system and analysis method

ActiveCN121410094AMaterial analysis by electric/magnetic meansOptical parametric amplifierElement analysis
The invention discloses a CARS-LA-ICP-MS combined multimode analysis system. The CARS-LA-ICP-MS combined multimode analysis system comprises a fusion light path module, a mass spectrometry gas path module, a three-dimensional mobile station, a CARS signal acquisition system and a computer control system, the fusion light path module comprises a laser source, an optical parameter amplifier, a light pulse time delay system, three dichroscopes, a light beam adjusting system, a galvanometer, a movable objective lens and a camera; the mass spectrometry gas circuit module comprises a carrier gas supply device, a multimode analysis sample cell and a mass spectrometry device; the multimode analysis sample cell comprises an upper cover, a lower cover and a side wall shell. On the other hand, the invention discloses an analysis method which comprises initialization setting, CARS analysis, mass spectrum analysis task parameter setting, laser ablation mass spectrum analysis and data synthesis. Through coaxial optical path design and an intelligent algorithm, deep fusion of chemical imaging and elemental analysis is realized, and a brand new research tool is provided for the fields of life science, material science and the like.
Owner:SHANGHAICHEMLABINSTRUMENTCO LTD

Differential privacy and comparative learning fused data synthesis method

The invention relates to the technical field of data privacy protection and artificial intelligence crossing, in particular to a differential privacy and contrast learning fused data synthesis method, which comprises the following steps: S1, data acquisition and clustering: acquiring a real training data set, and clustering the data set by using a DBSCAN algorithm; s2, establishing double discriminators: designing a double discriminator framework comprising a privacy discriminator and a utility discriminator, injecting adaptive differential privacy noise based on a clustering structure into gradient updating by the privacy discriminator to realize privacy protection, and focusing on keeping the quality and authenticity of generated data by the utility discriminator. Precise balance between privacy protection and data utility is realized through a double-discriminator architecture. The privacy discriminator focuses on differential privacy constraints to ensure that the generated data meet strict privacy requirements; the utility discriminator restrains the data distribution consistency through the Wasserstein distance, effectively reduces the damage of the noise to the data utility, and solves the problem that the privacy and the utility are difficult to consider in the prior art.
Owner:DATA SPACE RES INST

Vehicle rolling critical condition prediction method and device

The invention discloses a vehicle rolling critical condition prediction method and device. The method comprises the following steps: acquiring n groups of rolling data about vehicle rolling; based on the n groups of vehicle rollover parameters, performing new data synthesis processing to obtain m groups of enhanced rollover parameters corresponding to the vehicle rollover parameters; inputting the enhanced tumbling parameter into the implicit neural network, and outputting a tumbling result corresponding to the enhanced tumbling parameter; training the initial explicit neural network through the enhanced tumbling parameter and a tumbling result corresponding to the enhanced tumbling parameter to obtain a trained explicit neural network; and determining a critical condition for triggering the rolling of the target vehicle through the explicit neural network according to the input actual rolling parameter of the target vehicle. In this way, the implicit neural network is used to learn the rolling discrimination capability from the small-scale real vehicle rolling data; and then a large number of reasonable enhanced rolling parameters are generated to train an explicit neural network so as to output critical conditions of vehicle rolling expressed in a specific parameter threshold form.
Owner:CATARC AUTOMOTIVE TEST CENT TIANJIN CO LTD

Synthetic data generation method and device and storage medium

The embodiment of the invention provides a synthetic data generation method and device and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: in response to a data synthesis instruction initiated in a target page, acquiring scene data under multiple view angles, and performing three-dimensional Gaussian reconstruction on the scene data under the multiple view angles to obtain three-dimensional scene data; in response to an object editing instruction initiated in the target page, determining at least one scene object; determining data to be recorded according to the three-dimensional scene data and the at least one scene object; after a sensor is configured for a recording object, recording processing is carried out on the to-be-recorded data according to the configured sensor, and synthetic data of the recording object is obtained. According to the method, the efficiency, diversity, fidelity and automation degree of the generated synthetic data can be improved.
Owner:SZ ZHUOYU TECH CO LTD

Structural network fracture modeling method based on multi-scale factor constraint

PendingCN121831883ASeismic signal processingWell loggingMetric tensor
The invention relates to the technical field of oil and gas reservoir development and geological modeling, and discloses a multi-scale factor constraint-based tectonic network fracture modeling method, which comprises the following steps of: synthesizing a Riemannian metric tensor field on the basis of earthquake, logging and geomechanics data, and defining non-Euclidean distance cost of a fracture expanded in an anisotropic medium; initial seed points are screened according to the elastic strain energy density, and initial growth potential energy in a limited range is distributed; solving the eikonal equation by using an anisotropic fast marching algorithm to carry out wavefront competitive growth, and dividing a grid region into generalized Voronoi units; identifying a wavefront contact interface, and extracting gradient features to judge a fusion or truncation type so as to establish fracture topological connection; and finally, tracking a geodesic line path along an anti-gradient direction to generate a three-dimensional discrete fracture network. According to the method, macro and micro constraints are unified through Riemannian geometry, clear physical significance is given to the fracture by utilizing an energy mechanism, and automatic and accurate construction of the complex fracture network topology structure is realized.
Owner:CHINESE ACAD OF GEOLOGICAL SCI

Intelligent evaluation system and method for knowledge base question and answer application

The invention discloses an intelligent evaluation system and method for knowledge base question and answer application, and relates to the technical field of artificial intelligence, natural language processing and multi-agent collaborative systems. The system comprises a data preprocessing module, a data synthesis module, a data screening module, a data scoring module and a self-adaptive weight adjustment module. The data preprocessing module carries out data preprocessing on the enterprise original document and generates a data knowledge base; the data synthesis module adopts a large language model to perform repeated question and answer on each paragraph of the data knowledge base to generate question and answer pairs; the data screening module screens out high-quality samples through a screening agent collaborative screening MACA framework mechanism; the data scoring module performs question and answer scoring through a scoring agent collaborative scoring MACA framework mechanism; and the adaptive weight adjustment module dynamically updates the weight of each screening scoring dimension by adopting an exponential weighted moving average algorithm.
Owner:SHANGHAI PINJIAN INTELLIGENT TECH CO LTD

Road long-tail disease data synthesis method based on pavement degradation logic and linkage evolution

The invention provides a road long-tail disease data synthesis method based on pavement degradation logic and linkage evolution, and belongs to the technical field of computer vision and intelligent traffic. Comprising the following steps: preprocessing an original pavement image containing the long tail disease to obtain a core pixel block; determining a secondary crack area, and generating a secondary crack mask; the host pavement image generates a secondary disease pixel block according to the secondary crack mask; correcting the core pixel block in combination with the virtual depth and the self-shielding shadow intensity compensation coefficient to obtain a corrected core pixel block; consistency fusion of the synthetic disease block and the host road surface image is realized by utilizing feather mask and frequency domain and brightness alignment, and a final synthetic image is obtained; and automatically and synchronously generating a detection label according to the external contour of the feather mask. The problems that serious disease samples are scarce and traditional synthesis lacks physical reality are solved, the generated samples conform to the pavement mechanical degradation law, and the detection reliability of a deep learning model in a long-tail disease scene is remarkably improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Urban water supply pipe network leakage detection method, device, equipment, medium and product

The invention discloses an urban water supply pipe network leakage detection method, device and equipment, a medium and a product, and relates to the field of pipe network leakage detection, and the method comprises the steps: obtaining original radar image data; the method comprises the following steps: preprocessing original radar image data to obtain polarization decomposition component data and backscattering component data, and synthesizing the polarization decomposition component data and the backscattering component data into multichannel image data; constructing a deep convolutional neural network model, and training the deep convolutional neural network model through the multi-channel image data to obtain a prediction model; and inputting the to-be-predicted multi-channel image data of the target area into the prediction model, and outputting the leakage probability of the target point location, so that the timeliness of the leakage detection work of the water supply network can be improved.
Owner:PIPE NETWORK MANAGEMENT BRANCH OF BEIJING WATERWORKS GRP CO LTD +1

Oil-immersed transformer fault diagnosis method and system based on feature enhancement and deep isomerism

The invention relates to the technical field of oil-immersed transformer fault diagnosis, in particular to an oil-immersed transformer fault diagnosis method and system based on feature enhancement and deep isomerism. The method comprises the steps that a minority class oversampling method is used for synthesizing minority fault sample data into new samples conforming to physical constraints; an enhanced input space is constructed based on deep features generated by an auto-encoder VAE, and deep features of fault gas in a new sample are fully extracted; based on the extracted deep features, using a heterogeneous integrated model to simulate a lightweight gradient elevator Light GBM to capture a shallow relationship of the deep features; using a convolutional neural network 1D-CNN to identify local features of the fault gas; establishing a global dependency relationship on the basis of a Transform model; the robustness of the model is remarkably enhanced through a dynamic weighting mode, and therefore it is guaranteed that reliable and stable diagnosis results are continuously output in various complex and high-uncertainty practical application scenes.
Owner:YANTAI UNIV

Sound event detection data synthesis and sound event detection model training method

The invention discloses a semantic prompt-based sound event detection data synthesis and sound event detection model training method. The sound event detection task is converted into the semantic description information, and the structured semantic prompt instruction which accurately reflects the target sound event characteristics is generated in combination with the semantic constraint rule, so that automatic mapping from semantic description to instruction generation is realized, and the manual intervention cost is reduced. And inputting the structured semantic prompt instruction into the audio generation model, and synthesizing the audio data in a large scale, thereby reducing the data acquisition cost and improving the sample diversity and expandability. The structured semantic prompt instruction can guide the model to synthesize audios of various sound event types in batches, and is automatically generated by a large language model to ensure that the synthesized audios are strictly aligned with instruction semantics. When the label is generated, the sample event type can be obtained without manual labeling, and an efficient and reliable data source is provided for sound event detection model training.
Owner:SHANGHAI NORMAL UNIVERSITY +1

Ai model retraining with data distribution shift awareness

Embodiments herein use data synthesis to generate data distributions that predict how input data for a digital twin of a manufacturing process may drift as conditions change in the manufacturing process (e.g., tool deterioration, a change in materials, a design change, etc.). These different predicted data distributions can then be used, a priori, to build AI model variants (e.g., pre-trained AI models) for the digital twin. Thus, when input data drift is detected in the manufacturing process, the system can select one of the AI model variants to use which was built (or trained) using a data distribution that is similar to the new input data.
Owner:ADVANCED MICRO DEVICES INC

Artifact-driven data synthesis in computed tomography

ActiveUS12433560B2Image enhancementImage analysisNuclear medicineCT Image Artifact
Computer processing techniques are described for augmenting computed tomography (CT) images with synthetic artifacts for artificial intelligence (AI) applications. According to an example, a computer-implemented method can include generating, by a system comprising a processor, synthetic artifact data corresponding to one or more CT image artifacts, wherein the synthetic artifact data comprises anatomy agnostic synthetic representations of the one or more CT image artifacts. The method further includes generating, by the system, augmented CT images comprising the one or more CT image artifacts using the synthetic artifact data. In one or more examples, the method can further include training, by the system, a medical image inferencing model to perform an inferencing task using the augmented CT images as training images.
Owner:GE PRECISION HEALTHCARE LLC

Stainless steel corrosion rate prediction method based on virtual sample generation and transfer learning

The invention provides a stainless steel corrosion rate prediction method based on virtual sample generation and transfer learning, and relates to the technical field of data-driven prediction models, and the method comprises the steps: S1, obtaining target stainless steel material corrosion data and low alloy steel corrosion data, and carrying out the standardization processing; s2, determining the direction and range of virtual sample data generation based on an SMOTE virtual sample generation method, and generating a stainless steel material data synthesis sample; s3, constructing a cross-domain transfer learning model of a stainless steel material, training an artificial neural network model by using low alloy steel corrosion data, and transferring to a target domain model; and S4, constructing a corrosion performance prediction optimization model of the target stainless steel material, and performing optimization output to obtain a corrosion prediction result of the target stainless steel material. Cross-domain corrosion rule migration is realized through a virtual sample generation technology and migration learning, and an efficient and reliable solution is provided for stainless steel corrosion rate evaluation by increasing the basic data volume.
Owner:BEIJING JIAOTONG UNIV

A Method and System for Constructing an Industry Knowledge Base Based on Text-Based Data Synthesis

This invention relates to the field of knowledge base construction technology, specifically disclosing a method and system for constructing an industry knowledge base based on text-based data synthesis. The method includes entity recognition of multi-source text data, constructing a knowledge point sequence, and marking key knowledge points in the knowledge point sequence; expanding the key knowledge points based on a large language model to generate an expanded knowledge point set; constructing a knowledge graph based on the expanded knowledge point set, and constructing a preset number of knowledge connection paths in the knowledge graph; classifying the knowledge connection paths and inserting them as knowledge skeletons into the industry knowledge base. When dealing with massive amounts of multi-source text data, this invention expands the keywords based on a large language model to obtain expanded knowledge points, then constructs a knowledge graph, extracts node paths from the knowledge graph, uses the node paths as knowledge skeletons, classifies them, and stores them in the knowledge base. While the resulting knowledge base still contains a large amount of information, it significantly simplifies its size.
Owner:LANYUN NET

Mathematical question and answer method and device and computer program product

The invention discloses a mathematical question and answer method and device and a computer program product, and the method comprises the steps: firstly carrying out the tool integration cooperative reasoning of a to-be-answered sample mathematical question text proposed by a sample user through a tool integration data synthesis mode of an Actor agent and a Critic agent; and a high-quality tool integrated reasoning path sample data set is constructed by utilizing a reasoning result, so that good cross-model and cross-tool applicability is realized while the data quality is ensured. On the basis, after a fine-tuned question and answer model is obtained by utilizing a sample data set in a supervised fine-tuning manner, track-level and step-level layered optimization is performed on the fine-tuned question and answer model based on a reinforcement learning strategy, so that the mathematical reasoning ability of the optimized question and answer model is enhanced, and the question and answer model is optimized. And when the optimized question and answer model is used for answering the target mathematical question text proposed by the target user, the answering efficiency and accuracy can be effectively improved, and the question and answer experience of the target user is improved.
Owner:IFLYTEK CO LTD