Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

503 results about "Large scale data" patented technology

Large scale data analysis is the process of applying data analysis techniques to a large amount of data, typically in big data repositories. It uses specialized algorithms, systems and processes to review, analyze and present information in a form that is more meaningful for organizations or end users.

Parallel computing method and system suitable for large-scale data processing

PCT designated stageWO2026007489A1Resource allocationResource poolPathPing
The present application relates to the technical field of large-scale data processing, and particularly relates to a parallel computing method and system suitable for large-scale data processing. The system comprises a task management unit, a distributed load balancing module, an elastic expansion architecture, an intelligent communication optimization module, and a resource monitoring unit, wherein the task management unit divides large-scale data into a plurality of sub-tasks by means of a task decomposer, and distributes the sub-tasks to computing nodes by means of a task scheduler and a priority distributor; the distributed load balancing module achieves global load balancing by means of a load sensing unit, a dynamic adjustment unit and a balance optimization unit; the elastic expansion architecture dynamically adjusts system resources by means of a node manager, a resource pool controller and an expansion decision-making device; the intelligent communication optimization module optimizes inter-node communication by means of a communication path planning unit, a bandwidth distribution unit and a delay compensation unit; and the resource monitoring unit monitors the system performance in real time by means of a performance collector, a state analyzer and an anomaly detector.
Owner:CHONGQING COLLEGE OF FINANCE ECONOMICS

Data privacy protection method for industrial internet platform

The invention discloses a data privacy protection method for an industrial internet platform, and relates to the technical field of data security and privacy protection. According to the method, the sensitive information is accurately identified and deeply analyzed through the sensitive information feature library and the hierarchical matching algorithm, the problem of insufficient accuracy and flexibility during large-scale data processing is solved, the accuracy and reliability of data desensitization are remarkably improved, personal privacy is effectively protected, data availability is maximized, and the method is suitable for large-scale data processing. A hierarchical processing algorithm and a self-adaptive desensitization rule base are utilized, desensitization rules are dynamically adjusted according to dynamic access requirements and sensitivity levels of data, data security and availability are balanced, accurate protection under different scenes is ensured, system adaptability and flexibility are improved, and the data access process is monitored in real time, anomaly detection and rule verification are performed, so that the data access efficiency is improved. And the desensitization rule is dynamically adjusted, so that the security and reliability of the system are enhanced, the user credibility is improved, and the transparency and credibility of the data processing process are ensured.
Owner:GUANGDONG JIUBIAN TECH CO LTD

General electromyographic signal processing method and system based on large self-supervised model

The invention discloses a general electromyographic signal processing method and system based on a large self-supervised model. The general electromyographic signal processing method comprises the following steps: step 1, acquiring a multi-source original multi-electrode channel EMG signal X from an electromyographic acquisition device; and finally, performing data unification processing, and finally converting into a space-time activity diagram with a fixed size of 224 * 224. On the basis of the space-time activity diagram and the fatigue state mark, constructing an AEMG for training according to heterogeneous unlabeled EMG data collected by a collection device; performing light-weight Adapter layer fine adjustment on the pre-trained large myoelectricity model to adapt to gesture recognition muscle force regression gait analysis or rehabilitation evaluation downstream tasks; aiming at the problem that the dimension and the structure of myoelectricity data are not matched due to different acquisition devices, acquisition parts and acquisition tasks, original signals are converted into space-time activity diagrams in a unified format through data unification processing, device differences are represented by combining a sensor embedding module, effective alignment of cross-source data is achieved, and the accuracy of the data is improved. And a basis is provided for large-scale data utilization.
Owner:SOUTH CHINA UNIV OF TECH

Protein palmitoyl transferase prediction method and system based on multi-branch deep convolutional neural network

The invention discloses a protein palmitoyl transferase prediction method and system based on a multi-branch deep convolutional neural network, and belongs to the technical field of bioinformatics and artificial intelligence. The method comprises the following steps: S1, obtaining a to-be-detected protein sequence; s2, inputting the protein sequence into a pre-trained iPalmT model; and S3, judging whether the target protein is palmitoyl transferase or not according to a model output result. The iPalmT model comprises a coding module, two paths of parallel convolution branches, a feature fusion module and a classification module; and after the convolution layers of each convolution branch are stacked, an SE module is arranged and is used for channel weighting and feature re-calibration. The model extracts multi-level sequence features through convolution kernels of different scales, realizes high-precision prediction through feature fusion and a residual structure, can automatically learn multi-scale features from large-scale data, realizes end-to-end palmitoyl transferase recognition, and has high accuracy and good universality.
Owner:WENZHOU MEDICAL UNIV

Large language model long thinking chain verification method and device based on reasoning process abstract

The invention relates to a large language model long thinking chain verification method and device based on an inference process abstract. The method comprises the following steps: obtaining a to-be-verified inference thinking chain; abstracting the obtained reasoning thinking chain to obtain a linear path containing a key reasoning link; and gradually executing verification on the linear path by adopting a checker, judging whether each step is correct or not, positioning the first wrong step, regarding the subsequent steps as invalid, and not independently verifying. Compared with the prior art, the method has the advantages that an accurate and efficient verification scheme is provided, the preciseness of process verification is guaranteed, the verification and labeling cost is remarkably reduced, and large-scale data set construction and model training can be supported.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Multi-dimensional data asynchronous calculation method and system based on dynamic dependency graph

The invention discloses a multi-dimensional data asynchronous calculation method and system based on a dynamic dependency graph, and relates to the field of data processing. The method comprises the following steps: constructing a metadata, logic and instance three-layer separation storage model; analyzing the reference relationship to construct a directed acyclic graph, and generating a calculation priority; monitoring data change, and executing asynchronous serialization calculation through a message queue based on priority; the associated document is automatically updated based on anchor mapping. According to the method, the calculation deadlock and the performance bottleneck of large-scale data in the Web environment are solved, logic decoupling and dynamic expansion are realized, the final consistency of the data is guaranteed, and the high-concurrency throughput and the stability are remarkably improved.
Owner:XINJIANG UNIVERSITY

Teenager psychological sub-health intelligent early warning system based on multi-source heterogeneous data fusion

PendingCN121601241AHealth-index calculationBiological modelsOnline interventionSocial media
The invention discloses a teenager psychological sub-health intelligent early warning system based on multi-source heterogeneous data fusion, and the system comprises a data collection layer which collects the behavior, physiological and social three-dimensional data of teenagers through the multi-source channels of campus cards, wearable devices, social media and questionnaires, and builds an original data pool through cleaning, denoising and standardization; the data fusion layer is used for integrating multi-source data by adopting weighted average and Kalman filtering, mining psychological sub-health key indexes in combination with a feature selection and extraction technology, and forming a three-dimensional psychological portrait; the deep learning layer is used for constructing a multi-modal fusion early warning model based on a Transform architecture, and carrying out real-time prediction and dynamic tracking of psychological sub-health risks through large-scale data training and cross validation optimization; and the intelligent early warning layer is used for visually displaying an early warning result, integrating three-party linkage of a management end, a teacher end and a parent end, providing 24-hour online intervention by a built-in AI psychological counseling module, and automatically transferring high-risk cases to professional psychological consultants.
Owner:YICHUN UNIVERSITY

Visual intelligent reconstruction evaluation system for three-dimensional wave liquid level

The invention discloses a three-dimensional wave liquid level visual intelligent reconstruction evaluation system, and belongs to the field of ocean engineering monitoring. The system comprises an image acquisition module, an image stereoscopic vision processing module, an attention-enhanced reconstruction neural network module, a camera attitude evaluation module, a visualization and output module and a hydrodynamic parameter analysis module. An image is collected through a fixed baseline binocular camera system, after preprocessing, a neural network fused with a multi-scale attention mechanism is utilized to reconstruct a three-dimensional wave structure, coordinate system conversion is achieved in combination with self-supervised attitude evaluation, and finally hydrodynamic parameters such as significant wave height and a three-dimensional velocity field are extracted and visualized. The system does not need explicit calibration and large-scale data, has high automation, real-time performance and strong environmental adaptability, can be deployed on various platforms, and significantly improves the precision and engineering applicability of non-contact wave observation.
Owner:HARBIN INST OF TECH

Laboratory safety detection method, device, equipment and medium

The invention relates to a laboratory safety detection method and device, equipment and a medium. The method comprises the steps that a networking retrieval module is adopted to select a tool for large-scale data retrieval; guiding an intelligent agent RAG module to screen and verify information by adopting a thinking chain technology; in the generation stage, a contrast decoding method is used, the confidence difference between the reference lexical elements and the illusion lexical elements is calculated, and confidence reweighting is conducted on the reference lexical elements and the illusion lexical elements so as to strengthen real safety information and inhibit model illusion content. And inputting the laboratory scene image, the user input text instruction, the structured comprehensive query instruction and the effective information of all rounds into a multi-mode large language model reinforced by a thinking chain technology to generate a laboratory safety detection report, and proactively providing an optimization suggestion of the laboratory safety layout based on the laboratory safety detection report. According to the invention, the reasoning capability of the model in a complex security scene is enhanced, and the accuracy and reliability of a detection result are improved.
Owner:GUANGDONG UNIV OF TECH

Data processing method and device, electronic equipment and computer storage medium

The invention discloses a data processing method and device, electronic equipment and a computer storage medium, and relates to the technical field of data acquisition, the method comprises the following steps: acquiring an acquisition data packet sent by each data acquisition module, and analyzing according to the acquisition data packet to obtain successfully analyzed data, the successfully analyzed data at least comprising an acquisition node identifier; determining a memory address identifier of the successfully analyzed data according to the acquisition node identifier and a preset Hash mapping table, and after the successfully analyzed data is stored in a memory buffer area mapped by the memory address identifier, performing data quality verification according to the successfully analyzed data to obtain a quality verification result; after the data quality verification is completed, data reorganization is carried out according to the successfully analyzed data to obtain reorganized data, data restoration is carried out on the reorganized data according to the quality verification result to obtain final output data, and the method and device aim at improving the processing capacity of large-scale data collection so as to ensure the accuracy and timeliness of data collection.
Owner:HEFEI ZHONGKE CAIXIANG TECH CO LTD

Abnormality detection method and electronic equipment

The invention discloses an anomaly detection method and electronic equipment, and relates to the technical field of computers, time dependence and spatial topology characteristics can be captured at the same time by introducing running state characteristics changing along with time into nodes and describing an incidence relation between computing equipment through a dynamic graph structure; the dynamic graph neural network automatically learns the fused spatio-temporal features, can adapt to continuous evolution of network topology and node states, and improves the accuracy of anomaly recognition; and performing anomaly judgment based on the updated node embedding. Visibly, the technical problems that a statistical and machine learning method depends on artificial features and distribution hypothesis and is low in detection precision and poor in adaptability in a complex dynamic environment are solved, and the technical effect that better real-time performance and expandability are achieved in a large-scale data center and a computing network is achieved.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

System and method for language model architecture with dataset comparisons at scale

A computer-implemented natural language processing method can include storing a criteria embedding generated from a criteria document containing natural language or otherwise unstructured text in a vector database; for each individual document in the throughput of electronic documents: generating, with a first language model, a text summary of the individual document; conducting, using retrieval augmented generation, a semantic search to identify relevant passages from the vector database based on the text summary of the individual document; generating, with a second language model, retrieval-augmented text by performing semantic textual similarity with the identified passages and the text summary of the individual document; and generating, with a third language model, a set of assessment outputs based on the text summary of the individual document and the retrieval-augmented text.
Owner:ROYAL BANK OF CANADA

Data breakpoint recovery method and device, storage medium and program product

The embodiment of the invention provides a data breakpoint recovery method and device, a storage medium and a program product, and relates to the field of distribution. The method comprises the following steps: determining that transaction processing occurs, wherein the transaction processing is used for performing database increment change; generating a structured breakpoint identifier used for identifying the transaction processing position, wherein the structured breakpoint identifier comprises a transaction sequence number, a transaction fragment number and an offset in the fragment; the offset in the fragment is used for identifying the relative position of the SQL in the fragment; and determining the processing state of the message in the message queue based on the structured breakpoint identifier and the offset existing in the message queue. According to the method, the breakpoint recovery accuracy and the cross-platform compatibility are improved, and the processing efficiency and the system stability in a large-scale data synchronization scene are improved.
Owner:CETC JINCANG (BEIJING) TECH CO LTD

Vehicle fine granularity detection method based on three-dimensional grid and YOLOv11 transfer learning

The invention discloses a vehicle fine granularity detection method based on a three-dimensional grid and YOLOv11 transfer learning. The method mainly comprises three components: a source domain prediction network structure, a target domain prediction network structure and a transfer learning module. Wherein the source domain prediction network and the target domain prediction network are dual-channel deep networks fusing two-dimensional images and three-dimensional grids, and efficient migration of source domain knowledge in a target domain is realized by aligning feature distribution of the source domain and the target domain. Through deep fusion of three-dimensional grid information and two-dimensional image features, the method can maintain robust detection performance under adverse conditions of vehicle attitude change, illumination interference, shielding and the like, can effectively reduce large-scale data annotation and training cost, improves the rapid adaptation capability of the model in a new scene or a new vehicle type, and improves the robustness of the model. And a high-precision, extensible and rapid-iteration fine-grained identification and detection solution is provided for intelligent traffic monitoring, unmanned driving perception, military equipment identification and digital twin systems.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Privacy protection type data desensitization method for big data analysis platform

The invention discloses a privacy protection type data desensitization method for a big data analysis platform, and relates to the technical field of data security and privacy protection. According to the method, the sensitive information is accurately identified and deeply analyzed through the sensitive information feature library and the hierarchical matching algorithm, the problem of insufficient accuracy and flexibility during large-scale data processing is solved, the accuracy and reliability of data desensitization are remarkably improved, personal privacy is effectively protected, data availability is maximized, and the method is suitable for large-scale data processing. A hierarchical processing algorithm and a self-adaptive desensitization rule base are utilized, desensitization rules are dynamically adjusted according to dynamic access requirements and sensitivity levels of data, data security and availability are balanced, accurate protection under different scenes is ensured, system adaptability and flexibility are improved, and the data access process is monitored in real time, anomaly detection and rule verification are performed, so that the data access efficiency is improved. And the desensitization rule is dynamically adjusted, so that the security and reliability of the system are enhanced, the user credibility is improved, and the transparency and credibility of the data processing process are ensured.
Owner:TIBET TIANHE SHENGYU INFORMATION TECHNOLOGY CO LTD

Automatic construction method for CAD (Computer Aided Design) multi-modal data set

The invention belongs to the field of computer aided design systems, and relates to a CAD multi-modal data set automatic construction method, which comprises the following steps of: firstly, rendering a three-dimensional CAD model to be marked into a standardized multi-view two-dimensional image set through a data processing system, enhancing a multi-modal large model by utilizing a keyword vector and a low-rank adapter, and constructing a multi-modal data set; the method comprises the following steps of: extracting a geometric label from an image as a hard constraint, generating a text description of a natural language, then verifying a labeling result through a self-consistent strategy, and finally associating the multi-view two-dimensional image, the geometric label and the text description with an original model to construct a high-quality CAD multi-modal image-text pair data set. According to the method, visual feature extraction and language description generation are decoupled into two subtasks with clear purposes, so that the labeling consistency and the recognition precision are improved, and low-cost automatic construction of large-scale CAD data is realized.
Owner:ZHEJIANG UNIV

Performance management and control method and system for distributed system

The invention discloses a performance management and control method and system for a distributed system. The method comprises the following steps: receiving a plurality of link data packets of the distributed system within a preset time period; constructing a service dependency graph of the distributed system in a preset time period according to the plurality of link data packets; reasoning the service dependency graph by using a graph convolutional network model to obtain an abnormal probability of each graph node, and determining at least one target root cause path of the distributed system having an abnormality within a preset time period according to the abnormal probability of each graph node; and determining a target disposal strategy matched with each target root cause path from a preset disposal strategy library, and feeding back the target disposal strategy to the corresponding service node. The technical problems of low performance management and control efficiency and poor effect caused by incapability of efficiently processing large-scale data in a distributed system and incapability of accurately positioning a root path in a complex service scene in related technologies are solved.
Owner:CHINA TELECOM CORP LTD

Content auditing method and device, electronic equipment and storage medium

The invention provides a content auditing method and device, electronic equipment and a storage medium, and relates to the technical field of content auditing, multi-level feature extraction is performed on an original text to generate text vectors, and a self-adaptive hierarchical clustering algorithm is adopted to cluster the text vectors, so that the content auditing efficiency is improved. And meanwhile, key text features are extracted based on a clustering result to generate hierarchical violation labels, and the violation classification model is updated by utilizing new violation labels through an incremental learning mechanism. The problems that in the prior art, text variants are difficult to recognize due to insufficient keyword filtering semantic understanding, traditional hierarchical clustering calculation is high in complexity and cannot adapt to large-scale data, a machine learning classification model lacks a dynamic updating mechanism, so that new illegal content is missed, and the model is difficult to continuously adapt to new auditing requirements can be solved. The technical effects of improving the semantic recognition accuracy of the violation content, improving the real-time performance of large-scale text auditing, automatically discovering new violation categories, reducing leak detection and enhancing the expansibility and dynamic adaptability of a content auditing system are achieved.
Owner:PEOPLE CN CO LTD +1

Wireless environmental monitoring method and system based on internet of things

The present invention relates to the field of wireless environmental monitoring. Provided are a wireless environmental monitoring method and system based on the Internet of Things. The wireless environmental monitoring method based on the Internet of Things comprises: S1, deploying a plurality of wireless sensor nodes for collecting environment parameters, which comprise temperature, humidity, and air quality; S2, the sensor nodes transmitting collected data to a gateway device by means of low-power wireless communication technology; and S3, the gateway device performing preliminary processing on the received data, and transmitting the data to a cloud server by means of the Internet. The stability and anti-interference capability of data transmission are ensured by using Zigbee communication technology, adaptive frequency-hopping technology, and a dual-antenna system. In terms of data processing, the system uses edge computing technology to perform partial data processing at the gateway device, thereby reducing the computational pressure on the cloud server. Moreover, the cloud server uses a distributed computing architecture, thereby achieving efficient processing of large-scale data and providing good scalability.
Owner:HEBEI CHEM & PHARMA COLLEGE

Integrated artificial intelligence analysis and modeling platform supporting large-scale data processing

The invention belongs to the field of artificial intelligence, particularly relates to an integrated artificial intelligence analysis and modeling platform supporting large-scale data processing, and aims to solve the problems of low resource scheduling efficiency, model training reasoning splitting, data algorithm coupling, limited cross-modal modeling and the like under high-dimensional heterogeneous mass data. The platform comprises a distributed data access layer, a unified metadata governance engine, a heterogeneous computing resource scheduling center, a model life cycle management module, a multi-modal fusion modeling framework and a dynamic service orchestration interface, and realizes automatic data access governance, dynamic resource scheduling, model full-cycle management and control and cross-modal intelligent fusion. And the resource utilization rate, the modeling accuracy and the service deployment efficiency are remarkably improved.
Owner:BEIJING ANRUISHENG TECH CO LTD

Data blood relationship management analysis method and device, equipment and storage medium

The invention discloses a data blood relationship management analysis method, device and equipment and a storage medium, and the method comprises the steps: obtaining original data, carrying out the cleaning, conversion and standardization processing of the original data, and constructing a blood relationship model; storing the blood relationship model based on a JanusGraph graph database, and performing blood relationship analysis through a Gremlin query statement to obtain a blood relationship analysis result; when the change of the data entity is detected, the incremental update of the affected blood relationship is automatically triggered, so that the update delay of the blood relationship information can be effectively compressed from the traditional hour level to the minute level, and the technical bottlenecks of blood relationship information islands, low analysis efficiency and insufficient large-scale data expansibility are solved; the timeliness and the accuracy of data change influence evaluation and the data governance decision support capability are remarkably improved, and the speed and the efficiency of data blood relationship management analysis are improved.
Owner:CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD

Hepatocellular carcinoma prognosis model construction method based on immunogen cell death related gene and application

The invention relates to the technical field of hepatocellular carcinoma, in particular to a hepatocellular carcinoma prognosis model construction method based on immunogen cell death related genes and application, and an accurate prognosis prediction model is constructed by integrating the immunogen cell death related genes (IRGs) and molecular characteristics of hepatocellular carcinoma. A training set and a verification set provided by TCGA and ICGC databases are utilized, so that the model can perform effective sample analysis under a large-scale data background. Through consistency clustering analysis, the optimal clustering number is determined, the samples are orderly divided into different molecular subtypes, and it is ensured that the samples in each subtype have similar molecular characteristics. According to the invention, the capability of capturing liver cancer heterogeneity on the molecular level of the model is increased, so that the accuracy of prognosis prediction is improved.
Owner:THE FOURTH HOSPITAL OF HEBEI MEDICAL UNIVERSITY (HEBEI CANCER HOSPITAL)

Method and device for automatically generating AI drive front-end component for low-code platform

The invention relates to the technical field of front-end development and artificial intelligence, and discloses a low-code-platform-oriented AI-driven front-end component automatic generation method and device, the method supports automatic generation from a natural language to a front-end component, the use threshold of a low-code platform is remarkably reduced, and high intelligence is achieved; aST verification and a template completion mechanism are combined, so that the generated component is correct in grammar and complete in logic, and the code reliability is ensured; meanwhile, monitoring and dynamic optimization during operation are introduced, so that the components still keep efficient rendering in a large-scale data scene; by means of a feedback mechanism, the generation quality can be continuously improved, self-learning and personalized customization are achieved, and therefore the sustainable evolution ability is achieved; in addition, the method is wide in applicability and can be applied to various scenes such as low-code platforms, front-end development IDE and intelligent programming assistants.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Large-scale data analysis and prediction system based on machine learning

The invention discloses a large-scale data analysis and prediction system based on machine learning, and relates to the technical field of machine learning and big data. The system comprises a data acquisition module, a data preprocessing module, a feature engineering module, a machine learning prediction model module, a model training module, an analysis prediction output module and a visual display module. According to the method, the optimal BPNN architecture is selected according to the prediction precision returned by the test set, the optimal neural network architecture containing the optimal initialization weight and bias is acquired, the future state is predicted by using the machine learning model, and the guidance path is dynamically adjusted, so that accurate analysis of logistics information is realized, the logistics efficiency is improved, and the logistics operation cost is reduced.
Owner:BEIFANG UNIV OF NATITIES

Json file parallel analysis and data writing method and device based on DataX

The invention relates to a Json file parallel analysis and data writing method and device based on DataX. The method comprises the following steps: acquiring a source end Json file set, and dividing the source end Json file set into small file subsets and large file subsets according to file sizes; the small file aggregation path is configured with a json reader of the DataX, and parallel reading and analysis are carried out; splitting a root array of the large file into a plurality of independent small Json files through a Python script, aggregating paths, and reading and analyzing by using a DataX parallel channel; and finally executing the DataX Job to write the data into a target end. Through a classification processing strategy, the large-scale Json data analysis writing efficiency is remarkably improved, the large file memory occupation risk is reduced, the file scale is automatically adapted, a mature DataX frame and a Python script are reused, high performance and low maintenance cost are both considered, multi-type target end storage is supported, and the problems that a traditional tool is weak in parallel, high in resource consumption, poor in flexibility and the like are effectively solved.
Owner:CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD

Data center load balancing system and method based on SDN

The invention discloses a data center load balancing system and method based on an SDN (Software Defined Network). The system comprises a switch monitoring module, a load balancing module and a flow table issuing module. The method comprises the following steps: periodically obtaining network global state information through an in-band network telemetry INT technology; calculating a global average priority based on the INT data to distinguish high and low priority data streams; dynamically calculating an optimal path in K shortest paths for avoiding congestion nodes for the high-priority flow, and selecting a path for the low-priority flow by adopting an improved ECMP algorithm based on link utilization rate weighting; and finally compiling the path strategy into a flow table item and issuing the flow table item to the switch. Through a priority-aware dynamic load balancing mechanism, the Hash collision and static defects of the traditional ECMP are overcome, the transmission delay of the high-priority data flow is effectively reduced, the bandwidth utilization rate of the low-priority data flow and the whole network flow transmission quality are improved, and the method is suitable for a large-scale data center network.
Owner:JIANGSU YITONG HIGH TECH

Compression and decompression method based on generic genome representation

PendingCN121237234ADigital data information retrievalSpecial data processing applicationsReference genome sequenceHuman DNA sequencing
The invention discloses a compression and decompression method based on generic genome expression, and relates to the technical field of compression and decompression of DNA next-generation sequencing data, in particular to the compression and decompression method based on generic genome expression. The method aims at solving the problems that in the prior art, the capacity of processing population genetic diversity is insufficient, original sequencing quality information cannot be effectively restored during decompression, and memory occupation is too high during large-scale data processing. Obtaining a to-be-compressed sequencing sequence data file, a reference genome sequence and a thousand-person genome variation sample; obtaining a haplotype list, a variation list and a haplotype offset list corresponding to each window block; storing the window number, the haplotype number, the haplotype offset, the head and tail unmatched sequences, the current sequence name and the quality score character string into a single compression block; carrying out binding storage; completing the compression processing of the mass fraction; and obtaining each to-be-compressed sequencing sequence based on the result of the compressed part.
Owner:HARBIN INST OF TECH

Question and answer data construction method and device

A question and answer data construction method comprises the following steps: 1) reading samples containing image and text data in batches, constructing a unified instruction for each batch of samples, and calling a large language model to extract candidate questions which have discriminability and teaching value and can be answered from images to form multiple groups of candidate questions; (2) carrying out structured deduplication merging on the cross-batch candidate problems obtained in the step (1) to obtain a deduplication problem pool; 3) for each sample, identifying an available image mode of the sample, screening candidate problems matched with the mode from the problem pool, and randomly extracting candidates of which the number is greater than a target value q; 4) constructing a request only based on image answering, calling the multi-modal large language model by taking the sample image as input, and generating a question-answer pair; if reliable answers of the candidate questions cannot be obtained only through the images, replacing the candidate questions with new candidates and retrying until q question and answer pairs capable of being answered are accumulated and obtained; 5, the generated question and answer pairs are subjected to structured verification and safe disking.The method supports parallel processing and breakpoint continuous running so as to meet the large-scale data construction.The method has the advantages of being good in universality and expandability and standardized in result.
Owner:ZHEJIANG UNIV

Map reconstruction method, device and equipment, vehicle, storage medium and program product

The embodiment of the invention provides a map reconstruction method and device, equipment, a vehicle, a storage medium and a program product. The method comprises the following steps: acquiring image data of a target scene; performing semantic segmentation processing on the image data to obtain a semantic segmentation result; wherein the semantic segmentation result comprises a ground semantic segmentation map, a sky semantic segmentation map and a dynamic obstacle semantic segmentation map; performing point cloud reconstruction processing on the image data according to a semantic segmentation result to obtain point cloud data; wherein the point cloud data comprises sparse point cloud and ground dense point cloud; performing three-dimensional Gaussian processing on the point cloud data to obtain a three-dimensional Gaussian model; generating a scene map of the target scene according to the three-dimensional Gaussian model; wherein the scene map comprises annotation information of the passable area. According to the method, high-precision and automatic three-dimensional scene map reconstruction and passable area labeling can be realized, and the requirement of large-scale data acquisition is met.
Owner:NINGBO LOTUS ROBOTICS CO LTD

Data processing method and device, equipment and storage medium

The invention discloses a data processing method and device, equipment and a storage medium, and relates to the technical field of data processing. The method comprises the following steps: acquiring a first rule to be stored, wherein the first rule comprises a first matching condition and first rule content; storage strategy information corresponding to the first matching condition is generated in a rule storage structure table in the memory, the storage strategy information comprises first partition information used for indicating that the first matching condition is stored in the disk, and the first partition information is partition information of a first partition in the disk; and determining a first condition node corresponding to the first partition and a first rule node configured for the first condition node from a rule network diagram of the disk, storing the first matching condition in the first condition node, and storing the first rule content in the first rule node. Therefore, in a large-scale data storage scene, memory storage is replaced by disk storage, and the problem of high cost of memory storage rules can be effectively solved.
Owner:CHINA UNIONPAY