Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

97 results about "Dirty data" patented technology

Dirty data, also known as rogue data, are inaccurate, incomplete or inconsistent data, especially in a computer system or database. Dirty data can contain such mistakes as spelling or punctuation errors, incorrect data associated with a field, incomplete or outdated data, or even data that has been duplicated in the database. They can be cleaned through a process known as data cleansing.

Electrical connection safety management and control method based on intelligent monitoring

The invention relates to the technical field of artificial intelligence, in particular to an electrical connection safety management and control method based on intelligent monitoring, which comprises the steps of data acquisition, time synchronization, quality evaluation, feature engineering and the like. According to the method, electric, thermal, mechanical, acoustic and other channels are placed on the same time axis through unified alignment and quality marking of multi-modal data and two-stage time synchronization and resampling, explicit marking is carried out on missing, noise and abnormal points, the weight of a subsequent algorithm can be reduced or skipped according to marks, and misjudgment caused by phase errors and dirty data is reduced.
Owner:CHET ELECTRONIC TECH (SHENZHEN) CO LTD

Automatic customer complaint work order circulation system and method based on composite intention disassembly

The invention discloses a customer complaint work order automatic circulation system and method based on composite intention disassembly. The system receives and preprocesses the multi-mode customer complaint data and initializes a service circulation context; business intention identification and business key slot position extraction are carried out; disassembling the composite business intention, and constructing an ordered work order execution queue of a directed acyclic graph structure through cyclic dependence verification; evaluating the work order circulation action with the maximum expected cumulative income by using a reinforcement learning model; when business knowledge query is involved, compliance evidence fragments are retrieved, sorted and output; and generating a candidate reply and a business operation instruction by the fine-tuned large language model, outputting the candidate reply after business compliance verification, and calling an API (Application Program Interface) of an underlying business system to execute entity business handling operation in a cross-system manner. The problems of data interaction and state synchronization among multi-source heterogeneous systems are solved, the automatic execution efficiency of the composite work order at the bottom layer node is improved, and the system unauthorized and dirty data risks caused by an uncertain instruction are effectively avoided.
Owner:FUJIAN GOTOP XINGYI NETWORK TECH

Data synchronization method and system for intelligent dirty data detection and restoration based on DataX

The invention relates to a data synchronization method and system for intelligent dirty data detection and restoration based on DataX. According to the method, an intelligent dirty data processing engine is embedded between a Reader plug-in and a Writer plug-in of DataX, type matching, format verification, constraint conflict pre-detection and abnormal value recognition are completed online, and type conversion, format standardization, abnormal value replacement or flexible writing are carried out on dirty data according to a preset or self-adaptive strategy. And finally, the Writer executes insertion, updating, skipping or log recording only according to an instruction, and a traceable governance log is generated in the whole process. According to the method, the synchronization-cleaning capability is embedded into the DataX framework, so that the synchronization success rate and the data quality are remarkably improved, the external ETL dependence and the artificial script cost are reduced, and the method has the advantages of high intelligence, high throughput and flexible configuration.
Owner:CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD

Data cleaning method for ship information infrastructure fault diagnosis

The invention discloses a data cleaning method for ship information infrastructure fault diagnosis, and belongs to the field of data processing. The method specifically relates to a preprocessing method of ship information infrastructure monitoring data and a data cleaning method for distinguishing fault data and dirty data. Through data preprocessing, a consistent and complete data set is constructed for subsequent data cleaning. At the moment, abnormal values in the data set are divided into interference data and abnormal equipment states, the abnormal equipment state values have important significance for subsequent fault diagnosis, and the interference data can influence model training. According to the method, a mode of combining FP-Growth and DBSCAN algorithms is adopted, and the characteristics that cabin equipment is complex in coupling and one fault often generates derivative alarms are utilized, so that equipment fault points and interference points are reasonably distinguished. Through the ship information infrastructure fault diagnosis-oriented data cleaning method, the problems of fuzzy data quality and low model precision caused by application of traditional fault diagnosis to ship information infrastructures can be solved.
Owner:CHINA SHIPBUILDING RES INST (SEVENTH RES INST OF CHINA STATE SHIPBUILDING CORP)

Standardized cleaning method for multi-source alarm data access and related device

The invention discloses a standardized cleaning method for multi-source alarm data access and a related device, and relates to the technical field of IT operation and maintenance in the financial industry, and the method comprises the steps: obtaining synchronization task parameters configured by a user, including a task execution plan, a processing class adaptive to multiple protocols, and an analysis and mapping rule, starting a task to receive multi-source alarm data, and obtaining a synchronization task; dirty data are screened through processing class matching and then analyzed into indexable objects, the indexable objects are mapped into data of a unified structure, formatting processing (including type conversion, null value filling and integrity and consistency verification) is carried out, and standardized cleaning is completed. According to the method, multi-protocol compatibility and multi-source data unified access are realized, scripts do not need to be customized, and technical barriers and maintenance cost are reduced; through data screening and standardization processing, the data quality and operation and maintenance efficiency are improved, quick response to key alarms is facilitated, and the continuity of financial core services is guaranteed.
Owner:QIDIAN HAOHAN DATA TECH BEIJING CO LTD

Robot training data analysis method and system based on deep learning

The invention provides a robot training data analysis method and system based on deep learning. The method comprises the steps that operation parameters reflecting the work intensity of a robot are acquired; acquiring original sensor data acquired by the robot; deducing the physical state deviation of an end effector of the robot based on the operation parameters; according to the physical state deviation, adjusting the original sensor data to obtain adjusted sensor data; and optimizing a robot control strategy based on the adjusted sensor data. The method can effectively identify and quantify the sensor data error of the robot caused by the physical state deviation, and solves the problem that the model performance is reduced due to the fact that a traditional deep learning system learns in dirty data by correcting the original sensor data, thereby improving the accuracy and efficiency of robot control strategy optimization, and improving the user experience. And invalid adjustment caused by error attribution is avoided.
Owner:BEIJING OUYI INTELLIGENT TECH CO LTD

A Lexical-Level Entity Matching Method and System for Heterogeneous Attributes

This invention provides a word-level entity matching method for heterogeneous attributes, comprising: S1: obtaining a set of entity pairs and dividing the set into a training set and a test set; S2: constructing a cross-word matching matrix using the training set, and reconstructing word vectors using the cross-word matching matrix to obtain word-level matching vectors; S3: training a matching model using the word-level matching vectors to obtain a trained matching model; S4: matching the test set using the trained matching model to obtain entity matching results. This invention converts words in attributes into vectors, and constructs a cross-word matching matrix by comparing these vectors with the vectors of each word in the entity to be matched. The cross-word matching matrix contains the vector information of the entire entity, and can adaptively obtain suitable matching objects for each attribute. It has the advantages of high accuracy and robustness in entity matching between data, and can effectively handle the dirty data problem that occurs in entity matching.
Owner:CHINA UNIV OF GEOSCIENCES (WUHAN)

Method for automatic structuring of JSON data and incorporation into a database

The application relates to the field of data processing and provides a JSON data automatic structuring and warehousing method. The main idea is to solve the problem of how to store different JSON data structures of various data sources in a warehouse. The scheme includes judging the type of the accessed JSON data source, obtaining JSON data by using different methods according to the type, pre-processing the data, checking and processing dirty data to obtain standard JSON fields, parsing and processing the obtained JSON of different JSON data sources, agreeing on the format of the data file, generating a standard data file, investigating the structuring processing progress of the data, generating an ok file, checking the accuracy of the generated data file, obtaining the standard data file after the checking, and performing batch processing on the data file to store the data file in the warehouse.
Owner:WUHAN ZBANK CO LTD

Light-weight AI-based phase modifier edge diagnosis system and method

The invention discloses a phase modifier edge diagnosis system and method based on lightweight AI, relates to the technical field of power management, and aims to solve the technical problem of large automatic defect of current phase modifier diagnosis. Comprising a multi-mode sensing module, a lightweight AI edge deployment module, a network communication module, a data processing and storage module and an autonomous decision-making and fault diagnosis module, and the lightweight AI edge deployment module comprises an adaptive unit to avoid total model transmission. According to the system, significant technical gain is realized through the lightweight AI edge deployment module, the adaptive unit is embedded into the working condition feature mapping layer, model weight is dynamically adjusted in combination with working condition parameters such as the load rate of the phase modifier and the voltage of the power grid, misinformation caused by dirty data input is reduced, and the reliability of the system is improved. And the edge increment updating unit only transmits the deviation sample to perform local parameter updating, so that the diagnosis accuracy, the working condition adaptability and the operation efficiency are comprehensively improved, and efficient and reliable technical support is provided for the edge diagnosis of the phase modifier.
Owner:INNER MONGOLIA UHV BRANCH OF STATE GRID INNER MONGOLIA EASTERN ELECTRIC POWER CO LTD +1

Dirt regulation control method and control device for surface cleaning equipment

The invention relates to a smudginess regulation control method and control device for surface cleaning equipment, and the method comprises the steps: obtaining a plurality of preset monitoring time windows, collecting smudginess induction values of the surface cleaning equipment in a first monitoring time window, and storing the smudginess induction values to obtain a first smudginess data set; comparing data in the first smudginess data set with the obtained initial smudginess judgment threshold value to obtain the smudginess degree of the target area; adjusting the cleaning behavior according to the smudginess degree of the target area; based on the first smudginess data set, a first smudginess judgment threshold value is calculated through a preset statistical algorithm, the first smudginess judgment threshold value is used for replacing the initial smudginess judgment threshold value and is used for comparison of a next monitoring time window, and the steps are repeated. The method can solve the problem of systematic drifting of the reference reading of the sensor, and realizes the accuracy of smudginess detection and the reliability of control in the whole equipment life cycle.
Owner:SUZHOU EUP ELECTRIC CO LTD

A makeup transfer model training method based on cross-identity triplets

The application discloses a makeup transfer model training method and a makeup transfer method based on cross-identity triplets, and relates to the technical field of computer vision and generative artificial intelligence. The model training method is divided into two stages of offline data construction and model training. In the offline data construction stage, a makeup prompt word dictionary is generated by a large language model; a homologous local rendering is performed on a nude face image cluster to generate a made-up image; a face mask is used to perform structure, background and rendering effectiveness filtering and semantic consistency verification, and a large number of high-quality cross-identity triplets are constructed through cross combination; a pre-trained diffusion model is fine-tuned under the strong supervision of a fixed text prompt and a double-path image condition, a conversion from text control to pure image control is realized, and a makeup transfer model is obtained. The application solves the problems of cross-identity makeup data scarcity, text semantic ambiguity and dirty data interference, realizes precise, lossless and what-you-see-is-what-you-get makeup transfer, and is suitable for film and television makeup, digital person rendering and e-commerce virtual makeup testing scenes.
Owner:ZHAOYI INFORMATION TECH (SHANGHAI) CO LTD

Data update methods, apparatus, computer equipment and storage media

This application provides a data update method, apparatus, computer device, and storage medium, relating to cloud technology, big data, cloud storage, and other technical fields. By determining an access queue based on an access index when at least two access requests trigger a data update, multiple access requests are accurately and conveniently centralized in the queue for unified management. The head request in the access queue updates the index data, and the head thread corresponding to the head request wakes up the waiting threads corresponding to the other access requests, allowing the remaining access requests to directly access the updated index data without repeating the update process. This avoids duplicate updates by a large number of access requests. Especially in high-concurrency scenarios, it prevents massive concurrent requests from penetrating the backend device, avoiding update errors, dirty data generation, operating system crashes, and other problems caused by simultaneous duplicate updates, ensuring the accuracy and reliability of the data update process.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A method and system for the whole life cycle processing of a procurement contract based on multi-table value parameter unified stored procedure and AI protection verification

This invention discloses a method and system for processing the entire lifecycle of procurement contracts based on a unified stored procedure with multiple TVPs and AI protection verification, belonging to the field of industrial software MES / ERP technology. This invention uses three sets of table-valued parameters (TVPs) as a unified data entry point, integrating six core operations—adding, modifying, deleting, appending details, modifying details, and deleting details—in a single stored procedure, and returning unified exception information through a single output parameter. The system incorporates a two-level centralized parameter verification mechanism. The first level performs a one-time empty table check on the three sets of TVP tables; if any table is empty, the transaction is immediately rolled back and an exception is returned. The second level performs centralized integrity and legality verification on all business parameters, achieving security protection for AI and external calls, effectively intercepting dirty data and illegal requests. This invention features a simple architecture, high execution efficiency, and comprehensive business coverage. The transaction mechanism ensures strong consistency of master-slave table data, supports compatible operation with both SQL Server and GaussDB databases, and can securely integrate into the AI ​​ecosystem without business restructuring. It is particularly suitable for the high-stability and high-security procurement contract business processing of complex factories with more than 11 production lines. This invention has been fully implemented in a real industrial ERP system and can run stably in a real production environment for a long time, demonstrating mature engineering value and industrial promotion capabilities.
Owner:HANDAN DINGSHENG DIGITAL INTELLIGENCE TECHNOLOGY CO LTD

Cache and control method therefor, and computer system

Embodiments of this disclosure provide a cache, a control method therefor, and a computer system, and relate to the field of storage technologies, to improve memory access efficiency of the computer system. The cache is connected to a memory controller, and the cache includes a plurality of cache lines. The control method for a cache includes: storing write data of a received write command into cache lines, and before the cache lines are allocated to new memory addresses, sending dirty data stored in the cache lines to the memory controller.
Owner:HUAWEI TECH CO LTD

A data cleaning method and device

The embodiment of the present application provides a kind of data cleaning method and device, the method includes determining data attribute field from each business field of business scene, by data attribute field, determine the search field set from each database table, and determine the database table set matched with search field set, the source data associated with business scene is analyzed, determine the business data value of the business field to be cleaned, according to search field set and database table set, the structured query language sql data cleaning script corresponding to the business data value of the business field to be cleaned is constructed, and the relevant business data in the database table set associated with the business data value of the business field to be cleaned is cleaned by sql data cleaning script.It is thus, the scheme can realize the accurate cleaning of dirty data automatically by the constructed sql data cleaning script, so as to effectively reduce the time consumed by test personnel for cleaning dirty data.
Owner:WEBANK (CHINA)

Dirty tracking bit compression

PendingUS20260079840A1Memory systemsMemory hierarchyDirty data
A cache controller of a cache assigns a dirty tracking bit for each dirty byte of a cache line. Once a predetermined interval has elapsed without any accesses to the cache line or to a cache set that includes the cache line, the cache controller compresses contiguous dirty tracking bits for each portion of the cache line. Compressing the dirty tracking bits for contiguous dirty portions of the cache line allows the cache to store more dirty data using fewer dirty tracking bits, reducing area cost and bandwidth among levels of a memory hierarchy.
Owner:ADVANCED MICRO DEVICES INC

A data storage method, system, device, and storage medium

The application discloses a data storage method, comprising the following steps: in response to receiving a write IO, obtaining a current cache strategy; in response to the current cache strategy being a first cache strategy, judging whether there is unexpired dirty data in the current cache; in response to there being unexpired dirty data, copying the write IO to obtain two write IOs and judging whether the write IO hits the unexpired dirty data; in response to hitting the unexpired dirty data, directly writing one of the write IOs into a RAID and updating the hit unexpired dirty data by the other write IO; and in response to one of the write IOs being successfully written into the RAID, performing invalidation processing on the hit unexpired dirty data. The application further discloses a system, a computer device and a readable storage medium. When the cache strategy is switched from WRITE-BACK to WRITE-THROUGH, the scheme provided by the application does not cause permanent data loss.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

High-availability passenger flow data processing system

The invention discloses a high-availability passenger flow data processing system, and the system comprises a data collection and local caching module, a cloud data access and buffering module, a high-availability data cleaning process module, and a data storage and application module. According to the method, data are not lost, the accuracy, the consistency and the availability of the data are obviously improved, a solid foundation is provided for business decision making, the cloud cleaning service adopts a micro-service architecture, database decoupling and elastic telescopic design, a large-data-volume scene of a promotion activity can be easily dealt with, and the service experience is improved. Stable operation of the cleaning service and timeliness of data processing are guaranteed, various dirty data including repeated data, format errors, abnormal values and missing values can be effectively processed in the targeted cleaning stage, high-quality standardized data are output, and the system is easy to expand to adapt to business growth due to modularization and elastic expansion capacity of the cleaning service.
Owner:HANGZHOU HUMPBACK WHALE TECHNOLOGY CO LTD +1

Transaction detection method, device, equipment and storage medium

The application discloses an abnormality detection method, device and equipment and a storage medium, and comprises the following steps: obtaining people flow data of a target grid and preset reference data; determining first abnormality information of the target grid according to the people flow data and the preset reference data; in the case that the target grid is determined to be an abnormal grid according to the first abnormality information, determining an abnormal adjacent area corresponding to the target grid; determining second abnormality information according to the people flow data of the abnormal adjacent area; and determining actual abnormality information of the target grid according to the first abnormality information and the second abnormality information. The actual abnormality information of the target grid is determined through the first abnormality information of the target grid and the second abnormality information of the abnormal adjacent area, the technical problem of poor anti-interference capability of dirty data and data offset in the prior art is solved, and the accuracy of people flow abnormality detection is improved.
Owner:CHINA MOBILE GROUP JIANGSU +1

Power distribution network measurement data restoration method and system considering edge enhancement and multi-scale detail restoration

The invention discloses a power distribution network measurement data restoration method and system considering edge enhancement and multi-scale detail restoration, and relates to the technical field of power system data processing and industrial time series data governance. Real working condition abnormity and dirty data are distinguished, and a repair mask and morphological enhancement prior information are generated; a model is generated through WGAN-GP training with mask consistency and boundary constraint, and a preliminary repair result is output through latent variable inversion; then, repairing details are optimized through edge-oriented form remodeling and multi-scale residual error weighting; and finally, performing back-check iteration by adopting a homologous statistical standard. According to the scheme, the problems of excessive smoothness and edge fault of long block repair are effectively solved, the authenticity and reliability of repair data are improved, and accurate operation and maintenance requirements of power distribution network state monitoring, scheduling optimization and the like can be met.
Owner:ELECTRIC POWER RESEARCH INSTITUTE OF STATE GRID SHANDONG ELECTRIC POWER COMPANY

Excel data intelligent cleaning method and device

PendingCN122412764AOriginal dataDirty data
This application relates to the field of data processing technology, and in particular to an intelligent Excel data cleaning method and apparatus, comprising: preprocessing the original Excel data; based on the content characteristics of the preprocessed data, diverting it to a rule-based cleaning channel or an AI cleaning channel; in response to the preprocessed data being diverted to the rule-based cleaning channel, performing a preset rule-based cleaning operation on the preprocessed data; in response to the preprocessed data being diverted to the AI ​​cleaning channel, constructing task-specific prompts to invoke a large model, and performing semantic-level cleaning operations on the preprocessed data; performing result fusion and post-processing operations on the output results to obtain fused data; and processing the fused data according to a confidence score. Thus, simple data is diverted to rule-based processing, improving overall throughput and reducing economic costs. It can handle complex and dirty data while ensuring format consistency. The multi-layered quality self-checking mechanism can effectively screen out low-confidence results, prevent erroneous output, and improve result reliability.
Owner:LAUNCH TECH CO LTD

Multi-level cache implementation method and system suitable for nonvolatile storage data

The invention discloses a multi-level cache implementation method and system suitable for nonvolatile storage data, and relates to the technical field of data storage. The method is cooperatively executed based on a first-level cache managed by a hardware management module and a second-level cache managed by a software management module, and comprises the following steps: the hardware management module performs address analysis on a received host command and queries the state of a corresponding data block in the first-level cache; the states comprise invalid, in-loading, available, dirty data and in-unloading. And when the data block is in the invalid state, triggering a data loading operation from the nonvolatile memory or the second-level cache to the first-level cache. And when the data block in the dirty data state needs to be replaced from the first-level cache, triggering the operation of migrating the data block to the second-level cache for temporary storage. Through a software and hardware collaborative division architecture and a dirty data priority cache migration strategy, the data access delay is effectively reduced, and more stable overall performance is realized.
Owner:PENG TI STORAGE TECH (NANJING) CO LTD

Intelligent data governance platform and device based on multi-modal fusion and knowledge graph construction

PendingCN122654999ASolving the problem of multi-modal asynchronous spoofingAvoid computing power explosionTopology informationAlgorithm
The application discloses a multi-modal fusion and knowledge graph construction intelligent data governance platform and device, and relates to the technical field of data processing; through synchronous acquisition of time sequence sensing, image and text data, a half-life constant is given according to a physical decay rate, and a dynamic attenuation weight is calculated; the extracted heterogeneous features are mapped to a Poincare hyperbolic manifold space, the hyperbolic distance and time phase difference between nodes are calculated, cross-modal semantic tension is obtained, and a cause-effect tension graph is constructed in combination with a cause-effect contradiction threshold; conflict entanglement is calculated based on a local tension proportion, and logical isolation bubbles are used to block asynchronous fraudulent data; through counterfactual reasoning virtual erasing nodes, topological information entropy change values are calculated; for low-entropy change redundant data, downsampling and cold storage are performed, and for isolated high-entropy change key dirty data, a cause-effect generative adversarial network is used to reversely deduce missing features based on a health mode and to perform self-healing overwrite; the problems of multi-modal asynchronous fraud and data hoarding are solved.
Owner:HUBEI TIANCUN INFORMATION TECH CO LTD

Regression model-based target customer transaction probability prediction method and equipment

PendingCN121707735AFinanceOriginal dataInsurance life
The invention discloses a target customer transaction probability prediction method and device based on a regression model, and the method comprises the steps: S1, carrying out the dirty data elimination and unstructured data conversion of the original data of a customer of a life insurance company, and obtaining the standardized data; s2, dividing the standardized data into a low-net-value customer data source and a high-net-value customer data source based on the customer income level for subsequent prediction; s3, multiple machine learning models are selected as candidate models, a dual-drive strategy is adopted to screen customer transaction key features, and an initial prediction model is constructed; s4, training the initial prediction model, and determining an optimal parameter through regularization processing and cross validation to obtain an optimized regression prediction model; and S5, performing grouping verification and actual scene verification on the optimized regression prediction model, determining an optimal deal probability threshold, and establishing a long-acting iteration mechanism continuous optimization model at the same time.
Owner:CHINA LIFE INSURANCE CO LTD

Concurrent fill and byte merge

Systems and techniques for concurrently performing a fill and byte merge operation in a data processing system are described. An example technique includes receiving a memory access request from a user interface. A determination is made that the memory access request has encountered a cache miss within a cache directory in the computing system. In response to the determination, a fetch request is transmitted to an upper level cache within the computing system for a cache line associated with the memory access request. Dirty portions of the cache line are concurrently written and merged, based on the memory access request, with fill data of the cache line obtained from the upper level cache into a line buffer of a line engine within the computing system.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

A method for incrementally cleaning streaming data for a big data platform

PendingCN122346492AStreaming dataData stream
The present application relates to a kind of big data platform-oriented stream data incremental cleaning method, belong to data cleaning technical field, comprising the following steps: step 1: data stream is divided into sliding window according to the preset window time length, and adjacent relation graph is constructed according to leaf node co-occurrence relationship;Step 2: the data record in each near neighbor bucket is constructed field fingerprint string, and the data record that field fingerprint string is not identical with all leading mode fingerprints in leading mode fingerprint set is marked as dirty data mode seed;Step 3: with dirty data mode seed as the initial label of dirty data area, the data record of dirty data area is executed local cleaning operation according to original field unit and is recovered after inverse transformation and outputs cleaning result.The present application is applicable to real-time big data scene such as Internet of Things and financial transaction.

A data cleaning method based on a Bayesian formula, a terminal and a storage medium

The application discloses a data cleaning method based on a Bayesian formula, a terminal and a storage medium, and the method comprises the following steps: acquiring original data and pre-defined prior knowledge; constructing a Bayesian network and an association relationship according to the prior knowledge, performing causal reasoning on the original data according to the Bayesian network, and obtaining a probability score of the Bayesian network; determining an association relationship score of the original data according to the association relationship, and cleaning the original data according to the sum of the probability score and the association relationship score, to obtain a cleaned data version. The application combines the user knowledge prior information which is easy to master, the modeling capability of the Bayesian network on dirty data and the association relationship of mutual information existing in the data, performs scanning and cleaning on the original data, reduces the difficulty of data cleaning, and improves the accuracy and recall rate of data cleaning.
Owner:SHENZHEN UNIV

A standardized cleaning method and related apparatus for multi-source alarm data access.

This application discloses a standardized cleaning method and related apparatus for multi-source alarm data access, relating to the field of IT operations and maintenance technology in the financial industry. The method obtains user-configured synchronization task parameters, including task execution plans, multi-protocol compatible processing classes, parsing and mapping rules, and initiates a task to receive multi-source alarm data. Dirty data is filtered through processing class matching, parsed into indexable objects, mapped to a unified structure data, and then formatted (including type conversion, null value filling, integrity and consistency checks) to complete standardized cleaning. This method achieves multi-protocol compatibility and unified access to multi-source data without the need for customized scripts, reducing technical barriers and maintenance costs. Through data filtering and standardized processing, it improves data quality and operational efficiency, facilitates rapid response to critical alarms, and ensures the continuity of core financial businesses.
Owner:QIDIAN HAOHAN DATA TECH BEIJING CO LTD

Cache scheduling method and system for a file storage system

PendingCN122654086ADirty dataParallel computing
The application discloses a cache scheduling method and system of a file storage system, and relates to the technical field of file storage. The method comprises the following steps: receiving to-be-written data of a write object, determining data position information and object identification; determining a cache block identification based on a preset cache block size and the data position information, locating or allocating a target cache block and writing the to-be-written data; determining a target parallel writing unit based on the cache block identification, determining a target scheduling queue based on the object identification and associating the target cache block; polling multiple scheduling queues through the target parallel writing unit, writing to a back-end storage system when the cache block meets a data amount condition or a dirty data residence time condition, and switching to a next scheduling queue after writing one cache block; and recycling the cache block that has completed writing and meets a recycling condition. Thus, the cache writing delay of small data block parallel writing can be reduced, the delay stability of different write objects can be improved, and the write bandwidth is taken into account.
Owner:HUNAN TONGYOU FEIJI TECH CO LTD