An AI-assisted end-side dynamic data targeting processing method, system, device and medium for controllable data sharing

CN122570467BActive Publication Date: 2026-09-25CHENGDU UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611047333.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-25
Estimated Expiration
2046-07-15

AI Technical Summary

Technical Problem

[0004]本发明针对上述集中式加工的安全权属风险、治理冗余、适配性差、合规性不足的问题,一种面向可控数据共享的AI辅助端侧动态数据定向加工方法、系统、设备及介质;将AI智能治理能力下沉至数据源端侧:原始数据全程留存于权属用户本地,先经共享权限校验通过后,再由AI模块根据具体共享需求和数据质量情况,开展定向清洗、脱敏、格式适配、特征提取等加工处理;同时实时感知共享需求、数据状态变化,动态迭代加工规则,形成授权核验、需求解析、AI定向加工、动态调优、受控输出的闭环流程,全程数据不迁移、加工可管控、结果可追溯,既保障数据共享的安全性与可控性,又实现加工效率与适配度的双重提升

Benefits of technology

(1) 本发明原始数据端侧留存不出域,加工前权限核验、加工中全程管控、加工后审核输出,权属用户牢牢掌控加工全流程,从根源规避数据泄露风险。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122570467B_ABST
    Figure CN122570467B_ABST
Patent Text Reader

Abstract

The present application relates to big data governance and data security sharing technical field, specifically, it relates to an AI assisted end side dynamic data directional processing method, system, equipment and medium for controllable data sharing; Sink AI intelligent governance capability to data source end side: original data is kept in the whole process in the local of the right user, after passing through the sharing permission verification, then according to the specific sharing demand and data quality situation, directional cleaning, desensitization, format adaptation, feature extraction and other processing are carried out by AI module; At the same time, the sharing demand and data state change are perceived in real time, the processing rules are dynamically iterated, the closed loop process of authorized verification, demand analysis, AI directional processing, dynamic optimization and controlled output is formed, the data is not migrated in the whole process, the processing can be controlled, the result can be traced, which not only guarantees the safety and controllability of data sharing, but also realizes the double improvement of processing efficiency and adaptation degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data governance and data security sharing technology, specifically to an AI-assisted edge-side dynamic data-oriented processing method, system, device, and medium for controllable data sharing. Background Technology

[0002] Currently, multi-source data processing and governance mainly adopts a "platform-centralized processing model." The core process is as follows: all raw data from various data sources is collected and transmitted to a unified data platform. The platform centrally performs data cleaning, format conversion, and desensitization. The processed data is then shared and accessed as needed. Although some existing technologies have introduced AI-assisted processing to optimize processing efficiency, they have not deviated from the core paradigm of "centralized raw data, unified platform management, and pre-batch processing," and have not deeply integrated data sharing permission management with the processing process.

[0003] There are also a few edge-side governance technologies that only achieve basic data preprocessing. The processing rules are fixed and lack dynamic adaptation capabilities, which cannot meet the refined and customized needs of controllable data sharing. As a result, there are problems such as disconnect between processing and sharing, lagging access control, and weak security control. Summary of the Invention

[0004] This invention addresses the security risks, governance redundancy, poor adaptability, and insufficient compliance issues associated with centralized data processing. It provides an AI-assisted, edge-side dynamic data-oriented processing method, system, device, and medium for controllable data sharing. The invention decentralizes AI intelligent governance capabilities to the data source side: raw data is stored locally by the user with the rights throughout the process. After passing sharing permission verification, the AI ​​module performs targeted cleaning, desensitization, format adaptation, and feature extraction based on specific sharing needs and data quality. Simultaneously, it senses changes in sharing needs and data status in real time, dynamically iterating processing rules to form a closed-loop process of authorization verification, demand analysis, AI-oriented processing, dynamic optimization, and controlled output. The entire process involves no data migration, controllable processing, and traceable results, ensuring both the security and controllability of data sharing while simultaneously improving processing efficiency and adaptability.

[0005] The specific implementation details of this invention are as follows: A method for targeted processing of dynamic data on an AI-assisted edge device for controllable data sharing includes the following steps: Step S1: Obtain metadata of multi-source heterogeneous data, label the sensitivity level according to the preset sensitive word library, and store the labeling results in the local permission metadata database and verify them; Step S2: Train a lightweight language model based on the obtained historical shared application samples to obtain processing rules bound to the current session ID; Step S3: Based on the acquired raw data records and processing rules, call the rule parsing engine to generate an in-situ processing instruction sequence, and process the raw data records field by field in-situ. Step S4: Generate an intermediate quality report based on the batch quantity of in-situ processing, and call the triggerOptimizer() function to trigger the tuner to optimize and obtain new processing rules; Step S5: Obtain a temporary processing result object according to the new processing rules, generate a standardized processing dataset according to the review decision results, and call the encryptAndSend() function to transmit it to the target server via HTTPS POST request in an encrypted manner.

[0006] To better realize the present invention, step S1 further includes the following steps: Step S11: Read the metadata of the multi-source heterogeneous data obtained from the local machine, and mark the sensitivity level and sharing permission threshold according to the preset sensitive word library, and store the marking results in the local permission metadata database; Step S12: Based on the obtained shared processing request, call the request parsing function to extract the structured fields and save them to the current session context; Step S13: Query the local permission metadata database based on the extracted structured fields to determine whether access is authorized, and determine compliance based on the pre-configured compliance rule base; if unauthorized or non-compliant, return false, generate a denial message and log it; if the verification passes, return true and write the verification result to the current session context and authorization timestamp.

[0007] To better realize the present invention, step S2 further includes the following steps: Step S21: Train a lightweight language model based on the obtained historical shared application samples and manually annotated typical request statements to obtain the trained lightweight language model. Step S22: Obtain the application text from the current session context, and call the parseIntent() function and the trained lightweight language model to obtain the processing dimension parameter object; Step S23: Based on the processing dimension parameter object, the obtained data quality report, and the sensitive level tags obtained from the local permission metadata database, call the rule generator to generate processing rules bound to the current session ID.

[0008] To better realize the present invention, step S3 further includes the following steps: Step S31: Based on the table name and field list of the structured fields, use an SQL query to obtain the original data records and store them in the memory buffer; Step S32: Based on the original data records, call the rule parsing engine, traverse the rule entries of the processing rules, extract the operation type and parameter dictionary, and generate the in-situ processing instruction sequence; Step S33: Process the original data record field by field in situ according to the in-situ processing instruction sequence.

[0009] To better realize the present invention, step S4 further includes the following steps: Step S41: Generate an intermediate quality report based on the batch quantity of in-situ processing; Step S42: Based on the processing rules and intermediate quality reports, call the evaluateQuality() function to obtain the actual quality indicators and expected quality indicators; Step S43: Calculate the deviation based on the actual quality indicators and expected quality indicators, and call the triggerOptimizer() function according to the deviation type to trigger the tuner to optimize and obtain new processing rules.

[0010] To better realize the present invention, step S43 further includes the following steps: Step S431: Locate the corresponding adjustable parameters and adjustment rules based on the deviation type; Step S432: Modify the relevant parameters and operation type of the current processing rule according to the found adjustable parameters and adjustment rules to obtain the optimized rule fragment; Step S433: Merge the optimized rule fragment with the current processing rule to obtain a new processing rule.

[0011] To better realize the present invention, step S5 further includes the following steps: Step S51: According to the new processing rules, the columnar storage field array, field name, and session identifier are encapsulated into a temporary processing result object, and the pushToAudit() function is called to push the temporary processing result object to the audit queue through the local secure channel; Step S52: Based on the user's audit decision result (string type) and the current session context, call the handleAuditDecision() function to obtain the audit result; if the audit result is rejection, discard the temporary processing result object and send a retry signal to trigger reprocessing; if the audit result is approval, proceed to step S53. Step S53: Call the generateStandardDataset() function to serialize the approved temporary processing result object into a data file in the set format, and generate a data dictionary file according to the meaning of the fields, the desensitization method, and the data range of the temporary processing result object; Step S54: Combine the data file and the data dictionary file into a compressed package, call the encryptAndSend() function to generate a symmetric key, and call the AES-256-GCM algorithm to encrypt the compressed package to obtain a ciphertext package; Step S55: Combine the encrypted symmetric key and ciphertext packet, send them to the target server via an HTTPS POST request, and destroy the local temporary file and the key in memory.

[0012] Based on the aforementioned AI-assisted edge-side dynamic data orientation processing method for controllable data sharing, and to better realize this invention, a further AI-assisted edge-side dynamic data orientation processing system for controllable data sharing is proposed to execute the aforementioned AI-assisted edge-side dynamic data orientation processing method for controllable data sharing; it includes a shared permission verification unit, an orientation processing rule generation unit, an in-situ orientation processing execution unit, an optimization unit, and a shared output unit; The shared permission verification unit is used to obtain metadata of multi-source heterogeneous data, mark the sensitivity level according to the preset sensitive word library, and store the marking results in the local permission metadata database and verify them. The targeted processing rule generation unit is used to train a lightweight language model based on the acquired historical shared application samples to obtain processing rules bound to the current session ID; The in-situ directional processing execution unit is used to call the rule parsing engine to generate an in-situ processing instruction sequence based on the acquired original data records and processing rules, and process the original data records field by field in-situ. The optimization unit is used to generate an intermediate quality report based on the batch quantity of in-situ processing, and call the triggerOptimizer() function to trigger the optimizer to optimize and obtain new processing rules; The shared output unit is used to obtain a temporary processing result object according to the new processing rules, generate a standardized processing dataset according to the review decision results, and call the encryptAndSend() function to transmit it to the target server via an HTTPS POST request in an encrypted manner.

[0013] Based on the above-mentioned AI-assisted edge-side dynamic data orientation processing method for controllable data sharing, in order to better realize the present invention, an electronic device is further proposed, including a memory and a processor; the memory stores a computer program; when the computer program is executed on the processor, the above-mentioned AI-assisted edge-side dynamic data orientation processing method for controllable data sharing is realized.

[0014] Based on the above-mentioned AI-assisted edge-side dynamic data orientation processing method for controllable data sharing, in order to better realize the present invention, a computer-readable storage medium is further proposed, wherein computer instructions are stored on the computer-readable storage medium; when the computer instructions are executed on the above-mentioned electronic device, the above-mentioned AI-assisted edge-side dynamic data orientation processing method for controllable data sharing is realized.

[0015] The present invention has the following beneficial effects: (1) The original data of this invention is stored on the terminal side without leaving the domain. Before processing, the user has permission verification, during processing, full control, and after processing, the user has audit and output. The user has firm control over the entire processing process, thus avoiding the risk of data leakage from the source.

[0016] (2) This invention is based on on-demand targeted processing of shared needs, abandons full-scale redundant management, greatly reduces computing power and resource consumption, reduces management costs by about 80%, and significantly improves processing efficiency compared with the traditional model.

[0017] (3) The AI ​​adaptive parsing requirements and dynamic iterative processing rules of this invention are accurately adapted to multiple scenarios and differentiated sharing requirements. It has strong compatibility and adaptability with multi-source heterogeneous data, and the processed data fits the requirements of shared applications.

[0018] (4) The present invention encrypts and preserves evidence throughout the entire process, making the processing and sharing links traceable and clear in the division of rights and responsibilities, effectively solving the problems of disputes and accountability in data sharing and processing.

[0019] (5) This invention is adaptable to various multi-source heterogeneous data and various controllable sharing scenarios such as government affairs, enterprises, and medical care. It supports the expansion of processing rules as needed and can be seamlessly connected to the full-domain data sharing platform, making it highly practical. Attached Figure Description

[0020] Figure 1 This is a schematic flowchart of the AI-assisted edge-side dynamic data orientation processing method for controllable data sharing provided by the present invention.

[0021] Figure 2 This is a schematic diagram of the end-side AI directional processing module architecture provided by the present invention.

[0022] Figure 3 The diagram shows the hierarchical structure of the AI-assisted processing module provided by this invention. Detailed Implementation

[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments, and therefore should not be regarded as a limitation on the scope of protection. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set up," "connected," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0025] Example 1: This embodiment proposes an AI-assisted edge-side dynamic data-oriented processing method for controllable data sharing, specifically including the following steps: Specifically, the following steps are included: Step S1: Obtain metadata of multi-source heterogeneous data, label the sensitivity level according to the preset sensitive word library, and store the labeling results in the local permission metadata database and verify them; Step S1 specifically includes the following steps: Step S11: Read the metadata of the multi-source heterogeneous data obtained from the local machine, and mark the sensitivity level and sharing permission threshold according to the preset sensitive word library, and store the marking results in the local permission metadata database; Step S12: Based on the obtained shared processing request, call the request parsing function to extract the structured fields and save them to the current session context; Step S13: Query the local permission metadata database based on the extracted structured fields to determine whether access is authorized, and determine compliance based on the pre-configured compliance rule base; if unauthorized or non-compliant, return false, generate a denial message and log it; if the verification passes, return true and write the verification result to the current session context and authorization timestamp.

[0026] Step S2: Train a lightweight language model based on the obtained historical shared application samples to obtain processing rules bound to the current session ID; Step S2 specifically includes the following steps: Step S21: Train a lightweight language model based on the obtained historical shared application samples and manually annotated typical request statements to obtain the trained lightweight language model; Step S22: Obtain the application text from the current session context, and call the parseIntent() function and the trained lightweight language model to obtain the processing dimension parameter object; Step S23: Based on the processing dimension parameter object, the obtained data quality report, and the sensitive level tags obtained from the local permission metadata database, call the rule generator to generate processing rules bound to the current session ID.

[0027] Step S3: Based on the acquired raw data records and processing rules, call the rule parsing engine to generate an in-situ processing instruction sequence, and process the raw data records field by field in-situ. Step S3 specifically includes the following steps: Step S31: Based on the table name and field list in data_scope, use an SQL query or file reading interface to obtain the original data records and store them in the memory buffer; Step S32: Parse the customized processing rules, match the corresponding rule entries according to the current field name, extract the operation type and parameters, and generate a processing instruction sequence; Step S33: According to the processing instruction sequence, call the corresponding processing function for the corresponding field in the original data record to perform in-situ processing operation.

[0028] Step S4: Generate an intermediate quality report based on the batch quantity of in-situ processing, and call the triggerOptimizer() function to trigger the tuner to optimize and obtain new processing rules; Step S4 specifically includes the following steps: Step S41: Generate an intermediate quality report based on the batch quantity of in-situ processing; Step S42: Based on the processing rules and intermediate quality reports, obtain the actual quality indicators and expected quality indicators; Step S43: Calculate the deviation based on the actual quality indicators and expected quality indicators, and call the triggerOptimizer() function according to the deviation type to trigger the tuner to optimize and obtain new processing rules.

[0029] Step S43 specifically includes the following steps: Step S431: Locate the corresponding adjustable parameters and adjustment rules based on the deviation type; Step S432: Modify the relevant parameters and operation type of the current processing rule according to the found adjustable parameters and adjustment rules to obtain the optimized rule fragment; Step S433: Merge the optimized rule fragment with the current processing rule to obtain a new processing rule.

[0030] Step S5: Obtain a temporary processing result object according to the new processing rules, generate a standardized processing dataset according to the review decision results, and call the encryptAndSend() function to transmit it to requester_info.receiver_url via HTTPS POST request in encrypted form.

[0031] Step S5 specifically includes the following steps: Step S51: According to the new processing rules, the columnar storage field array, field name, and session identifier are encapsulated into a temporary processing result object, and the pushToAudit() function is called to push the temporary processing result object to the audit queue through the local secure channel; Step S52: Based on the user's audit decision result (string type) and the current session context, call the handleAuditDecision() function to obtain the audit result; if the audit result is rejection, discard the temporary processing result object and send a retry signal to trigger reprocessing; if the audit result is approval, proceed to step S53. Step S53: Call the generateStandardDataset() function to serialize the approved temporary processing result object into a data file in the set format and generate a data dictionary file; Step S54: Combine the data file and dictionary file into a compressed package, call the encryptAndSend() function to generate a symmetric key, and call the AES-256-GCM algorithm to encrypt the compressed package to obtain a ciphertext package; Step S55: Combine the encrypted symmetric key and ciphertext packet, send them to requester_info.receiver_url via an HTTPS POST request, and destroy the local temporary file and the key in memory.

[0032] Working principle: This embodiment extends AI intelligent governance capabilities to the data source side: the original data is stored locally on the user's local machine. After passing the sharing permission verification, the AI ​​module performs targeted cleaning, desensitization, format adaptation, feature extraction, and other processing based on specific sharing needs and data quality. At the same time, it senses changes in sharing needs and data status in real time, dynamically iterates processing rules, and forms a closed-loop process of "authorization verification - demand analysis - AI targeted processing - dynamic optimization - controlled output". The entire process does not involve data migration, the processing is controllable, and the results are traceable, which not only ensures the security and controllability of data sharing, but also achieves a dual improvement in processing efficiency and adaptability.

[0033] Example 2: This embodiment is based on the above embodiment 1, such as... Figure 1 , Figure 2 , Figure 3 As shown, a specific embodiment will be described in detail.

[0034] Step S1: Preliminary authorization confirmation and sharing permission verification; 1. End-side system initialization and data weighting classification The endpoint data targeting and processing system is deployed on the user's end. The system first loads the data ownership classification configuration: for multi-source heterogeneous data (such as database tables, files, and data streams) stored locally by the user, it reads their metadata (field names, data types, source identifiers), and based on a preset sensitive word library or classification model, labels each data object with a sensitivity level (e.g., L1 public, L2 internal, L3 sensitive, L4 highly sensitive) and sharing permission thresholds (e.g., access only allowed for specific roles, requiring approval, etc.). The labeling results are stored in the local permission metadata database.

[0035] 2. Receive and parse shared processing requests Data requesters initiate data sharing and processing applications through the sharing platform. The application format is structured data (JSON / XML) and must include at least the following fields: requester_id: A unique identifier for the requester; data_scope: The scope of the requested data (such as table name, field list, time range); purpose: Description of intended use (text, such as "for training regional disease statistical models"); format_requirements: Required output format (e.g., CSV, JSON, Parquet); privacy_level: The desired level of desensitization (optional "low / medium / high").

[0036] The client-side system receives the request through the shared platform interface and calls the request parsing function `parseRequest(request_json)` to extract the aforementioned fields and store them in the current session context `session_context`. `session_context` is a dictionary / object used to transmit and store context information for this shared session throughout the entire process, including but not limited to: requester identifier (`requester_id`), data scope (`data_scope`), purpose description (`purpose`), output format (`format_requirements`), privacy level (`privacy_level`), permission verification result (`permission_granted`), and authorization timestamp.

[0037] 3. Execution of permission verification The system calls the verifyPermission(session_context) function of the permission verification module, and performs the following sub-steps: Based on the requester_id, query the local permissions database (i.e., the database stored in step 1 or call the rule engine preset by the user with the permissions) to confirm whether the requester is authorized to access the requested data_scope. Based on the `purpose` field, semantic matching is performed using a pre-configured compliance rule base to determine whether the purpose is compliant. This compliance rule base is a set of rules pre-built according to the Data Security Law, industry standards, and user-defined policies, such as "medical data may not be used for marketing" and "de-identified data may be used for research purposes." If the requester is not authorized or the purpose is non-compliant, the function returns false, the system generates a rejection message and logs it, and the process terminates.

[0038] If the verification passes, the function returns true and writes the verification result to session_context with permission_granted=true and the authorization timestamp.

[0039] 4. Deployment instructions for AI processing nodes The AI ​​processing node, or AI-assisted processing module in the edge-side data-oriented processing system, runs as a resident process or container on the user's edge device (such as a server or edge gateway). This node starts during system initialization and loads a lightweight local language model (see step S2). AI-assisted processing module ( Figure 3 The hierarchy shown, along with the permission verification module and the in-situ processing execution module, are parallel functional modules within the same system. They interact with each other through inter-process communication (IPC) or API calls. Specifically, the permission verification module is responsible for authorization verification in step 1.3, the in-situ processing execution module is responsible for data processing in step 3, and the AI-assisted processing module is responsible for parsing, rule generation, and optimization in steps 2 and 4.

[0040] Step S2: AI-assisted demand analysis and targeted processing rule generation 1. Construction and Deployment of Lightweight Local Language Models This embodiment uses a lightweight local language model based on open-source pre-trained models (such as BERT-tiny, DistilBERT, TinyGPT) that undergo knowledge distillation and INT8 quantization compression, resulting in a parameter size of less than 50M and a memory footprint of less than 200MB, enabling real-time inference on the edge CPU. The specific construction method is as follows: Fine-tuning a lightweight model from a general corpus to enable it to identify entities and intents in the data-sharing domain (sensitization strength, format, geographic precision); The training dataset consists of pre-collected historical shared application samples or manually annotated typical request statements (e.g., "Statistical information on the age of patients in Beijing is needed, CSV format, medium level of anonymization"). These samples are independent of the weighting and classification in step 1 and are prepared in advance. The training objectives are sequence labeling and intent classification, and the output is structured processing dimension parameters.

[0041] This model, as a core component of the AI ​​semantic parsing unit, runs offline on the device and does not require an internet connection.

[0042] 2. Activate AI processing module and semantic parsing After successful permission verification, the client-side system calls `activateAIModule(session_context)` to activate the AI-assisted processing module. This module first calls the semantic parsing unit to read the pre-stored request text (i.e., the original content of the `purpose` field in step 1) from `session_context`, denoted as `request_text = session_context.purpose`. Then, it executes the `parseIntent(request_text)` function, which uses a loaded lightweight language model to perform entity recognition and intent parsing on `request_text`, outputting a processing dimension parameter, for example: json { "cleaning_granularity": "row_level", "desensitization_strength": "high", "output_format": "csv", "geographic_precision": "city_level" } If a parameter is not specified in the application, the default value will be used (e.g., the desensitization strength is "medium" by default).

[0043] 3. Input and rule generation of the rule generator The rule generator, as an independent decision-making unit within the AI ​​module (which can be implemented by a lightweight rule engine or driven by the same language model), receives the following three types of input: The processing dimension parameter object output by the semantic parsing unit; The source data quality report is obtained in real time by the data quality awareness interface: This interface calls the getDataQualityReport(data_scope) function to scan the requested data range, count indicators such as field missing rate, outlier ratio, and data type consistency, and generate a report (e.g., {"name": {"missing_rate":0.02}, "address": {"missing_rate":0.15}}). Sensitive level tags (e.g., {"address":"L4", "age":"L2"}) read from the local permissions metadata database.

[0044] Based on the above input, the rule generator selects and instantiates rules from a predefined atomic rule library. The atomic rule library is a configuration table, and each rule contains: applicable data type, operation type, and parameter template. For example: hash_replace:{field} → SHA256({field}) mask_phone: Keep the first three and last four digits, replace the middle digits with ****. generalize_to_city: Maps detailed addresses to city levels. add_laplace_noise: Adds Laplace(0, 1 / ε) noise to the numeric field. The rule generator executes the `generateRules(dimensions, qualityReport, sensitivity)` function. Where: dimensionParams are the processing dimension parameters output in step S2; qualityReport is a data quality report returned by the data quality awareness interface; sensitivityLabels are sensitivity level labels read from the local permissions metadata database.

[0045] The function execution logic is as follows: Based on the sensitivity level: If the field is L4, a strong desensitization rule (such as hash_replace or generalization) must be selected. Based on the desensitization intensity parameter: High → large noise scale or high generalization level; Medium → mask or moderate noise; Low → slight perturbation; According to the data quality report: if the missing rate of a certain field is >30%, then the completion rule should be "mean completion" instead of "model prediction"; Finally, a set of customized processing rules uniquely bound to the current session ID (session_id) is generated and stored in JSON format, for example: json { "session_id": "req_20250601_001", "rules": [ {"field": "name", "action": "hash_replace", "params": {"algorithm": "SHA256"}}, {"field": "address", "action": "generalize_to_city"}, {"field": "income", "action": "add_laplace_noise", "params": {"epsilon": 0.5}} ] } Step S3: In-situ orientation machining of the end side is performed; 1. Execution environment and invocation relationship The on-site processing execution module is an independent data processing engine that interacts with the AI-assisted processing module via a message queue. After the AI ​​module generates customized rules, it calls `submitRulesForExecution(rules, session_context)` to submit the customized processing rules generated in step S2 to the on-site processing execution module. Here, `rules` represents the customized processing rules output in step S2 (in JSON format, containing `session_id` and an array of specific rules); `session_context` is the session context object maintained since step S1.

[0046] 2. Acquisition of raw data Based on the table names and field lists in data_scope, raw data records are retrieved using SQL queries (SELECT fields FROM table WHERE conditions) or file reading interfaces (such as Pandas reading CSV), and stored one record at a time in a memory buffer, denoted as raw_record. The raw data refers to multi-source heterogeneous data (such as records in database tables, file content, etc.) stored on the user-side data ownership side, awaiting processing, and without any anonymization or cleaning.

[0047] 3. Rule parsing and processing instruction distribution The in-situ processing execution module embeds a rule parsing engine to parse the rules: For each raw_record to be processed, iterate through the rule entries in the rules array; Match the corresponding rule based on the current field name (e.g., field="name"); Extract the action (operation type, such as hash_replace) and params (parameter dictionary, such as {"algorithm":"SHA256"}) from the rule. Generate a sequence of processing instructions, in the form of {"field":"name", "action":"hash_replace", "params":{"algorithm":"SHA256"}}, to guide subsequent steps on which operation to perform on a specific field.

[0048] 4. Perform processing field by field according to the processing instruction sequence. The in-situ processing execution module calls the corresponding processing function for each field in raw_record according to the processing instruction sequence distributed in step 3. The following provides at least one specific implementation method for each processing operation (these methods are all known technologies, but this invention combines them and integrates them into a controllable shared architecture on the edge side): Data extraction: triggered by action="extract", based on the table name and field list in data_scope; Noise filtering: Triggered by action="noise_filter", it uses sliding window midpoint filtering for numeric fields; Redundant / erroneous data cleansing: triggered by action="deduplicate" or action="clean"; Intelligent data completion for missing data: triggered by action="impute", for numeric fields with a missing rate of <5%; Hash replacement: triggered by action="hash_replace", which calls hashlib.sha256(field_value.encode()).hexdigest(); Mask: Triggered by action="mask_phone", such as phone number "13812345678" → "138****5678"; Generalization: triggered by action="generalize_to_city", such as an exact address; Add noise: triggered by action="add_laplace_noise", which calls numpy.random.laplace(loc=0, scale=1 / epsilon); Data encryption: triggered by action="encrypt", using AES-256-GCM encryption; Unifying the format of heterogeneous data from multiple sources: triggered by action="format_unify", data from different sources is parsed into a unified row-based in-memory table; Business feature extraction: Triggered by action="feature_extract", extracts statistical features; All processing operations are completed in the local memory on the edge, and no network transmission occurs during processing. The processing execution module records the log of each operation (operation type, time, data volume) and stores it in a local log file.

[0049] Step S4: Dynamic feedback and processing strategy optimization; 1. Collection of data quality verification results; During the processing, the in-situ processing module generates an intermediate quality report, denoted as batch_quality, in batches (i.e., after processing a certain number of records in step 3, such as 1000 records). This report includes: The uniqueness ratio of each field after desensitization in the current batch (the duplication rate of desensitized values ​​should not be too high); Format compliance (e.g., whether the output field types are consistent with the rule requirements); Missing rate after completion.

[0050] These reports are pushed in real time to the strategy evaluation unit in the AI-assisted processing module via the callback interface onBatchComplete(batch_quality). This callback interface is pre-registered by the AI-assisted processing module and is called by the in-situ processing execution module after each batch is completed.

[0051] 2. Strategy Evaluation and Deviation Detection expected: Expected quality indicator objects, theoretical lower limits calculated based on current rule parameters, including expected_uniqueness (e.g., ≥0.85), expected_compliance (e.g., ≥0.95), and expected_missing_rate (e.g., ≤10%). currentRules: The set of processing rules currently in use (JSON format).

[0052] This function performs the following item-by-item comparison and deviation determination logic: (1) Determination of uniqueness deviation: compare actual.uniqueness_ratio < expected.expected_uniqueness, that is, whether the actual uniqueness ratio is lower than the lower expected limit (0.85). If yes, it is determined that there is a "too low uniqueness" deviation.

[0053] (2) Determination of format compliance deviation: compare actual.format_compliance < expected.expected_compliance, that is, whether the actual format compliance is lower than the lower expected limit (0.95). If yes, it is determined that there is a "non-compliant format" deviation.

[0054] (3) Determination of imputation effect deviation: compare actual.missing_rate_after_imputation > expected.expected_missing_rate, that is, whether the missing rate after imputation is still higher than the upper expected limit (10%). If yes, it is determined that there is an "imputation failure" deviation.

[0055] All the above comparisons are numerical comparisons, and the deviation determination results are calculated as shown in Table 1 below; Table 1 Comparison table of deviation determination results

[0056] The function returns a deviation determination result object, denoted as deviation_result, which includes: has_deviation: a Boolean value indicating whether there is any deviation; deviation_type_list: a string array that lists all triggered deviation types (e.g., ["uniqueness_too_low", "imputation_failure"]); deviation_details: detailed description of the specific values and thresholds of each deviation.

[0057] If has_deviation == true, the strategy evaluation unit uses the first deviation type in deviation_type_list and currentRules as parameters to call the triggerOptimizer function in step S4.3 to start the optimizer.

[0058] Meanwhile, the strategy evaluation unit also listens to the message channel of the shared platform. If the requester sends a "parameter change request" (such as {"new_strength":"higher"}) via API midway, it will also actively generate a deviation of type "user_request_change" and trigger optimization.

[0059] 3. Specific operation of the rule parameter tuner The policy evaluation unit calls `triggerOptimizer(deviation_type, current_rules)` to start the rule parameter tuner. Wherein: deviation_type is the deviation type identified in step S4, with values ​​such as "low uniqueness", "incompatible format", "completeness failure", etc., which are used for indexing optimization strategy tables; current_rules is the set of processing rules that are currently in effect (JSON format, with the same structure as the customized processing rules output in step S2; if optimization has been performed previously, it is the updated set of rules).

[0060] The tuner is a decision engine based on a pre-defined optimization strategy table. The optimization strategy table is shown in Table 2 below; Table 2 Optimization Strategy Table

[0061] The tuner performs the following steps: (1) Find the corresponding adjustable parameters and adjustment rules based on the deviation_type; (2) Modify the relevant parameter values ​​or operation types in the rule set; (3) Generate a new optimized rule fragment and merge it with the current rule set (i.e., the currentRules before optimization) to form updated_rules; (4) Resubmit the updated rules to the in-situ processing execution module to continue processing from the current batch (or roll back the previous batch).

[0062] For example, if the uniqueness of the address field is detected to be too low after anonymization (a large number of addresses are generalized to the same city), the tuner will adjust the generalization level from "street level" to "district level" and then reprocess the address field. The entire process is completed automatically without manual intervention.

[0063] Step S5: Processing result review and controlled shared output 1. Generation and delivery of processing results Once all data processing is complete (i.e., step S3 has finished processing all data and step S4 has not triggered further optimization), the in-situ processing module encapsulates the processed data (still a row-based memory table) in memory into a temporary processing result object, denoted as result_object. This object contains: data_content: A columnar array of fields; metadata: field name, data type, de-identification method, number of data rows, and generated timestamp; session_id: Session identifier.

[0064] Then, it can be used directly in subsequent calls to pushToAudit(result_object, session_context). The system call pushToAudit(result_object, session_context) pushes the result object to the audit output module's queue for review via a local secure channel (such as a Unix Domain Socket). Preview information (the first 20 lines of data and a statistical summary) is then displayed on the user interface (such as an administration console) for authorized users to review.

[0065] 2. Review and Decision Users with rights can click "Approve" or "Reject" on the interface. The audit output module calls handleAuditDecision(decision, session_context): The decision is a string representing the user's review decision, with values ​​of "approved" or "rejected". session_context is the session context object maintained starting from step S1.

[0066] The function execution logic is as follows: If decision == "rejected": the module discards the result object and sends a retry signal to the AI-assisted processing module to trigger reprocessing (rules can be modified or parameters adjusted). If decision == "approved": proceed to step 5.3, the output stage.

[0067] 3. Generate a standardized processing dataset After the review is approved, the review output module calls generateStandardDataset(result_object, format_requirements): Serialize the data in the in-memory table into the specified format according to the format_requirements in the application: CSV: Write to a file, fields are separated by commas, character set UTF-8; JSON: The output is a JSON array, with each record being an object; Parquet: Compresses storage using the Parquet library.

[0068] At the same time, a data dictionary file (YAML format) is generated, which describes the meaning of each field, the method of desensitization, and the data range.

[0069] The data files and dictionary files are packaged into a ZIP archive, and the SHA256 hash value of the entire archive is calculated and stored in the log.

[0070] 4. Encrypted transmission The audit output module calls `encryptAndSend(package, requester_info)`. The meanings of the parameters are as follows: package: The ZIP compressed package generated in step S5.3 contains standardized processing dataset files (such as .csv, .json or .parquet format) and data dictionary files (.yaml format). requester_info: A requester information object extracted from session_context, containing at least the following fields: receiver_url: The data receiving address specified by the requester, i.e., the HTTPS service endpoint URL that the requester has registered in advance on the sharing platform or provided in this application, used to receive the processed data packets; public_key: The RSA public key of the requester (used to encrypt the AES symmetric key); requester_id: A unique identifier for the requester (used for logging and tracing).

[0071] The function execution steps are as follows: (1) Generate a random AES-256 symmetric key; (2) Use the requester's public key (requester_info.public_key) to encrypt the AES key using the RSA algorithm to obtain the encrypted symmetric key ciphertext; (3) Encrypt the package data packet using the AES-256-GCM algorithm to obtain the encrypted data packet; (4) Combine the encrypted symmetric key ciphertext from step (2) with the data ciphertext from step (3) into a transmission payload object (JSON format), which contains two fields: encrypted_key and encrypted_data; (5) Send the payload object to requester_info.receiver_url via an HTTPS POST request. The HTTPS POST request refers to a network request submitted to the server using the POST method of the Hypertext Transfer Security Protocol (HTTPS), placing data in the request body. In this embodiment, this request is used to securely transmit the processed dataset from the data ownership user's end to the receiving server specified by the requester; (6) After the transmission is completed, destroy the local temporary file and the key in memory to ensure that no sensitive data remains.

[0072] Step S6: Full-process log storage and traceability control; The system generates structured log entries at each step (permission verification, semantic parsing, rule generation, processing and execution, batch optimization, review, and output), including timestamps, operation types, input parameter hashes, output result hashes, and participating module identifiers. Logs are stored in a local encrypted database (SQLite with AES) using append-only writing, and a Merkle tree hash chain is periodically generated to prevent tampering. The system supports reverse lookup of the entire process by session ID.

[0073] Working principle: This embodiment relies on mature and general technologies such as lightweight AI large model on the edge, data ownership classification and fine-grained permission control, dynamic rule iteration, edge data processing and encrypted computing. It has no special hardware dependency and is compatible with various conventional data source terminals. People skilled in the art can fully reproduce and implement it by following the above steps and combining conventional programming and AI inference methods. The solution fully meets the requirements of controllable data sharing and takes into account technical feasibility, security compliance and ease of implementation.

[0074] The other parts of this embodiment are the same as those in Embodiment 1 above, so they will not be described again.

[0075] Example 3: Based on any one of Embodiments 1-2 above, this embodiment proposes an AI-assisted edge-side dynamic data orientation processing system for controllable data sharing, used to execute the above-mentioned AI-assisted edge-side dynamic data orientation processing method for controllable data sharing; including a shared permission verification unit, an orientation processing rule generation unit, an in-situ orientation processing execution unit, an optimization unit, and a shared output unit; The shared permission verification unit is used to obtain metadata of multi-source heterogeneous data, mark the sensitivity level according to the preset sensitive word library, and store the marking results in the local permission metadata database and verify them. The targeted processing rule generation unit is used to train a lightweight language model based on the acquired historical shared application samples to obtain processing rules bound to the current session ID; The in-situ directional processing execution unit is used to call the rule parsing engine to generate an in-situ processing instruction sequence based on the acquired original data records and processing rules, and process the original data records field by field in-situ. The optimization unit is used to generate an intermediate quality report based on the batch quantity of in-situ processing, and call the triggerOptimizer() function to trigger the optimizer to optimize and obtain new processing rules; The shared output unit is used to obtain a temporary processing result object according to the new processing rules, generate a standardized processing dataset according to the review decision results, and call the encryptAndSend() function to transmit it to requester_info.receiver_url via HTTPS POST request.

[0076] This embodiment also proposes an electronic device, including a memory and a processor; the memory stores a computer program; when the computer program is executed on the processor, it implements the above-described AI-assisted edge-side dynamic data orientation processing method for controllable data sharing.

[0077] This embodiment also proposes a computer-readable storage medium storing computer instructions; when the computer instructions are executed on the aforementioned electronic device, the aforementioned AI-assisted edge-side dynamic data orientation processing method for controllable data sharing is realized.

[0078] The other parts of this embodiment are the same as any one of the above embodiments 1-2, so they will not be described again.

[0079] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Any simple modifications or equivalent changes made to the above embodiments based on the technical essence of the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for targeted processing of dynamic data on an AI-assisted edge device for controllable data sharing, characterized in that, Specifically, the following steps are included: Step S1: Obtain metadata of multi-source heterogeneous data, label the sensitivity level according to the preset sensitive word library, and store the labeling results in the local permission metadata database and verify them; Step S2: Train a lightweight language model based on the obtained historical shared application samples to obtain processing rules bound to the current session ID; Step S3: Based on the acquired raw data records and processing rules, call the rule parsing engine to generate an in-situ processing instruction sequence, and process the raw data records field by field in-situ. Step S4: Generate an intermediate quality report based on the batch quantity processed in situ, and call the triggerOptimizer() function to trigger the tuner and obtain new processing rules; Step S5: Obtain a temporary processing result object according to the new processing rules, generate a standardized processing dataset according to the review decision results, and call the encryptAndSend() function to transmit it to the target server via HTTPS POST request in an encrypted manner.

2. The AI-assisted edge-side dynamic data orientation processing method for controllable data sharing according to claim 1, characterized in that, Step S1 specifically includes the following steps: Step S11: Read the metadata of the multi-source heterogeneous data obtained from the local machine, and mark the sensitivity level and sharing permission threshold according to the preset sensitive word library, and store the marking results in the local permission metadata database; Step S12: Based on the obtained shared processing request, call the request parsing function parseRequest() to extract the structured fields and save them to the current session context; Step S13: Query the local permission metadata database based on the extracted structured fields to verify whether the access is authorized, and determine whether it is compliant based on the set compliance rule base; if it is not authorized or does not comply with the rules, return false, generate a denial message and record it in the log; if the verification passes, return true and write the verification result to the current session context and authorization timestamp.

3. The AI-assisted edge-side dynamic data orientation processing method for controllable data sharing according to claim 1, characterized in that, Step S2 specifically includes the following steps: Step S21: Train a lightweight language model based on the obtained historical shared application samples and manually annotated request statements to obtain the trained lightweight language model; Step S22: Obtain the application text from the current session context, and call the parseIntent() function and the trained lightweight language model to obtain the processing dimension parameters; Step S23: Based on the processing dimension parameters, the obtained data quality report, and the sensitive level tags obtained from the local permission metadata database, call the rule generator to generate processing rules bound to the current session ID.

4. The AI-assisted edge-side dynamic data orientation processing method for controllable data sharing according to claim 1, characterized in that, Step S3 specifically includes the following steps: Step S31: Based on the table name and field list of the structured fields, use an SQL query to obtain the original data records and store them in the memory buffer; Step S32: Based on the original data records, call the rule parsing engine, traverse the rule entries of the processing rules, extract the operation type and parameter dictionary, and generate the in-situ processing instruction sequence; Step S33: Process the original data record field by field in situ according to the in-situ processing instruction sequence.

5. The AI-assisted edge-side dynamic data orientation processing method for controllable data sharing according to claim 1, characterized in that, Step S4 specifically includes the following steps: Step S41: Generate an intermediate quality report based on the batch quantity of in-situ processing; Step S42: Based on the processing rules and intermediate quality reports, obtain the actual quality indicators and expected quality indicators; Step S43: Calculate the deviation based on the actual quality indicators and expected quality indicators, and call the triggerOptimizer() function according to the deviation type to trigger the tuner to optimize and obtain new processing rules.

6. The AI-assisted edge-side dynamic data orientation processing method for controllable data sharing according to claim 5, characterized in that, Step S43 specifically includes the following steps: Step S431: Locate the corresponding adjustable parameters and adjustment rules based on the deviation type; Step S432: Modify the parameters and operation type of the current processing rule according to the adjustable parameters and adjustment rules found, and obtain the optimized rule fragment; Step S433: Merge the optimized rule fragment with the current processing rule to obtain a new processing rule.

7. The AI-assisted edge-side dynamic data orientation processing method for controllable data sharing according to claim 1, characterized in that, Step S5 specifically includes the following steps: Step S51: According to the new processing rules, the columnar storage field array, field name, and session identifier are encapsulated into a temporary processing result object, and the pushToAudit() function is called to push the temporary processing result object to the audit queue through the local secure channel; Step S52: Based on the user's audit decision result (string type) and the current session context, call the handleAuditDecision() function to obtain the audit result; if the audit result is rejection, discard the temporary processing result object and send a retry signal to trigger reprocessing; if the audit result is approval, proceed to step S53. Step S53: Call the generateStandardDataset() function to serialize the approved temporary processing result object into a data file in the set format, and generate a data dictionary file according to the meaning of the fields, the desensitization method, and the data range of the temporary processing result object; Step S54: Combine the data file and the data dictionary file into a compressed package, call the encryptAndSend() function to generate a symmetric key, and call the AES-256-GCM algorithm to encrypt the compressed package to obtain a ciphertext package; Step S55: Combine the encrypted symmetric key and ciphertext packet, send them to the target server via an HTTPS POST request, and destroy the local temporary file and the key in memory.

8. A dynamic data orientation processing system for AI-assisted edge devices oriented towards controllable data sharing, used to execute the dynamic data orientation processing method for AI-assisted edge devices oriented towards controllable data sharing as described in claim 1; characterized in that, It includes a shared permission verification unit, a targeted processing rule generation unit, an in-situ targeted processing execution unit, an optimization unit, and a shared output unit; The shared permission verification unit is used to obtain metadata of multi-source heterogeneous data, mark the sensitivity level according to the preset sensitive word library, and store the marking results in the local permission metadata database and verify them. The targeted processing rule generation unit is used to train a lightweight language model based on the acquired historical shared application samples to obtain processing rules bound to the current session ID; The in-situ directional processing execution unit is used to call the rule parsing engine to generate an in-situ processing instruction sequence based on the acquired original data records and processing rules, and process the original data records field by field in-situ. The optimization unit is used to generate an intermediate quality report based on the batch quantity of in-situ processing, and call the triggerOptimizer() function to trigger the optimizer to optimize and obtain new processing rules; The shared output unit is used to obtain a temporary processing result object according to the new processing rules, generate a standardized processing dataset according to the review decision results, and call the encryptAndSend() function to transmit it to the target server via an HTTPS POST request in an encrypted manner.

9. An electronic device, characterized in that, It includes a memory and a processor; the memory stores a computer program; when the computer program is executed on the processor, it implements the AI-assisted edge-side dynamic data orientation processing method for controllable data sharing as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions; when the computer instructions are executed on the electronic device as described in claim 9, they implement the AI-assisted edge-side dynamic data orientation processing method for controllable data sharing as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Information scheduling scenarized service agent system

    CN121070566A

  • Enterprise-level user data full life cycle desensitization verification method based on block chain

    CN121615184A