An edge-end artificial intelligence model matching algorithm

CN122734451APending Publication Date: 2026-09-11NANJING GUOZHUN DATA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610909191.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0007]尽管学术界和工业界已在模型选择与匹配方面开展了一些初步探索,但面对边端场景的独特挑战如:资源受限、环境动态变化、任务多样性等,现有技术方案往往缺乏系统性,难以满足实战级应用的高要求,主要问题体现在以下几个方面:

Benefits of technology

[0036]1.通过构建任务、资源、模型的三维特征关联体系,并采用多维度加权融合的精细评分模型,能够全面、深度地评估模型与复杂需求的契合度,使匹配结果与实际部署性能高度相关,大幅提升Top-K命中率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122734451A_ABST
    Figure CN122734451A_ABST
Patent Text Reader

Abstract

The application discloses a kind of edge-end artificial intelligence model matching algorithm, belong to artificial intelligence and edge computing technology field, the method includes: obtaining task demand and edge-end equipment resource state information;Based on pre-constructed model feature library, the matching process of containing quick screening and fine screening is executed, and the fine screening generates the comprehensive score of candidate model by multidimensional scoring model;According to the score, determine and output final matching scheme, the system includes the feature management, hierarchical matching and scheme output module of implementation above-mentioned step, the present application overcomes the deficiency of traditional method in precision, real-time and expansibility by hierarchical matching and multidimensional fusion evaluation, realizes the precision, efficient and self-adapting of model selection in resource-limited edge-end environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and edge computing technology, specifically relating to a method and system for dynamic matching and scheduling of artificial intelligence models for resource-constrained edge devices, which is used to achieve fast, accurate and adaptive matching between task requirements and AI models in the model library. Background Technology

[0002] With the popularization of artificial intelligence technology and the rise of edge computing architecture, deploying AI models to terminals and edge devices has become an important trend for realizing real-time intelligent perception and decision-making. Against this backdrop, edge AI model matching, as a key hub connecting upper-layer applications and lower-layer model resources, is becoming increasingly important.

[0003] This technology can automatically and efficiently select and recommend the optimal model or model combination from a massive model library based on specific task requirements, target operating environment, and the stringent resource constraints (such as computing power, memory, and energy consumption) of edge devices, and directly deploy it to edge devices for inference. Its performance is directly related to the response speed, decision accuracy, and system reliability of key application scenarios such as intelligent manufacturing, autonomous driving, and intelligent security.

[0004] Chinese patent application CN114595023A discloses a method, system, and smart terminal for running a personalized interface based on a third-party service. The gateway device includes a system kernel and multiple storage modules connected to the system kernel. During the process of receiving data from the storage modules, the system kernel: obtains a service request instruction; extracts the feature identifier of the service request instruction and outputs a retrieval instruction based on the feature identifier; receives an image stream and a layout library returned by the storage modules, the layout library containing layout rules corresponding one-to-one with the image stream; extracts display element information from the image stream; matches the display element information and preset matching rules in the layout library to obtain a target layout rule; synthesizes a target display interface based on the target layout rule and display element information; and outputs the target display interface. Edge AI model matching, as a key hub connecting upper-layer applications and lower-layer model resources, directly affects the response speed, decision accuracy, and reliability of the application system.

[0005] US Patent 11232152B2 discloses a method and system for generating embeddings for nodes in a corpus graph. More specifically, the operation for generating aggregated embedding vectors for target nodes is effectively distributed across operations of a central processing unit (CPU) and a graphics processing unit (GPU). For a target node within the corpus graph, processing is performed by one or more CPUs to identify the relevant neighborhoods (nodes) of the target node in the corpus graph. This information is prepared and passed to one or more GPUs to determine the aggregated embedding vector of the target node based on the data of the relevant neighborhoods. However, this system lacks a hierarchical matching mechanism, cannot balance real-time performance and accuracy, and does not consider the dynamic changes in edge device resources.

[0006] Therefore, developing efficient, accurate, and robust edge AI model matching algorithms is of great significance for fully unleashing the potential of edge intelligence and ensuring the smooth execution of critical tasks.

[0007] Although academia and industry have conducted some preliminary explorations in model selection and matching, existing technical solutions often lack systematicity and fail to meet the high requirements of practical applications when faced with unique challenges in edge scenarios such as resource constraints, dynamic environmental changes, and task diversity. The main problems are reflected in the following aspects:

[0008] 1. Existing methods mostly use simple rule filtering or rely on a single similarity metric. These methods fail to build deep associations and ignore the coupling relationship between task constraints, real-time device resource status and model characteristics, resulting in matching results that are often one-sided and deviate significantly from the actual performance after deployment.

[0009] 2. Existing frameworks lack flexible layering and routing mechanisms, and usually adopt a one-size-fits-all matching process. Either they oversimplify the matching logic in pursuit of speed, sacrificing accuracy, or they introduce complex optimization algorithms in pursuit of accuracy, resulting in excessively high response latency, which cannot meet the real-time requirements of high concurrency or urgent tasks.

[0010] 3. Existing technologies lack effective fault tolerance, degradation strategies, and dynamic adjustment mechanisms for abnormal situations. When encountering incomplete data or resource mutations, they are prone to matching failures or recommending models that cannot be deployed, resulting in insufficient overall system robustness.

[0011] 4. Existing solutions often use hard-coded matching rules and similarity calculation methods, resulting in a rigid system architecture. When new heterogeneous models are added to the model library, or when a matching strategy needs to be customized for a specific task type, the core algorithm code often needs to be modified.

[0012] 5. Existing matching processes typically lack robust permission verification, data transmission encryption, sensitive information anonymization, and end-to-end audit logging. Once a problem occurs, it is impossible to trace the matching logic or reproduce the decision-making process, making it difficult to meet security audit and compliance requirements.

[0013] In summary, existing edge AI model matching technologies have inherent shortcomings in terms of feature engineering systematization, matching process flexibility, anomaly handling robustness, architectural scalability, and security mechanism completeness, and cannot yet provide reliable support for tactical-level edge intelligent deployment. Therefore, the industry urgently needs an innovative and systematic matching algorithm that can fundamentally solve the above-mentioned multi-dimensional challenges and achieve a balance between accuracy, speed, robustness, and security. Summary of the Invention

[0014] In response to the systemic deficiencies of existing edge AI model matching technologies in terms of accuracy, real-time performance, robustness, scalability, and security, as pointed out in the background art, this invention proposes a novel and systematic solution. By constructing a multi-dimensional feature association and hierarchical matching architecture and introducing a configurable intelligent scoring and dynamic adjustment mechanism, it achieves accurate, efficient, and reliable mapping from task requirements to model resources.

[0015] The core of this invention lies in an edge-based artificial intelligence model matching method, characterized by comprising the following steps:

[0016] S1: Obtain structured task requirement information and real-time resource status information of edge devices. The task requirement information includes at least task type, accuracy threshold, real-time level, and environmental conditions. The resource status information includes device hardware architecture, available memory, current computing load, and network bandwidth.

[0017] S2: Perform feature preprocessing steps: extract and standardize the features of the task requirement information, resource status information and metadata in the pre-built model feature library to generate standardized feature data; the standardization encoding includes: mapping text features to fixed-dimensional semantic vectors through a pre-trained semantic encoder, normalizing numerical features, and converting categorical features into one-hot encodings.

[0018] S3: Based on the standardized feature data, perform a hierarchical matching process:

[0019] S31: Based on hard constraint rules, perform rapid screening to eliminate models that do not meet the basic compatibility requirements and obtain a preliminary candidate model set; the hard constraint rules include at least: device architecture compatibility, minimum memory requirements, minimum computing power requirements, and functional type matching;

[0020] S32: Based on the real-time level in the task requirements information, automatically select the matching execution path: if the real-time level is "high", then enter the fast execution path and use a simplified version of the multidimensional scoring model to score and sort the preliminary candidate model set; if the real-time level is "medium / low", then enter the precise execution path and use a full version of the multidimensional scoring model to score and sort the preliminary candidate model set.

[0021] S33: For candidate models ranked by score, if the task requirement information indicates that multiple models need to collaborate, then perform the combinatorial optimization step: generate candidate model combinations based on heuristic algorithms, evaluate the functional complementarity, overall resource consumption and interface collaboration complexity of the combinations, and output the optimal model combination scheme; if multiple model collaboration is not required, then directly output the single model scheme with the highest score.

[0022] S4: Before outputting the final matching solution, perform a security audit step: desensitize sensitive fields in the task requirement information and generate an immutable audit log containing key matching nodes;

[0023] S5: Output the final model matching solution and deploy it to the edge device; after deployment, continuously monitor the actual operating indicators of the edge device. When the actual operating indicators deviate from the preset threshold by more than 20%, trigger the re-matching process.

[0024] The pre-built model feature library includes three sub-libraries: task feature library, resource feature library, and model feature library. A multi-level index system is built for each sub-library, with a global inverted index. For key fields of task type, device architecture, and model function, an approximate nearest neighbor index is used. For high-dimensional semantic vectors, a filter index is used, based on resource hard constraint fields.

[0025] The calculation process of the complete multidimensional scoring model is as follows: S321: Calculate the sub-scores of the candidate model in four dimensions, reduction compatibility score: calculated based on the degree of satisfaction of hard constraints, the score is 0 if any hard constraint is not satisfied, semantic similarity score: calculate the cosine similarity between the task semantic vector and the model functional semantic vector, resource suitability score: based on the historical performance baseline of the model on similar hardware, the comprehensive score of estimated inference latency, energy consumption and memory usage, performance matching score: calculate the score by comparing the difference between the measured accuracy of the model and the task accuracy threshold.

[0026] The combined optimization steps specifically include:

[0027] S331: Select the top K candidate models in the score ranking as the base set for combination, where K ranges from 5 to 10;

[0028] S332: Genetic algorithm is used to generate all possible model combinations, and combinations whose total resource consumption exceeds the resource budget of the edge device are eliminated;

[0029] S333: Calculate the functional complementarity score and interface coordination complexity score of the remaining combinations, and then weight and fuse them to obtain the comprehensive score of the combinations;

[0030] S334: Output the model combination scheme with the highest comprehensive score and the corresponding deployment strategy.

[0031] The security audit process includes: sensitive fields are desensitized using replacement or obfuscation methods, and sensitive information such as geographic location, time window, and task number are irreversibly transformed; audit logs are recorded in a structured format, which includes request ID, candidate model list, detailed scores for each dimension, final selection result and decision reasons, and the format records are digitally signed, and the logs are kept for no less than one year.

[0032] The actual operating indicators monitored in the dynamic adjustment steps include: model inference latency, inference accuracy, memory usage, CPU utilization, and energy consumption. When any of the monitored actual operating indicators deviates from the preset threshold by more than 20% for three consecutive sampling values, the rematching process is automatically triggered.

[0033] An edge AI model matching system is characterized by comprising: a feature management module, a hierarchical matching engine, and a solution output module; the feature management module acquires and processes task requirement information, edge device resource status information, and model feature library information to generate multi-dimensional feature data for matching; the hierarchical matching engine is connected to the feature management module, receives the multi-dimensional feature data, and executes the matching process; the matching process includes at least: rapid model screening based on hard constraints, and fine-tuning of the rapidly screened candidate models based on a multi-dimensional scoring model to generate a comprehensive score; the solution output module is connected to the hierarchical matching engine, determines and outputs the final model matching solution based on the comprehensive score.

[0034] The feature management module has a built-in feature encoder, index manager, and cache unit. The feature encoder supports incremental updates and batch pre-computation. The cache unit is used to store standardized feature vectors of commonly used task templates and high-frequency models. The hierarchical matching engine has a built-in rule engine, fast scoring unit, and precise scoring unit. The rule engine is used to perform fast filtering. The fast scoring unit corresponds to a fast execution path. The precise scoring unit corresponds to a precise execution path.

[0035] Compared with existing technologies, the edge AI model matching method and system provided by this invention can bring the following significant benefits:

[0036] 1. By constructing a three-dimensional feature association system of tasks, resources, and models, and adopting a fine scoring model with multi-dimensional weighted fusion, it is possible to comprehensively and deeply evaluate the fit between the model and complex requirements, making the matching results highly correlated with the actual deployment performance, and significantly improving the Top-K hit rate.

[0037] 2. Through a layered architecture that combines rapid and refined filtering, and a flexible dual-execution path design, the system can intelligently adapt to different scenarios, from millisecond-level emergency response to complex high-precision planning, maximizing matching quality while meeting real-time constraints.

[0038] 3. Built-in dynamic adjustment steps and exception handling mechanisms enable the system to effectively cope with abnormal situations such as fluctuations in edge device resources and incomplete input data, ensuring continuous service availability and reducing the matching failure rate.

[0039] 4. The modular and pluggable design enables the system to quickly integrate new models or adapt to new task types without modifying the core architecture, greatly improving the efficiency of technology iteration and scenario migration.

[0040] 5. An integrated security audit mechanism, from data anonymization to end-to-end log recording, ensures the security of sensitive information and makes the matching decision-making process transparent, auditable, and explainable, meeting the compliance needs of scenarios with high security requirements.

[0041] 6. By combining optimization steps and resource feasibility verification, it can intelligently recommend multi-model collaborative solutions and make full use of the limited computing, memory and energy resources at the edge while ensuring performance, thereby improving the overall efficiency of task execution.

[0042] In summary, this invention, through a systematic and innovative design, effectively solves the fragmentation and limitations of existing edge AI model matching technologies, providing key technical support for achieving efficient, accurate, reliable, and secure edge intelligent deployment. It has significant practical value and broad application prospects. Attached Figure Description

[0043] Figure 1 This is an overall flowchart of the edge artificial intelligence model matching algorithm provided in the embodiments of the present invention.

[0044] Figure 2 This is a schematic diagram of the three-dimensional feature association and index structure provided in the embodiment of the present invention.

[0045] Figure 3 This is a schematic diagram of the hierarchical matching and combination optimization structure provided in the embodiments of the present invention.

[0046] Figure 4 This is a schematic diagram of the system deployment architecture provided in an embodiment of the present invention. Detailed Implementation

[0047] To enhance understanding of the present invention, the invention will be further described in detail below with reference to embodiments and accompanying drawings. These embodiments are only for explaining the invention and do not constitute a limitation on the scope of protection of the invention.

[0048] This invention provides an edge-side artificial intelligence model matching method and system, aiming to solve systemic defects in existing technologies such as low matching accuracy, difficulty in balancing real-time performance and accuracy, poor system robustness, insufficient scalability, and lack of security auditing. The following will combine... Figures 1 to 4 The specific embodiments of the present invention will be described in detail below.

[0049] like Figure 1 As shown, the process begins when the system receives a matching request. The request includes at least the structured task requirement information as described in claim 1 and the real-time resource status information of the target edge device. The system then performs a feature preprocessing step, extracting and standardizing features from the input information and model feature library metadata to generate standardized feature data. Next, it initiates a dual-path judgment: automatically selecting the matching execution path based on the real-time level in the task requirement information; if the real-time level is "high", it enters the fast execution path; if the real-time level is "medium / low", it enters the precise execution path.

[0050] Regardless of the path chosen, the core process includes rapid screening based on hard constraint rules and refined screening based on multi-dimensional scoring models. For candidate models ranked by scoring, if the task requirements indicate the need for multi-model collaboration, a combined optimization step is performed. Before outputting the final matching solution, a security audit is conducted to anonymize sensitive information and generate an immutable audit log containing key matching nodes. The matching solution is then output and deployed to edge devices. After deployment, the actual operating metrics of the devices are continuously monitored, and a re-matching process is triggered when the metrics deviate from a preset threshold by more than 20%. This entire process ensures end-to-end accuracy, efficiency, and controllability from requirement input to solution output.

[0051] like Figure 2 As shown, the improvement of this invention lies in the construction of a systematic feature engineering system, which systematically maintains the three core feature libraries described in claim 2:

[0052] Task Feature Library: Stores feature vectors of historical and template tasks. Features include semantic vectors, numerical constraints, and category labels.

[0053] Resource feature library: records the static capabilities and dynamic status of each edge device;

[0054] Model Feature Library: Stores metadata for each AI model, including feature description vectors, interface specifications, performance baselines on different hardware, and minimum resource requirements.

[0055] To accelerate retrieval, the system constructs a dedicated multi-level index for each sub-database: a. Global inverted index: deployed in the task feature database, quickly locating relevant entries based on key fields such as task type and target category; b. Approximate nearest neighbor index: deployed in the model feature database, supporting millisecond-level similarity retrieval for high-dimensional semantic vectors; c. Filter index: deployed in the resource feature database, enabling rapid batch filtering based on resource hard constraint fields. This structure transforms the originally scattered data into a highly correlated and easily searchable knowledge network, providing a fundamental guarantee for improved matching accuracy and speed.

[0056] like Figure 3 As shown, the hierarchical design of the hierarchical matching engine is a key improvement of this invention. After the matching request is feature-encoded, it first enters the coarse screening layer. This layer mainly relies on the rule engine and the aforementioned filter index to execute the four hard constraint rules of claim 1: excluding models incompatible with the device hardware architecture, excluding models whose minimum memory requirement exceeds the device's currently available memory, excluding models whose minimum computing power requirement exceeds the device's currently available computing power, and excluding models whose function type does not match the task type. This stage aims to quickly eliminate a large number of mismatches with extremely low overhead, producing a small preliminary candidate set.

[0057] The system then proceeds to the refinement layer, where it invokes a multi-dimensional scoring model to quantitatively evaluate each model in the initial candidate set. The precise execution path employs the complete multi-dimensional scoring model described in claim 3, calculating sub-scores across four dimensions: 1. Reduction compatibility score: calculated based on the degree of hard constraint satisfaction; a score of 0 is awarded if any hard constraint is not met; 2. Semantic similarity score: calculating the cosine similarity between the task semantic vector and the model's functional semantic vector; 3. Resource suitability score: based on the model's historical performance baseline on similar hardware, estimating a comprehensive score for inference latency, energy consumption, and memory usage; 4. Performance matching score: calculating the score by comparing the difference between the model's measured accuracy and the task accuracy threshold.

[0058] Each sub-score is weighted and fused based on a preset or dynamically adjusted scenario-based weight configuration to obtain a comprehensive score, which is then used to rank the candidate models. The fast execution path uses a simplified version of the multidimensional scoring model, calculating sub-scores only for three dimensions: specification compatibility, semantic similarity, and resource adaptability.

[0059] For complex tasks requiring collaboration among multiple models, the system will activate a combinatorial optimization layer. This layer executes the combinatorial optimization steps described in claim 4: using the top K models ranked by the screening layer as input (K ranging from 5 to 10), a genetic algorithm is used to generate candidate model combinations. First, infeasible combinations whose overall resource consumption exceeds the resource budget of the edge devices are eliminated. Then, the functional complementarity and interface coordination complexity of the remaining combinations are evaluated, and a weighted fusion is performed to obtain a comprehensive combinatorial score. Finally, the model combination scheme with the highest comprehensive score is recommended. This progressive structure ensures that the optimal solution is gradually approached while keeping resource consumption under control.

[0060] like Figure 4 As shown, the system of the present invention can be deployed on edge servers or central clouds, and works in collaboration with edge devices distributed in various locations. Its modular design reflects good scalability.

[0061] The feature management module serves as the data entry point and incorporates the feature encoder, index manager, and cache unit as described in claim 8: it is responsible for connecting to task requests and device status reports from different sources and calling the feature encoder to perform feature extraction and standardized encoding; the index manager maintains multi-level indexes of three feature libraries and supports dynamic updates; the cache unit stores standardized feature vectors of commonly used task templates and high-frequency models to accelerate the feature processing flow.

[0062] The hierarchical matching engine is the core computing unit, which incorporates the rule engine, fast scoring unit, and precise scoring unit as described in claim 8: the rule engine is used to execute the fast filtering step; the fast scoring unit implements a simplified multidimensional scoring model, corresponding to a fast execution path; and the precise scoring unit implements a complete multidimensional scoring model, corresponding to a precise execution path.

[0063] The solution output module not only outputs model identifiers, but also generates detailed deployment lists and configuration scripts, supporting one-click deployment to edge devices.

[0064] In addition, the security audit agent that runs independently throughout the system executes the security audit mechanism described in claim 5: sensitive fields such as geographical location, time window, and task number in the task are replaced or obfuscated before entering the matching engine; the audit log records the request ID, candidate model list, score details of each dimension, final selection result and decision reasons for each match in a structured format, and uses digital signature to prevent tampering. The log is kept for no less than one year to meet compliance archiving requirements.

[0065] The dynamic adjustment monitor continuously collects the five operational metrics described in claim 6 from the edge devices of the deployed models: model inference latency, inference accuracy, memory usage, CPU utilization, and energy consumption. When any metric deviates from the preset threshold by more than 20% for three consecutive sampling values, a rematch request is sent to the matching engine along with the latest status of the current device, triggering a new round of matching process and forming a closed-loop feedback.

[0066] The key algorithms and mechanisms involved in this invention are detailed below:

[0067] Weight configuration of the multidimensional scoring model: It supports static configuration and dynamic adjustment. In the embodiment of the tactical edge scenario, the default weight configuration described in claim 3 is adopted: specification compatibility 35%, resource cost 30%, semantic similarity 18%, and performance matching degree 17%. The system also allows automatic selection of preset weight templates based on task tags.

[0068] Dynamic adjustment mechanism: The monitor executes according to the triggering conditions described in claim 6 to ensure that the model scheme can be adaptively adjusted in a timely manner when equipment resources fluctuate.

[0069] Security audit mechanism: Strictly follow the de-identification method, log format and storage requirements described in claim 5 to ensure data security and decision traceability.

[0070] Implementation Example

[0071] Example: Single-model matching scenario for container identification in smart ports

[0072] The claims in this embodiment demonstrate the complete matching process for a single model task.

[0073] Scenario: A smart port management system needs to dynamically assign container number identification tasks to smart cameras on the gantry cranes at the dock.

[0074] Task requirements: Task type = character recognition, accuracy threshold = 0.99, real-time performance level = medium (response time ≤ 500ms), environmental conditions = day / night cycle, light rain.

[0075] Edge device resource status: Hardware architecture = ARM Cortex-A76, available memory = 2GB, current computing load = 30%, network bandwidth = 10Mbps

[0076] Matching execution flow:

[0077] Information Acquisition and Feature Preprocessing (corresponding to claims 1-S1, S2): After acquiring the above structured information, the system performs feature preprocessing: converts text features such as "character recognition", "day and night alternation, light rain" into 768-dimensional semantic vectors through the BERT encoder; normalizes numerical features such as 2GB of available memory and 30% computing load to the [0,1] interval; converts the hardware architecture "ARMCortex-A76" into one-hot encoding to generate standardized feature data.

[0078] Quick filtering (corresponding to claims 1-S31): The hierarchical matching engine calls the rule engine to perform quick filtering based on the following hard constraint rules:

[0079] Device architecture compatibility: All x86 and CUDA architecture models were excluded.

[0080] Minimum memory requirements: Models with memory requirements > 2GB are excluded.

[0081] Minimum computing power requirement: Models with inference latency >500ms on ARM architecture will be excluded.

[0082] Functional type matching: After quickly filtering out non-character recognition models, 8 preliminary candidate models were obtained from the initial 1200+ model library, reducing the size of the candidate set to 0.67% of the original.

[0083] Execution path selection and multidimensional scoring: Based on the task's real-time level of "medium," the system automatically selects a precise execution path and uses a full-version multidimensional scoring model to calculate the comprehensive score of each candidate model.

[0084] All eight models met all hard constraints and received a score of 1.0 for specification compatibility.

[0085] Semantic similarity score: Calculate the cosine similarity between the task semantic vector and the semantic vector of each model's functional description, with a score range of 0.72-0.91.

[0086] Resource Adaptability Score: Based on the model's historical performance baseline on the ARM Cortex-A76 platform, this score is a comprehensive score estimating inference latency, power consumption, and memory usage, ranging from 0.68 to 0.89.

[0087] Performance matching score: The difference between the measured accuracy of the model in light rain and the task accuracy threshold of 0.99 was compared. The score range was 0.75-0.93. The default weight configuration of tactical scene (reduction compatibility 35%, resource cost 30%, semantic similarity 18%, performance matching 17%) was used for weighted fusion. Finally, the container recognition model based on YOLOv5s improved the highest comprehensive score (0.92 points).

[0088] Security Audit (corresponding to claims 1-S4 and 5): The security audit module performs hash-based desensitization processing on sensitive information such as "gantry crane number" and "operation area" in the task, generates a structured audit log, which includes the request ID, the scoring details of 8 candidate models, the final selection result and the decision reason for "highest comprehensive score", and attaches a SHA-256 digital signature. The log is stored in the system audit database and is kept for 1 year.

[0089] Deployment and Dynamic Monitoring (corresponding to claims 1-S5 and 6): The system outputs a matching scheme and deploys the model to the target camera. After deployment, the dynamic adjustment module monitors the operating indicators at a frequency of 10 seconds / time: actual model inference latency = 120ms, accuracy = 0.993, memory utilization = 45%, CPU utilization = 52%, power consumption = 2.1W. All indicators are within the preset threshold range, and the re-matching process is not triggered.

[0090] In summary, this example demonstrates how the present invention abstracts complex application scenarios into structured tasks, and through hierarchical screening and multi-dimensional quantitative evaluation, ultimately outputs accurate and executable model matching decisions. The entire process fully reflects the feasibility, effectiveness, and systematic advantages of the present invention in dealing with dynamic needs, heterogeneous resources, and diverse models in real industrial environments.

[0091] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for matching edge-end artificial intelligence models, characterized in that, Includes the following steps: S1: Obtain structured task requirement information and real-time resource status information of edge devices. The task requirement information includes at least task type, accuracy threshold, real-time level, and environmental conditions. The resource status information includes device hardware architecture, available memory, current computing load, and network bandwidth. S2: Perform feature preprocessing steps: extract and standardize the features of the task requirement information, resource status information and metadata in the pre-built model feature library to generate standardized feature data; The standardized encoding includes: mapping textual features to fixed-dimensional semantic vectors through a pre-trained semantic encoder, normalizing numerical features, and converting categorical features into one-hot encodings. S3: Based on the standardized feature data, perform a hierarchical matching process: S31: Based on hard constraint rules, perform rapid screening to eliminate models that do not meet the basic compatibility requirements and obtain a preliminary candidate model set; the hard constraint rules include at least: device architecture compatibility, minimum memory requirements, minimum computing power requirements, and functional type matching; S32: Based on the real-time level in the task requirements information, automatically select the matching execution path: if the real-time level is "high", then enter the fast execution path and use a simplified version of the multidimensional scoring model to score and sort the preliminary candidate model set; if the real-time level is "medium / low", then enter the precise execution path and use a full version of the multidimensional scoring model to score and sort the preliminary candidate model set. S33: For candidate models ranked by score, if the task requirement information indicates that multiple models need to collaborate, then perform the combinatorial optimization step: generate candidate model combinations based on heuristic algorithms, evaluate the functional complementarity, overall resource consumption and interface collaboration complexity of the combinations, and output the optimal model combination scheme; if multiple model collaboration is not required, then directly output the single model scheme with the highest score. S4: Before outputting the final matching solution, perform a security audit step: desensitize sensitive fields in the task requirement information and generate an immutable audit log containing key matching nodes; S5: Output the final model matching solution and deploy it to the edge device; after deployment, continuously monitor the actual operating indicators of the edge device. When the actual operating indicators deviate from the preset threshold by more than 20%, trigger the re-matching process.

2. The edge-end artificial intelligence model matching method according to claim 1, characterized in that, The pre-built model feature library includes three sub-libraries: task feature library, resource feature library, and model feature library. A multi-level index system is built for each sub-library, with a global inverted index. For key fields of task type, device architecture, and model function, an approximate nearest neighbor index is used. For high-dimensional semantic vectors, a filter index is used, based on resource hard constraint fields.

3. The edge-end artificial intelligence model matching method according to claim 1, characterized in that, The calculation process of the complete multidimensional scoring model is as follows: S321: Calculate the sub-scores of the candidate model in four dimensions, reduction compatibility score: calculated based on the degree of satisfaction of hard constraints, the score is 0 if any hard constraint is not satisfied, semantic similarity score: calculate the cosine similarity between the task semantic vector and the model functional semantic vector, resource suitability score: based on the historical performance baseline of the model on similar hardware, the comprehensive score of estimated inference latency, energy consumption and memory usage, performance matching score: calculate the score by comparing the difference between the measured accuracy of the model and the task accuracy threshold.

4. The edge-end artificial intelligence model matching method according to claim 1, characterized in that, The combined optimization steps specifically include: S331: Select the top K candidate models in the score ranking as the base set for combination, where K ranges from 5 to 10; S332: Genetic algorithm is used to generate all possible model combinations, and combinations whose total resource consumption exceeds the resource budget of the edge device are eliminated; S333: Calculate the functional complementarity score and interface coordination complexity score of the remaining combinations, and then weight and fuse them to obtain the comprehensive score of the combinations; S334: Output the model combination scheme with the highest comprehensive score and the corresponding deployment strategy.

5. The edge-end artificial intelligence model matching method according to claim 1, characterized in that, The security audit steps include: sensitive fields are desensitized using replacement or fuzzing methods, and sensitive information such as geographic location, time window, and task number are irreversibly transformed. Audit logs are recorded in a structured format, which includes the request ID, a list of candidate models, detailed scores for each dimension, the final selection result and the rationale for the decision. The logs are digitally signed and are kept for at least one year.

6. The edge-end artificial intelligence model matching method according to claim 1, characterized in that, The actual operating indicators monitored in the dynamic adjustment step include: model inference latency, inference accuracy, memory usage, CPU utilization, and energy consumption; when any of the monitored actual operating indicators deviates from the preset threshold by more than 20% for three consecutive sampling values, the rematching process is automatically triggered.

7. An edge-end artificial intelligence model matching system, characterized in that, include: Feature management module, hierarchical matching engine, and solution output module; The feature management module acquires and processes task requirement information, edge device resource status information, and model feature library information to generate multi-dimensional feature data for matching. The hierarchical matching engine is connected to the feature management module, receives the multi-dimensional feature data, and executes the matching process; The matching process includes at least: quickly screening models based on hard constraints, and finely screening the candidate models after quick screening based on a multi-dimensional scoring model and generating a comprehensive score; The scheme output module is connected to the hierarchical matching engine, and determines and outputs the final model matching scheme based on the comprehensive score.

8. The edge-end artificial intelligence model matching system according to claim 6, characterized in that, The feature management module has a built-in feature encoder, index manager, and cache unit. The feature encoder supports incremental updates and batch pre-computation. The cache unit is used to store standardized feature vectors of commonly used task templates and high-frequency models. The hierarchical matching engine has a built-in rule engine, fast scoring unit, and precise scoring unit. The rule engine is used to perform fast filtering. The fast scoring unit corresponds to a fast execution path. The precise scoring unit corresponds to a precise execution path.

Citation Information

Patent Citations

  • Method and system for operating personalized interface based on third-party service, and intelligent terminal

    CN114595023A

  • Efficient processing of neighborhood data

    US11232152B2