An AI Edge Computing Method, Device, Electronic Device, and Storage Medium
Through dynamic adjustment of model and resource allocation by edge servers, the problems of high demand for AI edge computing resources and insufficient intelligence are solved, and efficient and intelligent business processing is achieved, which is suitable for a variety of application scenarios.
Patent Information
- Application Number
- CN202510317444.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-18
AI Technical Summary
Existing AI edge computing has high demands on computing resources and storage resources and is not intelligent enough, resulting in inefficiency.
Through edge servers, multiple model sets corresponding to business types are obtained, the model is dynamically adjusted according to terminal attributes and environmental data, resource allocation and processing strategies are optimized, and the model is realized in real-time adaptation and efficient processing.
It improves the performance and intelligence of AI edge computing, reduces computing latency, adapts to changes in different business scenarios and environments, provides personalized services, and enhances system robustness and stability.
Smart Images

Figure CN119847765B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to an AI edge computing method, device, electronic device, and storage medium. Background Art
[0002] Edge computing is a new computing system and technology that sinks computing power from the cloud to the edge side of the network to achieve real-time services, efficient data processing, application intelligence, and security and privacy protection. Currently, with the development of artificial intelligence technology, artificial intelligence technology is gradually participating in all aspects of people's work and life. However, artificial intelligence technology has very high requirements for computing resources and storage resources, and it is relatively inefficient and not intelligent enough during edge computing. Summary of the Invention
[0003] Based on the above problems, the present invention proposes an AI edge computing method, device, electronic device, and storage medium. Through the solution of the present invention, not only can the performance of AI edge computing be improved, but also more efficient and intelligent business processing can be achieved in various application scenarios.
[0004] In view of this, one aspect of the present invention proposes an AI edge computing method, including:
[0005] The edge server obtains a plurality of first model sets corresponding to the business types of the services it processes, and the first model sets are generated by the cloud server according to specific business types and business contents;
[0006] The edge server selects corresponding first service processing terminals for connection from the service processing terminals according to the task plans and data processing requirements of a plurality of service processing terminals;
[0007] The edge server obtains the terminal attribute data and service request data of the first service processing terminal;
[0008] The edge server generates a second model set according to the terminal attribute data, a preset first model adjustment algorithm, and the first model sets;
[0009] The edge server determines a service processing party according to the terminal attribute data, its own service processing ability, and resource occupancy rate;
[0010] When it is determined that the service processing party is the first service processing terminal, the edge server determines a corresponding second model from the second model set according to the service request data, and sends the second model to the first service processing terminal;
[0011] The edge server obtains the current environment data of the environment where the first service processing terminal is located and the current operation data of the first service processing terminal;
[0012] The edge server sends the current environment data, the current operation data, and the first control instruction to the first service processing terminal, so as to control the first service processing terminal to adjust the second model according to the current environment data, the current operation data, and a preset second model adjustment algorithm to obtain a third model;
[0013] The edge server controls the first service processing terminal to process service data by using the third model, and controls the first service processing terminal to perform corresponding operations according to the processing results.
[0014] Optionally, the edge server obtains a plurality of first model sets corresponding to the service types of the services it processes. The steps for the cloud server to generate the first model sets according to specific service categories and service contents include:
[0015] The edge server sends a service type request to the cloud server. The service type request includes: the unique identification information of the edge server; the hardware configuration information of the edge server, including the processor type, memory size, and storage capacity; the geographical location information of the area served by the edge server; the list of service types currently processed by the edge server, and each service type includes a specific service category identifier and service content description;
[0016] The cloud server performs the following operations according to the service type request: confirm the authorization status of the edge server according to the unique identification information of the edge server; filter out the model types that match the hardware conditions according to the hardware configuration information of the edge server; select the training data set of the corresponding area according to the geographical location information of the area served by the edge server; for each service type: select the corresponding basic model from the model library according to the service category identifier; determine the specific parameter configuration of the model according to the service content description; use the selected training data set to train and optimize the basic model; generate an optimized model adapted to the service type;
[0017] The cloud server packages the generated multiple optimized models into a first model set, and sends a model distribution request to the edge server. The model distribution request includes: the metadata information of each optimized model, including the model version, applicable service type, and resource requirements; model deployment configuration information; model update strategy;
[0018] The edge server receives the model distribution request and performs the following operations: verify the integrity and security of the model set; create a model running environment according to the model deployment configuration information; store the first model set in a preset model library; set a regular inspection and update mechanism according to the model update strategy.
[0019] Optionally, the step of the edge server selecting a corresponding first service processing terminal from the service processing terminals according to the task plans and data processing requirements of multiple service processing terminals includes:
[0020] The edge server establishes a terminal management table, which includes: the device identifier of each service processing terminal, the network connection status of each service processing terminal, the task plan information of each service processing terminal, and the data processing requirement information of each service processing terminal; wherein, the task plan information includes: task type identifier, task execution time period, task priority, and task plan status; the data processing requirement information of each service processing terminal includes: processing type, data scale, real-time requirement, and accuracy requirement;
[0021] The edge server performs terminal classification according to the terminal management table, specifically including: grouping terminals by task type; grouping terminals by data processing type; sorting terminals according to task priority within each group; marking terminals with multitasking capabilities;
[0022] The edge server selects the first service processing terminal according to the following conditions, including: judging the task time overlap degree and selecting terminals with matching time resources; judging the matching degree between the data processing capabilities and requirements of the terminals; evaluating the stability of the network connection status; calculating the comprehensive score of task priority; wherein, the calculation formula for the comprehensive score of task priority is: score = w1×time matching degree + w2×processing capacity matching degree + w3×network stability + w4×task priority, where w1, w2, w3, and w4 are preset weight coefficients;
[0023] The edge server establishes a connection with the selected first service processing terminal, including: sending a connection request to the first service processing terminal; waiting for the terminal to respond and verifying the connection status; establishing a secure communication channel; synchronizing the task plan and data processing configuration; updating the connection status in the terminal management table;
[0024] The edge server establishes a task execution monitoring mechanism, including: regularly checking the connection status; monitoring the task execution progress; recording the data processing performance metrics; dynamically adjusting the terminal selection strategy according to the monitoring results.
[0025] Optionally, the step of the edge server generating a second model set according to the terminal attribute data, a preset first model adjustment algorithm, and the first model set includes:
[0026] The edge server parses and classifies the terminal attribute data to determine the terminal hardware attributes, terminal software attributes, and terminal function attributes; among them, the terminal hardware attributes include: processor architecture type, processor performance parameters, memory capacity and type, and whether there is a dedicated AI accelerator; the terminal software attributes include: operating system type and version, AI framework support, and device driver compatibility; the terminal function attributes include: supported sensor types, actuator types, and communication interface types;
[0027] The edge server performs model analysis on each model in the first model set to obtain model structure information, model running requirements, and model application scenarios; among them, the model structure information includes: number and type of network layers, parameter scale, and computational complexity; the model running requirements include: minimum computational resource requirements, memory occupancy requirements, and response time requirements; the model application scenarios include: applicable task types, input data requirements, and output format requirements;
[0028] The edge server executes the first model adjustment algorithm, which specifically includes: adjusting the implementation method of the model calculation layer according to the terminal processor architecture; optimizing the parameter storage structure according to the terminal memory capacity; adjusting the operator implementation method according to whether there is an AI accelerator; selecting the corresponding quantization precision from the predefined set of quantization precision levels according to the terminal performance, performing weight quantization, and performing activation value quantization; calculating the parameter importance score, removing low-importance parameters, and retraining the pruned model structure;
[0029] The edge server performs adaptability verification on the optimized model, including: simulating the terminal hardware environment and configuring the target software environment; testing the computational performance, memory occupancy, and response latency, and testing various input scenarios; verifying the output accuracy; checking the exception handling mechanism;
[0030] The edge server integrates the optimized models that pass the verification into a second model set, including: generating a model description file; packaging the model files; configuring the model loading parameters; setting the version control information.
[0031] Optionally, the steps for the edge server to determine the service processing party according to the terminal attribute data, its own service processing capabilities, and resource occupancy rate include:
[0032] The edge server establishes a resource evaluation index system, which includes computational resource indexes, network resource indexes, and storage resource indexes; among them, the computational resource indexes include: CPU usage rate and available core numbers, GPU / NPU usage, memory usage rate and available capacity; the network resource indexes include: network bandwidth occupancy rate, network latency level, connection stability; the storage resource indexes include: storage space usage rate, I / O performance indexes, cache hit rate;
[0033] The edge server constructs a processing capacity evaluation model, including:
[0034] Evaluating its own processing capacity: calculating the current task queue length of the edge server; counting the average processing delay; predicting the short-term load trend;
[0035] Evaluating the first service processing terminal: analyzing the computing power of the first service processing terminal; evaluating the resource margin of the first service processing terminal; calculating the task adaptation degree;
[0036] Evaluating other processing terminals: obtaining the real-time status of other terminals; evaluating the cooperative processing ability; calculating the load distribution ratio;
[0037] The edge server performs multi-dimensional decision-making calculations, including:
[0038] Calculating the local processing cost: Cost_local = q1 × computing resource occupancy + q2 × storage resource occupancy;
[0039] Calculating the terminal processing cost: Cost_terminal = q3 × terminal resource occupancy + q4 × network transmission overhead;
[0040] Calculating the cooperative processing cost: Cost_cooperative = q5 × task decomposition overhead + q6 × coordination communication overhead;
[0041] Calculating the processing benefits, including: calculating the processing time benefit; calculating the resource utilization benefit; calculating the system stability benefit;
[0042] Generating a decision score, including: Score = benefit metric - cost metric; adjusting with application constraints; calculating the final decision score;
[0043] The edge server establishes a dynamic load balancing mechanism, including:
[0044] Real-time monitoring of the load status, specifically: monitoring the resource changes of each processing party; recording the task processing performance; analyzing the load distribution;
[0045] Performing load adjustment, specifically: setting the load threshold; calculating the load deviation; triggering load migration;
[0046] Updating the processing strategy, specifically: recording the processing effect; optimizing the decision parameters; adjusting the weight coefficients;
[0047] The edge server determines the final service processing party according to the following rules, including:
[0048] When Score_local is the highest and meets the threshold requirements, select the edge server as the processing party; the edge server allocates and locks a specific number of processor resources, memory resources, and storage resources from the available computing resource pool according to the resource requirements of business processing, forms an independent resource space, and initializes the processing environment in the resource space;
[0049] When Score_terminal is the highest and meets the threshold requirements, select the first business processing terminal as the processing party; configure task parameters for the first business processing terminal and establish a monitoring mechanism;
[0050] When Score_cooperative is the highest and meets the threshold requirements, select the cooperative processing mode, divide the task boundary, and establish a cooperation mechanism.
[0051] Optionally, the step that when determining that the business processing party is the first business processing terminal, the edge server determines the corresponding second model from the second model set according to the service request data and sends the second model to the first business processing terminal includes:
[0052] The edge server parses the service request data, including:
[0053] Parse the service type information, specifically: identify the main service type; identify the sub-task type; determine the task priority;
[0054] Parse the data feature information, specifically: analyze the data scale; identify the data format; determine the data quality requirements;
[0055] Parse the performance requirement information, specifically: obtain the real-time requirement; obtain the accuracy requirement; obtain the resource constraint conditions;
[0056] The edge server filters candidate models from the second model set, including:
[0057] Execute feature matching, specifically: match the service type label; match the data format requirements; match the performance index requirements;
[0058] Evaluate the model applicability, specifically: calculate the feature matching degree; evaluate the performance satisfaction degree; calculate the resource adaptation degree;
[0059] Generate a candidate model list, specifically: sort by the matching degree; filter out models that do not meet the constraints; record the model score;
[0060] The edge server performs model combination optimization, including:
[0061] Analyze the task dependency relationship, specifically: construct a task dependency graph; identify the critical path; determine the execution order;
[0062] Computing resource constraints, specifically: counting available resources; calculating resource requirements; evaluating resource matching degree;
[0063] Generating an optimal combination plan, specifically: calculating combination benefits; evaluating combination feasibility; selecting the optimal plan;
[0064] The edge server performs model packaging operations, including:
[0065] Preparing model files, specifically: organizing model parameters; packaging model structures; generating configuration files;
[0066] Creating a deployment description, specifically: recording dependencies; setting deployment sequences; configuring operating parameters;
[0067] Performing integrity verification, specifically: verifying file integrity; checking configuration correctness; generating verification information;
[0068] The edge server performs model distribution, including:
[0069] Establishing a secure transmission channel, specifically: negotiating encryption methods; establishing a data channel; verifying connection security;
[0070] Performing batch transmission, specifically: setting transmission priorities; controlling transmission rates; monitoring transmission status;
[0071] Verifying distribution results, specifically: checking reception integrity; verifying model availability; confirming successful deployment.
[0072] Optionally, the edge server sends the current environment data, the current operation data, and the first control instruction to the first service processing terminal to control the first service processing terminal to adjust the second model according to the current environment data, the current operation data, and a preset second model adjustment algorithm to obtain a third model. The steps include:
[0073] The edge server collects and preprocesses environment data, including:
[0074] Collecting physical environment data such as temperature, humidity, light, vibration, noise, spatial position, and movement state;
[0075] Collecting service environment data such as the distribution of surrounding devices, network environment conditions, and electromagnetic interference conditions;
[0076] Performing data preprocessing on the physical environment data and the service environment data, specifically: data standardization processing; outlier detection and processing; data time series alignment;
[0077] The edge server collates the current running data, which includes resource usage status, network transmission status, and hardware operation parameters. Among them, the resource usage status includes CPU, GPU utilization rate, memory occupancy, and storage space usage; the network transmission status includes real-time bandwidth utilization rate, network latency data, and packet loss rate; the hardware operation parameters include processor temperature, power consumption data, and working status of each component.
[0078] The edge server generates a first control instruction, which includes a model adjustment strategy, a resource allocation strategy, and an execution control strategy. Among them, the model adjustment strategy includes determining the adjustment target, setting the adjustment range, and defining the adjustment step size; the resource allocation strategy includes allocating computing resources, setting memory limits, and controlling the energy consumption level; the execution control strategy includes setting the execution priority, defining the timeout mechanism, and configuring exception handling.
[0079] The edge server integrates and sends data packets, including:
[0080] Data packing, specifically: merging environmental data; integrating running data; including control instructions.
[0081] Data compression, specifically: selecting a compression algorithm; performing data compression; generating a checksum.
[0082] Sending in batches, specifically: determining the sending order; controlling the sending rate; monitoring the sending status.
[0083] The edge computing module of the first service processing terminal performs model adjustment, including:
[0084] Parsing the received data, specifically: decompressing the data packet; verifying data integrity; classifying and storing the data.
[0085] Evaluating the necessity of adjustment, specifically: analyzing the degree of environmental change; evaluating the degree of performance impact; calculating the adjustment benefit.
[0086] Performing model adjustment, specifically: applying the second model adjustment algorithm; performing parameter fine-tuning; optimizing the model structure.
[0087] The edge computing module verifies the adjusted third model, including:
[0088] Performing performance testing, specifically: testing computing performance; testing response time; testing resource occupancy.
[0089] Verifying the functional correctness, specifically: testing model accuracy; verifying output stability; checking for abnormal responses.
[0090] Generating a verification report, specifically: recording the adjustment effect; statistical performance indicators; providing optimization suggestions.
[0091] Another aspect of the present invention provides an AI edge computing device for performing an AI edge computing method, including: a data acquisition module and a control processing module; wherein,
[0092] The data acquisition module is configured to: acquire a plurality of first model sets corresponding to the service types of the services processed by itself, and the first model sets are generated by a cloud server according to specific service categories and service contents;
[0093] The control processing module is configured to: select a corresponding first service processing terminal for connection from the service processing terminals according to the task plans and data processing requirements of a plurality of service processing terminals;
[0094] The data acquisition module is further configured to: acquire the terminal attribute data and service request data of the first service processing terminal;
[0095] The control processing module is further configured to:
[0096] generate a second model set according to the terminal attribute data, a preset first model adjustment algorithm, and the first model sets;
[0097] determine a service processing party according to the terminal attribute data, its own service processing capabilities, and resource occupancy rate;
[0098] When it is determined that the service processing party is the first service processing terminal, determine a corresponding second model from the second model set according to the service request data, and send the second model to the first service processing terminal;
[0099] The data acquisition module is further configured to: acquire the current environment data of the environment where the first service processing terminal is located and the current operation data of the first service processing terminal;
[0100] The control processing module is further configured to:
[0101] send the current environment data, the current operation data, and a first control instruction to the first service processing terminal to control the first service processing terminal to adjust the second model according to the current environment data, the current operation data, and a preset second model adjustment algorithm to obtain a third model;
[0102] control the first service processing terminal to process service data using the third model, and control the first service processing terminal to perform corresponding operations according to the processing results.
[0103] Another aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of an AI edge computing method are implemented.
[0104] Another aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of an AI edge computing method are implemented.
[0105] Adopting the technical solution of the present invention, the AI edge computing method includes: the edge server obtains a plurality of first model sets corresponding to the service types of the services it processes, and the first model sets are generated by the cloud server according to specific service types and service contents; the edge server selects corresponding first service processing terminals for connection from the service processing terminals according to the task plans and data processing requirements of a plurality of service processing terminals; the edge server obtains the terminal attribute data and service request data of the first service processing terminal; the edge server generates a second model set according to the terminal attribute data, a preset first model adjustment algorithm and the first model sets; the edge server determines the service processing party according to the terminal attribute data, its own service processing capacity and resource occupancy rate; when it is determined that the service processing party is the first service processing terminal, the edge server determines a corresponding second model from the second model set according to the service request data, and sends the second model to the first service processing terminal; the edge server obtains the current environment data of the environment where the first service processing terminal is located and the current operation data of the first service processing terminal; the edge server sends the current environment data, the current operation data and a first control instruction to the first service processing terminal to control the first service processing terminal to adjust the second model according to the current environment data, the current operation data and a preset second model adjustment algorithm to obtain a third model; the edge server controls the first service processing terminal to process service data by using the third model, and controls the first service processing terminal to perform corresponding operations according to the processing results. Through the solution of the present invention, by dynamically selecting a suitable model according to the task plan and data processing requirements of the service processing terminal, the edge server can utilize computing resources more efficiently, reduce unnecessary calculations and delays; the edge server adjusts the model according to the terminal attribute data and current environment data, enabling the model to adapt to different service scenarios and environmental changes in real time, thereby improving the accuracy and reliability of processing; by real-time monitoring the computing resource occupancy rate and network transmission rate of the terminal, the edge server can intelligently allocate resources, avoid resource waste, and ensure the high efficiency of service processing; the implementation of edge computing makes data processing closer to the data source, reduces the time delay of data transmission, and is particularly suitable for application scenarios with high real-time requirements, such as autonomous driving, industrial monitoring, etc.; by obtaining the attribute data and service request data of the terminal, the edge server can provide more personalized services to meet the specific needs of different users and services; by dynamically adjusting the model and real-time monitoring the system state, it can effectively respond to emergencies and environmental changes, and improve the overall robustness and stability of the system. The solution of the present invention can not only improve the performance of AI edge computing, but also achieve more efficient and intelligent service processing in various application scenarios. Brief Description of the Drawings
[0106] Figure 1 It is a flowchart of the AI edge computing method provided by an embodiment of the present invention;
[0107] Figure 2 It is a schematic block diagram of the AI edge computing device provided by an embodiment of the present invention. Detailed implementation manners
[0108] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.
[0109] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0110] The terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0111] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0112] Next, refer to Figures 1 to 2 to describe an AI edge computing method, device, electronic device and storage medium provided according to some embodiments of the present invention.
[0113] As Figure 1 shown, an embodiment of the present invention provides an AI edge computing method, including:
[0114] The edge server obtains a plurality of first model sets corresponding to the service types of the services it processes, and the first model sets are generated by the cloud server according to specific service types and service contents;
[0115] The edge server selects corresponding one or more first service processing terminals for connection from the service processing terminals according to the task plans (such as navigation, inspection, translation, etc.) and data processing requirements (such as image recognition, speech recognition, etc.) of multiple service processing terminals;
[0116] The edge server obtains the terminal attribute data and service request data of the first service processing terminal;
[0117] It can be understood that in this step, first of all, the edge server needs to define the types of terminal attribute data required, such as device model, operating system version, hardware configuration, network status, etc.; the edge server requests the attribute data of the terminal through the communication interface (such as REST API or WebSocket) with the first service processing terminal, and the terminal can actively send these data to the edge server regularly or when a specific event occurs (such as startup, status change); ensure that the obtained data is transmitted in a standard format (such as JSON or XML) for easy parsing and storage by the edge server; the edge server clarifies the content of the service request data, including information such as request type, request parameters, timestamp, etc.; the edge server can set up a listening mechanism to receive the service request data from the first service processing terminal in real time, which can be achieved by establishing a persistent connection (such as WebSocket) or using a polling mechanism; whenever the service request data is received, the edge server records and stores it in the local database for subsequent processing and analysis; the edge server integrates the terminal attribute data and service request data to form a comprehensive data set. This process can be achieved through data processing algorithms to ensure the consistency and integrity of the data; the integrated data can be stored in the database of the edge server for subsequent query and analysis; the edge server can regularly request updated terminal attribute data and service request data from the first service processing terminal to ensure the timeliness of the data; the edge server can feedback the data processing results or status information to the terminal so that the terminal can make corresponding adjustments according to the feedback. Through this step, the edge server can effectively obtain the terminal attribute data and service request data of the first service processing terminal, providing necessary information support for subsequent service processing; by obtaining the terminal attributes and service request data in real time, the edge server can ensure the accuracy and timeliness of the data; the acquisition and processing of real-time data enable the system to quickly respond to user requests and improve the user experience; by analyzing the terminal attribute data, the edge server can better optimize resource allocation and improve the overall system performance.
[0118] The edge server generates a second model set according to the terminal attribute data, a preset first model adjustment algorithm, and the first model set;
[0119] The edge server determines the service processing party according to the terminal attribute data, its own service processing capabilities, and resource occupancy rate (that is, selects the service processing party from among the edge server, the first service processing terminal, or other processing terminals);
[0120] When it is determined that the service processing party is the first service processing terminal, the edge server determines the corresponding second model(s) from the second model set according to the service request data, and sends the second model to the first service processing terminal;
[0121] The edge server obtains the current environmental data of the environment where the first service processing terminal is located and the current operation data of the first service processing terminal (including computing resource occupancy rate, network transmission rate, operation parameters of each component, etc.);
[0122] It can be understood that the edge server obtains the current environmental data of the environment where the first service processing terminal is located through sensors or API interfaces (these data may include environmental parameters such as temperature, humidity, light intensity, noise level, etc., which can be collected by the first service processing terminal itself or other sensors); the edge server monitors the operation status of the first service processing terminal and obtains the current operation data, which includes: computing resource occupancy rate (by monitoring the usage of CPU, memory, and storage, the edge server can evaluate the computing resource occupancy rate of the terminal), network transmission rate (through network monitoring tools, the edge server can obtain the current network bandwidth usage and data transmission rate), operation parameters of each component (the edge server can obtain the operation status and performance parameters of each component (such as sensors, actuators, etc.) through the monitoring system of the terminal); the edge server integrates the collected environmental data and operation data to form a comprehensive data set. This data set can be used for subsequent analysis and decision support; the edge server updates the environmental data and operation data regularly or according to an event trigger mechanism (such as data change, task request, etc.) to ensure that the obtained information is up-to-date; the collected data can be transmitted to the edge server through a secure communication protocol (such as HTTPS) and stored in a local database for subsequent access and analysis. Through this step, the edge server can effectively obtain the environmental data and operation data of the first service processing terminal, thereby providing support for subsequent service processing and decision-making; it can obtain and monitor the environment and operation status of the terminal in real time, improving the response ability of the system; by analyzing the computing resource occupancy rate and network transmission rate, the edge server can optimize resource allocation and improve the overall system performance; the integrated environmental data and operation data provide an important basis for business decision-making, helping the system better adapt to changing environments and requirements.
[0123] The edge server sends the current environment data, the current operation data, and the first control instruction to the first service processing terminal, so as to control the first service processing terminal (an edge computing module is configured on the first service processing terminal) to adjust the second model according to the current environment data, the current operation data, and a preset second model adjustment algorithm to obtain a third model;
[0124] The edge server controls the first service processing terminal to process service data by using the third model, and controls the first service processing terminal to execute corresponding operations according to the processing results.
[0125] In this embodiment, by dynamically selecting an appropriate model according to the task plan and data processing requirements of the service processing terminal, the edge server can more efficiently utilize computing resources, reduce unnecessary calculations and delays; the edge server adjusts the model according to the terminal attribute data and the current environment data, so that the model can adapt to different service scenarios and environmental changes in real time, thereby improving the accuracy and reliability of processing; by real-time monitoring the computing resource occupancy rate and network transmission rate of the terminal, the edge server can intelligently allocate resources, avoid resource waste, and ensure the efficiency of service processing; the implementation of edge computing makes data processing closer to the data source, reduces the time delay of data transmission, and is particularly suitable for application scenarios with high real-time requirements, such as autonomous driving and industrial monitoring; by obtaining the attribute data and service request data of the terminal, the edge server can provide more personalized services to meet the specific needs of different users and services; by dynamically adjusting the model and real-time monitoring the system status, it can effectively respond to emergencies and environmental changes, and improve the overall robustness and stability of the system. The solution of the present invention can not only improve the performance of AI edge computing, but also achieve more efficient and intelligent service processing in various application scenarios.
[0126] In some possible implementation manners of the present invention, the edge server obtains a plurality of first model sets corresponding to the service types of the services processed by itself, and the steps for the first model set to be generated by the cloud server according to specific service types and service contents include:
[0127] The edge server sends a service type request to the cloud server, and the service type request includes: the unique identification information of the edge server; the hardware configuration information of the edge server, including the processor type, memory size, and storage capacity; the geographical location information of the area served by the edge server; the service type list currently processed by the edge server, and each service type includes a specific service type identifier and service content description;
[0128] The cloud server requests to perform the following operations according to the service type: confirm the authorization status of the edge server based on the unique identification information of the edge server; filter out the model types that match the hardware conditions according to the hardware configuration information of the edge server; select the training data set for the corresponding region according to the geographical location information of the region served by the edge server; for each service type: select the corresponding basic model from the model library according to the service type identifier; determine the specific parameter configuration of the model according to the service content description; use the selected training data set to train and optimize the basic model; generate an optimized model adapted to the service type;
[0129] The cloud server packages the generated multiple optimized models into a first model set and sends a model distribution request to the edge server. The model distribution request includes: the metadata information of each optimized model, including the model version, applicable service type, and resource requirements; the model deployment configuration information; the model update policy;
[0130] The edge server receives the model distribution request and performs the following operations: verify the integrity and security of the model set; create a model running environment according to the model deployment configuration information; store the first model set in a preset model library; set up a regular inspection and update mechanism according to the model update policy.
[0131] The solution of this embodiment ensures that the generated model set can best adapt to the actual application scenario by considering the hardware conditions, geographical location, and specific service requirements of the edge server; filters models based on the actual configuration of the edge server to avoid resource waste or insufficiency; ensures the reliability and security of the model distribution process through a complete request-response mechanism; establishes a model version control and update mechanism to ensure that the model always maintains the optimal state.
[0132] In some possible implementation manners of the present invention, the step of the edge server selecting the corresponding first service processing terminal for connection from the service processing terminals according to the task plans and data processing requirements of multiple service processing terminals includes:
[0133] The edge server establishes a terminal management table, which includes: the device identifiers of each service processing terminal, the network connection status of each service processing terminal, the task plan information of each service processing terminal, and the data processing requirement information of each service processing terminal; wherein, the task plan information includes: the task type identifier, the task execution time period, the task priority, and the task plan status (not started, in execution, completed); the data processing requirement information of each service processing terminal includes: the processing type (such as image recognition, voice recognition, etc.), the data scale, the real-time requirement, and the accuracy requirement;
[0134] The edge server performs terminal classification according to the terminal management table, specifically including: grouping terminals by task type; grouping terminals by data processing type; within each group, sorting terminals according to task priority; marking terminals with multitasking capabilities;
[0135] The edge server selects the first service processing terminal according to the following conditions, including: judging the task time overlap degree and selecting a terminal with matching time resources; judging the matching degree between the data processing ability of the terminal and the requirements; evaluating the stability of the network connection status; calculating the comprehensive score of task priority; where the calculation formula for the comprehensive score of task priority is: score = w1×time matching degree + w2×processing ability matching degree + w3×network stability + w4×task priority, and w1, w2, w3, and w4 are preset weight coefficients;
[0136] The edge server establishes a connection with the selected first service processing terminal, including: sending a connection request to the first service processing terminal; waiting for the terminal to respond and verifying the connection status; establishing a secure communication channel; synchronizing the task plan and data processing configuration; updating the connection status in the terminal management table;
[0137] The edge server establishes a task execution monitoring mechanism, including: regularly checking the connection status; monitoring the task execution progress; recording data processing performance metrics; dynamically adjusting the terminal selection strategy according to the monitoring results.
[0138] The solution of this embodiment realizes the optimal allocation of computing resources through the terminal classification and priority scoring mechanism; improves the overall processing efficiency of the system by identifying terminals with multitasking capabilities; ensures the dynamic balance of system load through real-time monitoring and strategy adjustment; and guarantees the stability and reliability of service processing through a complete connection establishment and monitoring mechanism.
[0139] In some possible implementation manners of the present invention, the step in which the edge server generates a second model set according to the terminal attribute data, a preset first model adjustment algorithm, and the first model set includes:
[0140] The edge server parses and classifies the terminal attribute data to determine the terminal hardware attributes, terminal software attributes, and terminal function attributes; where the terminal hardware attributes include: processor architecture type (such as ARM, x86, etc.), processor performance parameters (such as the number of cores, main frequency, etc.), memory capacity and type, and whether there is a dedicated AI accelerator; the terminal software attributes include: operating system type and version, AI framework support situation, and device driver compatibility; the terminal function attributes include: supported sensor types, actuator types, and communication interface types;
[0141] The edge server performs model analysis on each model in the first model set to obtain model structure information, model running requirements, and model application scenarios. Among them, the model structure information includes: the number and type of network layers, parameter scale, and computational complexity; the model running requirements include: minimum computational resource requirements, memory occupancy requirements, and response time requirements; the model application scenarios include: applicable task types, input data requirements, and output format requirements.
[0142] The edge server executes the first model adjustment algorithm, which specifically includes: adjusting the implementation method of the model calculation layer according to the terminal processor architecture; optimizing the parameter storage structure according to the terminal memory capacity; adjusting the operator implementation method according to whether there is an AI accelerator; selecting the corresponding quantization precision from a predefined set of quantization precision levels according to the terminal performance, performing weight quantization, and performing activation value quantization; calculating the parameter importance score, removing low-importance parameters, and retraining the pruned model structure.
[0143] It can be understood that in this step, according to the terminal performance metrics, the corresponding quantization precision (such as a specific numerical representation of the bit width) is selected from a predefined set of quantization precision levels. The set of quantization precision levels includes, but is not limited to, 8-bit integer, 4-bit integer, and 2-bit integer quantization precisions. Among them, the quantization precision with a lower bit width corresponds to lower computational resource requirements but may result in higher precision loss. The aforementioned selection operation is based on a preset balance relationship between the terminal processing capacity and the model precision requirements.
[0144] The edge server performs adaptability verification on the optimized model, including: simulating the terminal hardware environment and configuring the target software environment; testing the computational performance, memory occupancy, and response latency, and testing various input scenarios; verifying the output accuracy; checking the exception handling mechanism.
[0145] The edge server integrates the optimized models that pass the verification into a second model set, including: generating a model description file; packaging the model files; configuring the model loading parameters; setting the version control information.
[0146] The solution of this embodiment ensures that the generated model can run efficiently on the target terminal by comprehensively considering the hardware, software, and functional attributes of the terminal; significantly reduces the model resource occupancy while ensuring the function through model structure optimization, quantization, and pruning; ensures the correctness and stability of the optimized model through a complete verification mechanism; and can generate differentiated optimized models according to the characteristics of different terminals.
[0147] In some possible implementation manners of the present invention, it further includes the step of constructing the first model adjustment algorithm, which specifically includes:
[0148] Build a model evaluation index system, including: performance indicators (such as computing latency, memory occupancy, energy consumption level, etc.), accuracy indicators (such as task accuracy rate, error range, recall rate, etc.), and resource utilization indicators (such as CPU utilization rate, memory usage efficiency, bandwidth occupancy rate, etc.);
[0149] Establish a model structure analysis mechanism, including:
[0150] Analyze the model computation graph: identify computation-intensive nodes; mark memory-intensive nodes; calculate the data dependencies between nodes;
[0151] Evaluate the characteristics of operators: count the usage frequencies of various operators; calculate the resource consumption of operators; analyze the optimization space of operators;
[0152] Build an adaptive optimization strategy generator, including:
[0153] According to the architecture characteristics of the terminal processor: generate instruction set optimization strategies; design parallel computing strategies; formulate storage access strategies;
[0154] According to the memory capacity of the terminal: generate data caching strategies; design parameter compression strategies; formulate dynamic loading strategies;
[0155] According to the functional characteristics of the terminal: generate task decomposition strategies; design resource scheduling strategies; formulate load balancing strategies;
[0156] Implement a model converter, including:
[0157] Design a quantization conversion module: implement dynamic quantization algorithms; implement mixed-precision quantization; implement automatic adjustment of quantization parameters;
[0158] Design a pruning conversion module: implement structural importance evaluation; implement iterative pruning algorithms; implement automatic weight compensation;
[0159] Design a structure reorganization module: implement layer fusion optimization; implement branch merging optimization; implement computation graph rearrangement;
[0160] Build an optimization effect verification mechanism, including:
[0161] Establish a benchmark test set: different scales of input data; different scenario test cases; boundary condition tests;
[0162] Implement performance comparison and analysis: calculate the changes in indicators before and after optimization; evaluate the stability of the optimization effect; generate an optimization report;
[0163] Establish an optimization strategy feedback mechanism: record the historical optimization data; analyze the laws of the optimization effect; update the parameters of the optimization strategy;
[0164] Implement an online dynamic adjustment mechanism, including:
[0165] Build a real-time monitoring module: Monitor the status of terminal resources; Track the running metrics of the model; Detect performance anomalies;
[0166] Implement a dynamic optimization trigger: Set the optimization trigger threshold; Evaluate the necessity of optimization; Control the optimization frequency;
[0167] Establish an optimization strategy selector: Predict the optimization effect; Evaluate the optimization cost; Select the optimal strategy.
[0168] In this embodiment, through the adaptive optimization strategy generator, the intelligent adjustment of the model is realized; Through a complete evaluation and verification mechanism, the performance of the optimized model is ensured; Through multi-dimensional optimization strategies, the efficient utilization of terminal resources is realized; Through the online dynamic adjustment mechanism, the model can adapt to the changes in the terminal operating environment.
[0169] In some possible implementation manners of the present invention, the steps for the edge server to determine the service processing party according to the terminal attribute data, its own service processing ability and resource occupancy rate include:
[0170] The edge server establishes a resource evaluation index system, and the resource evaluation index system includes computing resource indexes, network resource indexes and storage resource indexes; Among them, the computing resource indexes include: CPU usage rate and available core number, GPU / NPU usage, memory usage rate and available capacity; The network resource indexes include: network bandwidth occupancy rate, network latency level, connection stability; The storage resource indexes include: storage space usage rate, I / O performance index, cache hit rate;
[0171] The edge server constructs a processing ability evaluation model, including:
[0172] Evaluate its own processing ability: Calculate the current task queue length of the edge server; Statistically analyze the average processing delay; Predict the short-term load trend;
[0173] Evaluate the first service processing terminal: Analyze the computing ability of the first service processing terminal; Evaluate the resource margin of the first service processing terminal; Calculate the task adaptability;
[0174] Evaluate other processing terminals: Obtain the real-time status of other terminals; Evaluate the collaborative processing ability; Calculate the load distribution ratio;
[0175] The edge server performs multi-dimensional decision-making calculations, including:
[0176] Calculate the local processing cost: Cost_local = q1×computing resource occupancy + q2×storage resource occupancy;
[0177] Calculate the terminal processing cost: Cost_terminal = q3 × Terminal resource occupancy + q4 × Network transmission overhead;
[0178] Calculate the cooperative processing cost: Cost_cooperative = q5 × Task decomposition overhead + q6 × Coordination communication overhead;
[0179] Calculate the processing benefits, including: Calculation processing time benefit; Calculation resource utilization benefit; Calculation system stability benefit;
[0180] Generate a decision score, including: Score = Benefit metric - Cost metric; Apply constraint conditions for adjustment; Calculate the final decision score;
[0181] It can be understood that q1 to q6 are preset weight values. Applying constraint conditions for adjustment is a process of considering the impact of various constraint conditions on the score and making corresponding adjustments based on the original decision score. The types of constraint conditions include hard constraints and soft constraints; Hard constraints, that is, conditions that must be met, such as: The use of computing resources cannot exceed the maximum system capacity, the network bandwidth cannot exceed the physical limit, the response time must meet the real-time requirement, the memory usage must be within the safe range, etc.; Soft constraints, that is, conditions that are expected to be met but can be appropriately violated, such as: The processor utilization rate is preferably maintained within a certain range, the energy consumption level is preferably controlled below the target value, the load balance degree is preferably maintained within the ideal range, etc.
[0182] The edge server establishes a dynamic load balancing mechanism, including:
[0183] Real-time monitor the load status, specifically: Monitor the resource changes of each processing party; Record the task processing performance; Analyze the load distribution;
[0184] Execute load adjustment, specifically: Set the load threshold; Calculate the load deviation; Trigger load migration;
[0185] Update the processing strategy, specifically: Record the processing effect; Optimize the decision parameters; Adjust the weight coefficients;
[0186] The edge server determines the final service processing party according to the following rules, including:
[0187] When Score_local is the highest and meets the threshold requirement, select the edge server as the processing party; The edge server allocates and locks a specific number of processor resources, memory resources, and storage resources from the available computing resource pool according to the resource requirements of the service processing, forms an independent resource space, and initializes the processing environment in the resource space;
[0188] When Score_terminal is the highest and meets the threshold requirements, select the first service processing terminal as the processing party; configure task parameters for the first service processing terminal and establish a monitoring mechanism;
[0189] When Score_cooperative is the highest and meets the threshold requirements, select the cooperative processing mode, divide the task boundary, and establish a cooperation mechanism.
[0190] The solution of this embodiment realizes the optimal allocation of system resources through a multi-dimensional evaluation and decision-making mechanism; avoids resource overload or idleness through a dynamic load balancing mechanism; improves the overall service processing efficiency through an intelligent selection of the processing party; and ensures the stability of system operation through reasonable resource allocation and load control.
[0191] In some possible implementation manners of the present invention, the step that when it is determined that the service processing party is the first service processing terminal, the edge server determines a corresponding second model from the second model set according to the service request data and sends the second model to the first service processing terminal includes:
[0192] The edge server parses the service request data, including:
[0193] Parsing service type information, specifically: identifying the main service type; identifying sub-task types; determining task priorities;
[0194] Parsing data feature information, specifically: analyzing data scale; identifying data formats; determining data quality requirements;
[0195] Parsing performance requirement information, specifically: obtaining real-time requirements; obtaining accuracy requirements; obtaining resource limitation conditions;
[0196] The edge server screens candidate models from the second model set, including:
[0197] Performing feature matching, specifically: matching service type tags; matching data format requirements; matching performance index requirements;
[0198] Evaluating model applicability, specifically: calculating feature matching degrees; evaluating performance satisfaction; calculating resource adaptation degrees;
[0199] Generating a list of candidate models, specifically: sorting by matching degree; filtering models that do not meet the constraints; recording model scores;
[0200] The edge server performs model combination optimization, including:
[0201] Analyzing task dependencies, specifically: constructing a task dependency graph; identifying critical paths; determining the execution order;
[0202] Computing resource constraints, specifically: counting the available resources; calculating the resource requirements; evaluating the resource matching degree;
[0203] Generating an optimal combination plan, specifically: calculating the combination benefits; evaluating the combination feasibility; selecting the optimal plan;
[0204] The edge server performs model packaging operations, including:
[0205] Preparing the model file, specifically: sorting out the model parameters; packaging the model structure; generating the configuration file;
[0206] Creating a deployment description, specifically: recording the dependencies; setting the deployment sequence; configuring the running parameters;
[0207] Performing integrity verification, specifically: verifying the file integrity; checking the configuration correctness; generating the verification information;
[0208] The edge server performs model distribution, including:
[0209] Establishing a secure transmission channel, specifically: negotiating the encryption method; establishing the data channel; verifying the connection security;
[0210] Performing batch transmission, specifically: setting the transmission priority; controlling the transmission rate; monitoring the transmission status;
[0211] Verifying the distribution result, specifically: checking the reception integrity; verifying the model availability; confirming the successful deployment.
[0212] The solution of this embodiment ensures the selection of the most suitable model through multi-dimensional feature matching; realizes the efficient utilization of resources through model combination optimization; ensures the reliability of model deployment through a complete packaging and distribution mechanism; and protects the security of model data through a secure transmission mechanism.
[0213] In some possible implementation manners of the present invention, the step of the edge server sending the current environment data, the current operation data, and the first control instruction to the first service processing terminal to control the first service processing terminal to adjust the second model according to the current environment data, the current operation data, and a preset second model adjustment algorithm to obtain a third model includes:
[0214] The edge server collects and preprocesses the environment data, including:
[0215] Collecting physical environment data such as temperature, humidity, light, vibration, noise, spatial position, and movement state;
[0216] Collecting service environment data such as the distribution of surrounding devices, the network environment status, and the electromagnetic interference situation;
[0217] Perform data preprocessing on the physical environment data and the business environment data, specifically: data standardization processing; outlier detection and processing; data time series alignment;
[0218] The edge server sorts out the current operation data, and the current operation data includes resource usage status, network transmission status, and hardware operation parameters; among them, the resource usage status includes: CPU, GPU usage rate, memory occupancy, and storage space usage; the network transmission status includes: real-time bandwidth utilization rate, network latency data, and packet loss rate; the hardware operation parameters include: processor temperature, power consumption data, and working status of each component;
[0219] The edge server generates a first control instruction, and the first control instruction includes a model adjustment strategy, a resource allocation strategy, and an execution control strategy; among them, the model adjustment strategy includes: determining an adjustment target, setting an adjustment range, and defining an adjustment step size; the resource allocation strategy includes: allocating computing resources, setting a memory limit, and controlling the energy consumption level; the execution control strategy includes: setting an execution priority, defining a timeout mechanism, and configuring exception handling;
[0220] The edge server integrates and sends data packets, including:
[0221] Data packaging, specifically: merging environment data; integrating operation data; including control instructions;
[0222] Data compression, specifically: selecting a compression algorithm; performing data compression; generating a checksum;
[0223] Sending in batches, specifically: determining the sending order; controlling the sending rate; monitoring the sending status;
[0224] The edge computing module of the first business processing terminal performs model adjustment, including:
[0225] Parse the received data, specifically: decompress the data packet; verify the data integrity; classify and store the data;
[0226] Evaluate the necessity of adjustment, specifically: analyze the degree of environmental change; evaluate the degree of performance impact; calculate the adjustment benefit;
[0227] Execute model adjustment, specifically: apply the second model adjustment algorithm; perform parameter fine-tuning; optimize the model structure;
[0228] The edge computing module verifies the adjusted third model, including:
[0229] Perform performance testing, specifically: test computing performance; test response time; test resource occupancy;
[0230] Verify the correctness of the function, specifically: test the model accuracy; verify the output stability; check the abnormal response;
[0231] Generate a verification report, specifically: record the adjustment effect; count the performance metrics; provide optimization suggestions.
[0232] The solution of this embodiment enables the model to adapt to environmental changes through the acquisition and processing of real-time environmental data; improves the model operation efficiency through the analysis and adjustment of operation data; ensures the reliability of model adjustment through a complete data processing and verification mechanism; and realizes the real-time adjustment of the model through an efficient data transmission and processing mechanism.
[0233] Please refer to Figure 2 , another embodiment of the present invention provides an AI edge computing device for executing an AI edge computing method, including: a data acquisition module and a control processing module; wherein,
[0234] The data acquisition module is configured to: acquire a plurality of first model sets corresponding to the service type of the service it processes, and the first model sets are generated by a cloud server according to specific service types and service contents;
[0235] The control processing module is configured to: select a corresponding first service processing terminal for connection from the service processing terminals according to the task plans and data processing requirements of a plurality of service processing terminals;
[0236] The data acquisition module is further configured to: acquire the terminal attribute data and service request data of the first service processing terminal;
[0237] The control processing module is further configured to:
[0238] Generate a second model set according to the terminal attribute data, a preset first model adjustment algorithm, and the first model sets;
[0239] Determine the service processing party according to the terminal attribute data, its own service processing ability, and resource occupancy rate;
[0240] When it is determined that the service processing party is the first service processing terminal, determine a corresponding second model from the second model set according to the service request data, and send the second model to the first service processing terminal;
[0241] The data acquisition module is further configured to: acquire the current environmental data of the environment where the first service processing terminal is located and the current operation data of the first service processing terminal;
[0242] The control processing module is further configured to:
[0243] Send the current environmental data, the current operation data, and the first control instruction to the first service processing terminal to control the first service processing terminal to adjust the second model according to the current environmental data, the current operation data, and a preset second model adjustment algorithm to obtain a third model;
[0244] Control the first service processing terminal to process service data by using the third model, and control the first service processing terminal to perform corresponding operations according to the processing results.
[0245] It should be known that Figure 2 The block diagram of the shown AI edge computing device is only for illustration, and the number of each module shown therein does not limit the protection scope of the present invention. The AI edge computing device provided in this embodiment can be used to execute the solutions of the respective embodiments of the corresponding AI edge computing method. For the specific implementation process, please refer to the descriptions of the respective method embodiments, which will not be elaborated herein.
[0246] Another embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of an AI edge computing method are implemented.
[0247] Another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of an AI edge computing method are implemented.
[0248] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0249] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0250] In several embodiments provided in the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical or other forms.
[0251] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0252] In addition, in each embodiment of the present application, the functional units can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0253] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in each embodiment of the present application. And the aforementioned memory includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs and other media that can store program codes.
[0254] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory. The memory can include: flash drives, read-only memories (abbreviation: ROM), random access memories (abbreviation: RAM), magnetic disks, or optical discs, etc.
[0255] The above has introduced the embodiments of the present application in detail. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
[0256] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions without departing from the spirit and scope of the present invention, and can make various changes and modifications, including combinations of the above different functions and implementation steps, including software and hardware implementation manners, all within the protection scope of the present invention.
Claims
1. An AI edge computing method, characterized in that, Including: The edge server obtains multiple first model sets corresponding to the service types of the services it processes, and the first model sets are generated by the cloud server according to specific service categories and service contents; The edge server selects corresponding first service processing terminals for connection from the service processing terminals according to the task plans and data processing requirements of multiple service processing terminals; The edge server obtains the terminal attribute data and service request data of the first service processing terminal; The edge server generates a second model set according to the terminal attribute data, a preset first model adjustment algorithm, and the first model sets; The edge server determines a service processing party according to the terminal attribute data, its own service processing capabilities, and resource occupancy rate; When it is determined that the service processing party is the first service processing terminal, the edge server determines a corresponding second model from the second model set according to the service request data, and sends the second model to the first service processing terminal; The edge server obtains the current environmental data of the environment where the first service processing terminal is located and the current running data of the first service processing terminal; The edge server sends the current environmental data, the current running data, and a first control instruction to the first service processing terminal to control the first service processing terminal to adjust the second model according to the current environmental data, the current running data, and a preset second model adjustment algorithm to obtain a third model; The edge server controls the first service processing terminal to process service data using the third model, and controls the first service processing terminal to perform corresponding operations according to the processing results.
2. The AI edge computing method according to claim 1, wherein The step that the edge server obtains multiple first model sets corresponding to the service types of the services it processes, and the first model sets are generated by the cloud server according to specific service categories and service contents, includes: The edge server sends a service type request to the cloud server, and the service type request includes: the unique identification information of the edge server; the hardware configuration information of the edge server, including the processor type, memory size, and storage capacity; the geographical location information of the area served by the edge server; the list of service types currently processed by the edge server, and each service type includes a specific service category identifier and service content description; The cloud server performs the following operations according to the service type request: confirm the authorization status of the edge server according to the unique identification information of the edge server; filter out the model types adapted to the hardware configuration information according to the hardware configuration information of the edge server; select the training data set of the corresponding area according to the geographical location information of the area served by the edge server; for each service type: select the corresponding basic model from the model library according to the service category identifier; determine the specific parameter configuration of the model according to the service content description; use the selected training data set to train and optimize the basic model; generate an optimized model adapted to the service type; The cloud server packages the generated multiple optimized models into a first model set and sends a model distribution request to the edge server. The model distribution request includes: metadata information of each optimized model, including model version, applicable business type, and resource requirements; model deployment configuration information; model update policy; The edge server receives the model distribution request and performs the following operations: verifying the integrity and security of the model set; creating a model running environment according to the model deployment configuration information; storing the first model set in a preset model library; setting up a regular inspection and update mechanism according to the model update policy.
3. The AI edge computing method according to claim 2, wherein The step of the edge server selecting a corresponding first business processing terminal for connection from the business processing terminals according to the task plans and data processing requirements of multiple business processing terminals includes: The edge server establishes a terminal management table, which includes: device identifiers of each business processing terminal, network connection status of each business processing terminal, task plan information of each business processing terminal, and data processing requirement information of each business processing terminal; wherein, the task plan information includes: task type identifier, task execution time period, task priority, and task plan status; the data processing requirement information of each business processing terminal includes: processing type, data scale, real-time requirement, and accuracy requirement; The edge server performs terminal classification according to the terminal management table, specifically including: grouping terminals by task type; grouping terminals by data processing type; sorting terminals by task priority within each group; marking terminals with multitasking capabilities; The edge server selects the first business processing terminal according to the following conditions, including: judging the task time overlap degree and selecting terminals with matching time resources; judging the matching degree between the data processing capabilities of the terminals and the requirements; evaluating the stability of the network connection status; calculating the comprehensive task priority score; wherein, the calculation formula of the comprehensive task priority score is: score = w1×time matching degree + w2×processing capacity matching degree + w3×network stability + w4×task priority, where w1, w2, w3, and w4 are preset weight coefficients; The edge server establishes a connection with the selected first business processing terminal, including: sending a connection request to the first business processing terminal; waiting for the terminal to respond and verifying the connection status; establishing a secure communication channel; synchronizing the task plan and data processing configuration; updating the connection status in the terminal management table; The edge server establishes a task execution monitoring mechanism, including: regularly checking the connection status; monitoring the task execution progress; recording data processing performance indicators; dynamically adjusting the terminal selection strategy according to the monitoring results.
4. The AI edge computing method according to claim 3, wherein, The step of the edge server generating a second model set according to the terminal attribute data, a preset first model adjustment algorithm, and the first model set includes: The edge server parses and classifies the terminal attribute data to determine the terminal hardware attributes, terminal software attributes, and terminal function attributes. Among them, the terminal hardware attributes include: processor architecture type, processor performance parameters, memory capacity and type, and whether there is a dedicated AI accelerator; the terminal software attributes include: operating system type and version, AI framework support, and device driver compatibility; the terminal function attributes include: supported sensor types, actuator types, and communication interface types. The edge server performs model analysis on each model in the first model set to obtain model structure information, model running requirements, and model application scenarios. Among them, the model structure information includes: number and type of network layers, parameter scale, and computational complexity; the model running requirements include: minimum computational resource requirements, memory occupancy requirements, and response time requirements; the model application scenarios include: applicable task types, input data requirements, and output format requirements. The edge server executes the first model adjustment algorithm, which specifically includes: adjusting the implementation method of the model calculation layer according to the terminal processor architecture; optimizing the parameter storage structure according to the terminal memory capacity; adjusting the operator implementation method according to the presence of an AI accelerator; selecting the corresponding quantization precision from the predefined set of quantization precision levels according to the terminal performance, performing weight quantization, and performing activation value quantization; calculating the parameter importance score, removing low-importance parameters, and retraining the pruned model structure. The edge server performs adaptability verification on the optimized model, including: simulating the terminal hardware environment and configuring the target software environment; testing the computational performance, memory occupancy, and response latency, and testing various input scenarios; verifying the output accuracy; and checking the exception handling mechanism. The edge server integrates the verified optimized models into a second model set, including: generating a model description file; packaging the model files; configuring the model loading parameters; and setting the version control information.
5. The AI edge computing method according to claim 4, wherein, The steps for the edge server to determine the business processing party according to the terminal attribute data, its own business processing capabilities, and resource occupancy rate include: The edge server establishes a resource evaluation index system, which includes computing resource indicators, network resource indicators, and storage resource indicators. Among them, the computing resource indicators include: CPU usage rate and available core count, GPU / NPU usage, memory usage rate and available capacity; the network resource indicators include: network bandwidth occupancy rate, network latency level, and connection stability; the storage resource indicators include: storage space usage rate, I / O performance indicators, and cache hit rate. The edge server constructs a processing capacity evaluation model, including: Evaluating its own processing capacity: calculating the current task queue length of the edge server; counting the average processing delay; predicting the short-term load trend; Evaluating the first business processing terminal: analyzing the computing power of the first business processing terminal; evaluating the resource margin of the first business processing terminal; calculating the task adaptability; Evaluating other processing terminals: obtaining the real-time status of other terminals; evaluating the collaborative processing ability; calculating the load distribution ratio; The edge server performs multi-dimensional decision-making calculations, including: Calculate the local processing cost: Cost_local = q1 × computing resource occupancy + q2 × storage resource occupancy; Calculate the terminal processing cost: Cost_terminal = q3 × terminal resource occupancy + q4 × network transmission overhead; Calculate the cooperative processing cost: Cost_cooperative = q5 × task decomposition overhead + q6 × coordination and communication overhead; Calculate the processing benefits, including: computing processing time benefit; computing resource utilization benefit; computing system stability benefit; Generate a decision score, including: Score = benefit metric - cost metric; adjust using constraint conditions; calculate the final decision score; The edge server establishes a dynamic load balancing mechanism, including: Monitor the load status in real time, specifically: monitor the resource changes of each processing party; record the task processing performance; analyze the load distribution; Execute load adjustment, specifically: set a load threshold; calculate the load deviation; trigger load migration; Update the processing strategy, specifically: record the processing effect; optimize the decision parameters; adjust the weight coefficients; The edge server determines the final service processing party according to the following rules, including: When Score_local is the highest and meets the threshold requirement, select the edge server as the processing party; the edge server allocates and locks a specific number of processor resources, memory resources, and storage resources from the available computing resource pool according to the resource requirements of service processing, forms an independent resource space, and initializes the processing environment in the resource space; When Score_terminal is the highest and meets the threshold requirement, select the first service processing terminal as the processing party; configure task parameters for the first service processing terminal and establish a monitoring mechanism; When Score_cooperative is the highest and meets the threshold requirement, select the cooperative processing mode, divide the task boundary, and establish a cooperation mechanism.
6. The AI edge computing method according to claim 5, characterized in that, The step that when determining that the service processing party is the first service processing terminal, the edge server determines the corresponding second model from the second model set according to the service request data and sends the second model to the first service processing terminal, includes: The edge server parses the service request data, including: Parse the service type information, specifically: identify the main service type; identify the sub-task type; determine the task priority; Parse the data feature information, specifically: analyze the data scale; identify the data format; determine the data quality requirements; Parse the performance requirement information, specifically: obtain the real-time requirement; obtain the accuracy requirement; obtain the resource limit conditions; The edge server filters candidate models from the second model set, including: Execute feature matching, specifically: match the service type label; match the data format requirements; match the performance metric requirements; Evaluate the model applicability, specifically: calculate the feature matching degree; evaluate the performance satisfaction; calculate the resource adaptation degree; Generate a candidate model list, specifically: sort by the matching degree; filter out models that do not meet the constraints; record the model scores; The edge server performs model combination optimization, including: Analyze task dependencies, specifically: construct a task dependency graph; identify the critical path; determine the execution order; Calculate resource constraints, specifically: count the available resource quantity; calculate the resource demand; evaluate the resource matching degree; Generate an optimal combination plan, specifically: calculate the combination benefit; evaluate the combination feasibility; select the optimal plan; The edge server performs model packaging operations, including: Prepare model files, specifically: organize model parameters; package the model structure; generate configuration files; Create a deployment description, specifically: record dependencies; set the deployment sequence; configure running parameters; Execute integrity verification, specifically: verify file integrity; check configuration correctness; generate verification information; The edge server performs model distribution, including: Establish a secure transmission channel, specifically: negotiate the encryption method; establish a data channel; verify connection security; Execute batch transmission, specifically: set the transmission priority; control the transmission rate; monitor the transmission status; Verify the distribution result, specifically: check the receiving integrity; verify the model availability; confirm the successful deployment.
7. The AI edge computing method according to claim 6, characterized in that, The step that the edge server sends the current environment data, the current operation data, and the first control instruction to the first service processing terminal to control the first service processing terminal to adjust the second model according to the current environment data, the current operation data, and a preset second model adjustment algorithm to obtain a third model includes: The edge server collects and preprocesses environment data, including: Collect physical environment data such as temperature, humidity, light, vibration, noise, spatial position, and movement state; Collect service environment data such as the distribution of surrounding devices, network environment status, and electromagnetic interference situation; Perform data preprocessing on the physical environment data and the service environment data, specifically: data standardization processing; outlier detection and processing; data time series alignment; The edge server organizes the current operation data, and the current operation data includes resource usage status, network transmission status, and hardware operation parameters; among them, the resource usage status includes: CPU, GPU usage rate, memory occupancy, and storage space usage; the network transmission status includes: real-time bandwidth utilization rate, network delay data, and packet loss rate; the hardware operation parameters include: processor temperature, power consumption data, and working status of each component; The edge server generates a first control instruction, and the first control instruction includes a model adjustment strategy, a resource allocation strategy, and an execution control strategy; among them, the model adjustment strategy includes: determining the adjustment target, setting the adjustment range, and defining the adjustment step size; the resource allocation strategy includes: allocating computing resources, setting memory limits, and controlling the energy consumption level; the execution control strategy includes: setting the execution priority, defining the timeout mechanism, and configuring exception handling; The edge server integrates and sends data packets, including: Data packaging, specifically: merge environment data; integrate operation data; include control instructions; Data compression, specifically: select a compression algorithm; execute data compression; generate a check code; Batch sending, specifically: determine the sending order; control the sending rate; monitor the sending status; The edge computing module of the first service processing terminal performs model adjustment, including: Parsing the received data, specifically: decompressing the data packet; verifying the data integrity; classifying and storing the data; Evaluating the necessity of adjustment, specifically: analyzing the degree of environmental change; evaluating the degree of performance impact; calculating the adjustment benefit; Performing model adjustment, specifically: applying the second model adjustment algorithm; performing parameter fine-tuning; optimizing the model structure; The edge computing module verifies the adjusted third model, including: Performing performance testing, specifically: testing the computing performance; testing the response time; testing the resource occupancy; Verifying the functional correctness, specifically: testing the model accuracy; verifying the output stability; checking for abnormal responses; Generating a verification report, specifically: recording the adjustment effect; statistically analyzing the performance metrics; providing optimization suggestions.
8. An AI edge computing device for performing the AI edge computing method according to any one of claims 1 to 7, characterized in that, Including: A data acquisition module and a control processing module; wherein, The data acquisition module is configured to: acquire a plurality of first model sets corresponding to the service type of the service processed by itself, and the first model sets are generated by a cloud server according to specific service types and service contents; The control processing module is configured to: select a corresponding first service processing terminal for connection from the service processing terminals according to the task plans and data processing requirements of a plurality of service processing terminals; The data acquisition module is further configured to: acquire the terminal attribute data and service request data of the first service processing terminal; The control processing module is further configured to: Generate a second model set according to the terminal attribute data, a preset first model adjustment algorithm, and the first model sets; Determine the service processing party according to the terminal attribute data, its own service processing capabilities, and resource occupancy rate; When it is determined that the service processing party is the first service processing terminal, determine a corresponding second model from the second model sets according to the service request data, and send the second model to the first service processing terminal; The data acquisition module is further configured to: acquire the current environmental data of the environment where the first service processing terminal is located and the current operating data of the first service processing terminal; The control processing module is further configured to: Send the current environmental data, the current operating data, and a first control instruction to the first service processing terminal, so as to control the first service processing terminal to adjust the second model according to the current environmental data, the current operating data, and a preset second model adjustment algorithm to obtain a third model; Control the first service processing terminal to process service data using the third model, and control the first service processing terminal to perform corresponding operations according to the processing results.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the AI edge computing method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the AI edge computing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Adaptive edge computing resource management method and system based on deep reinforcement learning
CN116340003A
Drug information storage method based on distributed edge calculation and multi-modal data
CN118152481A