Distributed reasoning method and device based on hybrid expert architecture, equipment and medium
By adopting a hybrid expert architecture in the edge computing environment, the expert inference model is dispersed among multiple intelligent vehicles, and the computing power gateway is used to coordinate tasks, solving the problem of idle computing power resources of intelligent vehicles, and achieving efficient inference task processing and resource utilization.
Patent Information
- Application Number
- CN202511086609.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-08-05
AI Technical Summary
In the edge computing scenario, how to make full use of scattered and heterogeneous computing resources to achieve an efficient inference system, especially how to effectively utilize idle intelligent vehicle computing resources to solve the problem of insufficient computing resources for a single device.
A distributed inference method based on a hybrid expert architecture is adopted to deploy the expert inference model to multiple intelligent vehicles, coordinate task allocation and result aggregation through edge computing power gateways, use idle intelligent vehicle computing power resources to perform inference tasks, and dynamically adjust resource allocation through computing power scoring and migration mechanisms.
It realizes efficient utilization of the computing power of idle smart vehicles, reduces the computing and storage burden of a single device, improves the response speed, reduces the data transmission distance and bandwidth requirements, and reduces overall energy consumption.
Smart Images

Figure CN120579646A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a distributed reasoning method, apparatus, device and medium based on a hybrid expert architecture. Background Art
[0002] With the development of artificial intelligence (AI) technology, the number of parameters in large-scale language models and other AI models is increasing, making it difficult for a single device to efficiently complete inference tasks. Hybrid Expert Systems (HESs) are an effective solution. By splitting large models into multiple expert models, the router determines which expert models to activate for inference, significantly reducing computing resource consumption and improving inference efficiency.
[0003] In traditional hybrid expert architectures, expert models are typically deployed across multiple servers within the same data center, relying on high-speed network connectivity. However, in edge computing scenarios, leveraging distributed, heterogeneous computing resources to implement efficient reasoning systems remains a challenge.
[0004] In the wave of electric and intelligent vehicles, smart electric vehicles are generally equipped with high-performance artificial intelligence computing units for autonomous driving, environmental perception, planning, and decision-making, with computing power reaching hundreds of teraflops (Tera FLOPS). When these vehicles are parked and charging, their computing resources are idle, resulting in a waste of computing power. Summary of the Invention
[0005] The present invention provides a distributed reasoning method, apparatus, device and medium based on a hybrid expert architecture, which is used to solve the defect of computing power resource limitation of intelligent vehicles in the prior art and fully utilize the computing power resources of idle intelligent vehicles.
[0006] The present invention provides a distributed reasoning method based on a hybrid expert architecture, which is applied to an edge computing gateway. The method includes: If a reasoning task is received, determining a plurality of target intelligent vehicles participating in the reasoning task; Obtaining an expert reasoning model and deploying the expert reasoning model to each of the target intelligent vehicles; wherein the expert reasoning model includes a plurality of expert reasoning modules, and at least one of the expert reasoning modules is deployed in each of the target intelligent vehicles; Sending the reasoning task to each of the target intelligent vehicles, so that each of the target intelligent vehicles performs the reasoning task based on the expert reasoning module to obtain a reasoning result; Aggregate the reasoning results of each of the target intelligent vehicles to obtain a target reasoning result.
[0007] According to a distributed reasoning method based on a hybrid expert architecture provided by the present invention, deploying the expert reasoning model to each of the target intelligent vehicles includes: Obtaining first computing power information of each of the target intelligent vehicles; Calculating a computing power score of each target intelligent vehicle according to the first computing power information; Determining the expert reasoning module corresponding to each of the target intelligent vehicles based on the computing power score; The expert reasoning module corresponding to each target intelligent vehicle is deployed in the computing power sandbox of each target intelligent vehicle.
[0008] According to a distributed reasoning method based on a hybrid expert architecture provided by the present invention, the step of determining multiple target intelligent vehicles participating in the reasoning task includes: Obtain the second computing power information and estimated charging time of multiple currently connected smart vehicles; Calculating an inference score for each of the smart vehicles based on the second computing power information and the estimated charging time; A plurality of target intelligent vehicles participating in the reasoning task are determined according to the reasoning scores.
[0009] According to a distributed reasoning method based on a hybrid expert architecture provided by the present invention, after sending the reasoning task to each of the target intelligent vehicles, the method further includes: Determining a migration intelligent vehicle for each of the target intelligent vehicles; If it is detected that the target intelligent vehicle is offline, the expert reasoning module and reasoning data in the offline target intelligent vehicle are migrated to the migration intelligent vehicle, so that the migration intelligent vehicle performs the reasoning task based on the expert reasoning module to obtain the reasoning result.
[0010] The present invention also provides a distributed reasoning method based on a hybrid expert architecture, which is applied to a computing power management center. The method includes: If an inference task issued by a task requester is received, the location information of the task requester is obtained; Determine the edge computing gateway corresponding to the inference task based on the location information; The reasoning task is sent to the edge computing gateway corresponding to the reasoning task, so that the edge computing gateway executes the reasoning task.
[0011] According to a distributed reasoning method based on a hybrid expert architecture provided by the present invention, determining the edge computing gateway corresponding to the reasoning task based on the location information includes: Determine a candidate edge computing gateway based on the location information; Obtain the computing power resources of the candidate edge computing gateway, and calculate the computing power resource score of the candidate edge computing gateway based on the computing power resources; Determine the edge computing power gateway corresponding to the inference task based on the computing power resource score.
[0012] According to a distributed reasoning method based on a hybrid expert architecture provided by the present invention, after sending the reasoning task to the edge computing gateway corresponding to the reasoning task so that the edge computing gateway performs the reasoning task, the method further includes: Record the computing power contribution of each target intelligent vehicle in executing the reasoning task; The incentive information of each target intelligent vehicle is calculated based on the computing power contribution.
[0013] The present invention also provides a distributed reasoning method based on a hybrid expert architecture, which is applied to intelligent vehicles. The method includes: If the expert reasoning module and reasoning task sent by the edge computing power gateway are received, the expert reasoning module is deployed in the computing power sandbox; The reasoning task is performed based on the expert reasoning module to obtain the reasoning result, and the reasoning result is sent to the edge computing power gateway.
[0014] The present invention also provides a distributed reasoning device based on a hybrid expert architecture, which is applied to an edge computing gateway, including: A first determining module is configured to, upon receiving a reasoning task, determine a plurality of target intelligent vehicles participating in the reasoning task; A first deployment module is configured to obtain an expert reasoning model and deploy the expert reasoning model to each of the target intelligent vehicles; wherein the expert reasoning model includes a plurality of expert reasoning modules, and at least one of the expert reasoning modules is deployed in each of the target intelligent vehicles; A first sending module is configured to send the reasoning task to each of the target intelligent vehicles, so that each of the target intelligent vehicles performs the reasoning task based on the expert reasoning module to obtain a reasoning result; The aggregation module is configured to aggregate the reasoning results of each of the target intelligent vehicles to obtain a target reasoning result.
[0015] The present invention also provides a distributed reasoning device based on a hybrid expert architecture, which is applied to a computing power management center and includes: an acquisition module configured to acquire location information of a task requester upon receiving an inference task issued by the task requester; A second determination module is configured to determine the edge computing gateway corresponding to the inference task based on the location information; The second sending module is configured to send the inference task to the edge computing gateway corresponding to the inference task, so that the edge computing gateway executes the inference task.
[0016] The present invention also provides a distributed reasoning device based on a hybrid expert architecture, which is applied to an intelligent vehicle and includes: A second deployment module is configured to deploy the expert reasoning module in the computing power sandbox upon receiving the expert reasoning module and reasoning task sent by the edge computing power gateway; An execution module is configured to execute the reasoning task based on the expert reasoning module, obtain the reasoning result, and send the reasoning result to the edge computing power gateway.
[0017] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the distributed reasoning method based on the hybrid expert architecture as described above is implemented.
[0018] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the distributed reasoning method based on the hybrid expert architecture as described above is implemented.
[0019] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-described distributed reasoning methods based on the hybrid expert architecture.
[0020] The distributed reasoning method, device, equipment and medium based on the hybrid expert architecture provided by the present invention, after receiving the reasoning task, determine the multiple target intelligent vehicles participating in the reasoning task, obtain the multiple expert reasoning modules in the expert reasoning model and deploy them to each target intelligent vehicle, send the reasoning task to each target intelligent vehicle, and each target intelligent vehicle performs the reasoning task based on the expert reasoning module to obtain the reasoning result, aggregate the reasoning results of each target intelligent vehicle, and obtain the target reasoning result. The edge computing power gateway divides the reasoning calculation into three stages: "routing → expert parallel → aggregation", and proposes an expert hybrid distributed reasoning architecture suitable for edge environments, which fully utilizes the computing power resources of idle intelligent vehicles to achieve efficient reasoning of large-scale AI models. The expert reasoning model is stored in a decentralized manner through the expert hybrid architecture, and a single intelligent vehicle only needs to store part of the parameters, which greatly reduces the computing and storage burden of a single device; compared with centralized reasoning, edge distributed reasoning reduces the data transmission distance and bandwidth requirements, improves the response speed, and reduces overall energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 This is a flow chart of the distributed reasoning method based on hybrid expert architecture applied to the edge computing gateway provided by the present invention.
[0023] Figure 2 This is a scenario diagram of the distributed reasoning method based on the hybrid expert architecture provided by the present invention.
[0024] Figure 3 It is a flow chart of the distributed reasoning method based on hybrid expert architecture applied to a computing power management center provided by the present invention.
[0025] Figure 4 It is a flow chart of a distributed reasoning method based on a hybrid expert architecture applied to intelligent vehicles provided by the present invention.
[0026] Figure 5 It is a structural diagram of a distributed reasoning device based on a hybrid expert architecture applied to an edge computing gateway provided by the present invention.
[0027] Figure 6 It is a structural diagram of a distributed reasoning device based on a hybrid expert architecture provided by the present invention and applied to a computing power management center.
[0028] Figure 7 It is a structural diagram of a distributed reasoning device based on a hybrid expert architecture and applied to intelligent vehicles provided by the present invention.
[0029] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0030] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0031] Figure 1 FIG. 1 is a flow chart showing a distributed reasoning method based on a hybrid expert architecture according to an exemplary embodiment. Figure 1As shown, in an exemplary embodiment, a distributed reasoning method based on a hybrid expert architecture is applied to an edge computing gateway. The method includes steps 110 to 140, which are described in detail as follows.
[0032] Step 110: If a reasoning task is received, determine a plurality of target intelligent vehicles that participate in the reasoning task.
[0033] In the embodiment of the present invention, the edge computing gateway is the central coordinator of the distributed reasoning system, which stores the full model parameters of the expert reasoning model. Figure 2 As shown, the smart vehicle is connected to the edge computing gateway through charging piles 1-4. If it receives the inference task issued by the inference service requester (task requester), it determines multiple target smart vehicles participating in the inference task, such as Figure 2 For intelligent vehicles 1-4 in the figure, two public experts (A and B) and two domain experts (X and Y) are required to complete the reasoning task. The full model parameters of the expert reasoning model are divided into four expert reasoning modules (model weights W1-4).
[0034] Step 120: Acquire an expert reasoning model and deploy the expert reasoning model in each of the target intelligent vehicles; wherein the expert reasoning model includes a plurality of expert reasoning modules, and at least one of the expert reasoning modules is deployed in each of the target intelligent vehicles.
[0035] In an embodiment of the present invention, an expert reasoning model is obtained, and the expert reasoning model is divided into multiple expert reasoning modules, and the multiple expert reasoning modules are respectively deployed in the computing power sandbox in each target intelligent vehicle.
[0036] Step 130: Send the reasoning task to each of the target intelligent vehicles, so that each of the target intelligent vehicles performs the reasoning task based on the expert reasoning module to obtain a reasoning result.
[0037] In this embodiment of the present invention, an inference task is sent to each target intelligent vehicle. Each target intelligent vehicle then executes the inference task based on the deployed expert inference module and obtains an inference result. Each target intelligent vehicle executes the inference task in a pipelined manner. That is, after obtaining the inference result, the first target intelligent vehicle feeds the inference result back to the edge computing gateway. The edge computing gateway then sends the inference result to the second target intelligent vehicle, which then executes the inference task based on the inference result and obtains the corresponding inference result. This continues until the last target intelligent vehicle completes the inference task.
[0038] Step 140: Aggregate the reasoning results of each target intelligent vehicle to obtain a target reasoning result.
[0039] In the embodiment of the present invention, the reasoning results of each target intelligent vehicle are aggregated, and the final target reasoning result is obtained through a fusion algorithm.
[0040] In an embodiment of the present invention, the edge computing gateway divides an inference into three stages: "routing → expert parallelization → aggregation", and broadcasts a scheduling frame through PLC (Power Line Communication, power line communication technology) to specify the input offset, output buffer address and time slice of each target intelligent vehicle.
[0041] In this embodiment, a hybrid-expert distributed inference architecture suitable for edge environments is proposed, fully utilizing the computing power of idle intelligent vehicles to achieve efficient inference of large-scale AI models. This hybrid-expert architecture distributes model parameters, requiring only a subset of these parameters for each intelligent vehicle, significantly reducing the computational and storage burden on a single device. Compared to centralized inference, edge distributed inference reduces data transmission distance and bandwidth requirements, improving response speed and reducing overall energy consumption.
[0042] In an exemplary embodiment of the present invention, deploying the expert reasoning model to each of the target intelligent vehicles includes: Obtaining first computing power information of each of the target intelligent vehicles; Calculating a computing power score of each target intelligent vehicle according to the first computing power information; Determining the expert reasoning module corresponding to each of the target intelligent vehicles based on the computing power score; The expert reasoning module corresponding to each target intelligent vehicle is deployed in the computing power sandbox of each target intelligent vehicle.
[0043] In an embodiment of the present invention, the edge computing power gateway stores all model parameters of the expert reasoning model and maintains a real-time "vehicle-computing power" table. The "vehicle-computing power" table records in real time the first computing power information of the smart vehicle currently registered in the edge computing power gateway, such as the smart vehicle NPU (Neural Network Processing Unit, neural network processor) / GPU (Graphics Processing Unit, graphics processor) peak FLOPS (Floating Point Operations Per Second, floating-point operations per second), remaining storage, communication delay, estimated remaining charging time, etc., and calculates the computing power score of each target smart vehicle based on the first computing power information.
[0044] The edge computing power gateway quantifies the first computing power information into a computing power score S compute , the specific calculation formula is as follows: S compute= w1·log10(FLOPS peak ) + w2·(MEM free / MEM total ) + w3·(1 / RTT avg ) +w4·(1 / T charge ) Among them, FLOPS peak Indicates the peak floating-point computing power of the target intelligent vehicle's NPU / GPU (unit: Tera-FLOPS); MEM free / MEM total Indicates the ratio of the target intelligent vehicle's remaining video memory to the total video memory; RTT avg Indicates the average round-trip delay of the last N communications between the target smart vehicle and the edge computing gateway (unit: ms), such as the average round-trip delay of the last 10 communications; T charge Indicates the estimated remaining charging time (unit: h) when the current SOC (State Of Charge) of the target smart vehicle is ≥80%; w1~w4 represent preset weights, for example, w1=0.4, w2=0.3, w3=0.2, w4=0.1, satisfying Σw i = 1, i=1-4.
[0045] The model parameters of the expert reasoning model are divided into multiple expert reasoning modules. Specifically, the expert reasoning model can be divided according to the number of target intelligent vehicles. For example, if there are eight target intelligent vehicles, a 70-byte expert reasoning model can be divided into eight expert reasoning modules, with the number of expert reasoning module parameters being 8.75 bytes. Partitioning can also be done based on a fixed parameter granularity. For example, a 70-byte expert reasoning model can be sliced into 70 slices at a 128-MB granularity. The expert reasoning modules corresponding to each target intelligent vehicle are allocated based on the computing power score of each target intelligent vehicle. For example, a heuristic allocation algorithm (such as greedy or K-means clustering) can be used to allocate the expert reasoning modules in descending order of computing power score. For example, multiple adjacent expert reasoning modules can be deployed in target intelligent vehicles with high computing power to reduce communication, while target intelligent vehicles with low computing power can be deployed with a single expert reasoning module.
[0046] Taking a 24-layer Transformer language model as an example, each four layers can be used as an expert reasoning module, forming six independent expert reasoning modules. Each expert reasoning module contains specific self-attention sublayer and feedforward sublayer weight parameters. During deployment, the edge computing gateway distributes these expert reasoning modules to target intelligent vehicles that meet the storage and computing power requirements, ensuring a reasonable burden on each vehicle and reserving some expert processing modules for redundancy.
[0047] In this embodiment of the present invention, target intelligent vehicles are re-evaluated at preset intervals, and the expert reasoning modules deployed in the target intelligent vehicles are dynamically adjusted based on the re-evaluated computing power scores. If the computing power score of a target intelligent vehicle falls below a computing power threshold, migration is triggered.
[0048] In an exemplary embodiment of the present invention, determining a plurality of target intelligent vehicles participating in the reasoning task includes: Obtain the second computing power information and estimated charging time of multiple currently connected smart vehicles; Calculating an inference score for each of the smart vehicles based on the second computing power information and the estimated charging time; A plurality of target intelligent vehicles participating in the reasoning task are determined according to the reasoning scores.
[0049] In this embodiment of the present invention, edge computing nodes function as expert inference routers, determining which expert inference modules should handle inference tasks. For example, they control which expert inference module a particular inference task is routed to based on parameters such as load balancing coefficient, latency penalty, and exploration rate.
[0050] On the edge computing gateway, the inference task (token) is output by the router to the activation probability of each expert, and then the target intelligent vehicle is determined. That is, the edge computing gateway combines the real-time load of the connected intelligent vehicle, the PLC link communication delay, the NPU / GPU peak FLOPS of the intelligent vehicle, the remaining storage and other second computing power information for weighted summation to calculate the inference score of each intelligent vehicle. Based on the inference score, the top-K intelligent vehicles are retained as the target intelligent vehicles. At the same time, the total delay of the target intelligent vehicle should be less than the delay threshold.
[0051] The edge computing power gateway quantifies the second computing power information into an inference score S infer , used to screen target intelligent vehicles participating in reasoning tasks: S infer = w5·log10(FLOPS peak ) + w6·(MEM free / MEM total ) + w7·(1 / RTT avg ) +w8·(T remain / T max ) Among them, T remain Indicates the estimated remaining charging time (unit: h); T max Indicates the maximum allowed time to complete this inference task (unit: h); w5~w8 represent preset weights, such as w5=0.35, w6=0.25, w7=0.2, w8=0.2, satisfying Σwi =1, i=5~8.
[0052] The edge computing gateway calculates the reasoning score S for all currently online intelligent vehicles in each decision cycle. infer , sort the reasoning scores in descending order and select the top-K intelligent vehicles as the target intelligent vehicles, ensuring the total delay Σ(1 / RTT avg )< delay threshold L max .
[0053] In an exemplary embodiment of the present invention, after sending the reasoning task to each of the target intelligent vehicles, the method further includes: Determining a migration intelligent vehicle for each of the target intelligent vehicles; If it is detected that the target intelligent vehicle is offline, the expert reasoning module and reasoning data in the offline target intelligent vehicle are migrated to the migration intelligent vehicle, so that the migration intelligent vehicle performs the reasoning task based on the expert reasoning module to obtain the reasoning result.
[0054] To address dynamic vehicle changes, this embodiment of the present invention implements redundant backup and dynamic migration mechanisms for the expert reasoning module. Specifically, each target intelligent vehicle has at least one migration intelligent vehicle, whose computing power is similar to that of the corresponding target intelligent vehicle. The expert reasoning module is redundantly stored in the migration intelligent vehicle. Before the target intelligent vehicle leaves the network, model parameters are proactively pushed to the migration intelligent vehicle. If the target intelligent vehicle suddenly loses its connection, the migration intelligent vehicle performs the reasoning task.
[0055] In the embodiment of the present invention, a dynamic expert allocation mechanism is implemented through redundant backup and dynamic migration mechanisms, which can adapt to the scenarios of vehicle access and departure and ensure system stability.
[0056] In one exemplary embodiment of the present invention, when executing inference tasks, inference orchestration can be performed. For example, a 512-token inference task can be split into four micro-batches, with eight expert inference modules running in parallel in each batch. This task is then transmitted to the target intelligent vehicle via a PLC communication slot of 0-3ms. The target intelligent vehicle performs inference calculations in 3-6ms, and the inference results are transmitted back in 6-7ms. The edge computing gateway initiates All-Reduce aggregation at 7ms and outputs the target inference results at 7.5ms. Inference tasks are executed in a pipelined manner, overlapping communication and computation. For example, within the aforementioned inference execution time, the second micro-batch can be transmitted starting at 3s.
[0057] Figure 3 FIG. 1 is a flow chart showing a distributed reasoning method based on a hybrid expert architecture according to an exemplary embodiment. Figure 3As shown, in an exemplary embodiment, a distributed reasoning method based on a hybrid expert architecture is applied to a computing power management center. The method includes steps 310 to 330, which are described in detail as follows.
[0058] Step 310: If an inference task issued by a task requester is received, the location information of the task requester is obtained.
[0059] In an embodiment of the present invention, the computing power management center is responsible for managing multiple edge computing power gateways, realizing global computing power resource scheduling, providing an inference computing power interface to inference service providers, and selecting appropriate edge computing power gateways based on business needs and vehicle distribution.
[0060] After receiving the reasoning task issued by the task requester, the location information of the task requester is determined.
[0061] Step 320: Determine the edge computing gateway corresponding to the inference task based on the location information.
[0062] In an embodiment of the present invention, an edge computing gateway suitable for the inference task is determined based on location information.
[0063] Step 330: Send the inference task to the edge computing gateway corresponding to the inference task, so that the edge computing gateway executes the inference task.
[0064] In an embodiment of the present invention, the inference task is sent to the edge computing gateway corresponding to the inference task, and the edge computing gateway executes the inference task. The specific scheme of the edge computing gateway executing the inference task has been explained and will not be repeated here.
[0065] In an exemplary embodiment of the present invention, determining the edge computing gateway corresponding to the inference task based on the location information includes: Determine a candidate edge computing gateway based on the location information; Obtain the computing power resources of the candidate edge computing gateway, and calculate the computing power resource score of the candidate edge computing gateway based on the computing power resources; Determine the edge computing power gateway corresponding to the inference task based on the computing power resource score.
[0066] In this embodiment of the present invention, after receiving an inference request, the computing power management center first screens candidate edge computing gateways based on location information and user geofences (e.g., distance less than 1 km). It then multiplies the number of available smart vehicles in each candidate edge computing gateway by the average FLOPS / average RTT (Round-Trip Time) of each smart vehicle to obtain a computing power resource score. The candidate edge computing gateway with the highest computing power resource score is selected to perform the inference task. If at least two candidate edge computing gateways have the same highest computing power resource score, the edge computing gateway with a high historical inference success rate and / or low electricity price is prioritized and returned to the inference service requester.
[0067] In an exemplary embodiment of the present invention, after sending the inference task to the edge computing gateway corresponding to the inference task so that the edge computing gateway performs the inference task, the method further includes: Record the computing power contribution of each target intelligent vehicle in executing the reasoning task; The incentive information of each target intelligent vehicle is calculated based on the computing power contribution.
[0068] In an embodiment of the present invention, the contribution of each target intelligent vehicle in participating in reasoning is recorded, and then the incentive information of the target intelligent vehicle is calculated, and the incentive mechanism is implemented through computing power points or charging discounts.
[0069] The embodiments of the present invention provide a contribution-based incentive mechanism, creating a win-win situation for computing power providers, computing power users, and system operators.
[0070] Figure 4 FIG. 1 is a flow chart showing a distributed reasoning method based on a hybrid expert architecture according to an exemplary embodiment. Figure 4 As shown, in an exemplary embodiment, a distributed reasoning method based on a hybrid expert architecture is applied to an intelligent vehicle. The method includes steps 410 to 420, which are described in detail as follows.
[0071] Step 410: If the expert reasoning module and reasoning task sent by the edge computing power gateway are received, the expert reasoning module is deployed in the computing power sandbox.
[0072] In this embodiment of the present invention, intelligent vehicles are equipped with high-performance NPU / GPU computing units, providing a highly isolated and secure computing sandbox. Charging stations are connected to the intelligent vehicles via charging cables, and a high-speed data transmission channel is provided via PLC, enabling high-speed network connectivity between intelligent vehicles and with edge computing gateways. This invention utilizes power line communication technology to achieve high-speed interconnection between inference nodes, eliminating the need for additional network infrastructure.
[0073] When a smart vehicle connects to a charging station, it establishes a connection with the edge computing gateway via the PLC, registers available computing resources, and initializes the computing sandbox environment. When a smart vehicle is identified as the target vehicle for an inference task, it receives the expert inference module and inference task from the edge computing gateway and deploys the expert inference module in the computing sandbox. Each smart vehicle stores only a subset of the model's parameters, acting as one or more expert nodes in the hybrid expert architecture.
[0074] Step 420: Execute the reasoning task based on the expert reasoning module, obtain the reasoning result, and send the reasoning result to the edge computing gateway.
[0075] In an embodiment of the present invention, the reasoning task is performed based on the deployed expert reasoning module to obtain the reasoning result, and the reasoning result is sent to the edge computing power gateway, which is aggregated by the edge computing power node to obtain the final target reasoning result.
[0076] In one exemplary embodiment of the present invention, an inference service provider decides whether to use edge inference to provide inference tasks requested by users. Specifically, the inference service provider maintains an edge / cloud cost model. When the number of available smart vehicles on the edge computing gateway is greater than or equal to a vehicle threshold and the estimated RTT is less than a latency threshold, the edge inference system is selected for execution. If there are insufficient available smart vehicles or the RTT is too high, the inference falls back to the cloud. Furthermore, the inference service provider allows users to specify either "low latency first" or "low cost first" modes for executing inference tasks.
[0077] In the embodiment of the present invention, the distributed reasoning method based on the hybrid expert architecture includes five stages: initialization and computing power registration stage, expert reasoning model allocation stage, reasoning service request and routing stage, expert selection and reasoning execution stage, and contribution recording and resource release stage. The relevant descriptions of each stage are as follows: During the initialization and computing power registration phase, the domain expert and public expert's smart vehicles connect to charging piles via charging ports to obtain power. The charging piles establish high-speed communication channels using PLC technology. The smart vehicles report information such as NPU / GPU computing power type, content inventory, and computing power sandbox status. They also register computing power resources with the edge computing power gateway via PLC. Simultaneously, the domain expert's smart vehicle also reports its specialized computing power characteristics. The edge computing power gateway reports its available distributed computing power resources to the computing power management center, which then confirms resource registration with the edge computing gateway.
[0078] During the expert reasoning model allocation phase, the computing power management center sends the completed expert reasoning model parameters to the edge computing power gateway, including the expert partitioning strategy, routing algorithm parameters, and inter-expert communication protocol. The edge computing power gateway evaluates the computing power characteristics and stability of the smart vehicle, performing a tiered assessment based on the estimated charging time and computing power information to determine the smart vehicle's reasoning score. Based on the expert partitioning strategy, the edge computing power gateway distributes the expert model parameters to the smart vehicle and initializes the smart vehicle's computing power sandbox environment. The edge computing power gateway then reports the completion of the expert reasoning module deployment to the computing power management center.
[0079] During the inference service request and routing phase, the inference service requester initiates an inference request to the inference service provider, uploading the inference task type and input data. The inference service provider then requests edge inference resources from the computing power management center, conveying the task requirements and service quality requirements. The computing power management center evaluates the global computing power distribution based on network latency, computing power load, and other factors, determines the edge computing power node that will execute the inference task, and returns the edge computing power gateway credentials to the inference service provider. The inference service provider then establishes a connection with the edge computing power gateway based on the edge computing power gateway credentials and forwards the inference request.
[0080] During the expert selection and inference execution phase, the edge computing gateway executes an expert routing algorithm. This algorithm analyzes input features to determine which expert inference models to activate and how they collaborate. When a public expert is required to handle an inference task, the input data (inference task) is distributed to the public expert, who then runs the inference task in a computing sandbox. When a domain expert is required to handle an inference task, the input data (inference task) is distributed to the domain expert, who then performs optimized calculations for the specific domain. The edge computing gateway integrates the results (inference results) of each expert using a fusion algorithm. The final inference result (target inference result) is then transmitted sequentially between the edge computing gateway, the computing management center, the inference service provider, and the inference service requester.
[0081] During the contribution recording and resource release phase, the computing power contribution of each intelligent vehicle is determined based on the computational effort and duration of the inference task, as well as the quality assessment of the inference results. The computing power management center generates and updates the contribution incentive mechanisms, such as charging coupons or computing power points. The edge computing power gateway monitors the vehicle connection status and triggers expert parameter migration and reallocation when a vehicle leaves.
[0082] In the embodiments of the present invention, common reasoning tasks include text-to-text, text-to-video, and video-to-text. The following describes the specific distributed reasoning process: The scenario is as follows: a commercial complex's underground parking lot has 80 charging stations. Approximately 60 smart vehicles are charging simultaneously between 10:00 PM and 6:00 AM. These smart vehicles are equipped with NPUs and computing sandboxes, enabling them to perform inference tasks. The computing management center receives an inference task from the mall's customer service system: providing customers with multi-round conversational question-and-answer services, including store recommendations and real-time coupon generation.
[0083] Decompose the reasoning task: This reasoning task uses a 7-billion-parameter Transformer MoE conversational model, pre-divided into four common experts (A: intent recognition; B: store knowledge; C: promotion strategies; D: language generation) and two domain experts (X: women's clothing; Y: digital products). A complete conversation proceeds through four phases: intent → knowledge → strategy → generation, with two experts activated concurrently in each phase.
[0084] Reasoning process: At t0 (0 ms), the computing power management center selects the edge computing power gateway GW-01 that is closest and has the highest computing power resource score based on the GPS location of the customer's mobile phone APP.
[0085] At t1 (1 ms), the edge computing gateway GW-01 selects 6 smart vehicles as target smart vehicles (V1~V6) from 60 smart vehicles by calculating the inference score formula. infer The values are 9.2, 8.9, 8.7, 8.5, 8.3, and 8.0 respectively.
[0086] During the t2 (2 ms) phase, the edge computing gateway GW-01 distributes four public expert modules (each with 1.75 B parameters) and two domain expert modules (each with 1.75 B parameters) according to the computing power scores: V1 deploys A+B; V2 deploys C+D; V3 deploys X; V4 deploys Y; V5 and V6 serve as redundant backups.
[0087] At t3 (3 ms), the customer enters: "I want to buy a birthday present for xx, and my budget is xxx yuan." At stage t4 (4 ms), GW-01 broadcasts the token sequence of the sentence to V1 (intent recognition expert A).
[0088] At t5 (7 ms), V1 returns the intent vector I, and the edge computing gateway GW-01 sends I + the original token to V3 (X, an expert in the women's clothing field).
[0089] At t6 (10 ms), V3 returns the store candidate list L1, and the edge computing gateway GW-01 sends I+L1 to V2 (discount strategy expert C).
[0090] At t7 (13 ms), V2 generates the coupon vector C1, and the edge computing gateway GW-01 sends I+L1+C1 to V2's language generation expert D.
[0091] At t8 (16 ms), V2's language generation expert D outputs the final response: "Recommend her a new scarf from ×× Women's Clothing Store. Original price xxx yuan, yyy yuan after coupon, limited to 2 hours." At t9 (17 ms), the edge computing gateway GW-01 performs weighted aggregation on the intermediate vectors returned by the six experts and transmits the results back to the customer's mobile app.
[0092] At stage t10 (18 ms), the customer continues the conversation and the system repeats t4-t9, forming multiple rounds of interaction.
[0093] After the inference task is completed, the edge computing gateway GW-01 calculates the contribution based on the number of tokens actually calculated by each smart vehicle multiplied by the weight coefficient: Contribution i = Token i × (S infer_i / ΣS infer ); The contribution will be converted into charging coupons and distributed to the car owner's APP by the computing power management center the next day.
[0094] The following describes the distributed reasoning device based on a hybrid expert architecture provided by the present invention. The distributed reasoning device based on the hybrid expert architecture described below can be used in conjunction with the distributed reasoning method based on the hybrid expert architecture described above. It should be noted that the device provided in the following embodiments and the method provided in the above embodiments share the same concept. The specific manner in which each module and unit performs operations has been described in detail in the method embodiments and will not be repeated here.
[0095] In an exemplary embodiment of the present invention, see Figure 5 , Figure 5 A distributed reasoning device based on a hybrid expert architecture is shown according to an exemplary embodiment and is applied to an edge computing gateway, including the following modules.
[0096] A first determination module 510 is configured to, upon receiving a reasoning task, determine a plurality of target intelligent vehicles participating in the reasoning task; A first deployment module 520 is configured to obtain an expert reasoning model and deploy the expert reasoning model to each of the target intelligent vehicles; wherein the expert reasoning model includes a plurality of expert reasoning modules, and at least one expert reasoning module is deployed in each of the target intelligent vehicles; A first sending module 530 is configured to send the reasoning task to each of the target intelligent vehicles, so that each of the target intelligent vehicles performs the reasoning task based on the expert reasoning module to obtain a reasoning result; The aggregation module 540 is configured to aggregate the reasoning results of each target intelligent vehicle to obtain a target reasoning result.
[0097] In an exemplary embodiment of the present invention, the first deployment module 520 includes: A first acquisition submodule is configured to acquire first computing power information of each target intelligent vehicle; A first computing submodule is configured to calculate a computing power score of each target intelligent vehicle based on the first computing power information; A first determination submodule is configured to determine the expert reasoning module corresponding to each target intelligent vehicle based on the computing power score; The deployment submodule is configured to deploy the expert reasoning module corresponding to each target intelligent vehicle in the computing power sandbox of each target intelligent vehicle.
[0098] In an exemplary embodiment of the present invention, the first determining module 510 includes: A second acquisition submodule is configured to obtain second computing power information and estimated charging time of multiple currently connected smart vehicles; a second computing submodule, configured to calculate an inference score of each of the smart vehicles based on the second computing power information and the estimated charging time; The second determination submodule is configured to determine a plurality of target intelligent vehicles participating in the reasoning task according to the reasoning scores.
[0099] In an exemplary embodiment of the present invention, the distributed reasoning device based on the hybrid expert architecture further includes: A third determining module is configured to determine a migrating intelligent vehicle for each of the target intelligent vehicles; The migration module is configured to migrate the expert reasoning module and reasoning data in the offline target intelligent vehicle to the migration intelligent vehicle if it is detected that the target intelligent vehicle is offline, so that the migration intelligent vehicle performs the reasoning task based on the expert reasoning module and obtains the reasoning result.
[0100] In an exemplary embodiment of the present invention, see Figure 6 , Figure 6A distributed reasoning device based on a hybrid expert architecture is shown according to an exemplary embodiment and is applied to a computing power management center, including: An acquisition module 610 is configured to acquire location information of a task requester upon receiving a reasoning task issued by the task requester; A second determination module 620 is configured to determine the edge computing gateway corresponding to the inference task based on the location information; The second sending module 630 is configured to send the inference task to the edge computing gateway corresponding to the inference task, so that the edge computing gateway executes the inference task.
[0101] In an exemplary embodiment of the present invention, the second determining module 620 includes: A third determination submodule is configured to determine a candidate edge computing gateway based on the location information; A third calculation submodule is configured to obtain the computing power resources of the candidate edge computing gateway and calculate the computing power resource score of the candidate edge computing gateway based on the computing power resources; The fourth determination submodule is configured to determine the edge computing power gateway corresponding to the inference task based on the computing power resource score.
[0102] In an exemplary embodiment of the present invention, the distributed reasoning device based on the hybrid expert architecture further includes: a recording module configured to record the computing power contribution of each target intelligent vehicle in executing the reasoning task; A calculation module is configured to calculate the incentive information of each target intelligent vehicle based on the computing power contribution.
[0103] In an exemplary embodiment of the present invention, see Figure 7 , Figure 7 A distributed reasoning device based on a hybrid expert architecture is shown according to an exemplary embodiment and is applied to an intelligent vehicle, including: The second deployment module 710 is configured to deploy the expert reasoning module in the computing power sandbox upon receiving the expert reasoning module and reasoning task sent by the edge computing power gateway; The execution module 720 is configured to execute the reasoning task based on the expert reasoning module, obtain the reasoning result, and send the reasoning result to the edge computing power gateway.
[0104] Figure 8 An example of a physical structure diagram of an electronic device is shown below. Figure 8As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call logic instructions in the memory 830 to execute a distributed reasoning method based on a hybrid expert architecture.
[0105] Furthermore, the logic instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0106] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the distributed reasoning method based on the hybrid expert architecture provided by the above methods.
[0107] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the distributed reasoning method based on the hybrid expert architecture provided by the above methods.
[0108] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0109] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A distributed reasoning method based on a hybrid expert architecture, characterized in that: Applied to an edge computing gateway, the method includes: If a reasoning task is received, determining a plurality of target intelligent vehicles participating in the reasoning task; Obtaining an expert reasoning model and deploying the expert reasoning model to each of the target intelligent vehicles; wherein the expert reasoning model includes a plurality of expert reasoning modules, and at least one of the expert reasoning modules is deployed in each of the target intelligent vehicles; Sending the reasoning task to each of the target intelligent vehicles, so that each of the target intelligent vehicles performs the reasoning task based on the expert reasoning module to obtain a reasoning result; Aggregate the reasoning results of each of the target intelligent vehicles to obtain a target reasoning result.
2. The distributed reasoning method based on hybrid expert architecture according to claim 1, characterized in that: The deploying the expert reasoning model to each of the target intelligent vehicles includes: Obtaining first computing power information of each of the target intelligent vehicles; Calculating a computing power score of each target intelligent vehicle according to the first computing power information; Determining the expert reasoning module corresponding to each of the target intelligent vehicles based on the computing power score; The expert reasoning module corresponding to each target intelligent vehicle is deployed in the computing power sandbox of each target intelligent vehicle.
3. The distributed reasoning method based on hybrid expert architecture according to claim 1, characterized in that: The determining of multiple target intelligent vehicles participating in the reasoning task includes: Obtain the second computing power information and estimated charging time of multiple currently connected smart vehicles; Calculating an inference score for each of the smart vehicles based on the second computing power information and the estimated charging time; A plurality of target intelligent vehicles participating in the reasoning task are determined according to the reasoning scores.
4. The distributed reasoning method based on hybrid expert architecture according to any one of claims 1 to 3, characterized in that: After sending the reasoning task to each of the target intelligent vehicles, the method further includes: Determining a migration intelligent vehicle for each of the target intelligent vehicles; If it is detected that the target intelligent vehicle is offline, the expert reasoning module and reasoning data in the offline target intelligent vehicle are migrated to the migration intelligent vehicle, so that the migration intelligent vehicle performs the reasoning task based on the expert reasoning module to obtain the reasoning result.
5. A distributed reasoning method based on a hybrid expert architecture, characterized in that: Applied to a computing power management center, the method includes: If an inference task issued by a task requester is received, the location information of the task requester is obtained; Determine the edge computing gateway corresponding to the inference task based on the location information; The reasoning task is sent to the edge computing gateway corresponding to the reasoning task, so that the edge computing gateway executes the reasoning task.
6. The distributed reasoning method based on hybrid expert architecture according to claim 5, characterized in that: The determining the edge computing gateway corresponding to the inference task based on the location information includes: Determine a candidate edge computing gateway based on the location information; Obtain the computing power resources of the candidate edge computing gateway, and calculate the computing power resource score of the candidate edge computing gateway based on the computing power resources; Determine the edge computing power gateway corresponding to the inference task based on the computing power resource score.
7. The distributed reasoning method based on hybrid expert architecture according to claim 5 or 6, characterized in that: After sending the inference task to the edge computing gateway corresponding to the inference task so that the edge computing gateway performs the inference task, the method further includes: Record the computing power contribution of each target intelligent vehicle in executing the reasoning task; The incentive information of each target intelligent vehicle is calculated based on the computing power contribution.
8. A distributed reasoning method based on a hybrid expert architecture, characterized in that: Applied to a smart vehicle, the method includes: If the expert reasoning module and reasoning task sent by the edge computing power gateway are received, the expert reasoning module is deployed in the computing power sandbox; The reasoning task is performed based on the expert reasoning module to obtain the reasoning result, and the reasoning result is sent to the edge computing power gateway.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the distributed reasoning method based on the hybrid expert architecture as claimed in any one of claims 1 to 8 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the distributed reasoning method based on the hybrid expert architecture as claimed in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Vehicle-mounted distributed computing power system and method and vehicle
CN114460923A
Vehicle-mounted computing power sharing system, method and equipment based on edge computing and medium
CN115080210A
Model reasoning method and device based on calculation unit deployment, equipment and medium
CN117494816A
Vehicle computing resource allocation system based on distributed unloading algorithm
CN118449977A
Calculation power sharing transaction method, system, device, medium and program product
CN118502930A
Cited By
Large model batch reasoning and data flow optimization system oriented to MOE architecture
CN120849141A
Data processing method based on hybrid expert model and related equipment
CN121072776A
Data communication method and device, electronic equipment and storage medium
CN121509528A
Model processing method, electronic equipment, storage medium and program product
CN122285303A