Service processing method and device, network equipment, storage medium and program product

By decomposing model service requests into subtasks in the 6G network, determining QoS policies and resource requirements, and selecting multiple subnets to perform together, the problem of cross-subnet resource collaboration is solved, task execution efficiency and resource utilization are improved, and the 6G network needs for intelligence and efficiency are met.

CN120343023APending Publication Date: 2025-07-18CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510527046.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In 6G distributed networks, the lack of efficient AI task decomposition and scheduling mechanisms makes it difficult to achieve cross-subnet resource coordination, complex AI tasks are slow to execute, idle resources and excessive occupation coexist, and resource utilization is not high, so the potential advantages of 6G network cannot be fully utilized.

Method used

Through the central SOF, the model service request is decomposed into multiple subtasks, the QoS policy and demand resources are determined, multiple second subnets are selected to perform inference tasks in collaboration, and the task execution results are integrated to build distributed resource scheduling and task processing mechanisms, and the orchestration and collaboration mechanism of network functional units are optimized.

Benefits of technology

It improves the overall performance and resource utilization efficiency of the system, and enhances the execution capabilities and system adaptability of the subnet in the case of insufficient data, computing power and computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343023A_ABST
    Figure CN120343023A_ABST
Patent Text Reader

Abstract

The invention provides a service processing method and device, network equipment, a storage medium and a program product, and relates to the technical field of wireless communication. The collaborative reasoning method comprises the following steps: in response to a model service request sent by a first sub-SOF, decomposing the model service request into a plurality of first sub-tasks based on the business logic of the model service request; determining a quality of service (QoS) strategy and a demand resource matched with the first subtask; selecting a plurality of second sub-networks based on the QoS strategy and the demand resources, so that the plurality of second sub-networks cooperatively execute the reasoning task based on the corresponding plurality of first sub-tasks; and receiving task execution results fed back by the plurality of second sub-networks based on the reasoning task, integrating the task execution results to obtain a service result, and feeding back the service result to the service demand side through the first sub-SOF. According to the technical scheme, cross-subnet AI task collaboration is achieved through unified scheduling and task decomposition of the central network, and the task execution problem when the subnets are insufficient in capacity is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the development of AI (Artificial Intelligence) technology, new challenges have been posed to the service capabilities of 6G network systems. The 6G network not only needs to meet basic communication requirements but also needs to deeply support complex AI tasks such as model training and AI inference. However, in the actual scenario of 6G distributed networks, due to the lack of an efficient AI task decomposition and scheduling mechanism, it is difficult to achieve cross-subnet resource collaborative allocation. This makes the execution of complex AI tasks slow, with coexistence of resource idleness and over-occupation, resulting in low task execution efficiency and low resource utilization rate, and unable to fully utilize the potential advantages of the 6G network.

[0003] It should be noted that the information disclosed in the above Background Art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a collaborative inference method, a configuration device, a network device, a storage medium, and a computer program product, which can at least overcome to a certain extent the problems of low AI task execution efficiency and low resource utilization rate in related technologies.

[0005] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be learned in part through the practice of the present disclosure.

[0006] According to one aspect of the present disclosure, a collaborative inference method is provided, which is applied to a central SOF. The central SOF is a service orchestration function network element of a central network, and includes: in response to a model service request sent by a first sub-SOF, decomposing the model service request into a plurality of first sub-tasks based on the business logic of the model service request, where the model service request is generated when the first sub-SOF estimates that local resources do not meet or support the service request sent by the service requester, the first sub-SOF is a service orchestration function network element of a first subnet managed by the central network, and the plurality of first sub-tasks include model training and model inference; determining a quality of service (QoS) policy matching the first sub-task, and determining the required resources for processing the first sub-task; selecting a plurality of second subnets for processing the plurality of first sub-tasks based on the QoS policy and the required resources, so that the plurality of second subnets collaboratively execute an inference task based on the corresponding plurality of first sub-tasks; receiving the task execution results feedback by the plurality of second subnets based on the inference task, and integrating the task execution results to obtain a service result and then feedbacking the service result to the first sub-SOF, so that the first sub-SOF feedbacks the service result to the service requester.

[0007] In one embodiment of the present disclosure, determining a quality of service (QoS) policy matching the first subtask includes: configuring different requirement parameters in the service requirement information carried by the first subtask into corresponding initial QoS metrics based on a configured mapping rule library; adjusting the initial QoS metrics based on the feedback network state and initial resource capabilities of the candidate subnet to obtain target QoS metrics; and generating the QoS policy based on the target QoS metrics and the collaboration requirements of the first subtask.

[0008] In one embodiment of the present disclosure, determining the required resources for processing the first subtask includes: estimating the resource requirements of the first subtask based on the target QoS metrics, and determining the required resources based on the estimation result and the transmission resources required between the first subtasks with data dependencies.

[0009] In one embodiment of the present disclosure, selecting multiple second subnets for processing the multiple first subtasks based on the QoS policy and the required resources includes: obtaining the first processing capabilities reported by the candidate subnets; and determining the second subnet for performing the model training and the second subnet for performing the model inference based on the first processing capabilities, the QoS policy, and the required resources.

[0010] In one embodiment of the present disclosure, selecting multiple second subnets for processing the multiple first subtasks based on the QoS policy and the required resources further includes: sending the task information of the multiple first subtasks to the second subnets, where the task information includes at least one of a task identifier, the required resources of the first subtask, the QoS policy, the data address of the data required for processing the model service request, and the receiving address of the task execution result.

[0011] In one embodiment of the present disclosure, before receiving the task execution results feedback by the multiple second subnets based on the inference task, it further includes: receiving first task execution information reported by the multiple second subnets during the execution of the first subtask, where the first task execution information includes task progress information and computing power resource usage information. Among them, within each of the second subnets, the second sub-SOF selects a matching data service management function network element (DSMF), and the DSMF selects multiple data plane function network elements (DPFs) to collaboratively execute the first subtask. Then, the multiple DPFs report their respective task-related data to the second sub-SOF through the DSMF, and the second sub-SOF integrates the task-related data to obtain the first task execution information; if it is determined to adjust the inference task based on the first task execution information, at least one of the task volume, task type, task priority, and inter-subnet collaboration relationship in the configuration information of the inference task is adjusted to obtain configuration change information; the configuration change information is sent to the multiple second subnets, so that the multiple second subnets continue to execute the inference task based on the configuration change information to obtain the task execution results.

[0012] In one embodiment of the present disclosure, it further includes: storing the service result in a first DSF, where the first DSF is a data storage function network element of the central network.

[0013] According to another aspect of the present disclosure, a collaborative inference method is provided, which is applied to a DSMF. The DSMF is a data service management function network element of a second subnet managed by a central network, and includes: in response to the task information of a second subtask sent by the second sub-SOF of the second subnet, selecting a data plane function network element (DPF) that matches the task information within the second subnet to deploy the second subtask to the multiple DPFs. The second subtask is obtained by the second sub-SOF decomposing the received first subtask, and the first subtask is obtained by the central SOF decomposing the model service request of the first subnet. The central SOF is a service orchestration function network element of the central network; receiving the sub-task execution results feedback by the primary DPF among the multiple DPFs, and feeding back the sub-task execution results to the second sub-SOF, so that the second sub-SOF integrates the sub-task execution results through the central SOF to obtain a service result and feeds it back to the first sub-SOF of the first subnet.

[0014] In one embodiment of the present disclosure, the first subtask includes a model training task. Selecting a data plane function network element DPF that matches the task information within the second subnet includes: selecting a first DPF and a second DPF that match the task information within the second subnet. The first DPF is used to perform data processing operations on model training data to obtain processed data, and the second DPF is used to perform model training operations based on the processed data.

[0015] In one embodiment of the present disclosure, selecting a first DPF and a second DPF that match the task information within the second subnet includes: screening out first alternative DPFs with model training capabilities in a preset information table, and selecting the first DPF and the second DPF based on the second processing capabilities reported by the first alternative DPFs.

[0016] In one embodiment of the present disclosure, deploying the second subtask to multiple DPFs includes: deploying the task of data processing to the first DPF, and the deployment information of the data processing includes at least one of data processing method information, processed data output address information, and first computing power configuration information; deploying the task of model training to the second DPF, and the deployment information of the model training includes at least one of QoS requirements, collaborative training node information, second computing power configuration information, and training model storage address information.

[0017] In one embodiment of the present disclosure, the first DPF obtains the model training data from a second DSF within the second intranet to which it belongs.

[0018] In one embodiment of the present disclosure, the first subtask includes a model inference task. Receiving the execution results of the subtask fed back by the primary DPF among multiple DPFs includes: receiving the inference results of the model inference task fed back by the primary DPF as the execution results of the subtask.

[0019] In one embodiment of the present disclosure, selecting a data plane function network element DPF that matches the task information within the second subnet includes: screening out second alternative DPFs with inference capabilities in a preset information table, and selecting the matching DPF based on the second processing capabilities reported by the second alternative DPFs.

[0020] In one embodiment of the present disclosure, deploying the second subtask to multiple DPFs includes: deploying the model inference task to the DPF, and the deployment information of the model inference includes at least one of model acquisition address information, model cutting method information, and result output address information.

[0021] In one embodiment of the present disclosure, before receiving the sub-task execution results fed back by the primary DPF among the multiple DPFs, it further includes: receiving second task execution information reported by the multiple DPFs during the execution of the second sub-task; if it is determined to adjust the first sub-task based on the second task execution information, adjusting the allocation parameters of the first sub-task; and sending the adjusted allocation parameters to the multiple DPFs so that the multiple DPFs continue to execute the second sub-task based on the adjusted allocation parameters to obtain the sub-task execution results.

[0022] In one embodiment of the present disclosure, if the first sub-task is a model training task, the allocation parameters include the computing power allocation parameters of the model training task and the corresponding algorithm parameters; if the first sub-task is a model inference task, the allocation parameters include the allocation information of the model slicing of the inference model used in the model inference task.

[0023] According to still another aspect of the present disclosure, there is provided a collaborative inference device applied to a central SOF, where the central SOF is a service orchestration function network element of a central network, including: a decomposition module, configured to respond to a model service request sent by a first sub-SOF, and decompose the model service request into multiple first sub-tasks based on the service logic of the model service request, where the model service request is generated when the first sub-SOF anticipates that local resources do not meet or support the service request sent by the service requester, the first sub-SOF is a service orchestration function network element of a first subnet managed by the central network, and the multiple first sub-tasks include model training and model inference; a determination module, configured to determine a quality of service (QoS) policy matching the first sub-task and determine the required resources for processing the first sub-task; a first selection module, configured to select multiple second subnets for processing the multiple first sub-tasks based on the QoS policy and the required resources, so that the multiple second subnets collaboratively execute an inference task based on the corresponding multiple first sub-tasks; and a first receiving module, configured to receive the task execution results fed back by the multiple second subnets based on the inference task, and integrate the task execution results to obtain a service result and then feedback the service result to the first sub-SOF, so that the first sub-SOF feeds back the service result to the service requester.

[0024] According to another aspect of the present disclosure, there is provided a collaborative inference device, which is applied to a DSMF. The DSMF is a data service management function network element of a second subnet managed by a central network, and includes: a second selection module, configured to, in response to task information of a second sub-task sent by a second sub-SOF of the second subnet, select a data plane function network element DPF that matches the task information within the second subnet, so as to deploy the second sub-task to multiple DPFs. The second sub-task is obtained by decomposing a first sub-task received by the second sub-SOF, and the first sub-task is obtained by decomposing a model service request of a first subnet by a central SOF. The central SOF is a service orchestration function network element of the central network; a second receiving module, configured to receive sub-task execution results fed back by a primary DPF among multiple DPFs, and feed the sub-task execution results back to the second sub-SOF, so that the second sub-SOF integrates the sub-task execution results through the central SOF to obtain a service result and feeds it back to a first sub-SOF of the first subnet.

[0025] According to another aspect of the present disclosure, there is provided a network device, including: a processor; and a memory for storing executable instructions of the processor; the processor is configured to execute the collaborative inference method in the first aspect above by executing the executable instructions.

[0026] According to another aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the collaborative inference method described above is implemented.

[0027] According to another aspect of the present disclosure, there is provided a computer program product, on which a computer program is stored. When the computer program is executed by a processor, the collaborative inference method described above is implemented.

[0028] The collaborative inference solution provided by the embodiments of the present disclosure generates a model service request by a first sub-SOF when local resources are insufficient, decomposes it into first sub-tasks based on business logic, determines QoS policies and required resources for each sub-task, selects multiple second subnets to collaboratively execute inference tasks accordingly, and finally integrates task execution results, constructing a set of distributed resource scheduling and task processing mechanisms. Thus, by optimizing the orchestration and collaboration mechanisms of network function units, the overall performance and resource utilization efficiency of the system are improved, and by introducing a cross-network AI service collaboration mechanism, subnets can make full use of resources of other networks to complete complex AI service tasks when data, computing power, and computing resources are insufficient, enhancing the overall service ability and adaptability of the system.

[0029] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0031] Figure 1 A schematic diagram showing a 6G network system supporting collaborative data services in an embodiment of the present disclosure;

[0032] Figure 2 A schematic diagram showing another 6G network system supporting collaborative data services in an embodiment of the present disclosure;

[0033] Figure 3 A flowchart showing a collaborative inference method in an embodiment of the present disclosure;

[0034] Figure 4 A flowchart showing another collaborative inference method in an embodiment of the present disclosure;

[0035] Figure 5 A flowchart showing yet another collaborative inference method in an embodiment of the present disclosure;

[0036] Figure 6 A flowchart showing yet another collaborative inference method in an embodiment of the present disclosure;

[0037] Figure 7 A flowchart showing yet another collaborative inference method in an embodiment of the present disclosure;

[0038] Figure 8 A flowchart showing yet another collaborative inference method in an embodiment of the present disclosure;

[0039] Figure 9 A flowchart showing yet another collaborative inference method in an embodiment of the present disclosure;

[0040] Figure 10 A schematic diagram showing a collaborative inference device in an embodiment of the present disclosure;

[0041] Figure 11 A schematic diagram showing another collaborative inference device in an embodiment of the present disclosure;

[0042] Figure 12 A block diagram showing the structure of a computer device in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.

[0044] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0045] With the rapid development of communication technologies, 6G network systems are gradually becoming a research hotspot. 6G networks need to be more efficient, flexible, and capable of meeting diverse data requirements. Existing research on 6G network technologies has begun to explore how to improve the flexibility and intelligence level of the network through a service-based architecture and distributed networking. With the wide application of AI technologies, 6G network systems need to have strong AI service capabilities to support complex tasks such as model training and AI inference. In a distributed network environment, subnets often lack sufficient data or computing power to complete these tasks, and existing technologies are difficult to efficiently decompose and schedule complex AI tasks, unable to achieve cross-subnet resource coordination, resulting in low task execution efficiency and low resource utilization. Therefore, an effective cross-network AI service coordination mechanism is needed to solve this problem, so as to efficiently perform cross-network coordination, resource scheduling, and the execution of AI services.

[0046] As Figure 1 and Figure 2 shown, an architecture diagram of a 6G network system supporting data services in the form of a service-based interface and an architecture diagram of a 6G network system supporting data services in the form of a point-to-point interface are respectively shown.

[0047] Among them, the network function modules include:

[0048] UE (User Equipment): User equipment, which is a terminal device for users to interact with the network, such as mobile phones, computers, etc., and is connected to the network through the N1 interface.

[0049] (R)AN ((Radio)Access Network): (Wireless) access network, which is responsible for wireless communication with the UE, provides wireless access functions, and is connected to the core network through the N2 interface.

[0050] UPF (User Plane Function): The user plane function is responsible for processing data packet forwarding, QoS (Quality of Service) handling, etc. in the user plane, and communicates with other network elements through the N3, N4, N9, and N6 interfaces.

[0051] DN (Data Network): The data network is the network through which the UE accesses external data resources, such as the Internet, etc., and is connected to the UPF through the N6 interface.

[0052] The control plane network functions (NFs) include:

[0053] eAMF (Evolved Access and Mobility Management Function): The evolved access and mobility management function is responsible for functions such as UE access control, mobility management, and session management, and interacts with the eNRF through the Nnrf interface.

[0054] eNRF (Evolved Network Repository Function): The evolved network repository function stores network-related data and configuration information, provides data query and storage services for other network elements, and communicates with other control plane NFs through the Nnrf interface.

[0055] The specific function modules include:

[0056] SOF (Service Orchestration Function): The service orchestration function receives data service requirements, orchestrates services, including computing power, data, and connection orchestration, etc., splits and translates service requirements into corresponding data service tasks, selects the DSMF network elements required for orchestration, and allocates tasks to appropriate areas for deployment, and interacts with other network elements through the Nnrf and Naaf interfaces.

[0057] DSMF (Data Service Management Function): The data service management function selects specific NFs with different data capabilities, such as DPF and DSF, according to the SOF orchestration result, forms an executable logical topology, controls the work chain to implement specific data service functions, and also supports the data service capability reporting function of DPF, DSF, etc., and interacts with other network elements through the Nnrf interface.

[0058] DPF (Data Plane Function): The data plane function is responsible for executing specific tasks, implementing functions such as data collection, transmission, preprocessing, and analysis, and interacts with other network elements through the data bus.

[0059] DSF (Data Storage Function): The data storage function is used to store the collected data, data service results, etc., and interacts with other network elements through the data bus.

[0060] SEF (Service Exposure Function): The service exposure function provides an open interface for 6G network services externally, enabling external applications to access and use network capabilities, and interacts with other network elements through the Nnef and Naaf interfaces.

[0061] AF (Application Function): The application function represents an external application, interacts with the network function through the Naaf interface, and submits service requests.

[0062] Such as Figure 1 The control bus and data bus shown: are respectively used to transmit data and signaling on the control plane and user plane. Interfaces such as N1 - N9: are standard interfaces between different network elements, used to implement communication and interaction between network elements, and each interface defines specific communication protocols and functions.

[0063] Such as Figure 2 As shown, different functional entities interact through specific connections to complete a series of processes from user service request access, service orchestration management, data processing and storage to service result provision, realizing various service functions of the 6G network. Among them, the marked "DC" represents characteristics or deployment methods related to the Data Center.

[0064] Next, each step of the collaborative reasoning method in this exemplary embodiment will be described in more detail in conjunction with the accompanying drawings and embodiments.

[0065] Figure 3 Shows a flowchart of a collaborative reasoning method in an embodiment of the present disclosure.

[0066] Such as Figure 3 As shown, according to an embodiment of the present disclosure, the collaborative reasoning method is applied to the central SOF. The central SOF is a service orchestration function network element of the central network, including:

[0067] Step S302, in response to a model service request sent by the first sub - SOF, decomposing the model service request into multiple first sub - tasks based on the business logic of the model service request. Among them, the model service request is generated when the first sub - SOF anticipates that local resources do not meet or support the service request sent by the service requester. The first sub - SOF is a service orchestration function network element of the first subnet managed by the central network, and the multiple first sub - tasks include model training and model inference.

[0068] In some embodiments, the model service request is an AI service request generated by the first sub-SOF when it determines that local resources cannot meet or support the service request proposed by the service requester.

[0069] In some embodiments, the business logic refers to the inherent rules and processes related to the model service request, which are constructed based on factors such as the goal of the service request, data processing requirements, and model operation mechanism.

[0070] In some embodiments, the first sub-task is a refined task unit obtained by disassembling the model service request according to the business logic, including but not limited to model training and model inference mentioned above.

[0071] Step S304, determine the quality of service (QoS) policy matching the first sub-task, and determine the required resources for processing the first sub-task.

[0072] In some embodiments, the QoS policy is a service quality guarantee policy formulated for each first sub-task, which can stipulate the standards to be achieved in terms of performance, reliability, timeliness, etc. during the task execution process.

[0073] In some embodiments, the required resources refer to various resources required to complete the first sub-task, including computing resources (such as the number of CPU cores, GPU models and quantities), storage resources (memory size, disk capacity), network resources (bandwidth, latency requirements), and software resources (specific algorithm libraries, model frameworks), etc.

[0074] Step S306, select multiple second subnets for processing multiple first sub-tasks based on the QoS policy and required resources, so that the multiple second subnets jointly execute the inference task based on the corresponding multiple first sub-tasks.

[0075] In some embodiments, jointly executing the inference task means that multiple second subnets jointly complete the inference task through mutual cooperation, data interaction, and task connection based on the allocated first sub-tasks. For example, in the target recognition inference task in the autonomous driving scenario, subnet A is responsible for collecting and preprocessing camera image data, subnet B uses the image feature extraction model for data processing, and subnet C uses the target classification model to complete the final target recognition.

[0076] Step S308, receive the task execution results feedback by multiple second subnets based on the inference task, and integrate the task execution results to obtain the service result and then feedback it to the first sub-SOF, so that the first sub-SOF can feedback the service result to the service requester.

[0077] In some embodiments, after the central network collects the task execution results feedback by multiple second subnets, it combines, verifies, and optimizes these results according to the logical relationship between the tasks and the expectations of the service requester to obtain the service result.

[0078] In this embodiment, when local resources are insufficient, a model service request is generated through the first sub-SOF, decomposed into first subtasks based on business logic, QoS policies and required resources are determined for each subtask, and multiple second subnets are selected to collaboratively execute the inference task. Finally, the task execution results are integrated to construct a distributed resource scheduling and task processing mechanism, thereby improving the overall performance and resource utilization efficiency of the system by optimizing the scheduling and collaboration mechanism of network function units, and enabling subnets to make full use of resources of other networks to complete complex AI service tasks when data, computing power, and computing resources are insufficient by introducing a cross-network AI service collaboration mechanism, enhancing the overall service ability and adaptability of the system.

[0079] In this embodiment, when a subnet lacks the data or computing power required for AI services, it sends an AI service request to the central network. The central network decomposes the AI service into tasks, selects appropriate subnets for task allocation based on task requirements and subnet resource information, and each subnet collaboratively executes the distributed AI task. During the task execution process, the progress and resource usage are reported in real time, and the central network dynamically adjusts the task configuration according to the reported information.

[0080] In an embodiment of the present disclosure, determining a quality of service QoS policy matching the first subtask includes: configuring different requirement parameters in the service requirement information carried by the first subtask as corresponding initial QoS metrics based on a configured mapping rule library; adjusting the initial QoS metrics based on the feedback network state and initial resource capabilities of candidate subnets to obtain target QoS metrics; and generating a QoS policy based on the target QoS metrics and the collaboration requirements of the first subtask.

[0081] In some embodiments, the service requirement information SLA (Service-Level Agreement), that is, the service level agreement, stipulates the specific content of the service, quality standards, rights and obligations of both parties, etc.

[0082] In some embodiments, the mapping rule library refers to a pre-configured database that stores the corresponding relationships between various service requirement parameters and QoS metrics.

[0083] In some embodiments, the initial QoS metrics refer to the QoS metrics obtained by mapping and converting different requirement parameters in the service requirement information according to the mapping rule library, which are preliminary set service quality standards. After considering the feedback network state and initial resource capabilities of candidate subnets, these initial metrics need to be adjusted to obtain the target QoS metrics.

[0084] In some embodiments, after obtaining the target QoS metrics, a complete QoS policy is formulated by combining the collaboration relationships and interaction requirements among the first subtasks. The collaboration requirements include the data transmission frequency between subtasks, data consistency requirements, task execution order constraints, etc. For example, for a natural language processing task involving multi-subnet collaboration, some subtasks are responsible for text tokenization, and some subtasks perform semantic analysis. When generating the QoS policy, it is necessary to not only clarify the QoS metrics of each subtask in terms of computing resources and processing latency, but also specify the format of data transmission between subnets, the latency tolerance, and the startup and connection conditions of each subtask to ensure the efficient and stable operation of the entire task process.

[0085] In this embodiment, the service demand information is converted into initial QoS metrics by constructing a mapping rule library, and the target QoS metrics are obtained by optimizing the metrics using the network status and initial resource capabilities fed back by the candidate subnets. Then, a QoS policy is generated by combining the collaboration requirements of the first subtasks, realizing the effective docking of service demands and network resources. The use of the mapping rule library ensures the standardization and rationality of QoS metric setting. Based on the dynamic adjustment of subnet status and resource capabilities, the QoS metrics can adapt to the actual network environment. Generating the policy by combining the collaboration requirements ensures that multiple subnets can cooperate efficiently when executing the first subtasks, meeting the multi-dimensional requirements of complex tasks for service quality.

[0086] In an embodiment of the present disclosure, determining the required resources for processing the first subtask includes: estimating the resource requirements of the first subtask based on the target QoS metrics, so as to determine the required resources based on the estimation results and the transmission resources required between the first subtasks with data dependencies.

[0087] In some embodiments, assuming that the target QoS metrics include the completion time of the task, the accuracy of model inference, the packet loss rate of data transmission, etc., the first subtasks include model training, model evaluation, model inference, etc., and the corresponding required resources include the number of CPUs, the number of GPUs, the memory size, and the size of the temporary storage area, etc. After the data preprocessing subtask is completed, it is necessary to transmit the processed data to the model training subtask, and the corresponding required resources also include the bandwidth required for data transmission.

[0088] In this embodiment, by estimating the resource requirements of the first subtask based on the target QoS metrics and combining the transmission resources required between the first subtasks with data dependencies to determine the required resources, the target QoS metrics provide a clear quantitative basis for resource estimation, enabling the resource allocation to be precisely adjusted according to the actual requirements of the task. At the same time, considering the transmission resources between data-dependent subtasks ensures the efficiency and reliability of data flow in the task process, and can effectively improve the execution efficiency and quality of model training and inference tasks.

[0089] In one embodiment of the present disclosure, multiple second subnets for processing multiple first subtasks are selected based on QoS policies and required resources, including: obtaining the first processing capacity reported by the candidate subnet; based on the first processing capacity, QoS policies and required resources, determining the second subnet for performing model training and the second subnet for performing model reasoning.

[0090] In this embodiment, a multi-dimensional evaluation model is established by obtaining the first processing capacity reported by the candidate subnet, combining the hard indicator requirements for service quality in the QoS policy (such as response time, data accuracy), and the quantitative requirements of the required resources for computing, storage, network and other resources. When determining the second subnet for performing model training and reasoning, on the one hand, for model training tasks, priority is given to screening subnets with abundant computing resources, large storage capacity and meeting the training time requirements to ensure efficient model training. On the other hand, for model reasoning tasks, emphasis is placed on selecting subnets with low network latency, fast response speed and guaranteed reasoning accuracy to achieve fast and accurate reasoning services.

[0091] In one embodiment of the present disclosure, multiple second subnets for processing multiple first subtasks are selected based on QoS policies and required resources, and also include: sending task information of the multiple first subtasks to the second subnet, the task information including at least one of a task identifier, required resources for the first subtask, QoS policy, a data address of data required for processing model service requests, and a receiving address for task execution results.

[0092] In some embodiments, the task identifier facilitates the second subnet to quickly identify the task attribution and category. The required resources and QoS strategy provide a clear basis for resource allocation and quality control for the subnet to execute the task, ensuring that it reasonably allocates computing, storage and other resources according to established standards, thereby ensuring the quality and efficiency of task execution. The data address guides the subnet to accurately obtain the data required to process the model service request, avoiding data search confusion and transmission delays, and improving data flow efficiency. The receiving address of the task execution result ensures that after the task is completed, the result can be accurately transmitted back to the specified location, thereby realizing closed-loop management of the task processing process.

[0093] In this embodiment, by sending the task information of multiple first subtasks to the second subnet, the task execution deviation and resource waste caused by information asymmetry between subnets are effectively reduced, the coordination and accuracy of task execution by each subnet are enhanced, and tasks such as model training and reasoning are ensured to be completed efficiently and stably in a distributed network environment.

[0094] like Figure 4 As shown, in one embodiment of the present disclosure, it also includes:

[0095] Step S402: Receive the first task execution information reported by multiple second subnets during the execution of the first subtask. The first task execution information includes task progress information and computing resource usage information. Among them, within each second subnet, the second sub-SOF selects a matching data service management function network element (DSMF), and the DSMF selects multiple data plane function network elements (DPFs) to collaboratively execute the first subtask. Then, the multiple DPFs report their task-related data to the second sub-SOF through the DSMF, and the second sub-SOF integrates the task-related data to obtain the first task execution information.

[0096] In some embodiments, the first task execution information refers to a set of key data fed back by multiple second subnets during the execution of the first subtask, which consists of task progress information and computing resource usage information.

[0097] Step S404: If it is determined to adjust the inference task based on the first task execution information, adjust at least one of the task volume, task type, task priority, and inter-subnet collaboration relationship in the configuration information of the inference task to obtain configuration change information.

[0098] In some embodiments, the configuration change information refers to a set of key instruction sets generated after dynamically adjusting the inference task based on the first task execution information, covering the change content of at least one configuration element such as task volume, task type, task priority, and inter-subnet collaboration relationship.

[0099] Step S406: Send the configuration change information to multiple second subnets so that the multiple second subnets continue to execute the inference task based on the configuration change information to obtain the task execution result.

[0100] In some embodiments, after receiving the configuration change information, the multiple second subnets will reconfigure and optimize the inference tasks they execute according to the adjustment instructions therein.

[0101] In some embodiments, if the change information requires adjusting the task volume, the subnet will reallocate computing resources and give priority to processing newly added or higher-priority task data; if the task type is changed, the subnet needs to load new models or algorithms and adjust the data processing flow; for changes in task priority, the subnet will suspend or defer low-priority tasks and give priority to ensuring the execution of high-priority tasks; when the inter-subnet collaboration relationship changes, the subnet will interact with other subnets and collaborate according to the new regulations, such as changing the data transmission protocol and adjusting the data sharing frequency, so as to continue to promote the inference task on the new configuration until the final task execution result is obtained.

[0102] Step S408: Receive the task execution results fed back by multiple second subnets based on the inference task.

[0103] In this embodiment, by receiving the first task execution information reported by the second subnet and generating configuration change information based on this information to be sent to the subnet, guiding the subnet to continue execution based on the change information, the real-time collection of the first task execution information enables the system to perceive the task execution status and resource usage. Based on the generated configuration change information, the dynamic adjustment of the inference task is realized, and problems such as resource bottlenecks and progress lags that occur during the task execution can be solved in a timely manner. The subnet flexibly adjusts the task execution strategy according to the configuration change information to ensure that the task is always executed in an optimal state.

[0104] In an embodiment of the present disclosure, it further includes: storing the service result in the first DSF, where the first DSF is a data storage functional network element of the central network.

[0105] Such as Figure 5 shown, the collaborative inference method according to another embodiment of the present disclosure is applied to the DSMF. The DSMF is a data service management functional network element of the second subnet managed by the central network, and includes:

[0106] Step S502, in response to the task information of the second subtask sent by the second sub-SOF of the second subnet, select a data plane functional network element DPF that matches the task information in the second subnet to deploy the second subtask to multiple DPFs. The second subtask is obtained by decomposing the first subtask received by the second sub-SOF. The first subtask is obtained by decomposing the model service request of the first subnet by the central SOF. The central SOF is a service orchestration functional network element of the central network.

[0107] In some embodiments, the second subtask is a further refined decomposition of the first subtask, and is obtained by the second sub-SOF after receiving the first subtask and splitting it according to the resource characteristics, task execution logic, and performance optimization requirements of the second subnet.

[0108] Step S504, receive the subtask execution result fed back by the primary DPF among multiple DPFs, and feed back the subtask execution result to the second sub-SOF, so that the second sub-SOF integrates the subtask execution result through the central SOF to obtain the service result and feed it back to the first sub-SOF of the first subnet.

[0109] In some embodiments, among the multiple DPFs responsible for executing the second subtask, the primary DPF refers to the DPF responsible for overall planning and coordination, and is used to summarize the execution results of the second subtask to obtain the subtask execution result.

[0110] In this embodiment, the first subtask is decomposed into the second subtask through the second sub-SOF, and the DPF is matched in the second subnet for task deployment. The main DPF then collects and feeds back the subtask execution results. Based on the decomposition mechanism of the second subtask, the resource characteristics of the second subnet are fully adapted to achieve refined task allocation. The setting of the main DPF can solve the data aggregation and communication coordination problems when multiple DPFs are executed collaboratively. The orderly feedback and layer-by-layer integration of subtask execution results ensure the continuity and accuracy from the bottom-level task execution to the upper-level service result generation, enhance the flexibility and reliability of task processing in a distributed network environment, optimize resource utilization efficiency, and ensure that complex model service requests can be completed efficiently and stably.

[0111] In this embodiment, the DSMF within the subnet selects multiple appropriate DPFs to collaboratively perform tasks based on the capabilities and information tables, establishes DPF data interaction paths, performs distributed model training, and dynamically adjusts task configurations during the training process. The DSMF deploys joint reasoning tasks to the DPFs, and the DPFs collaborate to perform joint reasoning and dynamically adjust model cutting points during the reasoning process to optimize reasoning efficiency.

[0112] In one embodiment of the present disclosure, the first subtask includes a model training task, and a data plane function network element DPF matching the task information is selected in the second subnet, including: selecting a first DPF and a second DPF matching the task information in the second subnet, the first DPF being used to perform data processing operations on the model training data to obtain processed data, and the second DPF being used to perform model training operations based on the processed data.

[0113] In some embodiments, the first DPF is responsible for data processing operations and performs a series of pre-processing tasks on the model training data, such as cleaning, denoising, normalizing and other operations on the original model training data, or performing word segmentation, part-of-speech tagging, word vector conversion and other processing on the original model training data. The second DPF is used to carry out model training operations based on the data processed by the first DPF, equipped with a deep learning framework and optimization algorithm, and using computing resources to learn and adjust parameters of the processed data to obtain a trained model.

[0114] In this embodiment, by selecting the first DPF and the second DPF adapted for data processing and model training respectively in the second subnet, and clarifying the division of labor and cooperation between the two, resource conflicts and performance bottlenecks that may be caused by a single DPF undertaking multiple tasks are avoided, efficient resource utilization and parallel acceleration of tasks are achieved, the execution efficiency of the model training task and the quality of the final training results are greatly improved, and the stability and reliability of model training task processing in a distributed network environment are enhanced.

[0115] In one embodiment of the present disclosure, selecting a first DPF and a second DPF that match the task information within a second subnet includes: screening out first alternative DPFs with model training capabilities in a preset information table, and selecting the first DPF and the second DPF based on the second processing capabilities reported by the first alternative DPFs.

[0116] In this embodiment, by screening out the first alternative DPFs with model training capabilities in the preset information table and selecting the first DPF and the second DPF according to the second processing capabilities reported by the first alternative DPFs, the screening mechanism of the preset information table can quickly locate the DPFs with model training capabilities from numerous DPFs, thereby improving the resource utilization rate of the entire second subnet.

[0117] In one embodiment of the present disclosure, deploying a second subtask to multiple DPFs includes: deploying the task of data processing to the first DPF, where the deployment information of data processing includes at least one of data processing method information, output address information of processed data, and first computing power configuration information; deploying the task of model training to the second DPF, where the deployment information of model training includes at least one of QoS requirements, collaborative training node information, second computing power configuration information, and training model storage address information.

[0118] In this embodiment, by deploying the data processing and model training tasks respectively according to different key information, for the deployment of the data processing task of the first DPF, the data processing method information is used to clarify the processing logic and algorithm, the output address information of processed data ensures a clear data flow path, and the first computing power configuration information reasonably allocates computing resources, enabling the data preprocessing work to be completed efficiently and accurately. For the deployment of the model training task of the second DPF, the QoS requirements ensure that the training process meets the service quality standards, the collaborative training node information realizes multi-node collaboration optimization, the second computing power configuration information accurately matches the computing resources, and the training model storage address information standardizes the saving path of the training results, achieving the adaptation of the data processing and model training tasks in terms of resource allocation, execution process, quality guarantee, etc., and improving the task execution efficiency and model training quality.

[0119] In one embodiment of the present disclosure, the first DPF obtains model training data from a second DSF within the second internal network to which it belongs.

[0120] In one embodiment of the present disclosure, the first subtask includes a model inference task, and receiving the execution results of the subtask fed back by the master DPF among multiple DPFs includes: receiving the inference result of the model inference task fed back by the master DPF as the execution result of the subtask.

[0121] In this embodiment, by receiving the inference result of the model inference task fed back by the master DPF as the sub-task execution result, the master DPF integrates the intermediate results generated by each DPF during the inference process, and through processing and verification, obtains the sub-task execution result. Instead of each of the multiple DPFs feeding back data to the upper layer separately, the data is aggregated to the master DPF for processing, which helps to improve the data transmission efficiency.

[0122] In an embodiment of the present disclosure, selecting a data plane function network element DPF that matches the task information within the second subnet includes: screening out second alternative DPFs with inference capabilities in a preset information table, and selecting a matching DPF based on the second processing capabilities reported by the second alternative DPFs.

[0123] In this embodiment, by screening out second alternative DPFs with inference capabilities in a preset information table and selecting a matching DPF according to the second processing capabilities reported by them, the preset information table serves as the basis for screening, enabling the rapid positioning of DPFs with inference capabilities from numerous DPFs, effectively narrowing the selection range. The second processing capabilities reported by the second alternative DPFs provide a quantitative basis for selection, comprehensively considering the performance of each DPF in terms of computing speed, memory capacity, processing accuracy, etc., to ensure that the model inference task can be assigned to a DPF with matching processing capabilities, thereby improving resource utilization.

[0124] In an embodiment of the present disclosure, deploying a second sub-task to multiple DPFs includes: deploying the model inference task to the DPFs, and the deployment information of the model inference includes at least one of model acquisition address information, model cutting method information, and result output address information.

[0125] In some embodiments, the model acquisition address information clarifies the source from which the DPF obtains the model required for inference, enabling the DPF to quickly and accurately obtain the corresponding model; the model cutting method information characterizes the reasonable segmentation and deployment of the model according to the processing capabilities and task characteristics of different DPFs; the result output address information provides a path for the feedback of the inference result, ensuring that the DPF can accurately send the result to the specified location after completing the inference, facilitating subsequent integration and application.

[0126] In this embodiment, by sending the deployment information of the model inference, the effective allocation and collaborative execution of the model inference task among multiple DPFs are realized, improving resource utilization and inference efficiency, ensuring the accurate feedback of the inference result, and being able to better meet the requirements for model inference in practical applications.

[0127] In one embodiment of the present disclosure, before receiving the sub-task execution results fed back by the primary DPF among multiple DPFs, it further includes: receiving the second task execution information reported by multiple DPFs during the execution of the second sub-task; if it is determined to adjust the first sub-task based on the second task execution information, adjusting the allocation parameters of the first sub-task; and sending the adjusted allocation parameters to multiple DPFs, so that multiple DPFs continue to execute the second sub-task based on the adjusted allocation parameters to obtain the sub-task execution results.

[0128] In some embodiments, the second task execution information refers to a set of dynamic data reported in real time by multiple DPFs during the execution of the second sub-task, including content such as task execution status and resource usage.

[0129] In some embodiments, the allocation parameters of the first sub-task refer to the key configuration parameters used to guide the DPF to execute the second sub-task, which vary according to the type of the first sub-task. For a model training task, the computing power allocation parameters specify the number of CPU cores and the number of GPUs that each DPF can use, etc., and the algorithm parameters specify the optimization algorithm, learning rate, batch size, etc. adopted for training; for a model inference task, the allocation information of model slicing clarifies the slicing method of the inference model among multiple DPFs, the part of the model responsible for inference by each DPF, and the data interaction rules between each part, etc. The reasonable setting of the allocation parameters directly affects the efficiency and quality of task execution.

[0130] In this embodiment, by collecting the second task execution information in real time, the system can timely grasp the actual situation of task execution. Based on this information, the allocation parameters of the first sub-task are adjusted, and reasonable decisions can be made according to the actual situation. Sending the adjusted parameters to the DPF and continuously adjusting according to the actual execution situation is beneficial to improving the overall quality of the task. For example, in model training, by adjusting the training parameters, the model can converge faster and the performance of the model can be improved.

[0131] In one embodiment of the present disclosure, if the first sub-task is a model training task, the allocation parameters include the computing power allocation parameters and the corresponding algorithm parameters of the model training task; if the first sub-task is a model inference task, the allocation parameters include the allocation information of model slicing of the inference model used in the model inference task.

[0132] In this embodiment, by differentially setting allocation parameters according to the first subtask type, for the model training task, combining the computing power allocation parameters with the algorithm parameters, the system can reasonably allocate computing power resources such as CPUs and GPUs according to the actual computing resources of the DPF. At the same time, by adjusting algorithm parameters (such as the learning rate and the type of optimization algorithm), the training process is optimized, the model convergence is accelerated, and the training efficiency and model quality are improved. For the model inference task, using the allocation information of model slicing, according to the performance characteristics of the DPF, the execution nodes of each part of the inference model can be reasonably divided to achieve parallel inference, reduce the inference latency, and improve the response speed, realizing the matching of resources and tasks and ensuring the resource utilization rate of model training and inference tasks in a distributed network environment.

[0133] As Figure 6 shown, according to another embodiment of the present disclosure, a collaborative inference method, prerequisite: the data service capabilities supported by each subnet have been registered / synchronized when the subnet joins the distributed network. The collaborative inference method includes:

[0134] Step S602, an AI service demand is detected, but the local resources do not support or meet the service.

[0135] In some embodiments, the first subnet (such as an in-network UE or other third-party application) generates an AI service demand and confirms that it lacks the data, computing power, or insufficient computing resources required for the service.

[0136] Step S604, send a collaborative service request.

[0137] In some embodiments, the first sub-SOF of the first subnet initiates an AI service request to the central network.

[0138] Step S606, decompose the service into tasks, generate QoS and resource requirements, and select a subnet based on the task requirements.

[0139] In some embodiments, the SOF of the central network decomposes the AI service into tasks, maps the SLA to task QoS, estimates the resource requirements of the tasks, and selects a suitable subnet based on the task requirements and the configuration information such as the capabilities and resources reported by each second subnet.

[0140] Step S608, send task information.

[0141] In some embodiments, the SOF of the central network sends the decomposed task information to the selected second subnet, including but not limited to task ID, task resource requirements, QoS requirements, task input data address, task output result address, etc.

[0142] Step S610, collaboratively execute the AI task.

[0143] In some embodiments, each second subnet collaboratively executes distributed AI tasks.

[0144] Step S612: Report the task execution progress and local resources.

[0145] In some embodiments, during the task execution process, each second subnet reports the progress and local resources to the central network in real time.

[0146] Step S614: Adjust the task.

[0147] In some embodiments, the central network dynamically adjusts the task configuration accordingly. For example, when the resources of the second subnet are insufficient, the task volume is reduced or the task is offloaded to other subnets, etc.

[0148] Step S616: Send configuration change information.

[0149] In some embodiments, if the task is adjusted, the central network sends the task configuration change information to each second subnet.

[0150] Step S618: Output the execution result.

[0151] In some embodiments, each second subnet continues to execute the task and outputs the task result to the central network after completing the task.

[0152] Step S620: Store the result.

[0153] In some embodiments, the central network stores the task execution result.

[0154] Step S622: Return the AI service result.

[0155] In some embodiments, the central network returns the AI service result to the first subnet.

[0156] As Figure 7 shown, for the collaborative inference method according to another embodiment of the present disclosure, the precondition is that the DPF and DSF nodes within each network have reported their respective resource and capability information to the DSMF, and the DSMF performs resource aggregation and reports the total resource and capability information of the DSMF to the SOF. The collaborative inference method includes:

[0157] Step S702: Send an AI service request.

[0158] In some embodiments, the AI service requester (which can be an NF, UE, third-party application, etc.) sends an AI service request to the first sub-SOF of the first subnet, carrying SLA requirements, etc.

[0159] Step S704: Decompose the service into tasks, estimate the task resource requirements, and find that the local resources are insufficient or do not support the service.

[0160] In some embodiments, the first sub-SOF decomposes the AI service into tasks, maps the SLA to the task QoS, estimates the resource requirements of the task, and finds that the local resources are lacking or the related service capabilities are not supported.

[0161] Step S706: Send an AI service request.

[0162] In some embodiments, the first sub-SOF initiates an AI service request to the central network through a service agent, and may also directly forward the decomposed task information to the SOF of the central network.

[0163] Step S708: Decompose the AI service into tasks, generate QoS and resource requirements, and select a subnet based on the task requirements.

[0164] In some embodiments, the SOF of the central network is orchestrated based on business logic and SLA, decomposes AI services into tasks, generates task QoS and resource requirements, and selects other subnets based on task requirements and resource information of each subnet.

[0165] Step S710, sending task information.

[0166] In some embodiments, the central network SOF sends the task information to the SOF of the second subnet through a service agent. The service agent that forwards the message in the central network can be the same agent as that in step 3, or can be a different agent.

[0167] Step S712: Select DSMF based on task requirements.

[0168] In some embodiments, the second sub-SOF selects an appropriate DSMF based on mission requirements.

[0169] Step S714, sending task information.

[0170] In some embodiments, the second sub-SOF sends task information to the selected DSMF, including but not limited to task ID, task resource requirements, QoS requirements, task input data address, task output result address, etc.

[0171] Step S716, collaboratively execute AI tasks.

[0172] In some embodiments, the DSMF of the second subnet and each AI service execution network element collaborate to execute tasks. During the task execution process, each execution node reports the task execution progress and its own resources to the DSMF in real time.

[0173] Step S718: resource reporting.

[0174] In some embodiments, the DSMF counts resource usage of each execution node and feeds back the resource status to the second sub-SOF.

[0175] Step S720, aggregating the resource usage within the network.

[0176] In some embodiments, the second sub-SOF aggregates the resource usage within the second subnet.

[0177] Step S722, reporting the subnet resources.

[0178] In some embodiments, the second sub-SOF reports the local resource status to the central network SOF.

[0179] Step S724, adjusting the tasks.

[0180] In some embodiments, the central network SOF adjusts the task allocation based on the reported information.

[0181] Step S726, sending task change information.

[0182] In some embodiments, the central network SOF sends the task change information to the SOF of the second subnet, and the SOF of the second subnet then sends the task information to the DSMF.

[0183] Step S728, continuing to execute the tasks.

[0184] In some embodiments, the DSMF and each task execution node continue to cooperate to complete the tasks.

[0185] Step S730, returning the task execution result.

[0186] In some embodiments, after the tasks are completed, the execution nodes of the second subnet feedback the task execution results to the DSF of the central network.

[0187] Step S732, storing the results.

[0188] In some embodiments, the DSF stores the task execution results.

[0189] Step S734, returning the AI service result.

[0190] In some embodiments, the central network returns the AI service result to the AI service requester of the first subnet.

[0191] Such as Figure 8 shown, according to another embodiment of the present disclosure, for the collaborative inference method, the precondition is that the data required for model training has been stored in the network DSF, and the collaborative inference method includes:

[0192] Step S802, selecting the DPF and DSF.

[0193] In some embodiments, the DSMF selects a suitable DPF based on capabilities and information tables to execute tasks and establish a DPF data interaction path.

[0194] Step S804, deploy data processing tasks.

[0195] In some embodiments, the data processing tasks are deployed to the first DPF, providing data processing methods, processed data output addresses, computing power configurations, etc.

[0196] Step S806, deploy AI model training tasks.

[0197] In some embodiments, the model training tasks are deployed to DPF-1 and DPF-2, providing QoS requirements (such as model accuracy), collaborative training node information, computing power algorithm configurations, model storage addresses, etc.

[0198] Step S808, deploy data storage tasks.

[0199] In some embodiments, the data storage tasks are deployed to the DSF, providing data storage methods, etc.

[0200] Step S810, obtain input data.

[0201] In some embodiments, the first DPF sends a request to the DSF to obtain the input data required for the task, and the DSF returns the data to the first DPF.

[0202] Step S812, data processing.

[0203] In some embodiments, the first DPF performs data processing such as denoising and privacy protection on the received data.

[0204] Step S814, send data.

[0205] In some embodiments, the first DPF sends the processed data to DPF-1 and DPF-2.

[0206] Step S816, execute model training.

[0207] In some embodiments, DPF-1 and DPF-2 perform distributed model training and interact with the data during the execution process.

[0208] Step S818, report model training progress and local resources.

[0209] In some embodiments,

[0210] Step S820, adjust computing power resource allocation.

[0211] In some embodiments, during the model training process, DPF-1 and DPF-2 report their own status, load, local resources, training progress (such as the current model accuracy), and other information to DSMF in real time.

[0212] Step S822, task configuration change information.

[0213] In some embodiments, DSMF dynamically adjusts the computing power allocation, algorithm parameters, etc. of the task according to the information reported by DPF and the task requirements.

[0214] DSMF sends the task configuration change information to each of DPF-1 and DPF-2.

[0215] Step S824, continue to execute the task.

[0216] In some embodiments, DPF-1 and DPF-2 continue to execute the model training task.

[0217] Step S826, output the result.

[0218] In some embodiments, after the training is completed, DPF-2 sends the result to DPF-1 for result aggregation. In this embodiment, DPF-1 is used as the main data plane functional unit, which is the designated address for the output result of DPF-2. DPF-1 aggregates the task execution results and outputs the final result to DSF.

[0219] Step S828, result storage.

[0220] In some embodiments, DSF stores the task execution results.

[0221] Such as Figure 9 As shown, according to another embodiment of the present disclosure, the collaborative inference method includes:

[0222] Step S902, select DPF.

[0223] In some embodiments, DSMF selects a suitable DPF to execute the task based on the capabilities and information table, and establishes a DPF data interaction path.

[0224] Step S904, deploy the joint inference task.

[0225] In some embodiments, DSMF deploys the joint inference task to DPF-1, DPF-2, and DPF-3, and provides the model acquisition address, model cutting method, result output address, etc.

[0226] Step S906, obtain the model.

[0227] In some embodiments, each DPF sends a request to DSF to download the model, and DSF returns the model to each DPF.

[0228] Step S908, AI joint inference.

[0229] In some embodiments, each DPF performs joint inference. DPF-3 performs model inference for a specified number of layers and sends the inference result to DPF-2. After DPF-2 finishes inferring the specified number of layers of the model, it sends the result to DPF-1, and DPF-1 continues the inference.

[0230] Step S910, reporting inference progress and computing power resources.

[0231] In some embodiments, during the inference process, each DPF reports information such as its own status, computing power resource usage, and inference progress to DSMF in real time.

[0232] Step S912, adjusting the model cut point.

[0233] In some embodiments, DSMF dynamically adjusts the model cut point based on this information. For example, when the status or progress of DPF-3 is poor, the number of its inference layers can be reduced, and the number of inference layers of other collaborating DPFs can be increased to reduce the amount of data processed by DPF-3.

[0234] Step S914, task configuration change information.

[0235] In some embodiments, DSMF sends the task configuration change information to each DPF.

[0236] Step S916, continuing to execute the inference task.

[0237] In some embodiments, each DPF continues to execute the AI inference task.

[0238] Step S918, outputting the inference result.

[0239] In some embodiments, after the inference is completed, DPF-1 returns the inference result to the third-party application through SOF.

[0240] In this embodiment, by introducing the service orchestration function (SOF) and the data service management function (DSMF), different network function units can be flexibly selected and orchestrated according to needs to complete data service tasks, thus improving the flexibility of the system. Through the unified scheduling and task decomposition of the central network, cross-subnet AI task collaboration can be achieved, solving the problem of task execution when subnets lack data, computing power, and computing resources. During the task execution process, the resource usage of each subnet can be monitored in real time, and the resource allocation can be dynamically adjusted according to the task requirements, improving the resource utilization rate and task execution efficiency. Through the collaborative work of multiple DPFs within the subnet, data services can be completed more efficiently, meeting the requirements of 6G networks for intelligence and high efficiency.

[0241] It should be noted that the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for restrictive purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0242] The following will refer to Figure 10 to describe the collaborative inference device 1000 according to an embodiment of the present disclosure. Figure 10 The shown collaborative inference device 1000 is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.

[0243] The collaborative inference device 1000 is presented in the form of a hardware module. The components of the collaborative inference device 1000 may include, but are not limited to: a decomposition module 1002, configured to decompose a model service request into a plurality of first subtasks based on the business logic of the model service request in response to the model service request sent by a first sub-SOF, where the model service request is generated when the first sub-SOF anticipates that the local resources do not meet or support the service request sent by the service requester, and the first sub-SOF is a service orchestration function network element of the first subnet managed by the central network management; a determination module 1004, configured to determine a quality of service (QoS) policy matching the first subtasks and determine the required resources for processing the first subtasks; a first selection module 1106, configured to select a plurality of second subnets for processing the plurality of first subtasks based on the QoS policy and the required resources, so that the plurality of second subnets collaboratively execute an inference task based on the corresponding plurality of first subtasks, where the plurality of first subtasks include model training and model inference; and a first receiving module 1108, configured to receive the task execution results fed back by the plurality of second subnets based on the inference task, integrate the task execution results to obtain a service result, and then feed the service result back to the first sub-SOF, so that the first sub-SOF feeds the service result back to the service requester.

[0244] The following will refer to Figure 11 to describe the collaborative inference device 1100 according to an embodiment of the present disclosure. Figure 11 The shown collaborative inference device 1100 is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.

[0245] The collaborative inference device 1100 is implemented in the form of a hardware module. The components of the collaborative inference device 1100 may include, but are not limited to: a second selection module 1102, which is configured to select, within the second subnet, a data plane function network element DPF that matches the task information in response to the task information of the second subtask sent by the second sub-SOF of the second subnet, so as to deploy the second subtask to multiple DPFs. The second subtask is obtained by decomposing the first subtask received by the second sub-SOF, and the first subtask is obtained by decomposing the model service request of the first subnet by the central SOF. The central SOF is a service orchestration function network element of the central network; a second receiving module 1104, which is configured to receive the subtask execution results fed back by the primary DPF among the multiple DPFs, and feed the subtask execution results back to the second sub-SOF, so that the second sub-SOF integrates the subtask execution results through the central SOF to obtain a service result and feeds it back to the first sub-SOF of the first subnet.

[0246] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to herein as "circuitry", "module", or "system".

[0247] The following refers to Figure 12 to describe the electronic device 1200 according to this embodiment of the present disclosure. It can be a network device or a terminal. Figure 12 The illustrated electronic device 1200 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0248] As Figure 12 shown, the electronic device 1200 is implemented in the form of a general-purpose computing device. The components of the electronic device 1200 may include, but are not limited to: the at least one processing unit 1210 described above, the at least one storage unit 1220 described above, and a bus 1230 connecting different system components (including the storage unit 1220 and the processing unit 1210).

[0249] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1210, so that the processing unit 1210 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 1210 can execute as Figure 3 described in the solution.

[0250] The storage unit 1220 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 12201 and / or a cache storage unit 12202, and may further include a read-only storage unit (ROM) 12203.

[0251] The storage unit 1220 may also include a program / utilities 12204 having a set (at least one) of program modules 12205. Such program modules 12205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0252] The bus 1230 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0253] The electronic device 1200 may also communicate with one or more external devices 1270 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1200, and / or may communicate with any device that enables the electronic device 1200 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 1250. Moreover, the electronic device 1200 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1260. As shown in the figure, the network adapter 1260 communicates with other modules of the electronic device 1200 through the bus 1230. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0254] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0255] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium having stored thereon a program product capable of implementing the above-described method of this specification. In some possible implementation manners, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0256] The program product for implementing the above method according to an embodiment of the present disclosure may be a portable compact disc read-only memory (CD-ROM) and include program code, and may run on an electronic device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0257] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0258] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0259] The program code contained on the readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0260] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user's computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0261] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.

[0262] In addition, although the various steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in that specific order, or that all of the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0263] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to cause a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the methods according to the embodiments of the present disclosure.

[0264] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of the disclosure that follow the general principles of the disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the disclosure are pointed out by the appended claims.

Claims

1. A collaborative reasoning method, characterized in that, Applied to the central SOF, where the central SOF is a service orchestration function network element of the central network, including: In response to a model service request sent by the first sub-SOF, decomposing the model service request into multiple first sub-tasks based on the business logic of the model service request, where the model service request is generated when the first sub-SOF anticipates that local resources do not meet or support the service request sent by the service requester, the first sub-SOF is a service orchestration function network element of the first subnet managed by the central network, and the multiple first sub-tasks include model training and model inference; Determine the service quality QoS policy matching the first sub-task, and determine the required resources for processing the first sub-task; Based on the QoS policy and the required resources, select multiple second subnets for processing the multiple first sub-tasks, so that the multiple second subnets cooperate to execute the inference task based on the corresponding multiple first sub-tasks; Receive the task execution results fed back by the multiple second subnets based on the inference task, and integrate the task execution results to obtain a service result and then feedback it to the first sub-SOF, so that the first sub-SOF feeds the service result back to the service requester.

2. The collaborative inference method according to claim 1, characterized in that Determine the service quality QoS policy matching the first sub-task, including: Based on the configured mapping rule library, configure different requirement parameters in the service requirement information carried by the first sub-task as corresponding initial QoS metrics; Adjust the initial QoS metrics based on the feedback network status and initial resource capabilities of the candidate subnets to obtain the target QoS metrics; Generate the QoS policy based on the target QoS metrics and the cooperation requirements of the first sub-task.

3. The collaborative reasoning method according to claim 2, wherein Determine the required resources for processing the first sub-task, including: Estimate the resource requirements of the first sub-task based on the target QoS metrics, so as to determine the required resources based on the estimation results and the transmission resources required between the first sub-tasks with data dependencies.

4. The collaborative reasoning method according to claim 2, wherein Based on the QoS policy and the required resources, select multiple second subnets for processing the multiple first sub-tasks, including: Obtain the first processing capabilities reported by the candidate subnets; Based on the first processing capabilities, the QoS policy and the required resources, determine the second subnet for executing the model training and the second subnet for executing the model inference.

5. The collaborative reasoning method according to claim 3, wherein Based on the QoS policy and the required resources, selecting multiple second subnets for processing the multiple first sub-tasks further includes: Send the task information of the multiple first sub-tasks to the second subnets, where the task information includes at least one of a task identifier, the required resources of the first sub-task, the QoS policy, the data address of the data required for processing the model service request, and the receiving address of the task execution result.

6. The collaborative reasoning method according to claim 1, wherein Before receiving the task execution results fed back by the multiple second subnets based on the inference task, further includes: Receiving first task execution information reported by the multiple second subnets during the execution of the first subtask, where the first task execution information includes task progress information and computing power resource usage information. Among them, within each of the second subnets, the second sub-SOF selects a matching data service management function network element (DSMF), and the DSMF selects multiple data plane function network elements (DPFs) to collaboratively execute the first subtask. Then, the multiple DPFs report their task-related data to the second sub-SOF through the DSMF, and the second sub-SOF integrates the task-related data to obtain the first task execution information; If it is determined to adjust the inference task based on the first task execution information, at least one of the task volume, task type, task priority, and inter-subnet collaboration relationship in the configuration information of the inference task is adjusted to obtain configuration change information; Sending the configuration change information to the multiple second subnets, so that the multiple second subnets continue to execute the inference task based on the configuration change information to obtain the task execution result.

7. The collaborative reasoning method according to claim 1, characterized in that, It further includes: Storing the service result in the first DSF, where the first DSF is the data storage function network element of the central network.

8. A collaborative reasoning method, characterized in that, Applied to the DSMF, the DSMF is the data service management function network element of the second subnet managed by the central network, and it includes: Responding to the task information of the second subtask sent by the second sub-SOF of the second subnet, selecting a data plane function network element (DPF) that matches the task information within the second subnet to deploy the second subtask to the multiple DPFs. The second subtask is obtained by the second sub-SOF decomposing the received first subtask, and the first subtask is obtained by the central SOF decomposing the model service request of the first subnet. The central SOF is the service orchestration function network element of the central network; Receiving the sub-task execution results fed back by the primary DPF among the multiple DPFs and feeding back the sub-task execution results to the second sub-SOF, so that the second sub-SOF integrates the sub-task execution results through the central SOF to obtain the service result and feeds it back to the first sub-SOF of the first subnet.

9. The collaborative reasoning method according to claim 8, wherein The first subtask includes a model training task. Selecting a data plane function network element (DPF) that matches the task information within the second subnet includes: Selecting a first DPF and a second DPF that match the task information within the second subnet. The first DPF is used to perform data processing operations on the model training data to obtain processed data, and the second DPF is used to perform model training operations based on the processed data.

10. The collaborative reasoning method according to claim 9, wherein Selecting a first DPF and a second DPF that match the task information within the second subnet includes: Screening out the first alternative DPFs with model training capabilities in a preset information table, and selecting the first DPF and the second DPF based on the second processing capabilities reported by the first alternative DPFs.

11. The collaborative reasoning method according to claim 9, wherein Deploying the second subtask to the multiple DPFs includes: Deploy the task of the data processing to the first DPF, where the deployment information of the data processing includes at least one of data processing method information, processed data output address information, and first computing power configuration information; Deploy the task of the model training to the second DPF, where the deployment information of the model training includes at least one of QoS requirements, collaborative training node information, second computing power configuration information, and training model storage address information.

12. The collaborative inference method according to claim 9, wherein: The first DPF obtains the model training data from a second DSF within the second internal network to which it belongs.

13. The collaborative inference method according to claim 8, wherein The first subtask includes a model inference task, and receiving the subtask execution results fed back by the master DPF among multiple DPFs includes: Receiving the inference result of the model inference task fed back by the master DPF as the subtask execution result.

14. The collaborative inference method according to claim 13, wherein Selecting a data plane functional network element DPF that matches the task information within the second subnet includes: Filtering out second alternative DPFs with inference capabilities in a preset information table, and selecting the matching DPF based on the second processing capabilities reported by the second alternative DPFs.

15. The collaborative inference method according to claim 13, wherein Deploying the second subtask to multiple DPFs includes: Deploying the model inference task to the DPF, where the deployment information of the model inference includes at least one of model acquisition address information, model cutting method information, and result output address information.

16. The collaborative reasoning method according to claim 8, wherein Before receiving the subtask execution results fed back by the master DPF among multiple DPFs, it further includes: Receiving second task execution information reported by multiple DPFs during the execution of the second subtask; If it is determined to adjust the first subtask based on the second task execution information, adjusting the allocation parameters of the first subtask; Sending the adjusted allocation parameters to the multiple DPFs so that the multiple DPFs continue to execute the second subtask based on the adjusted allocation parameters to obtain the subtask execution results.

17. The collaborative inference method according to claim 16, wherein: If the first subtask is a model training task, the allocation parameters include the computing power allocation parameters of the model training task and the corresponding algorithm parameters; If the first subtask is a model inference task, the allocation parameters include the allocation information of the model cutting of the inference model used in the model inference task.

18. A collaborative reasoning device, characterized in that, Applied to the central SOF, the central SOF is a service orchestration functional network element of the central network, and includes: A decomposition module, configured to respond to a model service request sent by a first sub-SOF, and decompose the model service request into multiple first subtasks based on the service logic of the model service request, where the model service request is generated when the first sub-SOF anticipates that local resources do not meet or support the service request sent by the service requester, the first sub-SOF is a service orchestration functional network element of a first subnet managed by the central network, and the multiple first subtasks include model training and model inference; A determination module, configured to determine a quality of service (QoS) policy matching the first subtask, and determine the required resources for processing the first subtask; A first selection module, configured to select multiple second subnets for processing the multiple first subtasks based on the QoS policy and the required resources, so that the multiple second subnets cooperate to execute an inference task based on the corresponding multiple first subtasks; A first receiving module, configured to receive the task execution results fed back by the multiple second subnets based on the inference task, integrate the task execution results to obtain a service result, and feed back the service result to the first sub-SOF, so that the first sub-SOF feeds back the service result to the service requester; 19. A collaborative reasoning device, characterized in that, Applied to a DSMF, where the DSMF is a data service management functional network element for central network management, including: A second selection module, configured to, in response to the task information of the second subtask sent by the second sub-SOF of the second subnet, select a data plane functional network element (DPF) matching the task information within the second subnet, and deploy the second subtask to the multiple DPFs. The second subtask is obtained by decomposing the first subtask received by the second sub-SOF, and the first subtask is obtained by decomposing the model service request of the first subnet by the central SOF. The central SOF is a service orchestration functional network element of the central network; A second receiving module, configured to receive the subtask execution results fed back by the primary DPF among the multiple DPFs, and feed back the subtask execution results to the second sub-SOF, so that the second sub-SOF integrates the subtask execution results through the central SOF to obtain a service result and feeds it back to the first sub-SOF of the first subnet; 20. A network device, characterized in that, Including: A processor; And A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the collaborative inference method according to any one of claims 1 to 7 or any one of claims 8 to 17 by executing the executable instructions; 21. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the collaborative inference method according to any one of claims 1 to 17; 22. A computer program product having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the collaborative inference method according to any one of claims 1 to 17.