Communication method and communication device
Through the multi-model collaboration technology in NWDAF, problems that cannot be met under strict analysis requirements are solved, efficient analysis results are provided under time and accuracy requirements, and network services are enhanced.
Patent Information
- Application Number
- CN202410013946.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art cannot meet the strict time and accuracy requirements of network consumers for analysis results. NWDAF can only reduce the requirements when it cannot meet the analysis requirements and cannot provide satisfactory analysis results under strict requirements.
The first network function in NWDAF uses multi-model collaboration technology to obtain another model to assist the local model to meet the analysis requirements and enhance network service guarantee capabilities by requesting the second network function collaborative reasoning.
Through multi-model collaborative reasoning, NWDAF can provide satisfactory analysis results under strict time and accuracy requirements, enhancing the guarantee capabilities of network services.
Smart Images

Figure CN120263678A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies. More specifically, it relates to a communication method and a communication device. Background Art
[0002] The network data analytics function (NWDAF) has functions of data collection, training, analysis, and reasoning. It can be used to collect relevant data from network elements, third-party servers, terminal devices, or network management systems, perform analysis and training based on the relevant data, and provide data analysis results to network elements, third-party servers, terminal devices, or network management systems. The network consumers served by this NWDAF (for example, network elements, third-party servers, terminal devices, or network management systems) expect the analysis results feedback by the NWDAF to meet given analysis requirements (such as time requirements and / or accuracy requirements). When the NWDAF cannot meet the analysis requirements given by the network consumers, the NWDAF will request the network consumers to lower the analysis requirements for the feedback analysis results. This method is only applicable to the situation where the network consumers have less strict analysis requirements for the current analysis. When the network consumers have strict requirements for the current analysis, the current technology cannot meet the analysis requirements of the network consumers. Summary of the Invention
[0003] This application provides a communication method and a communication device, which can achieve...
[0004] In a first aspect, a communication method is provided. This method can be executed by a first network function, or by a module (such as a chip or a circuit) in the first network function, or by a logical node, logical module, or software that can implement all or part of the first network function. This application does not make any limitations in this regard.
[0005] The method includes: the first network function receives a first message from a first device. The first message includes a first analysis requirement for feedback of a first analysis result, and the first network function is deployed with a first model; the first network function sends a second message to a second network function based on the first message. The second message is used to determine a second model, and the second model is used to assist the first model in obtaining the first analysis result before a first time.
[0006] Exemplarily, the first device is a network consumer. For example, the first device can be a network element, a third-party server, a terminal device, or a network management system, etc.
[0007] Exemplarily, the first network function and the second network function can be NWDAF.
[0008] Exemplarily, the above first analysis requirement includes a first time requirement and / or a first accuracy requirement.
[0009] Through the above method, when a single local model in the first network function cannot meet the analysis requirements given by the network consumer, the first network function will request to obtain another model for collaborative inference, so as to meet the analysis requirements of the network consumer for the analysis results and enhance the network service guarantee ability.
[0010] Combined with the first aspect, in some implementation manners of the first aspect, the above second message includes first indication information, and the first indication information is used to indicate the execution entity of the second model.
[0011] Exemplarily, if there are sufficient resources locally in the first network function to execute multiple models, the above first indication information may indicate that the execution entity of the second model is the first network function; if there are not enough resources locally in the first network function to execute multiple models, the above first indication information may indicate that the execution entity of the second model is other nodes or other devices, etc.
[0012] Combined with the first aspect, in some implementation manners of the first aspect, the above second message includes at least one of the following information:
[0013] The inference speed of the above second model, the analysis type of the above second model.
[0014] Through the above method, the first network function can determine the parameters of the other model based on the time requirement for the first device to request and feedback the analysis result and the parameters of the local model, and feedback the parameters of the other model to the second network function. The second network function can find the other model that meets the requirements, which can reduce the processing complexity of the second network function.
[0015] Combined with the first aspect, in some implementation manners of the first aspect, the above second message includes at least one of the following information:
[0016] The inference speed of the above first model, the analysis type of the above first model, the acceleration multiple for accelerating inference.
[0017] Through the above method, the first network function can directly feedback the parameters of the local model and the parameters for accelerating inference to the second network function. The second model can determine the parameters of the other model and find the other model that meets the requirements, which can reduce the processing complexity of the first network function.
[0018] Combined with the first aspect, in some implementation manners of the first aspect, the above method further includes: the first network function receives a third message from the above second network function, and the third message is used to indicate the second model; the first network function uses the first model and the second model for inference to obtain a first analysis result.
[0019] By the above method, when there are sufficient resources locally in the first network function to execute multiple models, allowing the first network function to execute multiple models can reduce the latency of signaling interaction between the multiple models.
[0020] Combined with the first aspect, in some implementation manners of the first aspect, the above method further includes: the above first network function obtains the requirements for the input data of the first model and the requirements for the output data of the second model.
[0021] Specifically, the requirements for the input data of the above first model include the input data of the above first model, and the requirements for the output data of the above second model include the number N of output values for each inference during the inference process of the second model, where N is a positive integer greater than or equal to 1.
[0022] Exemplarily, the first network function obtaining the requirements for the input data of the first model may specifically include: the above first network function may collect the input data of the first model according to the above first message, or the DCCF may schedule the input data of the first model to the above first network function, and this application does not limit this.
[0023] Exemplarily, the first network function obtaining the requirements for the output data of the second model may specifically include: the first network function generates the requirements for the output data of the second model, or the second device generates the requirements for the output data of the second model and indicates them to the first network function, and the second device is the execution entity of the second model.
[0024] Combined with the first aspect, in some implementation manners of the first aspect, the above method further includes: the first network function receives N first output values from the second device; the first network function uses the first model to verify the N first output values and determines to reject M of the N first output values; the first network function sends a fourth message to the second device, and the fourth message is used to indicate the M first output values.
[0025] Exemplarily, the above fourth message may indicate the M first output values by directly indicating the number of the M first output values, or the above fourth message may indicate the M first output values by indicating the first ratio.
[0026] By the above method, the execution entity of the first model and the execution entity of the second model can align the requirements for the input data of the model and the requirements for the output data of the model, which can ensure the smooth progress of the negotiation inference process of the multiple models.
[0027] Combined with the first aspect, in some implementation manners of the first aspect, the above method further includes: the above first network function obtains the requirements for the input data of the first model and the requirements for the output data of the first model.
[0028] Specifically, the requirements for the input data of the first model described above include the input data of the first model. The requirements for the output data of the first model include the number N of output values for each inference during the inference process of the first model, and N is a positive integer greater than or equal to 1.
[0029] Exemplarily, the first network function obtaining the requirements for the input data of the first model may specifically include: the first network function may collect the input data of the first model according to the first message, or DCCF may schedule the input data of the first model to the first network function, and the present application does not limit this.
[0030] Exemplarily, the first network function obtaining the requirements for the output data of the first model may specifically include: the first network function generates the requirements for the output data of the first model, or the second device generates the requirements for the output data of the first model and indicates them to the first network function, and the second device is the execution entity of the second model.
[0031] In combination with the first aspect, in some implementation manners of the first aspect, the method further includes: the first network function sends N first output values to the second device; the first network function receives a fifth message from the second device, and the fifth message is used to indicate M first output values, and the M first output values are the output values rejected by the second device among the N first output values.
[0032] Exemplarily, the fifth message may indicate the M first output values by directly indicating the number of the M first output values, or the fifth message may indicate the M first output values by indicating a first ratio.
[0033] Through the above method, the execution entities of the first model and the second model can align the requirements for the input data of the model and the requirements for the output data of the model, and can ensure the smooth progress of the negotiation inference process of multiple models.
[0034] In combination with the first aspect, in some implementation manners of the first aspect, before the first network function receives the first message from the first device, the method further includes: the first network function receives a seventh message from the first device, and the seventh message includes a second analysis requirement for feeding back a first analysis result; when the first network function does not meet the first analysis requirement, the first network function sends an eighth message to the first device, and the eighth message includes a third analysis requirement for feeding back the first analysis result, and the third analysis requirement is lower than the second analysis requirement and lower than the first analysis requirement.
[0035] Exemplarily, when the first network function fails to meet the second time requirement included in the second analysis requirement when using the first model for inference (i.e., the time taken by the first network function to use the first model for inference exceeds the second time requirement), the third analysis requirement fed back by the first network function to the first device includes a third time requirement, which is the time requirement that the first network function can meet when using the first model for inference. At this time, the first time requirement included in the first analysis requirement fed back by the first device is later than the time required by the second time requirement but earlier than the time required by the third time requirement.
[0036] Exemplarily, when the first network function fails to meet the second accuracy requirement included in the second analysis requirement when using the first model for inference, the third analysis requirement fed back by the first network function to the first device includes a third accuracy requirement, which is the accuracy requirement that the first network function can meet when using the first model for inference. At this time, the first accuracy requirement included in the first analysis requirement fed back by the first device can be lower than or equal to the second accuracy requirement included in the second analysis requirement, but the first accuracy requirement included in the first analysis requirement fed back by the first device is higher than the third accuracy requirement included in the third analysis requirement.
[0037] Through the above method, the first network function does not need to prejudge in advance whether it can meet the analysis requirements based on the first message. When the analysis requirements are not met, it requests the network consumer to lower the analysis requirements. The network consumer may moderately lower the analysis requirements or may not lower the analysis requirements. Based on this, the first network device triggers the multi-model collaborative inference process. Through the above method, the processing complexity of the first network function can be reduced.
[0038] In a second aspect, a communication method is provided. This method can be executed by a second network function, or by a module (such as a chip or a circuit) in the second network function, or by a logic node, a logic module, or software that can implement all or part of the second network function. This application does not make any limitations in this regard.
[0039] The method includes: the second network function receives a second message from the first network function; the second network function determines a second model according to the second message, and the second model is used to assist the first model in obtaining a first analysis result to meet the first analysis requirement, and the first model is deployed on the first network function.
[0040] Exemplarily, the above first network function and the above second network function can be NWDAF.
[0041] Exemplarily, the above first analysis requirement includes a first time requirement and / or a first accuracy requirement.
[0042] Through the above method, when the second network function can obtain the other party's model based on the request of the first network function to cooperate with the local model of the first network function for collaborative inference, the analysis requirements of network consumers for the analysis results can be met, and the network service guarantee ability can be enhanced.
[0043] Combined with the second aspect, in some implementation manners of the second aspect, the above second message includes first indication information, and the first indication information is used to indicate the execution entity of the second model.
[0044] Exemplarily, if the first network function has sufficient resources locally to execute multiple models, the above first indication information may indicate that the execution entity of the second model is the first network function; if the first network function does not have sufficient resources locally to execute multiple models, the above first indication information may indicate that the execution entity of the second model is other nodes or other devices, etc.
[0045] Combined with the second aspect, in some implementation manners of the second aspect, the above second message includes at least one of the following information:
[0046] The inference speed of the above second model, the analysis type of the above second model.
[0047] Through the above method, the first network function can determine the parameters of the other party's model based on the time requirement for the first device to request and feedback the analysis result and the parameters of the local model, and feedback the parameters of the other party's model to the second network function. The second network function can find the other party's model that meets the requirements, which can reduce the processing complexity of the second network function.
[0048] Combined with the second aspect, in some implementation manners of the second aspect, the above second message includes at least one of the following information:
[0049] The inference speed of the above first model, the analysis type of the above first model, the acceleration multiple for accelerating inference.
[0050] Through the above method, the first network function can directly feedback the parameters of the local model and the parameters for accelerating inference to the second network function. The second model can determine the parameters of the other party's model and find the other party's model that meets the requirements, which can reduce the processing complexity of the first network function.
[0051] Combined with the second aspect, in some implementation manners of the second aspect, the above method further includes: the second network function sends a third message to the first network function, and the third message is used to indicate the second model.
[0052] Through the above method, when the first network function has sufficient resources locally to execute multiple models, allowing the first network function to execute multiple models can reduce the delay of signaling interaction between multiple models.
[0053] In combination with the second aspect, in some implementations of the second aspect, the above method further includes: the second network function sends a sixth message to the second device, where the sixth message is used to indicate the first network function and the above-mentioned second model, and the second device is the execution entity of the second model.
[0054] Through the above method, the execution entity of the first model and the execution entity of the second model can be connected, and the process of multi-model negotiation and reasoning can be ensured to proceed smoothly.
[0055] In a third aspect, a communication method is provided. This method can be executed by the second device, or, alternatively, by a module (such as a chip or a circuit) in the second device, or, further, by a logical node, a logical module, or software that can implement all or part of the second device. This application does not make any limitations in this regard.
[0056] The method includes: the second device receives a sixth message from the second network function, where the sixth message is used to indicate the first network function and the second model, and the first network function deploys the first model; the second device performs reasoning using the second model, and the second model is used to assist the first model in obtaining a first analysis result to meet a first analysis requirement.
[0057] Exemplarily, the above first analysis requirement includes a first time requirement and / or a first accuracy requirement.
[0058] Through the above method, the execution entity of the first model and the execution entity of the second model can be connected, and the process of multi-model negotiation and reasoning can be ensured to proceed smoothly.
[0059] In combination with the third aspect, in some implementations of the third aspect, the above method further includes: the above-mentioned second device obtains requirements for the input data of the second model and requirements for the output data of the second model.
[0060] Specifically, the requirements for the input data of the second model include the input data of the second model, and the requirements for the output data of the second model include the number N of output values for each inference during the inference process of the second model, where N is a positive integer greater than or equal to 1.
[0061] Exemplarily, the second device obtaining the requirements for the input data of the second model may specifically include: the above-mentioned first network function can indicate the input data of the collected model to the second device, or DCCF can schedule the input data of the second model to the second device. This application does not make any limitations in this regard.
[0062] Exemplarily, the requirements for the second device to obtain the output data of the second model may specifically include: the requirement for the second device to generate the output data of the second model, or the requirement for the first network function to generate the output data of the second model and indicate it to the second device.
[0063] In combination with the third aspect, in some implementation manners of the third aspect, the above method further includes: the second device sends N first output values to the first network function; the second device receives a fourth message from the first network function, where the fourth message is used to indicate the M first output values, and the M first output values are the output values among the N first output values that are rejected by the first network function, where M is a positive integer less than or equal to N.
[0064] Exemplarily, the above fourth message may indicate the M first output values by directly indicating the number of the M first output values, or the above fourth message may indicate the M first output values by indicating a first ratio.
[0065] Through the above method, the execution entities of the first model and the second model can align the requirements for the input data and output data of the models, and can ensure the smooth progress of the negotiation and reasoning process of multiple models.
[0066] In combination with the third aspect, in some implementation manners of the third aspect, the above method further includes: the above second device obtains the requirements for the input data of the second model and the requirements for the output data of the first model.
[0067] Specifically, the requirements for the input data of the above second model include the input data of the above second model, and the requirements for the output data of the first model include the number N of output values for each inference during the inference process of the first model, where N is a positive integer greater than or equal to 1.
[0068] Exemplarily, the requirement for the second device to obtain the input data of the second model may specifically include: the above first network function may indicate the collected input data of the model to the second device, or DCCF may schedule the input data of the second model to the second device, and the present application does not limit this.
[0069] Exemplarily, the requirement for the second device to obtain the output data of the first model may specifically include: the requirement for the second device to generate the output data of the first model, or the requirement for the first network function to generate the output data of the first model and indicate it to the second device.
[0070] In combination with the third aspect, in some implementation manners of the third aspect, the above method further includes: The second device receives N first output values from the first network function; the second device uses the second model to verify the N first output values and determines to reject M first output values among the N first output values; the second device sends a fifth message to the first network function, and the fifth message is used to indicate the M first output values.
[0071] Exemplarily, the above fifth message may indicate the M first output values by directly indicating the number of the M first output values, or the above fifth message may indicate the M first output values by indicating a first ratio.
[0072] Through the above method, the execution entities of the first model and the second model can align the requirements for the input data and the output data requirements of the models, and can ensure the smooth progress of the negotiation and reasoning process of multiple models.
[0073] In a fourth aspect, a communication device is provided. The communication device includes: a transceiver unit, configured to receive a first message from a first device, where the first message includes a first analysis requirement for feeding back a first analysis result, and a first network function is deployed with a first model; the transceiver unit is further configured to send a second message to a second network function based on the first message, where the second message is used to determine a second model, and the second model is used to assist the first model in obtaining the first analysis result to meet the first analysis requirement.
[0074] In combination with the fourth aspect, in some implementation manners of the fourth aspect, the above second message includes first indication information, and the first indication information is used to indicate the execution entity of the second model.
[0075] In combination with the fourth aspect, in some implementation manners of the fourth aspect, the above second message includes at least one of the following information:
[0076] The inference speed of the above second model, the analysis type of the above second model.
[0077] In combination with the fourth aspect, in some implementation manners of the fourth aspect, the above second message includes at least one of the following information:
[0078] The inference speed of the above first model, the analysis type of the above first model, the acceleration multiple for accelerating inference.
[0079] In combination with the fourth aspect, in some implementation manners of the fourth aspect, the above transceiver unit is further configured to receive a third message from the above second network function, and the third message is used to indicate the second model; the communication device further includes: a processing unit, configured to perform inference using the first model and the second model to obtain a first analysis result.
[0080] In combination with the fourth aspect, in certain implementations of the fourth aspect, the above-mentioned transceiver unit is further configured to obtain requirements for the input data of the first model and requirements for the output data of the second model.
[0081] In combination with the fourth aspect, in certain implementations of the fourth aspect, the above-mentioned transceiver unit is further configured to receive N first output values from a second device; the above-mentioned processing unit is further configured to use the first model to verify the N first output values and determine to reject M of the N first output values; the above-mentioned transceiver unit is further configured to send a fourth message to the second device, and the fourth message is used to indicate the M first output values.
[0082] In combination with the fourth aspect, in certain implementations of the fourth aspect, the above-mentioned transceiver unit is further configured to obtain requirements for the input data of the first model and requirements for the output data of the first model.
[0083] In combination with the fourth aspect, in certain implementations of the fourth aspect, the above-mentioned transceiver unit is further configured to send N first output values to a second device; the above-mentioned transceiver unit is further configured to receive a fifth message from the second device, and the fifth message is used to indicate M first output values, and the M first output values are the output values rejected by the second device among the above-mentioned N first output values.
[0084] In combination with the fourth aspect, in certain implementations of the fourth aspect, before the above-mentioned transceiver unit is configured to receive a first message from a first device, the above-mentioned transceiver unit is further configured to receive a seventh message from the first device, and the seventh message includes a second analysis requirement for feeding back a first analysis result; when the above-mentioned communication device does not meet the first analysis requirement, the above-mentioned transceiver unit is further configured to send an eighth message to the first device, and the eighth message includes a third analysis requirement for feeding back a first analysis result, and the third analysis requirement is lower than the second analysis requirement, and the third analysis requirement is lower than the first analysis requirement.
[0085] In a fifth aspect, a communication device is provided, and the communication device includes: a transceiver unit configured to receive a second message from a first network function; the communication device further includes: a processing unit configured to determine a second model according to the second message, and the second model is used to assist the first model in obtaining a first analysis result to meet a first analysis requirement, and the first model is deployed on the first network function.
[0086] In combination with the fifth aspect, in certain implementations of the fifth aspect, the above-mentioned second message includes first indication information, and the first indication information is used to indicate the execution entity of the second model.
[0087] In combination with the fifth aspect, in certain implementations of the fifth aspect, the above-mentioned second message includes at least one of the following pieces of information:
[0088] The inference speed of the second model above, the analysis type of the second model above.
[0089] Combined with the fifth aspect, in some implementations of the fifth aspect, the second message above includes at least one of the following pieces of information:
[0090] The inference speed of the first model above, the analysis type of the first model above, the acceleration multiple for accelerating inference.
[0091] Combined with the fifth aspect, in some implementations of the fifth aspect, the transceiver unit above is further configured to send a third message to the first network function, and the third message is used to indicate the second model.
[0092] Combined with the fifth aspect, in some implementations of the fifth aspect, the transceiver unit above is further configured to send a sixth message to the second device, and the sixth message is used to indicate the first network function and the second model above, and the second device is the execution entity of the second model.
[0093] In a sixth aspect, a communication device is provided, and the communication device further includes: a transceiver unit, configured to receive a sixth message from a second network function, and the sixth message is used to indicate a first network function and a second model, and the first network function deploys a first model; the communication device further includes: a processing unit, configured to perform inference using the second model, and the second model is used to assist the first model in obtaining a first analysis result to meet a first analysis requirement.
[0094] Combined with the sixth aspect, in some implementations of the sixth aspect, the transceiver unit above is further configured to obtain requirements for input data of the second model, requirements for output data of the second model.
[0095] Combined with the sixth aspect, in some implementations of the sixth aspect, the transceiver unit above is further configured to send N first output values to the first network function; the transceiver unit is further configured to receive a fourth message from the first network function, and the fourth message is used to indicate the M first output values, and the M first output values are the output values rejected by the first network function among the N first output values, where M is a positive integer less than or equal to N.
[0096] Combined with the sixth aspect, in some implementations of the sixth aspect, the transceiver unit above is further configured to obtain requirements for input data of the second model, requirements for output data of the first model.
[0097] Combined with the sixth aspect, in some implementations of the sixth aspect, the transceiver unit above is further configured to receive N first output values from the first network function; the processing unit is further configured to use the second model to verify the N first output values and determine to reject M of the N first output values; the transceiver unit is further configured to send a fifth message to the first network function, and the fifth message is used to indicate the M first output values.
[0098] In a seventh aspect, a communication device is provided, including a processor, which is configured to cause the communication device to execute the method described in the first aspect and any possible implementation of the first aspect, or cause the communication device to execute the method described in the second aspect and any possible implementation of the second aspect, or cause the communication device to execute the method described in the third aspect and any possible implementation of the third aspect, by executing a computer program or instruction or through a logic circuit.
[0099] In a possible implementation, the communication device further includes a memory for storing the computer program or instruction.
[0100] In a possible implementation, the communication device further includes a communication interface for inputting and / or outputting signals.
[0101] In an eighth aspect, a communication device is provided, including a logic circuit and an input / output interface for inputting and / or outputting signals, and the logic circuit is configured to execute the method described in the first aspect and any possible implementation of the first aspect, or execute the method described in the second aspect and any possible implementation of the second aspect, or execute the method described in the third aspect and any possible implementation of the third aspect.
[0102] In a ninth aspect, a computer-readable storage medium is provided, on which a computer program or instruction is stored. When the computer program or the instruction runs on a computer, it causes the method described in the first aspect and any possible implementation of the first aspect to be executed, or causes the method described in the second aspect and any possible implementation of the second aspect to be executed, or causes the method described in the third aspect and any possible implementation of the third aspect to be executed.
[0103] In a tenth aspect, a computer program product is provided, including an instruction. When the instruction runs on a computer, it causes the method described in the first aspect and any possible implementation of the first aspect to be executed, or causes the method described in the second aspect and any possible implementation of the second aspect to be executed, or causes the method described in the third aspect and any possible implementation of the third aspect to be executed.
[0104] In an eleventh aspect, a communication system is provided, which includes the above-mentioned first network function and / or the above-mentioned second network function and / or the above-mentioned second device. The first network function is configured to execute the method described in the first aspect and any possible implementation of the first aspect, the second network function is configured to execute the method described in the second aspect and any possible implementation of the second aspect, and the second device is configured to execute the method described in the third aspect and any possible implementation of the third aspect.
[0105] For the related explanations and descriptions of the beneficial effects regarding the fourth to eleventh aspects, reference can be made to the descriptions of the first to third aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0106] Figure 1 FIG. 100 is a schematic diagram of a network architecture to which the technical solution of the present application can be applied.
[0107] Figure 2 FIG. 200 is a schematic flowchart of a communication method provided by an embodiment of the present application.
[0108] Figure 3 FIG. 300 is a schematic flowchart of a communication method provided by an embodiment of the present application.
[0109] Figure 4 FIG. 400 is a schematic flowchart of a communication method provided by an embodiment of the present application.
[0110] Figure 5 FIG. 500 is a schematic flowchart of a communication method provided by an embodiment of the present application.
[0111] Figure 6 FIG. 600 is a schematic flowchart of a communication method provided by an embodiment of the present application.
[0112] Figure 7 FIG. 700 is a schematic flowchart of a communication method provided by an embodiment of the present application.
[0113] Figure 8 FIG. 800 is a schematic block diagram of a communication device applicable to an embodiment of the present application.
[0114] Figure 9 FIG. 900 is a schematic block diagram of a communication device applicable to an embodiment of the present application.
[0115] Figure 10 FIG. 1000 is a schematic block diagram of a communication device applicable to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0116] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings.
[0117] To facilitate understanding of the embodiments of the present application, the following points are explained before introducing the embodiments of the present application.
[0118] In the present application, "for indicating" or "indicating" may include direct indication and indirect indication, or in other words, "for indicating" or "indicating" may explicitly and / or implicitly indicate. For example, when describing that a certain piece of information is for indicating information I, it may include that this piece of information directly indicates I or indirectly indicates I, and it does not necessarily mean that I is carried in this piece of information.
[0119] In the embodiments shown below, the first, second, third, fourth, and various numbers are only for the convenience of description and are not used to limit the scope of the embodiments of the present application. For example, to distinguish different messages, etc.
[0120] In the embodiments of the present application, words such as "exemplary", "for example", "exemplarily", "as (another) example", etc. are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" in the present application should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Rather, the use of the word "exemplary" is intended to present concepts in a specific manner.
[0121] The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0122] In the embodiments of the present application, the related descriptions regarding A sending a message, information, or data to B, and B receiving the message, information, or data from A are intended to illustrate which object the message, information, or data is to be sent to, and do not limit whether they are directly sent or indirectly sent via other nodes.
[0123] The technical solutions provided by the present application can be applied to various communication systems. For example, the fifth generation (5G) or NR system, LTE system, LTE frequency division duplex (FDD) system, LTE time division duplex (TDD) system, etc. The technical solutions provided by the present application can also be applied to non-terrestrial network (NTN) communication systems such as satellite communication systems. The technical solutions provided by the present application can also be applied to device-to-device (D2D) communication, vehicle-to-everything (V2X) communication, machine-to-machine (M2M) communication, machine type communication (MTC), and Internet of Things (IoT) communication systems or other communication systems. The technical solutions provided by the present application can also be applied to future communication systems, such as the sixth generation (6G) mobile communication system.
[0124] As an example, Figure 1 A schematic diagram of a network architecture 100 is shown.
[0125] As Figure 1 shown, this network architecture takes the 5th generation system (5GS) as an example. The network architecture may include a user equipment (UE) part, a (radio) access network ((R)AN) device, a user plane function (UPF), a unified data management (UDM), operations, administration and management (OAM), an access and mobility management function (AMF), a session management function (SMF), a network exposure function (NEF), a network repository function (NRF), a network data analytics function (NWDAF), an application function (AF), a policy control function (PCF), a unified data repository (UDR), and a data collection coordination function (DCCF), etc.
[0126] Next, a brief description will be given of Figure 1 each part involved in the network architecture.
[0127] 1. UE
[0128] The UE in this application can be any type of mobile terminal, fixed terminal or portable terminal. The UE in this application includes but is not limited to: user unit, user station, mobile station, mobile terminal, remote station, remote terminal device, mobile terminal device, user terminal device, wireless communication device, user agent, user device, cellular phone, cordless phone, session initiation protocol (SIP) phone, wireless local loop (WLL) station, personal digital assistant (PDA), handheld device with wireless communication function, computing device or other processing devices connected to a wireless modem, vehicle-mounted device, wearable device, terminal device in the internet of Things (IoT) system, household appliance, virtual reality device, user equipment in 2G / 3G / 4G / 5G / 6G network or user equipment in a future evolved public land mobile network (PLMN) or user equipment in a future vehicle-to-everything network, etc. This application does not limit this.
[0129] 2. (R)AN device
[0130] The (R)AN device of this application can manage radio resources, provide access services for the UE, and then complete the forwarding of control signals and UE data between the UE and the core network.
[0131] Exemplarily, (R)AN can be a node in a radio access network. (R)AN can be a base station, an evolved NodeB (eNodeB), a transmission reception point (TRP), a home base station (e.g., home evolved NodeB, or home Node B, HNB), a Wi-Fi access point (AP), a remote radio unit (RRU), a mobile switching center, a next generation NodeB (gNB) in a 5G mobile communication system, a next generation base station in a 6G mobile communication system, or a base station in a future mobile communication system, etc. The (R)AN device can also be a module or unit that completes some functions of the base station. For example, it can be a central unit (CU) or a distributed unit (DU). (R)AN can also be a device that undertakes the base station function in a D2D communication system, a V2X communication system, an M2M communication system, and an IoT communication system, etc. (R)AN can also be a network device in NTN, that is, (R)AN can be deployed on a high-altitude platform or a satellite. (R)AN can be a macro base station, a micro base station, or an indoor station, and can also be a relay node or a donor node, etc.
[0132] Embodiments of this application do not limit the specific technologies, device forms, and names adopted by (R)AN. For convenience of description, (R)AN will be uniformly referred to as an access network device hereinafter.
[0133] 3. UPF
[0134] UPF is used for packet routing and forwarding, and quality of service (QoS) processing of user plane data, etc. User data can access the data network (DN) through UPF.
[0135] 4. DN
[0136] DN is an operator network mainly used to provide data services for terminals. For example, the Internet, a third-party service network, or an IP multimedia service (IMS) network, etc.
[0137] 5. OAM
[0138] OAM is short for network management. OAM is mainly used to complete the analysis, prediction, planning, and configuration of daily networks and services, as well as the testing and fault management of networks and their services. OAM can interact with the RAN to obtain information such as radio channel conditions and radio resource utilization on the RAN side.
[0139] 6. AMF
[0140] The main functions of AMF include managing user registration, reachability detection, selection of SMF nodes, access authorization and authentication, mobility management, and management of mobile state transitions, etc.
[0141] 7. SMF
[0142] SMF is mainly used for session management, allocation and management of the Internet Protocol (IP) address of the UE, selection and management of the UPF, termination points of policy control and charging function interfaces, and downlink data notification, etc.
[0143] 8. NEF
[0144] NEF is mainly used to securely open the services and capabilities provided by the 3rd generation partnership project (3GPP) network functions to the outside, and support the secure interaction between the 3GPP network and third-party applications.
[0145] 9. UDM
[0146] It is used for unified data management, subscription data management of the UE, storage and management of UE identifiers, access authentication of the UE, registration or mobility management, etc.
[0147] 10. NRF
[0148] NRF is mainly responsible for providing the opening of the capabilities and events of the network to the outside, and receiving relevant external information.
[0149] 11. PCF
[0150] PCF is used for a unified policy framework to guide network behavior, and provides policy rule information for network functions (such as AMF, SMF, etc.) or UEs.
[0151] 12. UDR
[0152] UDR is mainly responsible for providing the storage capabilities of subscription data, policy data, and data related to capability opening.
[0153] 13. AF
[0154] The AF mainly supports the transmission of requirements from the application side to the network side. For example, quality of service (QoS) requirements or user status event subscriptions, etc. The AF can be the AF deployed by the operator network itself or a third-party AF.
[0155] 14. NWDAF
[0156] The NWDAF can have at least one of the following functions:
[0157] Data collection, model training, model feedback, analysis result inference, analysis result feedback, etc. Among them, the data collection function refers to collecting data from network elements, third-party servers, terminal devices, or network management systems; the model training function refers to performing analysis and training based on relevant input data to obtain a model; the model feedback function refers to sending the trained artificial intelligence (AI) model / machine learning (ML) model to the network element that supports the inference function; the analysis result inference function determines the data analysis result based on the trained AI model / ML model and the inference data; the analysis result feedback function can provide the data analysis result to the network element, third-party server, terminal device, or network management system.
[0158] According to different functions, the NWDAF can be further divided into an analytics logical function (AnLF) that supports analysis and inference and a model training logical function (MTLF) that supports model training. The AnLF can request model information from the MTLF and perform analysis result inference based on the model feedback by the MTLF.
[0159] The AnLF and the MTLF can each be a separate functional entity or can each be co-located with other functional entities. For example, the AnLF can be co-located with the AMF or co-located with the SMF.
[0160] 15. DCCF
[0161] The DCCF is responsible for connecting each network function (NF) of the data source and the NWDAF. The DCCF can collect data from each NF and schedule the collected data to the NWDAF.
[0162] It should be understood that the network functions included in the communication architecture 100 listed above are only for illustrative purposes, and the present application is not limited thereto.
[0163] In in Figure 1In the network architecture 100 shown, the N2 interface is the interface between the RAN device and the AMF, and is used for sending wireless parameters, non-access stratum (NAS) signaling, etc.; the N3 interface is the interface between the RAN device and the UPF, and is used for transmitting user plane data, etc.; the N4 interface is the interface between the SMF and the UPF, and is used for transmitting, for example, service policies, tunnel identification information of the N3 connection, data caching indication information, and downlink data notification messages, etc. The N6 interface is the interface between the DN and the UPF, and is used for transmitting user plane data, etc.
[0164] It should be understood that in the above network architecture 100, information interaction can be carried out between different network functions through service-based interfaces. For example, the NWDAF can collect data generated by the UE on these network functions from these network functions through service-based interfaces (such as Namf, Nsmf, etc.) provided by other network functions (such as AMF, SMF, etc.), and provide data analysis results (Analytics), models, data, etc. to other network functions (such as AMF, PCF, etc.) through the Nnwdaf interface.
[0165] It should be understood that the network architecture 100 applied to the embodiments of the present application is only an example of a network architecture illustrated from the perspectives of traditional point-to-point architectures and service-based architectures. The network architectures applicable to the embodiments of the present application are not limited thereto, and any network architecture capable of implementing the functions of the above-mentioned network elements is applicable to the embodiments of the present application. Additionally, other network architectures applicable to the embodiments of the present application may not include all the network functions shown in the above network architecture 100, or, other network architectures applicable to the embodiments of the present application may further include network functions other than those shown in the above network architecture 100. The present application does not make any limitations in this regard.
[0166] It should be noted that the names of the various network functions and the names of the interfaces in the present application are only examples. The present application does not exclude the possibility that the various network functions may have other names in the future, as well as the situation where the functions between the various network functions are merged. With the evolution of technology, any device or network element capable of implementing the above-mentioned various network functions is within the protection scope of the present application. Secondly, the above-mentioned network functions may also be referred to as instances, entities, devices, apparatuses, or modules, etc., and the present application does not make any particular limitations.
[0167] The network consumers served by the above NWDAF (e.g., each network element, third-party server, terminal device, or network management system in the network) expect that the analysis results fed back by the NWDAF can meet the given analysis requirements (e.g., time requirements and / or accuracy requirements). When the NWDAF fails to meet the analysis requirements given by the network consumers, the NWDAF will request the network consumers to lower the analysis requirements for the fed-back analysis results. This method is only applicable to the case where the network consumers have loose analysis requirements for the current analysis. When the network consumers have strict requirements for the current analysis, the current technology cannot meet the analysis requirements of the network consumers.
[0168] Based on the above technical problems, this application provides a communication method 200. When a single local model in the NWDAF cannot meet the analysis requirements given by the network consumers, another model is found to assist the local model in inferring the analysis results to meet the analysis requirements given by the network consumers, thereby enhancing the service guarantee ability of the NWDAF in the network.
[0169] Figure 2 It is a schematic flowchart of a communication method 200 provided by an embodiment of this application. In this embodiment, the first device, the first network function, and the second network function are used as the execution subjects for interactive illustration to illustrate this method, but this application does not limit the execution subjects of this interactive illustration. For example, Figure 2 the first device in can also be a chip, a chip system, or a processor that supports the method that the first device can implement, and can also be a logic module or software that can implement all or part of the first device; the first network function can also be a chip, a chip system, or a processor that supports the method that the first network function can implement, and can also be a logic module or software that can implement all or part of the first network function; the second network function can also be a chip, a chip system, or a processor that supports the method that the second network function can implement, and can also be a logic module or software that can implement all or part of the second network function.
[0170] The communication method 200 may include the following steps:
[0171] Step S210, the first device sends a first message to the first network function, and the first message includes a first analysis requirement for feeding back a first analysis result. Correspondingly, the first network function receives the first message from the first device.
[0172] Exemplarily, the above first analysis requirement may include a time requirement and / or an accuracy requirement.
[0173] Exemplarily, the first device may be the above network consumer.
[0174] Step S212, when the first model fails to meet the above first analysis requirement, the first network function sends a second message to the second network function, and this second message is used to request the second network function to determine a second model. Accordingly, the second network function receives the second message from the first network function.
[0175] Wherein, the above first model is a local model deployed on the first network function.
[0176] The above second message may include first indication information, and this first indication information is used to indicate the execution entity of the second model. Exemplarily, the execution entity of the second model may be the first network function or other entities, and the present application does not make any limitation thereto.
[0177] Exemplarily, the first network function's request for the second network function to determine the second model can be divided into the following five cases:
[0178] The first case: The above first analysis requirement includes a time requirement and the first network function determines that the local first model cannot meet this time requirement.
[0179] The first network function requests the second network function for a second model to accelerate the inference of the first analysis result. The above second message can carry information in the following two ways:
[0180] Way 1: The above second message may include at least one of the following pieces of information:
[0181] The inference speed of the second model, the analysis type of the second model, etc.
[0182] The above Way 1 is that the first network function determines the conditions that the second model needs to meet based on the information of the local model and the acceleration multiple for accelerating the inference and feeds them back to the second network function, and the second network function can find the second model that meets this condition.
[0183] Way 2: The above second message may include at least one of the following pieces of information:
[0184] The inference speed of the above first model, the analysis type of the above first model, the acceleration multiple for accelerating the inference, etc.
[0185] The above Way 2 is that the first network function provides the information of the local model and the acceleration multiple for accelerating the inference to the second network function, and it is required that the second network function determines the conditions that the second model needs to meet based on this and finds the second model that meets the requirements.
[0186] The second case: The above first analysis requirement includes an accuracy requirement and the first network function determines that the local first model cannot meet this accuracy requirement.
[0187] The first network function requests a second model that can meet the accuracy requirement from the second network function to assist the first model in obtaining a first analysis result to meet the accuracy requirement. The more parameters a model has, the higher the accuracy of the analysis result inferred by the model. At this time, the second model has more parameters than the first model, or, at this time, the second model is a large model and the first model is a small model.
[0188] The third case: The above first analysis requirement includes a time requirement and an accuracy requirement, and the first network function determines that the local first model can meet the time requirement but cannot meet the accuracy requirement.
[0189] The first network function requests a second model that can meet the accuracy requirement from the second network function to assist the first model in obtaining a first analysis result to meet the accuracy requirement. At this time, the second model has more parameters than the first model, or, at this time, the second model is a large model and the first model is a small model.
[0190] The fourth case: The above first analysis requirement includes a time requirement and an accuracy requirement, and the first network function determines that the local first model can meet the accuracy requirement but cannot meet the time requirement.
[0191] The first network function requests the second network function for the second model to accelerate the inference of the first analysis result. At this time, the second message can carry information in the above manner 1 and the above manner 2.
[0192] The fifth case: The above first analysis requirement includes a time requirement and an accuracy requirement, and the first network function determines that the local first model cannot meet the time requirement and the accuracy requirement.
[0193] The first network function requests the second network function for a second model that meets the accuracy requirement to accelerate the inference of the first analysis result. At this time, the second model has more parameters than the first model, or, at this time, the second model is a large model and the first model is a small model.
[0194] Step S214, the second network function determines the second model based on the above second message, and the second model is used to assist the above first model in obtaining the above first analysis result to meet the above first analysis requirement.
[0195] Exemplarily, the above first network function can be an NWDAF that supports AnLF, and the above second network function can be an NEDAF that supports MTLF.
[0196] Through the above communication method 200, when a single local model in the NWDAF cannot meet the analysis requirement for the network consumer to request an analysis result. Search for another model to assist the local model in obtaining the analysis result to meet the analysis requirement, thereby enhancing the service guarantee ability of the NWDAF in the network.
[0197] Specifically, in the following embodiments, taking the first network function as AnLF1 and the second network function as MTLF as an example, the technical solutions provided by the embodiments of the present application are described in detail. As Figure 3 shown in the schematic flowchart of the communication method 300.
[0198] The communication method 300 may include the following steps:
[0199] Step S310, the first device sends a first message to AnLF1. Correspondingly, AnLF1 receives the first message from the first device. The first message includes a first analysis requirement for feeding back a first analysis result.
[0200] Taking the time requirement included in the first analysis requirement that needs to be satisfied as an example, the communication method 300 is described in detail, which may correspond to the first or fourth case where the first network function requests the second network function to determine the second model. The first analysis requirement may include a first time requirement. Exemplarily, the first time requirement is to feed back the first analysis result before the first time, or the first time requirement is to feed back the first analysis result within the first time period.
[0201] Exemplarily, the above first message may be a subscription message (Nnwdaf_AnalyticsSubscription_Subscribe or Nnwdaf-_AnalyticsInfo_Request), and the subscription message further includes at least one of the following information:
[0202] Analysis identifier (Analytics ID), accuracy requirement of the first analysis result, etc.
[0203] Step S312, AnLF1 collects the input data of the model according to the first message.
[0204] Exemplarily, AnLF1 may collect the input data of the model from the corresponding network function, device, node, etc. according to the content included in the above first message. The input data of the model may also be referred to as the inference data of the model.
[0205] Step S314, AnLF1 determines that using the first model cannot meet the first time requirement according to the first message, the collected input data of the model, the time of collecting the input data of the model, and the inference speed of the first model.
[0206] Exemplarily, the first time requirement is to feedback the first analysis result within 500 seconds (s). AnLF1 used 100 s to collect the input data of 10,000 models, and the remaining time is 400 s. The inference speed of the first model is 10 per second (s). AnLF1 needs 1000 s to infer 10,000 inference data using the first model, which far exceeds the remaining 400 s. Therefore, AnLF1 can determine that using the first model cannot meet the first time requirement. At this time, AnLF1 triggers collaborative acceleration inference of multiple models.
[0207] Step S316, AnLF1 determines the inference speed of the second model, etc. according to the above first message, the collected input data of the model, the time for collecting the input of the collection model, and the inference speed of the first model.
[0208] Exemplarily, the first time requirement is to feedback the first analysis result within 500 seconds (s). AnLF1 used 100 s to collect the input data of 10,000 models, and the remaining time is 400 s. The current inference speed of the first model is 10 per second (s). If the first model can infer 10,000 inference data within the remaining 400 s, the inference speed of the first model needs to be accelerated to 25 per s. Therefore, the acceleration multiple required for the accelerated inference is 25 / 10 = 2.5. Exemplarily, in order to achieve the effect of 2.5 times of accelerated inference, the inference speed of the second model is about 100 per s.
[0209] Step S318, AnLF1 sends a second message to MTLF. Correspondingly, MTLF receives the second message from AnLF1.
[0210] Specifically, the above second message includes at least one of the following information:
[0211] The inference speed of the second model, the analysis type of the second model, the first indication information, the analysis identifier, the number of the second models, the filtering information of the second models, etc.
[0212] Exemplarily, when the number of the second models included in the second message is greater than 1, AnLF1 can request MTLF to determine multiple second models through the above second message, and this application does not limit this.
[0213] Exemplarily, the analysis type of the second model included in the second message may include parameters such as the model structure of the second model and the word segmentation type of the second model. The word segmentation type of the second model indicates the word segmentation algorithm used by the second model. The analysis type of the second model should be the same as the analysis type of the above first model.
[0214] Exemplarily, the first indication information included in the second message is used to indicate the execution entity of the second model. Exemplarily, if AnLF1 has sufficient local resources to execute multiple models, AnLF1 can indicate through the first indication information that the execution entity of the second model is AnLF1. Therefore, after determining the second model, MTLF needs to feedback the second model to AnLF1; if AnLF1 does not have sufficient local resources to execute multiple models, AnLF1 can indicate through the first indication information that the execution entity of the second model is the second device below. Therefore, after determining the second model, MTLF needs to feedback the second model to the second device. Optionally, AnLF1 may also not carry the above first indication information in the second message. AnLF1 not carrying the first indication information in the second message may default that MTLF feedbacks the second model to AnLF1 after determining the second model or default that MTLF does not feedback the second model to AnLF1 after determining the second model. This application does not make any limitations in this regard.
[0215] Exemplarily, the filtering information of the second model included in the second message is used to indicate the conditions satisfied for training the second model. When MTLF fails to find a second model that meets the requirements, MTLF can train a second model that meets the requirements based on the filtering information of the second model.
[0216] Step S320, MTLF searches for a second model that meets the requirements based on the second message.
[0217] After MTLF finds a second model that meets the requirements, it needs to feedback the second model to the execution entity of the second model.
[0218] 1. If the above first indication information indicates that the execution entity of the second model is AnLF1, or, by default, MTLF feedbacks the second model to AnLF1 after determining the second model, this communication method 300 further includes the following steps S322 to S324:
[0219] Step S322, MTLF sends a third message to AnLF1, and this third message is used to indicate the second model. Correspondingly, AnLF1 receives the third message from MTLF.
[0220] Exemplarily, the above third message includes at least one of the following information:
[0221] Analysis identifier, identifier (Model ID) of the second model, file address of the second model, etc.
[0222] Exemplarily, the above third message may be an Nnwdaf_MLModelProvision_Subscribe message.
[0223] Step S324, AnLF1 performs inference based on the local first model and the second model to obtain the above first analysis result.
[0224] II. If the above first indication information indicates that the execution entity of the second model is the second device, or by default MTLF does not feedback the second model to AnLF1 after determining the second model, the communication method 300 further includes the following steps S326 to S346:
[0225] Optionally, in step S326, MTLF determines that the execution entity of the second model is AnLF2 (an example of the second device).
[0226] It should be noted that when the above first indication information indicates that the execution entity of the second model is the second device, MTLF does not need to determine the execution entity of the second model. When the above second message does not include the above first indication information and by default MTLF does not feedback the second model to AnLF1 after determining the second model, MTLF needs to determine the execution entity of the second model.
[0227] Exemplarily, MTLF can determine AnLF2 as the execution entity of the second model from the multiple connected AnLFs. For example, the current computing power resources, algorithm resources, etc. of this AnLF2 can meet the resource requirements for executing the second model.
[0228] Optionally, in step S328, MTLF sends message #1 to AnLF1, and this message #1 is used to indicate AnLF2. Correspondingly, AnLF1 receives message #1 from MTLF.
[0229] It should be noted that when the above first indication information indicates that the execution entity of the second model is the second device, MTLF does not need to send message #1 to AnLF1. When the above second message does not include the above first indication information and by default MTLF does not feedback the second model to AnLF1 after determining the second model, MTLF needs to determine the execution entity of the second model and send message #1 to AnLF1.
[0230] Exemplarily, the above message #1 may include the identifier of AnLF2, the address of AnLF2, etc.
[0231] In step S330, MTLF sends the sixth message to AnLF2, and this sixth message is used to indicate the second model and AnLF1. Correspondingly, AnLF2 receives the sixth message from MTLF.
[0232] Exemplarily, the above sixth message may include the analysis identifier, the identifier (Model ID) of the second model, the file address of the second model, the identifier of AnLF1, the address of AnLF1, etc.
[0233] The above steps S326 to S330 can enable the execution entities of the first model and the second model to establish a connection.
[0234] After the execution entities of the first model and the second model establish a connection, it is necessary to negotiate the requirements for the input data and output data of the models, etc. Specifically, the following two methods are included:
[0235] Method 1: Assume that the first model is a large model and the second model is a small model.
[0236] Step S332, AnLF1 and AnLF2 obtain the requirements for the input data of the model and the requirements for the output data of the second model.
[0237] Exemplarily, the requirements for the input data of the model may include the input data of the first model and the input data of the second model. AnLF1 can indicate the collected input data of the model to AnLF2. For example, AnLF1 can send the label (tag) or the dataset itself of the collected input dataset of the model to AnLF2.
[0238] Exemplarily, the requirements for the output data of the second model may include the number N of output values (tokens) for each inference during the inference process of the second model, where N is a positive integer greater than or equal to 1. The requirements for the output data of the second model can be generated by AnLF1 and indicated to AnLF2, or the requirements for the output data of the second model can be generated by AnLF2 and indicated to AnLF1. This application does not make a limitation in this regard.
[0239] Step S334, AnLF2 uses the second model to perform an inference once and generates N first output values (tokens), and sends the N first output values to AnLF1. Correspondingly, AnLF1 receives the N first output values from AnLF2.
[0240] Step S336, AnLF1 uses the first model to verify the above N first output values and determines to reject M of the N first output values.
[0241] Optionally, the above step S336 may also be: AnLF1 uses the first model to verify the above N first output values and determines to reject the first proportion of the N first output values. This first proportion can also be called the selection ratio.
[0242] Step S338, AnLF1 sends a fourth message to AnLF2, and this fourth message is used to indicate the above M first output values. Correspondingly, AnLF2 receives this fourth message from AnLF1.
[0243] Optionally, the above fourth message may include the above first ratio or the number of the above M first output values, etc., which are not limited in this application.
[0244] Exemplarily, the above N first output values are 7 tokens. After AnLF1 verifies the 7 tokens, it determines to accept the first 5 tokens and reject the last 2 tokens; then AnLF1 instructs AnLF2 to reject the last 2 tokens among the 7 tokens, and AnLF2 can re-infer the last 2 tokens accordingly.
[0245] Alternatively, for another example, the above N first output values are 7 tokens. After AnLF1 verifies the 7 tokens, it determines to accept the first 5 tokens and reject the last 2 tokens; then AnLF1 instructs AnLF2 to reject 2 / 7 of the 7 tokens, and AnLF2 can re-infer the 2 / 7 of the tokens accordingly.
[0246] Execute multiple inferences through the above steps S334 to S338 until the first analysis result is inferred.
[0247] Method 2: Assume that the first model is a small model and the second model is a large model.
[0248] Step S340, AnLF1 and AnLF2 obtain the requirements for the input data of the model and the requirements for the output data of the first model.
[0249] Exemplarily, the requirements for the input data of the model may include the input data of the first model and the input data of the second model. AnLF1 can indicate the collected input data of the model to AnLF2. For example, AnLF1 can send the label (tag) or the dataset itself of the collected input dataset of the model to AnLF2.
[0250] Exemplarily, the requirements for the output data of the first model may include the number N of output values (tokens) for each inference during the inference process of the first model, where N is a positive integer greater than or equal to 1. The requirements for the output data of the first model may be generated by AnLF1 and indicated to AnLF2, or the requirements for the output data of the first model may be generated by AnLF2 and indicated to AnLF1, which are not limited in this application.
[0251] Step S342, AnLF1 uses the first model to perform one inference and generates N first output values (tokens), and sends the N first output values to AnLF2. Correspondingly, AnLF2 receives the N first output values from AnLF1.
[0252] Step S344. AnLF2 uses the second model to verify the above N first output values and determines to reject M first output values among the N first output values.
[0253] Optionally, the above step S344 may also be: AnLF2 uses the second model to verify the above N first output values and determines to reject the first output values of the first proportion among the N first output values. This first proportion may also be referred to as the selection ratio.
[0254] Step S346. AnLF2 sends a fifth message to AnLF1, and this fifth message is used to indicate the above M first output values. Correspondingly, AnLF1 receives this fifth message from AnLF2.
[0255] Optionally, the above fifth message may include the above first proportion or the number of the above M first output values, etc., and the present application does not limit this.
[0256] Exemplarily, the above N first output values are 7 tokens. After AnLF2 verifies the 7 tokens, it determines to accept the first 5 tokens and reject the last 2 tokens; then AnLF2 indicates to AnLF1 to reject the last 2 tokens among the 7 tokens, and AnLF1 can re - reason about the last 2 tokens accordingly.
[0257] Or, for another example, the above N first output values are 7 tokens. After AnLF2 verifies the 7 tokens, it determines to accept the first 5 tokens and reject the last 2 tokens; then AnLF2 indicates to AnLF1 to reject 2 / 7 of the tokens among the 7 tokens, and AnLF1 can re - reason about the 2 / 7 tokens accordingly.
[0258] By performing multiple inferences through the above steps S340 to S346 until the first analysis result is inferred.
[0259] Step S348. AnLF1 feeds back the first analysis result to the first device.
[0260] Specifically, AnLF feeding back the first analysis result to the first device can meet the above first - time requirement.
[0261] The technical solution of the above method 300 details the process of how to find another model to assist the local model to accelerate the acquisition of the analysis result, promotes the cooperation between different network functions of NWDAF, and thus can enhance the service guarantee ability of NWDAF in the network.
[0262] For the above communication method 300, AnLF1 determines the relevant parameters of the second model, and MTLF only needs to find a second model that meets the requirements. This application can also provide a communication method 400, in which MTLF determines the relevant parameters of the second model and finds a second model that meets the requirements, as Figure 4 shown in the schematic flowchart of the communication method 400.
[0263] The communication method 400 may include the following steps:
[0264] Step S410, the first device sends a first message to AnLF1. Correspondingly, AnLF1 receives the first message from the first device. The first message includes a first analysis requirement for feeding back a first analysis result.
[0265] Taking the time requirement included in the first analysis requirement that needs to be met as an example, the communication method 400 is described in detail, which can correspond to the first or fourth case where the first network function requests the second network function to determine the second model. The first analysis requirement may include a first time requirement. Exemplarily, the first time requirement is to feed back the first analysis result before a first time, or the first time requirement is to feed back the first analysis result within a first time period.
[0266] Exemplarily, the above first message may be a subscription message (Nnwdaf_AnalyticsSubscription_Subscribe or Nnwdaf-_AnalyticsInfo_Request), and the subscription message further includes at least one of the following information:
[0267] Analysis ID, accuracy requirement of the first analysis result, etc.
[0268] Step S412, AnLF1 collects the input data of the model according to the first message.
[0269] Exemplarily, AnLF1 may collect the input data of the model from the corresponding network function, device, node, etc. according to the content included in the above first message. The input data of the model may also be referred to as the inference data of the model.
[0270] Step S414, AnLF1 determines that using the first model cannot meet the first time requirement according to the first message, the collected input data of the model, the time for collecting the input data of the model, and the inference speed of the first model.
[0271] Exemplarily, the first time requirement is to feedback the first analysis result within 500 seconds (s). AnLF1 used 100s to collect the input data of 10,000 models, and the remaining time is 400s. The inference speed of the first model is 10 per second (s). It takes 1000s for AnLF1 to finish inferring 10,000 inference data using the first model, far exceeding the remaining 400s. Therefore, AnLF1 can determine that using the first model cannot meet the first time requirement. At this time, AnLF1 triggers collaborative acceleration inference of multiple models.
[0272] Step S416, AnLF1 sends a second message to MTLF, and this second message is used to request MTLF to determine the second model. Correspondingly, MTLF receives the second message from AnLF1.
[0273] Specifically, the above second message includes at least one of the following information:
[0274] The inference speed of the first model, the analysis type of the first model, the acceleration multiple of the acceleration inference, the first indication information, the analysis identifier, the number of the second models, the filtering information of the second models, etc.
[0275] Exemplarily, the analysis type of the first model included in the second message may include parameters such as the model structure of the first model and the word segmentation type of the first model. The word segmentation type of the first model indicates the word segmentation algorithm used by the first model.
[0276] Exemplarily, the acceleration multiple of the acceleration inference included in the second message can be calculated as follows: Assume that the first time requirement is to feedback the first analysis result within 500 seconds (s). AnLF1 used 100s to collect the input data of 10,000 models, and the remaining time is 400s. The current inference speed of the first model is 10 per second (s). If the first model can finish inferring 10,000 inference data within the remaining 400s, the inference speed of the first model needs to be accelerated to 25 per s. Therefore, the required acceleration multiple of the acceleration inference is 25 / 10 = 2.5.
[0277] Exemplarily, when the number of the second models included in the second message is greater than 1, AnLF1 can request MTLF to determine multiple second models through the above second message, and this application does not make any limitation on this.
[0278] Exemplarily, the first indication information included in the second message is used to indicate the execution entity of the second model. Exemplarily, if AnLF1 has sufficient local resources to execute multiple models, AnLF1 can indicate through the first indication information that the execution entity of the second model is AnLF1. Therefore, after determining the second model, MTLF needs to feedback the second model to AnLF1; if AnLF1 does not have sufficient local resources to execute multiple models, AnLF1 can indicate through the first indication information that the execution entity of the second model is the second device below. Therefore, after determining the second model, MTLF needs to feedback the second model to the second device. Optionally, AnLF1 may also not carry the above first indication information in the second message. AnLF1 not carrying the first indication information in the second message may default that MTLF feedbacks the second model to AnLF1 after determining the second model or default that MTLF does not feedback the second model to AnLF1 after determining the second model. This application does not make any limitations on this.
[0279] Exemplarily, the filtering information of the second model included in the second message is used to indicate the conditions satisfied for training the second model. When MTLF fails to find a second model that meets the requirements, MTLF can train a second model that meets the requirements based on the filtering information of the second model.
[0280] Step S418, MTLF determines the inference speed of the second model, the analysis type of the second model, etc. according to the second message.
[0281] Assume that the current inference speed of the first model included in the second message is 10 per second and the acceleration multiple for accelerated inference is 2.5. Based on this, the inference speed of the second model can be determined to be approximately 100 per second.
[0282] The analysis type of the second model should be the same as the analysis type of the first model included in the second message.
[0283] Step S420, MTLF searches for a second model that meets the requirements based on the determined inference speed of the second model, the analysis type of the second model, etc.
[0284] After MTLF finds a second model that meets the requirements, it needs to feedback the second model to the execution entity of the second model.
[0285] Steps S422 to S448 can refer to the above steps S322 to S348 and will not be elaborated here.
[0286] The technical solution of the above method 400 details the process of finding another model to assist the local model in accelerating the acquisition of analysis results, promoting the collaboration between different network functions of NWDAF, thereby enhancing the service guarantee ability of NWDAF in the network.
[0287] In the above communication method 300 and the above communication method 400, the execution entity of the first model first collects data, then pre-judges whether multi-model collaborative acceleration inference needs to be triggered, then determines the execution entity of the second model, and finally the execution entity of the first model instructs the collected data to the execution entity of the second model or the execution entity of the second model collects data by itself. This process may need to go through two data collection processes successively. If there are more models for collaborative acceleration inference, this process needs to go through more data collection processes. Based on this, the present application can also provide a communication method 500, which can reduce the number of times of collecting data from the data source in the above multi-model collaborative acceleration inference process, thereby reducing the data collection time.
[0288] As Figure 5 shown in the schematic flowchart of the communication method 500, the communication method 500 may include the following steps:
[0289] Step S510, the first device sends a first message to AnLF1. Correspondingly, AnLF1 receives the first message from the first device. The first message includes a first analysis requirement for feeding back a first analysis result.
[0290] The communication method 500 is described in detail by taking the time requirement included in the first analysis requirement as an example, which may correspond to the first case or the fourth case where the first network function requests the second network function to determine the second model. The first analysis requirement may include a first time requirement. Exemplarily, the first time requirement is to feed back the first analysis result before the first time, or the first time requirement is to feed back the first analysis result within the first time period.
[0291] Exemplarily, the above first message may be a subscription message (Nnwdaf_AnalyticsSubscription_Subscribe or Nnwdaf-_AnalyticsInfo_Request), and the subscription message further includes at least one of the following information:
[0292] Analysis identifier (Analytics ID), accuracy requirement of the first analysis result, etc.
[0293] Step S512, AnLF1 estimates the amount of input data of the model to be collected, the time for collecting the input data of the model, etc., and determines that the first model cannot meet the first time requirement according to the above first message, the estimated amount of input data of the model, the estimated time for collecting the input data of the model, and the inference speed of the first model.
[0294] Exemplarily, the first time requirement is to feedback the first analysis result within 500 seconds (s). AnLF1 estimates that it needs to collect the input data of 10,000 models, and the estimated time for collecting the input data of these 10,000 models is about 100 s, and the remaining time is 400 s. The inference speed of the first model is 10 per second. AnLF1 needs 1000 s to infer 10,000 inference data using the first model, which far exceeds the remaining 400 s. Therefore, AnLF1 can determine that using the first model cannot meet the first time requirement. At this time, AnLF1 triggers multi-model collaborative accelerated inference.
[0295] Step S514, AnLF1 determines the inference speed of the second model, etc. according to the above first message, the estimated amount of input data of the model, the estimated time for collecting the input data of the model, and the inference speed of the first model.
[0296] Exemplarily, the first time requirement is to feedback the first analysis result within 500 seconds (s). AnLF1 needs 100 s to collect the input data of 10,000 models, and the remaining inference time is 400 s. The current inference speed of the first model is 10 per second. If the first model can infer 10,000 inference data within the remaining 400 s, the inference speed of the first model needs to be accelerated to 25 per second. Therefore, the acceleration multiple required for accelerated inference is 25 / 10 = 2.5. Exemplarily, in order to achieve the effect of 2.5 times of accelerated inference, the inference speed of the second model is about 100 per second.
[0297] Step S516, AnLF1 sends a second message to MTLF. Correspondingly, MTLF receives the second message from AnLF1.
[0298] Specifically, the above second message includes at least one of the following information:
[0299] The inference speed of the second model, the analysis type of the second model, the first indication information, the analysis identifier, the number of the second models, the filtering information of the second models, etc.
[0300] Exemplarily, when the number of the second models included in the second message is greater than 1, AnLF1 can request MTLF to determine multiple second models through the above second message, which is not limited in this application.
[0301] Exemplarily, the analysis type of the second model included in the second message may include parameters such as the model structure of the second model and the word segmentation type of the second model. The word segmentation type of the second model indicates the word segmentation algorithm used by the second model. The analysis type of the second model should be consistent with the analysis type of the above first model.
[0302] Exemplarily, the first indication information included in the second message is used to indicate the execution entity of the second model. Exemplarily, if AnLF1 has sufficient local resources to execute multiple models, AnLF1 can indicate through the first indication information that the execution entity of the second model is AnLF1. Therefore, after determining the second model, MTLF needs to feedback the second model to AnLF1; if AnLF1 does not have sufficient local resources to execute multiple models, AnLF1 can indicate through the first indication information that the execution entity of the second model is the second device below. Therefore, after determining the second model, MTLF needs to feedback the second model to the second device. Optionally, AnLF1 may also not carry the above first indication information in the second message. AnLF1 not carrying the first indication information in the second message may default that MTLF feedbacks the second model to AnLF1 after determining the second model or default that MTLF does not feedback the second model to AnLF1 after determining the second model. This application does not make any limitations in this regard.
[0303] Exemplarily, the filtering information of the second model included in the second message is used to indicate the conditions satisfied for training the second model. When MTLF fails to find a second model that meets the requirements, MTLF can train a second model that meets the requirements based on the filtering information of the second model.
[0304] Step S320, MTLF searches for a second model that meets the requirements based on the second message.
[0305] After MTLF finds a second model that meets the requirements, it needs to feedback the second model to the execution entity of the second model.
[0306] 1. If the above first indication information indicates that the execution entity of the second model is AnLF1, or, by default, MTLF feedbacks the second model to AnLF1 after determining the second model, this communication method 500 further includes the following steps S520 to S524:
[0307] Step S322, MTLF sends a third message to AnLF1, and this third message is used to indicate the second model. Correspondingly, AnLF1 receives the third message from MTLF.
[0308] Exemplarily, the above third message includes at least one of the following information:
[0309] Analysis identifier, identifier (Model ID) of the second model, file address of the second model, etc.
[0310] Exemplarily, the above third message may be an Nnwdaf_MLModelProvision_Subscribe message.
[0311] Step S522: AnLF1 collects the input data of the first message collection model according to the above.
[0312] Exemplarily, AnLF1 can collect the input data of the model from the corresponding network functions, devices, nodes, etc. according to the content included in the above first message. The input data of this model can also be referred to as the inference data of the model.
[0313] Step S524: AnLF1 performs inference based on the local first model and the second model to obtain the above first analysis result.
[0314] Second, if the above first indication information indicates that the execution entity of the second model is the second device, or by default MTLF does not feedback the second model to AnLF1 after determining the second model, the communication method 500 further includes the following steps S526 to S548:
[0315] Optionally, in step S526, MTLF determines that the execution entity of the second model is AnLF2 (an example of the second device).
[0316] It should be noted that when the above first indication information indicates that the execution entity of the second model is the second device, MTLF does not need to determine the execution entity of the second model. When the above second message does not include the above first indication information and by default MTLF does not feedback the second model to AnLF1 after determining the second model, MTLF needs to determine the execution entity of the second model.
[0317] Exemplarily, MTLF can determine AnLF2 as the execution entity of the second model from the multiple AnLFs it is connected to. For example, the current computing power resources, algorithm resources, etc. of this AnLF2 can meet the resource requirements for executing the second model.
[0318] Optionally, in step S528, MTLF sends message #1 to AnLF1, and this message #1 is used to indicate AnLF2. Correspondingly, AnLF1 receives message #1 from MTLF.
[0319] It should be noted that when the above first indication information indicates that the execution entity of the second model is the second device, MTLF does not need to send message #1 to AnLF1. When the above second message does not include the above first indication information and by default MTLF does not feedback the second model to AnLF1 after determining the second model, MTLF needs to determine the execution entity of the second model and send message #1 to AnLF1.
[0320] Exemplarily, the above message #1 can include the identifier of AnLF2, the address of AnLF2, etc.
[0321] Step S530, MTLF sends a sixth message to AnLF2, and the sixth message is used to indicate the second model and AnLF1. Accordingly, AnLF2 receives the sixth message from MTLF.
[0322] Exemplarily, the above-mentioned sixth message may include an analysis identifier, an identifier (Model ID) of the second model, a file address of the second model, an identifier of AnLF1, an address of AnLF1, etc.
[0323] The above steps S526 to S530 can enable the execution entities of the first model and the second model to establish a connection.
[0324] After the execution entities of the first model and the second model establish a connection, it is necessary to negotiate requirements for the input data of the model, requirements for the output data of the model, etc.
[0325] Specifically, the requirements for the input data of the model negotiated by the execution entities of the first model and the second model may specifically be to negotiate the input data of the model. Including step S532, DCCF may collect the input data of the model from the corresponding network functions, devices, nodes, etc. according to the content included in the above first message, and schedule the collected input data of the model to the execution entity AnLF1 of the first model and the execution entity AnLF2 of the second model, so that AnLF1 obtains the input data of the first model and AnLF2 obtains the input data of the second model.
[0326] Exemplarily, DCCF may indicate the collected input data of the model to AnLF1 and AnLF2. For example, DCCF may send the tag of the collected input data set of the model or the data set itself to AnLF1 and AnLF2.
[0327] Specifically, the requirements for the output data of the model negotiated by the execution entities of the first model and the second model include the following two methods:
[0328] Method 1: Assume that the first model is a large model and the second model is a small model.
[0329] Step S534, AnLF1 and AnLF2 obtain the requirements for the output data of the second model.
[0330] Exemplarily, the requirements for the output data of the second model may include the number N of output values (tokens) for each inference during the inference process of the second model, and N is a positive integer greater than or equal to 1. The requirements for the output data of the second model may be generated and indicated to AnLF2 by AnLF1, or the requirements for the output data of the second model may be generated and indicated to AnLF1 by AnLF2. This application does not make a limitation on this.
[0331] Step S536, AnLF2 performs one inference using the second model and generates N first output values (tokens), and sends the N first output values to AnLF1. Correspondingly, AnLF1 receives the N first output values from AnLF2.
[0332] Step S538, AnLF1 uses the first model to verify the above N first output values and determines to reject M first output values among the N first output values.
[0333] Optionally, the above step S538 may also be: AnLF1 uses the first model to verify the above N first output values and determines to reject the first output values of the first ratio among the N first output values. This first ratio may also be referred to as the acceptance and rejection ratio.
[0334] Step S540, AnLF1 sends a fourth message to AnLF2, and the fourth message is used to indicate the above M first output values. Correspondingly, AnLF2 receives the fourth message from AnLF1.
[0335] Optionally, the above fourth message may include the above first ratio or the quantity of the above M first output values, etc., and the present application does not limit this.
[0336] Exemplarily, the above N first output values are 7 tokens. After AnLF1 verifies the 7 tokens, it determines to accept the first 5 tokens and reject the last 2 tokens; then AnLF1 indicates to AnLF2 to reject the last 2 tokens among the 7 tokens, and AnLF2 can re - perform inference on the last 2 tokens accordingly.
[0337] Or, for another example, the above N first output values are 7 tokens. After AnLF1 verifies the 7 tokens, it determines to accept the first 5 tokens and reject the last 2 tokens; then AnLF1 indicates to AnLF2 to reject 2 / 7 of the tokens among the 7 tokens, and AnLF2 can re - perform inference on the 2 / 7 tokens accordingly.
[0338] By performing the above steps S536 to S540 multiple times, the first analysis result is inferred until.
[0339] Method 2: Assume that the first model is a small model and the second model is a large model.
[0340] Step S542, AnLF1 and AnLF2 obtain the requirements for the output data of the first model.
[0341] Exemplarily, the requirements for the output data of the first model may include the number N of output values (tokens) for each inference during the inference process of the first model, where N is a positive integer greater than or equal to 1. The requirements for the output data of the first model may be generated by AnLF1 and indicated to AnLF2, or the requirements for the output data of the first model may be generated by AnLF2 and indicated to AnLF1. This application does not make any limitations in this regard.
[0342] Step S544, AnLF1 performs an inference using the first model and generates N first output values (tokens), and sends the N first output values to AnLF2. Correspondingly, AnLF2 receives the N first output values from AnLF1.
[0343] Step S546, AnLF2 uses the second model to verify the above N first output values and determines to reject M first output values among the N first output values.
[0344] Optionally, the above step S546 may also be: AnLF2 uses the second model to verify the above N first output values and determines to reject the first output values with a first ratio among the N first output values. This first ratio may also be referred to as the selection ratio.
[0345] Step S548, AnLF2 sends a fifth message to AnLF1, and this fifth message is used to indicate the above M first output values. Correspondingly, AnLF1 receives this fifth message from AnLF2.
[0346] Optionally, the above fifth message may include the above first ratio or the quantity of the above M first output values, etc. This application does not make any limitations in this regard.
[0347] Exemplarily, the above N first output values are 7 tokens. After AnLF2 verifies the 7 tokens, it determines to accept the first 5 tokens and reject the last 2 tokens; then AnLF2 indicates to AnLF1 to reject the last 2 tokens among the 7 tokens, and AnLF1 can re-perform the inference on the last 2 tokens accordingly.
[0348] Or, for another example, the above N first output values are 7 tokens. After AnLF2 verifies the 7 tokens, it determines to accept the first 5 tokens and reject the last 2 tokens; then AnLF2 indicates to AnLF1 to reject 2 / 7 of the tokens among the 7 tokens, and AnLF1 can re-perform the inference on the 2 / 7 of the tokens accordingly.
[0349] By performing the above steps S544 to S548 multiple times of inference until the first analysis result is inferred.
[0350] Step S550, AnLF1 feeds back the first analysis result to the first device.
[0351] Specifically, AnLF feeding back the first analysis result to the first device can meet the above first time requirement.
[0352] The technical solution of the above method 500 describes in detail the process of finding another model to assist the local model to accelerate the acquisition of the analysis result, promotes the cooperation between different network functions of NWDAF, and thus can enhance the service guarantee ability of NWDAF in the network. In addition, in the communication method 500, DCCF can be used to collect data from the data source and schedule it to the execution entities corresponding to multiple models, which can reduce the number of times of collecting data from the data source during the above multi-model collaborative acceleration inference process, thereby reducing the data collection time.
[0353] For the above communication method 500, it is only necessary for AnLF1 to determine the relevant parameters of the second model and for MTLF to find a second model that meets the requirements. The present application can also provide a communication method 600, in which MTLF determines the relevant parameters of the second model and finds a second model that meets the requirements, as Figure 6 shown in the schematic flowchart of the communication method 600.
[0354] The communication method 600 may include the following steps:
[0355] Step S610, the first device sends a first message to AnLF1. Correspondingly, AnLF1 receives the first message from the first device. The first message includes a first analysis requirement for feeding back the first analysis result.
[0356] Taking the time requirement included in the first analysis requirement that needs to be met as an example, the communication method 600 is described in detail, which can correspond to the first or fourth case where the first network function requests the second network function to determine the second model. The first analysis requirement may include a first time requirement. Exemplarily, the first time requirement is to feed back the first analysis result before the first time, or the first time requirement is to feed back the first analysis result within the first time period.
[0357] Exemplarily, the above first message may be a subscription message (Nnwdaf_AnalyticsSubscription_Subscribe or Nnwdaf-_AnalyticsInfo_Request), and the subscription message further includes at least one of the following information:
[0358] Analysis ID, accuracy requirement of the first analysis result, etc.
[0359] Step S612, AnLF1 estimates the amount of input data required for the model to be collected, the time for collecting the input data of the model, etc., and determines that the first model cannot meet the first time requirement based on the above first message, the estimated amount of input data of the model, the estimated time for collecting the input data of the model, and the inference speed of the first model.
[0360] Exemplarily, the first time requirement is to feedback the first analysis result within 500 seconds (s). AnLF1 estimates that 10,000 input data of the model need to be collected, and the estimated time for collecting these 10,000 input data of the model is about 100 s, and the remaining time is 400 s. The inference speed of the first model is 10 per second, and it takes 1000 s for AnLF1 to infer 10,000 inference data using the first model, far exceeding the remaining 400 s. Therefore, AnLF1 can determine that using the first model cannot meet the first time requirement. At this time, AnLF1 triggers multi-model collaborative accelerated inference.
[0361] Step S614, AnLF1 sends a second message to MTLF, and this second message is used to request MTLF to determine the second model. Correspondingly, MTLF receives the second message from AnLF1.
[0362] Specifically, the above second message includes at least one of the following information:
[0363] The inference speed of the first model, the analysis type of the first model, the acceleration multiple of the accelerated inference, the first indication information, the analysis identifier, the number of the second models, the filtering information of the second models, etc.
[0364] Exemplarily, the analysis type of the first model included in the second message may include parameters such as the model structure of the first model and the word segmentation type of the first model, and the word segmentation type of the first model indicates the word segmentation algorithm used by the first model.
[0365] Exemplarily, the acceleration multiple of the accelerated inference included in the second message can be calculated as follows: Assume that the first time requirement is to feedback the first analysis result within 500 seconds (s). AnLF1 estimates that 10,000 input data of the model need to be collected, and the estimated time for collecting these 10,000 input data of the model is about 100 s, and the remaining time is 400 s. The current inference speed of the first model is 10 per second (s). If the first model can infer 10,000 inference data within the remaining 400 s, the inference speed of the first model needs to be accelerated to 25 per second. Therefore, the required acceleration multiple of the accelerated inference is 25 / 10 = 2.5.
[0366] Exemplarily, when the number of second models included in the second message is greater than 1, AnLF1 may request MTLF to determine multiple second models through the above second message, and this application does not make any limitations in this regard.
[0367] Exemplarily, the first indication information included in the second message is used to indicate the execution entity of the second model. Exemplarily, if AnLF1 has sufficient local resources to execute multiple models, AnLF1 may indicate through the first indication information that the execution entity of the second model is AnLF1. Therefore, after determining the second model, MTLF needs to feedback the second model to AnLF1; if AnLF1 does not have sufficient local resources to execute multiple models, AnLF1 may indicate through the first indication information that the execution entity of the second model is the second device below. Therefore, after determining the second model, MTLF needs to feedback the second model to the second device. Optionally, AnLF1 may also not carry the above first indication information in the second message. AnLF1 not carrying the first indication information in the second message may default that MTLF feedbacks the second model to AnLF1 after determining the second model or default that MTLF does not feedback the second model to AnLF1 after determining the second model, and this application does not make any limitations in this regard.
[0368] Exemplarily, the filtering information of the above second model included in the second message is used to indicate the conditions satisfied for training the second model. When MTLF fails to find a second model that meets the requirements, MTLF may train a second model that meets the requirements based on the filtering information of the second model.
[0369] Step S616, MTLF determines the inference speed of the second model, the analysis type of the second model, etc. according to the second message.
[0370] Assume that the current inference speed of the first model included in the second message is 10 per second and the acceleration multiple for accelerated inference is 2.5. Based on this, the inference speed of the second model can be determined to be approximately 100 per second.
[0371] The analysis type of the second model should be consistent with the analysis type of the first model included in the second message.
[0372] Step S618, MTLF searches for a second model that meets the requirements based on the determined inference speed of the second model, the analysis type of the second model, etc.
[0373] After MTLF finds a second model that meets the requirements, it needs to feedback the second model to the execution entity of the second model.
[0374] Steps S620 to S650 may refer to the above steps S520 to S550, and will not be elaborated here.
[0375] The technical solution of the above method 600 details the process of finding another model to assist the local model in accelerating the acquisition of analysis results, which promotes the collaboration between different network functions of NWDAF, thereby enhancing the service guarantee ability of NWDAF in the network. In addition, in the communication method 600, DCCF can be used to collect data from the data source and schedule it to the execution entities corresponding to multiple models, which can reduce the number of times of collecting data from the data source during the above multi-model collaborative acceleration inference process, thereby reducing the data collection time.
[0376] In the above communication methods 300, 400, 500, and 600, AnLF1 determines whether the local model can meet the analysis requirements based on the analysis requirements in the above first message, thereby triggering multi-model collaborative inference. The present application can also provide another communication method 700. This communication method 700 does not require AnLF1 to pre-judge whether the local model can meet the analysis requirements in the first message, which can reduce the processing complexity of AnLF1.
[0377] As Figure 7 shown in the schematic flowchart of the communication method 700, the communication method 700 may include the following steps:
[0378] Step S710, the first device sends a seventh message to AnLF1, and the seventh message includes a second analysis requirement for feeding back the first analysis result. Correspondingly, AnLF1 receives the seventh message from the first device.
[0379] Exemplarily, the above seventh message may be a subscription message (Nnwdaf_AnalyticsSubscription_Subscribe or Nnwdaf-_AnalyticsInfo_Request).
[0380] The second analysis requirement in the communication method 700 may include a second time requirement (expected waitingtime / time when analytics is needed) and / or a second accuracy requirement. The above seventh message may further include an analysis identifier (Analytics ID), etc.
[0381] Step S712, AnLF1 collects the input data of the model according to the first message.
[0382] Step S714, AnLF1 performs inference using the first model.
[0383] When the inference of the first model by AnLF1 times out or the accuracy of the inference of the first model by AnLF1 is insufficient, in step S716, AnLF1 sends an eighth message to the first device, and the eighth message includes a third analysis requirement. Accordingly, the first device receives the eighth message from AnLF1.
[0384] The above-mentioned third analysis requirement may include a third time requirement (revised waiting time) and / or a third accuracy requirement. The third time requirement or the third accuracy requirement is the time requirement or accuracy requirement that can be satisfied by AnLF1 using the first model for inference.
[0385] In step S718, the first device sends the above-mentioned first message to AnLF1, and the first message includes a first analysis requirement for feeding back the first analysis result. Accordingly, AnLF1 receives the first message from the first device.
[0386] Exemplarily, the above-mentioned first analysis requirement may include a first time requirement (expected waiting time / time when analytics is needed) and / or a first accuracy requirement.
[0387] Exemplarily, if the inference of the first model by AnLF1 cannot meet the above-mentioned second time requirement, then the time required by the first time requirement is later than the time required by the above-mentioned second time requirement and the time required by the first time requirement is earlier than the time required by the above-mentioned third time requirement.
[0388] Exemplarily, if the inference of the first model by AnLF1 cannot meet the above-mentioned second accuracy requirement, then the accuracy required by the first accuracy requirement may be equal to or lower than the accuracy required by the above-mentioned second accuracy requirement, and the accuracy required by the first accuracy is higher than the accuracy required by the above-mentioned third accuracy requirement.
[0389] In step S720, AnLF1 triggers collaborative inference of dual models.
[0390] Exemplarily, if AnLF1 triggers collaborative inference of dual models to meet the above-mentioned first time requirement, then AnLF1 triggers collaborative accelerated inference of dual models, corresponding to the first case and the fourth case where the first network function requests the second network function to determine the second model.
[0391] Exemplarily, if AnLF1 triggers collaborative inference of dual models to meet the above-mentioned first accuracy requirement, then AnLF1 triggers collaborative inference of dual models, corresponding to the second case and the third case where the first network function requests the second network function to determine the second model.
[0392] Exemplarily, if AnLF1 triggers dual-model collaborative inference to meet the above-mentioned first time requirement and first accuracy requirement, then AnLF1 triggers dual-model collaborative inference, corresponding to the fifth case where the first network function requests the second network function to determine the second model.
[0393] The steps after AnLF1 triggers dual-model collaborative inference can refer to steps S316 to S348 in communication method 300. Alternatively, the steps after AnLF1 triggers dual-model collaborative inference can refer to steps S416 to S448 in communication method 400.
[0394] Through the above communication method 700, another triggering method for AnLF1 to trigger dual-model collaborative inference is provided. This method does not require AnLF1 to pre-judge whether the local model can meet the analysis requirements in the first message, which can reduce the processing complexity of AnLF1.
[0395] Finally, the device embodiments of this application are introduced.
[0396] To implement the various functions in the method provided in this application, the first device, the first network function, and the second network function can all include a hardware structure and / or a software module, and implement the above various functions in the form of a hardware structure, a software module, or a combination of a hardware structure and a software module. Whether a certain function among the above various functions is executed in the form of a hardware structure, a software module, or a combination of a hardware structure and a software module depends on the specific application and design constraint conditions of the technical solution.
[0397] Figure 8 It is a schematic block diagram of a communication device 800 according to an embodiment of this application. The communication device 800 includes a processor 810 and a communication interface 820. Optionally, the processor 810 and the communication interface 820 can be connected to each other through a bus 830. The communication device 800 can be the first network function, the second network function, the first device, or the second device.
[0398] Optionally, the communication device 800 can further include a memory 840. The memory 840 includes but is not limited to a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a compact disc read-only memory (CD-ROM). The memory 840 is used to store relevant instructions and data.
[0399] The processor 810 may be one or more central processing units (CPUs). When the processor 810 is a single CPU, it may be a single-core CPU or a multi-core CPU.
[0400] When the communication device 800 is a first network function, for example, the communication device 800 is used to perform the following operations: receiving a first message from a first device, or sending a second message to a second network function, etc.
[0401] When the communication device 800 is a second network function, for example, the communication device 800 is used to perform the following operations: receiving a second message from a first network function, or determining a second model according to the second message, etc.
[0402] When the communication device 800 is a first device, for example, the communication device 800 is used to perform the following operations: sending a first message to a first network function, etc.
[0403] The above content is only an exemplary description. When the communication device 800 is a first network function / second network function / first device, it will be responsible for executing the methods or steps related to the first network function / second network function / first device in the foregoing method embodiments.
[0404] The above description is only an exemplary description. For specific content, reference may be made to the content shown in the foregoing method embodiments. Figure 8 The implementation of each operation in Figures 2 to 7 may also correspondingly refer to the corresponding description of the method embodiment shown in
[0405] Figure 9 is a schematic block diagram of the communication device 900 according to an embodiment of the present application. The communication device 900 may be a first network function or a second network function or a first device, or may be a chip or module in a first network function or a second network function or a first device, and is used to implement the methods involved in the foregoing embodiments. The communication device 900 includes a transceiver unit 910 and a processing unit 920. The transceiver unit 910 and the processing unit 920 will be introduced exemplarily below.
[0406] The transceiver unit 910 may include a sending unit and a receiving unit. The sending unit is used to perform the sending action of the communication device 900, and the receiving unit is used to perform the receiving action of the communication device 900. For the sake of convenience of description, in the embodiments of the present application, the sending unit and the receiving unit are combined into a transceiver unit. This is explained uniformly here and will not be repeated later.
[0407] When the communication device 900 is a first network function, for example, the transceiver unit 910 is used to receive a first message from a first device, and the processing unit 920 is used to determine that the first model cannot meet the first time requirement according to the first message, the input data of the model, the time for collecting the input data of the model, and the inference speed of the first model, etc.
[0408] When the communication device 900 is a second network function, for example, the transceiver unit 910 is used to receive a second message from the first network function, and the processing unit 920 is used to determine a second model according to the second message.
[0409] When the communication device 900 is a first device, for example, the transceiver unit 910 is used to send a first message to the first network function.
[0410] The above content is only for exemplary description. When the communication device 900 is a first network function / second network function / first device, it will be responsible for executing the methods or steps related to the first network function or second network function or first device in the foregoing method embodiments.
[0411] Optionally, the communication device 900 further includes a storage unit 930, and the storage unit 930 is used to store programs or codes for executing the foregoing methods.
[0412] Figure 8 and Figure 9 The device embodiments shown are used to implement Figures 2 to 7 the content described above. Figure 8 and Figure 9 For the specific execution steps and methods of the device shown, reference may be made to the content described in the foregoing method embodiments.
[0413] Figure 10 It is a schematic block diagram of the communication device 1000 according to an embodiment of the present application. The communication device 1000 is used to implement the functions of a first network function / second network function / first device. The communication device 1000 may be a chip in a first network function / second network function / first device.
[0414] The communication device 1000 includes: an input / output interface 1020 and a processor 1010. The input / output interface 1020 may be an input / output circuit. The processor 1010 may be a signal processor, a chip, or other integrated circuits that can implement the methods of the present application. Among them, the input / output interface 1020 is used for input or output of signals or data.
[0415] For example, when the communication device 1000 is a first network function, the input / output interface 1020 is used to receive a first message from a first device. The processor 1010 is used to determine that the first model cannot meet the first time requirement according to the first message, the input data of the model, the time for collecting the input data of the model, and the inference speed of the first model.
[0416] For example, when the communication device 1000 is a second network function, the input / output interface 1020 is used to receive a second message from the first network function, and the processor 1010 is used to determine a second model according to the second message.
[0417] For example, when the communication device 1000 is a first device, the input / output interface 1020 is used to send a first message to the first network function.
[0418] In a possible implementation, the processor 1010 realizes the functions implemented by the first network function or the second network function or the first device by executing instructions stored in the memory.
[0419] Optionally, the communication device 1000 further includes a memory.
[0420] Optionally, the processor and the memory are integrated together.
[0421] Optionally, the memory is outside the communication device 1000.
[0422] In a possible implementation, the processor 1010 may be a logic circuit, and the processor 1010 inputs / outputs messages or signaling through the input / output interface 1020. The logic circuit may be a signal processor, a chip, or other integrated circuits that can implement the methods of the embodiments of the present application.
[0423] The above description of the communication device 1000 is only an exemplary description. The communication device 1000 can be used to execute the methods described in the foregoing embodiments. For specific content, reference can be made to the description of the foregoing method embodiments, which will not be elaborated herein.
[0424] The present application also provides a chip, including a processor, which is used to call and run instructions stored in the memory, so that a communication device installed with the chip executes the methods in the above examples.
[0425] The present application also provides a chip, including: an input interface, an output interface, and a processor. The input interface, the output interface, and the processor are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the methods in the above examples. Optionally, the chip further includes a memory, and the memory is used to store computer programs or code.
[0426] The present application also provides a processor for coupling with a memory and for executing the methods and functions related to the first network function, the second network function, or the first device in any one of the foregoing embodiments.
[0427] The present application provides a computer program product containing instructions. When the computer program product runs on a computer, the methods of the foregoing embodiments can be implemented.
[0428] The present application also provides a computer program. When the computer program runs on a computer, the methods of the foregoing embodiments can be implemented.
[0429] The present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the methods described in the foregoing embodiments are implemented.
[0430] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0431] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0432] In several embodiments provided by the present application, the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings, direct couplings, or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in an electrical, mechanical, or other form.
[0433] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the technical solution of the embodiments of the present application.
[0434] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit.
[0435] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of various method embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0436] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A communication method, characterized in that, The method includes: A first network function receives a first message from a first device, the first message including a first analysis requirement for feeding back a first analysis result, and the first network function is deployed with a first model; The first network function sends a second message to a second network function based on the first message, the second message being used to determine a second model, and the second model is used to assist the first model in obtaining the first analysis result to meet the first analysis requirement.
2. The method according to claim 1, wherein The second message includes first indication information, and the first indication information is used to indicate the execution entity of the second model.
3. The method according to claim 1 or 2, characterized in that, The second message includes at least one of the following: The inference speed of the second model, the analysis type of the second model.
4. The method according to claim 1 or 2, characterized in that, The second message includes at least one of the following: The inference speed of the first model, the analysis type of the first model, the acceleration multiple for accelerating inference.
5. The method according to any one of claims 1 to 4, characterized in that The method further includes: The first network function receives a third message from the second network function, the third message being used to indicate the second model; The first network function performs inference using the first model and the second model to obtain the first analysis result.
6. The method according to any one of claims 1 to 4, characterized in that The method further includes: The first network function obtains the requirements for the input data of the first model and the requirements for the output data of the second model.
7. The method according to claim 6, characterized in that, The requirements for the input data of the first model include the input data of the first model, and the requirements for the output data of the second model include the number N of output values for each inference during the inference process of the second model, where N is a positive integer greater than or equal to 1.
8. The method according to claim 7, wherein The method further includes: The first network function receives N first output values from a second device, and the second device is the execution entity of the second model; The first network function verifies the N first output values using the first model and determines to reject M first output values among the N first output values; The first network function sends a fourth message to the second device, and the fourth message is used to indicate the M first output values, where M is a positive integer less than or equal to N.
9. The method according to any one of claims 1 to 4, characterized in that The method further includes: The first network function obtains the requirements for the input data of the first model and the requirements for the output data of the first model.
10. The method according to claim 9, wherein The requirements for the input data of the first model include the input data of the first model, and the requirements for the output data of the first model include the number N of output values for each inference during the inference process of the first model, where N is a positive integer greater than or equal to 1.
11. The method according to claim 10, characterized in that, The method further includes: The first network function sends N first output values to the second device, and the second device is the execution entity of the second model; The first network function receives a fifth message from the second device, the fifth message being used to indicate M first output values, and the M first output values are the output values rejected by the second device among the N first output values, where M is a positive integer less than or equal to N.
12. The method according to any one of claims 1 to 11, characterized in that, Before the first network function receives the first message from the first device, the method further includes: The first network function receives a seventh message from the first device, and the seventh message includes a second analysis requirement for feeding back the first analysis result; When the first network function fails to meet the second analysis requirement, the first network function sends an eighth message to the first device, and the eighth message includes a third analysis requirement for feeding back the first analysis result, the third analysis requirement is lower than the second analysis requirement, and the third analysis requirement is lower than the first analysis requirement.
13. The method according to any one of claims 1 to 12, characterized in that, The first analysis requirement includes a time requirement and / or a precision requirement.
14. A communication method, characterized in that, The method includes: A second network function receives a second message from the first network function; The second network function determines a second model according to the second message, and the second model is used to assist the first model to obtain a first analysis result to meet a first analysis requirement, and the first model is deployed on the first network function.
15. The method according to claim 14, characterized in that, The second message includes first indication information, and the first indication information is used to indicate the execution entity of the second model.
16. The method according to claim 14 or 15, characterized in that, The second message includes at least one of the following: The inference speed of the second model, the analysis type of the second model.
17. The method according to claim 14 or 15, characterized in that The second message includes at least one of the following: The inference speed of the first model, the analysis type of the first model, the acceleration multiple for accelerating inference.
18. The method according to any one of claims 14 to 17, characterized in that The method further includes: The second network function sends a third message to the first network function, and the third message is used to indicate the second model.
19. The method according to claim 15 or 16, characterized in that, The method further includes: The second network function sends a sixth message to a second device, and the sixth message is used to indicate the first network function and the second model, and the second device is the execution entity of the second model.
20. The method according to any one of claims 14 to 19, characterized in that, The first analysis requirement includes a time requirement and / or a precision requirement.
21. A communication method, characterized in that, The method includes: A second device receives a sixth message from the second network function, and the sixth message is used to indicate a first network function and a second model, and the first network function deploys a first model; The second device performs inference using the second model, and the second model is used to assist the first model to obtain a first analysis result to meet a first analysis requirement.
22. The method according to claim 21, wherein The method further includes: The second device obtains the requirements for the input data of the second model and the requirements for the output data of the second model.
23. The method according to claim 22, wherein The requirements for the input data of the second model include the input data of the second model, and the requirements for the output data of the second model include the number N of output values for each inference during the inference process of the second model, where N is a positive integer greater than or equal to 1.
24. The method according to claim 23, wherein The method further includes: The second device sends N first output values to the first network function; The second device receives a fourth message from the first network function, and the fourth message is used to indicate M first output values, and the M first output values are the output values rejected by the first network function among the N first output values, where M is a positive integer less than or equal to N.
25. The method according to claim 21, characterized in that, The method further includes: The second device obtains the requirements for the input data of the second model and the requirements for the output data of the first model.
26. The method according to claim 25, wherein The requirements for the input data of the second model include the input data of the second model, and the requirements for the output data of the first model include the number N of output values for each inference during the inference process of the first model. Wherein, N is a positive integer greater than or equal to 1.
27. The method according to claim 26, wherein The method further includes: The second device receives N first output values from the first network function; The second device uses the second model to verify the N first output values and determines to reject M first output values among the N first output values; The second device sends a fifth message to the first network function, and the fifth message is used to indicate the M first output values. Wherein, M is a positive integer less than or equal to N.
28. The method according to any one of claims 21 to 27, characterized in that, The first analysis requirement includes a time requirement and / or an accuracy requirement.
29. A communication device, characterized in that, Including a processor, which is configured to, by executing a computer program or instruction, cause the communication device to execute the method according to any one of claims 1 to 13, or cause the communication device to execute the method according to any one of claims 14 to 20, or cause the communication device to execute the method according to any one of claims 21 to 28.
30. The communication device according to claim 29, wherein, The communication device further includes a memory for storing the computer program or instruction.
31. The communication device according to claim 29, characterized in that, The communication device further includes a communication interface for inputting and / or outputting signals.
32. A computer-readable storage medium, characterized in that, A computer program or instruction is stored on the computer-readable storage medium, and when the computer program or the instruction runs on a computer, it causes the method according to any one of claims 1 to 13 to be executed, or causes the method according to any one of claims 14 to 20 to be executed, or causes the method according to any one of claims 21 to 28 to be executed.
33. A computer program product, characterized in that, Containing instructions, when the instructions run on a computer, it causes the method according to any one of claims 1 to 13 to be executed, or causes the method according to any one of claims 14 to 20 to be executed, or causes the method according to any one of claims 21 to 28 to be executed.