Multi-modal analysis or prediction method, network node and storage medium

By receiving multimodal requests in network elements and using embedded models and knowledge bases to retrieve vectors to generate multimodal analysis or prediction results, the problem of insufficient accuracy in multimodal prediction in existing technologies is solved, and more efficient multimodal analysis or prediction is achieved.

CN120804395APending Publication Date: 2025-10-17ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411091026.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing network elements do not support multimodal analysis or prediction, resulting in insufficient accuracy of multimodal prediction or analysis in the core network.

Method used

By receiving multimodal analysis or prediction requests, the system converts multimodal data into vectors using an embedding model, retrieves vectors with similarity requirements from a knowledge base, and generates analysis or prediction results by combining the multimodal model. RAG technology is used to improve the accuracy of multimodal reasoning.

Benefits of technology

The accuracy and efficiency of core network multimodal analysis or prediction are improved, and network performance and user experience are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804395A_ABST
    Figure CN120804395A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal analysis or prediction method, a network node and a storage medium. The method comprises the following steps: receiving a multi-modal analysis or prediction request; converting multi-modal data into a first vector based on an embedded model according to the multi-modal analysis or prediction request; sending a retrieval request, wherein the retrieval request is used for indicating a second network element to retrieve a second vector from a knowledge base, wherein the similarity of the second vector and the first vector meets the requirement; and generating a multi-modal analysis or prediction result according to the second vector and a multi-modal model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of wireless communication, for example to a multi-modal analysis or prediction method, a network node and a storage medium. BACKGROUND

[0002] In recent years, the Artificial Intelligence (AI) architecture of the core network has emerged. With the increase in data volume and data types, the UE and the servers inside and outside the core network may need multi-modal analysis or prediction / prediction in some scenarios to enhance the performance of the network and the experience of the user. However, the current network element does not support providing multi-modal analysis or prediction. How to realize multi-modal prediction or analysis of the core network and guarantee the accuracy of the prediction or analysis is a problem to be solved. SUMMARY

[0003] The present application provides a multi-modal analysis or prediction method, a network node and a storage medium.

[0004] The present application provides a multi-modal analysis or prediction method, which is applied to a first network element and includes the following steps.

[0005] receiving a multi-modal analysis or prediction request;

[0006] According to the multi-modal analysis or prediction request, the multi-modal data is converted into a first vector based on an embedding model;

[0007] sending a retrieval request, the retrieval request being used to instruct a second network element to retrieve a second vector with a similarity meeting a requirement from a knowledge base;

[0008] generating a multi-modal analysis or prediction result according to the second vector and a multi-modal model.

[0009] The present application also provides a multi-modal analysis or prediction method, which is applied to a second network element and includes the following steps.

[0010] receiving a retrieval request of a first network element;

[0011] retrieving a second vector with a similarity meeting a requirement from a knowledge base according to the retrieval request, the first vector being converted from multi-modal data based on an embedding model

[0012] The present application also provides a multi-modal analysis or prediction method, which is applied to a user equipment and includes the following steps.

[0013] sending a multi-modal analysis or prediction request to a fourth network element through a user plane;

[0014] receiving a multi-modal analysis or prediction result sent by the fourth network element through the user plane

[0015] The embodiment of the present application further provides a multi-modal analysis or prediction method, which is applied to a fourth network element and comprises the following steps:

[0016] receiving a multi-modal analysis or prediction request of a user equipment;

[0017] sending a multi-modal analysis or prediction request to a first network element according to the multi-modal analysis or prediction of the user equipment.

[0018] The embodiment of the present application further provides a network node, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the multi-modal analysis or prediction method described above when executing the program.

[0019] The embodiment of the present application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the program is executable on a processor to implement the multi-modal analysis or prediction method described above. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 a flow chart of a multi-modal analysis or prediction method provided for an embodiment;

[0021] Figure 2 a flow chart of another multi-modal analysis or prediction method provided for an embodiment;

[0022] Figure 3 a flow chart of still another multi-modal analysis or prediction method provided for an embodiment;

[0023] Figure 4 a flow chart of yet another multi-modal analysis or prediction method provided for an embodiment;

[0024] Figure 5 a schematic diagram of registration or update of multi-modal capability information of a first network element provided for an embodiment;

[0025] Figure 6 a schematic diagram of a multi-modal analysis / prediction process provided for an embodiment;

[0026] Figure 7 a schematic diagram of registration of multi-modal capability information of a UE provided for an embodiment;

[0027] Figure 8 a schematic diagram of multi-modal analysis / prediction of a UE by a first network element provided for an embodiment;

[0028] Figure 9 a structural schematic diagram of a multi-modal analysis or prediction device provided for an embodiment;

[0029] Figure 10A structural schematic diagram of another multi-modal analysis or prediction apparatus provided for an embodiment;

[0030] Figure 11 A structural schematic diagram of still another multi-modal analysis or prediction apparatus provided for an embodiment;

[0031] Figure 12 A structural schematic diagram of yet another multi-modal analysis or prediction apparatus provided for an embodiment;

[0032] Figure 13 A hardware structural schematic diagram of a network node provided for an embodiment. DETAILED DESCRIPTION

[0033] The present application will be described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other arbitrarily without conflict. In addition, it should be noted that only parts related to the present application are shown in the drawings for convenience of description, rather than all structures.

[0034] The Network Data Analytics Function (NWDAF) is a 5G Core (5GC) Network Function (NF) located in the control plane that can perform statistical data and Machine Learning related tasks. The NWDAF can interact with different entities for different purposes, e.g. collect data from event subscriptions provided by the Access and Mobility Management Function (AMF), Session Management function (SMF), The User plane function (UPF), Policy Control function (PCF), The Unified Data Management (UDM), Network Slice Admission Control Function (NSACF), Application Function (AF), directly or through the Network Exposure Function (NEF), and Operation Administration and Maintenance (OAM); perform analytics and data collection using the Data Collection Coordination and Coordination and Delivery Function (DCCF); retrieve information from data repositories, e.g. UDR for user related information through the UDM or PFD information through the NEF (PFDF); collect location information data from LCS system; store and retrieve information from the Analytic Data Repository Function (ADRF); analyze and collect data from the Messaging Framework Adaptor Function (MFAF); retrieve information about NFs, e.g. NF related information from the NRF; provide analytics on demand to consumers, e.g. provide batch data related to an analytics ID; provide accuracy information for analytics IDs; provide Machine Learning (ML) model accuracy information or ML model accuracy degradation indication about ML models.A single instance or multiple instances of NWDAF can be deployed in a Public Land Mobile Network (PLMN). A NWDAF can contain the following logical functions:

[0035] Analytics Logical Function (AnLF): used to perform inference, derive analytics information (i.e. derive statistics and / or predictions from analytics consumer requests) and expose analytics services (i.e. Nnwdaf_AnalyticsSubscription or Nnwdaf_AnalyticsInfo);

[0036] Model Training Logical Function (MTLF): a logical function in NWDAF used to train ML models and expose new training services (e.g. provide trained ML models).

[0037] A NWDAF can contain one MTLF or one AnLF, or both.

[0038] DCCF is responsible for coordinating data collection and distribution requested by NF consumers, preventing data sources from handling multiple subscriptions to the same data, and preventing multiple notifications containing the same information from being sent due to uncoordinated requests from data consumers.

[0039] DCCF is applicable to: NWDAF requesting data from data sources (e.g. for computing analytics); NF consumer requesting analytics from NWDAF data sources; NF consumer requesting data from ADRF data sources; ADRF receiving data from NF data sources.

[0040] The 5G system architecture supports ADRF to store collected data and analytics. ADRF exposes Nadrf service for information storage and retrieval of data by other 5GC network functions (e.g. NWDAF).

[0041] Based on a request from a network function or configuration on DCCF, DCCF can determine ADRF and interact with it directly or indirectly to request or store data. The way of interaction can be:

[0042] Direct: DCCF requests data storage in ADRF through Nadrf service or Ndccf_DataManagement_Notify (e.g. when ADRF requests DCCF to notify data collection). In addition, DCCF retrieves data from ADRF through Nadrf service.

[0043] Indirect: The DCCF requires the messaging framework to store data in the ADRF via the Nadrf service or Nmfaf_3daDataManagement_Configure. The messaging framework can contain one or more adapters that convert between protocols defined by 3GPP.

[0044] The consumer network function can specify in the request sent to the DCCF that data provided by the data source needs to be stored in the ADRF.

[0045] The ADRF stores data received in the Nadrf_DataManagement_StorageRequest sent directly from the network function or in the Ndccf_DataManagement_Notify / Nmfaf_3caDataManagement_Notify or Nnwdaf_DataManagement_Notify sent from the DCCF, MFAF or NWDAF.

[0046] RAG (Retrieval-Augmented Generation) is an emerging method in natural language processing that combines the strengths of both retrieval and generation techniques. Specifically, RAG models rely on both the internal knowledge of pre-trained language models and the retrieval of relevant documents from external databases to enhance the accuracy and diversity of generated results when generating text.

[0047] The workflow of RAG mainly includes:

[0048] Retrieval phase: The model retrieves the most relevant document fragments from the external database based on the given input query. This step usually uses information retrieval techniques such as BM25 or vector retrieval.

[0049] Generation phase: Then, the input and the retrieved document fragments are input into the generation model. The generation model combines the input and the retrieved document fragments to generate more rich and accurate output.

[0050] When using large language models for inference, the parameter configurations used can significantly affect the quality and characteristics of the output content. The following are some common inference parameters and their roles: temperature, maximum length, prompt, short burst length, repetition penalty, Top-k sampling, Top-p (or Nucleus) sampling, batch size, and softness. These parameters can be used individually or in combination to improve the quality of the output analysis / prediction content. Depending on the specific application scenario, adjusting these parameters can significantly improve the output effect.

[0051] Figure 1 A flowchart of a multi-modal analysis or prediction method provided for an embodiment, which can be applied to a first network element, which can be a network element for multi-modal analysis or prediction, such as a NWDAF. As shown in Figure 1 The method provided by the embodiment includes the following steps:

[0052] Step 110, receiving a multi-modal analysis or prediction request.

[0053] Step 120, converting multi-modal data into a first vector based on an embedding model according to the multi-modal analysis or prediction request.

[0054] Step 130, sending a retrieval request, the retrieval request being used to instruct a second network element to retrieve a second vector with a similarity satisfying a requirement from a knowledge base.

[0055] Step 140, generating a multi-modal analysis or prediction result according to the second vector and a multi-modal model.

[0056] The embedding model can also be referred to as a vector conversion model, which is used to convert multi-modal data into a vector. In the construction process of the knowledge base in the second network element, the embedding model can be used to convert multi-modal data into a vector and store it in the knowledge base of the second network element. When a multi-modal analysis or prediction request is made, the first network element can use the same embedding model to convert multi-modal data into a first vector, then request the second network element to retrieve a second vector with a similarity satisfying a requirement (such as a similarity greater than a threshold or a highest similarity) from the stored vectors in the knowledge base, and then the first network element uses a multi-modal model to generate a multi-modal analysis or prediction result according to the retrieved second vector. On this basis, multi-modal analysis or prediction is performed based on RAG, improving the quality and accuracy of core network multi-modal inference.

[0057] In an embodiment, the multi-modal analysis or prediction request comprises at least one of the following: a knowledge classification identifier; a multi-modal analysis or prediction indication; and at least one parameter in the multi-modal analysis or prediction indication.

[0058] a temperature parameter for controlling randomness of the generated text, usually with a value between 0.0 and 1.0, a lower temperature (e.g. 0.2) will make the model output more deterministic content, a higher temperature (e.g. 0.8) will make the output more creative and diverse;

[0059] an output maximum length for limiting the maximum length of the generated text, the value depends on the specific model, usually from tens to thousands of tokens, setting a higher maximum length can avoid generating too short text, but may also cause the generated content to be lengthy and irrelevant;

[0060] a parameter amount of the model;

[0061] a prompt word, such as a prefix (Prompt), which can be used to provide context or a starting portion of the generated text, a suitable prefix can guide the model to generate content that better meets the user's expectations; and a prefix length (Short Burst Length), which can be the number of tokens generated each time, short generation can better control the generation process, but may reduce efficiency;

[0062] a repetition penalty for penalizing repeated content, usually with a value between 1.0 and 2.0, a higher repetition penalty value can reduce the repetition in the generated content;

[0063] sampling-related parameters, such as a Top-k sampling parameter for limiting the number of candidate tokens at each generation, with a positive integer value, by setting a k value, the model only selects from the top k tokens with the highest probability; and a Top-p (or Nucleus) sampling parameter for limiting the total probability mass of the generated tokens, with a value between 0.0 and 1.0, by selecting a probability threshold p, the model only selects from tokens with a cumulative probability of p; and a batch size, which can be the number of samples generated at a time, a larger batch size can improve generation efficiency, but requires more computing resources;

[0064] a softness for adjusting the diversity of the generated text, similar to temperature, but more focused on adjusting the smoothness of the generation distribution.

[0065] In an embodiment, the method further comprises:

[0066] sending a multi-modal model acquisition request to the target network element in the case of no multi-modal model locally;

[0067] receiving a link or a file of the multi-modal model sent by the target network element.

[0068] In an embodiment, the method further comprises: receiving model capability information of the multi-modal model sent by the target network element.

[0069] In an embodiment, the request for obtaining the multi-modal model comprises at least one of: a knowledge classification identifier; model capability information; the model capability information comprises at least one of: supported data types, whether to support RAG technology, whether to belong to a large model; the supported data types comprise at least one of: text, audio, video, and picture.

[0070] In an embodiment, the method further comprises:

[0071] sending a discovery request containing a network element of the multi-modal model to a third network element (such as NRF) and determining a target network element.

[0072] In an embodiment, the retrieval request comprises at least one of: a knowledge base identifier; a knowledge classification identifier.

[0073] In an embodiment, each copy further comprises at least one of: sending a multi-modal analysis or prediction result to a consumer;

[0074] sending a multi-modal analysis or prediction result to a fourth network element (such as AF), and sending the multi-modal analysis or prediction result to a user equipment after processing by the fourth network element according to a capability of the user equipment.

[0075] In an embodiment, the method further comprises: registering multi-modal capability information, the multi-modal capability information comprising at least one of: whether to have a multi-modal analysis or prediction capability; whether to have a multi-modal model; whether to have a multi-modal model supporting RAG technology; whether to have an embedding model; and whether to have a vector conversion capability.

[0076] In an embodiment, an embedding model for converting the multi-modal data into a first vector is consistent with an embedding model used for constructing a knowledge base of the second network element.

[0077] In an embodiment, the method further comprises:

[0078] sending a knowledge classification identifier to the second network element for requesting to obtain embedding model information used for constructing a knowledge base of the second network element;

[0079] receiving a knowledge base identifier of the second network element and embedding model information of an embedding model used for constructing a knowledge base.

[0080] Figure 2A flowchart of a multi-modal analysis or prediction method provided by an embodiment, which can be applied to a second network element, which can be a network element storing and retrieving vectors. As shown in Figure 2 the method provided by the embodiment includes the following steps:

[0081] Step 210, receiving a retrieval request of a first network element.

[0082] Step 220, retrieving a second vector from a knowledge base according to the retrieval request, the second vector meeting a requirement in terms of similarity with the first vector, the first vector being converted from multi-modal data based on an embedding model.

[0083] In an embodiment, the embedding model used to convert the multi-modal data into the first vector is consistent with an embedding model used to build the knowledge base of the second network element.

[0084] In an embodiment, the method further includes:

[0085] receiving a knowledge classification identifier;

[0086] locating the knowledge base and an embedding model used to build the knowledge base according to the knowledge classification identifier;

[0087] sending a knowledge base identifier and embedding model information of the embedding model used to build the knowledge base.

[0088] Figure 3 A flowchart of a multi-modal analysis or prediction method provided by an embodiment, which can be applied to a user equipment (UE). As shown in Figure 3 the method provided by the embodiment includes the following steps:

[0089] Step 310, sending a multi-modal analysis or prediction request to a fourth network element through a user plane.

[0090] Step 320, receiving a multi-modal analysis or prediction result sent by the fourth network element through the user plane.

[0091] In an embodiment, the multi-modal analysis or prediction request includes at least one of the following: a knowledge classification identifier; a multi-modal data type; a device identifier.

[0092] In an embodiment, the method further includes:

[0093] registering multi-modal capability information, the multi-modal capability information including a data type supported by the UE in terms of receiving and / or parsing;

[0094] the data type includes at least one of the following: text, audio, video, and picture.

[0095] In an embodiment, the multi-modal capability information is transparently transmitted by the RAN to the AMF, and stored in the AMF, UDM or UDR.

[0096] Figure 4 An embodiment provides a flowchart of a multi-modal analysis or prediction method, which can be applied to a fourth network element, which can be an AF. As shown in Figure 4 The method provided by the embodiment includes the following steps:

[0097] Step 410, receiving a multi-modal analysis or prediction request of a user equipment.

[0098] Step 420, sending a multi-modal analysis or prediction request to a first network element according to the multi-modal analysis or prediction of the user equipment.

[0099] In an embodiment, the method further includes:

[0100] receiving a multi-modal analysis or prediction result of the first network element;

[0101] processing the multi-modal analysis or prediction result according to the capability of the user equipment, and sending the processed multi-modal analysis or prediction result to the user equipment.

[0102] In an embodiment, the method further includes: sending a query request to the AMF, UDM or UDR, the query request carrying a device identifier, and the query request being used to query the capability of the user equipment.

[0103] In the embodiments of the present application, the first network element (such as NWDAF) can receive an analysis / prediction request of a consumer, provide multi-modal analysis / prediction to the consumer network element by using a multi-modal model (which can be a large model) and RAG technology, and return the generated analysis / prediction result. In addition, after receiving a query request of a fourth network element (such as AF), the UE related capability can be queried, and the analysis / prediction result can be generated according to the capability and sent to the AF. The models supporting multi-modal analysis / prediction can be transmitted between different first network elements (such as NWDAF containing MTLF and NWDAF containing AnLF). The first network element can also register its multi-modal analysis / prediction capability and vector conversion capability to the third network element (such as NRF).

[0104] The third network element can receive the analysis / prediction capability and vector conversion capability sent by the first network element, and also support discovering a target network element (NWDAF containing AnLF) for vector conversion, and also support discovering ADRF for vector storage and knowledge base construction.

[0105] The second network element (which can be ADRF or a new NF) is used to construct a knowledge base, store vectors, support retrieval of the knowledge base, and extract vectors with high similarity.

[0106] The consumer NF (such as an OTT server, 5GC network element and / or OAM, etc.) can send a multimodal analysis / prediction request to the first network element and receive the analysis / prediction results.

[0107] The user equipment can report its multimodal capabilities in the registration message, send multimodal analysis / prediction requests to the AF through the user plane, and receive analysis / prediction sent by the AF through the user plane.

[0108] The fourth network element can receive the UE's multimodal analysis / prediction request and send the multimodal analysis / prediction request to the NWDAF; it can also receive the analysis / prediction results returned by the NWDAF, and determine the analysis / prediction results finally sent to the UE by querying the UE-related capability information.

[0109] The multimodal analysis or prediction method of the present application is exemplified below through some embodiments.

[0110] Example 1

[0111] In this embodiment, the first network element (such as NWDAF) may register or update the multimodal capability to the third network element (such as NRF).

[0112] like Figure 5 As shown, registration or update mainly includes:

[0113] 1. NWDAF registers or updates its multimodal capabilities with NRF. The multimodal capabilities may include at least one of the following: whether NWDAF supports multimodal analysis / prediction, whether a local multimodal model exists, whether a multimodal model that supports RAG exists, whether an embedded model exists, data types that support multimodal analysis / prediction (such as text, video, audio, and / or images), and whether the capability to convert multimodal data to vectors exists, which may include embedded model information (such as the type, identifier, and / or manufacturer of the embedded model). The aforementioned multimodal capabilities may be sent via the NF profile.

[0114] 2.NRF stores NWDAF-related NF profiles.

[0115] 3. NRF returns the registration / update result indication, success or failure.

[0116] Example 2

[0117] In this embodiment, the first network element (such as NWDAF including AnLF) can request the target network element (such as NWDAF including MTLF) to obtain the multimodal model, thereby providing multimodal analysis or prediction for the internal network elements of 5GC. Figure 6 As shown in Figure 2, the multimodal analysis / prediction process mainly includes:

[0118] 1a. The consumer (i.e. a 5GC network element, an OAM or an OTT server, etc.) sends a discovery request to a third network element (e.g. NRF) to discover a first network element NWDAF that can provide multi-modal analysis / prediction. The request can contain data types supported by multi-modal analysis / prediction, such as text, audio, video and / or pictures, etc.

[0119] 1b. The NRF matches the appropriate NWDAF according to the pre-stored registration or update information of the NWDAF, and returns a list of NWDAFs that meet the requirements.

[0120] 2. The consumer determines the first network element (i.e. the NWDAF containing AnLF in the figure) and sends a multi-modal analysis / prediction request to it, which can contain knowledge classification ID, data types contained in the required output analysis / prediction (such as text, audio, video and / or pictures), an indication that a large model is used to provide multi-modal analysis / prediction, which can contain parameters of the large model, such as temperature parameter, maximum length of output, parameter quantity of the large model, prompt word of the large model, repetition penalty, sampling related parameters and / or softness, etc.

[0121] 3a. If the NWDAF containing AnLF does not have a model locally to support multi-modal analysis / prediction, it can send a multi-modal model acquisition request to the NWDAF containing MTLF, which can contain capability information of the multi-modal model (such as supported data types, whether RAG is supported, whether it belongs to a large model, etc.), and can also contain knowledge classification ID, etc.

[0122] The NWDAF containing AnLF can send a discovery request to the NRF, which contains multi-modal model related capability parameters (such as data types that can be provided by the multi-modal model, such as text, audio, video and / or pictures, etc.).

[0123] 3b. The NWDAF containing MTLF returns a link containing the multi-modal model or a file of the multi-modal model, which contains capability information of the multi-modal model (such as supported data types, whether RAG is supported, whether it belongs to a large model, etc.).

[0124] 4. The NWDAF decides that RAG technology needs to be used to generate analysis / prediction results. The NWDAF sends the knowledge classification ID provided by the consumer to the second network element (ADRF) to request to obtain embedding model information used for knowledge base construction.

[0125] 5. The ADRF uses the knowledge classification ID to locate the relevant knowledge base and embedding model information used for knowledge base construction.

[0126] 6. The ADRF sends the knowledge base ID associated with the knowledge classification ID and / or embedding model information (e.g., model ID and / or model type) to the NWDAF.

[0127] 7. Using the embedding model required by the ADRF in step 6, the relevant questions and / or requirements (text, audio, video, and / or pictures) of the multi-modal data received from the Consumer in step 2 are converted into vectors (i.e., first vectors). It should be noted that if the NWDAF does not support vector conversion or is insufficient in computing power, another NWDAF can be requested to perform vector conversion, and the converted vectors are returned.

[0128] It should be noted that during the process of building the knowledge base, the vectors stored by the ADRF are converted from multi-modal data by the embedding model. In order to ensure the accuracy and reliability of the output analysis / prediction results, the embedding model used in step 7 should be consistent with the embedding model used when building the knowledge base.

[0129] 8. The NWDAF initiates a retrieval process, and the knowledge base ID and knowledge classification ID can be carried in the retrieval request, the purpose of which is to extract a second vector with high similarity to the first vector in step 7 through the ADR.

[0130] 9. The NWDAF collects inference data from other network elements and OAM, performs inference using a local large model, and generates an analysis result in combination with the second vector obtained from the knowledge base in step 8.

[0131] 10. The NWDAF returns the analysis / prediction result to the Consumer, and the result is a multi-modal output, which can include text, audio, video, and / or pictures, etc.

[0132] Embodiment 3

[0133] In this embodiment, a user equipment (UE) can register multi-modal capability information. As shown in Figure 7 , the registration process mainly includes:

[0134] 1. The UE carries multi-modal capability information in the registration message, which includes the data types supported for receiving and / or parsing, such as text, audio, video, and / or pictures.

[0135] 2. The RAN transmits the multi-modal capability information to the AMF;

[0136] 3. The AMF stores the multi-modal capability information of the UE.

[0137] 4. If the AMF does not store the multi-modal capability of the UE, the multi-modal capability information can also be transmitted to the UDM, and the UDM can store the multi-modal capability information to the UDR (background of the UDM).

[0138] 5. The UDM / UDR stores the UE multimodal capability.

[0139] 6. The UDM returns a storage indication.

[0140] 7. The AMF returns a registration complete indication to the RAN

[0141] 8. The RAN returns a registration complete indication to the UE.

[0142] Embodiment 4

[0143] In this embodiment, the first network element (NWDAF) can provide multimodal analysis / prediction for a user equipment (UE). As shown in FIG. 1, the multimodal analysis / prediction procedure mainly includes: Figure 8

[0144] 0. The UE sends a multimodal analysis / prediction request to the AF through a user plane, which can contain a knowledge classification ID. This request can contain related questions and / or requirements of multimodal data types (text, audio, video, and / or pictures), and can also contain a UE ID.

[0145] 1a. The AF (OTT server) decides that the analysis needs to be assisted by the NWDAF, sends a discovery request to the NRF to discover the NWDAF that can provide multimodal analysis / prediction.

[0146] 1b. The NRF matches the appropriate NWDAF according to the pre-stored NWDAF information, and returns a list of NWDAFs that meet the requirements.

[0147] 2. The AF decides that the NWDAF is needed for multimodal analysis / prediction, determines the first network element (i.e., the NWDAF in the figure) and sends a multimodal analysis / prediction request to the NWDAF, which can contain a knowledge classification ID, a UE ID, a requirement of data types (such as text, audio, video, and / or pictures) contained in the output analysis / prediction, an indication of application of a large model to provide multimodal analysis / prediction, which can contain parameters of the large model, such as temperature parameters, maximum length of output, parameter quantity of the large model, prompt words of the large model, repetition penalty, sampling related parameters, and / or gentleness, etc. This request can contain related questions and / or requirements of multimodal data types (text, audio, video, and / or pictures), which can be consistent with those sent by the UE in step 0, or can be optimized and then sent to the NWDAF.

[0148] Before step 3, the NWDAF can perform the steps 3a / 3b of the above embodiment, i.e., obtain a model for multimodal analysis / prediction from another NWDAF.

[0149] ​3. The NWDAF decision needs to use the RAG technology for multi-modal analysis / prediction. The NWDAF sends the knowledge classification ID provided by the Consumer to the ADRF, requesting to obtain the embedding model used for knowledge base construction.

[0150] 4. The ADRF uses the knowledge classification ID to locate the relevant knowledge base and embedding model information used during knowledge base construction.

[0151] 5. The ADRF sends the knowledge base ID and / or embedding model information (such as model ID and / or model type) associated with the knowledge classification ID to the NWDAF.

[0152] 6. Use the embedding model required by the ADRF in step 5 to convert the relevant questions and / or requirements (text, audio, video, and / or pictures) of the multi-modal data received from the Consumer in step 2 into vectors (i.e., first vectors). If the NWDAF does not support vector conversion locally or has insufficient computing power, it can also request another NWDAF to perform vector conversion and return the converted vectors.

[0153] 7. The NWDAF initiates a retrieval process, and the retrieval request can carry the knowledge base ID and knowledge classification ID, thereby extracting vectors with high similarity to the first vectors in step 6.

[0154] 8. The NWDAF collects inference data from other network elements, OAM, and uses the local large model for inference, combined with the second vectors obtained from the knowledge base in step 7, to generate analysis / prediction results. In addition, before generating the analysis / prediction results, the NWDAF can send a query request to the AMF / UDM / UDR, carrying the UE ID, to query the relevant capabilities of the UE.

[0155] 9. The NWDAF returns the multi-modal analysis / prediction results to the AF.

[0156] 10. After receiving the multi-modal analysis / prediction results returned by the NWDAF, the AF can perform secondary processing. During secondary processing, the AF can send a query request to the AMF / UDM / UDR, carrying the UE ID, to query the relevant capabilities of the UE, based on which secondary processing is performed, for example, if the multi-modal analysis / prediction results returned by the NWDAF contain multi-modal data that the UE does not support, the AF can perform deletion.

[0157] 11. The AF returns the secondary-processed multi-modal analysis / prediction results to the UE through the user plane.

[0158] The embodiments of the present application also provide a multi-modal analysis or prediction device. Figure 9 A structural schematic diagram of a multi-modal analysis or prediction device provided by an embodiment is shown in FIG. 1. Figure 9As shown, the multi-modal analysis or prediction device comprises:

[0159] The request receiving module 510 is configured to receive a multi-modal analysis or prediction request.

[0160] The conversion module 520 is configured to convert multi-modal data into a first vector based on an embedding model according to the multi-modal analysis or prediction request.

[0161] The request retrieving module 530 is configured to send a retrieving request, the retrieving request being used to instruct a second network element to retrieve a second vector with a similarity satisfying a requirement from a knowledge base.

[0162] The analysis or prediction module 540 is configured to generate a multi-modal analysis or prediction result according to the second vector and a multi-modal model.

[0163] In an embodiment, the multi-modal analysis or prediction request comprises at least one of the following: a knowledge classification identifier; a multi-modal analysis or prediction indication.

[0164] The multi-modal analysis or prediction indication comprises at least one of the following parameters: a temperature parameter, an output maximum length, a parameter quantity, a prompt word, a repetition penalty, a sampling related parameter, and a softness.

[0165] In an embodiment, the device further comprises a model obtaining module configured to, in a case where there is no multi-modal model locally, send a multi-modal model obtaining request to a target network element; and receive a link or a file of the multi-modal model sent by the target network element.

[0166] In an embodiment, the device further comprises an information receiving module configured to receive model capability information of the multi-modal model sent by the target network element.

[0167] In an embodiment, the multi-modal model obtaining request comprises at least one of the following: a knowledge classification identifier; model capability information; the model capability information comprises at least one of the following: a supported data type, whether to support RAG technology, and whether to belong to a large model; and the supported data type comprises at least one of the following: text, audio, video, and picture.

[0168] In an embodiment, the device further comprises a discovery module configured to send a discovery request containing a network element of a multi-modal model to a third network element and determine a target network element.

[0169] In an embodiment, the retrieving request comprises at least one of the following: a knowledge base identifier; and a knowledge classification identifier.

[0170] In an embodiment, the device further comprises a result sending module configured to perform one of the following:

[0171] sending the multi-modal analysis or prediction result to the consumer;

[0172] sending the multi-modal analysis or prediction result to a fourth network element, the multi-modal analysis or prediction result being sent to the user equipment by the fourth network element after being processed according to the capability of the user equipment.

[0173] In an embodiment, the apparatus further comprises:

[0174] a registration module configured to register multi-modal capability information, the multi-modal capability information comprising at least one of:

[0175] whether having the capability of multi-modal analysis or prediction; whether having a multi-modal model; whether having a multi-modal model supporting RAG technology; whether having an embedding model; whether having a vector conversion capability.

[0176] In an embodiment, the embedding model for converting the multi-modal data into the first vector is consistent with the embedding model used by the second network element for building the knowledge base.

[0177] In an embodiment, the apparatus further comprises: an information request module configured to send a knowledge classification identifier to the second network element for requesting to obtain embedding model information used by the second network element for building the knowledge base; and receive the knowledge base identifier of the second network element and the embedding model information of the embedding model used by the second network element for building the knowledge base.

[0178] The multi-modal analysis or prediction apparatus proposed in the embodiment belongs to the same inventive concept as the multi-modal analysis or prediction method proposed in the above embodiments, and the technical details not described in the embodiment can be referred to any of the above embodiments, and the embodiment has the same beneficial effects as the multi-modal analysis or prediction method.

[0179] The embodiment of the present application further provides a multi-modal analysis or prediction apparatus. Figure 10 A structural schematic diagram of a multi-modal analysis or prediction apparatus provided by an embodiment is shown in FIG. 1. Figure 10 As shown in FIG. 1, the multi-modal analysis or prediction apparatus comprises:

[0180] a request receiving module 610 configured to receive a search request of a first network element;

[0181] a search module 620 configured to search a second vector from a knowledge base according to the search request, the second vector satisfying a requirement in terms of similarity with the first vector, the first vector being obtained by converting multi-modal data based on an embedding model.

[0182] In an embodiment, the embedding model for converting the multi-modal data into the first vector is consistent with the embedding model used by the second network element for building the knowledge base.

[0183] In an embodiment, the apparatus further comprises:

[0184] The receiving module is configured to receive the knowledge classification identifier.

[0185] The positioning module is configured to position the knowledge base and an embedding model used for constructing the knowledge base according to the knowledge classification identifier.

[0186] The sending module is configured to send a knowledge base identifier and embedding model information of the embedding model used for constructing the knowledge base.

[0187] The multi-modal analysis or prediction device provided in the embodiment belongs to the same inventive concept as the multi-modal analysis or prediction method provided in the above embodiments, and technical details not described in detail in the embodiment can be referred to any of the above embodiments, and the embodiment has the same beneficial effects as performing the multi-modal analysis or prediction method.

[0188] The embodiment of the present application also provides a multi-modal analysis or prediction device. Figure 11 A structural schematic diagram of a multi-modal analysis or prediction device provided by an embodiment is shown in FIG. 7. Figure 11 As shown in FIG. 7, the multi-modal analysis or prediction device includes:

[0189] The request module 710 is configured to send a multi-modal analysis or prediction request to a fourth network element through a user plane.

[0190] The receiving module 720 is configured to receive a multi-modal analysis or prediction result sent by the fourth network element through a user plane.

[0191] In an embodiment, the multi-modal analysis or prediction request includes at least one of the following: a knowledge classification identifier; a multi-modal data type; and a device identifier.

[0192] In an embodiment, the device further includes:

[0193] The registration module is configured to register multi-modal capability information, and the multi-modal capability information includes a data type supported by the device for receiving and / or analyzing; and the data type includes at least one of the following: text, audio, video, and picture.

[0194] In an embodiment, the multi-modal capability information is transparently transmitted by the RAN to the AMF, and stored in the AMF, UDM, or UDR.

[0195] The multi-modal analysis or prediction device provided in the embodiment belongs to the same inventive concept as the multi-modal analysis or prediction method provided in the above embodiments, and technical details not described in detail in the embodiment can be referred to any of the above embodiments, and the embodiment has the same beneficial effects as performing the multi-modal analysis or prediction method.

[0196] The embodiment of the present application also provides a multi-modal analysis or prediction device.Figure 12 A structural diagram of a multi-modal analysis or prediction device is provided in an embodiment. As shown in the figure, the multi-modal analysis or prediction device includes: Figure 12

[0197] The receiving module 810 is configured to receive a multi-modal analysis or prediction request of a user equipment;

[0198] The sending module 820 is configured to send a multi-modal analysis or prediction request to a first network element according to the multi-modal analysis or prediction of the user equipment.

[0199] In an embodiment, the device further includes:

[0200] The result receiving module is configured to receive a multi-modal analysis or prediction result of the first network element;

[0201] The result processing module is configured to process the multi-modal analysis or prediction result according to the capability of the user equipment, and send the processed multi-modal analysis or prediction result to the user equipment.

[0202] In an embodiment, the device further includes:

[0203] The query module is configured to send a query request to an AMF, a UDM or a UDR, the query request carrying a device identifier, and the query request being used to query the capability of the user equipment.

[0204] The multi-modal analysis or prediction device provided in the embodiment belongs to the same inventive concept as the multi-modal analysis or prediction method provided in the above embodiments, and the technical details not described in the embodiment can be referred to the above embodiments, and the embodiment has the same beneficial effects as the multi-modal analysis or prediction method.

[0205] The embodiment of the present application further provides a network node, Figure 13 A hardware structure diagram of a network node is provided in an embodiment, as shown in the figure, the network node provided by the present application includes a processor 910 and a memory 920; the processor 910 in the network node can be one or more, Figure 13 The processor 910 in the network node can be one or more, Figure 13 The memory 920 is configured to store one or more programs; the one or more programs are executed by the one or more processors 910, so that the one or more processors 910 implement the multi-modal analysis or prediction method as described in the embodiments of the present application.

[0206] The network node further includes a communication device 930, an input device 940 and an output device 950.

[0207] ​The processor 910, the memory 920, the communication device 930, the input device 940 and the output device 950 in the network node can be connected by a bus or other means, Figure 13 The connection by the bus is taken as an example.

[0208] The input device 940 can be used to receive input digital or character information, and generate key signal input related to user settings and function control of the network node. The output device 950 can include a display device such as a display screen.

[0209] The communication device 930 can include a receiver and a transmitter. The communication device 930 is configured to perform information receiving and transmitting communication according to the control of the processor 910.

[0210] The memory 920 as a kind of computer readable storage medium can be configured to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the multi-modal analysis or prediction method (for example, modules in the multi-modal analysis or prediction device) described in the embodiments of the present application. The memory 920 can include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required by at least one function; the data storage area can store data created according to the use of the network node and the like. In addition, the memory 920 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state memory device. In some examples, the memory 920 can further include a memory remotely arranged relative to the processor 910, which can be connected to the network node through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0211] The embodiment of the present application further provides a storage medium storing a computer program, and the computer program is executed by a processor to implement the multi-modal analysis or prediction method in any of the embodiments of the present application. The method is applied to a first network element, and includes: receiving a multi-modal analysis or prediction request; converting multi-modal data into a first vector based on an embedding model according to the multi-modal analysis or prediction request; sending a retrieval request, the retrieval request being used to instruct a second network element to retrieve a second vector with a similarity to the first vector meeting a requirement from a knowledge base; and generating a multi-modal analysis or prediction result according to the second vector and a multi-modal model. Alternatively, the method is applied to the second network element, and includes: receiving a retrieval request of the first network element; retrieving a second vector with a similarity to the first vector meeting a requirement from the knowledge base according to the retrieval request, the first vector being converted based on an embedding model according to multi-modal data. Alternatively, the method is applied to a user equipment, and includes: sending a multi-modal analysis or prediction request to a fourth network element through a user plane; and receiving a multi-modal analysis or prediction result sent by the fourth network element through the user plane. Alternatively, the method is applied to the fourth network element, and includes: receiving a multi-modal analysis or prediction request of the user equipment; and sending a multi-modal analysis or prediction request to the first network element according to the multi-modal analysis or prediction of the user equipment.

[0212] The embodiment of the present application further provides a storage medium storing a computer program, and the computer program is executed by a processor to implement the multi-modal analysis or prediction method in any of the embodiments of the present application. The method is applied to a first network element, and includes: receiving a multi-modal analysis or prediction request; converting multi-modal data into a first vector based on an embedding model according to the multi-modal analysis or prediction request; sending a retrieval request, the retrieval request being used to instruct a second network element to retrieve a second vector with a similarity to the first vector meeting a requirement from a knowledge base; and generating a multi-modal analysis or prediction result according to the second vector and a multi-modal model. Alternatively, the method is applied to the second network element, and includes: receiving a retrieval request of the first network element; retrieving a second vector with a similarity to the first vector meeting a requirement from the knowledge base according to the retrieval request, the first vector being converted based on an embedding model according to multi-modal data. Alternatively, the method is applied to a user equipment, and includes: sending a multi-modal analysis or prediction request to a fourth network element through a user plane; and receiving a multi-modal analysis or prediction result sent by the fourth network element through the user plane. Alternatively, the method is applied to the fourth network element, and includes: receiving a multi-modal analysis or prediction request of the user equipment; and sending a multi-modal analysis or prediction request to the first network element according to the multi-modal analysis or prediction of the user equipment.

[0213] The computer storage medium of the embodiments of the present application can adopt any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may, for example, but is not limited to, be: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples (non-exhaustive list) of the computer readable storage medium include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM), a flash memory, an optical fiber, a portable CD-ROM, an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, device or apparatus.

[0214] The computer readable signal medium can include a data signal propagating in baseband or propagated as a carrier wave in a propagation medium, in which computer readable program code is embodied. Such propagated data signal can take a variety of forms including, but not limited to, electro-magnetic, optical or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport program for use by or in connection with an instruction execution system, apparatus, or device.

[0215] The program code embodied on the computer readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire line, optical fiber cable, Radio Frequency (RF), etc., or any suitable combination thereof.

[0216] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0217] The embodiments of the present application further provide a computer program product, comprising computer programs / instructions, which, when executed by a processor, implement the video encoding method according to any of the above embodiments.

[0218] The above merely provides example embodiments of the present application, but is not intended to limit the protection scope of the present application.

[0219] Those skilled in the art will appreciate that the term user terminal encompasses any appropriate type of wireless user equipment, such as a mobile phone, a portable data processing portable network browser or a vehicle mounted mobile station.

[0220] Generally, the various embodiments of the present application can be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects can be implemented in hardware, while other aspects can be implemented in

[0221] Embodiments of the present application can be implemented by a data processor of a mobile device executing computer program instructions, for example in a processor entity, or by hardware, or by a combination of software and hardware. Computer program instructions can be in assemblies, Instruction Set Architecture (ISA), machine, machine-related, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages.

[0222] The block diagrams of any logical flow of the present application in the accompanying drawings can represent program steps or can represent interconnected logic circuits, modules, and functions, or can represent a combination of program steps and logic circuits, modules, and functions. The computer program can be stored on a memory. The memory can have any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as, but not limited to, a Read-Only Memory (ROM), a Random Access Memory (RAM), an optical storage device, and a system (a Digital Video Disc (DVD) or a Compact Disk (CD), etc.). The computer readable medium can include a non-transitory storage medium. The data processor can be of any type suitable for the local technical environment, and can include, but is not limited to, a general purpose computer, a special purpose computer, a microprocessor, a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FGPA), and a processor based on multi-core processor architecture.

[0223] A detailed description of exemplary embodiments of the present application has been provided above with reference to the accompanying drawings. However, various modifications and alterations of the above embodiments will be apparent to those skilled in the art without departing from the scope of the present application, in view of this disclosure. Thus, the proper scope of the present application will be determined by the following claims.

Claims

1. A multimodal analysis or prediction method, applied to a first network element, characterized in that: include: receiving multimodal analysis or prediction requests; According to the multimodal analysis or prediction request, convert the multimodal data into a first vector based on the embedding model; Sending a search request, where the search request is used to instruct the second network element to retrieve a second vector from a knowledge base that satisfies a requirement of similarity with the first vector; A multimodal analysis or prediction result is generated according to the second vector and the multimodal model.

2. The method according to claim 1, characterized in that The multimodal analysis or prediction request includes at least one of the following: a knowledge classification identifier; a multimodal analysis or prediction instruction; The multimodal analysis or prediction indication includes at least one of the following parameters: temperature parameter, maximum output length, parameter quantity, prompt word, repetition penalty, sampling related parameters, and mildness.

3. The method according to claim 1, characterized in that Also includes: If there is no multimodal model locally, a request for obtaining the multimodal model is sent to the target network element; Receive a link or file of the multimodal model sent by the target network element.

4. The method according to claim 3, characterized in that Also includes: Receive model capability information of the multimodal model sent by the target network element.

5. The method according to claim 3, characterized in that The request for obtaining the multimodal model includes at least one of the following: a knowledge classification identifier; model capability information; The model capability information includes at least one of the following: supported data types, whether RAG technology is supported, and whether it belongs to a large model; The supported data types include at least one of the following: text, audio, video, and picture.

6. The method according to claim 3, characterized in that Also includes: A discovery request for a network element including a multimodal model is sent to a third network element and a target network element is determined.

7. The method according to claim 1, characterized in that The search request includes at least one of the following: Knowledge base identifier; knowledge classification identifier.

8. The method according to claim 1, characterized in that Also include one of the following: Send multimodal analysis or prediction results to consumers; The multimodal analysis or prediction result is sent to the fourth network element, where the multimodal analysis or prediction result is processed by the fourth network element according to the capability of the user equipment and then sent to the user equipment.

9. The method according to claim 1, characterized in that Also includes: Register multimodal capability information, where the multimodal capability information includes at least one of the following: Whether it has the ability to perform multimodal analysis or prediction; whether it has a multimodal model; whether it has a multimodal model that supports RAG technology; whether it has an embedded model; whether it has vector conversion capabilities.

10. The method according to claim 1, characterized in that The embedding model for converting the multimodal data into the first vector is consistent with the embedding model used to construct the knowledge base of the second network element.

11. The method according to claim 10, characterized in that Also includes: Sending a knowledge classification identifier to the second network element to request the acquisition of embedded model information used for constructing a knowledge base of the second network element; Receive the knowledge base identifier of the second network element and the embedding model information of the embedding model used for constructing the knowledge base.

12. A multimodal analysis or prediction method, applied to a second network element, characterized in that: include: receiving a search request from a first network element; According to the search request, a second vector that meets the similarity requirement with the first vector is retrieved from the knowledge base, where the first vector is obtained by converting the multimodal data based on the embedding model.

13. The method according to claim 12, characterized in that The embedding model for converting the multimodal data into the first vector is consistent with the embedding model used to construct the knowledge base of the second network element.

14. The method according to claim 12, characterized in that Also includes: Receive knowledge classification identification; Locating a knowledge base and an embedding model used to construct the knowledge base according to the knowledge classification identifier; Send the knowledge base identifier and the embedding model information of the embedding model used to build the knowledge base.

15. A multimodal analysis or prediction method, applied to a user device, characterized in that: include: Sending a multimodal analysis or prediction request to the fourth network element through the user; Receive the multimodal analysis or prediction result sent by the fourth network element through the user plane.

16. The method according to claim 15, characterized in that The multimodal analysis or prediction request includes at least one of the following: a knowledge classification identifier; a multimodal data type; and a device identifier.

17. The method according to claim 15, characterized in that Also includes: Registering multimodal capability information, where the multimodal capability information includes data types supported for reception and / or parsing; The data type includes at least one of the following: text, audio, video, and picture.

18. The method according to claim 17, characterized in that The multimodal capability information is transparently transmitted from the RAN to the AMF and stored in the AMF, UDM or UDR.

19. A multimodal analysis or prediction method, applied to a fourth network element, characterized in that: include: receiving a multimodal analysis or prediction request from a user device; Send a multimodal analysis or prediction request to the first network element according to the multimodal analysis or prediction of the user equipment.

20. The method according to claim 19, characterized in that Also includes: receiving a multimodal analysis or prediction result of the first network element; The multimodal analysis or prediction result is processed according to the capability of the user equipment, and the processed multimodal analysis or prediction result is sent to the user equipment.

21. The method according to claim 19, wherein Also includes: Send a query request to the AMF, UDM or UDR, where the query request carries the device identifier and is used to query the capabilities of the user equipment.

22. A network node, characterized in that: include: memory, and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the multimodal analysis or prediction method according to any one of claims 1 to 21.

23. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the multimodal analysis or prediction method according to any one of claims 1 to 21 is implemented.