Multimodal analysis or prediction method, network node, and storage medium
By receiving multimodal requests in the core network element and using embedded models and knowledge bases to retrieve vectors to generate multimodal analysis or prediction results, the problem of the core network element's inability to effectively perform multimodal analysis or prediction is solved, thereby improving network performance and user experience.
Patent Information
- Application Number
- PCT/CN2025/087736
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-09
- Filing Date
- 2025-04-08
- Publication Date
- 2026-02-12
AI Technical Summary
Existing core network elements fail to effectively support multimodal analysis or prediction, resulting in insufficient network performance and user experience.
By receiving multimodal analysis or prediction requests, the system converts multimodal data into vectors using an embedding model, retrieves vectors with similarity requirements from a knowledge base, and generates analysis or prediction results by combining the multimodal model. RAG technology is used to improve the accuracy of multimodal reasoning.
It improves the accuracy and efficiency of multimodal analysis or prediction in the core network, enhancing network performance and user experience.
Smart Images

Figure CN2025087736_12022026_PF_FP_ABST
Abstract
Description
Multi-modal analysis or prediction method, network node and storage medium TECHNICAL FIELD
[0001] The present application relates to the technical field of wireless communication, for example, to a multi-modal analysis or prediction method, a network node and a storage medium. BACKGROUND
[0002] In recent years, the Artificial Intelligence (AI) architecture of the core network has emerged. With the increase in data volume and data types, the UE and the servers inside and outside the core network may need multi-modal analysis or prediction / prediction in some scenarios to enhance the performance of the network and the experience of the user. However, the current network element does not support providing multi-modal analysis or prediction. How to realize multi-modal prediction or analysis of the core network and ensure the accuracy of the prediction or analysis is a problem to be solved. SUMMARY
[0003] The present application provides a multi-modal analysis or prediction method, a network node and a storage medium.
[0004] The present application provides a multi-modal analysis or prediction method, which is applied to a first network element and includes the following steps.
[0005] Receiving a multi-modal analysis or prediction request;
[0006] According to the multi-modal analysis or prediction request, converting multi-modal data into a first vector based on an embedding model;
[0007] Sending a retrieval request, the retrieval request being used to instruct a second network element to retrieve a second vector with a similarity meeting a requirement from a knowledge base;
[0008] Generating a multi-modal analysis or prediction result according to the second vector and a multi-modal model.
[0009] The present application also provides a multi-modal analysis or prediction method, which is applied to a second network element and includes the following steps.
[0010] Receiving a retrieval request of a first network element;
[0011] Retrieving a second vector with a similarity meeting a requirement from a knowledge base according to the retrieval request, the first vector being converted based on an embedding model from multi-modal data
[0012] The present application also provides a multi-modal analysis or prediction method, which is applied to a user equipment and includes the following steps.
[0013] Sending a multi-modal analysis or prediction request to a fourth network element through a user plane;
[0014] receive the multi-modal analysis or prediction result sent by the fourth network element through a user plane
[0015] The embodiment of the present application further provides a multi-modal analysis or prediction method, which is applied to a fourth network element and comprises the following steps:
[0016] receiving a multi-modal analysis or prediction request of a user equipment;
[0017] sending a multi-modal analysis or prediction request to a first network element according to the multi-modal analysis or prediction of the user equipment.
[0018] The embodiment of the present application further provides a network node, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the multi-modal analysis or prediction method described above when executing the program.
[0019] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the program is executable on a processor to implement the multi-modal analysis or prediction method described above. BRIEF DESCRIPTION OF DRAWINGS
[0020] Fig. 1 is a flow chart of a multi-modal analysis or prediction method according to an embodiment;
[0021] Fig. 2 is a flow chart of another multi-modal analysis or prediction method according to an embodiment;
[0022] Fig. 3 is a flow chart of still another multi-modal analysis or prediction method according to an embodiment;
[0023] Fig. 4 is a flow chart of yet another multi-modal analysis or prediction method according to an embodiment;
[0024] Fig. 5 is a schematic diagram of registration or update of multi-modal capability information of a first network element according to an embodiment;
[0025] Fig. 6 is a schematic diagram of a multi-modal analysis / prediction process according to an embodiment;
[0026] Fig. 7 is a schematic diagram of registration of multi-modal capability information of a UE according to an embodiment;
[0027] Fig. 8 is a schematic diagram of multi-modal analysis / prediction of a UE by a first network element according to an embodiment;
[0028] Fig. 9 is a structural schematic diagram of a multi-modal analysis or prediction apparatus according to an embodiment;
[0029] Fig. 10 is a structural schematic diagram of another multi-modal analysis or prediction apparatus according to an embodiment;
[0030] Fig. 11 is a structural schematic diagram of still another multi-modal analysis or prediction apparatus according to an embodiment;
[0031] Fig. 12 is a structural schematic diagram of yet another multi-modal analysis or prediction apparatus according to an embodiment;
[0032] Fig. 13 is a hardware structural schematic diagram of a network node according to an embodiment. DETAILED DESCRIPTION
[0033] The present application will be described below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other arbitrarily without conflict. In addition, it should be noted that only the parts related to the present application are shown in the drawings for convenience of description, rather than all the structures.
[0034] The Network Data Analytics Function (NWDAF) is a 5G Core (5GC) Network Function (NF) located in the control plane that can perform statistical data and machine learning related tasks. The NWDAF can interact with different entities for different purposes, e.g. collect data from event subscriptions provided by the Access and Mobility Management Function (AMF), the Session Management function (SMF), The User plane function (UPF), the Policy Control function (PCF), the Unified Data Management (UDM), the Network Slice Admission Control Function (NSACF), the Application Function (AF), directly or through the Network Exposure Function (NEF), and the Operation Administration and Maintenance (OAM); perform analytics and data collection using the Data Collection Coordination and Coordination and Delivery Function (DCCF); retrieve information from data repositories, e.g. the Unified Data Repository (UDR) to retrieve user related information through the UDM or retrieve Packet Flow Descriptions (PFD) information through the NEF (Packet Flow Descriptions Function (PFDF)); collect location information data from the Location Service (LCS) system; store and retrieve information from the Analytic Data Repository Function (ADRF); analyze and collect data from the Messaging Framework Adaptor Function (MFAF);retrieving information about NFs (e.g., retrieving NF related information from Network Repository Function (NRF)); providing analytics on-demand to consumers, such as providing batch data related to an analytics ID; providing accuracy information for an analytics ID; providing Machine Learning (ML) model accuracy information or ML model accuracy degradation indication about ML models. A single instance or multiple instances of NWDAF can be deployed in a Public Land Mobile Network (PLMN). NWDAF can contain the following logical functions:
[0035] Analytics Logical Function (AnLF): used to perform inference, derive analytics information (i.e., derive statistics and / or predictions based on analytics consumer requests), and expose analytics services (i.e., Nnwdaf_AnalyticsSubscription or Nnwdaf_AnalyticsInfo);
[0036] Model Training Logical Function (MTLF): a logical function in NWDAF used to train ML models and expose new training services (e.g., provide trained ML models).
[0037] One NWDAF can contain one MTLF or one AnLF, or both.
[0038] DCCF is responsible for coordinating data collection and distribution requested by NF consumers, preventing data sources from handling multiple subscriptions for the same data, and preventing multiple notifications containing the same information from being sent due to uncoordinated requests from data consumers.
[0039] DCCF is applicable to: NWDAF requesting data from data sources (e.g., for computing analytics); NF consumers requesting analytics from NWDAF data sources; NF consumers requesting data from ADRF data sources; ADRF receiving data from NF data sources.
[0040] The 5G system architecture supports ADRF to store collected data and analytics. ADRF exposes Nadrf service for information storage and retrieval of data by other 5GC network functions (e.g., NWDAF).
[0041] Based on a request from a network function or configuration on DCCF, DCCF can determine ADRF and interact with it directly or indirectly to request or store data. The way of interaction can be:
[0042] Direct: The DCCF requests data to be stored in the ADRF via the Nadrf service or Ndccf_DataManagement_Notify (e.g., when the ADRF requests the DCCF for data collection notifications). In addition, the DCCF retrieves data from the ADRF via the Nadrf service.
[0043] Indirect: The DCCF requests the message framework to store data in the ADRF via the Nadrf service or Nmfaf_3daDataManagement_Configure. The message framework can contain one or more adapters that convert between protocols defined by the 3rd Generation Partnership Project (3GPP).
[0044] The consumer network function can specify in the request sent to the DCCF that data provided by the data source needs to be stored in the ADRF.
[0045] The ADRF stores data received in the Nadrf_DataManagement_StorageRequest sent directly from network functions or in the Ndccf_DataManagement_Notify / Nmfaf_3caDataManagement_Notify or Nnwdaf_DataManagement_Notify sent from the DCCF, MFAF, or NWDAF.
[0046] RAG (Retrieval-Augmented Generation) is an emerging approach in natural language processing that combines the strengths of both retrieval and generation techniques. Specifically, RAG models rely on both the internal knowledge of pre-trained language models and the retrieval of relevant documents from external databases to enhance the accuracy and diversity of generated results when generating text.
[0047] The workflow of RAG mainly includes:
[0048] Retrieval phase: The model retrieves the most relevant document snippets from an external database based on the given input query. This step typically uses information retrieval techniques such as BM25 or vector retrieval.
[0049] Generation phase: Then, the input and retrieved document snippets are input into a generation model. The generation model combines the input and retrieved document snippets to generate more rich and accurate outputs.
[0050] When using large language models for inference, the parameter configurations used can significantly affect the quality and characteristics of the output content. The following are some common inference parameters and their roles: temperature, maximum length, prompt, short burst length, repetition penalty, Top-k sampling, Top-p (or Nucleus) sampling, batch size, and softness. These parameters can be used individually or in combination to improve the quality of the output analysis / prediction content. Depending on the specific application scenario, adjusting these parameters can significantly improve the output effect.
[0051] FIG. 1 is a flowchart of a multi-modal analysis or prediction method provided by an embodiment, which can be applied to a first network element, which can be a network element for multi-modal analysis or prediction, such as a NWDAF. As shown in FIG. 1, the method provided by the embodiment includes the following steps:
[0052] Step 110, receiving a multi-modal analysis or prediction request.
[0053] Step 120, converting multi-modal data into a first vector based on an embedding model according to the multi-modal analysis or prediction request.
[0054] Step 130, sending a retrieval request, the retrieval request being used to instruct a second network element to retrieve a second vector with a similarity to the first vector meeting a requirement from a knowledge base.
[0055] Step 140, generating a multi-modal analysis or prediction result according to the second vector and a multi-modal model.
[0056] The embedding model can also be referred to as a vector conversion model, which is used to convert multi-modal data into a vector. In the construction process of the knowledge base in the second network element, the embedding model can be used to convert multi-modal data into a vector and store it in the knowledge base of the second network element. When a multi-modal analysis or prediction request is made, the first network element can use the same embedding model to convert multi-modal data into a first vector, then request the second network element to retrieve a second vector with a similarity to the first vector meeting a requirement (such as a similarity greater than a threshold or a highest similarity) from the vectors stored in the knowledge base, and then the first network element uses a multi-modal model to generate a multi-modal analysis or prediction result according to the retrieved second vector. On this basis, multi-modal analysis or prediction is performed based on RAG, improving the quality and accuracy of core network multi-modal inference.
[0057] In an embodiment, the multi-modal analysis or prediction request includes at least one of the following: a knowledge classification identifier; a multi-modal analysis or prediction indication; the multi-modal analysis or prediction indication includes at least one parameter:
[0058] temperature parameter for controlling randomness of generated text, usually between 0.0 and 1.0, lower temperature (e.g. 0.2) will make the model output more deterministic content, higher temperature (e.g. 0.8) will make the output more creative and diverse;
[0059] maximum length of output for limiting the maximum length of generated text, the value depends on the specific model, usually from tens to thousands of tokens, setting a higher maximum length can avoid generating too short text, but may also cause the generated content to be lengthy and irrelevant;
[0060] the number of parameters of the model;
[0061] prompt words, such as prompt, can be used to provide context or initial part of generated text, appropriate prompt can guide the model to generate more content that meets user expectations; for example, the length of the prompt (Short Burst Length) can be the number of tokens generated each time, short generation can better control the generation process, but may reduce efficiency;
[0062] repetition penalty, for penalizing repeated content, usually between 1.0 and 2.0, higher repetition penalty value can reduce repetition in generated content;
[0063] sampling related parameters, such as Top-k sampling parameter, for limiting the number of candidate tokens at each generation, the value is a positive integer, by setting a k value, the model only selects from the top k tokens with the highest probability; for example, Top-p (or Nucleus) sampling parameter, for limiting the total probability mass of generated tokens, the value is between 0.0 and 1.0, by selecting a probability threshold p, the model only selects from tokens with cumulative probability reaching p; for example, batch size, which can be the number of samples generated at a time, larger batch size can improve generation efficiency, but requires more computing resources;
[0064] softness, for adjusting the diversity of generated text, similar to temperature, but more focused on adjusting the smoothness of the generation distribution.
[0065] In an embodiment, the method further comprises:
[0066] In the case of no multi-modal model locally, sending a multi-modal model acquisition request to the target network element;
[0067] Receiving the link or file of the multi-modal model sent by the target network element.
[0068] In an embodiment, the method further comprises: receiving the model capability information of the multi-modal model sent by the target network element.
[0069] In an embodiment, the acquisition request of the multi-modal model comprises at least one of the following: knowledge classification identification; model capability information; the model capability information comprises at least one of the following: supported data types, whether supporting RAG technology, whether belonging to a large model; the supported data types comprise at least one of the following: text, audio, video, and picture.
[0070] In an embodiment, the method further comprises:
[0071] sending a discovery request containing the network element of the multi-modal model to a third network element (such as NRF) and determining a target network element.
[0072] In an embodiment, the retrieval request comprises at least one of the following: knowledge base identification; knowledge classification identification.
[0073] In an embodiment, the method further comprises one of the following: sending the multi-modal analysis or prediction result to the consumer;
[0074] sending the multi-modal analysis or prediction result to a fourth network element (such as AF), and sending the multi-modal analysis or prediction result to the user equipment after being processed by the fourth network element according to the capability of the user equipment.
[0075] In an embodiment, the method further comprises: registering multi-modal capability information, the multi-modal capability information comprising at least one of the following: whether having the capability of multi-modal analysis or prediction; whether having a multi-modal model; whether having a multi-modal model supporting RAG technology; whether having an embedding model; and whether having a vector conversion capability.
[0076] In an embodiment, the embedding model for converting the multi-modal data into the first vector is consistent with the embedding model used for constructing the knowledge base of the second network element.
[0077] In an embodiment, the method further comprises:
[0078] sending the knowledge classification identification to the second network element for requesting to acquire the embedding model information used for constructing the knowledge base of the second network element;
[0079] receiving the knowledge base identification of the second network element and the embedding model information of the embedding model used for constructing the knowledge base.
[0080] FIG. 2 is a flowchart of a multi-modal analysis or prediction method provided by an embodiment, which can be applied to a second network element, which can be a network element for storing vectors and retrieving vectors. As shown in FIG. 2, the method provided by the embodiment comprises the following steps:
[0081] Step 210, receiving a retrieval request of a first network element.
[0082] Step 220, retrieving a second vector that meets the requirement of the similarity degree of the first vector from the knowledge base according to the retrieval request, the first vector being converted from the multi-modal data based on the embedding model.
[0083] In an embodiment, the embedding model used to convert the multi-modal data into the first vector is consistent with the embedding model used to build the knowledge base of the second network element.
[0084] In an embodiment, the method further comprises:
[0085] receiving a knowledge classification identifier;
[0086] locating the knowledge base and the embedding model used to build the knowledge base according to the knowledge classification identifier;
[0087] sending a knowledge base identifier and embedding model information of the embedding model used to build the knowledge base.
[0088] FIG. 3 is a flowchart of a multi-modal analysis or prediction method provided by an embodiment, which can be applied to a user equipment (UE). As shown in FIG. 3, the method provided by the embodiment comprises the following steps:
[0089] Step 310, sending a multi-modal analysis or prediction request to a fourth network element through a user plane.
[0090] Step 320, receiving a multi-modal analysis or prediction result sent by the fourth network element through the user plane.
[0091] In an embodiment, the multi-modal analysis or prediction request comprises at least one of the following: a knowledge classification identifier; a multi-modal data type; and a device identifier.
[0092] In an embodiment, the method further comprises:
[0093] registering multi-modal capability information, the multi-modal capability information comprising a data type supported by the UE for receiving and / or parsing;
[0094] The data type comprises at least one of the following: text, audio, video, and picture.
[0095] In an embodiment, the multi-modal capability information is transparently transmitted by a radio access network (RAN) to an AMF, and stored in the AMF, UDM, or UDR.
[0096] FIG. 4 is a flowchart of a multi-modal analysis or prediction method provided by an embodiment, which can be applied to a fourth network element, which can be an AF. As shown in FIG. 4, the method provided by the embodiment comprises the following steps:
[0097] Step 410, receiving a multi-modal analysis or prediction request of a user equipment.
[0098] Step 420, sending a multi-modal analysis or prediction request to a first network element according to the multi-modal analysis or prediction of the user equipment.
[0099] In an embodiment, the method further comprises:
[0100] receiving a multi-modal analysis or prediction result of the first network element;
[0101] processing the multi-modal analysis or prediction result according to the capability of the user equipment, and sending the processed multi-modal analysis or prediction result to the user equipment.
[0102] In an embodiment, the method further comprises: sending a query request to an AMF, a UDM or a UDR, the query request carrying a device identifier, the query request being used to query the capability of the user equipment.
[0103] In an embodiment of the application, the first network element (such as NWDAF) can receive an analysis / prediction request of a consumer, provide multi-modal analysis / prediction to the consumer by using a multi-modal model (which can be a large model) and RAG technology, and return the generated analysis / prediction result. In addition, after receiving a query request of a fourth network element (such as AF), the UE-related capability can be queried, and the analysis / prediction result can be generated according to the capability and sent to the AF. The models supporting multi-modal analysis / prediction can be transmitted between different first network elements (for example, NWDAF containing MTLF and NWDAF containing AnLF). The first network element can also register its multi-modal analysis / prediction capability and vector conversion capability to a third network element (such as NRF).
[0104] The third network element can receive the analysis / prediction capability and vector conversion capability sent by the first network element, and also support discovering a target network element (NWDAF containing AnLF) for vector conversion, and also support discovering ADRF for vector storage and knowledge base construction.
[0105] The second network element (which can be ADRF or a new NF) is used to construct a knowledge base, store vectors, support retrieval of the knowledge base, and extract vectors with high similarity.
[0106] The consumer (Consumer) NF (such as an Internet company Over-The-Top (OTT) server, a 5th Generation Core Network (5GC) network element and / or OAM, etc.) can send a multi-modal analysis / prediction request to the first network element, and receive an analysis / prediction result.
[0107] The user equipment can report its multi-modal capability in a registration message, send a multi-modal analysis / prediction request to the AF through the user plane, and also receive an analysis / prediction sent by the AF through the user plane.
[0108] The fourth network element can receive a multi-modal analysis / prediction request of the UE and send the multi-modal analysis / prediction request to the NWDAF; and also receive an analysis / prediction result returned by the NWDAF and determine the final analysis / prediction result sent to the UE by querying the UE-related capability information.
[0109] The multi-modal analysis or prediction method of the present application is exemplarily described below through some embodiments.
[0110] Embodiment 1
[0111] In this embodiment, the first network element (such as the NWDAF) can register or update the multi-modal capability to the third network element (such as the NRF). As shown in FIG. 5, the registration or update mainly includes:
[0112] 1. The NWDAF registers or updates the multi-modal capability to the NRF, which can include at least one of the following: whether the NWDAF supports multi-modal analysis / prediction, whether there is a multi-modal model locally, whether there is a multi-modal model supporting RAG, whether there is an embedded model, data types (such as text, video, audio and / or picture, etc.) supporting multi-modal analysis / prediction, and whether there is a multi-modal data-to-vector conversion capability, which can include embedded model information (such as the type, identification and / or manufacturer of the embedded model). The above multi-modal capability can be sent through the NF profile.
[0113] 2. The NRF stores the NF profile related to the NWDAF.
[0114] 3. The NRF returns a registration / update result indication, success or failure.
[0115] Embodiment 2
[0116] In this embodiment, the first network element (such as the NWDAF containing the AnLF) can request to obtain a multi-modal model from the target network element (such as the NWDAF containing the MTLF), thereby providing multi-modal analysis or prediction for the 5GC internal network element. As shown in FIG. 6,
[0117] The multi-modal analysis / prediction process mainly includes:
[0118] 1a. The consumer (i.e. Consumer, which can be a 5GC network element, OAM or OTT server, etc.) sends a discovery request to a third network element (such as NRF) to discover a first network element NWDAF that can provide multi-modal analysis / prediction. The request can contain data types supported by multi-modal analysis / prediction, such as text, audio, video and / or pictures, etc.
[0119] 1b. The NRF matches the appropriate NWDAF according to the pre-stored registration or update information of the NWDAF, and returns a list of NWDAFs that meet the requirements.
[0120] 2. The consumer determines the first network element (i.e. NWDAF containing AnLF in the figure) and sends a multi-modal analysis / prediction request to it, which can contain knowledge classification ID, data types (such as text, audio, video and / or pictures) contained in the required output analysis / prediction, an indication that a large model is used to provide multi-modal analysis / prediction, which can contain parameters of the large model, such as temperature parameter, maximum length of output, parameter amount of the large model, prompt word of the large model, repetition penalty, sampling related parameters and / or softness, etc.
[0121] 3a. If the NWDAF containing AnLF does not have a model locally to support multi-modal analysis / prediction, it can send a multi-modal model acquisition request to the NWDAF containing MTLF, which can contain capability information of the multi-modal model (such as supported data types, whether RAG is supported, whether it belongs to a large model, etc.), and can also contain knowledge classification ID, etc.
[0122] The NWDAF containing AnLF can send a discovery request to the NRF, containing multi-modal model related capability parameters (such as data types that the multi-modal model can provide, such as text, audio, video and / or pictures, etc.).
[0123] 3b. The NWDAF containing MTLF returns a link containing the multi-modal model or a file of the multi-modal model, which contains capability information of the multi-modal model (such as supported data types, whether RAG is supported, whether it belongs to a large model, etc.).
[0124] 4. The NWDAF decides that RAG technology needs to be used to generate analysis / prediction results. The NWDAF sends the knowledge classification ID provided by the consumer to the second network element (ADRF), thereby requesting to obtain embedding model information used for knowledge base construction.
[0125] 5. The ADRF uses the knowledge classification ID to locate the relevant knowledge base and embedding model information used for knowledge base construction.
[0126] 6. The ADRF sends the knowledge base ID associated with the knowledge classification ID and / or embedding model information (e.g., model ID and / or model type) to the NWDAF.
[0127] 7. Using the embedding model required by the ADRF in step 6, the relevant questions and / or requirements (text, audio, video, and / or pictures) of the multi-modal data received from the Consumer in step 2 are converted into vectors (i.e., first vectors). It should be noted that if the NWDAF does not support vector conversion or is insufficient in computing power, another NWDAF can be requested to perform vector conversion, and the converted vectors are returned.
[0128] It should be noted that during the process of building the knowledge base, the vectors stored by the ADRF are converted from multi-modal data by the embedding model. In order to ensure the accuracy and reliability of the output analysis / prediction results, the embedding model used in step 7 should be consistent with the embedding model used when building the knowledge base.
[0129] 8. The NWDAF initiates a retrieval process, and the knowledge base ID and knowledge classification ID can be carried in the retrieval request, the purpose of which is to extract a second vector with high similarity to the first vector in step 7 through the ADR.
[0130] 9. The NWDAF collects inference data from other network elements and OAM, performs inference using a local large model, and generates an analysis result in combination with the second vector obtained from the knowledge base in step 8.
[0131] 10. The NWDAF returns the analysis / prediction result to the Consumer, and the result is a multi-modal output, which can include text, audio, video, and / or pictures, etc.
[0132] Embodiment 3
[0133] In this embodiment, the user equipment (UE) can register multi-modal capability information. As shown in FIG. 7, the registration process mainly includes:
[0134] 1. The UE carries multi-modal capability information in the registration message, and the multi-modal capability information includes the types of data supported for receiving and / or parsing, such as text, audio, video, and / or pictures.
[0135] 2. The RAN transmits the multi-modal capability information to the AMF;
[0136] 3. The AMF stores the multi-modal capability information of the UE.
[0137] 4. If the AMF does not store the multi-modal capability of the UE, the multi-modal capability information can also be transmitted to the UDM, and the UDM can store the multi-modal capability information to the UDR (background of the UDM).
[0138] 5. The UDM / UDR stores the UE multimodal capability.
[0139] 6. The UDM returns a storage indication.
[0140] 7. The AMF returns a registration complete indication to the RAN
[0141] 8. The RAN returns a registration complete indication to the UE.
[0142] Embodiment 4
[0143] In this embodiment, the first network element (NWDAF) can provide multimodal analysis / prediction for a user equipment (UE). As shown in FIG. 8, the multimodal analysis / prediction procedure mainly includes:
[0144] 0. The UE sends a multimodal analysis / prediction request to the AF through a user plane, which can contain a knowledge classification ID. This request can contain related questions and / or requirements of multimodal data types (text, audio, video, and / or pictures), and can also contain a UE ID.
[0145] 1a. The AF (OTT server) decides that the analysis needs to be assisted by the NWDAF, sends a discovery request to the NRF to discover the NWDAF that can provide multimodal analysis / prediction.
[0146] 1b. The NRF matches the appropriate NWDAF according to the pre-stored NWDAF information, and returns a list of NWDAFs that meet the requirements.
[0147] 2. The AF decides that the NWDAF is needed for multimodal analysis / prediction, determines the first network element (i.e., the NWDAF in the figure) and sends a multimodal analysis / prediction request to the NWDAF, which can contain a knowledge classification ID, a UE ID, a requirement of data types (such as text, audio, video, and / or pictures) contained in the output analysis / prediction, an indication of application of a large model to provide multimodal analysis / prediction, which can contain parameters of the large model, such as temperature parameters, maximum length of output, parameter quantity of the large model, prompt words of the large model, repetition penalty, sampling related parameters, and / or gentleness, etc. This request can contain related questions and / or requirements of multimodal data types (text, audio, video, and / or pictures), which can be consistent with those sent by the UE in step 0, or can be optimized and then sent to the NWDAF.
[0148] Before step 3, the NWDAF can perform the steps 3a / 3b of the above embodiment, i.e., obtain a model for multimodal analysis / prediction from another NWDAF.
[0149] 3. The NWDAF decision needs to use the RAG technology for multi-modal analysis / prediction. The NWDAF sends the knowledge classification ID provided by the Consumer to the ADRF, requesting to obtain the embedding model used for knowledge base construction.
[0150] 4. The ADRF uses the knowledge classification ID to locate the relevant knowledge base and the embedding model information used during knowledge base construction.
[0151] 5. The ADRF sends the knowledge base ID and / or embedding model information (such as model ID and / or model type) associated with the knowledge classification ID to the NWDAF.
[0152] 6. Use the embedding model required by the ADRF in step 5 to convert the relevant questions and / or requirements (text, audio, video, and / or pictures) of the multi-modal data received from the Consumer in step 2 into vectors (i.e., first vectors). If the NWDAF does not support vector conversion locally or has insufficient computing power, it can also request another NWDAF to perform vector conversion and return the converted vectors.
[0153] 7. The NWDAF initiates a retrieval process, and the retrieval request can carry the knowledge base ID and knowledge classification ID, thereby extracting vectors with high similarity to the first vectors in step 6.
[0154] 8. The NWDAF collects inference data from other network elements, OAM, and uses the local large model for inference, combined with the second vectors obtained from the knowledge base in step 7, to generate analysis / prediction results. In addition, before generating the analysis / prediction results, the NWDAF can send a query request to the AMF / UDM / UDR, carrying the UE ID, to query the relevant capabilities of the UE.
[0155] 9. The NWDAF returns the multi-modal analysis / prediction results to the AF.
[0156] 10. After receiving the multi-modal analysis / prediction results returned by the NWDAF, the AF can perform secondary processing. During secondary processing, the AF can send a query request to the AMF / UDM / UDR, carrying the UE ID, to query the relevant capabilities of the UE, based on which secondary processing is performed, for example, if the multi-modal analysis / prediction results returned by the NWDAF contain multi-modal data that the UE does not support, the AF can perform deletion.
[0157] 11. The AF returns the secondary-processed multi-modal analysis / prediction results to the UE through the user plane.
[0158] Embodiments of the present application also provide a multi-modal analysis or prediction device. FIG. 9 is a structural schematic diagram of a multi-modal analysis or prediction device according to an embodiment. As shown in FIG. 9, the multi-modal analysis or prediction device comprises:
[0159] The request receiving module 510 is configured to receive a multi-modal analysis or prediction request;
[0160] The conversion module 520 is configured to convert multi-modal data into a first vector based on an embedding model according to the multi-modal analysis or prediction request;
[0161] The request retrieving module 530 is configured to send a retrieving request, the retrieving request being used to instruct a second network element to retrieve a second vector with a similarity satisfying a requirement from a knowledge base according to the first vector;
[0162] The analysis or prediction module 540 is configured to generate a multi-modal analysis or prediction result according to the second vector and a multi-modal model.
[0163] In an embodiment, the multi-modal analysis or prediction request comprises at least one of the following: a knowledge classification identifier; a multi-modal analysis or prediction indication;
[0164] The multi-modal analysis or prediction indication comprises at least one of the following parameters: a temperature parameter, an output maximum length, a parameter quantity, a prompt word, a repetition penalty degree, a sampling related parameter, and a softness.
[0165] In an embodiment, the apparatus further comprises a model obtaining module configured to, in a case where there is no multi-modal model locally, send an obtaining request of a multi-modal model to a target network element; and receive a link or a file of the multi-modal model sent by the target network element.
[0166] In an embodiment, the apparatus further comprises an information receiving module configured to receive model capability information of the multi-modal model sent by the target network element.
[0167] In an embodiment, the obtaining request of the multi-modal model comprises at least one of the following: a knowledge classification identifier; model capability information; the model capability information comprises at least one of the following: a supported data type, whether to support RAG technology, and whether to belong to a large model; and the supported data type comprises at least one of the following: text, audio, video, and picture.
[0168] In an embodiment, the apparatus further comprises a discovery module configured to send a discovery request of a network element comprising a multi-modal model to a third network element and determine a target network element.
[0169] In an embodiment, the retrieving request comprises at least one of the following: a knowledge base identifier; and a knowledge classification identifier.
[0170] In an embodiment, the apparatus further comprises a result sending module configured to perform one of the following:
[0171] send the multi-modal analysis or prediction result to a consumer;
[0172] sending the multimodal analysis or prediction result to a fourth network element, the multimodal analysis or prediction result being sent to the user equipment by the fourth network element after being processed according to the capability of the user equipment.
[0173] In an embodiment, the apparatus further comprises:
[0174] a registration module configured to register multimodal capability information, the multimodal capability information comprising at least one of:
[0175] whether having the capability of multimodal analysis or prediction; whether having a multimodal model; whether having a multimodal model supporting RAG technology; whether having an embedding model; whether having a vector conversion capability.
[0176] In an embodiment, the embedding model for converting the multimodal data into the first vector is consistent with an embedding model used for building the knowledge base of the second network element.
[0177] In an embodiment, the apparatus further comprises: an information request module configured to send a knowledge classification identifier to the second network element, for requesting to obtain embedding model information used for building the knowledge base of the second network element; and receive the knowledge base identifier of the second network element and the embedding model information of the embedding model used for building the knowledge base.
[0178] The multimodal analysis or prediction apparatus proposed in the embodiment belongs to the same inventive concept as the multimodal analysis or prediction method proposed in the above embodiments, and the technical details not described in the embodiment can be referred to any of the above embodiments, and the embodiment has the same beneficial effects as the multimodal analysis or prediction method.
[0179] Embodiments of the present application also provide a multimodal analysis or prediction apparatus. FIG. 10 is a structural schematic diagram of a multimodal analysis or prediction apparatus provided in an embodiment. As shown in FIG. 10, the multimodal analysis or prediction apparatus comprises:
[0180] a request receiving module 610 configured to receive a search request of a first network element;
[0181] a search module 620 configured to search, according to the search request, a second vector from a knowledge base, the second vector satisfying a requirement in terms of similarity with the first vector, the first vector being obtained by converting multimodal data based on an embedding model.
[0182] In an embodiment, the embedding model for converting the multimodal data into the first vector is consistent with an embedding model used for building the knowledge base of the second network element.
[0183] In an embodiment, the apparatus further comprises:
[0184] a receiving module configured to receive a knowledge classification identifier;
[0185] a positioning module configured to position a knowledge base according to the knowledge classification identifier and an embedding model used for constructing the knowledge base;
[0186] a sending module configured to send the knowledge base identifier and embedding model information of the embedding model used for constructing the knowledge base.
[0187] The multi-modal analysis or prediction device proposed in this embodiment belongs to the same inventive concept as the multi-modal analysis or prediction method proposed in the above embodiments, and technical details not described in detail in this embodiment can be referred to any of the above embodiments, and this embodiment has the same beneficial effects as performing the multi-modal analysis or prediction method.
[0188] This application also provides a multi-modal analysis or prediction device. Figure 11 is a structural schematic diagram of a multi-modal analysis or prediction device provided by an embodiment. As shown in Figure 11, the multi-modal analysis or prediction device comprises:
[0189] The request module 710 is configured to send a multi-modal analysis or prediction request to the fourth network element through the user plane.
[0190] The receiving module 720 is configured to receive a multi-modal analysis or prediction result sent by the fourth network element through the user plane.
[0191] In an embodiment, the multi-modal analysis or prediction request comprises at least one of the following: a knowledge classification identifier; a multi-modal data type; and a device identifier.
[0192] In an embodiment, the device further comprises:
[0193] The registration module is configured to register multi-modal capability information, wherein the multi-modal capability information comprises a data type supported by the device for receiving and / or parsing; and the data type comprises at least one of the following: text, audio, video, and picture.
[0194] In an embodiment, the multi-modal capability information is transparently transmitted by the RAN to the AMF, and stored in the AMF, UDM, or UDR.
[0195] The multi-modal analysis or prediction device proposed in this embodiment belongs to the same inventive concept as the multi-modal analysis or prediction method proposed in the above embodiments, and technical details not described in detail in this embodiment can be referred to any of the above embodiments, and this embodiment has the same beneficial effects as performing the multi-modal analysis or prediction method.
[0196] This application also provides a multi-modal analysis or prediction device. Figure 12 is a structural schematic diagram of a multi-modal analysis or prediction device provided by an embodiment. As shown in Figure 12, the multi-modal analysis or prediction device comprises:
[0197] The receiving module 810 is configured to receive a multimodal analysis or prediction request of a user equipment.
[0198] The sending module 820 is configured to send a multimodal analysis or prediction request to a first network element according to the multimodal analysis or prediction of the user equipment.
[0199] In an embodiment, the apparatus further includes:
[0200] The result receiving module is configured to receive a multimodal analysis or prediction result of the first network element.
[0201] The result processing module is configured to process the multimodal analysis or prediction result according to the capability of the user equipment, and send the processed multimodal analysis or prediction result to the user equipment.
[0202] In an embodiment, the apparatus further includes:
[0203] The querying module is configured to send a querying request to an AMF, a UDM or a UDR, the querying request carrying a device identifier, and the querying request being used to query the capability of the user equipment.
[0204] The multimodal analysis or prediction apparatus proposed in the embodiment belongs to the same inventive concept as the multimodal analysis or prediction method proposed in the above embodiments, and the technical details not described in the embodiment can be referred to the above embodiments, and the embodiment has the same beneficial effects as the multimodal analysis or prediction method.
[0205] Embodiments of the present application further provide a network node. FIG. 13 is a schematic diagram of a hardware structure of a network node according to an embodiment. As shown in FIG. 13, the network node provided by the present application includes a processor 910 and a memory 920. The processor 910 in the network node can be one or more, and FIG. 13 takes one processor 910 as an example. The memory 920 is configured to store one or more programs. The one or more programs are executed by the one or more processors 910, so that the one or more processors 910 implement the multimodal analysis or prediction method according to the embodiments of the present application.
[0206] The network node further includes a communication device 930, an input device 940 and an output device 950.
[0207] The processor 910, the memory 920, the communication device 930, the input device 940 and the output device 950 in the network node can be connected through a bus or other means, and FIG. 13 takes the connection through the bus as an example.
[0208] The input device 940 can be used to receive input digital or character information, and to generate key signal inputs related to user settings of the network node and function controls. The output device 950 can include a display screen or the like display device.
[0209] The communication device 930 can include a receiver and a transmitter. The communication device 930 is configured to perform information transceiving communication under the control of the processor 910.
[0210] The memory 920, as a kind of computer readable storage medium, can be configured to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the multi-modal analysis or prediction method described in the embodiments of the present application (for example, modules in the multi-modal analysis or prediction device). The memory 920 can include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required by at least one function; the data storage area can store data created according to the use of the network node, etc. In addition, the memory 920 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state memory device. In some examples, the memory 920 can further include a memory remotely arranged with respect to the processor 910, which can be connected to the network node through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0211] The embodiments of the present application also provide a storage medium, which stores a computer program, and the computer program is executed by a processor to implement the multi-modal analysis or prediction method described in any of the embodiments of the present application. The method is applied to a first network element, including: receiving a multi-modal analysis or prediction request; converting multi-modal data into a first vector based on an embedding model according to the multi-modal analysis or prediction request; sending a retrieval request, the retrieval request being used to instruct a second network element to retrieve a second vector with a similarity satisfying a requirement from a knowledge base; generating a multi-modal analysis or prediction result according to the second vector and a multi-modal model. Alternatively, the method is applied to a second network element, including: receiving a retrieval request of a first network element; retrieving a second vector with a similarity satisfying a requirement from a knowledge base according to the retrieval request, the first vector being converted based on an embedding model from multi-modal data. Alternatively, the method is applied to a user equipment, including: sending a multi-modal analysis or prediction request to a fourth network element through a user plane; receiving a multi-modal analysis or prediction result sent by the fourth network element through the user plane. Alternatively, the method is applied to a fourth network element, including: receiving a multi-modal analysis or prediction request of a user equipment; sending a multi-modal analysis or prediction request to a first network element according to the multi-modal analysis or prediction of the user equipment.
[0212] The embodiments of the present application also provide a computer program product, comprising computer programs / instructions, which, when executed by a processor, implement the multi-modal analysis or prediction method of any of the embodiments of the present application. The method is applied to a first network element and comprises: receiving a multi-modal analysis or prediction request; converting multi-modal data into a first vector based on an embedding model according to the multi-modal analysis or prediction request; sending a retrieval request, the retrieval request being used to instruct a second network element to retrieve a second vector with a similarity satisfying a requirement from a knowledge base; and generating a multi-modal analysis or prediction result according to the second vector and a multi-modal model. Alternatively, the method is applied to the second network element and comprises: receiving a retrieval request of the first network element; retrieving a second vector with a similarity satisfying a requirement from a knowledge base according to the retrieval request, the first vector being converted from multi-modal data based on an embedding model. Alternatively, the method is applied to a user equipment and comprises: sending a multi-modal analysis or prediction request to a fourth network element through a user plane; and receiving a multi-modal analysis or prediction result sent by the fourth network element through the user plane. Alternatively, the method is applied to the fourth network element and comprises: receiving a multi-modal analysis or prediction request of a user equipment; and sending a multi-modal analysis or prediction request to a first network element according to the multi-modal analysis or prediction of the user equipment.
[0213] The computer storage medium of the embodiments of the present application can adopt any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may, for example, be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device.
[0214] A computer readable signal medium can include a propagated data signal with computer executable prograrn code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport programming for use by or in connection with an instruction execution system, apparatus, or device.
[0215] Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, Radio Frequency (RF) etc., or any suitable combination of the foregoing.
[0216] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0217] The embodiments of the present application also provide a computer program product, including computer programs / instructions, which, when executed by a processor, implement the video encoding method according to any of the above embodiments.
[0218] The above descriptions are merely used to illustrate the exemplary embodiments of the present application, but not intended to limit and construe the present application.
[0219] Those skilled in the art should understand that the term user terminal covers any suitable type of wireless user equipment, such as a mobile phone, a portable data processing portable network browser or a vehicle mounted mobile station.
[0220] In general, the various embodiments of the application can be implemented in hardware or special-purpose circuits, software, logic or any combination thereof. For example, some aspects can be implemented in hardware, while other aspects can be implemented in
[0221] Embodiments of the application can be implemented by computer program instructions being executed by a data processor of a mobile device, for example in processor entities, or by hardware, or by a combination of software and hardware. The computer program instructions can be assembly instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, or be written in any combination of one or more programming languages, either declared or undeclared, to create source code or object code.
[0222] The block diagrams of any logical flow of the present application in the accompanying drawings can represent program steps or can represent interconnecting logic circuit, modules, and functions, or a combination of program steps and logic circuit, modules, and functions. The computer program can be stored in a memory. The memory can be of any type suitable to the local technical environment and can be implemented using any suitable data storage technology, such as a semiconductor-based memory device, or a system and / or a medium of a computer-readable medium. The computer-readable medium can include non-transitory storage media. The data processor can be of any type suitable to the local technical environment, and can include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), field- programmable gate arrays (FPGAs), and processors of multi-core processor architectures, as examples.
[0223] A detailed description of exemplary embodiments of the application has been provided above with reference to the accompanying drawings, but various modifications and changes can be made to the above embodiments by those skilled in the art without departing from the scope of the application, which is defined by the following claims. Accordingly, the proper scope of the application is to be determined not only by the embodiments disclosed above, but also by the reasonable equivalents thereof, and therefore the proper scope of the application is to be determined in accordance with the claims.
Claims
1. A multi-modal analysis or prediction method applied to a first network element, comprising: receiving a multi-modal analysis or prediction request; converting multi-modal data into a first vector based on an embedding model according to the multi-modal analysis or prediction request; sending a retrieval request for instructing a second network element to retrieve a second vector with a similarity satisfying a requirement from a knowledge base; generating a multi-modal analysis or prediction result according to the second vector and a multi-modal model.
2. The method of claim 1, wherein, The multi-modal analysis or prediction request comprises at least one of the following: a knowledge classification identifier; a multi-modal analysis or prediction indication; The multi-modal analysis or prediction indication comprises at least one of the following parameters: a temperature parameter, an output maximum length, a parameter quantity, a prompt word, a repetition penalty, a sampling related parameter, and a gentleness.
3. The method of claim 1, further comprising: in a case where there is no multi-modal model locally, sending a multi-modal model acquisition request to a target network element; receiving a link or a file of the multi-modal model sent by the target network element.
4. The method of claim 3, further comprising: receiving model capability information of the multi-modal model sent by the target network element.
5. The method of claim 3, wherein, The multi-modal model acquisition request comprises at least one of the following: a knowledge classification identifier; model capability information; The model capability information comprises at least one of the following: a supported data type, whether to support retrieval enhancement generation (RAG) technology, and whether to belong to a large model; The supported data type comprises at least one of the following: text, audio, video, and picture.
6. The method of claim 3, further comprising: sending a discovery request containing a network element of a multi-modal model to a third network element and determining a target network element.
7. The method of claim 1, wherein, The retrieval request comprises at least one of the following: a knowledge base identifier; and a knowledge classification identifier.
8. The method of claim 1, further comprising one of the following: sending a multi-modal analysis or prediction result to a consumer; sending a multi-modal analysis or prediction result to a fourth network element, which sends the multi-modal analysis or prediction result to a user device after processing according to a capability of the user device.
9. The method of claim 1, further comprising: registering multi-modal capability information, which comprises at least one of the following: whether to have a multi-modal analysis or prediction capability; whether to have a multi-modal model; whether to have a multi-modal model supporting RAG technology; whether to have an embedding model; and whether to have a vector conversion capability.
10. The method of claim 1, wherein, The embedding model used for converting the multi-modal data into the first vector is consistent with an embedding model used for constructing a knowledge base of the second network element.
11. The method of claim 10, further comprising: Sending a knowledge classification identifier to the second network element for requesting to acquire embedding model information used for constructing the knowledge base of the second network element; Receiving a knowledge base identifier of the second network element and embedding model information of the embedding model used for constructing the knowledge base.
12. A multi-modal analysis or prediction method applied to a second network element, comprising: receiving a retrieval request from a first network element; retrieving a second vector with a similarity satisfying a requirement from a knowledge base according to the retrieval request, the first vector being converted from multi-modal data based on an embedding model.
13. The method of claim 12, wherein, The embedding model used to convert the multi-modal data into the first vector is consistent with the embedding model used to build the knowledge base of the second network element.
14. The method of claim 12, further comprising: receiving a knowledge classification identifier; locating a knowledge base and an embedding model used to build the knowledge base according to the knowledge classification identifier; sending an embedding model information of the knowledge base identifier and the embedding model used to build the knowledge base.
15. A multi-modal analysis or prediction method applied to a user equipment, comprising: sending a multi-modal analysis or prediction request to a fourth network element through a user plane; receiving a multi-modal analysis or prediction result sent by the fourth network element through the user plane.
16. The method of claim 15, wherein, The multi-modal analysis or prediction request comprises at least one of the following: a knowledge classification identifier; a multi-modal data type; and a device identifier.
17. The method of claim 15, further comprising: registering multi-modal capability information, the multi-modal capability information comprising a data type supported by the user equipment for receiving and / or parsing; the data type comprises at least one of the following: text, audio, video, and picture.
18. The method of claim 17, wherein, The multi-modal capability information is transparently transmitted by a radio access network (RAN) to an access and mobility management function (AMF), and stored in the AMF, a unified data management (UDM), or a unified data repository (UDR).
19. A multi-modal analysis or prediction method applied to a fourth network element, comprising: receiving a multi-modal analysis or prediction request from a user equipment; sending a multi-modal analysis or prediction request to a first network element according to a multi-modal analysis or prediction of the user equipment.
20. The method of claim 19, further comprising: receiving a multi-modal analysis or prediction result from the first network element; processing the multi-modal analysis or prediction result according to a capability of the user equipment, and sending the processed multi-modal analysis or prediction result to the user equipment.
21. The method of claim 19, further comprising: sending a query request to an access and mobility management function (AMF), a unified data management (UDM), or a unified data repository (UDR), the query request carrying a device identifier, the query request being used to query a capability of the user equipment.
22. A network node, comprising: a memory, and one or more processors; the memory is configured to store one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the multi-modal analysis or prediction method according to any one of claims 1-21.
23. A computer readable storage medium having stored thereon a computer program, characterized in that, The program, when executed by a processor, implements the multi-modal analysis or prediction method according to any one of claims 1-21.
Citation Information
Patent Citations
Relay communication method and device, communication equipment, communication system and storage medium
CN117099326A
Intelligent question and answer method, system and device, computer equipment and readable storage medium
CN118377881A
Method and system for providing service experience analysis based on network data analysis
US20200358670A1
Methods, architectures, apparatuses and systems for multi-modal communication including multiple user devices
WO2023167979A1
Information processing method and apparatus, related devices, and storage medium
WO2023179604A1