Multi-modal data interaction method, system and device and storage medium

By adopting IMS registration, federated learning and computing power scheduling in 5G networks, dynamic allocation of resources is solved, and the problem of insufficient resource dispersion and intelligence is achieved, efficient multimodal data interaction is achieved, and the communication experience of the elderly group is improved.

CN120583433APending Publication Date: 2025-09-02IPLOOK NETWORKS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510515114.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing 5G networks have shortcomings in resource dispersion, insufficient intelligence, and energy consumption and delay issues, and cannot effectively support the multimodal interaction needs of the elderly population.

Method used

Device access and registration are carried out through the IMS defined by 3GPP, combining federated learning and computing power dynamic scheduling, dynamic distribution of network slices, and deploying multimodal converged servers on the IMS media surface, integrating multimodal data and interacting with the terminal rendering engine.

Benefits of technology

It realizes dynamic adjustment of network resources, improves data interaction efficiency, supports natural and efficient multimodal interaction among the elderly, and reduces equipment energy consumption and delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120583433A_ABST
    Figure CN120583433A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data interaction method, system and device and a storage medium. The method comprises the following steps: accessing and registering a plurality of devices through an IMS defined by a 3GPP; determining an event priority according to the identifier of the quality of service flow, and distributing network slices; network resources are allocated through federal learning cooperative training and dynamic computing power scheduling; and deploying a multi-modal fusion server on an IMS media plane, integrating multi-modal data, and performing data interaction and display with a rendering engine deployed by a terminal. According to the embodiment of the invention, the network slices and the resources are dynamically adjusted, so that the data interaction efficiency is improved. The method can be widely applied to the technical field of 5G.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of 5G technology, and in particular to a multimodal data interaction method, system, device and storage medium. Background Art

[0002] 5G networks not only offer higher transmission rates, lower latency, and greater reliability, but can also combine AI, smart devices, and immersive communication technologies to create more natural and efficient interactions for seniors and their families. For example, by integrating 5G networks with devices like smart glasses and smartwatches, users can engage in real-time voice and video interactions in an immersive communication environment, while enjoying features like real-time translation and the fusion of virtual and physical reality, thereby improving the quality of life and communication experience for these groups.

[0003] Related technical solutions suffer from resource silos, with computing, perception, and communication resources dispersed, unable to be dynamically coordinated and scheduled, and lacking dynamic adjustment capabilities. They also suffer from insufficient intelligence. The computing power on the edge where AI models are deployed is limited, making it difficult to support real-time multimodal interactions (such as gesture recognition and voice synchronization). Energy consumption and latency are also issues. Traditional network slicing is rigid, resulting in high energy consumption and latency jitter. There is a lack of detailed analysis and management of energy consumption across different services, and a failure to intelligently optimize the balance between high-energy consumption and critical services. Summary of the Invention

[0004] The purpose of the present invention is to solve one of the technical problems existing in the prior art to at least a certain extent.

[0005] To this end, an object of the present invention is to provide an efficient multimodal data interaction method, system, device and storage medium.

[0006] To achieve the above technical objectives, one aspect of an embodiment of the present invention provides a multimodal data interaction method, comprising the following steps: accessing and registering a number of devices through the IMS defined by 3GPP; determining event priorities and allocating network slices based on the identifier of the quality of service flow; allocating network resources through federated learning collaborative training and dynamic scheduling of computing power; deploying a multimodal fusion server on the IMS media plane to integrate multimodal data and interact and display data with the rendering engine deployed on the terminal. The embodiment of the present application dynamically adjusts network slices and resources, which is conducive to improving the efficiency of data interaction.

[0007] In some embodiments, the method of allocating network resources through federated learning collaborative training, dynamic computing power scheduling, and network resource allocation includes:

[0008] Based on the distributed intelligent architecture defined by 3GPP, federated learning is performed on edge nodes and the cloud, and lightweight models are deployed on edge nodes;

[0009] Embed computing power scheduling functions into the network protocol stack through distributed intelligent agents, and build local decision-making models based on federated learning;

[0010] Network resources are allocated through the lightweight model and the local decision model.

[0011] In some embodiments, deploying a multimodal fusion server on the IMS media plane to integrate multimodal data includes:

[0012] Integrate voice data and gesture data through voice burst control protocol;

[0013] Identify the user's dynamic behavior and confirm the user's status through sensors;

[0014] Based on the 3GPP AR call standard, AR media stream channels are configured through SIP signaling to integrate virtual images with physical scenes.

[0015] In some embodiments, accessing and registering a plurality of devices through the IMS defined by 3GPP includes:

[0016] Receiving registration requests sent by the plurality of devices via the SIP protocol;

[0017] Authentication and service triggering are performed through the service call session control function.

[0018] In some embodiments, the method further comprises:

[0019] Establishing a native fusion architecture to integrate functions related to the multimodal data into the network protocol stack;

[0020] Provide dynamic loading and collaborative scheduling of multi-factor capabilities through hardware and software platforms, and connect the data of the devices.

[0021] In some embodiments, the method performs collaborative training by:

[0022] The first edge node is trained based on historical data to obtain a trained lightweight model;

[0023] According to the trained lightweight model, the model weights are exchanged with the adjacent second edge node to obtain an updated lightweight model.

[0024] In some embodiments, the method further comprises:

[0025] Reuse the hardware and spectrum resources of communication base stations into sensing nodes to form a distributed sensing network;

[0026] Collaborative networking and data sharing are carried out through the distributed sensing network.

[0027] On the other hand, an embodiment of the present invention provides a multimodal data interaction system, including:

[0028] The first module is used to access and register several devices through the IMS defined by 3GPP;

[0029] The second module is used to determine the event priority and allocate network slices based on the identifier of the quality of service flow;

[0030] The third module is used to allocate network resources through federated learning collaborative training, dynamic computing power scheduling;

[0031] The fourth module is used to deploy a multimodal fusion server on the IMS media plane, integrate multimodal data, and interact and display data with the rendering engine deployed on the terminal.

[0032] On the other hand, an embodiment of the present invention provides a multimodal data interaction device, including:

[0033] at least one processor;

[0034] at least one memory for storing at least one program;

[0035] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned multimodal data interaction method.

[0036] On the other hand, an embodiment of the present invention provides a storage medium storing a program executable by a processor. When the program is executed by the processor, it is used to implement the above-mentioned multimodal data interaction method.

[0037] The embodiments of the present application include at least the following beneficial effects: The method provided by the embodiments of the present invention includes: accessing and registering multiple devices through the IMS defined by 3GPP; determining event priorities and allocating network slices based on quality of service flow identifiers; allocating network resources through federated learning collaborative training and dynamic computing power scheduling; deploying a multimodal fusion server on the IMS media plane to integrate multimodal data and interact and display data with the rendering engine deployed on the terminal. The embodiments of the present application dynamically adjust network slices and resources, which is conducive to improving the efficiency of data interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following introduction is made to the drawings of the embodiments of the present invention or the related technical solutions in the prior art. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative work.

[0039] Figure 1 A schematic diagram of a flow chart of an embodiment of the multimodal data interaction method provided by the present invention;

[0040] Figure 2 A schematic diagram of a flow chart of another embodiment of the multimodal data interaction method provided by the present invention;

[0041] Figure 3 A schematic structural diagram of an embodiment of the multimodal data interaction system provided by the present invention;

[0042] Figure 4 A schematic structural diagram of an embodiment of the multimodal data interaction device provided by the present invention. DETAILED DESCRIPTION

[0043] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and are not to be construed as limiting the present invention. The step numbers in the following embodiments are provided for ease of explanation only and do not limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0044] First, the terms used in this application are explained:

[0045] UE User Equipment: User equipment;

[0046] 5GS 5G System: 5G system;

[0047] Immersive RTC: immersive real-time communication service;

[0048] MEC: Multi-access edge computing;

[0049] NWDAF: Network data analysis function;

[0050] QoS (Quality of Service): Quality of Service, which ensures service performance (such as low latency for video calls) through priority marking (5QI).

[0051] 5QI: QoS flow identifier;

[0052] IMS: IP Multimedia Subsystem;

[0053] SIP Protocol (Session Initiation Protocol): short for Session Initiation Protocol, is a text-based application layer signaling protocol, mainly used to establish, modify and terminate real-time multimedia communication sessions (such as voice, video, instant messaging, etc.).

[0054] S-CSCF: Service Call Session Control Function;

[0055] AR: Augmented Reality;

[0056] VR: Virtual Reality;

[0057] DIA: Distributed Intelligent Agent;

[0058] DPS: Dynamic Member Selection Algorithm.

[0059] Traditional communication technologies can no longer fully meet the elderly's needs for efficient communication and life support, and the emergence of 5G technology provides an innovative solution to this problem.

[0060] 5G networks not only offer higher transmission rates, lower latency, and greater reliability, but can also combine AI, smart devices, and immersive communication technologies to create more natural and efficient interactions for seniors and their families. For example, by combining devices like smart glasses and smartwatches with 5G networks, users can engage in real-time voice and video interactions in an immersive communication environment, while enjoying features like real-time translation and the fusion of virtual and physical reality, thereby improving the quality of life and communication experience for seniors.

[0061] In 5G networks, MEC moves computing power to the edge of the network, supporting low-latency local data processing such as real-time video rendering. It analyzes network traffic and user behavior through the Network Data Analysis Function (NWDAF) to optimize resource allocation. However, it relies on external AI models and lacks native integration with the network architecture. The single perception function and communication system are designed independently, resulting in low resource utilization.

[0062] Specifically, the shortcomings of the existing technology are:

[0063] 1. Resource siloing: 5G's MEC and NWDAF are external modules, with scattered computing, perception, and communication resources. They cannot be dynamically coordinated and scheduled, and lack dynamic adjustment capabilities.

[0064] 2. Insufficient intelligence: AI model deployment relies on centralized training in the cloud, and the computing power on the terminal side is limited, making it difficult to support real-time multimodal interaction (such as gesture recognition and voice synchronization).

[0065] 3. Energy consumption and latency issues: Traditional network slicing is fixed and cannot dynamically adjust QoS according to business needs (such as AR rendering), resulting in high energy consumption and latency jitter. There is a lack of detailed analysis and management of the energy consumption of different services, and the balance between high-energy consumption and critical services cannot be intelligently optimized.

[0066] The present invention aims to propose an immersive real-time communication service architecture based on 5G synaesthesia, computing and intelligence fusion, including the following improvements:

[0067] 1. Native integration: Embeds perception, computing, and intelligent functions into the network protocol stack to achieve dynamic resource orchestration;

[0068] 2. Distributed intelligence: Federated learning combined with edge computing optimizes real-time processing and privacy protection of multimodal data;

[0069] 3. Adaptive QoS: AI-based prediction of user needs, dynamic allocation of network slices and computing resources, and reduction of end-to-end latency and energy consumption.

[0070] The multimodal data interaction method and system proposed in embodiments of the present invention will be described in detail below with reference to the accompanying drawings. First, the multimodal data interaction method proposed in embodiments of the present invention will be described with reference to the accompanying drawings.

[0071] Reference Figure 1 , a multimodal data interaction method is provided in an embodiment of the present invention. The multimodal data interaction method in the embodiment of the present invention can be applied to a terminal, can be applied to a server, or can be software running in a terminal or a server. The terminal can be a tablet computer, a laptop computer, a desktop computer, etc., but is not limited to this. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The multimodal data interaction method in the embodiment of the present invention mainly includes the following steps:

[0072] S100: Access and register several devices through the IMS defined by 3GPP;

[0073] S200: Determine event priority and allocate network slices based on the identifier of the quality of service flow;

[0074] S300: Allocates network resources through federated learning collaborative training, dynamic computing power scheduling, and network resource allocation.

[0075] S400: Deploy a multimodal fusion server on the IMS media plane to integrate multimodal data and interact and display data with the rendering engine deployed on the terminal.

[0076] In some embodiments, the method of allocating network resources through federated learning collaborative training, dynamic computing power scheduling, and network resource allocation includes:

[0077] Based on the distributed intelligent architecture defined by 3GPP, federated learning is performed on edge nodes and the cloud, and lightweight models are deployed on edge nodes;

[0078] Embed computing power scheduling functions into the network protocol stack through distributed intelligent agents, and build local decision-making models based on federated learning;

[0079] Network resources are allocated through the lightweight model and the local decision model.

[0080] In some embodiments, deploying a multimodal fusion server on the IMS media plane to integrate multimodal data includes:

[0081] Integrate voice data and gesture data through voice burst control protocol;

[0082] Identify the user's dynamic behavior and confirm the user's status through sensors;

[0083] Based on the 3GPP AR call standard, AR media stream channels are configured through SIP signaling to integrate virtual images with physical scenes.

[0084] In some embodiments, accessing and registering a plurality of devices through the IMS defined by 3GPP includes:

[0085] Receiving registration requests sent by the plurality of devices via the SIP protocol;

[0086] Authentication and service triggering are performed through the service call session control function.

[0087] In some embodiments, the method further comprises:

[0088] Establishing a native fusion architecture to integrate functions related to the multimodal data into the network protocol stack;

[0089] Provide dynamic loading and collaborative scheduling of multi-factor capabilities through hardware and software platforms, and connect the data of the devices.

[0090] In some embodiments, the method performs collaborative training by:

[0091] The first edge node is trained based on historical data to obtain a trained lightweight model;

[0092] According to the trained lightweight model, the model weights are exchanged with the adjacent second edge node to obtain an updated lightweight model.

[0093] In some embodiments, the method further comprises:

[0094] Reuse the hardware and spectrum resources of communication base stations into sensing nodes to form a distributed sensing network;

[0095] Collaborative networking and data sharing are carried out through the distributed sensing network.

[0096] The following is a detailed description of the interaction method provided by this application using a specific embodiment:

[0097] The present invention proposes an immersive real-time communication (Immersive RTC) system based on the deep integration of 5G synaesthesia, computing and intelligence, which achieves innovation through native fusion architecture, distributed intelligent optimization, dynamic resource orchestration, synaesthesia integration and immersive communication. The 5G network and the immersive real-time communication service (Immersive RTC) it supports are not only limited to communication between people, but also support seamless interaction between people and smart devices (such as robots and home smart devices). These interactions are supported by multimodal technologies (such as comprehensive recognition of voice, gestures, and actions) and intelligent computing optimization functions provided by the 3GPP system, making communications and services more efficient and closer to user needs. In addition, by deeply integrating virtual and physical scenes, immersive communication technology can predict and adjust user needs in real time, making the interaction between users and devices more natural, while significantly reducing the energy consumption of devices. The combination of these technologies provides a solid technical foundation for building future communication services for multi-device collaboration, and also brings new possibilities for solving the problem of social aging.

[0098] Native framework integration:

[0099] A native converged architecture deeply integrates communication, perception, computing, and intelligent functions into the network protocol stack. This unified hardware and software platform enables dynamic loading and coordinated scheduling of multiple capabilities, breaking down protocol barriers between communication, perception, and computing modules and enabling unified resource scheduling. Its core goal is to break down traditional technology boundaries and avoid the overlay of "plug-in" functions, thereby optimizing system efficiency and resource utilization. Within this architecture, various data sources, such as sensors, user terminals, and cloud services, can be seamlessly connected, enabling efficient data flow and sharing. This convergence not only improves overall service efficiency but also makes network management simple and intuitive. Furthermore, through APIs, various applications can easily connect to the service layer, enabling cross-platform data exchange and function invocation.

[0100] Distributed intelligent optimization:

[0101] Distributed intelligent optimization leverages edge computing and federated learning technologies to enable real-time data processing and model training across multiple nodes while protecting data privacy. This approach aims to reduce latency and improve system robustness, thereby enhancing network performance and meeting evolving user needs. For example, edge node A trains a lightweight model based on historical load data to predict future computing power requirements within 5ms. It then exchanges model weights with neighboring edge node B, improving regional collaboration efficiency. This process not only improves network efficiency but also reduces operating costs. By proactively identifying and addressing issues, operations teams can resolve network outages more quickly, achieve self-healing, and reduce manual intervention. The ultimate goal of distributed intelligent optimization is to achieve intelligent network management, enabling adaptive and self-recovery capabilities to further enhance the user experience.

[0102] Dynamic resource orchestration:

[0103] Dynamic resource orchestration, based on AI predictions and real-time network status, flexibly allocates computing, spectrum, and storage resources to ensure end-to-end service quality on demand. Its core goal is to break the static resource allocation model, supporting elastic expansion, energy-saving optimization, and intelligent scheduling and allocation of network resources to achieve optimal performance and service quality. By using AI to predict user needs (such as the activity patterns of elderly users), cloud-edge-end computing power is dynamically scheduled. For example, local GPU resources are prioritized during the morning health monitoring peak. This enables dedicated resource allocation for different business scenarios, ensuring consistent service quality. This functionality is continuously evolving with the application of AI and machine learning technologies, enabling the system to predict resource needs and allocate resources more intelligently. The ultimate goal of dynamic resource orchestration is to provide a flexible and scalable network environment that can quickly respond to market and technological changes.

[0104] Synaesthesia Integration:

[0105] By sharing spectrum and hardware resources, communication and perception functions can be coordinated within the same system. Multi-base station joint perception analyzes perception data and adjusts network resources in a timely manner to optimize user experience. Multi-base station joint perception: This fusion technology uses communication signals (such as 5G radio waveforms) through collaborative networking and data sharing, simultaneously enabling target detection, positioning, tracking, and environmental reconstruction. Its core approach is to reuse the hardware and spectrum resources of traditional communication base stations as perception nodes, forming a distributed perception network. For example, multi-base station joint perception uses a hybrid transmit-receive mode (base station A transmits signals, base stations B / C receive reflected signals) to build a distributed radar perception network, reducing perception coverage blind spots by 80%. This intelligent processing approach significantly enhances the network's ability to respond to emergencies and promotes efficient resource utilization. The ultimate goal of inter-sensory integration is to create a responsive, resource-efficient intelligent network that can provide users with safe and convenient services while promoting broader social and economic activities.

[0106] Immersive Communications:

[0107] Immersive communication is a form of communication that uses technologies such as AR (augmented reality), VR (virtual reality), and mixed reality to provide users with a highly interactive and immersive experience. With the development of 5G and future networks, immersive communication has garnered widespread attention. Its core goal is to enhance the realism and engagement of users in virtual environments, making them as natural as face-to-face communication.

[0108] Reference Figure 2 The interaction process shown in the figure is based on the premise that the 5G network and IMS services are operating normally. It specifically includes the following steps:

[0109] Step S01: Device access and registration. The 3GPP-defined IMS (IP Multimedia Subsystem) enables unified registration and authentication of multiple devices (such as smart glasses, wearables, and home robots). The device sends a registration request to the IMS core network via the SIP protocol, and the S-CSCF (Serving Call Session Control Function) completes authentication and service triggering.

[0110] Step S02: Network slicing and QoS assurance: Based on service requirements (such as real-time video and AR rendering), 5QI is dynamically allocated to higher-priority network slices to ensure low latency and high reliability. Through 3GPP-WLAN interconnection authentication optimization (such as the adaptive K selection mechanism), seamless switching between cellular networks and local Wi-Fi is achieved, authentication signaling overhead is reduced, and heterogeneous network convergence is realized.

[0111] Step S03: Intelligent computing optimization and resource allocation are achieved through federated learning collaborative training, dynamic computing power scheduling, and energy consumption optimization. Federated learning collaborative training: Based on the distributed intelligent architecture (TS23.288) defined in 3GPP R18, federated learning is performed on edge nodes and the cloud. Lightweight models are deployed on edge nodes, and only gradients, not raw data, are uploaded. For example, the dynamic member selection algorithm (DPS) is used to optimize model training efficiency and improve gesture recognition accuracy. Dynamic computing power scheduling: Through native protocol stack fusion design and distributed intelligent collaboration, the fragmentation between modules is completely eliminated to achieve global resource optimization. For example, the distributed intelligent agent (DIA) replaces the NWDAF, deeply embeds computing power scheduling functions into the network protocol stack, builds local decision models based on federated learning, and reduces cross-domain signaling interactions. For example, in high-density AR scenarios, the local GPU cluster is prioritized. Energy consumption optimization: Through lightweight AI model parameters (such as pruning and quantization) and transmission protocol optimization, device-side energy consumption is reduced.

[0112] Step S04: Deploy a multimodal fusion server on the IMS media plane to integrate voice, visual, and sensor data. For example, synchronized voice and gesture control is achieved through the voice burst control protocol. Voice processing: Users interact with the system through voice commands. The device's built-in high-performance microphone captures ambient sound and uses noise reduction technology to clearly receive user commands. Voice commands are parsed by the natural language processing module to understand user intent, such as initiating a video call or querying information. Gesture recognition: The device captures and analyzes user gestures using a camera or dedicated sensors. The touchless interaction feature allows users to flexibly use the system in different environments. Gestures such as waving, clicking, and swiping are instantly converted into control commands, such as zooming the virtual screen or switching applications. Motion recognition: Utilizing the accelerometer and gyroscope, the system can identify dynamic user behavior, such as standing up, sitting down, and walking. Simultaneously confirming the user's status, for example, determining whether smart reminders are needed during an activity, can also trigger the device's automatic adjustment function. Virtual and real-world mapping: Based on the 3GPP R18 AR call standard, AR media stream channels (5QI=76) are configured through SIP signaling to achieve dynamic superposition of virtual images and physical scenes.

[0113] Step S05: Deploy a lightweight rendering engine on the terminal side, combined with network transmission QoS guarantees, to achieve real-time content generation, low-latency rendering, and projection. For example, through devices such as smart glasses and smart watches combined with 5G networks, users can engage in real-time voice and video interaction in an immersive communication environment, while enjoying features such as real-time translation and the fusion of virtual and physical reality.

[0114] This application provides a native framework fusion: native fusion architecture refers to the deep integration of communication, perception, computing and intelligent functions into the network protocol stack, and the dynamic loading and coordinated scheduling of multiple factors through a unified hardware and software platform. Its core is to break the boundaries of traditional technologies and avoid the superposition of "plug-in" functions, thereby optimizing system efficiency and resource utilization.

[0115] This application provides distributed intelligent optimization: Distributed intelligent optimization uses edge computing and federated learning technology to achieve multi-node collaborative real-time data processing and model training while protecting data privacy, aiming to reduce latency and improve system robustness.

[0116] This application provides dynamic resource orchestration: Dynamic resource orchestration is based on AI prediction and real-time network status, flexibly allocating computing, spectrum and storage resources to achieve on-demand guarantee of end-to-end service quality. Its core is to break the static resource allocation model and support elastic expansion and energy-saving optimization.

[0117] This application provides synergy integration: by sharing spectrum and hardware resources, communication and perception functions can be coordinated in the same system.

[0118] This application provides immersive communication: a communication method that provides users with a highly interactive and immersive experience through technologies such as AR (augmented reality), VR (virtual reality) and mixed reality.

[0119] This application achieves end-to-end efficiency improvements: the native integration of synaesthesia and computing intelligence reduces protocol conversion overhead and reduces latency. This application achieves resource utilization optimization: dynamic computing power scheduling combined with federated learning improves computing power utilization while reducing device energy consumption. This application achieves user experience enhancement: the accuracy of intent recognition in multimodal interactions is improved, supporting natural operation for the elderly (such as automatic correction of accidental gestures).

[0120] In other embodiments, distributed computing based on blockchain can be used: computing power resources are managed through a decentralized ledger, but this may increase communication overhead. Heterogeneous network collaboration can be achieved: satellite networks can be used to supplement ground coverage, but satellite-to-ground protocol compatibility issues need to be resolved. Lightweight AI models can be deployed end-to-end: small AI models can be run directly on terminal devices, sacrificing some accuracy in exchange for low latency, which is suitable for resource-constrained scenarios. It should be noted that blockchain solutions are difficult to meet real-time requirements; satellite networks are expensive and lack standardization; and end-side models have limited accuracy and need to be updated frequently.

[0121] In summary, the method provided by the embodiment of the present application includes: accessing and registering multiple devices through the IMS defined by 3GPP; determining event priorities and allocating network slices based on the identifier of the quality of service flow; allocating network resources through federated learning collaborative training and dynamic computing power scheduling; deploying a multimodal fusion server on the IMS media plane to integrate multimodal data and interact and display data with the rendering engine deployed on the terminal. The embodiment of the present application dynamically adjusts network slices and resources, which is conducive to improving the efficiency of data interaction.

[0122] Secondly, refer to the attached Figure 3 A multimodal data interaction system according to an embodiment of the present invention is described, and the system specifically includes:

[0123] The first module 310 is used to access and register a number of devices through the IMS defined by 3GPP;

[0124] The second module 320 is configured to determine an event priority and allocate a network slice based on an identifier of the quality of service flow;

[0125] The third module 330 is used to allocate network resources through federated learning collaborative training, dynamic scheduling of computing power;

[0126] The fourth module 340 is used to deploy a multimodal fusion server on the IMS media plane, integrate multimodal data, and interact and display data with the rendering engine deployed on the terminal.

[0127] It can be seen that the contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0128] Reference Figure 4 , an embodiment of the present invention provides a multimodal data interaction device, comprising:

[0129] at least one processor 410;

[0130] at least one memory 420, for storing at least one program;

[0131] When the at least one program is executed by the at least one processor 410, the at least one processor 410 implements the multimodal data interaction method.

[0132] Similarly, the contents of the above method embodiments are applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0133] An embodiment of the present invention further provides a computer-readable storage medium storing a program executable by a processor. The program executable by the processor is used to execute the above-mentioned multimodal data interaction method when executed by the processor.

[0134] Similarly, the contents of the above method embodiments are applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0135] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operation and logic flow presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.

[0136] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the ordinary skill of an engineer. Therefore, a person skilled in the art will be able to implement the present invention set forth in the claims using ordinary skill without undue experimentation. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0137] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several programs for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0138] The logic and / or steps represented in a flowchart or otherwise described herein, for example, may be considered as an ordered list of executable programs for implementing the logical functions, and may be embodied in any computer-readable medium for use by, or in conjunction with, a program execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can retrieve and execute a program from a program execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" may be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, a program execution system, apparatus, or device.

[0139] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0140] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable program execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0141] In the above description of this specification, reference to the terms "one embodiment / example," "another embodiment / example," or "certain embodiments / examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0142] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

[0143] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.

Claims

1. A multimodal data interaction method, characterized in that: The following steps are involved: Access and register several devices through the IMS defined by 3GPP; Determine event priority and allocate network slices based on the identifier of the quality of service flow; Allocate network resources through federated learning collaborative training and dynamic computing power scheduling; A multimodal fusion server is deployed on the IMS media plane to integrate multimodal data and interact and display data with the rendering engine deployed on the terminal.

2. The multimodal data interaction method according to claim 1, characterized in that: The aforementioned collaborative training through federated learning, dynamic scheduling of computing power, and allocation of network resources include: Based on the distributed intelligent architecture defined by 3GPP, federated learning is performed on edge nodes and the cloud, and lightweight models are deployed on edge nodes; Embed computing power scheduling functions into the network protocol stack through distributed intelligent agents, and build local decision-making models based on federated learning; Network resources are allocated through the lightweight model and the local decision model.

3. The multimodal data interaction method according to claim 2, characterized in that: The multimodal fusion server is deployed on the IMS media plane to integrate multimodal data, including: Integrate voice data and gesture data through voice burst control protocol; Through several sensors, it integrates and identifies the user's dynamic behavior and confirms the user's status; Based on the 3GPP AR call standard, AR media stream channels are configured through SIP signaling to integrate virtual images with physical scenes.

4. The multimodal data interaction method according to claim 1, characterized in that: The IMS defined by 3GPP is used to access and register several devices, including: Receiving registration requests sent by the plurality of devices via the SIP protocol; Authentication and service triggering are performed through the service call session control function.

5. The multimodal data interaction method according to claim 1, characterized in that: The method further comprises: Establishing a native fusion architecture to integrate functions related to the multimodal data into the network protocol stack; Provide dynamic loading and collaborative scheduling of multi-factor capabilities through hardware and software platforms, and connect the data of the devices.

6. The multimodal data interaction method according to claim 1, characterized in that: The method performs collaborative training through the following steps: The first edge node is trained based on historical data to obtain a trained lightweight model; According to the trained lightweight model, the model weights are exchanged with the adjacent second edge node to obtain an updated lightweight model.

7. The multimodal data interaction method according to claim 1, characterized in that: The method further comprises: Reuse the hardware and spectrum resources of communication base stations into sensing nodes to form a distributed sensing network; Collaborative networking and data sharing are carried out through the distributed sensing network.

8. A multimodal data interaction system, characterized in that: include: The first module is used to access and register several devices through the IMS defined by 3GPP; The second module is used to determine the event priority and allocate network slices based on the identifier of the quality of service flow; The third module is used to allocate network resources through federated learning collaborative training, dynamic computing power scheduling; The fourth module is used to deploy a multimodal fusion server on the IMS media plane, integrate multimodal data, and interact and display data with the rendering engine deployed on the terminal.

9. A multimodal data interaction device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the multimodal data interaction method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to implement the multimodal data interaction method according to any one of claims 1 to 7 when executed by the processor.

Citation Information

Patent Citations

  • Immersive multimedia service control system and method, electronic equipment and storage medium

    CN116418789A

  • 6G communication network architecture generation method and system, electronic equipment and medium

    CN117749407A

  • Cross-modal knowledge fusion calculation method based on federated learning and big and small model collaboration

    CN118568666A

  • Internet of Things gateway control method fusing local communication and AI computing power

    CN119324848A

  • Extended reality (XR) enhancement

    WO2024097065A1