Method and apparatus for updating artificial intelligence model in wireless communication system

The method and device address the challenge of updating AI models in wireless communication systems by adapting to context changes and optimizing performance through efficient model configuration and transfer in 6G networks.

WO2025183448A1PCT designated stage Publication Date: 2025-09-04SAMSUNG ELECTRONICS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/002671
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-26
Filing Date
2025-02-26
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing wireless communication systems face challenges in efficiently updating and optimizing artificial intelligence models for diverse user needs and context changes, particularly in the transition from 5G to 6G networks, where terahertz band signals require advanced coverage and connectivity technologies.

Method used

A method and device for updating AI models in wireless communication systems that involve identifying context changes, requesting and receiving optimized AI models from a service provider, and configuring or combining existing models to ensure efficient inference performance.

Benefits of technology

Enables seamless adaptation of AI models to changing contexts, optimizing inference performance and reducing data transfer inefficiencies by utilizing existing models and minimizing redundant data transfers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025002671_04092025_PF_FP_ABST
    Figure KR2025002671_04092025_PF_FP_ABST
Patent Text Reader

Abstract

In one embodiment of the present disclosure, a method by which a device operates in a wireless communication system is provided. The method may comprise the steps of: identifying a context change of input data; transmitting, to a service provider, a request message that requests an artificial intelligence model for inferring of the input data; receiving, from the service provider in response to the request message, information related to the artificial intelligence model for inferring of a changed context; identifying the artificial intelligence model on the basis of the received information; and inferring the changed context through the artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for updating artificial intelligence models in wireless communication systems

[0001] The present invention relates to a method and device for updating an artificial intelligence (AI) model in a wireless communication system.

[0002] Looking back at the evolution of wireless communication over successive generations, technologies have primarily been developed for human-facing services such as voice, multimedia, and data. With the commercialization of 5G (5th-generation) communication systems, an explosive increase in connected devices is expected to be connected to communication networks. Examples of networked objects include vehicles, robots, drones, home appliances, displays, smart sensors installed in various infrastructures, construction equipment, and factory equipment. Mobile devices are expected to evolve into diverse form factors, including augmented reality glasses, virtual reality headsets, and holographic devices. In the 6th-generation (6G) era, efforts are being made to develop improved 6G communication systems to connect hundreds of billions of devices and objects and provide diverse services. For this reason, 6G communication systems are often referred to as "beyond 5G."

[0003] The 6G communication system, expected to be realized around 2030, will have a maximum transmission speed of terabytes per second (i.e., 1,000 gigabits per second) and a wireless latency of 100 microseconds (μsec). In other words, compared to 5G, the transmission speed in a 6G communication system will be 50 times faster, while the wireless latency will be reduced to one-tenth.

[0004] To achieve these high data rates and ultra-low latency, 6G communication systems are being considered for implementation in the terahertz band (e.g., from 95 gigahertz (GHz) to 3 terahertz (THz)). Compared to the millimeter wave (mmWave) band introduced in 5G, the terahertz band is expected to experience more severe path loss and atmospheric absorption, making it more crucial to ensure signal reach, or coverage, in this band. Key technologies to ensure coverage include radio frequency (RF) components, antennas, new waveforms that offer better coverage than OFDM (orthogonal frequency division multiplexing), beamforming, and multiple antenna transmission technologies such as massive multiple-input and multiple-output (MIMO), full-dimensional MIMO (FD-MIMO), array antennas, and large-scale antennas. In addition, new technologies such as metamaterial-based lenses and antennas, high-dimensional spatial multiplexing using orbital angular momentum (OAM), and reconfigurable intelligent surfaces (RIS) are being discussed to improve the coverage of terahertz band signals.

[0005] In addition, in order to improve frequency efficiency and system network, 6G communication systems are developing full duplex technology that utilizes the same frequency resources for uplink and downlink at the same time; network technology that integrates satellites and high-altitude platform stations (HAPS); network structure innovation technology that supports mobile base stations and enables optimization and automation of network operation; dynamic spectrum sharing technology through collision avoidance based on spectrum usage prediction; AI-based communication technology that utilizes artificial intelligence (AI) from the design stage and internalizes end-to-end AI support functions to realize system optimization; and next-generation distributed computing technology that realizes services with complexity that exceeds the limits of terminal computing capabilities by utilizing ultra-high-performance communication and computing resources (mobile edge computing (MEC), cloud, etc.). In addition, efforts are being made to further strengthen connectivity between devices, further optimize networks, promote softwareization of network entities, and increase the openness of wireless communications through the design of new protocols to be used in 6G communication systems, the implementation of hardware-based security environments, the development of mechanisms for the safe use of data, and the development of technologies for maintaining privacy.

[0006] Research and development of these 6G communication systems are expected to enable a new level of hyper-connected experience through the hyper-connectivity of 6G communication systems, which encompass not only connections between things but also connections between people and things. Specifically, 6G communication systems are expected to enable services such as truly immersive extended reality (Truly Immersive XR), high-fidelity mobile holograms, and digital replicas. Furthermore, services such as remote surgery, industrial automation, and emergency response, which are provided through enhanced security and reliability, will find application in diverse fields such as industry, healthcare, automotive, and home appliances.

[0007] One embodiment of the present disclosure can provide a method and device for effectively providing a service in a wireless communication system.

[0008] A method for a terminal to operate in a wireless communication system disclosed as a technical means for achieving the above-described technical task may include a step of identifying a change in the context of input data, a step of transmitting a request message requesting an artificial intelligence model for inferring the input data to a service provider, a step of receiving information related to the artificial intelligence model for inferring the changed context from the service provider in response to the request message, a step of identifying the artificial intelligence model based on the received information, and a step of performing inference of the changed context through the artificial intelligence model.

[0009] A method for a service provider to operate in a wireless communication system, disclosed as a technical means for achieving the above-described technical task, may include a step of receiving a request message requesting an artificial intelligence model for inferring input data from a terminal, a step of identifying an artificial intelligence model for inferring input data, and a step of transmitting information related to the artificial intelligence model for inferring input data to the terminal.

[0010] A terminal operating in a wireless communication system disclosed as a technical means for achieving the above-described technical task may include a transceiver and at least one processor. The at least one processor may identify a change in the context of input data, transmit a request message requesting an artificial intelligence model for inferring the input data to a service provider through the transceiver, receive information related to the artificial intelligence model for inferring the changed context from the service provider in response to the request message through the transceiver, identify the artificial intelligence model based on the received information, and perform inference of the changed context through the artificial intelligence model.

[0011] A service provider operating in a wireless communication system disclosed as a technical means for achieving the above-described technical task may include a transceiver and at least one processor. The at least one processor may receive a request message requesting an artificial intelligence model for inferring input data from a terminal via the transceiver, identify the artificial intelligence model for inferring input data, and transmit information related to the artificial intelligence model for inferring input data to the terminal via the transceiver.

[0012] FIG. 1 is a flowchart illustrating a process of a service provider and a terminal according to one embodiment of the present disclosure.

[0013] FIG. 2 is a diagram illustrating a network entity according to one embodiment of the present disclosure.

[0014] In describing the embodiments herein, descriptions of technical details that are well-known in the technical field to which the present invention pertains and are not directly related to the present invention will be omitted. This is to avoid obscuring the gist of the present invention by omitting unnecessary explanations and to convey the gist more clearly.

[0015] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below, along with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below and may be implemented in various different forms.

[0016] The terms used in the following description to identify connection nodes, terms referring to network entities, terms referring to messages, terms referring to interfaces between network entities, terms referring to various identification information, etc. are provided as examples for convenience of explanation. Therefore, the present disclosure is not limited to the terms described below, and other terms referring to objects with equivalent technical meanings may be used.

[0017] For convenience of explanation, this disclosure uses terms and names defined in the 3rd Generation Partnership Project New Radio (3GPP NR) standards. However, this disclosure is not limited to these terms and names and can be equally applied to systems conforming to other standards. In this disclosure, a base station may refer to a gNB. Furthermore, the term "terminal" may refer to mobile phones, NB-IoT devices, sensors, and other wireless communication devices.

[0018] Hereinafter, a base station is an entity that performs resource allocation for a terminal, and may be at least one of a gNode B, an eNode B, a Node B, a BS (Base Station), a wireless access unit, a base station controller, or a node on a network. The terminal may include a UE (User Equipment), an MS (Mobile Station), a cellular phone, a smartphone, a computer, or a multimedia system capable of performing a communication function. Of course, the above examples are not limited thereto.

[0019] This disclosure can be applied to 3GPP NR (the 5th generation mobile communications standard). Furthermore, this disclosure can be applied to intelligent services (e.g., smart homes, smart buildings, smart cities, smart or connected cars, healthcare, digital education, retail, security, and safety-related services) based on 5G communication technology and IoT-related technologies.

[0020] Wireless communication systems are evolving from providing voice-oriented services in the early days to broadband wireless communication systems that provide high-speed, high-quality packet data services, such as communication standards such as 3GPP's HSPA (High Speed ​​Packet Access), LTE (Long Term Evolution or E-UTRA (Evolved Universal Terrestrial Radio Access)), LTE-Advanced (LTE-A), LTE-Pro, 3GPP2's HRPD (High Rate Packet Data), UMB (Ultra Mobile Broadband), and IEEE's 802.16e.

[0021] As a representative example of a broadband wireless communication system, the LTE system adopts the Orthogonal Frequency Division Multiplexing (OFDM) method in the downlink (DL) and the Single Carrier Frequency Division Multiple Access (SC-FDMA) method in the uplink (UL). The uplink refers to a wireless link in which a user equipment (UE, or mobile station, MS) transmits data or control signals to a base station (eNode B or base station, BS), and the downlink refers to a wireless link in which a base station transmits data or control signals to a user equipment (UE). The above multiple access method distinguishes the data or control information of each user by allocating and operating the time-frequency resources for transmitting data or control information to each user so that they do not overlap, that is, so that orthogonality is achieved.

[0022] As the future communications system beyond LTE, 5G communication systems must be able to freely reflect the diverse needs of users and service providers. Therefore, they must support services that simultaneously satisfy these diverse requirements. Services being considered for 5G communication systems include Enhanced Mobile Broadband (eMBB), Massive Machine Type Communication (mMTC), and Ultra Reliability Low Latency Communication (URLLC).

[0023] Furthermore, while embodiments of the present invention will be described below using LTE, LTE-A, LTE Pro, or 5G (or NR, next-generation mobile communication) systems as examples, the embodiments of the present invention may also be applied to other communication systems with similar technical backgrounds or channel types. Furthermore, the embodiments of the present invention may be applied to other communication systems with some modifications, as determined by a person skilled in the art, without significantly departing from the scope of the present invention.

[0024] In the following description of the present disclosure, detailed descriptions of related known functions or configurations will be omitted if they are deemed to unnecessarily obscure the gist of the present disclosure. Hereinafter, embodiments of the present disclosure will be described with reference to the attached drawings.

[0025] Artificial intelligence (AI) and machine learning (ML) technologies (hereinafter collectively referred to as AI / ML) are being introduced and generalized in media-related applications, from traditional applications such as image classification and voice / facial recognition to cutting-edge applications such as video quality enhancement.

[0026] As research in this area matures, advanced applications requiring processing larger amounts of data are being considered. Research is ongoing to incorporate AI architectures for media applications that collaborate with servers in the network as well as terminals into 5G systems. For example, several use cases are being discussed, including object recognition in images and videos, quality enhancement of streaming video, and natural language processing for speech.

[0027] Scenarios for AI inference and learning include providing pre-trained AI models from the network to the UE, and splitting inference on the UE or between the UE and the network.

[0028] The structure of an AI model consists of layers, layers contain nodes, and nodes contain functions and weight values.

[0029] The AI ​​model can be identified by its composition of layers and nodes, and the weight values ​​can vary depending on the data used to train the model and the number of learning cycles.

[0030] The size of the data used to store weight values ​​is proportional to accuracy and processing performance. That is, while larger data sizes lead to higher accuracy, they also require more processing power.

[0031] The architecture for AI-enabled scenarios combines several core components, including an AI model repository, AI model provisioning capabilities, AI model access capabilities, AI model inference engines, and intermediate data provision capabilities, to efficiently and effectively provide AI models and related data over 5G networks.

[0032] AI services provided by AI service providers via 5G networks may be comprised of a step of determining which service to provide and receive based on communication between the application of the service provider server and the application of the terminal, a step of determining an AI model suitable for the service, and a step of determining a version (or variant) of the AI ​​model suitable for operation on the terminal (e.g., a high-accuracy model with high bit depth, a low-accuracy model with quantized bits, etc.).

[0033] In some embodiments, a single AI model may not be suitable for all general-purpose use. Therefore, to receive optimal service, it is necessary to receive an AI model optimized for the intended use from the service provider.

[0034] According to one embodiment of the present disclosure, a method for communicating between a terminal and a server to select an optimal AI model that matches a desired result of a user can be provided, and information to be exchanged between the terminal and the server can be standardized.

[0035] In one embodiment, a different version of an AI model, with only some differences from the target AI model, may already exist on the terminal. Even if only a portion of the target AI model is transferred, it may be usable in combination with portions of other AI models already present on the terminal. Since the selection of AI models cannot be determined arbitrarily on the terminal, the terminal's AI model holding status may be provided to the server and analyzed by the server to facilitate more efficient transfer operations.

[0036] A wireless communication system according to one embodiment of the present disclosure may be configured with an AI model (hereinafter, “model”) capable of performing inference according to a data context from input data, a service provider having models capable of inferring various data contexts and capable of providing the models to a service subscriber, and a terminal capable of receiving at least one model from the service provider and performing inference according to a data context using the model.

[0037] In one embodiment, the terminal can directly determine the target or context when the user's inference target changes using the first model and request a new model from the service provider. Alternatively, the terminal can preprocess the acquired input (e.g., image, audio, or other sensing information) into a format requested by the service provider and then transmit it to the service provider. The service provider can then select an appropriate second or third model and notify the terminal.

[0038] According to one embodiment of the present disclosure, a method for a terminal to update an AI model may include a step of determining a changed context when a target that a user was inferring using a first model changes or changes, a step of selecting an appropriate second model capable of inferring the changed context, and a step of receiving the second model. When a service provider provides the second model to the terminal, a method may be provided for the terminal to configure a second-compatible model that provides inference performance close to that of the second model based on information of the first model or at least one third model that the terminal has already received and possesses.

[0039] Each artificial intelligence model can have the following purposes:

[0040] Model 1: For first context inference purposes

[0041] Model 2: For second context inference

[0042] Third Model: For Third Context Inference

[0043] Second-compatible model: For use in providing inference performance equal to or close to that of the second model, based on parts of the first or third model and parts of the second model.

[0044] In this disclosure, context can refer to a unit where the connection and context between information intended and received by a user are maintained. For example, if the type of flower is one context, images of various flowers can be one context. If the type of car is another context, images of various cars can be one context. For example, since a model that determines the type of flower is not suitable for use in determining the type of car, the AI ​​model used by the terminal can be changed to a model that can determine the context of the changed object when the object to be inferred changes.

[0045] A context change may be intentional by the user, or it may be identified when the currently used AI model fails to find an appropriate inference result for the target of inference. If the user changes the context, the changed context can be explicitly requested from the service provider, and the service provider can find the most appropriate model for inferring the changed context from its internal storage and provide it to the terminal. If the context is determined and changed by the service provider, all or part of the input data acquired by the terminal for inference may be transmitted to the service provider. The input data may be preprocessed, such as by compressing or extracting features. The service provider can decode the preprocessed input data after receiving it and determine the context contained in the input data from the decoded information. The terminal can determine whether the input data can be preprocessed in the format requested by the service provider and can report to the service provider at least one format that it can support among multiple preprocessing methods requested by the service provider. The service provider can decrypt input data preprocessed and transmitted by the terminal, determine the context of the input data, and select a second model appropriate for the changed context if the context is determined to have changed. The service provider can, if possible, instruct the terminal to construct a model that is the same as or close to the second model from models it already possesses. The service provider can also transmit the selected second model to the terminal and instruct the terminal to use it. The terminal can receive the entire second model from the service provider, or a portion of the second model and information for constructing a second-compatible model.

[0046] In one embodiment, the terminal may receive a model after establishing a transmission session between the terminal's model receiving unit and the network's model transmitting unit. The terminal may identify a second model capable of performing inference on the modified context based on all or part of the received model, or construct a second-compatible model. The terminal may use the identified or constructed model to perform inference on input data through an inference engine.

[0047] In one embodiment of the present disclosure, the configuration of an artificial intelligence model may refer to an operation of rearranging layers constituting the artificial intelligence model and combining them with at least one layer of another artificial intelligence model to create a new model.

[0048] For example, a terminal may have a first model with 100 layers for general-purpose purposes such as identifying everyday objects. If a second model with 100 layers is used for inference limited to a specific category, a second compatible model (100 layers) may be constructed by combining some layers of the first model (e.g., layers 1-90) and some layers of the second model (e.g., layers 91-100), and the second compatible model may be trained in advance. Then, when the second model is needed at the terminal, only the 10 layers to be used for constructing the second compatible model at the terminal may be transmitted and combined with the first model that the terminal already has, thereby allowing the terminal to restore the second compatible model.

[0049] Figure 1 is a flowchart illustrating a process between a service provider (network) and a terminal according to one embodiment of the present disclosure. Not all steps illustrated in Figure 1 are essential components, and some steps may be omitted in one embodiment.

[0050] In step 1, an overview of services provided by a service provider is delivered through communication between a terminal application and a network application, and services can be selected by the user.

[0051] In step 2, the terminal and the network can mutually determine whether the exchange of context and the configuration of the artificial intelligence model are supported.

[0052] In step 3, if context exchange between the terminal and the network is possible, the terminal application can generate context information from the input data. For example, the terminal can perform context preprocessing from the input data.

[0053] In step 4, the generated context information can be transmitted to the network.

[0054] In step 5, the network application can restore the preprocessed context information, use the network's inference engine to determine the context contained in the input data, and select an appropriate artificial intelligence model capable of inferring the context.

[0055] In step 6, the network application may request information about the AI ​​model already owned by the terminal from the terminal. In one embodiment, the terminal may report AI model information to the network as needed, even without a network request.

[0056] In step 7, the terminal can collect information on the artificial intelligence models it possesses and generate a report.

[0057] In step 8, the terminal can report information about the models it has to the network application.

[0058] In step 9, the network application can query the model store to collect a list of models that can infer context.

[0059] In step 10, the network application and model repository can generate instructions on how to construct an optimal artificial intelligence model or a compatible model for context inference based on the list of models held by the terminal.

[0060] In step 11, the generated instructions can be passed from the network application to the model repository, and from the model repository to the model transmission unit.

[0061] In step 12, the network application can convey an instruction method to the terminal application.

[0062] In step 13, the terminal application can instruct the model receiver to establish a model transmission session to receive the artificial intelligence model.

[0063] In step 14, the model receiving unit of the terminal can communicate with the model transmitting unit of the network to establish a model transmission session.

[0064] In step 15, the model receiver can receive at least a portion of the model specified in the received instruction.

[0065] In step 16, the model receiving unit can generate a new model by combining at least a part of the received model with the existing model according to the received instructions.

[0066] In step 17, the model receiver can transmit the generated new model to the terminal's inference engine.

[0067] At step 18, new inferences can be made from the input data.

[0068] Hereinafter, the following will be specifically described with respect to a method of using an AI model for inference by replacing it according to the context of input data in a terminal according to one embodiment of the present disclosure.

[0069] - How to request context

[0070] - Terminal reporting information

[0071] - Server judgment procedure, method of instructing combination model configuration

[0072] How to request context

[0073] In one embodiment, a message may be used to request a context exchanged between a service provider (network, server) and a terminal. The message may include context format systems supported by the service provider and context format identifiers included thereunder. The context format system may be identified by a context format system identifier for identifying it. For example, if the context format system identifier is urn:3gpp:ai4media:context:cset=15, the context format identifier is one of the context format system identifiers defined as a context supported by the 3GPP ai4media specification, and the version of the context set may be understood to be 15. A higher version context set indicated by a higher number means that it includes additional contexts compared to a lower version context set with a lower number. If a terminal has searched and requested a context based on a lower version context set, the service provider may propose a higher version context set, and the terminal may search and request a new context after receiving the higher version context system.

[0074] A service application may provide a list of context format system identifiers that the service application can support, upon request at the time of service initiation or upon request by the terminal's needs.

[0075] The terminal application can receive a list of available context format identifiers from the context format system identifier received from the service application and request the service application to transmit a model that supports each or more of the contexts included in the list.

[0076] To enable terminals and their users to search for desired contexts, context format identifiers can provide human-readable text-based context format descriptors. For example, a user can select a desired context by reading it from the descriptor, and the terminal application can request the context format identifier corresponding to the selected descriptor from the service application.

[0077] The components of a context request message can be represented as shown in Table 1 below:

[0078] Name Meaning Context Type Scheme Identifier The type of the declared context, such as a URN (e.g., urn:3gpp:ai4media:context:cset=15) Context Type Descriptor A human-readable text identifier for the context (e.g., Flower) Context Type Identifier An explicitly requested context (e.g., Flower, Car, Facial Expression) Context Profile Descriptor A human-readable text profile of the context (e.g., Korean flowers) Context Profile Identifier A type or scope within the context (e.g., Korean flowers, Flowers of the World)

[0079] The Service Requirement Information (SRI) in 3GPP TR 26.927 includes information about the type of service provided by the service provider. Context-related information can be represented as shown in Table 2 below.

[0080] Metadata categoryMetadata typeDefinitionMetadata type description (Examples)Service requirement informationMaximum service inference latencyThe maximum inference latency requirement specified for a given AI media service, in milliseconds. In the case of split inferencing, this requirement includes the delivery latency of the intermediate data between the first and second split inference entities.100msMinimum service inference accuracyThe minimum accuracy specified for a given AI media service.80%Service type identifierAn identifier for the service type to be supported by the AI / ML model, such as ASR (Automatic Speech Recognition), TTS (Text To Speech), Translation (with the indication of input and output languages).TTS, ASR, Trans-EN-to-ZHContext type scheme identifier (컨텍스트 형식 체계 식별자)An identifier specifying a list of supportable contexts by the AI / ML model. A URN is used as the identifier of the list and the corresponding list is managed by service provider.URNContext Schemestructured information of context scheme. It may include multiple context_types where each context_type has multiple context_profiles.structured textService accuracyThe expected service accuracy85%.

[0081] Table 2: Service requirement information

[0082] The composition of the context system expressed in the form of a structure can be represented as shown in Table 3 below. Context_scheme can contain multiple context_types (identifiers and descriptors), and context_type can contain multiple context_profiles (identifiers and descriptors).

[0083] metadata:context_schemecontext_scheme_idenfier: stringcontext_type{context_type_idenfier: valuecontext_type_descriptor: stringcontext_profile{context_profile_identifier: valuecontext_profile_descriptor: string}}

[0084] In one embodiment, the context of input data may not be explicitly indicated by the user, but rather determined by the service provider. To this end, the service provider may instruct the terminal application to transmit all or part of the input data, either as is or after preprocessing. The service application may use the inference engine contained within the server to infer the context contained in the input data received from the terminal.

[0085] A service application can select and instruct the type of preprocessing that a terminal application should perform. The context preprocessing system identifier can indicate the identifier of a list system containing context preprocessing methods that are requested to be performed on input data in the terminal, and the context preprocessing method can be a method belonging to the list system. Preprocessing can be performed on all or part of the input data, and in this case, identifiers representing one or more preprocessing methods and attribute values ​​describing the identifiers can be provided to the terminal.

[0086] During the process of negotiating a context preprocessing scheme identifier with a service application, the terminal application may select a supportable context preprocessing scheme identifier from among the contexts included in the system and report it to the service application, or report that no supportable context preprocessing scheme identifier exists. If the terminal reports to the service application that there is no supportable context preprocessing method, the service application may present a different context preprocessing scheme identifier to the terminal.

[0087] After negotiation, the terminal application can perform preprocessing on the input data. Preprocessing can be performed on all or part of the input data, depending on the preprocessing-specific identifier and attribute values. The preprocessed input data transmitted from the terminal to the server may include the type of preprocessing, the media time corresponding to each input data processing unit, and the preprocessing start and completion time.

[0088] The terminal application can transmit preprocessed context information, preprocessed information, and the transmission time of the preprocessed information to the service application. The service application can then pass the received context information to the network's inference engine to infer and determine the context. The network's inference engine can then derive the context from the context information and send it back to the service application.

[0089] The service application can select one or more AI models by judging the context based on the input data media time. The service application can adjust the number of transmissions per second, resolution, and preprocessing file size of the preprocessing information to be processed by the terminal application by considering the interval between preprocessing start times. The service application can determine the preprocessing transmission time using the transmission time when the terminal transmits the preprocessing information, the reception time when the preprocessing information is received from the network, and the size of the preprocessing information. In addition, the service application can adjust the terminal's transmission volume per second and the number of transmissions per second by considering factors such as network resource management and QoS allocated to the terminal. When a changed preprocessing policy is received from the service application, the terminal can change the type of preprocessing, the number of executions, etc. according to the changed policy.

[0090] The service application can present different AI models for different media times within the preprocessed context information reported by the terminal. That is, the first model can be presented for the first media time by judging it as the first context, and the second model can be presented for the second media time by judging it as the second context. The response of the service application can be delivered for each inference, for each media time unit, or by bundling multiple inferences within a certain time interval. Therefore, when the terminal generates and transmits context information for multiple media times, the service application responds by indicating the use of multiple AI models, and at this time, one or more different models can be specified for each media time.

[0091] The components of a message related to context preprocessing can be represented as shown in Table 4 below:

[0092] Name MeaningContext Preprocessing Scheme Identifier A scheme that defines preprocessing methods that can be used for context preprocessing on the terminal (e.g. urn:3gpp:ai4media:context-preprocessings)Context Preprocessing Method Identifier An identifier for a method to preprocess input data on the terminal. (e.g. image compression, feature extraction, outline extraction, texture extraction, embedding, execution of the first n layers of a specific AI model, etc.)Preprocessing Control Parameter Identifier An item that limits preprocessing on the terminal. (e.g. resolution, preprocessing file size, transfers per second, transfers per second, color space range (RGB, Luma brightness))Input Data Media Timestamp Indicates the media time of the input data.Preprocessing Start Timestamp Indicates the time when preprocessing started on the terminal.Preprocessing Transmission Timestamp Indicates the time when the terminal started transmitting context information after preprocessing.

[0093] Endpoint capability information in 3GPP TR 26.927 may include functional and performance requirements for the terminal or network's inference engine. Information related to context preprocessing can be represented as shown in Table 5 below.

[0094] Metadata categoryMetadata typeDefinitionMetadata type description (Examples)Endpoint capability informationProcessing capabilitiesThe available resources for processing AI / ML model including the computational power (in FLOPS), the memory to store model parameters and perform the inference.NPU 10TFLOPS, MEM 10GBSupported AI FrameworkThe AI framework(s) supported by the endpoint.TensorFlow 2.0Supported compression algorithmsThe supported compression algorithm(s) for intermediate data compression.NONE, FC_VCM, SNAPPY, ...Context preprocessing scheme identifierAn identifier specifying a list of supportable context preprocessing methods by UE or Network. A URN is used as the identifier of the list and the corresponding list is managed by service provider.URNContext preprocessing method identifierA list of identifiers on preprocessing methods those are supported.Video compression, feature extraction, edge detection, embedding, split inferencing of specific modelContext preprocessing control parameterPreprocessing parameter to control the size, frequency and so on.512 by 512 resolutionConnection capabilitiesThis indicates the available bandwidth in bit / s between the UE and the network for transmitting the AI ​​model and / or the intermediate data.256 kb / s.

[0095] Table 5: Endpoint capability information

[0096] Preprocessed data can be transmitted as intermediate data and may contain information related to the preprocessed input media.

[0097] Metadata categoryMetadata typeDefinitionMetadata type description (Examples)Intermediate data informationTensor structure informationThe exact underlying tensor structure of the intermediate data tensors including the exact version of it.PyTorch 2.0,Tensor flow v2.13.0, NumPy v1 .25Tensor shapeThe tensor shape(s) when the output is intermediate data. Tensor shape is a tuple of positive integers, where the size of the tuple represents the dimension of the tensor, and each value represents the size in each dimension.[1,64,64,64].Tensor element data typeThe data type of each output intermediate data tensor:int64, Float32Data directionThis defines the direction of transmitted data, either uplink (from UE endpoint to network endpoint) or downlink (From a network endpoint to the UE endpoint).This information may be useful to configure an intermediate data delivery sessionUpstream, DownstreamCompression algorithmIdentifies the compression algorithm(s) that can be applied to the intermediate data. When the connectivity condition between the UE and the network is insufficient to transmit the original intermediate data, a compression algorithm may be applied.NONE, FC_VCM, SNAPPY, ...Input data media timestampThe media playback time of input data. Server may suggest multiple different contexts per each media timestamp. UE may select one of them, or use different AI models per frame. This timestamp can be relative to media encoder or media playback timeline.timestampPreprocessing start at timestamp (T1)The timestamp when preprocessing of input data is started at. Multiple T1 can be included as many as the number of media frames. UE and server may refer same clock to generate the timestamp.timestampPreprocessing sent at timestamp (T1')The timestamp when preprocessed data is sent to server. The server may calculate the time difference between sending and receiving to suggest smaller or larger preprocessing result to be sent. UE and server may refer same clock to generate the timestamp.timestamp.

[0098] Table 6: Intermediate data information for split AI / ML operations

[0099] Terminal reporting information

[0100] In one embodiment, the service application may review a list of AI models already available on the terminal to select the optimal AI model for a context explicitly requested by the terminal or for an inferred context. To achieve this, the service application may request the terminal application to transmit a list of models, or the terminal application may voluntarily transmit the list of models to the service application.

[0101] The model list may include information such as the AI ​​model identifier and AI model version. More specifically, the AI ​​model may provide an identifier for each layer it contains and a version for that layer. The service provider can update the model layer by layer through additional training on the same model, and may also suppress updates for certain layers. For example, the values ​​for layers 1, 2, 3, and 4 may be fixed to prevent updates, while values ​​for layers 5 and above may be updated based on training. In this case, the identifier and version of the model may be fixed or changed for each layer.

[0102] After receiving information about the model held by the terminal, the service application can determine the suitability of the model, i.e., the accuracy of inference for the context, determine whether there is a new version of the same model in the network's model storage, and instruct a partial layer update for the model stored in the terminal accordingly.

[0103] In addition, after receiving information about the model held by the terminal, the service application can determine whether there is a model that can be utilized among the models held by the terminal to support inference about a new context, and if so, which layer of the model and which layer of a model held in the network's model storage to combine to support inference about a new context with minimal data transfer.

[0104] In one embodiment, information components that can be transmitted from the terminal in relation to a model stored in the terminal may be represented as shown in Table 7 below:

[0105] Name MeaningModel identifier: An identifier that distinguishes a model (e.g., VGG, mobileNet, etc.)Model version: Different version information for the same model (e.g., mobile version, quantization step (float / 32 / 16 / 8 bit) version, operation set version, size, etc.)Model graph: A description of the model's input, output, intermediate nodes, operations, and their parameter specifications.Model layer identifier: An identifier for each layer of the model (e.g., layer-01)Model layer version: A version for each layer of the model (e.g., layer-01-16bit-v01), hash code, etc.Model layer URL: Path information for receiving only a specific layer of the model.

[0106] 3GPP TR 26.927's AI model information for split AI / ML operations includes information for identifying each layer of an AI model. The model graph and layer-related information can be represented as shown in Table 8 below.

[0107] Metadata categoryMetadata typeDefinitionMetadata type description (Examples)Split model informationSplit pointsThe number of predefined split points at which a certain model can be divided into two for split inferencing.2Model graphThe graph of model which provides outline of the model.ONNX graphLayer identifierAn identifier of the layer in a description of a graph. UE may distinguish whether the same operation and same parameter will be processed as the service provider intended.layer-01Layer versionA version information of a layer. The version number changes with training of new input. UE may distinguish whether the same result will be calculated as the service provider intended.16bit-v1-build3322, hash codeLayer URLURL to download specific layer from networkURLSplit point informationSplit point identifierAn identifier of the split point in a description of a computing graph, may be generated by a neural network description language such as ONNX / NNEF.Identifiers must guarantee unique identification of a specific split point.Nb:10, 75 Name: Layer_10,Split point intermediate data sizeThe size of the intermediate data resulting from the give split point, in kilobytes. Intermediate data size is typically dependent on the tensor size at the given split point.1086KBSplit point numberThe number of the split point where the split occurs. The number may belong to set of identified numbers defined at the configuration stage.10Split point nameThe name of the split point where the split occurs. The name may belong to set of identified split point names defined at the configuration stage.conv2d_1234Split point flagAn information on whether to consider the split point before the split point identifier or after. The convention on whether it is before or after may be defined at the configuration stage.before, after.

[0108] 표 8: AI model information for split AI / ML operations

[0109] 서버 판단 절차, 조합 모델 구성 지시 방법

[0110] Service applications can use the main models provided by the service to index the usage frequency and reusability of each model, and train a new model combining two or more models based on these metrics. Users can change contexts frequently or request multiple contexts. However, if the terminal does not have a model capable of inferring the context, downloading begins only after determining that no model is available, resulting in service delays and inconvenience. Furthermore, AI models are typically several hundred MB in size, making it inefficient to download them for a one-time service and then delete them after use.

[0111] Since AI models are trained on common images (e.g., the ImageNet dataset), the final layer can be fine-tuned or transfer-trained to learn about a new context.

[0112] According to one embodiment of the present disclosure, a service provider or service application can first check the terminal's existing models and determine how to utilize the existing models before proposing the optimal model for the context to be provided. In other words, the service provider or service application can provide the terminal with a method for utilizing the terminal's existing models, or provide the terminal with a new model.

[0113] A service application can determine the context it wishes to provide, select one or more primary models to be provided by the service, select a fine-tuning or transfer learning layer for the primary models, and perform the tuning / transfer. The service application can then identify information about the corresponding model, the combination of the tuning / transfer layer, and the attributes of the combined model (context, processing time, accuracy, size).

[0114] Information on the identified model, combination information of models, and attribute information of the combination model are stored in a storage of a network or service provider server, and can be retrieved and provided upon request of a service application.

[0115] The combined model information may include at least one of the following information: a model identifier, which is an attribute of the model, a context identifier that is supported, performance (e.g., processing time, accuracy), model size, etc. In addition, as information on the hierarchy of the models to be combined, at least one of the identifier of the model to be imported for each model or adjustment / transfer layer, the scope of the layer to be combined (model layer scope), the size of the target model layer, and reception information (e.g., URL, etc.) may be further included.

[0116] The service application can retrieve combined model information for a context determined from preprocessing information requested or received from a terminal, determine whether a model held by the terminal can be reused, check the size of the model layer from combined layer information to be transmitted to the terminal, and check the accuracy and processing time that can be provided when combined at the terminal.

[0117] If one or more combination models are available based on the results of the inquiry and availability determination, information on one or more combination models can be selected and transmitted to the terminal based on service priority. For example, in an embodiment where the smallest data size to be transmitted has the highest priority, the selected combination model information may include the model layer with the smallest size to be transmitted to the terminal.

[0118] The elements of the combination model configuration information related to the combination model can be represented as shown in Table 9 below:

[0119] Name Meaning Combination Layer Information () - Model Identifier Identifier of the partial model to be used by partially referencing the layer - Model Version Version of the model - Model Layer Range (n, m) Indicates that the range of the layer: from layer n to layer m, or from layer identifier n to layer identifier m - Model Layer Identifier List List of identifiers for each layer of the model - Model Layer Version List List of versions for each layer of the model - Model Layer Size File size from layer n to layer m - Model Layer URL Address that can receive data corresponding to layer n to layer m

[0120] When the terminal application receives combined model configuration information in addition to the model information received from the service application, the terminal application can determine a single combined model based on the information of the model already received in the terminal, based on the priority of whether a model can be reused, the size of the layer to be received, the accuracy of the combined model, etc., and can decide to configure the model by receiving the entire model or only some layers of the model.

[0121] The terminal application may instruct the model receiving portion of the terminal to establish a transmission session with the model transmitting portion of the network for reception of all or part of the layers.

[0122] The model receiving unit of the terminal can receive model combination information decided to be used in the terminal application, identify the model layer to be received, receive the corresponding model layer range from the model layer URL, and create a new combination model by combining it with the received model.

[0123] The model receiving unit of the terminal transmits the generated combination model to the inference engine, and the inference engine can initiate a new inference.

[0124] The Common AI model information in 3GPP TR 26.927 includes general model attribute information. The combination-related information needed to create a combination model on a terminal can be represented as shown in Table 10 below.

[0125] Metadata categoryMetadata typeDefinitionMetadata type description (Examples)Model informationModel identifierAn identifier for an AI model (or variants of it) specified for a certain AI media service. The identifier may be a name, a number, a combination thereof, a hash value. The identifier is defined during the configuration stage.model_1, model_2Number of parametersTotal number of parameters in the neural network.11 millionModel sizeThe size of the AI model file in megabytes.40MBInput sizeThe maximum size of the input data supported by the AI model in kilobytes.256 KBOutput sizeThe maximum size of the output data supported by the AI model in kilobytes.256 KBAccuracyThe trained accuracy of the AI model as a percentage.85%Target inference latencyThe target inference latency specified for a given AI model in milliseconds. Such latency is measured between the input and output layers of the AI model at inference.This value is related to the service inference latency requirement of the service for which the AI model is provided, as well as the typical hardware capabilities of an entity performing the inference of the model.20msFormat / frameworkThe format or framework used to express the AI model, including its version number.Pytorch 2.0 ONNX 1.15.0Processing capabilitiesEstimated capabilities for processing the model including the computational power such as the computational cost (in FLOPS), the computational complexity (in MAC operations). It also includes the temporary memory to store model parameters.NPU 10TFLOPS, MEM 10GBModel compositionReferenced model identifierAn identifier of model which to be referenced.Model identifierModel layer identifierA list of identifiers of layers, or order number of layers of the referenced model.[1:3], which means from layer 1 to 3 OR [layer_identifer1:layer_identifier3]Model layer sizeThe size of model layer(s) to be referenced.10MB, 500KBModel layer URLURL to download specific layer from networkURLType of model compositionMethod used for model performance improvementfine_tuning OR transfer_learning.

[0126] Table 10: Common AI model information

[0127] In Table 10, the Referenced model identifier indicates the identifier of the target model that will be referenced and become the composition of the combination.

[0128] A model layer identifier can be a range of model layers or a list of model layer identifiers. The referenced model layers can be represented as a range, such as [n,m], from the nth to the mth, or each identifier for each layer can be listed.

[0129] Model layer size indicates the size of the transmitted data for each model layer or range to be referenced.

[0130] Model layer URL indicates the data address for transmission for each model layer or range to be referenced.

[0131] FIG. 2 is a diagram illustrating a network entity (200) according to one embodiment of the present disclosure. In one embodiment, the network entity (200) may correspond to the terminal (UE) or network (service provider, server, etc.) of FIG. 1 described above.

[0132] Referring to FIG. 2, the network entity (200) may be composed of a transceiver (210), a processor (220), and a memory (230). Depending on the communication method of the network entity (200) described above, the transceiver (210), the processor (220), and the memory (230) of the network entity (200) may operate. However, the components of the network entity (200) are not limited to the examples described above. For example, the network entity (200) may include more components than the components described above. In one embodiment, the transceiver (210), the processor (220), and the memory (230) may be implemented in the form of a single chip. In addition, the processor (220) may include one or more processors.

[0133] The transceiver (210) collectively refers to the receiver of the network entity (200) and the transmitter of the network entity (200), and can transmit and receive signals with a terminal, a network, or a base station. The signals transmitted and received with the terminal, network, or base station may include control information and data.

[0134] Additionally, the transceiver (210) can perform functions for transmitting and receiving signals via a wireless channel. For example, the transceiver (210) can receive a signal via a wireless channel, output it to the processor (220), and transmit the signal output from the processor (220) via the wireless channel.

[0135] The memory (230) can store programs and data required for the operation of the network entity (200). In addition, the memory (230) can store control information or data included in a signal acquired from the network entity (200). The memory (230) can be configured as a storage medium such as a ROM, a RAM, a hard disk, a CD-ROM, a DVD, or a combination of storage media. In addition, the memory (230) may not exist separately but may be configured as included in the processor (220). The memory (230) can be configured as a volatile memory, a nonvolatile memory, or a combination of volatile and nonvolatile memories. In addition, the memory (230) can provide stored data upon request of the processor (220).

[0136] The processor (220) may control a series of processes so that the network entity (200) may operate according to the above-described embodiment of the present disclosure. For example, the processor (220) may receive control signals and data signals through the transceiver (210) and process the received control signals and data signals. The processor (220) may transmit the processed control signals and data signals through the transceiver (210). In addition, the processor (220) may write or read data to or from the memory (230). The processor (220) may perform functions of a protocol stack required by a communication standard. For this purpose, the processor (220) may include at least one processor or microprocessor. In one embodiment, a portion of the transceiver (210) or the processor (220) may be referred to as a communication processor (CP).

[0137] The above description of the present disclosure is provided for illustrative purposes only, and those skilled in the art will readily appreciate that modifications to other specific forms can be made without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.

[0138] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.

Claims

1. A method for operating a terminal in a wireless communication system, A step for identifying context changes in input data; A step of transmitting a request message requesting an artificial intelligence model for inference of the above input data to a service provider; In response to the above request message, a step of receiving information related to an artificial intelligence model for inferring a changed context from the service provider; A step of identifying the artificial intelligence model based on the received information; and A method comprising a step of performing inference of the changed context through the artificial intelligence model.

2. In paragraph 1, the step of transmitting the request message to the service provider comprises: A step of preprocessing input data; and A method comprising the step of transmitting the preprocessed input data to the service provider.

3. In the second paragraph, the step of preprocessing the input data is: a step of compressing the above input data; or A method comprising the step of extracting feature points of the above input data.

4. A method according to claim 1, wherein the request message includes a list of artificial intelligence models stored in the terminal.

5. In paragraph 4, The artificial intelligence model for inference of the above changed context is determined as the first model included in the list of artificial intelligence models stored in the terminal, A method wherein information related to an artificial intelligence model for inferring the changed context includes an identifier of the first model.

6. In paragraph 4, information related to the artificial intelligence model for inference of the changed context is An identifier of at least one first model selected from a list of artificial intelligence models stored in the terminal; At least one third model not stored in the terminal; and A method comprising information on a method for constructing a second model based on at least one first model and at least one third model.

7. In the 6th paragraph, the step of identifying the artificial intelligence model based on the received information is: A method comprising the step of repositioning and combining at least some of the layers constituting at least one first model and at least one third model.

8. In the fourth paragraph, the artificial intelligence model for inference of the changed context is determined as the fourth model not stored in the terminal, A method further comprising the step of receiving the fourth model from the service provider.

9. A method by which a service provider operates in a wireless communication system, A step of receiving a request message requesting an artificial intelligence model for inference of input data from a terminal; A step of identifying an artificial intelligence model for inference of the above input data; and A method comprising a step of transmitting information related to an artificial intelligence model for inferring the input data to the terminal.

10. A method according to claim 9, wherein the request message includes preprocessed input data.

11. A method according to claim 10, further comprising a step of decrypting the preprocessed input data to obtain input data.

12. A method according to claim 9, wherein the request message includes a list of artificial intelligence models stored in the terminal.

13. In paragraph 12, A step of identifying whether an artificial intelligence model for inference of the identified input data is included in the list of artificial intelligence models stored in the terminal; and A method further comprising the step of transmitting a corresponding identifier to the terminal when an artificial intelligence model for inference of the identified input data is included in the list, and transmitting a corresponding artificial intelligence model to the terminal when the artificial intelligence model for inference of the identified input data is not included in the list.

14. In paragraph 12, information related to the artificial intelligence model for inference of the input data is An identifier of at least one first model selected from a list of artificial intelligence models stored in the terminal; At least one third model not stored in the terminal; and A method comprising information on a method for constructing a second model based on at least one first model and at least one third model.

15. In the 14th paragraph, the artificial intelligence model for inference of the input data is a second model configured by rearranging and combining at least some of the layers constituting the at least one first model and the at least one third model.

Citation Information

Patent Citations

  • Model updating method and device

    CN113765957A

  • Model configuration method and device

    CN116541088A

  • Semiconductor package

    KR1020240134648A

  • Access method, access apparatus, and storage medium

    US20230217366A1