Method and device for selectively processing input data for ai media service in wireless communication system

The method of selectively performing AI inference on partial input data based on predefined conditions addresses inefficiencies in existing systems, enhancing efficiency and resource conservation while maintaining accuracy in wireless communication systems.

WO2025174169A1PCT designated stage Publication Date: 2025-08-21SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/099192
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-16
Filing Date
2025-02-03
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Performing inference on all input data using AI models in wireless communication systems leads to performance degradation, battery drain, increased usage time, heat generation, and network resource congestion, especially when continuous data results in consistently identical inference outcomes, reducing efficiency.

Method used

Implement a method and device for selectively performing inference on a portion of input data based on predefined conditions, such as time, space, or other perspectives, using selective inference execution condition information to optimize AI model usage in terminals and networked architectures.

Benefits of technology

Enhances efficiency by reducing unnecessary processing, conserving resources, and maintaining accuracy in AI services by allowing selective inference on relevant data portions, thereby minimizing performance degradation and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025099192_21082025_PF_FP_ABST
    Figure KR2025099192_21082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a 5G or 6G communication system for supporting a higher data transmission rate. A method performed by a terminal of a wireless communication system, according to embodiments of the present disclosure, may comprise the steps of: receiving first information about selective inference of an artificial intelligence model from a base station; selecting input data for the selective inference of the artificial intelligence model on the basis of the first information; and performing the selective inference of the artificial intelligence model on the basis of the selected input data.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for selectively processing input data for AI media services in a wireless communication system

[0001] The present disclosure relates to a communication method and device for a wireless communication system, and more particularly, to a method and device for selectively processing input data for an AI media service.

[0002] 5G mobile communication technology defines a wide frequency band to enable fast transmission speeds and new services, and can be implemented not only in the sub-6GHz frequency band such as 3.5 gigahertz (3.5GHz), but also in the ultra-high frequency band called millimeter wave (mmWave) such as 28GHz and 39GHz ('Above 6GHz'). In addition, for 6G mobile communication technology, which is called the system after 5G communication (Beyond 5G), implementation in the terahertz band (for example, the 3 terahertz (3THz) band at 95GHz) is being considered to achieve a transmission speed that is 50 times faster than 5G mobile communication technology and an ultra-low latency time that is reduced to one-tenth.

[0003] In the early stages of 5G mobile communication technology, the goal is to support services and satisfy performance requirements for enhanced Mobile Broadband (eMBB), Ultra-Reliable Low-Latency Communications (URLLC), and massive Machine-Type Communications (mMTC). These include beamforming and massive MIMO to mitigate path loss of radio waves in ultra-high frequency bands and increase the transmission distance of radio waves, support for various numerologies (such as operation of multiple subcarrier intervals) and dynamic operation of slot formats for efficient use of ultra-high frequency resources, initial access technology to support multi-beam transmission and wideband, definition and operation of BWP (Bidth Part), new channel coding methods such as LDPC (Low Density Parity Check) codes for large-capacity data transmission and Polar Code for reliable transmission of control information, and L2 pre-processing (L2). Standardization has been made for network slicing, which provides dedicated networks specialized for specific services, and pre-processing.

[0004] Currently, discussions are underway to improve and enhance the initial 5G mobile communication technology in consideration of the services that 5G mobile communication technology was intended to support, and physical layer standardization is in progress for technologies such as V2X (Vehicle-to-Everything) to help autonomous vehicles make driving decisions and increase user convenience based on their own location and status information transmitted by vehicles, NR-U (New Radio Unlicensed) for the purpose of system operation that complies with various regulatory requirements in unlicensed bands, NR terminal low power consumption technology (UE Power Saving), Non-Terrestrial Network (NTN), which is direct terminal-satellite communication to secure coverage in areas where communication with terrestrial networks is impossible, and Positioning.

[0005] In addition, standardization of wireless interface architecture / protocols is in progress for technologies such as intelligent factories (Industrial Internet of Things, IIoT) to support new services through linkage and convergence with other industries, Integrated Access and Backhaul (IAB) that provides nodes for expanding network service areas by integrating wireless backhaul links and access links, Mobility Enhancement technology including Conditional Handover and Dual Active Protocol Stack (DAPS) handover, and 2-step random access (2-step RACH for NR) that simplifies random access procedures. Standardization is also in progress for system architecture / services such as 5G baseline architecture (e.g., Service-based Architecture, Service-based Interface) for grafting Network Functions Virtualization (NFV) and Software-Defined Networking (SDN) technologies, and Mobile Edge Computing (MEC) that provides services based on the location of the terminal.

[0006] Once these 5G mobile communication systems are commercialized, an explosive increase in connected devices will be connected to the communication network, necessitating enhanced functionality and performance of 5G mobile communication systems and integrated operation of these connected devices. To this end, new research will be conducted on improving 5G performance and reducing complexity, supporting AI services, supporting metaverse services, and drone communications by utilizing eXtended Reality (XR), Artificial Intelligence (AI), and Machine Learning (ML) to efficiently support Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR).

[0007] In addition, the development of these 5G mobile communication systems includes new waveforms to ensure coverage in the terahertz band of 6G mobile communication technology, multi-antenna transmission technologies such as Full Dimensional MIMO (FD-MIMO), Array Antenna, and Large Scale Antenna, metamaterial-based lenses and antennas to improve the coverage of terahertz band signals, high-dimensional spatial multiplexing technology using Orbital Angular Momentum (OAM), Reconfigurable Intelligent Surface (RIS) technology, as well as full duplex technology to improve the frequency efficiency and system network of 6G mobile communication technology, satellite, AI (Artificial Intelligence) from the design stage and AI-based communication technology that realizes system optimization by internalizing end-to-end AI support functions, and ultra-high-performance communication and computing resources to provide services with complexity that exceeds the limits of terminal computing capabilities. It can serve as a basis for the development of next-generation distributed computing technologies that can be realized by utilizing them.

[0008] Embodiments of the present disclosure may have as a first purpose a method and device for effectively providing AI services.

[0009] A method performed by a terminal of a wireless communication system according to embodiments of the present disclosure may include the steps of receiving first information for selective inference of an artificial intelligence model from a base station, selecting input data for selective inference of the artificial intelligence model based on the first information, and performing selective inference of the artificial intelligence model based on the selected input data.

[0010] The present disclosure provides a device and method capable of effectively providing AI services in a wireless communication system.

[0011] FIG. 1 illustrates the architecture of a terminal and a network of a wireless communication system according to embodiments of the present disclosure.

[0012] FIG. 2 illustrates information transfer between a terminal application and a service application of a wireless communication system according to embodiments of the present disclosure.

[0013] FIG. 3 illustrates a method for transmitting optional inference execution condition information of a wireless communication system according to embodiments of the present disclosure.

[0014] FIG. 4 illustrates a method for transmitting optional inference execution condition information of a wireless communication system according to embodiments of the present disclosure.

[0015] FIG. 5 illustrates a method for transmitting optional inference execution condition information of a wireless communication system according to embodiments of the present disclosure.

[0016] FIG. 6 illustrates the operation of a monitoring process unit in a wireless communication system according to embodiments of the present disclosure.

[0017] FIG. 7 illustrates the operation of an inference input setting unit in a wireless communication system according to embodiments of the present disclosure.

[0018] FIG. 8 illustrates a method for performing selective inference in a wireless communication system according to embodiments of the present disclosure.

[0019] FIG. 9 is a diagram illustrating the structure of a user terminal according to embodiments of the present disclosure.

[0020] FIG. 10 is a diagram illustrating the structure of a network entity according to embodiments of the present disclosure.

[0021] Recent artificial intelligence (AI) and machine learning technologies (ML, hereinafter collectively referred to as AI / ML) are being introduced and generalized in media-related applications, from traditional applications such as image classification and voice / face recognition to cutting-edge applications such as video quality enhancement.

[0022] As research in this area matures, advanced applications requiring processing larger amounts of data are being considered. Research is ongoing to incorporate AI architectures for media applications that collaborate with servers in the network as well as terminals into 5G systems. Several use cases are being identified and discussed, including object recognition in images and videos, quality enhancement of streaming video, and natural language processing for speech.

[0023] Scenarios for AI inference and learning include delivering pre-trained AI models from the network to user equipment (UE), and inference performed either on the UE or split between the UE and the network. An architecture for these scenarios combines several core components: an AI model repository, AI model provisioning capabilities, AI model access capabilities, an AI model inference engine, and intermediate data provisioning capabilities to efficiently and effectively deliver AI models and associated data over 5G networks.

[0024] AI services provided by AI service providers via 5G networks may be comprised of a step for determining which services to provide and receive based on communication between the application of the service provider server and the application of the terminal, a step for determining an AI model suitable for the service, and a step for determining a version (or variant) of the AI ​​model suitable for operation on the terminal (e.g., a high-accuracy model with high bit depth, a low-accuracy model with quantized bits, etc.).

[0025] The inference results of the AI ​​engine can be transmitted to the data processing unit of the terminal and displayed to the user or applied to an appropriate processing unit for use.

[0026] The operating principles of embodiments of the present disclosure are described in detail below with reference to the attached drawings. The terms described below are defined based on the functions of the present disclosure. These terms may vary depending on the intent or custom of the user or operator, and therefore, their definitions should be determined based on the overall content of this specification.

[0027] The terms used in this disclosure to refer to network entities, network functions, and terminal system objects, terms referring to messages, terms referring to identification information, and the like are provided for convenience of explanation. Therefore, this disclosure is not limited to the terms described below, and other terms referring to objects with equivalent technical meanings may be used.

[0028] For convenience, the present disclosure may use terms and names defined in the 5G system specifications, but is not limited by the terms and names, and may be equally applied to systems conforming to other specifications.

[0029] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the attached drawings. The attached drawings are provided to help understand embodiments of the present disclosure, and it should be noted that the contents of the present disclosure are not limited to the forms or arrangements illustrated in the drawings. In addition, detailed descriptions of well-known functions and configurations that may obscure the gist of the embodiments in the present disclosure will be omitted. It should be noted that in the following description, only the parts necessary for understanding the operation according to various embodiments of the present disclosure will be described, and the description of other parts will be omitted so as not to obscure the gist of the present disclosure. In addition, although the present disclosure describes various embodiments using terminology used in some communication standards (e.g., 3rd Generation Partnership Project (3GPP)), this is merely an example for the purpose of explanation. The various embodiments of the present disclosure can be easily modified and applied to other communication systems.

[0030] Meanwhile, performing inference using an AI model on all input data that can be collected from the terminal, such as each frame, can result in performance degradation, battery drain, increased usage time due to inference, heat generation due to inference, and network resource congestion. Therefore, it is efficient to perform inference on input data at appropriate intervals. In particular, for inference based on the user's gaze, if the user continuously looks at the same object, the same processing is performed on dozens or hundreds of consecutive frames, resulting in consistently identical inference results, which reduces efficiency.

[0031] Inference can be performed by inputting a portion of continuous or discontinuous data based on criteria arbitrarily set by the terminal. However, depending on the AI ​​model, there may be a difference in accuracy between inferences on continuous and discontinuous data, and discontinuous data may not have been learned. Furthermore, since the AI ​​model does not provide information on the continuity of learned data or the resulting accuracy difference, unexpected errors may occur when the terminal arbitrarily selects the data to be inferred.

[0032] Embodiments of the present disclosure relate to a method and device for enabling a service provider that provides an inference service using an AI model and a terminal that receives models from the service provider and performs inference using the models, to selectively perform inference on a part, rather than the entire, of a continuous data input of a target for which the terminal wishes to perform inference.

[0033] Embodiments of the present disclosure may provide a method for a terminal to receive functional information, reception information, and selective inference information of an AI model that can be provided by a service provider. Furthermore, a method may be provided for providing information regarding the types of selection conditions permitted by the service provider for an AI model provided by the service provider and the performance of the AI ​​model for each selection condition. Furthermore, a method may be provided for determining whether the terminal can determine the selection conditions permitted by the AI ​​model it wishes to receive and selecting the AI ​​model and selection conditions. Furthermore, a method may be provided for performing selective inference on data input based on the selection conditions selected by the terminal.

[0034] In addition, embodiments of the present disclosure can provide a method for notifying an entity (terminal or network) performing pre-stage processing of selection condition information selected by the terminal to an entity (network or terminal) performing post-stage processing in a split AI processing architecture that splits AI processing into pre-stage and post-stage processing, and a method for maintaining a split AI processing session.

[0035] Feature information can indicate the functions provided by an AI model. For example, feature information can include providing inferences or image enlargement.

[0036] The received information may indicate information for receiving an AI model from a server. For example, the received information may include a URL, a manifest file, etc.

[0037] Performance information can indicate performance metrics for AI model functions. For example, received information can include inference accuracy, inference time, etc.

[0038] Optional inference execution condition information may indicate information related to the range of selection when allowing partial information compared to the entire set as input to an AI model according to embodiments of the present disclosure. For example, the optional inference execution condition information may include selection condition formats, utterance condition values, etc.

[0039] Selection criteria formats can indicate criteria for selecting a portion of an image over the whole. For example, selection criteria formats can indicate a portion of a frame or a portion of a space.

[0040] The firing condition value can represent information obtained for the selection condition.

[0041] Hereinafter, an embodiment of the present invention will be described with reference to the attached drawings.

[0042] FIG. 1 illustrates the architecture of a terminal and a network of a wireless communication system according to embodiments of the present disclosure.

[0043] Referring to FIG. 1, the architecture of a terminal and a network for supporting embodiments of the present disclosure is illustrated.

[0044] Components for providing and using AI inference services according to embodiments of the present disclosure are described below. Specifically, to support embodiments of the present disclosure, an inference input configuration unit (Inference Input Configuration) and a monitoring process unit (Status Monitor) may be included in the terminal, and optional inference execution condition information (Input Inference Configuration Metadata) may be provided by the service provider.

[0045] In an embodiment of the present disclosure, the monitoring process unit may correspond to the Status Monitor of the drawing. Furthermore, the inference input configuration unit may correspond to the Inference Input Configuration of the drawing. Furthermore, optional inference execution condition information may correspond to the Input Inference Configuration Metadata (IICM) of the drawing.

[0046] Service applications (Network Applications) can provide inference services and AI models for inference to service subscribers.

[0047] Terminal applications (UE Applications) can provide AI media services by utilizing AI inference engines and AI model access functions, discover services that provide AI inference, and receive AI models and various information from the services.

[0048] The AI ​​model repository can store AI models that can be provided by the service.

[0049] The AI ​​model delivery function can deliver an AI model from the AI ​​model storage to a terminal by terminal selection and request.

[0050] The AI ​​model access function can receive an AI model from a network server.

[0051] The input data source can generate or preprocess data to be inferred from the terminal.

[0052] Input data can represent data generated from the input data section.

[0053] The AI ​​model inference engine can perform inference based on input data received from the input data section.

[0054] The data processing unit (Data destination) can process the inference results of the AI ​​inference engine.

[0055] Intermediate data is the result data of partial inference performed from a terminal or a server of a network in a partition inference configuration, and can be used as input for the remaining partition inference to be performed later.

[0056] In a split inference configuration, the UE intermediate data delivery function can transmit intermediate data generated by first performing partial inference on the terminal to the network, or receive intermediate data generated by first performing partial inference on the network server from the network.

[0057] The network intermediate data delivery function can receive intermediate data generated by first performing partial inference on the terminal in a split inference configuration from the terminal, or can transmit intermediate data generated by first performing partial inference on the network server to the terminal.

[0058] Selective inference execution condition information (Input Inference Configuration Metadata) enables selective inference for AI models that can infer using input data selected according to conditions such as time, space, or other perspectives, and can provide information on the quality of inference for each selection condition.

[0059] The Inference Input Configuration unit can control data transmission between the input data unit and the AI ​​inference engine based on optional inference execution conditions selected in the terminal application.

[0060] The monitoring process unit (Status monitor) can report the status of the terminal and input data to the inference input setting unit for determining the optional inference execution conditions selected by the terminal application.

[0061] The roles and operations of components of the terminal and network according to embodiments of the present disclosure are as follows.

[0062] A terminal may discover a service that provides AI inference, receive functional information about AI models from the service, and receive information for receiving a selected AI model among the AI ​​models, and may include a terminal application (UE application) for receiving the AI ​​model. The terminal application may communicate with a service application on a network to discover a service that provides AI inference, and exchange information for receiving functional information and AI models. The AI ​​model and service application provided by a service provider according to the present disclosure may include optional inference execution condition information that allows the AI ​​model to operate conditionally.

[0063] The terminal's AI model access function receives AI models stored in the network's storage from the AI ​​model delivery function. The network's service application determines which AI model to provide to the terminal based on negotiation with the terminal application and transmits the determined AI model from the storage (AI model repository) to the network's AI model delivery function.

[0064] The AI ​​model received on the terminal is transmitted to the AI ​​model inference engine. Selective inference execution condition information can be transmitted to the terminal together with or separately from the AI ​​model, and the terminal application can set up and execute an inference input configuration unit (Inference Input Configuration) that determines whether there are conditions that the terminal can determine among the selection condition information of the received selective inference execution condition information, selects a possible condition among them, and decides whether to perform selective inference, and a monitoring process unit (Status Monitor) that can provide information. The monitoring process unit responds to whether the selection condition described in the selective inference execution condition information can be determined on the terminal based on analysis of status information and input data that can be acquired from inside and outside the terminal, and generates status information according to the judgment and settings of the terminal application and provides it to the inference input configuration unit. The inference input configuration unit controls whether the input data is transmitted to the inference engine depending on whether the selection conditions are met. In other words, the inference input configuration unit may or may not transmit input data to the inference engine depending on the conditions.

[0065] The input data unit can generate input data that is the subject of inference from inputs external or internal to the terminal. For example, the camera device of the terminal can capture the external surroundings of the terminal and generate them in the form of continuous images. Furthermore, for example, a motion sensor within the terminal can continuously track the movements of a user wearing the terminal and generate spatial position or changes in the position in the form of continuous coordinate data. The input data can collectively refer to raw data from the aforementioned devices (e.g., a camera or motion sensor) or data preprocessed for application to the inference engine.

[0066] The AI ​​inference engine performs inference when input data is received from the input data unit. Otherwise, it waits. In other words, it does not operate. Once inference is performed, the results are transmitted to the terminal's data processing unit and can then be displayed to the user or applied to an appropriate processing unit for use.

[0067] For split AI processing architectures that divide AI processing into pre- and post-stages, a split inference architecture can be considered, where AI inference engines on both terminals and network servers execute AI models separately. If split inference is performed rather than all AI inference performed on a single terminal or network server, intermediate data can be transferred between the UE intermediate data delivery function and the network intermediate data delivery function to connect the inference from the terminal or network that performed the inference first to the network or terminal that performs the inference later.

[0068] When the terminal first performs inference using some layers of the AI ​​model, passes intermediate data to the AI ​​inference engine of the network, and the AI ​​inference engine of the network completes inference using the remaining layers of the AI ​​model, when the inference input setting unit passes the input data to the AI ​​inference engine of the terminal, the intermediate data generated by the AI ​​inference engine can be passed to the AI ​​inference engine of the network.

[0069] According to embodiments of the present disclosure, if the inference input setting unit does not transmit input data to the inference engine of the terminal, a notification that selective inference was not performed and thus intermediate data was not transmitted may be transmitted to the AI ​​inference engine of the network. The AI ​​inference engine of the network may recognize that a specific section of the entire flow of input data has been blocked by the inference input setting unit and operate the AI ​​inference engine of the network accordingly. For example, the AI ​​inference engine of the network may reuse the results of previous inference for the time interval during which input was not performed for an AI model that allows discontinuous inputs on the time axis. In addition, the AI ​​inference engine of the network may determine whether the inference session has ended or whether it is a discontinuous time interval during which inference is not performed when intermediate data is not transmitted for a certain period of time or longer, and may terminate or maintain the session.

[0070] When the network first performs inference using some layers of the AI ​​model, passes intermediate data to the AI ​​inference engine of the terminal, and the AI ​​inference engine of the terminal completes inference using the remaining layers of the AI ​​model, the inference input setting unit determines whether to pass the intermediate data passed from the network based on the status information of the terminal, and then the intermediate data can be passed to the AI ​​inference engine of the terminal.

[0071] If the inference input configuration unit does not transmit intermediate data to the terminal's inference engine, a notification may be sent to the network's AI inference engine indicating that the intermediate data has not been used. The network's AI inference engine may recognize that inference has not been completed based on conditions determined by the terminal and operate accordingly. For example, the terminal may identify time intervals during which intermediate data is not used and, accordingly, may not send intermediate data during corresponding time intervals.

[0072] AI model discovery

[0073] Negotiation is initiated between the terminal application and the service application to request a list of AI inference services available from the service provider's service application and to initiate the service.

[0074] FIG. 2 illustrates information transfer for service negotiation between a terminal application and a service application of a wireless communication system according to embodiments of the present disclosure.

[0075] Information transmitted for service negotiation between terminal applications and service applications (network applications) may include the following:

[0076] Types of inference functions provided by AI inference services (e.g., speech recognition, object recognition in images, image enlargement, etc.)

[0077] Feature human description

[0078] Feature profile

[0079] Solution human description

[0080] Solution Profile

[0081] Input data information that can be used for inference provided by the AI ​​inference service (e.g., type, properties, codec, and profile)

[0082] - Type (e.g. video, 3D video, audio, etc.)

[0083] - Attributes (e.g. resolution, FPS)

[0084] - Codec / Profile

[0085] Specifications for inference by AI model (e.g., number of objects that can be recognized, accuracy)

[0086] - Performance of AI models by version (e.g. accuracy, inference execution time, number of model layers, number of nodes, etc.)

[0087] ㆍ Whether or not information on the conditions for performing selective inference is provided (e.g., flag)

[0088] ㆍ Method of receiving information on optional inference execution conditions (e.g. URL)

[0089] ㆍ Information on conditions for implementing selective inference

[0090] Referring to Figure 2, the terminal application can receive the above information and the types of inferences available from the AI ​​inference service provider. The terminal application can select one or multiple inferences. If multiple inferences are selected, the service application can determine whether to provide one AI model or multiple AI models that provide the multiple inferences.

[0091] Feature human descriptions provide AI model functions in a form that users can read and interpret. They can include information that users can read and understand, such as "pet identification."

[0092] A feature profile provides the capabilities of an AI model in a form that applications can evaluate. Predefined profiles can be used to determine values ​​and then be used.

[0093] Since there may be multiple different AI models that provide the same inference, the terminal application can select the version to receive by considering the specifications of each AI model and the performance of each version of each AI model.

[0094] A human description provides the name or number of a specific AI model in a form that users can read and interpret. It can include information that users can read and understand, such as "VGG32 version 1.0 extended for pet."

[0095] A solution profile provides the name of an AI model in a form that applications can interpret. It is determined by the solution provider and may be determined by the service provider's rules.

[0096] The input data format can dictate the format of data that an AI model can process. Since each inference may require different input data types for inference, the terminal application can consider whether the terminal can generate input data that the inference engine can understand when selecting an inference function. For example, when the inference function is object identification, language translation, or image augmentation, the input data types that can be used for inference may include video or audio (speaker identification) for object identification, audio or video (text identification within an image) for language translation, and video or audio (sound enhancement) for image augmentation.

[0097] According to embodiments of the present disclosure, in order to allow selective inference on input data by AI model or version of AI model, information such as whether the AI ​​model can process data discontinuously, the types of conditions under which selective inference can be performed, and the quality difference between inference on continuous data and inference on discontinuous data may be provided.

[0098] The above-mentioned information may also be provided as metadata for Common AI model information, Service requirement information, and Endpoint capability information of 3GPP TR 26.927.

[0099]

[0100]

[0101]

[0102]

[0103] Composition of information on conditions for performing selective inference

[0104] According to embodiments of the present disclosure, [selective inference execution condition information] can provide information on whether some or all of the input data can be used for inference using an AI model, as well as judgment criteria information for that portion. Before receiving the AI ​​model, the terminal can receive the selective inference execution condition information, and then determine whether the terminal can understand the judgment criteria and make a judgment based on them before deciding whether to receive the AI ​​model.

[0105] Optional inference execution condition information may include:

[0106] 1) Whether optional input data can be processed

[0107] 2) Selection conditions

[0108] a. Selection condition format

[0109] b. Ignition condition value

[0110] - Policy: Policy value

[0111] - Accuracy: Accuracy value

[0112] [Whether optional input data processing is possible] can be determined at the pre-learning stage and provided as information on whether there is a difference in results and accuracy between inference on all input data and inference on a portion of the input data in the operation of the AI ​​model.

[0113] [Selection Condition] may be provided with one or more [Selection Condition Types] used to determine whether the AI ​​model can selectively process data, and [Speech Condition Values] allowed for each selection condition type, i.e., one or more values ​​allowed for the selection condition type.

[0114] For example, when the [selection condition format] is the number of frames per second for the video image and the [speech condition values] are 5, 10, and 15, the terminal can receive the AI ​​model by determining that it can understand the [selection condition format] of the number of frames per second and make a judgment accordingly, and then perform inference on the entire input data, and if necessary, change it to perform inference every 5 frames, which is one of the presented [speech condition values].

[0115] According to embodiments of the present disclosure, things that can be provided in the form of [selection conditions] for input data include trigger_frequency, trigger_area, trigger_pose, trigger_foveatedPoint, trigger_object, trigger_motionVector, trigger_viewport, etc., and it can be considered that they mean frequency, area, pose, object, motion vector, and viewport, respectively, but are not limited thereto, and all criteria that can be used to select some, not all, of the input data for inference to an AI model can be included.

[0116] The Triggering Condition Value can vary depending on the type of selection condition. If multiple values ​​are allowed, they can be provided as a list of values ​​(e.g., 5, 10, 15, etc.).

[0117] trigger_frequency indicates that selective inference is possible for some frames among all input frames. Depending on the nature of the AI ​​inference, if the AI ​​model includes a step of comparing with previous frames or comparing with the inference results from previous frames, it may be necessary to use one or more consecutive input frames. For example, an AI model may not produce results for the first 1 or 2 frames, but instead recognizes motion through comparison with previous frames from the subsequent frames and identifies the skeleton accordingly.

[0118] Accordingly, when allowing only a portion of the input frames from the entire set, additional detailed conditions, such as ignition condition values, may be specified. These ignition condition values ​​are provided to ensure a minimum level of accuracy for the AI ​​model's inference results. Using values ​​other than these may indicate that the accuracy of the AI ​​model's inference results is not guaranteed.

[0119] The trigger condition value according to the embodiments may indicate that every Nth frame can be selectively used as inference input (trigger_frequency_every), that non-consecutive frames can be used as inference input but at least N consecutive frames must be input (trigger_frequency_frameAtLeast), or that inference must be performed at least once every few milliseconds or seconds (trigger_frequency_timeAtLeast).

[0120] The trigger_area indicates that selective inference is possible for information within a specific area within the two-dimensional space of a video frame. If allowed, inference can be performed on a portion of the entire area, which can reduce processing time. The trigger condition value for the corresponding selection condition format according to embodiments may indicate a minimum size requirement (trigger_area_sizeAtLeast), etc. By combining with a separate AI model, the terminal can select an area expected to be meaningful within the input data and perform inference only on that area, thereby reducing resource waste.

[0121] The trigger_pose is the terminal's position and orientation information representing the user's field of view in 3D space. If the pose changes rapidly, the inference results may not be used even if inference is performed on the input data. For example, if the head is turned sharply, even if inference is performed on the image between the start and end of the head turn, the target image for the results may not be displayed because the user is no longer looking at the corresponding position after the head has already turned. Another example is that if the user continues to look in a certain direction, continuous inference may only produce the same results, resulting in unnecessary resource waste due to inference. Therefore, the trigger condition value for pose in the form of a selection condition can indicate that inference is not performed for poses outside a certain threshold compared to the previous pose (trigger_pose_inBound) or that inference is not performed for poses within a certain threshold (trigger_pose_outOfBound).

[0122] trigger_foveatedPoint is the location of the user's gaze in 3D space, identified by combining the eye tracking sensor and motion sensor values ​​attached to the HMD and glasses-type terminals. For example, even if the user's pose changes, that is, even if the head is turned, the gaze can remain fixed on a certain object. Also, even if the user's pose does not change, the gaze can fixate on another object. Therefore, the firing condition value for the foveated point as a selection condition type can indicate that the terminal capable of identifying it allows not to re-perform inference after N ms or N seconds fixed within a certain threshold for the foveated point (trigger_foveatedPoint_inBound), or allows not to perform inference when a movement outside of a certain threshold occurs (trigger_foveatedPoint_outOfBound).

[0123] trigger_object refers to objects around the user identified by an object identification device attached to the terminal. For example, by utilizing a separate on-board AI model equipped on the terminal, objects of a certain category (e.g., roads, buildings, signs, pedestrians, etc. in the case of a car) can be identified, and inference about the object can be allowed only when a new object (e.g., sign 4) appears in addition to already identified objects (e.g., signs 1, 2, 3). In order to support the trigger_object selection condition format, the terminal may include a function to track objects in the image. As a selection condition format, the trigger condition value for the object can indicate that it is limited to new objects and excludes existing objects (trigger_object_onlyNew).

[0124] trigger_motionVector can indicate a conditional judgment based on motion flow information such as macroblocks in the video identified by the terminal's encoder. For example, when the terminal encodes a camera-captured video for the purpose of storing it in the terminal itself or playing it back on another terminal, the internal structure of the video encoder can calculate or predict the motion of the video elements in the previous frame and the video elements in the current frame, and calculate the change amount accordingly. When receiving this information from the encoder, new inference may need to be performed according to changes in the screen (e.g., a car's black box) even when the terminal, the user, and the user's gaze are fixed. Therefore, the trigger condition value for the motion vector as a selection condition format can indicate that inference may not be performed when the change amount of the motion vector is within a certain threshold (trigger_motionVector_inBound).

[0125] trigger_viewport can indicate a conditional judgment based on the viewport being viewed, such as when the entire screen range is composed of multiple viewports, such as in cases such as 180-degree or 360-degree VR playback / recording devices. For example, in the case of cube projection, each face of a cube is designated as a viewport, and each face is displayed by mapping it to a sphere, and the origin of the user's gaze direction belongs to any of the six viewports. Inference is performed when the selected viewport changes, and may not be performed if there is no change thereafter. Therefore, as a selection condition type, the trigger condition value for viewport can indicate that inference is performed on the input when a viewport change occurs, and that inference may not be performed thereafter (trigger_viewport_limit).

[0126] According to embodiments, one or more [policies] or multiple policies may be listed and indicated for each case. Policies may be indicated as policy_minimum_required, policy_timeout, etc.

[0127] policy_minimum_required can be indicated when the minimum input required for a single inference is N frames (policy_minimum_required_frames) or N milliseconds (policy_minimum_required_mseconds). That is, it indicates that accurate inference may not be possible for frames less than the specified N. Accordingly, when the policy is indicated as policy_minimum_required and the policy value is N, the terminal can be instructed to add at least N frames as input data and use the inference obtained from the Nth and subsequent frames.

[0128] policy_timeout can be specified to force at least one inference to be performed after N frames (policy_timeout_frames) or N milliseconds (policy_timeout_mseconds) of input have elapsed, even if inference is interrupted by the terminal's choice. If inference is not performed even after the timeout has elapsed, the transmission of intermediate data is interrupted in the case of a split inference configuration, so the server may consider the terminal operation to have ended, and thus the maintenance of the server's split inference service session may not be guaranteed. The server may transmit a timeout value for maintaining the session as policy_timeout and terminate the session after the timeout.

[0129] Policies can be specified in combination. For example, if policy_minimum_required_frame and policy_timeout_mseconds are specified in combination, and N and M are respectively specified as policy values, this indicates that at least N frames are to be used as input for a single inference, but inference can optionally be stopped after that, but at least one inference must be performed after M milliseconds.

[0130] policy_timeout_mseconds is provided in AI model information for split AI / ML operations, which is metadata for split inference configuration in 3GPP TR 26.927, and can be used as a requirement for the minimum operation of the terminal during model splitting.

[0131]

[0132]

[0133] Section 6.6.3 of TR 26.927 (AI model information for split AI / ML operations)

[0134] Additionally, policy_timeout_mseconds may be provided as metadata for Service requirement information in 3GPP TR 26.927.

[0135]

[0136] According to embodiments of the present disclosure, accuracy items and accuracy values ​​may be provided when accuracy values ​​change for each utterance condition value. For example, if accuracy changes between inferences performed every 5 frames and inferences performed every 10 frames, the accuracy for every 5 frames and the accuracy for every 10 frames may be described as sub-items of the corresponding utterance condition values, 5 and 10. If there is no change, the accuracy item and accuracy value may not be included in the optional inference execution condition information.

[0137] Method of conveying information on conditions for performing selective inference

[0138] The terminal can receive selective inference execution condition information from the service server. The selective inference execution condition information may be (1) received by the terminal application by determining the AI ​​service to be received from the service application and included in information for receiving the corresponding AI model, (2) received by being included in structural information indicating the structure of the AI ​​model, (3) included in the AI ​​model so that it can be identified during the parsing process after receiving the AI ​​model, or (4) received after the terminal application requests the service application to transmit the corresponding selective inference execution condition information by separately specifying the AI ​​model.

[0139] (1) In the case of service information provided by the service application to the terminal application, the service application may be included in the list of all AI inference service types (e.g., object identification, language translation, etc.), or in the list of AI model types (e.g., VGG32, PoseNet, etc.) belonging thereto, or in the list of versions (e.g., full version, 32-bit quantified version, mobile version, etc.) of AI models belonging thereto.

[0140] FIG. 3 illustrates a method for transmitting optional inference execution condition information of a wireless communication system according to embodiments of the present disclosure.

[0141] Referring to Figure 3, the following operations can be performed.

[0142] 1. Service subscription and authentication procedures can be performed between terminal applications and service applications.

[0143] 2. The terminal application can request information about available inference services.

[0144] 3. The service application can provide information about the inference service that can be provided.

[0145] 4. The terminal application can select the inference service it wants to use.

[0146] 5. The terminal application can request AI model information that can provide the inference service it wants to use.

[0147] a. Common AI model information may describe the types of inferences provided by the AI ​​model (Feature profile and Feature human description).

[0148] 6. The service application can provide the available AI models and their respective versions and their respective optional inference execution condition information (input interference configuration metadata, IICM).

[0149] a. You can determine whether selective inferencing is possible based on the Flag on selective inferencing in the Common AI model information and Service requirement information. If it is determined to be possible, you can receive it via the URL of Input Inference Configuration Metadata or read the embedded Input Inference Configuration Metadata.

[0150] 7. The terminal application can check whether the terminal's monitoring process unit can check the conditions specified in the optional inference execution condition information.

[0151] 8. The terminal application can determine the AI ​​model and its version with testable conditions.

[0152] 9. The terminal application can notify the service application of the determined AI model and version.

[0153] 10. The service application can identify the AI ​​model selected from the terminal in the repository.

[0154] 11. The terminal application can establish a transmission session between the AI ​​model receiving unit and the AI ​​model transmitting unit to receive the AI ​​model.

[0155] 12. The AI ​​model can be transmitted from the storage to the AI ​​model receiving unit of the terminal via the AI ​​model transmission unit.

[0156] When the terminal application receives the inference service and the corresponding AI model and its corresponding version and performs inference, it can select a different version or a different AI model depending on whether the selective inference execution condition information is supported by the terminal. For example, after the terminal application determines an inference service, it determines a first AI model among the first and second AI models that support it, and in the step of determining the first version of the first AI model, if the selective inference execution condition information is supported by the terminal, the first version of the first AI model can be determined, and if it is not supported, if the first version of the selective inference execution condition information of the first and second versions of the second AI model is supported by the terminal, it can be determined to select the first version of the second AI model by considering this.

[0157] Cases (2), (3), and (4) according to embodiments of the present disclosure are described below.

[0158] (2) is the case where the optional inference execution condition information is received and included within the structural information representing the structure of the AI ​​model. (3) is the case where the optional inference execution condition information is included within the AI ​​model and identified during the parsing process after receiving the AI ​​model. (4) is the case where the terminal application separately specifies the AI ​​model and requests the service application to transmit the corresponding optional inference execution condition information and then receives it.

[0159] The terminal's AI model receiving unit can receive the AI ​​model and then transmit the optional inference execution condition information contained within it to the inference input setting unit. The AI ​​model is then transmitted to the AI ​​inference engine.

[0160] Address information or identifier information for requesting optional inference execution condition information may be included and received in the process in which the terminal application receives information about the inference service from the service application, may be included and received in address information or identifier information for requesting reception of an AI model, or may be included in information of an AI model and determined after receiving the AI ​​model and reading the model information.

[0161] When selective inference execution condition information is received as a separate message, the direct recipient may be the AI ​​model receiver or a separate terminal component. The AI ​​model receiver may request selective inference execution condition information for the AI ​​model from the AI ​​model transmission component of the network and receive it. If a separate terminal component receives AI-related information according to a separate protocol, the terminal component may forward the received information to the AI ​​model receiver, and the AI ​​model receiver may forward the AI ​​model among the received information to the AI ​​inference engine and forward information identified as selective inference execution condition information to the inference input setting component.

[0162] FIG. 4 illustrates a method for transmitting optional inference execution condition information of a wireless communication system according to embodiments of the present disclosure.

[0163] Referring to FIG. 4, a transmission method is illustrated in which selective inference execution condition information is received as included within structural information representing the structure of an AI model or is identified in the process of parsing the AI ​​model after receiving it.

[0164] 13. The terminal application can request AI model information that can provide the inference service it wants to use from the service application (network application).

[0165] 14. The service application can provide the terminal application with information about the AI ​​models available and their respective versions.

[0166] 15. The terminal application can decide which AI model and version it wants to use.

[0167] 16. The terminal application can notify the service application of the determined AI model and version.

[0168] 17. The service application can identify the AI ​​model selected (determined) from the terminal in the storage.

[0169] 18. The terminal application can establish a transmission session between the AI ​​model receiving unit and the AI ​​model transmitting unit to receive the AI ​​model.

[0170] 19. The AI ​​model is transmitted from the storage to the AI ​​model receiving unit of the terminal via the AI ​​model transmission unit.

[0171] 20. The AI ​​model receiving unit can read optional inference execution condition information included in the AI ​​model.

[0172] 21. The AI ​​model receiving unit can transmit optional inference execution condition information to the terminal application.

[0173] 22. The terminal application can check whether the terminal's monitoring process unit can check the conditions specified in the optional inference execution condition information.

[0174] 23. The terminal application can select conditions to be checked and transmit optional inference execution condition information to the inference input setting section.

[0175] FIG. 5 illustrates a method for transmitting optional inference execution condition information of a wireless communication system according to embodiments of the present disclosure.

[0176] Referring to FIG. 5, a method is illustrated for initiating reception of an AI model by determining whether a terminal application can detect a condition for performing optional inference after receiving AI model information.

[0177] 24,25: The terminal application can request AI model information from the service application (network application), and the service application can provide information about available AI models to the terminal application.

[0178] 26,27: The terminal application can request information on optional inference execution conditions of a specific AI model from the service application and receive information on optional inference execution conditions from the service application.

[0179] 28. The terminal application can check whether the condition information can be checked through the monitoring process unit.

[0180] 29. If it is confirmed that the condition information can be checked, the terminal application selects the AI ​​model and version.

[0181] 30. The terminal application notifies the service application of the selected AI model and version information.

[0182] 31. The service application identifies the AI ​​model selected from the terminal in the repository.

[0183] 32. The terminal application can establish a transmission session between the AI ​​model receiving unit and the AI ​​model transmitting unit to receive the AI ​​model.

[0184] 33. The AI ​​model is transmitted from the storage to the AI ​​model transmission unit and then to the AI ​​model reception unit of the terminal.

[0185] Whether to check conditions and how to select them

[0186] A terminal may be equipped with various means for detecting processes running on the terminal and internal / external state changes. For example, as shown in Figure 6, information such as the CPU's process occupancy rate and the encoder's current frame information can be collected for processes; information such as the temperature of various components, such as the CPU, GPU, RAM, storage, and battery, can be collected for internal state changes; and information such as the terminal's spatial position, movement speed, and movement angle can be collected for external state changes.

[0187] The terminal application can review a list of devices received from the terminal OS, a list of status information received from the terminal OS, or a list of information that can be received from the monitoring process unit to determine whether the condition types and firing condition values ​​specified in the optional inference execution condition information can be monitored at the terminal.

[0188] FIG. 6 illustrates the operation of a monitoring process unit in a wireless communication system according to embodiments of the present disclosure.

[0189] Referring to FIG. 6, the monitoring process unit can provide a list of information that can be detected and generated based on information such as selection condition formats and ignition condition values ​​that can be applied within the optional inference execution conditions. The list can be managed statically or dynamically added or deleted. For example, the terminal's battery information may be unavailable during charging and then monitored from the moment charging is complete. Additionally, pose information may be disabled the moment the HMD is removed and enabled the moment the HMD is put back on.

[0190] When the selection condition format and ignition condition value that the monitoring process unit can provide are added or deleted, the monitoring process unit can notify the terminal application of this and cause the terminal application to determine the inference execution conditions again. At this time, the notification can include the occurrence of the change, a list of added conditions, a list of deleted conditions, etc. The terminal application can ignore the list of added or deleted conditions if it does not correspond to the conditions that were previously set. Alternatively, the terminal application can restart the selective inference setting process if it determines that the conditions are more suitable. If the terminal application is simply notified of the occurrence of a change, it can request the list of conditions that can be monitored again, compare it with the list of conditions previously received, and restart the selective inference setting process with the more suitable conditions.

[0191] When the selection condition format and ignition condition value information that the terminal application has decided to use are transmitted to the monitoring process unit, the monitoring process unit can install a dispatcher inside to receive information from terminal components (e.g., motion sensors, charging cable sensors, etc.) that provide information that can determine the corresponding conditions, and continuously monitor them. The monitoring process unit acquires and processes information from the terminal components through the dispatcher to generate ignition condition values. The monitoring process unit can substitute various pieces of information collected from the terminal with the selection condition format and ignition condition values. The substitution can be instructed to the corresponding information generation unit within the monitoring process unit or during the process of setting up the monitoring process unit to acquire information. For example, if the condition is that the motion vectors of one frame generated by the encoder rise or fall by m% or m threshold states compared to the previous frame or the past n frames, the dispatcher can receive the n and m values ​​and generate them as ignition condition information values.

[0192] FIG. 7 illustrates the operation of an inference input setting unit in a wireless communication system according to embodiments of the present disclosure.

[0193] The inference input configuration unit is configured to perform or not perform inference based on the ignition condition value from the terminal application, and a callback that reports the ignition condition value can be installed in the monitoring process unit. The monitoring process unit can report the value to the inference input configuration unit by calling the callback whenever the ignition condition value changes or at regular intervals. For example, the inference input configuration unit can determine whether the n and m values ​​reported by the monitoring process unit have passed a critical state.

[0194] The inference input setting unit can determine whether to perform inference based on a critical state and transmit input data or intermediate data received from the network to the AI ​​inference engine.

[0195] FIG. 8 illustrates a method for performing selective inference in a wireless communication system according to embodiments of the present disclosure.

[0196] Referring to FIG. 8, the method for performing selective inference according to embodiments can be performed as follows.

[0197] 1. Analyze the optional inference execution condition information in the terminal application.

[0198] 2. The terminal application can communicate with the monitoring process unit to determine what information can be detected on the terminal.

[0199] 3. The terminal application can install a dispatcher to obtain status from the monitoring process unit regarding information that is determined to be detectable.

[0200] 4. The terminal application can set conditions for performing selective inference in the inference input setting section.

[0201] 5. The inference input setting section can install a callback in the monitoring process section to obtain utterance information for triggering conditions.

[0202] 6. The monitoring process can update the status or call a callback periodically based on the information collected by the installed dispatcher.

[0203] 7. The inference input setting section can determine whether to perform inference based on information.

[0204] 8. When inference is performed, input data or intermediate data is passed to the AI ​​inference engine.

[0205] 9. Inference is performed in the AI ​​inference engine.

[0206] 10. The results of the inference are communicated.

[0207] 11. If inference is not performed, input data or intermediate data is not passed to the AI ​​inference engine.

[0208] 12. Inference is not performed in the AI ​​inference engine.

[0209] A method for selectively exchanging inference state between split processing architectures.

[0210] A service provider or service application that provides information on conditions for performing selective inference with an AI model may know that the terminal can use all or only a portion of the input data for inference, and may accordingly request information related to the terminal's selective inference operation.

[0211] In one embodiment, there may be an AI inference engine in each of the terminal and the network, and part of the inference may be performed first in the terminal or the network (first half inference), and then the remaining inference may be performed in the network or the terminal (second half inference).

[0212] Intermediate data of AI models that are executed in a divided manner are transmitted between the terminal and the network, and thus, logically, one AI model is executed.

[0213] 1) At this time, selective inference can be performed on the terminal based on conditions. If the terminal is first executed, intermediate data may or may not be transmitted depending on the conditions of the selective inference. Therefore, the remaining inference is not performed on the network.

[0214] 2) If executed first on the network, the remaining inference may or may not be performed on the terminal depending on the conditions of the optional inference. Therefore, intermediate data may or may not be used.

[0215] In case 1), the network needs to determine whether the terminal failed to perform inference based on the triggering of the optional inference execution condition, or whether there was a problem with the transmission of intermediate data due to issues such as network transmission paths. If the terminal's AI inference service session has ended, the resources used by the network's AI inference engine can be immediately released and used for other users. If the terminal's AI inference session is maintained, separate inference can be performed, taking into account the interval where intermediate data was not transmitted.

[0216] Accordingly, in split processing, if the terminal performs AI inference first and inference is selectively performed, it is necessary to notify the network's AI inference engine that inference is not being performed. This notification maintains the session and allows for the recognition of units (e.g., time intervals, space, etc.) where inference is not performed.

[0217] The intermediate data of the terminal can have the following structure.

[0218] Intermediate data message

[0219] - Data type: Intermediate data status notification

[0220] - Selection condition type (frequency) selected from the terminal

[0221] - Status value (frame number=n)

[0222] - Whether or not it is uttered (e.g. 0: not inferred, 1: inferred)

[0223] - Media timestamp (media time applied to determine whether or not a utterance was spoken)

[0224] - Terminal Timestamp (terminal time at which the determination of whether or not to speak was made)

[0225] A server response or request may have the following structure:

[0226] Intermediate data message

[0227] - Server-required non-fire timeout (fire at least once before timeout)

[0228] - Policy after timeout (0: session reset required, 1: session maintained, 2: move to lower priority)

[0229] The network's AI inference engine can receive intermediate data messages and determine whether inference has been performed on the terminal and what conditions and states the terminal has selected.

[0230] In case 2), the first half of the inference performed on the network is transmitted to the terminal as intermediate data. However, since the terminal may not use it, it is necessary to determine whether inference and the transmission of intermediate data are necessary on the network. If conditions are met where the second half of the inference cannot be performed on the terminal, performing the first half of the inference on the network is unnecessary from an energy and resource perspective.

[0231] Therefore, terminals need to report whether or not they will perform inference based on the occurrence of the optional inference execution condition. If the units where inference is not performed are provided as ranges or intervals, the network's AI inference engine and computing management system can more dynamically reclaim and reallocate resources.

[0232] Accordingly, in split processing, when the network first performs AI inference for the first half and the terminal selectively performs AI inference for the second half, it is necessary to notify the network's AI inference engine that inference is not being performed. This notification maintains the session, and allows for discontinuous first-half inference to be performed, or not performed, for the sections where second-half inference is not performed.

[0233] The terminal notification can have the following structure:

[0234] Intermediate data message

[0235] - Data type: Intermediate data status notification

[0236] - Selection condition type (frequency) selected from the terminal

[0237] - Status value (frame number=n)

[0238] - Whether or not it is uttered (e.g. 0: not inferred, 1: inferred)

[0239] - Media timestamp (media time applied to determine whether or not a utterance was spoken)

[0240] - Terminal Timestamp (terminal time at which a decision was made on whether or not to speak)

[0241] - Expected non-firing range (e.g. next x frames, etc.)

[0242] The network's AI inference engine can receive intermediate data messages and determine how many frames the terminal will not perform inference on in the future, as well as the conditions and states selected by the terminal.

[0243] After the nth intermediate message is predicted to be unspoken for the next m frames and transmitted, the n+1th intermediate message can be predicted to be unspoken for the next m-1 frames under the same outlook. If the outlook changes, for example, to predict unspoken for p frames, the network's AI inference engine can schedule the inference of the first half to begin p frames after receiving that message.

[0244] FIG. 9 is a diagram illustrating the structure of a user equipment (UE) according to embodiments of the present disclosure.

[0245] A terminal according to embodiments of the present disclosure may include a processor (930) that controls the overall operation of the terminal, a transceiver (910) including a transmitter and a receiver, and a memory (920). Of course, the present invention is not limited to the examples, and the terminal may include more or fewer components than those illustrated in FIG. 9. The terminal of FIG. 9 may correspond to the terminals described in FIGS. 1, 3 to 5, and 8.

[0246] According to embodiments of the present disclosure, the transceiver (910) can transmit and receive signals with network entities or other terminals. The signals transmitted and received with the network entities may include control information and data. In addition, the transceiver (910) can receive signals via a wireless channel, output them to the processor (930), and transmit the signals output from the processor (930) via the wireless channel.

[0247] According to embodiments of the present disclosure, the processor (930) can control the terminal to perform any one of the operations described above. Meanwhile, the processor (930), the memory (920), and the transceiver (910) do not necessarily have to be implemented as separate modules, and can of course be implemented as a single component in the form of a single chip. In addition, the processor (930) and the transceiver (910) can be electrically connected. In addition, the processor (930) can be an Application Processor (AP), a Communication Processor (CP), a circuit, an application-specific circuit, or at least one processor.

[0248] According to embodiments of the present disclosure, the memory (920) can store data such as basic programs, application programs, and setting information for the operation of the terminal. In particular, the memory (920) provides the stored data upon request of the processor (930). The memory (920) can be configured as a storage medium or a combination of storage media such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD. In addition, there can be a plurality of memories (920). In addition, the processor (930) can perform the above-described embodiments based on a program for performing the above-described embodiments of the present disclosure stored in the memory (920).

[0249] FIG. 10 is a diagram illustrating the structure of a network entity according to embodiments of the present disclosure.

[0250] A network entity according to embodiments of the present disclosure may include a processor (1030) that controls the overall operation of the network entity, a transceiver (1010) including a transmitter and a receiver, and a memory (1020). Of course, the present invention is not limited to the examples, and the network entity may include more or fewer components than those illustrated in FIG. 10 . The network entity of FIG. 10 may correspond to the networks illustrated in FIGS. 1, 3, and 5 .

[0251] According to embodiments of the present disclosure, the transceiver (1010) can transmit and receive signals with at least one of other network entities or terminals. The signals transmitted and received with at least one of the other network entities or terminals may include control information and data.

[0252] According to embodiments of the present disclosure, the processor (1030) can control a network entity to perform any one of the operations described above. Meanwhile, the processor (1030), the memory (1020), and the transceiver (1010) do not necessarily have to be implemented as separate modules, and can of course be implemented as a single component in the form of a single chip. In addition, the processor (1030) and the transceiver (1010) can be electrically connected. In addition, the processor (1030) can be an Application Processor (AP), a Communication Processor (CP), a circuit, an application-specific circuit, or at least one processor.

[0253] According to embodiments of the present disclosure, the memory (1020) may store data such as basic programs, application programs, and setting information for the operation of a network entity. In particular, the memory (1020) provides the stored data upon request of the processor (1030). The memory (1020) may be configured as a storage medium or a combination of storage media such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD. In addition, there may be a plurality of memories (1020). In addition, the processor (1030) may perform the aforementioned embodiments based on a program for performing the aforementioned embodiments of the present disclosure stored in the memory (1020).

[0254] It should be noted that the aforementioned configuration diagrams, examples of control / data signal transmission methods, examples of operational procedures, and configuration diagrams are not intended to limit the scope of the present disclosure. That is, not all components, entities, or operations described in the embodiments of the present disclosure should be construed as essential components for implementing the disclosure, and implementations may be made within a scope that does not detract from the essence of the disclosure even if only some components are included. Furthermore, each embodiment may be combined and operated as needed. For example, parts of the methods proposed in the present disclosure may be combined to operate network entities and terminals.

[0255] The operations of the network entity or terminal described above can be realized by providing a memory device storing the corresponding program code within any component of the network entity or terminal device. That is, the control unit of the network entity or terminal device can execute the operations described above by reading and executing the program code stored in the memory device using a processor or CPU (Central Processing Unit).

[0256] The various components and modules of the entity or terminal device described in this specification may be operated using hardware circuits, such as logic circuits based on complementary metal oxide semiconductors, firmware, software, and / or hardware and firmware and / or software embedded in a machine-readable medium. For example, various electrical structures and methods may be implemented using electrical circuits such as transistors, logic gates, and application-specific semiconductors.

[0257] When implemented in software, a computer-readable storage medium storing one or more programs (software modules) may be provided. The one or more programs stored in the computer-readable storage medium are configured for execution by one or more processors within an electronic device. The one or more programs include instructions that cause the electronic device to execute methods according to embodiments described in the claims or specification of the present disclosure.

[0258] These programs (software modules, software) may be stored in random access memory, non-volatile memory including flash memory, read only memory (ROM), electrically erasable programmable read only memory (EEPROM), magnetic disc storage device, compact disc ROM (CD-ROM), digital versatile discs (DVDs) or other forms of optical storage device, magnetic cassette. Or, they may be stored in a memory configured as a combination of some or all of these. In addition, each configuration memory may be included in multiple numbers.

[0259] Additionally, the program may be stored on an attachable storage device that is accessible via a communication network, such as the Internet, an intranet, a local area network (LAN), a wide local area network (WLAN), a storage area network (SAN), or a combination thereof. Such a storage device may be connected to a device implementing an embodiment of the present disclosure via an external port. Additionally, a separate storage device on the communication network may be connected to a device implementing an embodiment of the present disclosure.

[0260] In the specific embodiments of the present disclosure described above, components included in the disclosure are expressed singularly or plurally, depending on the specific embodiment presented. However, the singular or plural expressions are selected to suit the presented situation for convenience of explanation, and the present disclosure is not limited to singular or plural components. Components expressed in plural may be composed of singular elements, or components expressed in singular may be composed of plural elements.

[0261] A method according to embodiments of the present disclosure is a method performed by a terminal of a wireless communication system, the method including: receiving first information regarding selective inference of an artificial intelligence model from a base station; selecting input data for selective inference of the artificial intelligence model based on the first information; and performing selective inference of the artificial intelligence model based on the selected input data. In this case, the first information may include second information indicating whether the artificial intelligence model supports selective inference. In addition, the first information may include third information regarding at least one condition for selecting input data in relation to the selective inference of the artificial intelligence model. In addition, the first information may include fourth information indicating a prediction accuracy for at least one condition related to the selective inference of the artificial intelligence model.

[0262] In one embodiment, the input data may include any one of data about a video frame, data about a location of the terminal, data about an orientation of the terminal, data about a user's gaze, data about an identified object, data about a motion of the terminal, or data about a viewport.

[0263] Additionally, in one embodiment, the method may further include a step of selecting an artificial intelligence model by the terminal based on the first information; a step of requesting the artificial intelligence model from the base station; and a step of receiving the artificial intelligence model from the base station.

[0264] Additionally, in one embodiment, the method may further include transmitting a message to the base station containing intermediate data generated from the artificial intelligence model through selective inference. The message may include any one of: information regarding the type of intermediate data, information indicating whether selective inference is to be performed on the intermediate data, or information regarding a period during which selective inference is not performed.

[0265] While the detailed description of the present disclosure has described specific embodiments, it should be understood that various modifications are possible without departing from the scope of the present disclosure. Therefore, the scope of the present disclosure should not be limited to the described embodiments, but should be determined not only by the scope of the claims described below but also by equivalents thereof. In other words, it will be apparent to those skilled in the art that other modifications based on the technical idea of ​​the present disclosure are possible. In addition, each embodiment can be combined and operated as needed. For example, parts of the methods proposed in the present disclosure can be combined to operate a base station and a terminal. In addition, although the embodiments have been presented based on a 5G, NR system, other modifications based on the technical idea of ​​the embodiments can be implemented in other systems such as LTE, LTE-A, and LTE-A-Pro systems.

[0266] While the detailed description of this disclosure has described specific embodiments, it should be understood that various modifications are possible without departing from the scope of this disclosure. Therefore, the scope of this disclosure should not be limited to the described embodiments, but should be defined not only by the scope of the claims described below, but also by equivalents thereof.

Claims

1. In a method performed by a terminal of a wireless communication system, A step of receiving first information about selective inference of an artificial intelligence model from a base station; A step of selecting input data for selective inference of the artificial intelligence model based on the first information; and A method comprising the step of performing selective inference of the artificial intelligence model based on the selected input data.

2. In paragraph 1, A method wherein the first information includes second information indicating whether the artificial intelligence model supports selective inference.

3. In paragraph 1, A method wherein the first information includes third information about at least one condition for selecting input data in relation to selective inference of the artificial intelligence model.

4. In paragraph 3, A method wherein the first information includes fourth information indicating the prediction accuracy for at least one condition related to selective inference of the artificial intelligence model.

5. In paragraph 3, A method wherein the input data includes any one of data about a video frame, data about a location of the terminal, data about a direction of the terminal, data about a user's gaze, data about an identified object, data about a motion of the terminal, or data about a viewport.

6. In paragraph 1, A step of selecting the artificial intelligence model based on the first information; A step of requesting the artificial intelligence model to the base station; and A method further comprising the step of receiving the artificial intelligence model from the base station.

7. In paragraph 1, The above method, A method further comprising the step of transmitting, to the base station, a message including intermediate data generated from the artificial intelligence model by selective inference.

8. In paragraph 7, The above message is, A method comprising any one of information about the type of the intermediate data, information indicating whether to perform selective inference on the intermediate data, or information about a period during which selective inference is not performed.

9. In the terminal of a wireless communication system, Transmitter and receiver; and Includes a control unit that is constantly connected to the transmitter and receiver, The above control unit: Receive first information about the selective inference of the artificial intelligence model from the base station, Based on the first information, input data for selective inference of the artificial intelligence model is selected, and A terminal configured to perform selective inference of the artificial intelligence model based on the selected input data.

10. In paragraph 9, A terminal wherein the first information includes second information indicating whether the artificial intelligence model supports selective inference.

11. In paragraph 9, A terminal, wherein the first information includes third information about at least one condition for selecting input data in relation to selective inference of the artificial intelligence model.

12. In paragraph 11, A terminal, wherein the first information includes fourth information indicating the prediction accuracy for at least one condition related to the selective inference of the artificial intelligence model.

13. In paragraph 11, A terminal, wherein the input data includes any one of data about a video frame, data about the position of the terminal, data about the direction of the terminal, data about the user's gaze, data about an identified object, data about the motion of the terminal, or data about a viewport.

14. In paragraph 9, the control unit: Selecting the artificial intelligence model based on the first information, Requesting the artificial intelligence model to the above base station, and A terminal configured to receive the artificial intelligence model from the base station.

15. In paragraph 9, the control unit is set to transmit a message including intermediate data generated from the artificial intelligence model by selective inference to the base station, The terminal, wherein the message includes any one of information about the type of the intermediate data, information indicating whether to perform selective inference on the intermediate data, or information about a period during which selective inference is not performed.

Citation Information

Patent Citations

  • Model reasoning method and device, electronic equipment and storage medium

    CN113139660A

  • Optimized co-inference for a pluralty of ai agents in a mobile communication network

    EP4261742A1

  • Method and system for artificial intelligence based medical image segmentation

    US20210110135A1

  • Artificial intelligence with explainability insights

    US20230281426A1

  • Conditional artificial intelligence, machine learning model, and parameter set configurations

    US20230412470A1