Resource allocation method and device for real-time video analysis in wireless communication system

The method addresses computing limitations in real-time video analysis by employing joint scheduling and adaptive techniques to manage network and computing variability, ensuring stable latency and accuracy in wireless communication systems, particularly for applications like mixed reality and autonomous driving.

WO2026089398A1PCT designated stage Publication Date: 2026-04-30SAMSUNG ELECTRONICS CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-10-17
Publication Date
2026-04-30

Smart Images

  • Figure KR2025016490_30042026_PF_FP_ABST
    Figure KR2025016490_30042026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a 5G or 6G communication system for supporting a data transmission rate higher than that of a 4G communication system such as LTE. The present disclosure relates to a method and a device therefor, the method comprising: determining a target bitrate and a neural network model on the basis of an E2E latency requirement, network state information, and computing state information; dividing the E2E latency requirement into a first latency requirement and a second latency requirement on the basis of the target bitrate, the neural network model, a first latency distribution, and a second latency distribution; and transmitting the first latency requirement and encoding frame information to a base station, wherein a terminal is allocated, from the base station, a resource block determined on the basis of the first latency requirement, the encoding frame information, and the network state information, and a GPU is allocated to the terminal on the basis of the second latency requirement and the computing state information.
Need to check novelty before this filing date? Find Prior Art

Description

Resource allocation method and device for real-time video analysis in a wireless communication system

[0001] The present disclosure relates to a wireless communication system, and more specifically, to a method and apparatus for allocating resources for real-time video analysis in a wireless communication system.

[0002] Looking back at the evolution of wireless communication through successive generations, technologies have been developed primarily for human-oriented services, such as voice, multimedia, and data. Following the commercialization of 5G (5th-generation) communication systems, connected devices, which have been increasing explosively, are expected to be connected to communication networks. Examples of networked objects include vehicles, robots, drones, home appliances, displays, smart sensors installed in various infrastructures, construction machinery, and factory equipment. Mobile devices are expected to evolve into various form factors, such as augmented reality glasses, virtual reality headsets, and holographic devices. In the 6G (6th-generation) era, efforts are underway to develop improved 6G communication systems to connect hundreds of billions of devices and objects to provide diverse services. For this reason, 6G communication systems are being referred to as "beyond 5G" systems.

[0003] In the 6G communication system predicted to be realized around 2030, the maximum transmission speed is tera (i.e., 1,000 gigabit) bps, and the wireless latency is 100 microseconds (μsec). In other words, compared to the 5G communication system, the transmission speed in the 6G communication system is 50 times faster, and the wireless latency is reduced to one-tenth.

[0004] To achieve such high data transmission speeds and ultra-low latency, 6G communication systems are being considered for implementation in the terahertz band (e.g., the 95 GHz to 3 terahertz (3 THz) band). In the terahertz band, due to more severe path loss and atmospheric absorption compared to the millimeter wave (mmWave) band introduced in 5G, the importance of technology capable of guaranteeing signal reach, or coverage, is expected to increase. As key technologies to ensure coverage, radio frequency (RF) devices, antennas, new waveforms that offer better coverage than orthogonal frequency division multiplexing (OFDM), beamforming, and multi-antenna transmission technologies such as massive multiple-input and multiple-output (massive MIMO), full-dimensional MIMO (FD-MIMO), array antennas, and large-scale antennas must be developed. In addition, new technologies such as metamaterial-based lenses and antennas, high-dimensional spatial multiplexing technology using orbital angular momentum (OAM), and reconfigurable intelligent surface (RIS) are being discussed to improve coverage of terahertz band signals.

[0005] In addition, to improve frequency efficiency and system network, development is underway in 6G communication systems for full duplex technology, in which uplink and downlink simultaneously utilize the same frequency resources at the same time; network technology that integrates satellites and HAPS (high-altitude platform stations); network structure innovation technology that supports mobile base stations and enables network operation optimization and automation; dynamic spectrum sharing technology through collision avoidance based on spectrum usage prediction; AI-based communication technology that utilizes AI (artificial intelligence) from the design stage and internalizes end-to-end AI support functions to realize system optimization; and next-generation distributed computing technology that realizes services of complexity exceeding the limits of terminal computing capabilities by utilizing ultra-high performance communication and computing resources (mobile edge computing (MEC), cloud, etc.). In addition, attempts are continuing to further strengthen connectivity between devices, further optimize networks, promote the softwareization of network entities, and increase the openness of wireless communication through the design of new protocols to be used in 6G communication systems, the implementation of hardware-based security environments, the development of mechanisms for the safe utilization of data, and the development of technologies regarding privacy maintenance methods.

[0006] Due to the research and development of such 6G communication systems, it is expected that a new dimension of hyper-connected experience will become possible through the hyper-connectivity of 6G communication systems, which encompasses not only connections between objects but also connections between people and objects. Specifically, it is projected that 6G communication systems will enable the provision of services such as truly immersive extended reality (truly immersive XR), high-fidelity mobile holograms, and digital replicas. Furthermore, services such as remote surgery, industrial automation, and emergency response, which are provided through 6G communication systems with enhanced security and reliability, will be applied in various fields including industry, healthcare, automotive, and home appliances.

[0007] A method performed by a server in a wireless communication system according to one embodiment may include: a step of determining a target bitrate for streaming real-time video and a neural network model for analyzing real-time video based on an end-to-end (E2E) latency requirement, network status information received from a base station, and computing status information; a step of dividing the E2E latency requirement into a first latency requirement at a base station and a second latency requirement at a server based on the target bitrate, the neural network model, a first latency distribution at a base station, and a second latency distribution at a server; transmitting the first latency requirement and encoding frame information received from a terminal to a base station, wherein the terminal is allocated a resource block determined based on the first latency requirement, encoding frame information, and network status information from the base station; and a step of allocating a graphic processing unit (GPU) to the terminal based on the second latency requirement and computing status information.

[0008] In a wireless communication system according to one embodiment, the server includes a transceiver; and at least one processor coupled to the transceiver. The at least one processor determines a target bitrate for streaming real-time video and a neural network model for analyzing real-time video based on an end-to-end (E2E) latency requirement, network status information received from a base station, and computing status information. Based on the target bitrate, the neural network model, a first latency distribution at the base station, and a second latency distribution at the server, the E2E latency requirement is divided into a first latency requirement at the base station and a second latency requirement at the server. The processor transmits the first latency requirement and encoding frame information received from the terminal to the base station. The terminal is allocated a resource block determined based on the first latency requirement, encoding frame information, and network status information from the base station. Based on the second latency requirement and computing status information, a graphic processing unit (GPU) can be allocated to the terminal.

[0009] FIG. 1 is a diagram illustrating the operational structure of a real-time video analysis system according to one embodiment of the present disclosure.

[0010] FIG. 2 is a diagram illustrating the correlation between resource variability according to one embodiment of the present disclosure.

[0011] FIG. 3 is a diagram illustrating an example of performing joint scheduling in consideration of resource variability according to one embodiment of the present disclosure.

[0012] FIG. 4 is a diagram illustrating latency variability according to one embodiment of the present disclosure.

[0013] FIG. 5 is a diagram illustrating a resource allocation operation for analyzing real-time video according to one embodiment of the present disclosure.

[0014] FIG. 6 is a drawing illustrating a method for protecting sensitive information according to one embodiment of the present disclosure.

[0015] FIG. 7 is a diagram illustrating a scheduling method based on latency probability modeling and a latency requirement partitioning method according to one embodiment of the present disclosure.

[0016] FIG. 8 is a diagram illustrating a resource allocation operation that takes into account a change in latency distribution according to one embodiment of the present disclosure.

[0017] FIG. 9 is a flowchart illustrating a resource allocation operation for real-time video analysis according to one embodiment of the present disclosure.

[0018] FIG. 10 is a block diagram illustrating the configuration of a terminal, a base station, and a server according to one embodiment of the present disclosure.

[0019] The operating principles of the present disclosure will be described in detail below with reference to the attached drawings. In describing the present disclosure below, specific descriptions of related known functions or configurations will be omitted if it is determined that such detailed descriptions would unnecessarily obscure the essence of the present disclosure. Furthermore, the terms described below are defined in consideration of their functions in the present disclosure, and these may vary depending on the intentions or practices of the user or operator. Therefore, their definitions should be based on the content throughout this specification.

[0020] The advantages and features of the present disclosure and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below but may be implemented in various different forms. The embodiments provided are merely to ensure that the disclosure is complete and to fully inform those skilled in the art of the scope of the invention, and the present disclosure is defined only by the scope of the claims. Throughout the specification, the same reference numerals refer to the same components.

[0021] At this point, it will be understood that each block of the process flow diagrams and combinations of the flow diagrams can be executed by computer program instructions. Since these computer program instructions can be loaded into the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment, the instructions executed through the processor of the computer or other programmable data processing equipment create means to perform the functions described in the flow diagram block(s). Since these computer program instructions can also be stored in computer-available or computer-readable memory that can be directed toward the computer or other programmable data processing equipment to implement the function in a specific way, the instructions stored in computer-available or computer-readable memory can also produce a manufactured item containing instruction means to perform the function described in the flow diagram block(s). Since computer program instructions can be loaded onto a computer or other programmable data processing equipment, instructions that perform a series of operation steps on the computer or other programmable data processing equipment to create a process executed by the computer can also provide steps for executing the functions described in the flowchart block(s).

[0022] Additionally, each block may represent a module, segment, or part of code containing one or more executable instructions for executing a specific logical function(s). It should also be noted that in some alternative execution examples, the functions mentioned in the blocks may occur out of order. For example, two blocks described in succession may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order according to their corresponding functions.

[0023] In this embodiment, the term "part" refers to a software or hardware component such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the "part" performs certain roles. However, the meaning of "part" is not limited to software or hardware. The "part" may be configured to reside in an addressable storage medium or configured to run one or more processors. Thus, as an example, the "part" includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided within the components and "parts" may be combined into a smaller number of components and "parts" or further separated into additional components and "parts." In addition, the components and 'parts' may be implemented to utilize one or more CPUs within the device or secure multimedia card. Also, in the embodiments, 'parts' may include one or more processors.

[0024] In describing the present disclosure below, specific descriptions of related known functions or configurations will be omitted if it is determined that such detailed descriptions would unnecessarily obscure the essence of the present disclosure. Embodiments of the present disclosure will be described below with reference to the attached drawings.

[0025] The terms used in this invention have been selected based on currently widely used general terms, taking into account their functions within the invention; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this invention should be defined not merely by their names, but based on their meanings and the overall content of the invention.

[0026] When a part of a specification is described as "including" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, terms such as "...part" or "module" as used in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware or software, or as a combination of hardware and software.

[0027] Additionally, the description 'at least one of A, B, and C' means that it may be any one of 'A', 'B', 'C', 'A and B', 'A and C', 'B and C', and 'A, B, and C'.

[0028] It should be understood that the combinations of blocks and flowcharts in each flowchart can be executed by one or more computer programs containing computer-executable instructions. One or more computer programs may be stored entirely in a single memory or may be partitioned and stored in multiple different memories.

[0029] All functions or operations described in this document may be processed by a single processor or a combination of processors. A single processor or a combination of processors is a circuitry that performs processing and may include circuitry such as an AP (Application Processor), CP (Communication Processor), GPU (Graphical Processing Unit), NPU (Neural Processing Unit), MPU (Microprocessor Unit), SoC (System on Chip), IC (Integrated Chip), etc.

[0030] A processor may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including in the claims, may include at least one processor and various processing circuits. In at least one processor, one or more processors may be configured to perform the various functions described herein in a distributed manner, individually and / or collectively. As used herein, "processor," "at least one processor," and "one or more processors" may be configured to perform various functions. However, these terms cover, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor can perform all functions. Additionally, at least one processor may include a combination of processors performing various functions of the disclosed functions in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.

[0031] The following describes embodiments with reference to the attached drawings so that those skilled in the art can easily implement the present invention. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present invention in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.

[0032] The present disclosure will be described below with reference to the attached drawings.

[0033] FIG. 1 is a diagram illustrating the operational structure of a real-time video analysis system according to one embodiment of the present disclosure.

[0034] Referring to FIG. 1, the real-time video analysis system may include a terminal (110), a base station (network, 120), and a server (edge / cloud server, 130).

[0035] According to one embodiment of the present disclosure, the terminal (110) may be a device of various types that streams a video stream collected in real time through a camera via a base station (120) and transmits the video stream to a server (130). For example, the terminal (110) may be implemented as an electronic device of various types and forms including a processor and a transceiver. The terminal (110) may include devices that communicate with the base station (120) and the server (130), such as a smart TV, smartphone, tablet PC, laptop PC, digital camera, smart watch, glasses-type display, head-mounted display (HMD), etc., but is not limited thereto. The terminal (110) may be expressed as an edge terminal, edge device, end user, end terminal, user, etc., but is not limited thereto.

[0036] According to one embodiment of the present disclosure, a base station (120) may be a device that allocates resources (e.g., resource blocks) to a terminal (110) for streaming a real-time video stream. The terminal (110) may stream a video stream through the base station (120). In one embodiment, the base station (120) may communicate with a server (130) to transmit network status information and receive computing status information. In one embodiment, the base station (120) may be represented as a network, a base station (BS), a radio access network (RAN), a network operator, a network stage, a network stage, etc., but is not limited thereto.

[0037] According to one embodiment of the present disclosure, the server (130) may be a device that analyzes and infers a video stream received from a terminal (110) using a neural network model. The server (130) may be a device capable of processing complex operations and tasks using large-scale data, such as training, inference, management, and distribution of the neural network model. According to one embodiment of the present disclosure, the training of the neural network model executed on the server (130) may be performed by another computing device. The server (130) may receive a video stream from the terminal (110) and transmit the results of analyzing the video stream to the terminal (110). According to one embodiment of the present disclosure, the server (130) may be represented as an edge server, an edge computing server, a cloud server, a service operator, a compute stage, a computing unit, etc., but is not limited thereto.

[0038] Here, "deployment" refers to the state in which a model has been moved to an actual production environment to perform inference operations, rather than to an environment for training (e.g., pre-training and / or fine-tuning training) or testing. In other words, "deployment" indicates a state where the model is ready to provide a service after training is complete and performance verification has been conducted. However, since the service may be provided to individuals and / or limited groups, deployment does not necessarily mean that it must be publicly released.

[0039] According to one embodiment of the present disclosure, the server (130) may include a GPU (graphic processing unit). The GPU is a component used for deep learning and artificial intelligence analysis using a neural network model, and may have more efficient computing power than a CPU (central processing unit) in parallel processing, processing of large-scale matrix operations, and training and inference of neural network models.

[0040] Real-time video analysis applications are becoming increasingly important in various fields, such as mixed reality, autonomous driving, and video surveillance. The core of these applications lies in the high-precision analysis of video streams collected in real-time from edge device cameras using multiple neural network models (e.g., Deep Neural Networks (DNNs)). However, due to limitations in the computing performance, power consumption, and heat generation of the edge devices themselves, there are difficulties in performing computationally intensive tasks, such as high-definition video analysis, or immersive user interactions requiring ultra-low latency real-time processing.

[0041] Mobile edge computing (MEC) architecture deploys servers at the edge of the network to overcome the limitations of edge devices and enables precise real-time analysis with high processing speed and low latency. This system can stream video over the cellular RAN and perform real-time neural network model inference on GPU-equipped cloud servers near base stations.

[0042] FIG. 2 is a diagram illustrating the correlation between resource variability according to one embodiment of the present disclosure.

[0043] Referring to FIG. 2, the resources at the base station (120) and the resources at the server (130) may not be constant and may have variability. The resources at the base station (120) may include bandwidth at the network level, and the resources at the server (130) may include a GPU (graphic processing unit) at the computing level.

[0044] Referring to Graph 210, network bandwidth can change over time. For example, variability in network bandwidth may occur due to reasons such as user movement or the presence of obstacles. Depending on the variability in network bandwidth, the video frame transmission time, the streaming of video frames, and the quality of the encoded video may be affected. For example, when a terminal (110) is encoding and transmitting video at a bitrate of 5 Mbps, if the network bandwidth drops below 5 Mbps, a bottleneck may occur in the video frame transmission. That is, the video frame transmission may be delayed and the video quality may be degraded.

[0045] A bottleneck refers to a bottleneck within a system that limits processing speed or performance. It can be used to describe a situation in a physical or technical system where a specific step slows down the flow of the entire process. When a bottleneck occurs, the overall performance of the system degrades, and unless the bottleneck is resolved, improvement may not be possible, no matter how fast the rest of the system operates.

[0046] Referring to graph 220, the workload at the computing stage may change over time. The workload may refer to the amount of work that the server (130) must process. For example, the workload may increase as the number of objects to be analyzed increases.

[0047] According to one embodiment of the present disclosure, the analysis time and analysis quality using a neural network model may be affected by the variability of the workload. For example, when the number of GPUs available for allocation to the terminal (110) is 6, if the number of GPUs required for video frame analysis is greater than 6, a bottleneck may occur in the analysis using a neural network model. That is, video frame analysis may be delayed, and the analysis quality may be degraded.

[0048] In one embodiment, variations in network bandwidth and variations in workload may have a low correlation with each other. For example, variations in network bandwidth and variations in workload may have a Pearson correlation coefficient of -0.15. That is, variations in network bandwidth and variations in workload may occur independently of each other and may not be affected by each other's variability.

[0049] The Pearson correlation coefficient is a statistical indicator that measures the linear correlation between two variables. The Pearson correlation coefficient ranges from -1 to 1, and values ​​closer to 0 indicate that there is almost no linear relationship (i.e., correlation) between the two variables.

[0050] For systems where end users stream video in real-time over a network and the server analyzes it in real-time using artificial intelligence, adaptive analysis technology is crucial to maintain stable latency and analysis accuracy even under real-world constraints such as network bandwidth fluctuations or competition for GPU resources.

[0051] Adaptive analysis techniques based on resource variability may include Adaptive Bitrate Streaming and DNN Model Adaptation.

[0052] Adaptive Bitrate Streaming monitors real-time network bandwidth and can adjust video encoding quality below available bandwidth. In this case, to minimize the decrease in DNN analysis accuracy caused by video quality degradation due to low bandwidth, video encoding can be performed adaptively according to the video content and purpose. For example, frame rate, resolution, and quantization can be adaptively applied considering the size and movement speed of objects, or Region-of-Interest (RoI) within video frames can be predicted considering object detection, and only those RoIs can be selectively maintained at high quality.

[0053] DNN Model Adaptation allows for adjusting the complexity of the DNN model required for analysis when GPU resources are scarce on the server. In such cases, an adaptive approach can be utilized to maintain analysis accuracy without exceeding GPU usage limits. For instance, GPU usage can be reduced by progressively performing operations layer by layer of the DNN model and monitoring the results, then skipping remaining operations once sufficient analysis accuracy is achieved. Additionally, the complexity of the DNN model can be adaptively selected by predicting the difficulty of analyzing objects within a video frame (e.g., object pose, obstacles).

[0054] However, the aforementioned adaptive analysis technique performs adaptive scheduling limited to a single stage, either the network stage or the computing stage. In real-world environments, bottlenecks may appear independently or in a mixed manner at both stages. For example, in the case of a real-time video analysis workload that detects people within a video and analyzes individual characteristics (e.g., behavior, gender), network and computing bottlenecks can occur independently of variations in network bandwidth and workload (i.e., the number of people in the video).

[0055] FIG. 3 is a diagram illustrating an example of performing joint scheduling in consideration of resource variability according to one embodiment of the present disclosure.

[0056] Referring to FIG. 3, the server (130) can determine the bitrate and the neural network model by taking into account the variability of network bandwidth and computing workload. The server (130) can perform joint scheduling to determine the target bitrate for streaming and the neural network model used for video analysis by monitoring the variability of network bandwidth and computing workload.

[0057] Referring to Graph 310, for example, it can be assumed that a terminal (110) transmits video frames at a bit rate of 1 Mbps, and a neural network model analyzes the video frames using D0 among deep neural network (DNN) models. Here, D0 and D5 may differ in the size of the DNN models. D5 may mean a DNN model that is larger (i.e. heavier) than D0, and D5 may have higher accuracy and take more time to analyze than D0.

[0058] When network bandwidth is reduced, the terminal (110) can transmit video frames at a low bit rate, for example, lower than 0.1 Mbps, depending on the reduced bandwidth. Video frames transmitted at a low bit rate may contain video frames of low quality, and the accuracy of video frame analysis may be reduced.

[0059] Analysis accuracy can be represented by the value of mIoU (mean Intersection of Union), and the mIoU value may decrease depending on the low analysis accuracy. The server (130) may use the DNN model of D5 instead of D0 to increase analysis accuracy, thereby increasing the analysis accuracy (i.e., mIoU).

[0060] Here, mIoU may refer to the average value of the overlap ratio between the actual object region and the predicted object region, and can serve as an indicator to determine whether accuracy requirements are met. For example, mIoU can be a numerical indicator representing how accurately object boundaries are predicted when performing object detection.

[0061] According to one embodiment of the present disclosure, the server (130) may adjust the neural network model to overcome the degradation of analysis accuracy caused by fluctuations in network bandwidth. Conversely, the server (130) may adjust the bitrate to overcome the degradation of analysis accuracy caused by fluctuations in computing workload.

[0062] According to one embodiment of the present disclosure, the server (130) may determine or select a heavier neural network model to analyze video quality reduced due to network bandwidth reduction with high accuracy. However, when determining the neural network model, the server (130) may determine the neural network model to satisfy the latency requirements per frame.

[0063] According to one embodiment of the present disclosure, the terminal (110) may transmit low-quality video frames due to a reduction in network bandwidth. The server (130) may select a heavier neural network model that can perform video analysis more accurately during the idle time resulting from this, and may select a neural network model such that the time required for video analysis satisfies the latency requirements.

[0064] According to one embodiment of the present disclosure, a server (130) can select a combination of video bitrate and neural network model that satisfies latency requirements based on available network resources (i.e., bandwidth) and available computing resources (i.e., GPU). The operation of selecting a combination of target bitrate and neural network model in consideration of resource variability can be expressed as joint scheduling.

[0065] According to one embodiment of the present disclosure, a server (130) can monitor network resources and computing resources in real time for joint scheduling. The server (130) can monitor available resources for joint scheduling in real time to determine whether fluctuations occur or bottlenecks occur.

[0066] According to one embodiment of the present disclosure, a server (130) can monitor network resources by receiving network status information from a base station (120). The server (130) can monitor available network resources by receiving information related to network resources from the base station (120) for joint scheduling.

[0067] According to one embodiment of the present disclosure, the server (130) can monitor network bandwidth. The server (130) can monitor analysis latency or inference latency for each neural network model running on a GPU within the server (130). The server (130) can monitor whether latency occurs in video frame analysis using the neural network model.

[0068] According to one embodiment of the present disclosure, a server (130) may perform profiling, which is an operation to determine a combination of selectable bitrates and neural network models that satisfy specific accuracy requirements for a video frame. The server (130) may select an optimal bitrate and neural network model based on the profiling results and resource status information.

[0069] According to one embodiment of the present disclosure, the server (130) can perform profiling based on scene changes of video content. To minimize profiling overhead, the server (130) can perform profiling only on frames with scene changes of video content.

[0070] According to one embodiment of the present disclosure, the server (130) can determine a combination of a target bit rate and a neural network model per terminal that minimizes the total network and computing resource usage cost based on latency requirements, available network resources, available computing resources, and real-time network and computing resource change amounts.

[0071] According to one embodiment of the present disclosure, a server (130) can schedule a GPU to satisfy end-to-end (E2E) latency requirements. A base station (120) can schedule a resource block (RB) to satisfy end-to-end (E2E) latency requirements. The server (130) and the base station (120) can coordinate the GPU and RB to satisfy end-to-end (E2E) latency requirements.

[0072] According to one embodiment of the present disclosure, a server (130) and a base station (120) may allocate GPUs and RBs based on a determined target bit rate and neural network model combination, a latency requirement per frame per terminal, and an encoded frame size. At this time, the server (130) and the base station (120) may determine the order of resources allocated to terminals according to time slots.

[0073] However, the joint scheduling and resource allocation method performed considering the variability of the above resources has a problem in that sensitive information is not protected during the information exchange process between the server (130) and the base station (120), and the satisfaction of E2E latency requirements is low because it does not consider the latency variability occurring between the server (130) and the base station (120).

[0074] FIG. 4 is a diagram illustrating latency variability according to one embodiment of the present disclosure.

[0075] Referring to FIG. 4, the latency at the network end and the latency at the computing end may vary depending on time or frame. The latency occurring at the base station (120) and the latency occurring at the server (130) may vary depending on time or frame.

[0076] Referring to Graph 410, network latency can vary on a frame-by-frame basis. Network latency is the delay time that occurs when transmitting and streaming video frames, depending on network bandwidth or video frame size, and can vary on a frame-by-frame basis.

[0077] According to one embodiment of the present disclosure, the latency at the computing unit may have variability on a frame-by-frame basis. The latency at the computing unit is a delay time that occurs depending on the analysis workload, such as the number of objects to be analyzed, when the server (130) infers and analyzes video frames, and may vary on a frame-by-frame basis.

[0078] Referring to graph 420, for example, if the E2E latency requirement is 200ms, the probability of satisfying the latency requirement indoors is 0.3, and the probability of satisfying the latency requirement on a train is 0.66. Scheduling performed without considering latency variability, which may vary depending on the environment in which the terminal (110) is located, can reduce the satisfaction of the latency requirement.

[0079] Previously, scheduling and resource allocation were performed without considering latency variability at the frame level, assuming that the same variability occurs within a certain range of chunks.

[0080] Resource scheduling performed without considering frame-level latency variability occurring at the network level and frame-level latency variability occurring at the computing level can reduce the satisfaction of E2E latency requirements.

[0081] Here, E2E latency may refer to the delay time that occurs from when the terminal (110) transmits a video frame until the server (130) analyzes the video frame using a neural network model. The E2E latency requirement may indicate a target value of the delay time that the E2E latency must satisfy.

[0082] Inaccurate RB and GPU scheduling that fails to account for frame-level latency variability existing at both the network and computing ends may result in a decrease in the satisfaction of application latency requirements. To resolve the above problem, a signaling structure that enables bidirectional information sharing between the RAN and the server and precise resource scheduling will be described in more detail through the drawings and explanations described below.

[0083] FIG. 5 is a diagram illustrating a resource allocation operation for analyzing real-time video according to one embodiment of the present disclosure.

[0084] Referring to FIG. 5, a signaling structure and scheduling method between a terminal (110), a base station (120), and a server (130) to improve the satisfaction of latency and accuracy requirements in analyzing real-time video are described.

[0085] To improve the satisfaction of frame-unit latency requirements, the server (130) can share information bidirectionally with the base station (120), divide E2E latency requirements based on computing latency distribution and network latency distribution, and perform precise resource scheduling.

[0086] The objectives to be achieved through the method proposed in this disclosure are as follows.

[0087] First, we aim to minimize the total operational cost, which is the sum of the cost due to video bitrate at the network level and the cost due to inference latency using neural network models at the computing level.

[0088] The formula for minimizing total operational costs can be expressed as follows.

[0089]

[0090] Here, b i can refer to the bitrate for streaming and transmitting user i's video stream. l i may refer to the latency (i.e., delay) required for the analysis and inference of a video stream using user i's neural network model.

[0091] Here, Cost Network (b) can be expressed as follows.

[0092]

[0093]

[0094] Here, Cost Compute (l) can be expressed as follows.

[0095]

[0096]

[0097] Secondly, we aim to satisfy the latency and accuracy requirements per frame for each terminal. For example, the latency requirement per frame may be 100ms for autonomous driving, and the accuracy requirement may be an mIoU of 0.5 or higher.

[0098] Thirdly, we aim to develop a signaling structure that enables the base station (120) and the server (130) to exchange information for resource scheduling without disclosing sensitive information (proprietary information), thereby allowing for mutually cooperative and separated joint scheduling. In particular, considering the probabilistic distribution and variability of network and computing latency, E2E latency requirements can be divided into computing latency requirements and network latency requirements. This aims to improve resource usage efficiency and the satisfaction of E2E application latency requirements for each terminal.

[0099] In the following, specific embodiments for achieving the above objectives are described.

[0100] In operation S505, the base station (120) may transmit network status information to the server (130). The network status information may include at least one of the total number of resource blocks (RB), the number of resource blocks (RB) allocated per terminal (110), the transport block size (TBS), and the first latency distribution. The network status information is information about the network status, which the base station (120) may provide to the server (130).

[0101] The total number of RBs may refer to the total number of available RBs available to the base station (120). The number of RBs allocated per terminal (110) may refer to the number of RBs allocated to each terminal (110) out of the total number of RBs. TBS is the size of a data transmission unit in a wireless communication system and may refer to the amount of data transmitted through the resources allocated to the terminal (110) in the network (120) during a single transmission cycle. The first latency distribution may refer to a probability distribution that probabilistically models the delay time of data transmission occurring according to the network bandwidth at the base station (120).

[0102] According to one embodiment of the present disclosure, network status information transmitted by a base station (120) to a server (130) may be transmitted in the RIC (RAN Intelligent Controller) message format.

[0103] RIC is a software-defined component of the Open-RAN architecture that provides multi-vendor interoperability, intelligence, agility, and programmability to wireless access networks, and can play a role in controlling and optimizing RAN functions.

[0104] RIC messages can refer to data messages for communication between the RIC, network elements, and RAN components. RIC messages can be transmitted and received through the RIC API (Application Programming Interface).

[0105] A RIC API can refer to an interface that enables external applications or RAN nodes to communicate with the RIC. Through the RIC API, the RIC can provide the exchange of data that allows applications to coordinate RAN functions.

[0106] In operation S510, the server (130) can determine a target bitrate for streaming real-time video and a neural network model for analyzing real-time video based on E2E latency requirements, network status information received from a base station, and computing status information. The entity determining the target bitrate and the neural network model may be a joint scheduler (550).

[0107] A joint scheduler (550) is a hardware device or software included in the server (130) and can perform joint scheduling operations. Here, joint scheduling may refer to the operation of determining or selecting a combination of a target bit rate and a neural network model based on a combination of a target bit rate and a neural network model that satisfies network status information, computing status information, E2E latency requirements, and accuracy requirements received from the base station (120).

[0108] Latency requirements may refer to criteria for the maximum allowable delay time for a system to achieve user experience or specific performance goals. E2E latency requirements may refer to a target value of delay time that must be satisfied by the delay time occurring from the transmission of a video stream from the terminal (110) to the analysis of the video stream at the server (130).

[0109] Requirements can be expressed as Service Level Objectives (SLOs). An SLO can refer to the goal of the service level that a service provider offers to a customer. Requirements may include accuracy requirements and latency requirements.

[0110] According to one embodiment of the present disclosure, computing state information may include at least one of the total number of GPUs and a second latency distribution. The computing state information may be state information regarding compute at the server (130) and may be information held by the server (130).

[0111] The second latency distribution may refer to a probability distribution that probabilistically models the analysis and inference latency of a neural network model occurring on the server (130) depending on the analysis workload or GPU resources.

[0112] A neural network model can refer to an artificial intelligence (AI) model that processes temporal and spatial information of video frames. Deep neural network models may be used for video frame analysis and inference. For example, neural network models may include, but are not limited to, Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Long Short-Term Memory (LSTM), and Spatial-Temporal Graph Neural Networks (ST-GNN).

[0113] According to one embodiment of the present disclosure, the server (130) can select a combination of a target bit rate and a neural network model that satisfies E2E latency requirements based on network state information and computing state information from among combinations of a target bit rate and a neural network model that satisfy accuracy requirements.

[0114] A combination of target bitrate and neural network model that satisfies the accuracy requirements may be a combination determined by considering only the accuracy requirements without considering the E2E latency requirements, and may exist in the form of a list, table, or set.

[0115] According to one embodiment of the present disclosure, a server (130) (or a joint scheduler (550)) may receive as input a combination of a target bitrate and a neural network model that satisfies an E2E latency requirement, network status information, computing status information, and accuracy requirement, and output a combination of a target bitrate and a neural network model for each terminal.

[0116] According to one embodiment of the present disclosure, a server (130) (or joint scheduler (550)) can determine a target bitrate for video frame streaming for each terminal (110) by considering E2E latency requirements, RB resource information within network status information and a first latency distribution.

[0117] According to one embodiment of the present disclosure, a server (130) (or joint scheduler (550)) can determine a neural network model for video frame analysis for each terminal (110) by considering E2E latency requirements, GPU resource information within computing state information, and a second latency distribution.

[0118] According to one embodiment of the present disclosure, a server (130) (or joint scheduler (550)) can determine a combination of a target bit rate and a neural network model that satisfies the E2E latency requirement for each terminal (110) by considering RB resource information and a first latency distribution in network state information and GPU resource information and a second latency distribution in computing state information.

[0119] According to one embodiment of the present disclosure, the server (130) (or joint scheduler (550)) may select a combination of a determined target bit rate and a neural network model from among combinations of a target bit rate and a neural network model that satisfy accuracy requirements.

[0120] In operation S515, the server (130) can transmit a target bitrate to the terminal (110). The server (130) can transmit a target bitrate determined or selected through the joint scheduler (550) to the terminal (110). The terminal (130) can encode a video frame based on the target bitrate. The target bitrate transmitted by the server (130) to the terminal (110) may be transmitted in the form of a frame transmission request message (or a frame size transmission request message, an encoding frame information transmission request message), but the name or form of the transmitted message is not limited thereto.

[0121] In operation S520, the server (130) can divide the E2E latency requirements into the first latency requirements at the base station and the second latency requirements at the server based on the target bit rate, the neural network model, the first latency distribution at the base station, and the second latency distribution at the server. The entity dividing the E2E latency requirements may be a latency requirement splitter (560).

[0122] A latency requirement splitter (560) is a hardware device or software included in the server (130) and can divide the E2E latency requirement into a first latency requirement and a second latency requirement by considering the latency distribution at the network end and the latency distribution at the computing end.

[0123] The first latency requirement may refer to a target value of latency that must be satisfied by the delay time occurring at the base station (120) (i.e., the network stage). The second latency requirement may refer to a target value of latency that must be satisfied by the delay time occurring at the server (130) (i.e., the computing stage).

[0124] According to one embodiment of the present disclosure, a server (130) (or a latency requirement splitter (560)) may receive a target bit rate and a neural network model determined by a joint scheduler (550), a first latency distribution and a second latency distribution as inputs, and output a first latency requirement and a second latency requirement.

[0125] According to one embodiment of the present disclosure, a server (130) (or a latency requirement splitter (560)) can determine a first latency requirement by calculating the time a video stream is streamed at a determined target bitrate and calculating how much time should be allocated to the network end by taking into account the network latency distribution (i.e., variability) accordingly.

[0126] According to one embodiment of the present disclosure, a server (130) (or a latency requirement splitter (560)) can determine a second latency requirement by calculating the time at which a video stream is analyzed (or inferred) using a determined neural network model, and calculating how much time should be allocated to the computing unit by considering the computing latency distribution (i.e., variability) accordingly.

[0127] In operation S525, the terminal (110) can encode a video frame based on a target bitrate received from the server (130). The terminal (110) can encode a video to match the target bitrate.

[0128] In operation S530, the terminal (110) can transmit encoding frame information to the server (130). According to one embodiment of the present disclosure, the encoding frame information may include the frame size of a video frame encoded by the terminal (110) based on a target bitrate.

[0129] According to one embodiment of the present disclosure, the encoding frame information transmitted by the terminal (110) to the server (130) may be transmitted in the form of a response to a frame transmission request message (or a frame size transmission request message, an encoding frame information transmission request message).

[0130] In operation S535, the server (130) can transmit a first latency requirement and encoding frame information received from the terminal (110) to the base station (120). The first latency requirement and the encoding frame information transmitted from the server (130) to the base station (120) can be transmitted in an RIC message format.

[0131] In operation S540, the base station (120) can determine a resource block to be allocated to the terminal (110) based on the first latency requirement, encoding frame information, and network status information. The base station (120) can allocate the resource block determined based on the first latency requirement, encoding frame information, and network status information to the terminal (110). The entity allocating the resource block to the terminal (110) may be an RB scheduler (570).

[0132] The RB scheduler (570) is a hardware device or software included within the base station (120) and can allocate resource blocks to the terminal (110) based on the first latency requirement, encoding frame information, and network status information. For example, the RB scheduler (570) may include a PF (Proportional Fair) scheduler, but is not limited thereto.

[0133] According to one embodiment of the present disclosure, a base station (120) (or RB scheduler (570)) can determine which resource block to allocate to which terminal (110) per time slot based on a first latency requirement, encoding frame information, and network state information. The base station (120) (or RB scheduler (570)) can determine which resource block to allocate to which terminal (110) for which transmission time interval (TTI) based on a first latency requirement, encoding frame information, and network state information.

[0134] According to one embodiment of the present disclosure, a base station (120) (or RB scheduler (570)) can determine the order of which resource blocks to allocate to which terminal (110) in which time slot (or TTI) based on a first latency requirement, encoding frame information, and network state information.

[0135] According to one embodiment of the present disclosure, a base station (120) (or RB scheduler (570)) receives a first latency requirement and encoding frame information from a server (130) and can allocate a resource block to a terminal (110) to satisfy the first latency requirement.

[0136] According to one embodiment of the present disclosure, a base station (120) (or RB scheduler (570)) may receive a first latency requirement, encoding frame information, and network status information as inputs and output a resource block to be allocated to a terminal (110).

[0137] According to one embodiment of the present disclosure, a base station (120) (or RB scheduler (570)) can determine the number of RBs required for each terminal (110) by considering video frame size information encoded by each terminal (110) according to the target bit rate, streaming processing time according to the video frame size, RB resource information in network status information, a first latency distribution in network status information, and how much latency requirement must be satisfied for each terminal (110), and can allocate resource blocks determined for each terminal (110).

[0138] According to one embodiment of the present disclosure, a base station (120) may transmit an RB allocation result to a terminal (110). The RB allocation result may include information regarding the number of resource blocks allocated to the terminal (110) by the base station (120) (or the RB scheduler (570)). The RB allocation result may include information regarding which resource blocks the base station (120) (or the RB scheduler (570)) allocated to which terminal (110) in which time slot (or TTI).

[0139] According to one embodiment of the present disclosure, a base station (120) (or RB scheduler (570)) may transmit an RB allocation result to each terminal (110) that includes information on which resource block is allocated to each terminal (110) in which time slot (or TTI).

[0140] In operation S545, the terminal (110) may transmit an encoded video frame to the server (130). According to one embodiment of the present disclosure, the video frame transmitted by the terminal (110) to the server (130) may be transmitted in the form of a video analysis request message (or video inference request message), but the name or form of the transmitted message is not limited thereto.

[0141] In operation S550, the server (130) may allocate a GPU to the terminal (110) based on the second latency requirement and computing state information. The entity allocating the GPU to the terminal (110) may be a GPU scheduler (580).

[0142] The GPU scheduler (580) is a hardware device or software included in the server (130) and can allocate a GPU to the terminal (110) based on the second latency requirement, computing state information, and video frames. For example, the GPU scheduler (580) may include a finish-time fairness scheduler, but is not limited thereto.

[0143] According to one embodiment of the present disclosure, a server (130) (or a GPU scheduler (580)) may allocate a GPU to a terminal (110) based on a video frame received from the terminal (110). The server (130) (or a GPU scheduler (580)) may allocate a GPU to a terminal (110) based on a second latency requirement, computing state information, and a video frame.

[0144] According to one embodiment of the present disclosure, a server (130) (or GPU scheduler (580)) may determine which GPU to allocate to which terminal (110) per time slot based on a second latency requirement, computing state information, and video frames. The server (130) (or GPU scheduler (580)) may determine which GPU to allocate to which terminal (110) for which TTI based on a second latency requirement, computing state information, and video frames.

[0145] According to one embodiment of the present disclosure, the server (130) (or GPU scheduler (580)) may allocate a GPU to the terminal (110) to satisfy a second latency requirement based on computing state information and a video frame received from the terminal (110).

[0146] According to one embodiment of the present disclosure, a server (130) (or a GPU scheduler (580)) may receive a second latency requirement, computing state information, and a video frame as inputs and output a GPU to be allocated to a terminal (110).

[0147] According to one embodiment of the present disclosure, a server (130) (or a GPU scheduler (580)) can determine the number of GPUs required for each terminal (110) by considering a video frame, a computing workload according to the video frame, GPU resource information in computing state information, a second latency distribution in computing state information, and how much latency requirement must be satisfied for each terminal (110), and can allocate the determined GPUs for each terminal (110).

[0148] According to one embodiment of the present disclosure, the server (130) (or GPU scheduler (580)) can determine the order of which GPU to allocate to which terminal (110) in which time slot (or TTI) based on a second latency requirement, computing state information, and a video frame.

[0149] According to one embodiment of the present disclosure, a server (130) (or GPU scheduler (580)) can analyze or infer video frames based on the GPU assigned to each terminal (110). According to one embodiment of the present disclosure, a server (130) (or GPU scheduler (580)) can analyze or infer video frames received from each terminal (110) based on the GPU assigned to each terminal (110) per time slot (or TTI).

[0150] FIG. 6 is a drawing illustrating a method for protecting sensitive information according to one embodiment of the present disclosure.

[0151] Referring to FIG. 6, information transmitted and received between the server (130) and the base station (120) can be converted into non-sensitive information and transmitted and received.

[0152] There is a problem in that the proprietary information of different operators (i.e., network / service operators) is not protected during the signaling process for the combination of target bitrate and neural network model and integrated RB and GPU scheduling.

[0153] When the server (130) and the base station (120) exchange information, the unconverted raw data can be transmitted by operating for the purpose of minimizing costs and average throughput.

[0154] A network operator may refer to an organization or company that owns and manages a communication network and is responsible for providing network-based services such as mobile communication, the internet, and data services, and may be a role performed at a base station (120). For example, a network operator may include AT&T, Verizon, etc., but is not limited thereto.

[0155] Referring to 610, the proprietary information of a network operator transmitted from a base station (120) to a server (130) may include, for example, information such as a physical layer protocol, SINR prediction and modulation coding scheme (MCS) modulation / selection algorithm, RB scheduling algorithm, and channel state estimation information. However, it is not limited thereto.

[0156] A service operator may refer to an organization or enterprise that provides and manages various application services to users based on network infrastructure, and may be a role performed on a server (130). A service operator may cooperate with a network operator or provide services through its own infrastructure. For example, a service operator may include, but is not limited to, Microsoft and AWS.

[0157] Referring to 620, the proprietary information of a service operator transmitted from the server (130) to the base station (120) may include, for example, information such as a DNN backbone, video codec information, GPU scheduling algorithm, video content, video encoding algorithm, DNN task type, and DNN model. However, it is not limited thereto.

[0158] According to one embodiment of the present disclosure, a server (130) and a base station (120) can convert confidential information into non-sensitive information and transmit or receive it in order to protect proprietary information. The server (130) and the base station (120) can protect proprietary information by transmitting only the output of the corresponding implementation instead of the implementation information which is proprietary information.

[0159] According to one embodiment of the present disclosure, network status information, a first latency requirement, and encoding frame information transmitted and received between a server (130) and a base station (120) can be transmitted and received in an RIC message format.

[0160] According to one embodiment of the present disclosure, information transmitted and received between a server (130) and a base station (120) can be implemented on an RIC. According to one embodiment of the present disclosure, information transmitted and received between a server (130) and a base station (120) can be transmitted and received through an RIC API.

[0161] According to one embodiment of the present disclosure, information transmitted and received between a server (130) and a base station (120) can be transmitted and received in the form of an RIC message by mapping sensitive information (proprietary information) to non-sensitive information.

[0162] RIC messages implemented on the O-RAN RIC can guarantee latency of less than 1-10ms through server (130) side RAN (120) monitoring and control, eliminate the need for firmware updates, and allow for the configuration of flexible message formats and interfaces. For example, they can be implemented using RIC message formats based on the Key Performance Metric (KPM) and RAN Control (RC) Service Model (SM) of the E2 interface, or based on Microsoft’s Programmable RAN with dynamic SM.

[0163] FIG. 7 is a diagram illustrating a scheduling method based on latency probability modeling and a latency requirement partitioning method according to one embodiment of the present disclosure.

[0164] Referring to Fig. 7, network / computing latency can be modeled probabilistically. Network / computing latency can be represented by a probability distribution that is modeled probabilistically.

[0165] Network latency may refer to the delay time of data transmission that occurs at a base station (120) depending on network bandwidth, video frame size, etc. Network latency may be expressed as a first latency.

[0166] Computing latency may refer to the delay in analysis and inference of a neural network model that occurs depending on the computing workload or GPU resources at the server (130). Computing latency may be expressed as a second latency.

[0167] According to one embodiment of the present disclosure, network / computing latency can be modeled by a Gaussian distribution. However, it is not limited thereto, and can be modeled by a binomial distribution, a Poisson distribution, an exponential distribution, a gamma distribution, a Bernoulli distribution, a multinomial distribution, a chi-square distribution, a beta distribution, a Dirichlet distribution, a logistic distribution, and a Weibull distribution, etc.

[0168] According to one embodiment of the present disclosure, modeling of network / computing latency per terminal (110) may include receiving measured values ​​of network latency and computing latency as input and modeling the network latency and computing latency as a Gaussian distribution. The network / computing latency Gaussian distribution can probabilistically express the variability of the latency.

[0169] According to one embodiment of the present disclosure, a network latency distribution or a network latency Gaussian distribution may be represented by a first latency distribution. A computing latency distribution or a computing latency Gaussian distribution may be represented by a second latency distribution.

[0170] According to one embodiment of the present disclosure, since the network latency Gaussian distribution and the computing latency Gaussian distribution are independent of each other, the E2E latency, which is the sum of the two latencies, can also be modeled as a Gaussian distribution. The distribution modeled as a Gaussian distribution for the E2E latency, which is the sum of the two latencies, can be expressed as an E2E latency distribution or an E2E latency Gaussian distribution.

[0171] Referring to 710, based on the modeled network / computing latency distribution, a target bitrate and a neural network model that maximize the satisfaction of E2E latency requirements can be determined.

[0172] Based on the modeled network / computing latency distribution, the formula for determining the target bitrate and neural network model that maximizes the satisfaction of E2E latency requirements can be expressed as follows.

[0173]

[0174] Here, L i can mean the E2E latency of user i, and L max This may mean an E2E latency requirement.

[0175] Referring to 710, the combination of two bitrates and DNN models is a combination that satisfies the accuracy requirements, and the server (130) can select a combination of bitrates and DNN models that minimizes variability. By selecting a combination that minimizes variability, the satisfaction of the E2E latency requirements can be maximized.

[0176] Referring to 720, the network / computing latency distribution for each terminal (110) can be taken into account to maximize the satisfaction of the latency requirements of each network / computing stage, and the E2E latency requirements can be divided into network latency requirements and computing latency requirements.

[0177] The formula for dividing E2E latency requirements to maximize the satisfaction of latency requirements for each network / computing stage can be expressed as follows.

[0178]

[0179] Here, L i,net can mean the delay time in user i's network (120), and L i,net,max This may mean the latency requirement in the network (120) of user i.

[0180] Here, L i,comp may mean the delay time at the computing unit (130) of user i, and L i,comp,max This may mean the latency requirement at the computing unit (130) of user i.

[0181] Referring to 720, the server (130) can assign more latency requirements to stages with more estimated variability. If a computing stage is estimated to have more variability, the satisfaction of latency requirements at each stage can be maximized by assigning more computing latency requirements.

[0182] According to one embodiment of the present disclosure, a base station (120) and a server (130) can schedule RBs and GPUs based on a determined bit rate, a combination of DNN models, and a determined network / computing latency requirement.

[0183] FIG. 8 is a diagram illustrating a resource allocation operation that takes into account a change in latency distribution according to one embodiment of the present disclosure.

[0184] Referring to FIG. 8, the base station (120) and the server (130) can detect the latency distribution of the network and computing stages in real time and perform scheduling and resource allocation considering latency variability.

[0185] In operation S810, the server (130) can measure a first latency and a second latency. The first latency may refer to network latency, and the second latency may refer to computing latency. The server (130) can measure and observe the first latency and the second latency for each terminal (110).

[0186] According to one embodiment of the present disclosure, a server (130) can measure and observe a first latency per hour and a second latency per hour. The server (130) can measure and observe the first latency and the second latency per hour for each time slot.

[0187] In operation S820, the server (130) can model the first latency and the second latency as a probability distribution. The server (130) can model the first latency and the second latency for each measured terminal (110) as a probability distribution. The first latency and the second latency modeled as a probability distribution can be expressed as a first latency distribution and a second latency distribution. For example, the first latency and the second latency can be modeled as a Gaussian distribution.

[0188] According to one embodiment of the present disclosure, the modeling of the first latency and the second latency for each terminal (110) may include receiving the measured values ​​of the first latency and the second latency as input and modeling the first latency and the second latency as a Gaussian distribution. The latency Gaussian distribution can probabilistically express the variability of the latency.

[0189] According to one embodiment of the present disclosure, since the first latency distribution and the second latency distribution are independent of each other, the sum of the two distributions can be expressed as an E2E latency distribution.

[0190] Here, the E2E latency distribution may refer to a probability distribution that probabilistically models the delay time occurring from when the terminal (110) transmits the video stream until the server (130) analyzes the video stream, and may include a first latency distribution and a second latency distribution.

[0191] According to one embodiment of the present disclosure, a server (130) can receive a network latency measurement, a computing workload measurement, a target bitrate, and a neural network model as inputs and output a network latency distribution, a computing latency distribution, and an E2E latency distribution.

[0192] According to one embodiment of the present disclosure, the server (130) has a latency distribution of 1 Mbps through a network latency measurement value. Calculate and the network latency distribution It can be calculated. Here, μ can represent the mean value, σ can represent the standard deviation (variance), and b can represent the bitrate.

[0193] According to one embodiment of the present disclosure, the server (130) has a computing workload distribution through a computing workload measurement value. Calculate and the computing latency distribution It can be calculated. Here, l can mean the latency value.

[0194] According to one embodiment of the present disclosure, the E2E latency distribution is the sum of the network latency distribution and the computing latency distribution. It can be calculated. Since the network latency Gaussian distribution and the computing latency Gaussian distribution are independent, the E2E latency distribution, which is the sum of the two, can also follow a Gaussian distribution.

[0195] According to one embodiment of the present disclosure, the server (130) measures the video frame size ( By checking the Gaussian distribution of ), and using a QQ plot (Quantile-Quantile Plot), the Gaussian distribution of the measured network latency can be checked, and it can be determined whether the network latency distribution follows a Gaussian distribution based on the similarity between the measured value and the actual Gaussian distribution.

[0196] The Shapiro-Wilk test is one of the methods used in statistics to test for normality, and it refers to a method for verifying whether data follows a normal distribution.

[0197] A QQ plot can be described as a tool that allows one to visually check whether two distributions are similar—that is, whether one distribution closely follows the other—by comparing them.

[0198] According to one embodiment of the present disclosure, the server (130) can use a QQ plot to check the Gaussian distribution of the measured inference workload and determine whether the computing latency distribution follows a Gaussian distribution based on the similarity between the measured value and the actual Gaussian distribution.

[0199] In operation S830, the server (130) can detect changes in the first latency distribution and the second latency distribution. According to one embodiment of the present disclosure, the server (130) can determine whether to update the target bit rate, the neural network model, the first latency requirement, and the second latency requirement based on the changes in the latency distribution.

[0200] According to one embodiment of the present disclosure, the server (130) can compare a previous latency distribution with a current latency distribution. If a change in the latency distribution is detected, the server (130) may decide to update the target bitrate, the neural network model, the first latency requirement, and the second latency requirement.

[0201] According to one embodiment of the present disclosure, the server (130) may determine whether to update by comparing the new network / computing latency distribution with the network / computing latency distribution of the previous time slot. The server (130) may determine whether to update based on whether the difference between the newly modeled latency distribution and the previous latency distribution, using the measured latency values, exceeds a threshold value.

[0202] According to one embodiment of the present disclosure, the server (130) can detect a change by comparing the distribution of network / computing latency during a given time slot based on a certain time slot and determine whether to update.

[0203] For example, the decision to update can be made based on the number of measured latency values ​​that deviate by more than a certain range from the median of the latency distribution of the previous time slot. Additionally, the decision to update can be made based on the Kullback-Leibler (KL) divergence value between the new distribution modeled with the measured latency and the latency distribution of the previous time slot.

[0204] Kullback-Leibler (KL) divergence is an asymmetric distance concept that measures the difference between two probability distributions, and it can be described as a tool that quantitatively expresses how different one probability distribution is from another.

[0205] The method of comparison based on the Kullback-Leibler divergence value can be expressed by the following formula.

[0206]

[0207]

[0208] In operation S840, if the server (130) decides to update the target bitrate and neural network model, it may update the target bitrate and neural network model based on the changed first latency distribution and the changed second latency distribution. Based on the changed first latency distribution and the changed second latency distribution, the server (130) may determine the target bitrate and neural network model to maximize the satisfaction of E2E latency requirements.

[0209] According to one embodiment of the present disclosure, the server (130) can determine and update the target bitrate and neural network model to maximize the satisfaction of E2E latency requirements based on the E2E latency distribution, which is the sum of the changed first latency distribution and the changed second latency distribution.

[0210] The method for determining the target bitrate and neural network model to maximize E2E latency requirements can be expressed by the following formula.

[0211]

[0212] st

[0213]

[0214] L i,j can refer to the inference latency of user i in frame j. L max can refer to E2E latency requirements. Fps can refer to the video frame rate. l detection can refer to the detector's inference latency. n i can mean the workload of user i.

[0215] t util can mean available utilization time. N GPU can mean the number of GPUs. T can mean the total time. Ri can refer to the frequency efficiency (Mbps per RB) of user i. TTI can refer to the transmission time interval. N RB,total can mean the number of RBs in the scheduling window at time T.

[0216] According to one embodiment of the present disclosure, the server (130) can update the target bitrate and the neural network model based on a combination of the target bitrate and the neural network model satisfying the E2E latency requirement, network state information, computing state information, the changed first latency distribution and the changed second latency distribution, and the accuracy requirement.

[0217] According to one embodiment of the present disclosure, the server (130) can compute an approximate solution using iterative gradient descent to determine the target bitrate and the neural network model. Since the problem of maximizing the product of the probability distribution function (PDF) is not in a closed-form, it can be obtained through an approximate solution.

[0218] According to one embodiment of the present disclosure, the iterative gradient descent method may include the process of determining an optimal bit rate and an optimal neural network model that maximize the probability of satisfying latency requirements for each terminal (110), and, when the amount of available resources is exceeded, adjusting the bit rate and neural network model such that the reduction amount of the probability of satisfying latency requirements is minimized.

[0219] According to one embodiment of the present disclosure, the reduction in the probability of satisfying a latency requirement can be expressed by the following formula representing the difference in the probability of satisfying a latency requirement between two bitrate and neural network model combinations.

[0220]

[0221] According to one embodiment of the present disclosure, the entity determining (or updating) the target bitrate and neural network model based on the changed latency distribution may be a joint scheduler. The server (130) (or joint scheduler) may receive as input a combination of the target bitrate and neural network model satisfying the E2E latency requirement, network state information, computing state information, the changed first latency distribution, the changed second latency distribution, and the accuracy requirement, and output an updated target bitrate and an updated neural network model.

[0222] The updated target bitrate and the updated neural network model may differ from the target bitrate and the neural network model of the previous time slot, and can maximize the satisfaction of E2E latency requirements based on changes in the network / computing latency distribution.

[0223] In operation S850, if the server (130) decides to update the first latency requirement and the second latency requirement, it can divide the E2E latency requirement into the updated first latency requirement and the updated second latency requirement based on the changed first latency distribution and the changed second latency distribution.

[0224] According to one embodiment of the present disclosure, the server (130) can divide the E2E latency requirements based on the changed first latency distribution and the changed second latency distribution to maximize the probability of satisfying the first latency requirement and the second latency requirement.

[0225] According to one embodiment of the present disclosure, the server (130) can divide the E2E latency requirement into an updated first latency requirement and an updated second latency requirement based on an updated target bit rate, an updated neural network model, a changed first latency distribution, and a changed second latency distribution.

[0226] According to one embodiment of the present disclosure, the server (130) can calculate an approximate solution using iterative gradient descent to divide the E2E latency requirements. Since the problem of maximizing the product of the probability distribution function (PDF) is not in a closed-form, it can be obtained through an approximate solution.

[0227] For example, a method for splitting E2E latency requirements is the variability of network latency and computing latency ( It may include a method of dividing in proportion to ).

[0228] In one embodiment, a formula for dividing E2E latency requirements in proportion to the variability of network latency and computing latency can be expressed as follows.

[0229]

[0230] where ,

[0231] ,

[0232] According to one embodiment of the present disclosure, the entity splitting the E2E latency requirements may be a latency requirement splitter. The server (130) (or latency requirement splitter) may receive an updated target bitrate, an updated neural network model, a changed first latency distribution, a changed second latency distribution, and an E2E latency requirement as inputs, and output an updated first latency requirement and an updated second latency requirement.

[0233] The updated first latency requirement and the updated second latency requirement may differ from the first latency requirement and the previous second latency requirement of the previous time slot, and the satisfaction of the network latency requirement and computing latency requirement per terminal can be maximized based on changes in the network / computing latency distribution.

[0234] In operation S860, if the server (130) decides not to update the target bit rate, neural network model, first latency requirement, and second latency requirement, the target bit rate, neural network model, first latency requirement, and second latency requirement may be reused. If the server (130) decides not to update, the previously determined target bit rate, previously determined neural network model, previously divided first latency requirement, and previously divided second latency requirement may be reused to perform a resource allocation operation.

[0235] In operation S870, if an update is performed, the base station (120) and the server (130) can allocate RB and GPU to the terminal (110) based on the updated target bit rate, the updated neural network model, the updated first latency requirement, and the updated second latency requirement.

[0236] In operation S870, if no update is performed, the base station (120) and the server (130) may allocate RB and GPU to the terminal (110) based on a previously determined target bit rate, a previously determined neural network model, a previously divided first latency requirement, and a previously divided second latency requirement.

[0237] According to one embodiment of the present disclosure, operations S810 through S870 can be repeated.

[0238] FIG. 9 is a flowchart illustrating a resource allocation operation for real-time video analysis according to one embodiment of the present disclosure.

[0239] With reference to FIG. 9, the operations for real-time video analysis of the terminal (110), base station (120), and server (130) are schematically described, and since detailed descriptions of each operation have been described in previous drawings, duplicate content is omitted.

[0240] In operation S910, the server (130) can determine a target bitrate for streaming real-time video and a neural network model for analyzing real-time video based on end-to-end latency requirements, network status information received from a base station, and computing status information.

[0241] According to one embodiment of the present disclosure, a server (130) may receive network status information from a base station (120). According to one embodiment of the present disclosure, the server (130) may transmit a target bitrate to a terminal (110). According to one embodiment of the present disclosure, the encoding frame information may include the frame size of a video frame encoded by the terminal (110) based on the target bitrate.

[0242] According to one embodiment of the present disclosure, network state information may include at least one of the total number of resource blocks, the number of resource blocks allocated per terminal (110), the transport block size (TBS), and the first latency distribution. According to one embodiment of the present disclosure, computing state information may include at least one of the total number of GPUs and the second latency distribution.

[0243] According to one embodiment of the present disclosure, network status information, a first latency requirement, and encoding frame information transmitted and received between a server (130) and a base station (120) may be transmitted and received in the form of an RIC (Radio Access Network Intelligent Controller) message.

[0244] According to one embodiment of the present disclosure, the server (130) can detect changes in the first latency distribution and the second latency distribution. According to one embodiment of the present disclosure, the server (130) can determine whether to update the target bit rate, the neural network model, the first latency requirement, and the second latency requirement based on changes in the first latency distribution and the second latency distribution.

[0245] According to one embodiment of the present disclosure, when the server (130) decides to update the target bitrate and the neural network model, it can update the target bitrate and the neural network model based on the changed first latency distribution and the changed second latency distribution.

[0246] According to one embodiment of the present disclosure, the server (130) can select a combination of a target bit rate and a neural network model that satisfies E2E latency requirements based on network state information and computing state information from among combinations of a target bit rate and a neural network model that satisfy accuracy requirements.

[0247] In operation S920, the server (130) can divide the E2E latency requirements into a first latency requirement at the base station and a second latency requirement at the server based on the target bit rate, the neural network model, the first latency distribution at the base station, and the second latency distribution at the server.

[0248] According to one embodiment of the present disclosure, when the server (130) decides to update the first latency requirement and the second latency requirement, it may divide the E2E latency requirement into the updated first latency requirement and the updated second latency requirement based on the changed first latency distribution and the changed second latency distribution.

[0249] According to one embodiment of the present disclosure, if the server (130) decides not to update the target bit rate, neural network model, first latency requirement, and second latency requirement, the target bit rate, neural network model, first latency requirement, and second latency requirement may be reused.

[0250] In operation S930, the server (130) can transmit the first latency requirement and the encoding frame information received from the terminal (110) to the base station (120).

[0251] In operation S940, the terminal (110) may be allocated a resource block determined based on the first latency requirement, encoding frame information, and network status information from the base station (120). The base station (120) may determine the resource block based on the first latency requirement, encoding frame information, and network status information and allocate the determined resource block to the terminal (110).

[0252] According to one embodiment of the present disclosure, the server (130) can receive a video frame from the terminal (110).

[0253] In operation S950, the server (130) may allocate a GPU (graphic processing unit) to the terminal based on the second latency requirement and computing state information. According to one embodiment of the present disclosure, the server (130) may allocate a GPU to the terminal (110) based on a video frame received from the terminal (110).

[0254] Through the present disclosure, information security issues between different operators that occur when cooperatively performing network-computing joint scheduling can be resolved through a signaling structure in which the RAN (120) and the server (130) do not share sensitive information with each other. In addition, by probabilistically modeling the variability of latency, high stability is provided in various real-world network situations, and it is expected that this will enable precise resource usage optimization and satisfy frame-unit latency requirements for each terminal (110).

[0255] FIG. 10 is a block diagram illustrating the configuration of a terminal (110), a base station (120), and a server (130) according to one embodiment of the present disclosure.

[0256] Referring to FIG. 10, the terminal (110) may include a processor (1010) and a transceiver (1020). However, the components of the terminal (110) are not limited to the examples described above. For example, the terminal (110) may include more components or fewer components than the components described above. In addition, the processor (1010) and the transceiver (1020) may be implemented in the form of a single chip.

[0257] According to one embodiment of the present disclosure, a processor (1010) can control a series of processes that allow a terminal (110) to operate according to an embodiment of the present disclosure. For example, in a wireless communication system according to an embodiment of the present disclosure, components of the terminal (110) can be controlled to transmit and receive a target bit rate, encoding frame information, RB allocation result, and video frame, and to encode a video frame. There may be a plurality of processors (1010), and the processor (1010) can perform the operation of transmitting and receiving signals and encoding video frames in the wireless communication system of the present disclosure by executing a program stored in memory (not shown).

[0258] According to one embodiment of the present disclosure, the transceiver (1020) can transmit and receive signals with a base station (120) and a server (130). The signals transmitted and received with the base station (120) and the server (130) may include control information and data. The transceiver (1020) may be composed of an RF transmitter that up-converts and amplifies the frequency of a transmitted signal, and an RF receiver that low-noise amplifies a received signal and down-converts the frequency. However, such a transceiver (1020) is merely one embodiment, and the components of the transceiver (1020) are not limited to an RF transmitter and an RF receiver. Additionally, the transceiver (1020) may receive a signal through a wireless channel and output it to a processor (1010), and transmit the signal output from the processor (1010) through a wireless channel.

[0259] According to one embodiment of the present disclosure, a base station (120) may include a processor (1030) and a transceiver (1040). However, the components of the base station (120) are not limited to the examples described above. For example, the base station (120) may include more components or fewer components than the components described above. In addition, the processor (1030) and the transceiver (1040) may be implemented in the form of a single chip.

[0260] According to one embodiment of the present disclosure, a processor (1030) can control a series of processes that allow a base station (120) to operate according to an embodiment of the present disclosure. For example, in a wireless communication system according to an embodiment of the present disclosure, components of the base station (120) can be controlled to transmit and receive network status information, encoding frame information, a first latency requirement, RB allocation results, etc., and to allocate resource blocks. There may be multiple processors (1030), and the processor (1030) can perform operations of transmitting and receiving signals and allocating resource blocks in the wireless communication system of the present disclosure by executing a program stored in memory (not shown).

[0261] According to one embodiment of the present disclosure, the transceiver (1040) can transmit and receive signals with the terminal (110) and the server (130). The signals transmitted and received with the terminal (110) and the server (130) may include control information and data. The transceiver (1040) may be composed of an RF transmitter that up-converts and amplifies the frequency of a transmitted signal, and an RF receiver that low-noise amplifies a received signal and down-converts the frequency. However, such a transceiver (1040) is merely one embodiment, and the components of the transceiver (1040) are not limited to an RF transmitter and an RF receiver. Additionally, the transceiver (1040) may receive a signal through a wireless channel and output it to a processor (1030), and transmit the signal output from the processor (1030) through a wireless channel.

[0262] According to one embodiment of the present disclosure, the server (130) may include a processor (1050) and a transceiver (1060). However, the components of the server (130) are not limited to the examples described above. For example, the server (130) may include more components or fewer components than the components described above. In addition, the processor (1050) and the transceiver (1060) may be implemented in the form of a single chip.

[0263] According to one embodiment of the present disclosure, a processor (1050) can control a series of processes that allow a server (130) to operate according to an embodiment of the present disclosure. For example, in a wireless communication system according to an embodiment of the present disclosure, components of the server (130) can be controlled to transmit and receive network status information, a target bit rate, encoding frame information, a first latency requirement, a video frame, etc., perform joint scheduling, divide latency requirements, and allocate a GPU. There may be multiple processors (1050), and the processor (1050) can perform operations of transmitting and receiving signals, performing joint scheduling, dividing latency requirements, and allocating a GPU in the wireless communication system of the present disclosure by executing a program stored in memory (not shown).

[0264] According to one embodiment of the present disclosure, the transceiver (1060) can transmit and receive signals with the terminal (110) and the base station (120). The signals transmitted and received with the terminal (110) and the base station (120) may include control information and data. The transceiver (1060) may be composed of an RF transmitter that up-converts and amplifies the frequency of the transmitted signal, and an RF receiver that low-noise amplifies the received signal and down-converts the frequency. However, such a transceiver (1060) is merely one embodiment, and the components of the transceiver (1060) are not limited to an RF transmitter and an RF receiver. Additionally, the transceiver (1060) may receive a signal through a wireless channel and output it to a processor (1050), and transmit the signal output from the processor (1050) through a wireless channel.

[0265] The present disclosure relates to a method and apparatus for allocating resources for real-time video analysis in a wireless communication system. The technical problems to be solved by the present disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art from the description in this specification.

[0266] According to one aspect of the present disclosure, a method performed by a server in a wireless communication system may be provided.

[0267] The above method may include the step of transmitting the target bit rate to the terminal.

[0268] The above encoding frame information may include the frame size of a video frame encoded by the terminal based on the target bitrate.

[0269] The above method may include the step of receiving a video frame from the terminal.

[0270] The above method may include the step of allocating the GPU to the terminal based on the video frame received from the terminal.

[0271] The above network status information may include at least one of the total number of resource blocks, the number of resource blocks allocated per terminal, the transport block size (TBS), and the first latency distribution.

[0272] The computing state information may include at least one of the total number of GPUs and the second latency distribution.

[0273] The network status information, the first latency requirement, and the encoding frame information transmitted and received between the server and the base station can be transmitted and received in the form of an RIC (Radio Access Network Intelligent Controller) message.

[0274] The above method may include the step of detecting a change in the first latency distribution and the second latency distribution.

[0275] The above method may include a step of determining whether to update the target bit rate, the neural network model, the first latency requirement, and the second latency requirement based on the change in the first latency distribution and the second latency distribution.

[0276] The above method may include the step of updating the target bitrate and the neural network model based on the changed first latency distribution and the changed second latency distribution when it is decided to update the target bitrate and the neural network model.

[0277] The above method may include the step of dividing the E2E latency requirement into the updated first latency requirement and the updated second latency requirement based on the changed first latency distribution and the changed second latency distribution when it is decided to update the first latency requirement and the second latency requirement.

[0278] The above method may further include the step of reusing the target bit rate, the neural network model, the first latency requirement, and the second latency requirement when it is decided not to update the target bit rate, the neural network model, the first latency requirement, and the second latency requirement.

[0279] The above method may include the step of selecting a combination of a target bitrate and a neural network model that satisfies the E2E latency requirement based on the network state information and the computing state information among combinations of a target bitrate and a neural network model that satisfy the accuracy requirement.

[0280] According to one aspect of the present disclosure, a server may be provided in a wireless communication system.

[0281] The above server may include a transmitting and receiving unit; and at least one processor coupled to the transmitting and receiving unit.

[0282] The above at least one processor can transmit the target bit rate to the terminal.

[0283] The above at least one processor can receive a video frame from the terminal.

[0284] The above at least one processor can allocate the GPU to the terminal based on the video frame received from the terminal.

[0285] The above at least one processor can detect changes in the first latency distribution and the second latency distribution.

[0286] The above at least one processor can determine whether to update the target bit rate, the neural network model, the first latency requirement, and the second latency requirement based on the change in the first latency distribution and the second latency distribution.

[0287] When the above at least one processor decides to update the target bit rate and the neural network model, it can update the target bit rate and the neural network model based on the changed first latency distribution and the changed second latency distribution.

[0288] If the above at least one processor decides to update the first latency requirement and the second latency requirement, it may divide the E2E latency requirement into the updated first latency requirement and the updated second latency requirement based on the changed first latency distribution and the changed second latency distribution.

[0289] If the above at least one processor decides not to update the target bit rate, the neural network model, the first latency requirement, and the second latency requirement, the target bit rate, the neural network model, the first latency requirement, and the second latency requirement may be reused.

[0290] The above at least one processor can select a combination of a target bit rate and a neural network model that satisfies the E2E latency requirement based on the network state information and the computing state information among combinations of a target bit rate and a neural network model that satisfy the accuracy requirement.

[0291] In the specific embodiments of the present disclosure described above, the components included in the invention are expressed in a singular or plural form according to the specific embodiments presented. However, the singular or plural expression is selected to suit the situation presented for convenience of explanation, and the present disclosure is not limited to singular or plural components; even if a component is expressed in the plural form, it may be composed of a singular form, or even if a component is expressed in the singular form, it may be composed of a plural form.

[0292] Meanwhile, although specific embodiments have been described in the detailed description of the present disclosure, it is understood that various modifications are possible within the scope of the present disclosure. Therefore, the scope of the present disclosure should not be limited to the described embodiments, but should be defined by the claims set forth below as well as equivalents thereof.

[0293] Methods according to the embodiments described in the claims or specification of the present disclosure may be implemented in the form of hardware, software, or a combination of hardware and software.

[0294] When implemented in software, a computer-readable storage medium may be provided for storing one or more programs (software modules). One or more programs stored in the computer-readable storage medium are configured for execution by one or more processors within an electronic device. One or more programs include instructions that cause the electronic device to execute methods according to the embodiments described in the claims or specification of this disclosure.

[0295] Such programs (software modules, software) may be stored in random access memory, non-volatile memory including flash memory, ROM (Read Only Memory), Electrically Erasable Programmable Read Only Memory (EEPROM), magnetic disc storage devices, Compact Disc-ROM (CD-ROM), Digital Versatile Discs (DVDs), or other forms of optical storage devices, magnetic cassettes. Alternatively, they may be stored in memory composed of some or all of these. Additionally, each constituent memory may include multiple units.

[0296] Additionally, the above program may be stored on an attachable storage device that can be accessed via a communication network such as the Internet, Intranet, Local Area Network (LAN), Wide LAN (WLAN), or Storage Area Network (SAN), or a combination thereof. Such a storage device may be connected to a device performing an embodiment of the present disclosure through an external port. Additionally, a separate storage device on a communication network may be connected to a device performing an embodiment of the present disclosure.

[0297] In the specific embodiments of the present disclosure described above, the components included in the disclosure are expressed in a singular or plural form according to the specific embodiments presented. However, the singular or plural expression is selected to suit the situation presented for convenience of explanation, and the present disclosure is not limited to singular or plural components; even if a component is expressed in the plural form, it may be composed of a singular form, or even if a component is expressed in the singular form, it may be composed of a plural form.

[0298] Meanwhile, although specific embodiments have been described in the detailed description of the present disclosure, it is understood that various modifications are possible within the scope of the present disclosure. Therefore, the scope of the present disclosure should not be limited to the described embodiments, but should be defined by the claims set forth below as well as equivalents thereof. In other words, it is obvious to those skilled in the art that other modifications based on the technical concept of the present disclosure are possible. Furthermore, each of the above embodiments may be combined and operated as needed.

Claims

1. In a method performed by a server in a wireless communication system, A step of determining a target bitrate for streaming real-time video and a neural network model for analyzing real-time video based on end-to-end (E2E) latency requirements, network status information received from a base station, and computing status information; A step of dividing the E2E latency requirement into a first latency requirement at the base station and a second latency requirement at the server based on the target bitrate, the neural network model, the first latency distribution at the base station, and the second latency distribution at the server; Transmitting the first latency requirement and encoding frame information received from the terminal to the base station, and the terminal receiving a resource block determined based on the first latency requirement, the encoding frame information, and the network state information from the base station; and A method comprising the step of allocating a GPU (graphic processing unit) to the terminal based on the second latency requirement and the computing state information.

2. In Paragraph 1, The method further includes the step of transmitting the above target bit rate to the terminal, A method in which the above encoding frame information includes the frame size of a video frame encoded by the terminal based on the target bitrate.

3. In Paragraph 1, The method further includes the step of receiving a video frame from the terminal, The step of allocating a GPU to the above terminal is, A method comprising the step of allocating the GPU to the terminal based on the video frame received from the terminal.

4. In Paragraph 1, The above network state information includes at least one of the total number of resource blocks, the number of resource blocks allocated per terminal, the transport block size (TBS), and the first latency distribution. The above computing state information comprises at least one of the total number of GPUs and the second latency distribution.

5. In Paragraph 1, A step of detecting a change in the first latency distribution and the second latency distribution; and A method further comprising the step of determining whether to update the target bit rate, the neural network model, the first latency requirement, and the second latency requirement based on the changes in the first latency distribution and the second latency distribution.

6. In Paragraph 5, The step of determining the above target bitrate and the above neural network model is, A method comprising the step of updating the target bitrate and the neural network model based on the changed first latency distribution and the changed second latency distribution when it is decided to update the target bitrate and the neural network model.

7. In Paragraph 6, The step of dividing the above E2E latency requirement into the above first latency requirement and the above second latency requirement is, A method comprising the step of, when it is decided to update the first latency requirement and the second latency requirement, dividing the E2E latency requirement into the updated first latency requirement and the updated second latency requirement based on the changed first latency distribution and the changed second latency distribution.

8. In a wireless communication system, the server, Transmitter / receiver; and It includes at least one processor coupled to the above-mentioned transmitting and receiving unit, and The above server is, Based on end-to-end (E2E) latency requirements, network status information received from a base station, and computing status information, a target bitrate for streaming real-time video and a neural network model for analyzing the real-time video are determined, and Based on the above target bitrate, the above neural network model, the first latency distribution at the base station, and the second latency distribution at the server, the above E2E latency requirement is divided into the first latency requirement at the base station and the second latency requirement at the server, and The above first latency requirement and encoding frame information received from the terminal are transmitted to the base station, and the terminal is allocated a resource block determined based on the above first latency requirement, the encoding frame information, and the network state information from the base station, and A server that allocates a GPU (graphic processing unit) to the terminal based on the second latency requirement and the computing state information.

9. In Paragraph 8, The above server is, Transmit the above target bitrate to the terminal, and The above encoding frame information includes the frame size of a video frame encoded by the terminal based on the target bitrate, for a server.

10. In Paragraph 8, The above server is, Receive a video frame from the above terminal, and A server that allocates the GPU to the terminal based on the video frame received from the terminal.

11. In Paragraph 8, The above network state information includes at least one of the total number of resource blocks, the number of resource blocks allocated per terminal, the transport block size (TBS), and the first latency distribution. The above computing state information comprises at least one of the total number of GPUs and the second latency distribution, a server.

12. In Paragraph 8, The above server is, Detecting changes in the first latency distribution and the second latency distribution, A server that determines whether to update the target bit rate, the neural network model, the first latency requirement, and the second latency requirement based on the changes in the first latency distribution and the second latency distribution.

13. In Paragraph 12, The above server is, A server that updates the target bitrate and the neural network model based on the changed first latency distribution and the changed second latency distribution when it is decided to update the target bitrate and the neural network model.

14. In Paragraph 13, The above server is, A server that, when it is decided to update the first latency requirement and the second latency requirement, divides the E2E latency requirement into the updated first latency requirement and the updated second latency requirement based on the changed first latency distribution and the changed second latency distribution.

15. In Paragraph 12, The above server is, A server that reuses the target bitrate, the neural network model, the first latency requirement, and the second latency requirement when it is decided not to update the target bitrate, the neural network model, the first latency requirement, and the second latency requirement.

Citation Information

Patent Citations

  • Optimization of resource allocation based on received sensory quality information

    JP2022000961A

  • Decoding device and program

    JP2022163805A

  • Base station device, control method, and program for performing communication control based on communication service requests

    JP2023178860A

  • Communication system and communication method

    JP2024016565A

  • KR20240045903A