Method, computer readable medium and system for automatic service profile generation
By monitoring resource utilization and performance indicators to generate service profiles, the challenge of determining optimal resource requirements in distributed applications is solved, thereby improving application performance and resource utilization efficiency.
Patent Information
- Application Number
- CN202510378981.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2025-03-28
- Publication Date
- 2025-09-30
AI Technical Summary
How to accurately determine the optimal resource requirements for distributed applications to ensure optimal application performance without requiring developers to have detailed knowledge of the underlying network architecture.
By monitoring resource utilization and performance indicators, multiple candidate service profiles are generated, and the final service profile is generated based on performance measurements and quality of experience feedback.
It enables accurate determination of the optimal resource requirements of distributed applications without relying on developers to have detailed knowledge of the underlying network architecture, thereby improving application performance and resource utilization efficiency.
Smart Images

Figure CN120729933A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to communication systems and, more particularly, to techniques for automatic service profile generation, ie, service profile generation. Background Art
[0002] The statements in this section merely provide background information related to the present disclosure and may not constitute prior art.
[0003] Wireless communication systems are widely deployed to provide a variety of telecommunication services, such as telephony, video, data, messaging, and broadcasting. A typical wireless communication system may employ multiple-access technologies capable of supporting communication with multiple users by sharing available system resources. Examples of such multiple-access technologies include code division multiple access (CDMA) systems, time division multiple access (TDMA) systems, frequency division multiple access (FDMA) systems, orthogonal frequency division multiple access (OFDMA) systems, single-carrier frequency division multiple access (SC-FDMA) systems, and time division synchronous code division multiple access (TD-SCDMA) systems.
[0004] These multiple-access technologies have been adopted in various telecommunications standards to provide a common protocol that enables different wireless devices to communicate at a municipal, regional, and even global level. An example of a telecommunications standard is 5G New Radio (NR). 5G NR is part of the ongoing mobile broadband evolution promoted by the 3rd Generation Partnership Project (3GPP) and is designed to meet new requirements related to latency, reliability, security, scalability (for example, related to the Internet of Things (IoT)), and other requirements. Some aspects of 5G NR may be based on the 4G Long Term Evolution (LTE) standard. There is a demand for further improvements to 5G NR technology. These improvements may also be applicable to other multiple-access technologies and the telecommunications standards that adopt them. Summary of the Invention
[0005] The following provides a simplified overview of one or more aspects in order to provide a basic understanding of these aspects. This overview is not an extensive overview of all contemplated aspects and is not intended to identify key or core elements of all aspects nor to delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description provided later.
[0006] In one aspect of the present disclosure, a method, a computer-readable medium, and a system are provided. The method can be implemented by one or more computing devices. One or more computing devices obtain an initial deployment configuration of connectivity between multiple microservices of a specified distributed application. One or more computing devices deploy multiple microservices in a high-capacity computing environment. One or more computing devices monitor resource utilization and performance indicators when executing the distributed application in the high-capacity computing environment and providing sufficient resources. One or more computing devices generate multiple candidate service profiles by changing the resource allocation of multiple microservices. One or more computing devices collect performance measurements and experience quality feedback for each candidate service profile. One or more computing devices generate a final service profile based on the collected performance measurements and experience quality feedback. A fundamental challenge is solved: how to accurately determine the resource requirements for optimal application performance without requiring developers to have detailed knowledge of the underlying network architecture.
[0007] To accomplish the foregoing and related ends, one or more aspects comprise the features fully described and particularly pointed out in the claims. The following description and drawings set forth in detail certain illustrative features of one or more aspects. However, these features are indicative of but a few of the various ways in which the principles of the various aspects may be employed, and this description is intended to include all such aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 A wireless communication system and an access network are shown.
[0009] Figure 2 A base station is shown communicating with user equipment (UE) in an access network.
[0010] Figure 3 Shows an example logical architecture of a distributed access network.
[0011] Figure 4 Shows an example physical architecture of a distributed access network.
[0012] Figure 5 A time slot centered around the downlink is shown.
[0013] Figure 6 A time slot centered on the uplink is shown.
[0014] Figure 7 A service profile and distributed process communication diagram are shown.
[0015] Figure 8 Shows an example deployment of microservices.
[0016] Figure 9 An exemplary visualization of a service profile is shown.
[0017] Figure 10 Presents a high-level protocol architecture for distributed applications.
[0018] Figure 11A and Figure 11B Shows an example of an automated service configuration approach.
[0019] Figure 12 A flowchart showing the process of automatic service configuration generation. DETAILED DESCRIPTION
[0020] The following detailed description is presented in conjunction with the accompanying drawings and is intended to describe various configurations. It is not intended to imply that these concepts can only be practiced in the described configurations. The detailed description includes specific details for the purpose of providing a deeper understanding of the various concepts. However, those skilled in the art will appreciate that these concepts can be practiced without these specific details. In some cases, known structures and components are shown in block diagram form to avoid obscuring these concepts.
[0021] Several aspects of telecommunications systems will now be described, with reference to various devices and methods. These devices and methods will be described in the following detailed description and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively, "elements"). These elements can be implemented using electronic hardware, computer software, or any combination of the two. Whether these elements are implemented as hardware or software depends on the specific application and design constraints imposed on the overall system.
[0022] For example, an element, any portion of an element, or any combination of elements may be implemented as a "processing system" that includes one or more processors. Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, systems on a chip (SoC), baseband processors, field programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functions described herein. One or more processors in a processing system may execute software. Software should be broadly understood to mean instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software components, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0023] Thus, in one or more example aspects, the functions described may be implemented in hardware, software, or any combination thereof. If implemented in software, the functions may be stored or encoded as one or more instructions or codes stored on a computer-readable medium. Computer-readable media include computer storage media. A storage medium may be any available medium that can be accessed by a computer. For example, and without limitation, such computer-readable media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), optical disk storage, magnetic disk storage, other magnetic storage devices, combinations of the foregoing types of computer-readable media, or any other medium that can be used to store computer-executable code in the form of instructions or data structures.
[0024] Figure 1 Figure 1 is an example diagram illustrating a wireless communication system and access network 100. The wireless communication system (also known as a wireless wide area network (WWAN)) includes base stations 102, user equipment (UEs) 104, an evolved packet core (EPC) 160, and another core network 190 (e.g., a 5G core (5GC)). Base stations 102 may include macro cells (high-power cellular base stations) and / or small cells (low-power cellular base stations). Macro cells include base stations. Small cells include femtocells and micro cells.
[0025] Base stations 102 configured for 4G LTE (collectively referred to as the Evolved Universal Mobile Telecommunications System Terrestrial Radio Access Network (E-UTRAN)) can interface with EPC 160 via backend links 132 (e.g., an S1 interface). Base stations 102 configured for 5G NR (collectively referred to as the Next Generation RAN (NG-RAN)) can interface with core network 190 via backend links 184. Base stations 102 can perform one or more of the following functions, among other things: transmission of user data, radio channel encryption and decryption, integrity protection, header compression, mobility control functions (e.g., handover, dual connectivity), inter-cell interference coordination, connection establishment and release, load balancing, distribution of non-access stratum (NAS) messages, NAS node selection, synchronization, radio access network (RAN) sharing, multimedia broadcast multicast service (MBMS), user and device tracking, RAN information management (RIM), paging, positioning, and delivery of warning messages. Base stations 102 can communicate with each other directly or indirectly (e.g., via EPC 160 or core network 190) via backend links 134 (e.g., an X2 interface). The backend link 134 may be wired or wireless.
[0026] Base stations 102 can wirelessly communicate with user equipment (UE) 104. Each base station 102 provides communication coverage for a corresponding geographic coverage area 110. There may be overlapping geographic coverage areas 110. For example, a small base station 102' may have a coverage area 110' that overlaps with the geographic coverage area 110 of one or more macro base stations 102. A network that includes small base stations and macro base stations may be referred to as a heterogeneous network. Heterogeneous networks may also include Home eNodeBs (HeNBs), which can provide service to a restricted group known as a Closed Subscriber Group (CSG). The communication link 120 between the base station 102 and the user equipment 104 may include uplink (UL) (also known as reverse link) transmissions from the user equipment 104 to the base station 102 and / or downlink (DL) (also known as forward link) transmissions from the base station 102 to the user equipment 104. The communication link 120 may utilize multiple-input multiple-output (MIMO) antenna technology, including spatial multiplexing, beamforming, and / or transmit diversity. The communication link may be conducted over one or more carriers. The base station 102 / user equipment 104 can use spectrum with a bandwidth of up to 7 MHz (e.g., 5, 10, 15, 20, 100, 400, etc. MHz), and the total bandwidth used for transmission on each carrier in each direction can be up to Yx MHz (x component carriers). The carriers may be adjacent or non-adjacent. The allocation of carriers may be asymmetric for DL and UL (e.g., more or fewer carriers may be allocated for DL than for UL). The component carriers may include one primary component carrier and one or more secondary component carriers. The primary component carrier may be referred to as a primary cell (PCell), and the secondary component carriers may be referred to as a secondary cell (SCell).
[0027] Certain user devices 104 may communicate with each other using device-to-device (D2D) communication links 158. The D2D communication links 158 may utilize the DL / UL wireless wide area network (WWAN) spectrum. The D2D communication links 158 may utilize one or more sidelink channels, such as a physical sidelink broadcast channel (PSBCH), a physical sidelink discovery channel (PSDCH), a physical sidelink shared channel (PSSCH), and a physical sidelink control channel (PSCCH). D2D communication may be performed using various wireless D2D communication systems, such as FlashLinQ, WiMedia, Bluetooth, ZigBee, Wi-Fi based on the IEEE 802.11 standard, LTE, or NR.
[0028] The wireless communication system may also include a Wi-Fi access point (AP) 150 that communicates with Wi-Fi stations (STAs) 152 in the 5 GHz unlicensed spectrum via a communication link 154. When communicating in the unlicensed spectrum, the stations 152 / AP 150 may perform a clear channel assessment (CCA) before communicating to determine whether the channel is available.
[0029] Small base station 102' can operate in licensed and / or unlicensed spectrum. When operating in unlicensed spectrum, small base station 102' can use NR and use the same 5 GHz unlicensed spectrum as Wi-Fi AP 150. Small base station 102' using NR in unlicensed spectrum can enhance access network coverage and / or increase capacity.
[0030] Base station 102, whether a small base station 102' or a large base station (e.g., a macro base station), may include an evolved Node B (eNB), a gNode B (gNB), or other types of base stations. Some base stations, such as gNB 180, may communicate with user equipment 104 in the traditional sub-6 GHz spectrum, millimeter wave (mmW) frequencies, and / or near-mmW frequencies. When gNB 180 operates at or near mmW frequencies, gNB 180 may be referred to as a mmW base station. Extremely high frequency (EHF) is a portion of the radio frequency spectrum in the electromagnetic spectrum. EHF ranges from 30 GHz to 300 GHz, with wavelengths between 1 mm and 10 mm. Radio waves in this frequency band may be referred to as millimeter waves. Near-mmW frequencies extend to 3 GHz, with wavelengths of 100 mm. Super high frequency (SHF) bands range from 3 GHz to 30 GHz and are also known as centimeter waves. Communications using the mmW / near-mmW radio frequency bands (e.g., 3 GHz-300 GHz) have extremely high path loss and short range. The mmW base station 180 may communicate with the user equipment 104 using beamforming 182 to compensate for the extremely high path loss and short range.
[0031] Base station 180 may transmit beamformed signals in one or more transmit directions 108a to user equipment 104. User equipment 104 may receive beamformed signals in one or more receive directions 108b from base station 180. User equipment 104 may also transmit beamformed signals in one or more transmit directions to base station 180. Base station 180 may receive beamformed signals in one or more receive directions from user equipment 104. Base station 180 / user equipment 104 may perform beam training to determine optimal receive and transmit directions for base station 180 / user equipment 104. The transmit and receive directions of base station 180 may be the same or different. The transmit and receive directions of user equipment 104 may be the same or different.
[0032] The EPC 160 may include a Mobility Management Entity (MME) 162, other MMEs 164, a Serving Gateway 166, a Multimedia Broadcast Multicast Service (MBMS) Gateway 168, a Broadcast Multicast Service Center (BM-SC) 170, and a Packet Data Network (PDN) Gateway 172. The MME 162 may communicate with a Home Subscriber Server (HSS) 174. The MME 162 is the control node that handles signaling between the User Equipment (UE) 104 and the EPC 160. Typically, the MME 162 provides bearer and connection management. All user Internet Protocol (IP) packets are transmitted through the Serving Gateway 166, which itself is connected to the PDN Gateway 172. The PDN Gateway 172 provides UE IP address allocation, among other functions. The PDN Gateway 172 and BM-SC 170 are connected to IP services 176. IP services 176 may include the Internet, an intranet, an IP Multimedia Subsystem (IMS), PS streaming services, and / or other IP services. The BM-SC 170 may provide configuration and delivery functions for MBMS user services. The BM-SC 170 may serve as an entry point for content providers' MBMS transmissions, may be used to authorize and initiate MBMS bearer services within a public land mobile network (PLMN), and may be used to schedule MBMS transmissions. The MBMS Gateway 168 may be used to distribute MBMS traffic to base stations 102 belonging to a multicast broadcast single frequency network (MBSFN) area broadcast by a specific service, and may be responsible for session management (start / stop) and collecting billing information related to the enhanced Multimedia Broadcast Multicast Service (eMBMS).
[0033] The core network 190 may include an access and mobility management function (AMF) 192, other AMFs 193, a location management function (LMF) 198, a session management function (SMF) 194, and a user plane function (UPF) 195. The AMF 192 may communicate with a unified data management (UDM) 196. The AMF 192 is the control node that handles signaling between the user equipment (UE) 104 and the core network 190. Typically, the SMF 194 provides QoS flow and session management. All user Internet Protocol (IP) packets are transmitted through the UPF 195. The UPF 195 provides UE IP address allocation and other functions. The UPF 195 connects to the IP services 197. The IP services 197 may include the Internet, an intranet, an IP multimedia subsystem (IMS), PS streaming services, and / or other IP services.
[0034] A base station may also be referred to as a gNB, Node B, evolved Node B (eNB), access point, base transceiver station, wireless base station, radio transceiver, transceiver functionality, basic service set (BSS), extended service set (ESS), transmission reception point (TRP), or other appropriate terminology. Base station 102 provides an access point to EPC 160 or core network 190 for user equipment (UE) 104. Examples of user equipment (UE) 104 include a cellular phone, smartphone, Session Initiation Protocol (SIP) phone, laptop, personal digital assistant (PDA), satellite radio, global positioning system, multimedia device, video device, digital audio player (e.g., MP3 player), camera, game console, tablet, smart device, wearable device, vehicle, utility meter, gas pump, large or small kitchen appliance, medical device, implant, sensor / actuator, display, or any other similarly functional device. Some user equipment (UE) 104 may be referred to as Internet of Things (IoT) devices (e.g., parking meter, gas pump, toaster, vehicle, heart monitor, etc.). The user equipment (UE) 104 may also be referred to as a site, mobile site, subscriber site, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber site, access terminal, mobile terminal, wireless terminal, remote terminal, handheld device, user agent, mobile client, client, or other suitable terminology.
[0035] Although the present disclosure may refer to 5G New Radio (NR), the present disclosure may also be applicable to other similar areas such as LTE, LTE-Advanced (LTE-A), Code Division Multiple Access (CDMA), Global System for Mobile Communications (GSM), or other wireless / radio access technologies.
[0036] Figure 2 The following is a block diagram of a base station 210 communicating with a user equipment (UE) 250 in an access network. In the downlink (DL), IP packets from the evolved packet core (EPC) 160 may be provided to the controller / processor 275. The controller / processor 275 implements Layer 3 and Layer 2 functions. Layer 3 includes the radio resource control (RRC) layer, and Layer 2 includes the packet data convergence protocol (PDCP) layer, the radio link control (RLC) layer, and the medium access control (MAC) layer. The controller / processor 275 provides RRC layer functions related to broadcasting of system information (e.g., primary information blocks (MIBs), secondary information blocks (SIBs)), RRC connection control (e.g., RRC connection paging, RRC connection establishment, RRC connection modification, and RRC connection release), inter-radio access technology (RAT) mobility, and measurement configuration for UE measurement reporting; PDCP layer functions related to header compression / decompression, security (ciphering, deciphering, integrity protection, integrity verification), and handover support functions; RLC layer functions related to transmission of upper layer packet data units (PDUs), error correction through automatic repeat request (ARQ), splicing, segmentation, and reassembly of RLC service data units (SDUs), re-segmentation of RLC data PDUs, and reordering of RLC data PDUs; and mapping between logical channels and transport channels, multiplexing of MAC SDUs to transport blocks (TBs), MAC MAC layer functions related to demultiplexing of SDUs from TBs, scheduling information reporting, error correction through Hybrid Automatic Repeat Request (HARQ), priority handling and logical channel priority.
[0037] The transmit (TX) processor 216 and receive (RX) processor 270 implement Layer 1 functions associated with various signal processing functions. Layer 1, including the physical (PHY) layer, may include error detection for the transmission channel, forward error correction (FEC) encoding / decoding of the transmission channel, interleaving, rate matching, mapping to physical channels, modulation / demodulation of the physical channels, and multiple-input multiple-output (MIMO) antenna processing. The TX processor 216 handles the mapping of various modulation schemes (e.g., binary phase shift keying (BPSK), quadrature phase shift keying (QPSK), M-phase shift keying (M-PSK), M-quadrature amplitude modulation (M-QAM)) to signal constellations. The coded and modulated symbols may then be split into parallel streams. Each stream may then be mapped to an orthogonal frequency division multiplexing (OFDM) subcarrier, multiplexed with a reference signal (e.g., a pilot) in the time and / or frequency domain, and then combined using an inverse fast Fourier transform (IFFT) to produce a physical channel carrying a time-domain OFDM symbol stream. The OFDM stream is spatially precoded to generate multiple spatial streams. Channel estimates from a channel estimator 274 may be used to determine the coding and modulation schemes, as well as for spatial processing. The channel estimates may be derived from a reference signal and / or channel condition feedback transmitted by the UE 250. Each spatial stream may then be provided to a different antenna 220 via a separate transmitter 218TX. Each transmitter 218TX may modulate a radio frequency carrier for transmission using the corresponding spatial stream.
[0038] At the UE 250, each receiver 254RX receives a signal via its corresponding antenna 252. Each receiver 254RX recovers the information modulated onto the RF carrier and provides the information to a receive (RX) processor 256. The TX processor 268 and the RX processor 256 implement layer 1 functions associated with various signal processing functions. The RX processor 256 may perform spatial processing on the information to recover any spatial streams intended for the UE 250. If multiple spatial streams are intended for the UE 250, they may be combined by the RX processor 256 into a single OFDM symbol stream. The RX processor 256 then converts the OFDM symbol stream from the time domain to the frequency domain using a fast Fourier transform (FFT). The frequency domain signal includes a separate OFDM symbol stream for each subcarrier of the OFDM signal. By determining the most likely signal constellation point transmitted by the base station 210, the symbols and reference signals on each subcarrier are recovered and demodulated. These soft decisions may be based on channel estimates calculated by the channel estimator 258. The soft decisions are then decoded and deinterleaved to recover the data and control signals originally transmitted on the physical channel by base station 210. The data and control signals are then provided to controller / processor 259, which implements Layer 3 and Layer 2 functionality.
[0039] The controller / processor 259 may be associated with a memory 260 that stores program codes and data. The memory 260 may be referred to as a computer-readable medium. In the uplink (UL), the controller / processor 259 provides demultiplexing between transport and logical channels, packet reassembly, decryption, header decompression, and control signal processing to recover IP packets from the EPC 160. The controller / processor 259 is also responsible for error detection using an ACK and / or NACK protocol to support HARQ operations.
[0040] Similar to the functional description related to DL transmission of the base station 210, the controller / processor 259 provides RRC layer functions related to system information (e.g., MIB, SIBs) acquisition, RRC connection and measurement reporting; PDCP layer functions related to header compression / decompression and security (encryption, decryption, integrity protection, integrity verification); RLC layer functions related to transmission of upper layer PDUs, error correction through ARQ, splicing, segmentation and reassembly of RLC SDUs, re-segmentation of RLC data PDUs and reordering of RLC data PDUs; and MAC layer functions related to mapping between logical channels and transport channels, multiplexing of MAC SDUs to TBs, demultiplexing of MAC SDUs from TBs, scheduling information reporting, error correction through HARQ, priority handling and logical channel priority.
[0041] Channel estimates derived from a reference signal or feedback transmitted from the base station 210 by the channel estimator 258 are used by the TX processor 268 to select the appropriate coding and modulation scheme and facilitate spatial processing. The spatial streams generated by the TX processor 268 may be provided to different antennas 252 via different transmitters 254TX. Each transmitter 254TX may modulate a radio frequency carrier with a corresponding spatial stream for transmission. Uplink (UL) transmissions are processed in the base station 210 in a manner similar to the receive function in the user equipment (UE) 250. Each receiver 218RX receives a signal via its corresponding antenna 220. Each receiver 218RX recovers the information modulated onto the radio frequency carrier and provides the information to the RX processor 270.
[0042] The controller / processor 275 may be associated with a memory 276 that stores program codes and data. Memory 276 may be referred to as a computer-readable medium. In the uplink, the controller / processor 275 provides demultiplexing between transport and logical channels, packet reassembly, decryption, header decompression, and control signal processing to recover IP packets from the user equipment 250. The IP packets from the controller / processor 275 may be provided to the EPC 160. The controller / processor 275 is also responsible for error detection using an ACK and / or NACK protocol to support HARQ operations.
[0043] New Radio (NR) may refer to a radio configured to operate according to a new air interface (e.g., non-orthogonal frequency division multiple access (OFDMA)-based air interface) or a fixed transport layer (e.g., non-Internet Protocol). NR may utilize orthogonal frequency division multiplexing (OFDM) with cyclic prefixes (CPs) in both uplink and downlink, and may include support for half-duplex operation using time division duplexing (TDD). NR may include enhanced mobile broadband (eMBB) services for wide bandwidths (e.g., over 80 MHz), millimeter wave (mmW) for high carrier frequencies (e.g., 60 GHz), massive MTC (mMTC) for non-backward-compatible MTC technologies, and / or mission-critical ultra-reliable low-latency communication (URLLC) services.
[0044] A single component carrier bandwidth of 100MHz may be supported. In one example, NR resource blocks (RBs) may span 12 subcarriers with a subcarrier bandwidth of 60kHz for 0.25ms, or a bandwidth of 30kHz for 0.5ms (similarly, for 15kHz SCS, 50MHz BW for 1ms). Each radio frame may consist of 10 subframes (10, 20, 40, or 80 NR time slots) of length 10ms. Each time slot may indicate the link direction of data transmission (i.e., downlink or uplink), and the link direction of each time slot may be dynamically switched. Each time slot may include downlink / uplink data and downlink / uplink control data. The uplink and downlink time slots of NR may be as follows: Figure 5 and Figure 6 As described in more detail.
[0045] The NR RAN may include a central unit (CU) and distributed units (DUs). An NR base station (e.g., gNB, 5G Node B, Node B, transmission reception point (TRP), access point (AP)) may correspond to one or more base stations. NR cells can be configured as access cells (ACells) or data-only cells (DCells). For example, the RAN (e.g., a central unit or a distributed unit) may configure the cell. DCells may be cells used for carrier aggregation or dual connectivity and may not be used for initial access, cell selection / reselection, or handover. In some cases, DCells may not transmit synchronization signals (SS), and in some cases DCells may transmit SS. The NR base station may transmit a downlink signal to the UEs indicating the cell type. Based on the cell type indication, the UE may communicate with the NR base station. For example, the UE may determine the NR base station to be considered for cell selection, access, handover, and / or measurement based on the indicated cell type.
[0046] Figure 3According to aspects of the present disclosure, an example logical architecture of a distributed RAN 300 is shown. A 5G access node 306 may include an access node controller (ANC) 302. The ANC may be a central unit (CU) of the distributed RAN. The backhaul interface to the next generation core network (NG-CN) 304 may terminate at the ANC. The backhaul interface to adjacent next generation access nodes (NG-ANs) 310 may terminate at the ANC. The ANC may include one or more TRPs 308 (which may also be referred to as BSs, NR BSs, Node Bs, 5G NBs, APs or other terms). As mentioned above, TRP may be used interchangeably with "cell".
[0047] TRPs 308 may be a distribution unit (DU). TRPs may be connected to one ANC (ANC 302) or multiple ANCs (not shown). For example, for RAN sharing, Radio as a Service (RaaS), and service-specific ANC deployments, a TRP may be connected to multiple ANCs. A TRP may include one or more antenna ports. TRPs may be configured to serve UE traffic individually (e.g., dynamically selected) or collectively (e.g., joint transmission).
[0048] The local architecture of the distributed RAN 300 can be used to illustrate the fronthaul definition. The architecture may be defined to support fronthaul solutions across different deployment types. For example, the architecture may be based on transport network capabilities (e.g., bandwidth, latency, and / or jitter). The architecture may share features and / or components with LTE. According to aspects, the next-generation AN (NG-AN) 310 may support dual connectivity with NR. The NG-AN may share a common fronthaul for LTE and NR.
[0049] The architecture may enable collaboration between TRPs 308 and between TRPs. For example, collaboration may be provisioned within a TRP and / or across TRPs via ANC 302. According to certain aspects, an inter-TRP interface may not be required / present.
[0050] According to certain aspects, there may be dynamic configuration of split logical functions in the architecture of the distributed RAN 300. PDCP, RLC, and MAC protocols may be flexibly placed in the ANC or TRP.
[0051] Figure 4According to certain aspects of the present disclosure, an example physical architecture of a distributed RAN 400 is shown. A centralized core network unit (C-CU) 402 may host core network functions. The C-CU may be centrally deployed. C-CU functions may be offloaded (e.g., to Advanced Wireless Services (AWS)) in an effort to handle peak capacity. A centralized RAN unit (C-RU) 404 may host one or more ANC functions. Optionally, the C-RU may host core network functions locally. The C-RU may have a distributed deployment. The C-RU may be closer to the network edge. A distributed unit (DU) 406 may host one or more TRPs. The DU may be located at the network edge with radio frequency (RF) functions.
[0052] Figure 5 An example of a downlink (DL) centric time slot is shown. The DL centric time slot may include a control portion 502. The control portion 502 may be present at the initial or beginning portion of the DL centric time slot. The control portion 502 may include various scheduling information and / or control information corresponding to various portions of the DL centric time slot. In some configurations, the control portion 502 may be a physical downlink control channel (PDCCH), such as Figure 5 As shown. The DL-centric timeslot may also include a DL data portion 504. The DL data portion 504 may sometimes be referred to as the payload of the DL-centric timeslot. The DL data portion 504 may include communication resources used to communicate DL data from a scheduling entity (e.g., a UE or a base station) to a subordinate entity (e.g., a UE). In some configurations, the DL data portion 504 may be a physical downlink shared channel (PDSCH).
[0053] The DL-centric timeslot may also include a common uplink (UL) portion 506. Common UL portion 506 may sometimes be referred to as a UL burst, a common UL burst, and / or various other suitable terms. Common UL portion 506 may include feedback information corresponding to various portions of the DL-centric timeslot. For example, common UL portion 506 may include feedback information corresponding to control portion 502. Non-limiting examples of feedback information may include ACK signals, NACK signals, HARQ indicators, and / or various other suitable types of information. Common UL portion 506 may include additional or alternative information, such as information related to random access channel (RACH) procedures, scheduling requests (SRs), and various other suitable types of information.
[0054] like Figure 5As shown, the end of the DL data portion 504 may be separated in time from the beginning of the common UL portion 506. This time separation may sometimes be referred to as a gap, a guard period, a guard interval, and / or various other suitable terms. This separation provides time to switch from DL communications (e.g., reception by a slave entity (e.g., a UE)) to UL communications (e.g., transmission by a slave entity (e.g., a UE)). Those skilled in the art will appreciate that the foregoing is merely an example of a DL-centric time slot, and alternative structures with similar features may exist without necessarily departing from the aspects described herein.
[0055] Figure 6 An example of an uplink (UL) centric timeslot is shown. The UL centric timeslot may include a control portion 602. The control portion 602 may be present at the beginning or start of the UL centric timeslot. Figure 6 The control portion 602 may be the same as that in the above reference Figure 5 The control portion 502 is similar. The UL-centric timeslot may also include a UL data portion 604. The UL data portion 604 may sometimes be referred to as the payload of the UL-centric timeslot. The UL portion may refer to the communication resources used to communicate UL data from a dependent entity (e.g., a UE) to a scheduling entity (e.g., a UE or a base station). In some configurations, the control portion 602 may be a physical downlink control channel (PDCCH).
[0056] like Figure 6 As shown, the end of the control portion 602 may be separated in time from the beginning of the UL data portion 604. This temporal separation may sometimes be referred to as a gap, a guard period, a guard interval, and / or various other suitable terms. This separation provides time to switch from DL communications (e.g., reception by the scheduling entity) to UL communications (e.g., transmission by the scheduling entity). The UL-centric timeslot may also include a common UL portion 606. Figure 6 The common UL portion 606 in the above reference may be Figure 5 The common UL portion 506 is similar to the common UL portion 506. The common UL portion 606 may also include information related to a channel quality indicator (CQI), sounding reference signals (SRSs), and various other suitable types of information. Those skilled in the art will appreciate that the foregoing is merely an example of a UL-centric timeslot, and alternative structures with similar features may exist without necessarily deviating from the aspects described herein.
[0057] In some cases, two or more subordinate entities (e.g., UEs) may communicate with each other using sidelink signals. Real-world applications of such sidelink communications may include public safety, proximity services, UE-to-network relay, vehicle-to-vehicle (V2V) communications, Internet of Things (IoE) communications, Internet of Things (IoT) communications, mission-critical grids, and / or other appropriate applications. Generally, sidelink signals may refer to communications from one subordinate entity (e.g., UE1) to another subordinate entity (e.g., UE2) without relaying the communications through a scheduling entity (e.g., UE or BS), even though the scheduling entity may be used for scheduling and / or control purposes. In some examples, sidelink signals may use licensed spectrum for communication (unlike wireless local area networks that typically use unlicensed spectrum).
[0058] Distribution and communication are fundamental concepts in distributed application architectures. This architecture uses distributed computing resources to improve performance, so performance considerations are crucial. Effective application distribution requires interactions between application modules, microservices, or remote functions, ranging from simple point-to-point interactions to complex, large-scale clusters and dynamic service-oriented architectures. Furthermore, communication across system boundaries is crucial for scaling software systems and improving their availability.
[0059] As distributed computing systems become more dynamic, their architectural complexity increases. In this context, managing quality of service (QoS) involves allocating network resources to provide optimal performance. QoS is important for maintaining the responsiveness and reliability of high-priority services, especially in heavily loaded or shared network environments.
[0060] For developers, understanding the impact of service deployments in production can be challenging. This complexity, combined with a lack of understanding of the underlying system capabilities, can lead to increased resource consumption and degraded performance. This is because application performance is significantly affected by the configuration of the infrastructure and the distribution of remote application modules. Therefore, it is important to make the details of the underlying communication and infrastructure between remote nodes transparent to application developers.
[0061] When it comes to distributed nodes, connection management and threading models are key considerations. Establishing a connection can take time, and the threading model determines how requests are handled—either synchronously, which blocks the thread until a response is received, or asynchronously, in which a callback function is invoked when the response arrives. Beyond the node level, the network itself is a core component of distributed applications, influencing scalability and impacting performance.
[0062] This disclosure introduces a mechanism for optimizing orchestration to achieve high performance in distributed microservices and / or distributed function environments. It involves communicating application and network-related performance requirements, defined as "service profiles," from application developers to the system orchestrator. A "service profile" in a distributed resource-sharing environment defines a set of parameters for managing resource allocation and quality of service. These parameters may include, but are not limited to, priority, bandwidth allocation, latency sensitivity, jitter control, traffic shaping, congestion management, service availability, fairness, policy compliance, scalability, adaptive QoS, resource reservation, monitoring, and other related parameters.
[0063] The present disclosure also provides a mechanism for application developers to generate service configuration files for distributed applications.
[0064] A "service profile" is an important component of addressing service requirements in a distributed microservice, application module, or remote functional environment. The term "microservice" in this disclosure may represent any type of application module that can be distributed and work together remotely. A service profile should include a variety of factors to meet the diverse needs of microservices to provide optimal performance, reliability, and user experience. QoS in a microservice architecture involves managing and monitoring network resources to ensure performance for different types of traffic, as well as controlling the utilization of computing resources such as central processing units (CPUs) and memory. QoS is critical to maintaining the responsiveness and reliability of high-priority services, especially in the face of heavy network loads (e.g., system overload conditions) or in environments with shared computing resources and networks.
[0065] To deliver optimally distributed applications, developers should consider three aspects from the start of implementation: design, operations, and performance. In some configurations, "design aspects" include: decomposing services into microservices, application programming interface (API) gateways for communicating with remote modules, synchronous or asynchronous communication, traffic shaping, caching, etc.
[0066] In some configurations, "operational aspects" include service discovery, load balancing, scalability, and other aspects. These aspects are integral to application design and implementation. Parameters for the following performance aspects are contained in service configuration files and used by the orchestrator during the real-time lifecycle of the distributed application. The detailed text structure and format of the service configuration file are described in further detail below.
[0067] In some configurations, "performance aspects" include:
[0068] (1) Latency: Minimize network latency (delay) between microservices, especially for synchronous communication. Low latency is critical to achieving fast response times, which is especially important in user-facing applications or when microservices communicate frequently.
[0069] (2) Bandwidth: Provide sufficient network bandwidth to meet the data transmission needs of microservices. This is also important, especially when microservices exchange large amounts of data or high-resolution media files.
[0070] (3) Packet loss: Implement strategies to minimize packet loss. High packet loss can lead to data corruption, retransmissions, and increased latency. Using reliable transport protocols and optimizing network configuration can help reduce packet loss.
[0071] (4) Adaptive QoS: Implementing an adaptive QoS mechanism can dynamically adjust priorities and resource allocation based on real-time network performance and service requirements.
[0072] (5) Priority: Define different priorities for different types of traffic. For example, critical service requests may be given higher priority than batch jobs or data backup activities.
[0073] (6) Bandwidth allocation: Allocate bandwidth based on the importance and needs of each microservice. High-priority services may require guaranteed bandwidth to provide consistent performance.
[0074] (7) Delay Sensitivity: Identify services that are delay-sensitive and require low response times. QoS policies can prioritize these services so that they are less affected by network congestion.
[0075] (8) Jitter: Jitter refers to the variability of packet delay. For services that require real-time data processing or streaming, jitter is important to application quality.
[0076] (9) Service availability: Provide high availability for critical services by prioritizing their traffic, especially when network resources are limited.
[0077] (10) Resource reservation: For very critical services, consider reserving network resources to ensure uninterrupted performance.
[0078] (11) Scheduling: While prioritizing certain types of traffic, it is important to prevent low-priority traffic from being excessively downgraded. This is determined by the scheduler policy implemented by the orchestrator (e.g., a form of proportional fairness that also takes into account service history).
[0079] (12) QoS Policy: Define and enforce QoS policies to prioritize network traffic. This is especially important for time-sensitive applications. Critical traffic (e.g., real-time data) receives the necessary bandwidth and low latency.
[0080] (13) Monitoring, logging, and tuning: Implement comprehensive monitoring and logging to track the health, performance, and usage of microservices, and adjust QoS policies in a timely manner based on changing network conditions and service requirements. This facilitates rapid debugging and performance tuning.
[0081] Once a distributed application is implemented, application designers, including developers, need to prepare a list of performance requirements for network nodes and user devices for each microservice or distributed module. Furthermore, they need to specify the communication performance requirements between microservices. This structured information is called a "service profile."
[0082] Figure 7 The service profile and distributed process communication diagram are presented. Specifically, Figure 7 A user device running a distributed application, consisting of six microservices, is shown as an exemplary embodiment. Figure 7 A logical view of the service profile is provided by the distributed process communication diagram 710 depicted in FIG. The diagram 710 shows that three microservices (μs1, μs2, μs3) are instantiated individually, while the remaining three (μs4, μs5, μs6) are instantiated as a group. Figure 7 Communication arrows between microservices are also shown, representing the communication paths between them (ε1 to ε6).
[0083] The service configuration file for each distributed application may contain the following information, presented in a structured format and deployed in real time when the application starts, including:
[0084] Each microservice in a service profile may be defined based on its compute resource requirements. Example parameters include, but are not limited to, CPU, memory, and storage resource requirements for each node or microservice.
[0085] Each communication between remote processes or microservices in a service profile may be defined based on specific communication resource requirements or conditions and perceived link utilization. Example metrics include, but are not limited to, source and destination microservices for traffic direction, message rate, message size, link latency, link latency variation, message response time latency, and packet loss between connected nodes or microservices.
[0086] Service configuration files may be prepared in a structured format for use as input by an orchestrator. Subsequently, the service configuration files may be converted to another format specifically for use by a specific orchestrator for application deployment. This enables the orchestrator to optimally distribute microservices across worker nodes and configure communication channels to meet application QoS and / or Quality of Experience (QoE), performance, and reliability requirements throughout the real-time lifecycle of the distributed application. This allows the orchestrator to efficiently meet performance and resource requirements.
[0087] The service configuration file or the location of the service configuration file for a distributed application may be installed on the device when the main application is installed.
[0088] When launching a distributed application, a device cluster, which may include user devices and network devices, is formed and a master orchestrator is elected among the devices in the device cluster. The service profile is sent to the orchestrator, which uses it as a decision source for application cluster creation.
[0089] When the orchestrator forms an application cluster with remote devices, which may include user devices and network devices, only the applicable portion of the service profile corresponding to the role of the worker node may be transferred to the target worker node.
[0090] When a distributed application runs, each microservice or the underlying worker nodes (compute nodes) supporting it continuously monitors the application's behavior or performance. This may include each node's CPU, memory, and storage utilization; network link utilization, such as packet rate, bandwidth, message rate, and traffic patterns; and network performance metrics between connected nodes, such as packet latency, latency variation, packet error rate, and message error rate. Performance monitoring results are fed back to the orchestrator to maintain or update optimal application cluster performance.
[0091] The orchestrator continuously evaluates the application QoS, if the application performance is objectively measurable, which includes performance measurements of each microservice and performance measurements of connections between microservices.
[0092] Alternatively, the orchestrator continuously predicts the QoE of the application from a user-aware perspective based on any algorithm or approach, including machine learning (ML) models, continuous measurements of the performance of each microservice and communication measurements between microservices.
[0093] If the observed application performance (QoS and / or QoE) does not meet the expected performance, the orchestrator may actively adjust the application cluster, such as excluding some low-performing devices, adding new devices, changing network connectivity, or interacting with the systems on the devices in the cluster to reserve required resources or configure the network scheduler and protocols to improve application performance.
[0094] A service profile contains the compute and communication resource and performance requirements for a distributed application and is independent of the physical topology of user devices and network equipment. A single worker node might host all distributed processes, and communication between these processes can still occur within the same physical device. There may be multiple physical worker nodes in a cluster, and the orchestrator is responsible for assigning microservices to the most appropriate worker device.
[0095] Figure 8 Shows an example deployment of microservices, reflecting Figure 7 The distributed process communication diagram shown in Figure 1 is as follows. Figure 8 As shown, the system includes multiple devices configured to operate in a distributed computing environment, forming a device cluster. For example, device A can be implemented as a smartphone, tablet, or personal computer (PC) associated with a user. In a crowded environment, such as a shopping mall or public area, there may be multiple users, each with a corresponding device (e.g., device A, device B, device C, and device D).
[0096] Microservice 1 (μs1) is instantiated on device A, microservice 2 (μs2) is instantiated on device B, microservice 3 (μs3) is instantiated on device D, and the remaining microservices (μs4, μs5, and μs6) are instantiated on the same device E. This topology and list of active devices can change dynamically based on network and / or device conditions.
[0097] exist Figure 8 In the figure, device AD represents user equipment, while device F and device G represent network nodes, such as base stations (BSs), access points (APs), or gateways (GWs), which provide connectivity between user equipment or between user equipment and the operator's network. Device E is a network device within the operator's network. Therefore, device AG is also referred to as a network entity in the network.
[0098] Distributed computing resource sharing is enabled in a device AG because it includes the Device Compute Orchestrator (DCO) and the Device Distribute Compute Function (DDCF) and resides in the same device cloud cluster. Devices A, B, D, and E provide computing resources for the microservices installed on their nodes. Devices C and F provide connectivity to other devices, and the orchestrator can potentially utilize devices C and F when necessary to maintain application performance.
[0099] Specifically, user devices A and D support three radio access technologies (RATs), while user devices B and C support two RATs. User device B connects to network node F, which is further connected to the network operator's core cloud via network node E in the edge cloud. User device D connects to another network node G, which is also connected to the core cloud of the same network operator. Therefore, user devices B and D are subscribers to the same network operator, while user devices A and C may be unsubscribed user devices, and in this example embodiment, they are not subscribed to a network operator.
[0100] In architecture 800, if the network supports granting remote computing resources to user devices through the device cloud, the DDCF in the core network configures network nodes in the core cloud, edge cloud, hyperlocal cloud, and subscribing user devices (e.g., user devices B and D). The dashed lines from the DDCF in the core cloud to network nodes E and F and subscribing user devices B and D represent a distributed device cloud function instantiation and / or configuration scenario.
[0101] The network nodes supporting the remote computing resources (e.g., network nodes E and F) then communicate their intent (via a request) to participate in the sharing of computing resources. Intermediate user device B adds its device ID B to the forwarded message to indicate the traffic path. At this stage, a service bearer is created between the tenant and the proxy device (e.g., user device B). If the network requires a GTP or IP tunnel, an intermediate GTP or IP tunnel may also be established. The network tunnel is transparent to the tenant device (e.g., user device A).
[0102] The following steps describe the Figure 8 The embodiment described in the embodiment of the present invention describes the construction of the device cloud frame exchange table. In process (1), user device A runs an application. User device A estimates the amount of computing resources that the application may require (main application). User device A may further estimate that the amount of computing resources is higher than the amount that user device A can provide. In process (2), the DCO (main orchestrator) hosted by user device A triggers the DDCF of user device A to discover remote computing resources.
[0103] In process (3), if user device A is not currently in the subnet, all RATs of user device A attempt to connect to neighboring user devices. Subsequently, neighboring user devices that are capable of cooperating (e.g., user devices B, C, and D) use corresponding RAT-specific protocols to establish connectivity with user device A, including security mechanisms. In process (4), after the subnet is established, the DDCF of user device A broadcasts a subnet message, i.e., a resource query message, through all connected RATs. This resource query message contains information about the destination device and the source device (i.e., user device A). For example, the resource query message may contain information such as (destination device ID = X, source device ID = A), where the device ID "X" represents any subnet device ID.
[0104] In process (5), after receiving the resource query message sent by user device A, user device B forwards the resource query message to network node F. In process (6), after receiving the resource query message forwarded by user device B, network node F forwards it to network node E in the edge cloud. Each intermediate device along the path between the tenant and the lessee (e.g., user device B and network node F) updates the device cloud frame forwarding table containing neighboring information (destination device ID, destination device IP address, RAT output port, and next hop device ID). If network node E is willing to provide the requested computing resources, it will send a confirmation message with its IP address back to user device A (through intermediate devices F and B). After receiving the confirmation message, user device A has information about network node E, including information on how to contact network node E.
[0105] In processes (7) and (8), another exemplary traffic flow is between user devices A and D, which shows a potential frame exchange loop problem because user device D receives duplicate resource query messages, including one message forwarded by user device C through process (7) (i.e., user device A to user device C and then to user device D), and another message sent directly from user device A through process (8) (i.e., user device A to user device D), via different RATs. In this case, the duplicate resource query messages can be identified by the message sequence number, and the forwarding loop can be identified by the path vector in the device-cloud message. User device D should select the best link to avoid the forwarding loop.
[0106] When device A executes an application that requires significant computational resources (e.g., high CPU or memory usage), the system is configured to discover nearby devices within communication range. These devices may include other user devices (e.g., smartphones, tablets, or personal computers), access points (APs), base stations, or network servers. For example, if a user is sitting near other devices, the system can identify and utilize these devices for resource sharing. These devices, in a device cluster, collaborate to execute distributed applications, potentially forming an application cluster, such as Figure 8 Devices AB and DF are shown in FIG.
[0107] The system is further integrated with network infrastructure components, such as base stations, customer premises equipment (CPE), access points (APs), edge cloud servers, and core cloud servers. These components may be operated by network service providers (e.g., AT&T, Verizon) and can be used as part of a resource-sharing ecosystem. For example, if a network operator provides computing resources (e.g., central processing units, memory) on its network servers, Device A can offload part of the application workload to these resources.
[0108] The system dynamically selects resources based on application requirements, such as latency or proximity. For low-latency applications, the system might prioritize nearby access points or base stations. If higher computing power is required, the system might use a nearby user device or network-side machine with greater computing power.
[0109] The system is configured to form clusters or groups of available devices, including user devices, access points, and network servers. Based on application requirements, the system distributes the portion of the application running on device A across multiple devices within the cluster, achieving efficient resource utilization and improved performance.
[0110] exist Figure 8 In the system shown in Figure 1, each circle represents a function or process of the application. To execute the application, six microservices (μs1 to μs6) are required, and all six microservices should run simultaneously. If device A has sufficient computing resources, all six microservices (represented by the six circles) can be executed locally on device A, as centralized execution is generally more efficient.
[0111] However, if device A has limited resources, the system is configured to distribute parts of the application to other devices. For example, some microservices might be offloaded to a nearby device with greater computing power. The arrows between the circles represent communication paths between microservices (e.g., the communication between μs1 and μs2, denoted as ε1), where messages are exchanged to facilitate coordinated execution.
[0112] The system operates based on application requirements and resource availability. It dynamically forms clusters or groups of devices and distributes parts of the application to remote devices for execution. This dynamic clustering provides optimal resource utilization and performance.
[0113] In scenarios involving user mobility, the system considers the potential for devices to leave the cluster. For example, if the owner of device D moves out of communication range, microservice 3 (μs3) previously executing on device D needs to be relocated to another device within the cluster (e.g., device C or device G) to maintain application continuity. All required microservices remain running and accessible to the application.
[0114] The DCO is configured to manage the distribution and execution of microservices across multiple devices, including smartphones, servers, and other computing resources. The DCO determines which device handles each microservice, the number of microservices assigned to each device, and when to migrate a microservice to a different device. These decisions are based on the specific requirements of the application.
[0115] Application requirements determine the resource requirements of each microservice. For example, microservice 2 (μs2) may require high CPU utilization, while microservice 3 (μs3) may be memory-intensive, and another microservice may be visualization-oriented. Each microservice performs different functions and has unique resource requirements.
[0116] Communication constraints and placement optimization
[0117] DCO also considers communication constraints between microservices. For example, communication between microservice 1 (μs1) and microservice 2 (μs2) may require low latency or short distance communication, while communication between microservice 1 (μs1) and microservice 3 (μs3) may tolerate longer distances or higher latency. If μs1 and μs3 require low latency communication, DCO provides the option of placing μs3 near μs1.
[0118] DCO leverages application requirements, including resource demands and communication constraints, to determine the optimal placement of microservices across devices. This decision-making process is an integral part of distributed systems because it provides efficient resource utilization and application performance. DCO continuously monitors the system and dynamically adjusts the placement of microservices as needed.
[0119] Service Profile Parameters
[0120] By default, service profile parameters are categorized into three types: service, microservice, and communication. However, more types can be added. The "service" type contains parameters for the entire application or service, rather than parameters for a specific microservice or communication connection.
[0121] The "Microservice" type includes the compute-related parameters required to run a specific microservice on a device. Examples include the number of central processing units (CPUs), CPU clock rate, microservice image repository location, image size, required memory and local storage size, scalability, monitoring and feedback capabilities, and service availability.
[0122] The "Communication" type includes parameters related to communication and connectivity between specific microservices. Example parameters include source microservice, destination microservice, maximum tolerable packet or message delay, maximum tolerable delay jitter, delay and jitter sensitivity, bandwidth requirements, expected message rate, required bandwidth reservation, maximum tolerable message / packet error rate, etc.
[0123] If adaptive Quality of Service (QoS) is feasible for the application, the service profile may contain multiple levels of sub-service profiles (ie, layered service profiles).
[0124] The "Service" type may include parameters such as adaptive Quality of Service (QoS) in the service profile. This refers to a tiered service profile with multiple levels of QoS requirements. If the network environment is not adaptable, a lower level of QoS requirements can be used. An example of a two-tiered service profile is a "Desired" service profile for optimal performance and a "Minimal" service profile for applications running at the lowest performance.
[0125] The "Service" type may also include a "Cluster Identifier (ID)" parameter, where multiple cluster IDs represent different or parallel clusters. These can be used to deploy services with multiple application clusters. Alternatively, each cluster can run independently.
[0126] The "Service" type may also include a "Quality of Experience Type (QoE Type)" parameter. There are at least two types of Quality of Experience: objective and subjective. The so-called "Objective Quality of Experience" means that the performance of the application can be evaluated based on quantifiable measurements, such as observed data rate, packet delay, delay variation, and packet loss rate (implicit feedback). Quality of Experience estimation for this type of application can be achieved through application Quality of Service (QoS) measurements without the need for explicit feedback from humans / users. Some real-time applications may fall into this category. The so-called "Subjective Quality of Experience" category requires explicit human feedback because the level of Quality of Experience is determined by the user's perception of availability when using the service.
[0127] The "Microservice" type in a service configuration file may include parameters such as CPU, which specifies the minimum computational processing power (for example, millions of instructions per second (MIPS)), and CPU clock frequency, which defines the required processing speed. The "Microservice" type may also include the URL of the image repository location where the microservice images are stored. In addition, parameters for minimum, average, and maximum memory utilization, as well as minimum, average, and maximum local storage utilization are defined to provide adequate resource allocation. Scalability parameters determine whether the microservice can be scaled independently.
[0128] Each microservice should have monitoring and feedback capabilities, enabling it to monitor resource utilization and performance through self-monitoring and reporting or through host devices. This capability provides feedback for dynamic orchestration. Furthermore, computing resource reservation requirements ensure that critical microservices have guaranteed resources. Service availability, expressed on a scale of 0-100%, defines the required uptime for a microservice. The priority of microservices in a distributed application determines the relative importance of each microservice.
[0129] The "Communication" type in a service profile might include parameters such as the communication source microservice and the communication destination microservice, which define the endpoints of the communication path. The maximum tolerable message delay between distributed modules specifies the allowed delay, while the maximum tolerable delay jitter (including packet retransmissions, if applicable) defines the acceptable variation in delay. Latency and jitter sensitivity, typically rated on a scale of 0-10, indicates the severity of the impact if the latency requirement is violated.
[0130] Additionally, the bandwidth requirements (expressed in bits per second (bps)) and / or the expected message rate (expressed in messages per second) between distributed modules provide sufficient network capacity. Bandwidth reservation requirements ensure that critical communication paths have dedicated resources. The maximum tolerable message / packet error rate defines the acceptable level of data corruption or loss. The priority of communication connections within a distributed application determines the relative importance of each communication path.
[0131] Service Profile Data File Format
[0132] A service profile contains the application's computational and communication performance requirements. These requirements are recorded in a structured data file format as key-value pairs. This is a lightweight data exchange format that is easy for humans to read and write, and easy for machines to parse and generate. Common standard text formats for representing structured data include, but are not limited to, JSON (JavaScript Object Notation), XML (Extensible Markup Language), YAML (YAML Ain't Markup Language), TOML (Tom's Obvious, Minimal Language), INI (Initialization File Format), and Protocol Buffers (Protobuf).
[0133] Table 1 shows an example of a JSON-based text file for a service configuration file containing two microservices (ms1 and ms2) and two communications (ms1 to ms2 and ms2 to ms1).
[0134] Table 1
[0135]
[0136]
[0137] This disclosure designs a virtualization framework for application developers and organizations to test, evaluate, and configure their distributed applications. This framework, known as the Service Configuration System, enables the identification of resource and communication requirements for each microservice in an application. The resulting configuration files, known as "service profiles," provide crucial guidance for deploying applications in real-world environments.
[0138] The service configuration system allows application developers to deploy their distributed applications into the framework. The system semi-automatically analyzes and identifies the resource requirements of each microservice, as well as the communication requirements between microservices. Specifically, the system determines parameters such as CPU utilization, memory allocation, latency constraints, and the amount of data exchanged between microservices.
[0139] Once a service profile is generated, it serves as a guideline for the orchestrator in real-world deployment environments. The orchestrator uses the service profile to make intelligent decisions about allocating microservices across available computing resources to provide optimal performance and resource utilization.
[0140] A service profile may include various parameters that define the resource requirements and capabilities of each microservice. For example, a CPU requirement specifies the number of CPU cores and their clock rates, taking into account variations in core performance (e.g., faster or slower cores). Additionally, the system considers other CPU features, such as processing power and architecture, to provide the optimal allocation.
[0141] A microservice can be implemented as a software package. This package can be stored locally on the device, on a remote network server, or at a specific location identified by a URL or IP address.
[0142] The service profile may also include parameters related to memory requirements, storage requirements, and scalability. For example, some microservices may be designed to be replicated, similar to web servers (e.g., google.com) that are replicated across multiple locations to handle high traffic. The system evaluates whether a microservice can be replicated and whether resource reservations are required to guarantee performance.
[0143] Microservices may have different priorities, some of which require higher priority due to their critical roles in the application. A service profile defines these priorities, as well as communication parameters between microservices such as data transfer rates, latency constraints, and bandwidth requirements.
[0144] A service profile might define the communication parameters between microservices, including the source (sending microservice) and the destination (receiving microservice). Key parameters include message latency requirements, latency jitter tolerance, and sensitivity to violations (e.g., whether a latency or jitter violation is critical or non-critical). Bandwidth requirements, such as data volume and transmission speed, are also specified. In addition, the system considers the number of messages per second, packet error rate, and tolerance for packet loss. For example, in voice communication, the loss of a few packets might be acceptable, while in data transmission, even a single lost packet might be critical.
[0145] The service configuration file is stored in a text file format, which may include well-known formats such as JSON, XML, YAML, TOML, or INI. The specific format is not limited as long as it contains the required parameters. In one embodiment, JSON is used as an example format.
[0146] The service profile includes detailed parameters for each microservice. For example, microservice 1 (μs1) may require 10 MIPS (million instructions per second), a central processing unit clock speed of 3 GHz, and a memory allocation of 100 megabytes. The microservice is stored and can be downloaded at a specified URL, with a file name and image format (e.g., IMG). The storage requirement may be set to 10 megabytes, and the service availability may be specified as 100% (i.e., continuously running). Similarly, microservice 2 (μs2) includes its own set of parameters, as defined in the service profile.
[0147] Service profiles also consider QoE requirements to provide optimal performance and user satisfaction. These requirements are incorporated into microservice parameters to guide resource allocation and system behavior.
[0148] A service profile might define communication parameters for each type of communication between microservices. For example, communication type 1 (from microservice 1 (MS1) to microservice 2 (MS2)) might specify a data rate of 100 megabits per second, an average message size of 500 bytes, link utilization, latency, jitter, and packet loss. In some cases, communication might be unidirectional, depending on the functional requirements of the microservice. Therefore, these communication parameters are for one-way communication, and if applicable, corresponding parameters need to be defined for the reverse direction (e.g., from MS2 to MS1).
[0149] In scenarios involving multiple microservices, the communication paths can be more complex. For example, Microservice 1 (MS1) might send a message to Microservice 2 (MS2), which might send a message to Microservice 3 (MS3), which might respond to MS1. Service profiles capture these communication patterns and their associated parameters.
[0150] At the end of the analysis process, the system generates a service profile text file. This file contains all defined parameters, including the resource requirements for each microservice and the communication parameters between multiple interoperating microservices. The service profile serves as a comprehensive guide for deploying and orchestrating applications in a distributed environment.
[0151] Figure 9 An example visualization of a service profile is shown, where computing resource requirements are shown within six ovals 902 - 912 representing six microservices ( μs1 through μs6 ), and communication requirements are indicated along ten arrowed lines connecting the microservices.
[0152] Figure 9 It shows detailed information about six microservices and their associated resource requirements, such as CPU utilization, memory allocation, and the URL locations of the microservices. The arrowed lines in the visualization represent the communication directions between microservices and their associated parameters, such as data rate, latency, jitter, and packet loss.
[0153] A service profile serves as a comprehensive template that outlines the necessary information needed to deploy and orchestrate a distributed application. It includes the resource requirements of each microservice (e.g., CPU, memory) and communication parameters between microservices (e.g., latency, bandwidth).
[0154] Automatic service profile generation (service profiling)
[0155] A distributed application is an application or software that runs on multiple computers in independent devices, appearing to the user as a single system. This differs from traditional applications, which typically run on a single physical system. In the development of distributed applications, abstractions are often used to hide the physical separation of each software module as much as possible.
[0156] Figure 10 The high-level protocol architecture for distributed applications, where software modules may be running on remote devices, is presented. From the application developer's perspective, communication between client module 1010 and server module 1020 may appear no different than a monolithic application, as this is achieved through client API 1030 and server API 1040, effectively "hiding" the underlying network communication.
[0157] Therefore, application developers do not need to pay attention to the bottom two layers, namely, the client connection layer 1050, the server connection layer 1060 and the network layer 1070 and the network layer 1080. Figure 10 shown.
[0158] Therefore, building service profiles for distributed applications is a challenge for typical application developers, especially in determining network performance requirements. These requirements can be a source of significant performance issues and may be beyond the domain knowledge of some developers.
[0159] To address the difficulty of building service profiles for distributed applications and promote the adoption of distributed computing resource sharing systems, an automated service configuration mechanism is proposed. Together with the Service Profiling Development Kit (SDK), this mechanism aims to automatically generate service profiles for application developers.
[0160] Figure 11A and Figure 11B An example of an automated service configuration method is presented. The following text describes the steps and procedures involved in this method.
[0161] Step 1: Application developers develop distributed applications
[0162] In this step, application developers develop distributed applications in the usual way. Figure 11A As shown in Step 1 of the , developers develop their software by implementing individual microservice packages, such as microservice 1, microservice 2, microservice 3, microservice 4, microservice 5, and microservice 6, denoted as μs1, μs2, μs3, μs4, μs5, and μs6, respectively. Once these individual service packages are implemented, developers can establish communication links between them, allowing the six microservices to run together as a single application. It's important to note that although the application consists of multiple parts, it may still operate as a unified system.
[0163] Step 2: Configure the initial deployment architecture (pre-service configuration file)
[0164] In this step, the application developer prepares an initial deployment configuration for the distributed application, called a pre-service configuration file. This configuration includes the connectivity between distributed modules or microservices, such as a connectivity graph between microservices, such as Figure 11A The arrowed lines are shown in step 2. However, there is no need to specify the computational and communication resource requirements of the microservices or the communication links between them.
[0165] In other words, developers know the direction of communication between microservices. For example, Microservice 1 communicates with Microservice 2, Microservice 2 or Microservice 3 sends messages to Microservice 1, and so on. However, in this step, developers do not have detailed knowledge of specific resource utilization requirements, such as message rates, latency requirements, or throughput requirements, which are useful for providing optimal application performance.
[0166] Specifically, from the user's perspective, the performance of an application is evaluated based on its functionality and responsiveness. For example, in a gaming application, if a user shoots a bullet, it should hit the target as expected, not jitter or disappear unexpectedly. If another user's bullet arrives first despite being fired later, the performance is considered unsatisfactory. This highlights the importance of understanding the communication requirements (represented as arrows) and resource requirements of microservices.
[0167] If a microservice is mistakenly placed far away from other microservices, or if the network connection between devices is poor (for example, a low-bandwidth communication channel), delays will occur, leading to poor performance and user dissatisfaction. Although application developers understand the structural relationships between microservices, they lack the detailed information necessary to establish the required performance levels.
[0168] Step 3: Runtime Software Analysis / Profiling
[0169] In this step, if Figure 11AAs shown, automated service provisioning begins in a high-capacity computing environment. This environment might include a single powerful server, a cluster of multiple servers, virtual machines on a single machine, or a cloud network that provides sufficient computing and communication resources to meet the Quality of Service (QoS) or Quality of Experience (QoE) requirements of distributed applications. Users install applications or input pre-service configuration files into a service provisioning development kit (SDK) installed on a high-capacity platform. This platform might run in a virtual, physical, or hybrid environment.
[0170] The SDK may then create an application cluster consisting of multiple virtual nodes or containers, or a combination of both, including physical nodes. It then deploys microservices based on the configuration in the pre-service configuration file. For example, Figure 11A As shown, microservices μs1, μs2, μs3, μs4, μs5, and μs6 are deployed in virtual machines (VMs) / containers 1110-1160, respectively.
[0171] In this phase, sufficient computing resources, including CPU, memory, and storage, as well as communication resources such as high link capacity, possibly with the lowest or no packet error rate, and lowest or no packet latency, are allocated. These resources are provided so that the application achieves optimal performance levels under ideal computing and communication conditions for benchmark evaluation.
[0172] Each virtual node may contain one or more virtual network interfaces, depending on the target cluster network environment to be evaluated, such as Figure 7 and Figure 8 For example, Figure 8 As shown, devices A and D contain three network interfaces, while devices B, C, E, and F each contain two. If multiple network interfaces are required for each node and static connections are required between nodes, the pre-service configuration file should also include the network configuration so that the orchestrator can build the application cluster accordingly. The network infrastructure may be entirely virtual, configured through network emulation, or combined with real network equipment such as 4G, 5G, Wi-Fi, and Bluetooth.
[0173] While your application is running, the SDK monitors and records its behavior and performance. This includes CPU, memory, and storage utilization per node; network link utilization, such as packet rate, bandwidth, message rate, and traffic patterns; and network performance metrics between connected nodes, such as packet latency, latency variation, packet error rate, and message error rate. Monitoring and result feedback are supported externally by the SDK. Additionally, applications may provide internal monitoring and result feedback.
[0174] Step 4: Runtime Software Analysis / Profiling
[0175] In step 4, if Figure 11B As shown, multiple service profiles containing various combinations of computing and communication resources were deployed sequentially. During this process, Quality of Service (QoS) measurements and User Experience (QoE) feedback for the applications were collected. There are two types of QoE: objective and subjective.
[0176] Objective QoE: Application performance can be evaluated based on quantifiable measurements such as observed data rate, packet latency, latency variation, and packet loss rate (implicit feedback). If the performance of an application can be objectively evaluated based on these measurements, multiple levels of QoE guidelines can be configured, and service profiles can be automated based on these guidelines. This scenario does not require explicit feedback from humans / users. Some real-time applications may fall into this category, and application QoS measurement may be good enough because it can be mapped to QoE.
[0177] Subjective QoE: This category requires explicit human feedback because the level of QoE is determined by the user's perception of service availability during their operation. While the application is running, users may provide feedback in the form of a score (e.g., on a scale of 0 to 10) based on how the availability of CPU, memory, and storage resources on each node, as well as the network communication conditions of each connection, affect application performance and QoE. Evaluating application QoE under this category may require a group of users to conduct performance evaluations to obtain representative QoE information.
[0178] The multiple service profiles contain various combinations of computing resource amounts and communication capacity resources, which will be included in the final service profile. The results obtained in step 3 are used as upper bounds (if larger values are preferred) or lower bounds (if smaller values are preferred) for evaluating the allocation of computing and communication resources in the various service profiles.
[0179] We introduced a service profiling system to evaluate the performance of six microservices. In this system, the microservices are hosted on high-capacity machines, such as supercomputers with large numbers of CPU cores (e.g., thousands of cores) and multiple terabytes of memory, all within a single machine. Because all components are internal, communication latency between microservices is minimal, typically on the order of a few microseconds, enabling extremely fast communication.
[0180] These microservices are virtually separated internally through virtualization technologies such as virtual machines (VMs) or containers. For example, in a cloud environment like Amazon, six VMs can be created, each running on the same physical machine but logically isolated from each other. This isolation enables them to behave as independent machines while using the high-performance resources of the underlying hardware.
[0181] In this virtual environment, each virtual machine is allocated a large number of CPU cores and a generous amount of memory, sufficient to support any type of application. Furthermore, the system provides ultra-fast inter-process communication because all microservices reside on the same physical machine. This setup allows applications to run without resource constraints, enabling each microservice to reach its maximum potential.
[0182] By running the application under these ideal conditions, the system can measure the resource utilization of each microservice and monitor all communications between them. This process establishes an upper bound on resource utilization that represents the best performance achievable under unconstrained conditions.
[0183] The detailed procedure in step 4 is described below:
[0184] (i) The orchestrator prepares a service profile that contains a different combination of compute and communication resource allocations than the previously evaluated service profile. In an exemplary embodiment, the new service profile may change the number of worker nodes in the cluster (a change in cluster information / topology), the amount of CPU and memory per node, and other parameters of the service profile. The service configuration system provides virtual network simulation for each communication connection (channel) between microservices for packet latency, latency variation, packet loss, and link bandwidth. Network link and / or network protocol simulation can be applied to both virtual and physical network connections. For parameters that use the upper bound result in step 3, a lower bound constraint needs to be provided to the service configuration SDK to select an evaluation value within a specified range. Conversely, for parameters that use the lower bound result in step 3, an upper bound constraint needs to be provided to the service configuration SDK to select an evaluation value within a specified range. Alternatively, a range value can also be provided as input.
[0185] (ii) A candidate service profile, which includes the selected combination of evaluation values as a resource allocation guide, is recorded in database 1170 along with cluster information. The selected service profile is then input into the orchestrator to modify the existing cluster, adjust the computational and communication resource allocation, and update the network simulation based on the resource allocation guide provided in the new service profile. These operations are performed in real time.
[0186] (iii) Continuously monitor the application's behavior and performance while it runs. This includes CPU, memory, and storage resource utilization for each node or microservice, as well as real-time network link utilization and performance metrics such as message rate, message size, link latency, link latency variation, and packet loss between connected nodes or microservices. These monitoring results are recorded in a database. Real-time performance monitoring and measurements should correspond to the service profile used during the measurement period. The frequency of recording depends on the desired granularity of analysis. Real-time measurements should be designed in a way that does not significantly impact the utilization or performance of computing and communication resources within the service configuration system environment.
[0187] (iv) While the application is running, user experience (QoE) feedback is collected from application users and recorded in a database along with the corresponding service profile and real-time application performance monitoring results. These records, including cluster information, service profiles, application runtime behavior, performance measurements, and user feedback, serve as the source for generating the final service profile. For applications in the objective QoE category, QoE feedback collection can be performed directly within the application or in the service configuration system environment. However, for applications in the subjective QoE category, feedback is collected through a user input mechanism in the service configuration system environment, which may operate outside the application.
[0188] (v) Repeat the procedure in steps (i) to (iv) until all prepared candidate service profiles are evaluated.
[0189] Automating the construction of service profiles can be done in at least two different ways:
[0190] Service Profile Type 1: A series of statically configured candidate service profiles are constructed, each with varying amounts of compute and communication resource allocation. These profiles are applied sequentially, and measurement results and user feedback are collected. Once all candidate service profiles have been evaluated, a single-tier or multi-tier (layered) service profile is created.
[0191] Service Profile Type 2: Within the upper and lower bounds determined in step 3, or within the range information provided, a series of dynamically or automatically configured candidate service profiles are constructed, with different amounts of compute and communication resource allocations. Changes in a large amount of compute and communication capabilities will be applied dynamically in real time, and users will provide QoE feedback as needed. Once a large amount of data is collected, a single-layer or multi-layer (hierarchical) service profile is created. This process can use machine learning (ML) modeling, and the orchestrator can use the ML model to predict QoE based on the available compute and communication resources and the state of the cluster environment, which is done before actual deployment in the actual environment.
[0192] Step 5: Figure 11B As shown, the runtime behavior and performance of monitored applications under various computational and communication resource allocations along with user feedback are analyzed and evaluated.
[0193] Step 6: Figure 11B As shown, a single-layer or multi-layer (hierarchical) service profile is generated, and / or an ML model for QoE prediction is built. The final service profile, ML model, or both are deployed with the corresponding distributed application. During application runtime, the orchestrator deploys the distributed application according to the service profile to maintain the application's QoS and / or QoE. If the current cluster network conditions change, the ML model can be used to predict the application's QoE in real time. The ML model can be used to dynamically build or modify the cluster to maintain or improve the application's QoE throughout the application's lifecycle.
[0194] Method for verifying automated service profile building system
[0195] Once the automated service profile building system and service profile development kit (SDK) are implemented according to the service profile building steps and procedures described above, it may be necessary to verify or test the accuracy and functionality of the automated service profile building system before implementing or deploying real distributed applications. To do this, at least the following two functions need to be implemented:
[0196] Load Generator: Implements a general-purpose microservice, such as an application load generator, to adjust CPU load (static or distributed), memory utilization (static or distributed), and network link load and utilization, including message size (static or distributed), message rate, and message rate variation (e.g., CBR, VBR). Load generators can also handle one-way and two-way message patterns, as well as message service types such as request / response or publish / subscribe between destination nodes. The same microservice-based load generator can be instantiated in multiple different containers, pods, or nodes with different configurations. They can effectively simulate distributed applications with configurable communication and compute resource requirements. CPU and memory load may include multiple factors, such as application processes, kernel processes, network processes, etc. These factors can be configured independently as contributions to CPU and memory load or consumption, or, for simplicity, they can be expressed as an aggregated load without considering the individual factors.
[0197] Performance metrics or visualization modules / functions: Performance or measurement metrics may include real-time monitoring of CPU and memory resource utilization per node, as well as real-time network link utilization metrics such as message rate, message size, link latency, link latency variation, and inter-node packet loss. Providing visualization capabilities or metrics of real-time monitoring results allows users to easily view or evaluate the performance of potential simulated applications. This visualization capability allows users to effectively provide QoE feedback.
[0198] One or more load generation modules / microservices can be configured with a hypothetical application microservice for CPU, memory, and storage load in step 1, and a hypothetical communication connection load and utilization in step 3.
[0199] Verification or testing can be performed iteratively. That is, multiple test cycles can be performed, with resource allocations varied during each cycle. For example, microservice 1 might be allocated fewer compute resources than the maximum observed during initial testing. Similarly, throughput might be reduced, or latency and delay variations might be introduced in communication between microservices.
[0200] By systematically varying the mix of allocated resources and running the same application under these varying conditions, the system simulates real-world deployment scenarios. In real deployments, microservices might run on different devices with varying computing resources, and performance can vary significantly between user devices. The goal is to replicate these diverse conditions in a virtual environment.
[0201] During each testing cycle, monitor and record the application's behavior. Because applications are designed for human use, incorporate human involvement into the testing process. For example, users might play a game and provide feedback about their experience. This feedback might include subjective responses such as "I didn't feel anything," "It was okay," "Very good," or "Terrible."
[0202] Some applications have quantifiable performance metrics. For example, in a voice call application, performance is considered poor if voice packet latency exceeds 150 or 200 milliseconds. However, most applications rely on subjective performance assessments based on user perception. To capture this, users can provide feedback using a rating system, such as a scale from 0 to 10, where 10 represents excellent performance and 0 represents poor performance.
[0203] The information collected during each test cycle is recorded and the resource mix is adjusted for the next cycle. Feedback is collected repeatedly and the process is iterated to accumulate sufficient data. Once sufficient data is collected, the results can be processed in two ways:
[0204] Rule-based policies: If data is insufficient to train a machine learning model, create rule-based policies to define application behavior based on available resources.
[0205] Machine Learning Model: If sufficient data is available, a machine learning model is generated. This model represents the behavior of the application under various resource conditions.
[0206] The testing process can involve different combinations of virtual devices. Although the system uses six devices as an example, real-world testing might use a single machine, two machines, three machines, or more, depending on available time and resources. This flexibility allows for testing a wide range of configurations.
[0207] The testing process follows these steps:
[0208] (1) Select a resource combination.
[0209] (2) Apply a selected set of computing and communication resources.
[0210] (3) Deploy the application and record relevant information.
[0211] (4) Execute the test and collect feedback corresponding to the current conditions.
[0212] (5) Repeat this process using different combinations of resources until enough data has been collected. Once all the information has been collected, the analysis phase begins.
[0213] Although the system is automated, depending on the application, human involvement may be required to test it from start to finish.
[0214] During testing, resource combinations can be adjusted in real time to dynamically capture user feedback.
[0215] The output of this process includes, but is not limited to, a single service configuration file as a text file. When using a machine learning model, it may include:
[0216] (1) A layered service profile containing multiple levels of requirements.
[0217] (2) A series of service profile requirements based on different resource conditions.
[0218] Machine learning models may predict application behavior under different resource allocations.
[0219] Results are not binary (e.g., “good” or “bad”), but are expressed as a range of scores, such as from 0 to 10, that reflect the detailed performance characteristics of the application.
[0220] The quality of application performance can range widely. If a component fails completely, it can cause the system to crash. However, if performance is merely degraded, users may still consider the application acceptable. Machine learning models can provide resource requirements to the orchestrator, enabling the system to provide services based on resource availability, even under suboptimal conditions.
[0221] During testing, virtual machines are used to simulate the communication links between different microservices. This simulation allows for manipulation of communication parameters, such as reducing communication speed or introducing latency. Furthermore, the system can simulate various communication technologies, including wireless communication and different speed configurations. These parameters are incorporated into the simulation process.
[0222] Simulations may include multiple performance metrics, including, for example:
[0223] (1) Message rate: the number of messages per second.
[0224] (2) Packet Error Rate: The rate at which packets are lost or damaged.
[0225] (3) Bandwidth: Link capacity, which can be limited to simulate low-bandwidth conditions (e.g., 1 Mbps or 100 Mbps).
[0226] (4) Latency, delay, and jitter: These parameters are simulated to replicate real-world network distortions.
[0227] By incorporating these factors into link emulation, applications experience realistic performance variations, making testing and evaluation accurate.
[0228] As mentioned above, the automated service profile system addresses a fundamental challenge: how to accurately determine the resource requirements for optimal application performance without requiring developers to have detailed knowledge of the underlying network architecture.
[0229] In some implementations, the service profile process involves deploying the distributed application in a high-capacity computing environment that provides sufficient resources for benchmark evaluation. Such an environment can be implemented as a powerful server with thousands of central processing unit cores and multiple terabytes of memory, a cluster of multiple servers, or a cloud-based infrastructure with sufficient computing and communication resources to meet the Quality of Service (QoS) and Quality of Experience (QoE) requirements of the distributed application.
[0230] In this environment, the Service Profiling Development Kit (SDK) creates an application cluster consisting of multiple virtual nodes or containers, possibly including physical nodes, and deploys microservices based on pre-configured service profiles provided by the developer. Each virtual node may contain one or more virtual network interfaces, configured to match the target cluster network environment to be evaluated. The network infrastructure can be entirely virtual, configured through network emulation, or mixed with real network equipment, such as 4G, 5G, Wi-Fi, and Bluetooth interfaces.
[0231] In this phase, sufficient computing resources (central processing units, memory, and storage) and communication resources (high link capacity, minimal packet error rate, and latency) are allocated to allow applications to run at optimal levels under ideal conditions. This establishes an upper limit on resource utilization, representing the maximum performance achievable if resources were not constrained.
[0232] After the initial baseline assessment, the service profiling process moves into a phase where the system varies resource allocations to evaluate application performance under different conditions. This involves creating multiple service profiles with different resource combinations and deploying them sequentially. For each profile, the system monitors application behavior and collects performance metrics, which may include CPU utilization, memory usage, network traffic patterns, and user experience feedback.
[0233] Resource allocation testing can follow two different approaches. In the first approach (Service Profile Type 1), a set of statically configured candidate service profiles with various resource allocations are prepared in advance and applied sequentially. In the second approach (Service Profile Type 2), the system dynamically generates candidate service profiles with resource allocations within a specified range, applies them in real time, and collects performance data. The second approach is particularly suitable for machine learning model development because it can generate larger and more diverse datasets.
[0234] During each test cycle, the orchestrator prepares a service profile with a unique combination of compute and communication resource allocations. The service profile system provides virtual network emulation capabilities that manipulate parameters such as packet latency, latency variation (jitter), packet loss, and link bandwidth for each communication channel between microservices. These parameters are systematically varied within defined ranges to simulate different network conditions and resource constraints.
[0235] As the application runs in each resource allocation scenario, the SDK continuously monitors and records the application's behavior and performance. This monitoring covers several dimensions:
[0236] Computing resource utilization: CPU, memory, and storage usage per node or microservice;
[0237] Network link utilization: packet rate, bandwidth consumption, message rate, and traffic patterns;
[0238] Network performance metrics: packet latency, latency variation, packet error rate, and message error rate between connected nodes.
[0239] These measurements are recorded in a database along with the corresponding service profile used during the measurement. The frequency of recording depends on the desired granularity of analysis. The measurement process itself does not significantly impact the computational and communication resource utilization or performance within the service profile system environment.
[0240] In some implementations, the system supports two different approaches to quality of experience (QoE) evaluation: objective and subjective. For applications with objectively measurable performance metrics, such as observed data rate, packet delay, delay variation, and packet loss rate, the system can automatically evaluate QoE without human intervention. This approach is suitable for certain real-time applications where quality of service (QoS) measurements can be directly mapped to QoE, eliminating the need for explicit user feedback.
[0241] For applications where performance evaluation is inherently subjective, the system incorporates a human feedback mechanism. Users interact with the application under various resource conditions and provide ratings based on their perceived experience (e.g., on a scale of 0 to 10). This feedback is recorded along with the corresponding service profiles and performance measurements, creating a comprehensive dataset that links resource allocation to user satisfaction. Subjective QoE evaluation may require evaluations from multiple users to obtain representative QoE information.
[0242] Whether objective or subjective, QoE feedback is collected and stored in a database along with the corresponding service profiles and real-time application performance monitoring results. These records serve as source data for generating final service profiles or training machine learning models for QoE prediction.
[0243] After sufficient data has been collected through multiple testing cycles, with resource allocations constantly changing, the system analyzes the application's runtime behavior and performance, along with any user feedback. Based on this analysis, the system generates either a single-tier service profile or a multi-tier (layered) service profile that defines the resource requirements for optimal application performance.
[0244] When sufficient data is collected through the dynamic service configuration approach (Type 2), the system can build a machine learning model for QoE prediction. This model captures the relationship between resource allocation and application performance, enabling the orchestrator to predict how changes in resource allocation will affect application performance before implementing these changes.
[0245] Machine learning models are very useful for applications with subjective QoE requirements, as it may not be possible to directly measure performance quality. They allow orchestrators to make informed resource allocation and microservice placement decisions during actual deployments, promoting proactive optimization rather than reactive tuning.
[0246] The final service profile, machine learning model, or both are deployed along with the corresponding distributed application. During application runtime, the orchestrator uses the service profile to distribute microservices across available worker nodes and configure communication channels to meet the application's QoS and QoE requirements.
[0247] If a machine learning model has been developed, it can be used to predict the QoE of an application under different cluster network conditions in real time. As network conditions change, the model helps the orchestrator dynamically build or modify the cluster to maintain or improve the application's QoE over its lifetime.
[0248] Service profiles enable continuous monitoring and tuning of applications during operation. Each microservice or its underlying worker nodes continuously monitors the application's behavior and performance, feeding this information back to the orchestrator. If observed performance falls short of the expected level specified in the service profile, the orchestrator can take corrective actions, such as removing underperforming devices, adding new ones, changing network connections, or configuring network schedulers and protocols to improve application performance.
[0249] The Service Profile Development Kit (SDK) implements a framework for automating the generation of service profiles. The SDK creates a controlled test environment that can evaluate application performance under various resource conditions. This environment is implemented on a high-capacity computing platform, which can be a powerful server with thousands of CPU cores and multiple terabytes of memory, a cluster of multiple servers, or a cloud-based infrastructure with sufficient computing and communication resources.
[0250] The SDK architecture consists of several components:
[0251] Virtualization layer: Creates and manages virtual machines or containers for each microservice, providing complete isolation while maintaining high-speed inter-process communication.
[0252] Network Simulation Module: Simulates various network conditions by manipulating parameters such as bandwidth, latency, jitter, and packet loss between each communication link.
[0253] Resource Allocation Manager: Controls the allocation of compute resources (CPU, memory, storage) to each virtual node or container based on the current service profile being tested.
[0254] Performance monitoring system: Collects detailed metrics on resource utilization, network performance, and application behavior during testing.
[0255] Data collection and storage: All performance measurements, service profiles, and user feedback are recorded in a structured database for analysis.
[0256] Profile Generation Engine: Analyzes collected data to generate service profiles or train machine learning models.
[0257] In some implementations, service configuration can simulate various network conditions between microservices. The SDK incorporates network emulation capabilities that can manipulate communication parameters such as bandwidth, latency, latency variation (jitter), and packet error rate. These parameters can be adjusted for each communication link between microservices, allowing for comprehensive testing of application performance under diverse network conditions.
[0258] Network emulation can be applied to both virtual and physical network connections. For virtual connections, emulation occurs entirely in software. For physical connections, emulation may involve actual network equipment, such as 4G, 5G, Wi-Fi, and Bluetooth interfaces. This flexibility enables testing applications in a wide range of potential deployment scenarios, from ideal high-bandwidth, low-latency environments to constrained or degraded network conditions.
[0259] The Network Emulation module supports simulation of various communication technologies and their characteristic performance parameters. For example, it can simulate the higher latency and jitter typically associated with cellular networks, or Wi-Fi networks, which may have higher contention but lower latency. This capability is crucial for understanding the performance of applications in different connection scenarios that may be encountered in real-world deployments.
[0260] During the profiling process, the system employs monitoring mechanisms to capture detailed performance metrics across multiple dimensions. For compute resources, the system tracks CPU utilization patterns, memory consumption behavior, and storage access patterns at fine-grained intervals. Network performance monitoring includes detailed measurement of inter-microservice communication, capturing metrics such as message rate, payload size, end-to-end latency, and jitter variation.
[0261] These measurements are time-stamped and associated with a specific resource allocation configuration, enabling precise analysis of performance characteristics under a variety of conditions. The data collection process is designed to be as minimally intrusive as possible. The monitoring itself does not significantly affect the performance being measured.
[0262] Performance data is stored in a structured database that maintains the relationships between service profiles, resource allocations, performance measurements, and user feedback. This comprehensive dataset is the basis for service profile generation and machine learning model training.
[0263] Once sufficient performance data has been collected, the system can use machine learning techniques to build predictive models of application behavior. These models are trained on the collected performance data to learn the relationship between resource allocation, communication patterns, and resulting performance indicators.
[0264] Machine learning models can predict the quality of experience (QoE) of applications under different resource conditions, enabling orchestrators to make proactive decisions about resource allocation and microservice deployment. This is very useful for applications with subjective quality of experience requirements, where it may not be possible to directly measure performance quality.
[0265] Machine learning methods can use a variety of algorithms depending on the nature of the application and the available data. Regression models can be used to predict continuous performance metrics, while classification models might categorize performance into discrete quality levels. More complex methods, such as reinforcement learning, can be used to optimize resource allocation strategies over time based on observed performance results.
[0266] To verify the accuracy and functionality of the automatic service profile building system before deploying real distributed applications, a dedicated testing component was implemented: the load generator, a general-purpose microservice that adjusts CPU load, memory utilization, and network link characteristics.
[0267] The load generator supports both static and distribution-based resource utilization modes, allowing it to simulate a wide range of application behaviors. It can handle different messaging patterns (unidirectional / bidirectional) and service types (request / response or publish / subscribe) between target nodes. Multiple load generator instances can be deployed with different configurations to simulate complex distributed applications with specific resource requirements.
[0268] The load generator can create configurable workload patterns that stress different aspects of the system, including CPU-intensive operations, memory-intensive tasks, and various communication patterns. This flexibility allows a wide range of application scenarios to be tested on the service profile system.
[0269] Complementing the load generator is a performance visualization module that provides real-time monitoring of resource utilization and network performance metrics. This visualization capability enables testers to observe system behavior under different conditions and provide quality of experience feedback, simulating the evaluation process of real applications.
[0270] The system implements a dynamic adjustment mechanism for service profiles based on observed performance and available resources. When the orchestrator detects that the current resource allocation cannot meet the performance requirements, it can automatically adjust the service profile within a predefined range.
[0271] This adjustment may involve switching to a lower tier of service quality requirements or modifying resource allocation patterns while maintaining the application's essential functionality. The adjustment process takes into account both immediate performance requirements and long-term resource availability trends.
[0272] For applications with multiple tiers of service profiles, the system can dynamically switch between different profile tiers based on current conditions. For example, if network congestion increases, the system might switch from a "desirable" profile to a "minimal" profile, which requires fewer resources while still maintaining acceptable performance.
[0273] The adjustment mechanism is guided by the quality of experience prediction capabilities of machine learning models that can predict the impact of resource changes on the user experience. This predictive capability enables the system to make intelligent decisions about when and how to adjust the service profile to maintain optimal performance under changing conditions.
[0274] Figure 12 A flowchart for automatic service profile generation is shown 1200. The process may be implemented by one or more computing devices.
[0275] At block 1202 , one or more computing devices obtain an initial deployment configuration specifying connectivity between a plurality of microservices of a distributed application.
[0276] At block 1204 , one or more computing devices deploy a plurality of microservices in a high-capacity computing environment.
[0277] At block 1206 , one or more computing devices monitor resource utilization and performance metrics while executing the distributed application in the high-capacity computing environment and providing sufficient resources.
[0278] At block 1208 , the one or more computing devices generate a plurality of candidate service profiles by changing resource allocations of the plurality of microservices.
[0279] At block 1210 , one or more computing devices collect performance measurements and quality of experience feedback for each candidate service profile.
[0280] At block 1212 , one or more computing devices generate a final service profile based on the collected performance measurements and quality of experience feedback.
[0281] In some configurations, deploying the plurality of microservices may include: creating virtual nodes in a high-capacity computing environment; and instantiating each of the plurality of microservices in at least one of the virtual nodes.
[0282] In some configurations, monitoring resource utilization and performance metrics may include: measuring at least one of central processing unit utilization, memory utilization, and storage utilization of each microservice; and measuring at least one of network performance metrics including message rate, message size, link latency, and packet loss rate between communicating microservices.
[0283] In some configurations, generating the multiple candidate service profiles may include: selecting different combinations of computing and communication resource allocations; applying network simulation to simulate different network conditions between the microservices; and recording the resource allocations and network conditions in a database.
[0284] In some configurations, applying network emulation may include simulating at least one of: packet delay; delay variation; bandwidth limitation; and packet loss rate.
[0285] In some configurations, collecting quality of experience feedback may include at least one of: collecting objective measurements of applications with quantifiable performance metrics; and collecting subjective user feedback ratings of applications requiring human evaluation.
[0286] In some configurations, the process may further include: training a machine learning model using the collected performance measurements and quality of experience feedback; and using the trained machine learning model to predict application performance under different resource conditions.
[0287] In some configurations, generating the final service profile may include: generating a tiered service profile comprising one or more tiers of resource requirements; and specifying minimum and desired performance levels for each tier.
[0288] In some configurations, the process may further include validating the generation of the service profile by: implementing a load generator to simulate a configurable workload pattern; and monitoring application behavior under the simulated workload pattern.
[0289] In some configurations, the load generator may be configured to: generate tunable CPU and memory utilization patterns; and simulate different messaging patterns between microservices.
[0290] In some configurations, the final service profile may include: the computational resource requirements for each microservice; and the communication requirements between microservices.
[0291] In some configurations, communication requirements may specify: maximum tolerable message delay; bandwidth requirements; message rate; and packet loss rate tolerance.
[0292] In some configurations, the process may further include: monitoring runtime performance of the distributed application using the final service profile; detecting when performance requirements are not met; and dynamically adjusting resource allocation to maintain application performance.
[0293] In certain configurations, a high-capacity computing environment may include at least one of: a single server with multiple central processing unit cores; a cluster of multiple servers; and a cloud computing infrastructure.
[0294] In some configurations, the process may further include: storing the final service configuration file in a structured data format; and deploying the final service configuration file with the distributed application for use by an orchestrator in managing resource allocation.
[0295] It is understood that the specific order or hierarchy of blocks in the disclosed processes / flowcharts is illustrative of example methods. Based on design preferences, it is understood that the specific order or hierarchy of blocks in the processes / flowcharts may be rearranged. Furthermore, certain blocks may be combined or omitted. The accompanying method claims present elements of the various blocks in an example order and are not meant to be limited to the specific order or hierarchy presented.
[0296] The foregoing description is provided to enable those skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Accordingly, the claims are not intended to be limited to the aspects shown herein, but rather to the full scope consistent with the claim language, wherein a singular reference to an element does not mean "only one" unless specifically stated as such, but rather "one or more." The word "example" is used herein to mean "serving as an example, instance, or illustration." Any aspect described as an "example" is not necessarily to be considered superior or advantageous over other aspects. Unless otherwise expressly stated, the term "some" refers to one or more. Combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "A, B, C, or any combination" include any combination of A, B, and / or C, and may include multiple A's, multiple B's, or multiple C's. Specifically, combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "A, B, C, or any combination" may be A only, B only, C only, A and B, A and C, B and C, or A and B and C, where any such combination may contain one or more members of A, B, or C. All structural and functional equivalents to the various aspects described in this document that are known or later become known to those skilled in the art are hereby incorporated herein by reference and are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be dedicated to the public regardless of whether expressly recited in the claims. The words "module," "mechanism," "element," "device," etc. may not be used in place of the word "means." Therefore, no claim element should be construed as having a means for function unless the element is expressly recited using the phrase "means for."
Claims
1. A method, implemented by one or more computing devices, comprising: Get the initial deployment configuration, specifying the connectivity between multiple microservices of the distributed application; Deploy the multiple microservices in a high-capacity computing environment; monitoring resource utilization and performance metrics while executing distributed applications in this high-capacity computing environment and providing sufficient resources; By changing the resource allocation of the multiple microservices, multiple candidate service configuration files are generated; Collect performance measurements and quality of experience feedback for each candidate service profile; as well as A final service profile is generated based on the collected performance measurements and quality of experience feedback.
2. The method of claim 1, wherein deploying multiple microservices comprises: creating a virtual node in the high-capacity computing environment; as well as Each of the plurality of microservices is instantiated in at least one of the virtual nodes.
3. The method of claim 1 , wherein monitoring resource utilization and performance indicators comprises: Measuring at least one of central processing unit utilization, memory utilization, and storage utilization of each microservice; as well as Measuring at least one of the network performance indicators including message rate, message size, link delay and packet loss rate.
4. The method of claim 1 , wherein generating a plurality of candidate service profiles comprises: Select different combinations of computing and communication resource allocation; Apply network simulation to simulate different network conditions between microservices; as well as The resource allocation and the network condition are recorded in a database.
5. The method of claim 4, wherein the applying network simulation comprises simulating at least one of: Packet latency; Delayed changes; Bandwidth limitations; and Packet loss rate.
6. The method of claim 1 , wherein collecting quality of experience feedback comprises at least: Collect objective measurements of applications with quantifiable performance metrics; as well as Collect subjective user feedback ratings for applications that require human evaluation.
7. The method of claim 1, further comprising: Using the collected performance measurements and quality of experience feedback to train a machine learning model; as well as Use trained machine learning models to predict application performance under varying resource conditions.
8. The method of claim 1 , wherein generating the final service configuration file comprises: generating a hierarchical service profile comprising one or more resource requirement tiers; as well as Specify minimum and desired performance levels for each tier.
9. The method of claim 1 , further comprising verifying the generation of the service configuration file by: Implementing load generators to simulate configurable workload patterns; and Monitor application behavior under this simulated workload pattern.
10. The method of claim 9, wherein the load generator is configured to: Generates tunable CPU and memory utilization patterns; and Simulate different messaging patterns between microservices.
11. The method of claim 1 , wherein the final service configuration file comprises: Computational resource requirements for each microservice; as well as Communication requirements between microservices.
12. The method of claim 11, wherein the communication requirement specifies: Maximum tolerable message delay; Bandwidth requirements; message rate; and Packet loss rate tolerance.
13. The method of claim 1, further comprising: monitoring runtime performance of the distributed application using the final service profile; When the test performance requirements are not met; as well as Dynamically adjust resource allocation to maintain application performance.
14. The method of claim 1 , wherein the high-capacity computing environment comprises at least one of: A single server with multiple central processing unit cores; Clusters of multiple servers; and Cloud computing infrastructure.
15. The method of claim 1, further comprising: storing the final service configuration file in a structured data format; as well as The final service profile is deployed with the distributed application for use by the orchestrator in managing resource allocation.
16. A system comprising one or more computing devices, wherein the system is configured to Get the initial deployment configuration, specifying the connectivity between multiple microservices of the distributed application; Deploy the multiple microservices in a high-capacity computing environment; monitoring resource utilization and performance metrics while executing the distributed application in the high-capacity computing environment and providing sufficient resources; By changing the resource allocation of the multiple microservices, multiple candidate service configuration files are generated; Collect performance measurements and quality of experience feedback for each candidate service profile; as well as A final service profile is generated based on the collected performance measurements and quality of experience feedback.
17. The system of claim 16, wherein deploying the plurality of microservices comprises: Creating virtual nodes in the high-capacity computing environment; and Each of the plurality of microservices is instantiated in at least one of the virtual nodes.
18. The system of claim 16, wherein monitoring resource utilization and performance indicators comprises: measuring central processing unit utilization, memory utilization, and storage utilization for at least one of each microservice; and Measuring network performance metrics includes at least one of message rate, message size, link latency, and packet loss rate between communicating microservices.
19. The system of claim 16, wherein generating a plurality of candidate service profiles comprises: Select different combinations of computing and communication resource allocations; Apply network simulation to simulate different network conditions between microservices; and The resource allocation and network conditions are recorded in a database.
20. A computer-readable medium storing computer-executable code for a process implemented by one or more computing devices, comprising code to: Get the initial deployment configuration of connectivity between multiple microservices in a specified distributed application; Deploy the multiple microservices in a high-capacity computing environment; monitoring resource utilization and performance metrics while executing the distributed application with sufficient resources in the high-capacity computing environment; Generate multiple candidate service configuration files by changing resource allocation of the multiple microservices; Collect performance measurements and quality of experience feedback for each candidate service profile; and A final service profile is generated based on the collected performance measurements and quality of experience feedback.