Application-specific reconfigurable network
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NEWPHOTONICS LTD
- Filing Date
- 2026-02-01
- Publication Date
- 2026-08-06
Smart Images

Figure IL2026050102_06082026_PF_FP_ABST
Abstract
Description
GA REF.:766-20APPLICATION-SPECIFIC RECONFIGURABLE NETWORKCROSS-REFERENCE TO RELATED APPLICATIONSThis application claims the benefit of provisional patent application No. 63 / 753,411, 5 titled “Application-Specific Reconfigurable Al Cluster Network” filed February 3, 2025, and provisional patent application No. 63 / 814,148, titled “Application-Specific Reconfigurable Al Cluster Network” filed May 29, 2025, which are hereby incorporated by reference in their entirety without giving rise to disavowment.TECHNICAL FIELD
[0001] The present disclosure relates to photonic systems in general, and to a method and system for managing communication in data centers for sharing workload, in particular.BACKGROUND15
[0002] Photonics is the physical science of light (photon) generation, detection, and manipulation through emission, transmission, modulation, signal processing, switching, amplification, and sensing.
[0003] Photonic systems are gaining more and more popularity in all areas, such as but not limited to light detection, telecommunications, information processing, photonic 20 computing, lighting, metrology, spectroscopy, holography, medicine (such as surgery, vision correction, endoscopy, health monitoring), biophotonics, military technology, laser material processing, art diagnostics, material processing, art diagnostics involving InfraRed Reflectography Xrays, UltraViolet fluorescence, XRF), agriculture, robotics, and others.
[0004] Some important uses of photonic systems include transmitting and receiving information, multiplexing and demultiplexing information, or the like. Photonic devices may include but are not limited to photo detectors including photo diodes or photo transistors, laser diodes, light-emitting diodes, solar and photovoltaic cells, displays and optical amplifiers. Other examples include devices for modulating a beam of light and for 30 combining and separating beams of light of different wavelength.GA REF.:766-20
[0005] The need for photonic devices arises from the limits and limitations of electronic devices. A first limit relates to the transfer rate of information, and is due to electron speed saturation. A second limitation arises from the high power consumption of electronic devices, and thus the generated heat, the footprint and cost of heat dissipation. The use of 5 photonic devices provides for higher rates, with little heating, thus curing or easing these problems.
[0006] Of special importance is the area of communication in data centers, which may amount to as much as 75% of the total volume of digital communication in the world. Thus, making this communication efficient may have significant effect on multiple processing tasks.BRIEF SUMMARY
[0007] One exemplary embodiment of the disclosed subject matter is to be applied in a communication system comprising a data center comprising one or more first computing platforms and a plurality of Artificial Intelligence (Al) Radio Access Networks (RANs), 15 wherein each of the plurality of Al RANs comprises at least one second computing platform, and is in communication with the data center or with another Al RAN from the plurality of Al RANs, a computing platform comprising a processor adapted to: obtaining information related to one or more applications to be executed within the communication system and acceptable performance parameters thereof; determining according to the 20 information one or more Al RANs from the plurality of Al RANs to participate in execution of the one or more applications; determining operational parameters for a connection between the Al RANs and one or more second Al RANs, or between the one or more Al RANs the one or more first computing platform; and adapting one or more optical transmitter and one or more optical receiver for transmitting and receiving information over a channel connecting the one or more Al RANs to the one or more second Al RANs or one or more Al RANs to the one or more first computing platforms, in accordance with the operational parameters, such that the acceptable performance parameters are achieved. Within the communication system, the performance parameters optionally include at least bandwidth and latency. Within the communication system, the 30 operational parameters optionally comprise time and frequency multiplexing parameters.Within the communication system, the adaptation optionally comprises selecting aGA REF.:766-20plurality of wavelengths to be transmitted over an optical channel. Within the communication system, the adaptation optionally comprises selecting parameters for time multiplexing of wavelengths to be transmitted over an optical channel. Within the communication system, the one or more applications optionally comprise a plurality of 5 applications. Within the communication system, the applications optionally comprise an application scheduled predicted to be executed. Within the communication system, determining the operational parameters optionally comprises: determining whether configuration of one or more communication channels within the communication system needs to be adapted to provide the acceptable performance parameters; and in response to a configuration adaptation being required, calculating a required adaptation of the operational parameters. Within the communication system, determining whether the configuration needs to be adapted optionally also considers improving resource usage. Within the communication system, the connection between the Al RANs and the second Al RANs or the first computing platform is optionally an Ethernet connection.15
[0008] Another exemplary embodiment of the disclosed subject matter is a method to be applied in a communication system comprising a data center comprising one or more first computing platforms and a plurality of Artificial Intelligence (Al) Radio Access Networks (RANs), wherein each of the plurality of Al RANs is in communication with the data center or with another Al RAN from the plurality of Al RANs, the method 20 comprising: obtaining information related to one or more applications to be executed within the communication system and acceptable performance parameters thereof; determining according to the information one or more Al RANs from the plurality of Al RANs to participate in execution of the one or more applications; determining operational parameters for a connection between the one or more Al RANs and one or more second 25 Al RANs, or between the one or more Al RANs and the one or more first computing platforms; and adapting one or more optical transmitter for transmitting information over a channel connecting the one or more Al RANs to the one or more second Al RANs or to the one or more first computing platforms, in accordance with the operational parameters, such that the acceptable performance parameters are achieved. Within the 30 method, the performance parameters optionally include bandwidth and latency. Within the method, the operational parameters optionally comprise time and frequency multiplexing parameters. The method can further comprise using non-retum-to-zeroGA REF.:766-20transmission when transfer rate below a predetermined threshold is acceptable, thereby eliminating errors and reducing latency. Within the method, the adaptation optionally comprises selecting a plurality of wavelengths to be transmitted over an optical channel. Within the method, the adaptation optionally comprises selecting parameters for time 5 multiplexing of wavelengths to be transmitted over an optical channel. Within the method, the one or more applications optionally comprise a plurality of applications. Within the method, the one or more applications optionally comprise an application scheduled or predicted to be executed. Within the method, said determining the operational parameters optionally comprises: determining whether configuration of at 10 least one communication channel within the communication system needs to be adapted to provide the acceptable performance parameters; and in response to a configuration adaptation being required, calculating a required adaptation of the operational parameters. Within the method, determining whether the configuration needs to be adapted optionally also includes a requirement to improve resource usage.15GA REF.:766-20THE BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS
[0010] The present disclosed subject matter will be understood and appreciated more fully from the following detailed description taken in conjunction with the drawings in which corresponding or like numerals or characters indicate corresponding or like 5 components. Unless indicated otherwise, the drawings provide exemplary embodiments or aspects of the disclosure and do not limit the scope of the disclosure. In the drawings:
[0011] Fig. 1 is a schematic illustration of a structure common in communication networks, in accordance with some exemplary embodiments of the disclosure;
[0012] Fig. 2 is a schematic flow chart of steps in a method for adapting a network,; 10
[0013] Fig. 3 is a block diagram of a system for adapting a network, in accordance with some exemplary embodiments of the disclosure;
[0014] Fig. 4 is a diagram of the transmitting side of an optical link, in accordance with some exemplary embodiments of the disclosure; and
[0015] Fig. 5 is a diagram of the receiving side of an optical link, in accordance with 15 some exemplary embodiments of the disclosure.GA REF.:766-20DETAILED DESCRIPTION
[0017] The term “Data Center” is to be widely construed to cover a facility used to house computer systems and associated components, such as telecommunication and storage systems. The computer systems and associated components may be connected 5 therebetween using communication channels operating under any required protocol(s).
[0018] It is estimated that up to 75% of the total digital communication in the world is within data centers, and a lot more communications are to and from the data centers. This includes communication required for sharing workloads, or distributed computing.
[0019] Artificial Intelligence (Al) and Machine Learning (ML) tasks encompass a wide range of tasks and applications that leverage Al and ML techniques to analyze and make predictions upon data. These workloads are at the heart of many modem technological advancements and applications, and are typically communication, storage, and computeintensive. Some of the more common tasks include:
[0020] Supervised Learning is used for training models using labeled datasets, where 15 each input data point is associated with a known target or label.
[0021] Unsupervised Learning involves unlabeled data, where the model learns patterns and structures in the data without explicit guidance.
[0022] Reinforcement Learning involves training agents to make a sequence of decisions in an environment, to maximize a reward signal.
[0023] Deep Learning is a subset of ML that focuses on neural networks with a plurality of layers (deep neural networks - (DNNs)), for example simple feed-forward DNN with 3-10 layers, Convolutional Neural Network (CNN) with 5-100 layers, Modem Large Language Model with 100 to over 200 , or the like.
[0024] Natural Language Processing (NLP) involves processing and understanding 25 human language, enabling machines to interact with text or audio speech data.
[0025] Computer vision relates to understanding and interpreting visual data, such as images or videos, to recognize objects, patterns, and scenes.
[0026] Time series analysis focuses on data that varies over time, involving modeling and predicting future values based on historical data.GA REF.:766-20
[0027] Recommendation systems use AI / ML to suggest items or content to users based on their preferences, behavior, or historical data.
[0028] Generative models aim to generate new data that resembles existing data.
[0029] Anomaly detection focuses on identifying rare or unusual patterns in data that 5 deviate from expected behavior.
[0030] The term “Distributed Al” (DAI) is to be widely construed to cover numerous agents, such as computing platforms or clusters of computing platforms cooperating or competing to solve problems and achieve goals. These agents can function alone or jointly to improve system performance. A significant challenge in DAI relates to how agents can exchange or share knowledge, resources, and duties to solve complicated issues across edge devices.
[0031] The rapid advancement of DAI has revolutionized the development of large- scale generative models, and especially the training of such models. However, achieving optimal performance remains challenging due to the non-uniformity of the participating 15 computing platforms and network latency. Obstacles such as hardware inefficiencies, communication bottlenecks, resource contention, and increased latency caused by geographic dispersing of training data and computing platforms, hinder the scalability and efficiency of distributed systems.
[0032] Some distributed training strategies have evolved to overcome the obstacles and meet the computational demands associated with DAI. These strategies include:
[0033] Hybrid Parallelism: combining various parallelization techniques, such as data parallelism, tensor parallelism, pipeline parallelism, and expert parallelism. Different parallelism strategies impose varying demands on communication channels. For instance, tensor parallelism typically requires high -bandwidth connections, such as NVLink, to 25 facilitate frequent communication. In contrast, pipeline parallelism involves exchanging intermediate tensors only at designated cut points, leading to less frequent communication requirements.
[0034] Auto and Heterogeneous Parallelism: including automation techniques for distributing computations dynamically across diverse hardware.GA REF.:766-20
[0035] Parallelism may support the Single Program Multiple Data (SPMD) programming model, where the same program is executed by multiple processors, each processing different pieces of data. However, some strategies attempt to break the restriction of SPMD and further improve resource utilization with the Multiple Program 5 Multiple Data (MPMD) model, where different programs (or different parts of a program) executed by different processors, handle different parts of the data or model.
[0036] Theoretically, distributing computations across N machines can yield a performance improvement of xN. However, in practice, the actual performance improvement often falls short of this expectation. This decline in efficiency can be attributed to various factors, including identifying stragglers and high-latency issues as significant causes.
[0037] Stragglers are tasks that execute significantly slower than other workers. Their low performance can arise from various causes, including faulty hardware, resource contention on shared infrastructure in data centers, pre-emption by competing jobs, or 15 other reasons.
[0038] Latency refers to the time it takes for data to travel over the network to a destination where it is supposed to be processed or stored. The growing demand for computational power often outpaces its availability, leading to scenarios where training data is performed at locations that are geographically distant from the root aggregator and 20 workers. This geographic separation results in increased latency, requiring reliance on the communication channel's maximum bandwidth, which is inherently limited.
[0039] Additionally or alternatively, optimizations in communication, such as the Ring AllReduce and Double Binary Trees-based algorithms that mitigate bandwidth and latency challenges, may be beneficial in all-to-all data exchanges. The Ring AllReduce takes advantage of bandwidth-optimal communication algorithms without loosening synchronization constraints.
[0040] In order to distribute the workload, for example in training an LLM, efficient training and inference processes can utilize strategic network architecture. Technologies such as bandwidth-optimal algorithms and overlapping communication with computation 30 platforms, may address latency issues while improving scalability.GA REF.:766-20
[0041] It is appreciated that any or all of the above techniques may be combined with the disclosure, to enhance the performance and provide better bandwidth, shorter latency, better power utilization, or other benefits.
[0042] One technical problem associated with the disclosure is the need to optimize the 5 performance of a data center in executing tasks assigned to the data center. IN other words, the problem relates to how to distribute workloads, and in particular a plurality of Al or ML-related tasks, in a data center comprising a plurality of connected computing platforms. Each task may be associated with its own difficulty degree and resource requirements, such as memory, CPU, bandwidth, latency and other resources and constraints. The distribution should comply with the computations’ bandwidth, performance, latency and power consumption, as required for a plurality of tasks, each having its own difficulty degree, volume of communication and other resources, acceptable latency, or the like, and optimize the same such that more processing can be achieved with the same time, communication and processing resources.15
[0043] Referring now to Fig. 1, showing a structure which is currently common in communication networks. The structure, generally referenced 100, comprises a data center 104, which may comprise a plurality of computing platforms such as servers. The computing platforms may be connected therebetween, and may share information, distribute workloads, or the like.20
[0044] The structure may also comprise a plurality of Artificial Intelligence (Al) clusters 108, also referred to as “Al Radio Access Networks” (Al RANs), or “satellite data center”. The Al RANs are typically located close to cellular antennas so as to reduce the latency. Each cluster may comprise one or more computing platforms, which may be at a small geographical distance from each other, for example up to about 100km, and may be interconnected therebetween. Al clusters 108 may be co-located with an antenna or at a proximity such as up to 20km from a cellular antenna, such that the cluster may receive or transmit data from the antenna with minimal delay. Each Al cluster may be connected to and may communicate with one or more data centers and / or one or more other Al clusters.30
[0045] The computing platforms within each Al cluster 108 may cooperate to enable workload sharing therebetween.GA REF.:766-20
[0046] Another technical problem handled by the disclosure relates to the communication protocol(s) used within a data center and / or between the data center and other entities, and the communication protocol used within Al clusters, and the interface between the protocols.5
[0047] In some situations, data centers have traditionally used the Ethernet protocol, while traditional solutions for building high-performance computing (HPC) clusters have relied upon InfiniBand protocol.
[0048] The required gateway between the two systems implies higher latency and higher power consumption, and thus a barrier to performance improvement.
[0049] Yet another technical problem handled by the disclosure relates to characteristics with which the used architecture and infrastructure need to comply, including but not limited to:
[0050] Network scaling: with current Al models growing from billions to one trillion parameters, the volume of data exchanged is so significant that any slowdown due to a 15 poor network can critically impact the Al application performance. Thus, the network needs to be scalable to allow for the ever growing demands.
[0051] Predictable, deterministic latency: Al workload is most dependent on the timely completion of the separate processing steps and their synchronization, to allow for completion of the whole task with minimal stalls.
[0052] Management: integration between existing infrastructures.
[0053] Congestion management: a common “incast” problem occurs in Al network, when multiple senders are sending data to a single receiver, resulting in drastic decrease in throughput, which creates congestion.
[0054] Bandwidth and speeds: it is essential to have higher radix switches and higher 25 port speeds.
[0055] Performance: flow completion time is a critical measure for AL / ML training performance.
[0056] Power: lower power consumption for the overall solution.GA REF.:766-20
[0057] Telemetry - network fabric capabilities to troubleshoot link failures, anomalies, link utilization, traffic monitoring.
[0058] One technical solution of the disclosure comprises distributing the workload of a data center between the computing platforms within the data center and additional 5 smaller clusters, such as the Al clusters or Al RANs described in association with Fig. 1 above.
[0059] Using this scheme, in addition to the workload they may perform some computing resources are located close to cellular antennas, thereby also reducing communication time.
[0060] However, the distribution is non-trivial as static allocation may not always be optimal. There may be peaks of workload that need to be handled differently from quiet hours, certain task may require more resources of certain types than others, or the like. Thus, the distribution needs to be dynamic in accordance with the tasks that need to be performed. Moreover, such distribution may also take into account tasks that are predicted 15 to be executed. It is appreciated that the exact characteristics of the task may not be obtained, but rather its type or general characteristics.
[0061] Thus, a system in accordance with the disclosure may obtain the collection of tasks to be performed at a data center, wherein one or more tasks may be associated with a type or purpose, such as training an Al model for classifying images, training an LLM model, or the like. For each such task, a set of requirements such as processing resources including memory and GPU and / or CPU, required bandwidth, allowed latency, or the like may be calculated or otherwise obtained. It is appreciated that the needs of each such task may change during its execution time, and the set of tasks to be executed may also change dynamically, as tasks are completed and new ones are received.25
[0062] Once the collection of tasks to be executed and the respective requirements are obtained, the availability of the different resources within the data center and at one or more available Al RANs may be determined. It is appreciated that a significant part of the resource availability is expressed by the connectivity between the data center and the Al RANs, within the data center, and among the Al RANs.GA REF.:766-20
[0063] Depending on the requirements and resources, at least one Al RAN may be determined, which may be assigned to execute at least a part of at least one task.
[0064] For example, for a computation -intensive tasks, a distant but powerful Al RAN may be selected, and the communication channel(s) connected to the Al RAN may be 5 prioritized over communication channels connected to other Al RANs. In another example, a short task may be assigned to computing platforms within the data center, to eliminate excess communication, and no precedence may be given to communication links for the purpose of executing this task. In yet another example, a communicationintensive task may be assigned to an Al RAN which is close to an antenna, thereby reducing the communication latency,
[0065] Another technical solution of the disclosure relates to how the communication channels are to be adapted once the at least one Al RAN is determined in accordance with the changing tasks to be executed and their respective priorities, the priorities and the capacities of the data center and / or the Al RANs. The operational parameters of at least 15 one connection between a data center and the Al RAN may be configured accordingly thus creating a dynamic network, which will enable appropriate allocation of the communication channels and better utilization of the computing resources.
[0066] Thus, in some embodiments of the disclosure, the network may be reconfigurable according to the needs, where one or more connections within the data 20 center, and / or between the data center and the Al RANs and / or among the Al RANs may be adapted dynamically. Such adaptation may be performed, for example, when the connections use photonic systems. The communication channels may receive light from an array waveguide adapted to transmit a plurality of wavelengths. By adapting the time and frequency multiplexing, the parameters of each such connection may be adapted, thereby enabling the assignment of different tasks and providing more or less data to one or more computing platforms, according to the required and available resources.
[0067] In accordance with some embodiments of the disclosure, a comb laser or another one or more laser sources may emit light at different wavelengths. An Arrayed Waveguide Grating (AWG) may be used for multiplexing the different wavelengths such 30 that they can be carried over a single optical fiber. Changing the number and time division of the various wavelengths changes the characteristics of the optic fiber, and can thusGA REF.:766-20make some computing platforms receive and process more data, and thereby comply with the required workload distribution.
[0068] Yet another technical solution of the disclosure relates to taking into account not only the current computation needs, but also prediction about future needs, such as 5 expected tasks, such that the allocation of tasks to the computing platforms also adapts to the future needs. Such forward-facing allocation may further improve the performance, for example by eliminating unnecessary communication between platforms, context switches, or the like.
[0069] For example, at the end of a workweek, more heavy-computing tasks may be expected, when programmers or other users use the time to train Al engines. Similarly, more tasks can be expected early in the week, after drawing conclusions from the previously training sessions.
[0070] Yet another technical solution of the disclosure relates to unification of the used network hardware and protocols.15
[0071] When data centers and the Al RANs use different protocols, such as Ethernet in the data centers and InfiniBand in the Al Rans, the gateway bridging therebetween consumes energy and incurs latency. Therefore, by using the same protocol all over the network, including within the data center, between the data center and the Al RANs and inside and among the Al RANS, such gateway is not required, thereby providing for higher efficiency and less used resources.
[0072] In some embodiments, the Ethernet protocol has proven more efficient than others, including InfiniBand. By using the Ethernet protocol in all communication links, the data center and the Al RANs are unified for optimized networking, the workloads may be unified for compute, network and storage, all intermediate means such as 25 gateways are not required, thereby reducing the power consumption and the delays.
[0073] One technical effect of the disclosure provides for higher utilization of computing resources in a data center and associated Al RANs, according to the computing needs and the resources available at every platform, including its location and the distance from an antenna, its GPU and / or CPU and the available memory. Thus, each task fromGA REF.:766-20the tasks that need to be earned out is assigned to one or more computing platforms within the data center or at one or more RANs.
[0074] Another technical effect is the dynamic nature of the assignment, wherein the utilization of the computing platforms may change in accordance with the changing needs 5 of the various tasks, thereby providing an application-specific reconfigurable network.
[0075] These changes are enabled by changing dynamically the characteristics of the optic fibers connecting the various computing platforms, by adapting the number, the wavelength selection and the timing of the light pulses carrying data.
[0076] Yet another technical effect of the disclosure relates to using the same communication protocol within the data center, between the data center and the Al RANs, and among the Al RANs. Specifically, using the Ethernet protocol for all the abovementioned communication segments provides better utilization of the communication links and thus of the whole network.
[0077] Yet another technical effect of the disclosure relates to further improving the 15 performance of the network, by taking into account in the network adaptation not only the current needs, but also prediction about future needs, such as expected tasks or at least the types and volume of expected tasks, such that the allocation of tasks to the computing platforms also adapts to the future needs. Such dynamic but forward-facing allocation may help reduce unnecessary communication, context switch, or the like.20
[0078] All the above effects provide for better network scaling, predictable and deterministic latency, ease of network management, congestion management, bandwidth and speed improvement, performance improvement, reduced power consumption and enabling telemetry.
[0079] Referring now to Fig. 2, showing a flow chart of steps in a method for adapting a network, and to Fig. 3, showing a block diagram of a system for the same, in accordance with some exemplary embodiments of the disclosure.
[0080] At step 204, information related to at least one application to be executed within the data center may be received. The information may be associated with acceptable parameters thereof, such as priority, urgency, amount of data, computing complexity, 30 communication requirements, repetitions, telemetry, or the like. The performanceGA REF.:766-20parameters may be received together with the identification of the application, retrieved from another source such as a database, or dynamically calculated or predicted.
[0081] The information is collected and added to a collection of information items related to other applications executed within the communication system, whether to be 5 initiated, during execution, scheduled or predicted.
[0082] Referring now also to Fig. 3, the system may comprise a network management platform 300. Network management platform 300 may be adapted to manage a network such as network 100 comprising data center 104 and one or more Al RANs 108, distribute workloads, collect performance statistics such that the distribution can be learned and improved over time, make recommendations on preferred locations of Al RANs, or the like.
[0083] Network management platform 300 may be any one or more computing platform within data center 104, or external thereto. In some embodiments, the functionality of network management platform 300 may be distributed among two or more computing 15 platforms, whether either belongs to data center 104 or not.
[0084] Network management platform 300 may comprise a processor 304 which may be one or more Central Processing Units (CPU), a microprocessor, an electronic circuit, an Integrated Circuit (IC) or the like. Processor 304 may be configured to provide the required functionality, for example by loading to memory and activating the modules stored on storage device 316 detailed below.
[0085] Network management platform 300 may comprise one or more Input / Output (I / O) devices 308, which may be display, a touch screen, a speakerphone, a microphone, a headset, a pointing device, a keyboard, or the like. I / O device 308 may be utilized to receive input from and provide output to a user, such as a system administrator, a network 25 manager, or the like. I / O device 308 may be used for instructing about system changes or constraints, failures, override configurations, or the like, and for outputting reports about the system behavior and performance, workloads performed by each computing platform, possible failures, or the like.GA REF.:766-20
[0086] Network management platform 300 may comprise a communication device 312 for communicating with other computing platforms or storage devices, such as one or more computing platforms of data center 104, any of Al RANs 108 or other servers.
[0087] Network management platform 300 may comprise one or more Storage Devices 5 316, such as a hard disk drive, a Flash disk, a Random Access Memory (RAM), a memory chip, or the like. In some exemplary embodiments, Storage Device 316 may retain program code operative to cause processor 304 to perform acts associated with any of the modules listed below, or execute steps of the flowchart of Fig. 2. The program code may comprise one or more executable units, such as functions, libraries, standalone programs or the like, adapted to execute instructions as detailed below. It is appreciated that Storage Device 316 may comprise a plurality of operatively connected storage devices.
[0088] Storage Device 316 may retain Application Layer 320 for obtaining information about all the applications currently being executed or that are scheduled to be executed by the data center and / or associated Al RANs.15
[0089] Application Layer 320 may analyze the information for determining the computing requirements of the applications. Application Layer 320 may also retain information about applications that are currently being executed and their computing and communication utilization such as CPU, memory, communication volume, bandwidth and latency of the involved communication links, or the like.20
[0090] Moreover, Application Layer 320 may also obtain predictions about applications that will be required to execute. Such prediction may be obtained, for example, from an engine trained upon the applications that have been executed in the past, their timing, repeatability, time, computing and communication requirements, the resulting performance, or the like. Such engine may utilize unsupervised learning, or other techniques.
[0091] Thus, Application Layer 320 may have “visibility” of all the newly received, currently existing, scheduled and predicted applications that may require resources from the data center and / or associated Al RANs.
[0092] Application Layer 320 may therefore perform information obtaining step 204 of 30 Fig. 2.GA REF.:766-20
[0093] At step 208, an assignment of a newly received application to a computing platform, may be determined. In some embodiments, the computing platform may be an Al RAN or one or more computing platforms thereof. Step 208 may be performed, for example, by Task and Flow Manager 324 retained by storage device 316 of network 5 management platform 300. Task and Flow Manager 324 may thus be configured to determine the assignment, using the information collected regarding the currently executed, planned or predicted applications, their timetables, and their computing and communication requirements. Task and Flow Manager 324 may further consider the available resources and properties of each site, including their memory, CPU, distance from other computing platforms, bandwidth and latency of each connection, distance from antenna, or the like. In some embodiments, Task and Flow Manager 324 may determine to reassign one or more applications to a different computing platform within the network.
[0094] Determining the Al RAN may take into account the application and its 15 parameters as obtained as step 200: the date and time, for example whether it is during the weekend, night hours or other less busy hours or not, the other applications currently being executed and their respective information, the computing capabilities of each computing platform in the data center, the distance of each Al RN from an antenna which affects the latency, or the like. The determination may take into account the full mapping 20 of the executed, scheduled and predicted applications, such that the overall performance is enhanced. For example, tasks that can withstand longer latency may be assigned to Al RANs that are farther from antennas and from the data center, while other tasks may be executed by computing platforms that are closer to antennas. In another example, applications with heavy computing requirements may be assigned to the data center or to 25 Al RANs that have sufficient computing resources to handle the task, or the like.
[0095] Task and Flow Manager 324 may use any deterministic or heuristic algorithm for determining the assignment for the new application and for planned or predicted ones, and / or redistribute applications that are currently being executed to other computing platforms. In further embodiments, Task and Flow Manager 324 may use one or more Al 30 engines trained on past and / or fabricated schedules, whether using supervised or unsupervised learning.GA REF.:766-20
[0096] The assignment may imply the requirements from the communication channels within the computing network, e.g. the communication links between the data center and the Al RANs, or within any of them.
[0097] At step 212, based on the determined assignment, operational parameters may 5 be determined for one or more connections, between the date center and any of the Al RANs, within the data center, within an Al RAN, or the like, such that the communication channels may be able to transmit all required input and output to any of the processing units.
[0098] Said determination may comprise step 216 for determining whether a change or adaptation is required to the configuration of any of the communication channels.
[0099] If such adaptation is required, then at step 220 the required adaptation may be calculated. For example, the number of wavelengths to be emitted by a comb laser and their frequencies, and their multiplexing parameters may be calculated for one or more communication links.15
[0100] In some embodiments, the operational parameters may be determined also by Task and Flow Manager 324, or by another dedicated component.
[0101] At step 224, the operational parameters as determined may be applied to the communication links. For example, Storage Device 316 may retain Orchestrator 328 configured to transmit corresponding commands to one or more controllers of 20 Reconfigurable Optical Network 332. By the controllers performing the commands, the characteristics of the communication channels may change, thereby enabling the assignment of the different tasks to the corresponding computing platforms in Data Center 104 and Al RANs 108. Thus dynamic utilization of the computing network may be achieved, in accordance with the changing needs, including the executed, planned or predicted tasks, the day, date and time, the available computing resources and parameters of the network. Orchestrator 328 may transmit, and the controllers at the transmitting side and the receiving side of a link may receive and execute commands for changing the operational parameters and thereby the behavior and characteristics of the link.
[0102] Referring now to Fig. 4, showing a diagram of the transmitting side of an optical 30 link, in accordance with some exemplary embodiments of the disclosure.GA REF.:766-20
[0103] The link receives light from Laser Source 400. The light source can be a Passively mode locked laser generating a frequency comb, a Micro ring based laser generating a frequency comb, or any other type of frequency comb source. The frequency comb can comprise a few tens of modes (wavelengths). The power of each mode may depend on 5 the laser structure and can be about few milliwatts. The spacing between the modes may depend on the laser structure and can be 50GHz, 100GHz, 150GHz, 200GHz, or the like. In some embodiments, the wavelengths are selected such that they create minimal interference when transmitted over the same fiber.
[0104] The light may be fed into Power Splitter 404, for splitting the power of Laser Source 400 among the produced light signals.
[0105] The signals may be fed into an Optical Amplifier 408, such as a semiconductor fiber or waveguide based optical amplifier for amplifying the signals.
[0106] In an example, Laser Source 400 may output up to 192 different data channels.
[0107] Each signal group may then be fed into a separate channel. The grouping criteria 15 may be according to the minimum number of modes required to achieve the desired pulse shape, the number of required groups, the number of wavelengths that can be created by the laser source, or the like. All the above criteria, and possibly additional ones may be subject to the link budget, including power and noise. In the example above, there may be up to eight (8) channels, each comprising up to 24 frequencies. Each signal may be 20 provided to an Arrayed Waveguide grating (AWG) 412 for demultiplexing the different wavelengths in the group, each wavelength may then be adapted to carry optical information converted from electric information. For example, a first wavelength carries Datall 420, and the 24th wavelength carries Datal24 422 (and similarly for all frequencies between 2 and 23).
[0108] The 24 signals carrying the data may then be fed into AWG 424 for multiplexing the signals into a common signal thereby providing frequency multiplexing, and to Phase Reshaping 428 for reshaping the time pulses before interleaving the signals, and using empty slots between pulses of the same wavelengths such as 432 and 436 for transmitting pulses of other wavelengths, thereby providing for time multiplexing. For example, pulses 30 432 of IpSec width with 20pSec repetition rate enable the transmission of a plurality of pulses of other wavelengths.GA REF.:766-20
[0109] An analogous route may be taken by the other signal groups, transmitting additional data such as Data8i 440...Data824444. It is appreciated that at least some of the multiplexed signals need to be delayed, for example by Delay Line 448, in order to adjust their timing relative to the other signals, thereby enabling further time multiplexing 5
[0110] The signals may then be summed using Adder 452, thereby providing a signal to be transmitted over optic fiber 456, the signal having a data rate of up to 8*24=192 from each original signal such as 420, 422, 440 and 444.
[0111] This exemplary arrangement enables to dynamically vary the data rates, modulation schemes with a factor of up to 192 over the same fiber, such that each link may be utilized with a high or low data rate as required, in accordance with the computing platforms intended to be used for the collection of tasks to be executed by the network. For example, if the volume of data to be transmitted is low, low data transfer rate may be acceptable, for example below a threshold value such as for example 26Gbps, non-retum- to-zero transmission may be used for improving the signal quality. The signal can then be 15 transmitted error-free, thereby eliminating the need to forward error correction, and significantly reducing the latency. In some applications high data rate may be required, and another configuration may be applied. The reconfigurable network thus allows for optimal assignment of the data channels to the application requirements.
[0112] Referring now to Fig. 5, showing a diagram of the receiving side of an optical 20 link, in accordance with some exemplary embodiments of the disclosure.
[0113] Signal 456 may be fed into Optical Passive (OSP) Demultiplexer 504, for multiplexing the signals according to their timing, up to 8 channels in accordance with the number of channels used by the transmitting side. Thus, OSP 504 may output Channeh 508, Channels 510 and optionally additional Channels2-7.
[0114] Channeh 508 may then be fed into OSP 512 for frequency domain demultiplexing, to split the channel into up to 24 data streams, such as Data1! 516... DataX24520.
[0115] An analogous route may be taken by up to seven more channel, such as Channels 510, to output from each channel up to corresponding 42 data streams such as Data8i 30 524... Data824 528.GA REF.:766-20
[0116] All output data channels may be provided to OZE converters (not shown) for retrieving the electronic data.
[0117] The exemplary transmitter and receiver thus provide for time and frequency multiplexing and demultiplexing, thereby enabling significantly higher data rate over a 5 single channel, and thereby dynamically varying the data rates. This in turn provides for dynamic usage of computing platforms to support varying computing requirements by the same existing platforms.
[0100] Implementation details of the optical transmission may be found, for example, in US Patent No. 12,034,527 titled “System for Pulsed Laser Optical Data Transmission with Receiver Recoverable Clock”, filed on January 5, 2022, assigned to the same assignee as the current application, and incorporated herein by reference in its entirety and for all purposes.
[0118] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer readable storage 15 medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present invention.
[0119] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, 20 a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non- exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be 30 construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through aGA REF.:766-20waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0120] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to 5 an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0121] Computer readable program instructions for carrying out operations of the present invention may be assembler instructions, instruction-set-architecture (ISA) 15 instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, such as "C", C#, C++, Java, Phyton, Smalltalk, or others. The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, 20 partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field- programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present invention.30
[0122] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computerGA REF.:766-20program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.5
[0123] These computer readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the 15 function / act specified in the flowchart and / or block diagram block or blocks.
[0124] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which 20 execute on the computer, other programmable apparatus, or other device implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0125] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially 30 concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagramsGA REF.:766-20and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.5
[0126] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0127] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements in the claims below are intended to include any structure, 15 material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention. The embodiment was chosen and described in order to best explain the principles of the invention and the practical application, and to enable others of ordinary skill in the art to understand the invention for various embodiments with various modifications as are suited to the particular use contemplated.25
Claims
GA REF.:766-20CLAIMSWhat is claimed is:
1. In a communication system comprising a data center comprising at least one first computing platform and a plurality of Artificial Intelligence (Al) Radio Access 5 Networks (RANs), wherein each of the plurality of Al RANs comprises at least one second computing platform, and is in communication with the data center or with another Al RAN from the plurality of Al RANs, a computing platform comprising a processor adapted to:obtaining information related to at least one application to be executed within the communication system and acceptable performance parameters thereof;determining according to the information at least one Al RAN from the plurality of Al RANs to participate in execution of the at least one application;determining operational parameters for a connection between the at least 15 one Al RAN and at least one second Al RAN, or between the at least one Al RAN the at least one first computing platform; andadapting at least one optical transmitter and at least one optical receiver for transmitting and receiving information over a channel connecting the at least one Al RAN to the at least one second Al RAN or to the at least one first 20 computing platform, in accordance with the operational parameters, such that the acceptable performance parameters are achieved.
2. The communication system of Claim 1, wherein the performance parameters include at least bandwidth and latency.
3. The communication system of Claim 1, wherein the operational parameters comprise time and frequency multiplexing parameters.
4. The communication system of Claim 1, wherein the adaptation comprises selecting a plurality of wavelengths to be transmitted over an optical channel.
5. The communication system of Claim 1, wherein the adaptation comprises selecting parameters for time multiplexing of wavelengths to be transmitted over an optical 30 channel.GA REF.:766-206. The communication system of Claim 1, wherein the at least one application comprises a plurality of applications.
7. The communication system of Claim 6, wherein the at least one application comprises an application scheduled predicted to be executed.5 8. The communication system of Claim 1, wherein said determining the operational parameters comprises:determining whether configuration of at least one communication channel within the communication system needs to be adapted to provide the acceptable performance parameters; andin response to a configuration adaptation being required, calculating a required adaptation of the operational parameters.
9. The communication system of Claim 8, wherein determining whether the configuration needs to be adapted also considers improving resource usage.
10. The communication system of Claim 1 , wherein the connection between the at least 15 one Al RAN and the at least one second Al RAN or the at least one first computing platform is an Ethernet connection.
11. A method to be applied in a communication system comprising a data center comprising at least one first computing platform and a plurality of Artificial 20 Intelligence (Al) Radio Access Networks (RANs), wherein each of the plurality of Al RANs is in communication with the data center or with another Al RAN from the plurality of Al RANs, the method comprising:obtaining information related to at least one application to be executed within the communication system and acceptable performance parameters thereof;determining according to the information at least one Al RAN from the plurality of Al RANs to participate in execution of the at least one application;determining operational parameters for a connection between the at least one Al RAN and at least one second Al RAN, or between the at least one Al 30 RAN and the at least one first computing platform; andGA REF.:766-20adapting at least one optical transmitter for transmitting information over a channel connecting the at least one Al RAN to the at least one second Al RAN or to the at least one first computing platform, in accordance with the operational parameters, such that the acceptable performance parameters are achieved. 5 12. The method of Claim 11, wherein the performance parameters include at least bandwidth and latency.
13. The method of Claim 11, wherein the operational parameters comprise time and frequency multiplexing parameters.
14. The method of Claim 11, further comprising using non-retum-to-zero transmission when transfer rate below a predetermined threshold is acceptable, thereby eliminating errors and reducing latency.
15. The method of Claim 11, wherein the adaptation comprises selecting a plurality of wavelengths to be transmitted over an optical channel.
16. The method of Claim 11, wherein the adaptation comprises selecting parameters 15 for time multiplexing of wavelengths to be transmitted over an optical channel.
17. The method of Claim 11, wherein the at least one application comprises a plurality of applications.
18. The method of Claim 17, wherein the at least one application comprises an application scheduled or predicted to be executed.
19. The method of Claim 11, wherein said determining the operational parameters comprises:determining whether configuration of at least one communication channel within the communication system needs to be adapted to provide the acceptable performance parameters; and25 in response to a configuration adaptation being required, calculating a required adaptation of the operational parameters.
20. The method of Claim 19, wherein determining whether the configuration needs to be adapted also includes a requirement to improve resource usage.