Computing device computing power optical interconnection interface construction method and computing device computing power optical interconnection interface system
By constructing optical interconnect communication channels and direct interconnect links for computing power, the problems of signal attenuation and delay of electrical interfaces were solved, realizing efficient and low-latency direct connection of computing units and improving the efficiency of distributed computing power collaborative inference.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UMU TECH CO LTD
- Filing Date
- 2026-04-30
- Publication Date
- 2026-05-29
AI Technical Summary
Existing interface protocols based on electrical signals suffer from signal attenuation and crosstalk issues under high-frequency transmission, making it difficult to meet the high bandwidth requirements between modern computing units. Furthermore, the communication latency is relatively high, making it impossible to achieve real-time direct signal connection between computing units, which limits the efficiency of distributed computing power collaborative inference.
By constructing a computing power optical interconnection communication channel and establishing a direct interconnection link, the hardware inference unit of the second device is registered as the remote inference engine of the first device. By utilizing the low-loss characteristics of optical signals, the direct mapping and mapping of the original bus signals are realized, forming a deeply integrated 'physical connection-logical registration-dynamic scheduling' mechanism.
It significantly improves the transmission bandwidth between devices, reduces communication latency, and achieves efficient aggregation of cross-device resources, meeting the high-performance interconnection needs of complex AI application scenarios.
Smart Images

Figure CN122111922A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer hardware interface and optical communication technology, and more specifically, to a method for constructing a computing power optical interconnect interface for computing devices and a computing power optical interconnect interface system for computing devices. Background Technology
[0002] With the rapid development of artificial intelligence technology, consumer devices such as laptops and high-performance new energy vehicles are increasingly demanding collaborative inference among computing units (CPU / NPU / GPU). In multi-device computing power aggregation scenarios, devices need to frequently distribute large-scale tensor data and send back inference results, which requires interconnect interfaces to have extremely high bandwidth and extremely low communication latency.
[0003] Existing consumer device interconnection solutions typically employ electrical signal-based interface protocols, such as Thunderbolt 4, USB4, or PCIe. This approach first connects the electrical interfaces of the two devices via physical cables; then, a data transmission link is established through a handshake protocol; finally, data exchange occurs at the operating system layer by simulating network devices or peripheral buses.
[0004] However, this electrical interface-based solution has significant technical drawbacks. Electrical signals are highly susceptible to severe signal attenuation and crosstalk during high-frequency transmission, limiting further increases in single-channel bandwidth and making it difficult to meet the hundreds of GB / s memory bandwidth requirements of modern NPUs. Furthermore, traditional interfaces suffer from high communication latency between devices due to significant protocol stack overhead during cross-device communication, hindering real-time direct signal connection at the computing unit level and severely restricting end-to-end latency performance in distributed computing collaborative inference. Summary of the Invention
[0005] This application provides a method for constructing a computing power optical interconnect interface for computing devices and a computing power optical interconnect interface system for computing devices, so as to at least alleviate the above-mentioned technical problems.
[0006] A method for constructing a computing power optical interconnect interface for computing devices includes: Step 1: Construct a computing power optical interconnection communication channel containing link parameter information between the first device and the second device; Step 2: Establish a direct interconnection link between the first device and the second device based on the computing power optical interconnection communication channel; Step 3: Based on the direct interconnection link, the first device registers the hardware inference unit of the second device as a locally schedulable remote inference engine.
[0007] Optionally, step 1 includes: a first device sensing the physical access of a second device through a CCOI connector; in response to sensing the physical access of the second device, triggering the establishment of an optical physical layer link between the first device and the second device; the first device obtaining the link parameter information of the second device through the optical physical layer link; and the first device establishing the computing power optical interconnection communication channel based on the link parameter information.
[0008] Optionally, the link parameter information includes single-channel rate, number of parallel channels, operating wavelength, and module power consumption level.
[0009] Optionally, step 2 includes: the first device maps the original bus signal of its internal computing unit to the computing power optical interconnection communication channel, so as to encapsulate the original bus signal into a computing power optical interconnection protocol data frame in accordance with the CCOI transport layer protocol; the first device sends the computing power optical interconnection protocol data frame to the second device through the computing power optical interconnection communication channel, so as to establish the direct interconnection link between the first device and the second device.
[0010] Optionally, the original bus signals often include memory bus extension signals, NPU inference data stream signals, and PCIe extension signals.
[0011] Optionally, step 3 includes: the first device extracting the memory bus extension signal from the computing power optical interconnection protocol data frame, mapping the memory space of the second device with which it has established the direct interconnection link to a remote memory node of a local non-unified memory access architecture (NUMA) based on the memory bus extension signal; and registering the hardware inference unit of the second device as the remote inference engine that is locally schedulable by the first device based on the remote memory node of the NUMA.
[0012] Optionally, step 3 includes: the first device identifying all second devices with which it has established the direct interconnect link; forming a remote memory node queue based on the remote memory nodes corresponding to the local non-unified memory access architecture (NUMA) of all second devices, so that after registering the hardware inference unit of the second device as the remote inference engine that is locally schedulable by the first device, all second devices corresponding to all the remote inference engines that are registered locally by the first device form a cross-device high-speed computing power aggregation resource pool.
[0013] Optionally, it further includes: the first device calling its MDCCP protocol layer to sense the link status of the direct interconnect link in real time, so as to determine the computing power scheduling mode of the first device for the second device based on the link status, and in the computing power scheduling mode, distributing the computing tasks to be processed by the first device to the cross-device high-speed computing power aggregation resource pool for consumption through the computing power optical interconnection communication channel.
[0014] Optionally, during the process of mapping the memory space of the second device to a remote memory node of the Local Non-Unified Memory Access Architecture (NUMA), the first device maintains memory address consistency between the first device and the second device through the CCOI transport layer protocol. This memory address consistency enables physical addressing of the memory space of the second device when the first device accesses the remote memory node.
[0015] Optionally, when distributing the computing tasks to be processed based on the computing power scheduling mode, if the computing power scheduling mode is a high-speed optical interconnect mode, the first device will distribute the heavy tasks in the computing tasks to be processed to the cross-device high-speed computing power aggregation resource pool to match the second device through the computing power optical interconnect communication channel, so as to build a pipelined collaborative inference cluster.
[0016] Optionally, it further includes: the first device calling its MDCCP protocol layer to construct a wireless charging channel, and obtaining wireless charging status parameters of the charging device based on the constructed wireless charging channel through a configured wireless control unit, the wireless charging status parameters including real-time input power, coil alignment quality, and charging area status; the first device calculating a real-time available computing power budget based on the wireless charging status parameters, establishing a scheduling permission level based on the real-time available computing power budget, and then implementing dynamic constraints on the dispatch scale of the computing tasks to be processed based on the scheduling permission level.
[0017] Optionally, the first device and the charging device transmit a computing power access handshake command through the in-band communication channel of the wireless control unit to trigger the acquisition of the wireless charging status parameters of the charging device based on the constructed wireless charging channel.
[0018] Optionally, the computing power access handshake instruction is encapsulated in an extended field of a wireless charging standard message frame, the extended field including message type, computing power level, expected dwell time, and task priority mask.
[0019] Optionally, the method further includes: the first device generating a computing power reservation request, splitting the computing task to be processed into node-level microtasks based on the computing power reservation request, and dividing the computing task to be processed into multiple execution windows corresponding to the charging coverage area; when implementing dynamic constraints on the dispatch scale of the computing task to be processed based on the scheduling permission level, executing the node-level microtasks segment by segment through the multiple execution windows, wherein the charging coverage area consists of multiple charging devices that have established a wireless charging channel with the first device.
[0020] A computing device optical interconnect interface system includes: a first device and a second device, wherein a computing power optical interconnect communication channel is constructed between the first device and the second device; a direct interconnect link is established between the first device and the second device based on the computing power optical interconnect communication channel; and the first device registers the hardware inference unit of the second device as a locally schedulable remote inference engine based on the direct interconnect link.
[0021] Technical advantages of the technical solution provided in this application The computing power optical interconnect interface construction method provided in this application addresses the technical shortcomings of traditional electrical interface interconnection schemes, such as large signal attenuation, limited bandwidth, high communication latency, and difficulty in supporting direct connections between computing units. By constructing a computing power optical interconnection communication channel containing link parameter information and establishing a direct interconnection link, it solves the bandwidth bottleneck problem caused by the physical characteristics of electrical interfaces in traditional schemes. Compared to the traditional scheme's reliance on copper cable electrical signal transmission, this application uses optical physical layer links for high-speed signal interconnection. Leveraging the low-loss characteristics of optical signals at high frequencies, it significantly improves the effective transmission bandwidth between devices, providing a solid high-speed physical foundation for the real-time distribution of large-scale computing power data.
[0022] Based on a direct interconnect link, this application registers the hardware inference unit of the second device as a locally schedulable remote inference engine of the first device, solving the problems of high overhead and poor real-time performance in cross-device resource call protocols in traditional solutions. Traditional solutions typically require complex multi-layer network protocol encapsulation, while this application achieves direct mapping of raw bus signals through an optical interconnect channel, enabling remote hardware resources to access local resources in a manner similar to local extended resources. Compared to the virtualization call mode of traditional solutions, this application significantly reduces the logical latency of computation task distribution and feedback, achieves lower communication clock skew, and thus significantly improves the efficiency of cross-device collaborative inference.
[0023] Furthermore, by mapping remote memory space to local NUMA remote memory nodes and sensing link status in real time at the MDCCP protocol layer to adjust the computing power scheduling mode, a deeply integrated "physical connection-logical registration-dynamic scheduling" mechanism is formed. This solves the problem in traditional solutions that cannot dynamically optimize scheduling strategies based on the characteristics of the interconnect medium, enabling heavy computing tasks to be preferentially matched with high-speed optical interconnect channels for pipelined processing. Thus, while ensuring memory consistency, it achieves efficient aggregation of cross-device computing power resource pools, better meeting the high-performance interconnection needs of consumer devices in complex AI application scenarios. Attached Figure Description
[0024] Figure 1 This is a flowchart illustrating a method for constructing a computing power optical interconnect interface for a computing device, according to an embodiment of this application.
[0025] Figure 2 This is a schematic diagram of a computing device optical interconnect interface system according to an embodiment of this application. Detailed Implementation
[0026] like Figure 1 As shown, this application provides a method for constructing a computing power optical interconnect interface for a computing device, which includes the following steps: Step 1: Construct a computing power optical interconnection communication channel containing link parameter information between the first device and the second device; Step 2: Establish a direct interconnection link between the first device and the second device based on the computing power optical interconnection communication channel; Step 3: Based on the direct interconnection link, the first device registers the hardware inference unit of the second device as a locally schedulable remote inference engine.
[0027] Optionally, step 1 includes: a first device sensing the physical access of a second device through a CCOI connector; in response to sensing the second device in physical access, triggering the establishment of an optical physical layer link between the first device and the second device; the first device obtaining the link parameter information of the second device through the optical physical layer link; and the first device establishing the computing power optical interconnection communication channel based on the link parameter information.
[0028] Preferably, in the specific technical implementation of step 1, the first device contacts the second device through a consumer-grade computing optical interconnect interface connector with blind-mating characteristics. This consumer-grade computing optical interconnect interface connector is adapted to the side of the portable computing device and is physically compatible with the universal Universal Serial Bus (USB-C) connector specification, supporting adaptive switching between optical signal mode and electrical signal mode within the same physical interface. When the second device is physically connected, the first device extracts the level offset characteristics of the interface circuit through the sensing pins inside the consumer-grade computing optical interconnect interface connector, thereby determining the physical connection sensing result. This physical connection sensing result serves as the starting signal to trigger the underlying hardware power-on and activate microcode execution, initiating the subsequent photoelectric conversion preparation process.
[0029] Preferably, after the first device responds to the physical access sensing result, it activates a miniaturized optical transceiver module integrated within the main chip package. This miniaturized optical transceiver module can be integrated with the main processor in a chip co-package form, or it can be deployed as an external miniaturized module (such as an M.2 interface module with a size no larger than 2280), and its operating power consumption is strictly constrained to a preset low power consumption threshold (such as 5 watts). After receiving the physical access sensing result, the miniaturized optical transceiver module attempts to handshake and establish an optical physical layer link between the first and second devices by emitting a laser carrier with a wavelength in the near-infrared band (such as 850 nm or 1310 nm). This process achieves synchronization of the original physical signal through the optical channel in the consumer-grade computing optical interconnect interface connector, providing a low-loss, interference-resistant physical carrier medium for subsequent data interaction.
[0030] Preferably, the process of obtaining link parameter information in step 1 is achieved through a handshake using the established optical physical layer link protocol. After the first device synchronizes the optical physical layer link, it sends a link capability probe frame to the second device via the optical physical layer link and receives a link capability description message from the second device. The first device performs message parsing processing on the link capability description message to extract the hardware interconnection characteristics of the second device. These characteristics together constitute the link parameter information, which specifically includes, but is not limited to, single-channel transmission rate (e.g., not less than 100Gbps), number of parallel channels (e.g., 4 parallel channels), operating wavelength, and module power consumption level. By extracting this link parameter information, the first device can grasp the physical transmission limits of the second device, thereby providing data support for the fine-grained configuration of logical channels.
[0031] Preferably, based on the extracted link parameter information, the first device divides the corresponding logical time slots and frequency bands on the optical physical layer link through link management logic, thereby establishing a computing power optical interconnection communication channel. During this process, the first device calculates the total bandwidth limit (e.g., not less than 400Gbps) according to the number of parallel channels and the single-channel transmission rate in the link parameter information, and configures the error correction coding and flow control strategy of the transport layer based on the calculation results. The established computing power optical interconnection communication channel enables transparent transmission of key signals of the computing unit on the optical path. This processing method can convert parallel signals, which are originally limited by the electrical characteristics of wiring length and signal attenuation, into high-speed serial optical signals, thus fundamentally solving the bandwidth bottleneck when computing units are directly connected between the first and second devices.
[0032] Preferably, after the optical interconnection communication channel for computing power is established, the first device notifies the upper-level computing power aggregation protocol layer to dynamically switch the interconnection mode. Specifically, when the optical interconnection communication channel for computing power is sensed to be in a ready state, the upper-level computing power aggregation protocol layer switches the communication mode from a low-speed wireless compatibility mode (such as wireless LAN or Bluetooth transmission mode) to a high-speed optical interconnection mode. In the high-speed optical interconnection mode, the first device updates the bandwidth weight table and latency parameter table of the internal scheduler contained in the computing power aggregation scheduling layer of the upper-level computing power aggregation protocol layer to ensure that it can carry large-scale tensor data distribution tasks. This mode switching process is achieved by redirecting the protocol stack transmission path, enabling computing tasks to directly utilize the high bandwidth characteristics of the optical interconnection communication channel for computing power without going through cumbersome network protocol encapsulation, thereby reducing end-to-end scheduling latency.
[0033] Preferably, to address stability requirements in mobile scenarios, step 1, during the establishment of the optical interconnect communication channel, also incorporates hardening processing based on sensed alignment quality characteristics. Specifically, the first device monitors the coupling strength of the optical signal in real time through the direct connection protocol layer of the computing unit, converts the coupling strength into alignment quality characteristics, and then maps these characteristics to a normalized alignment score. If the normalized alignment score is high, the current single-channel transmission rate is maintained; if the normalized alignment score decreases due to external disturbances, the connectivity of the optical interconnect communication channel is maintained by dynamically adjusting the modulation format. This feedback adjustment based on physical connection quality ensures that the optical interconnect communication channel can still provide reliable computing power transmission support in complex consumer application scenarios (such as device movement or loose connectors), thus providing a stable communication foundation for subsequently registering the hardware resources of the second device as local extended resources.
[0034] Optionally, the link parameter information includes single-channel rate, number of parallel channels, operating wavelength, and module power consumption level.
[0035] Preferably, in the specific technical implementation of acquiring link parameter information, feature extraction is performed on the underlying handshake protocol messages of the optical physical layer link feedback to determine the set of key physical dimension indicators reflecting the underlying hardware capabilities of the miniaturized optical transceiver module. First, the single-channel transmission rate reflecting the single carrier transmission capability (e.g., not less than 100 gigabits per second) is extracted, and the number of parallel channels (e.g., 4 parallel channels) is determined in combination with the spatial multiplexing dimension of the physical channels. Then, the product relationship between the single-channel transmission rate and the number of parallel channels is mapped to the theoretical peak bandwidth corresponding to the consumer-grade computing optical interconnect interface connector (e.g., not less than 400 gigabits per second), providing a stable throughput expectation for the high-speed distribution of large-scale tensor data. Second, the carrier frequency characteristics used in the optical signal modulation are identified, i.e., the operating wavelength is obtained (e.g., 850 nm or 1310 nm in the near-infrared band is selected), to guide the optical transmission control unit to implement targeted dispersion compensation and gain adjustment for the associated optical fiber medium, reducing signal attenuation in long-distance transmission. Finally, considering the sensitivity of consumer-grade portable devices to battery life, the power consumption level of the miniaturized optical transceiver module under full-speed operation (e.g., set to no more than 5 watts) is extracted and transmitted back to the computing power aggregation scheduling layer in the upper-level computing power aggregation protocol layer as an input constraint for calculating the real-time available computing power power budget. The aforementioned link parameter information is compared and correlated with a preset link configuration table to jointly participate in the logical time slot allocation and dynamic hardening of the flow control protocol for the computing power optical interconnect communication channel. This ensures a clear functional support relationship between the physical layer transmission capability of the optical physical layer link and the upper-level computing power scheduling logic of the upper-level computing power aggregation protocol layer.
[0036] Optionally, step 2 includes: the first device maps the original bus signal of its internal computing unit to the computing power optical interconnection communication channel, so as to encapsulate the original bus signal into a computing power optical interconnection protocol data frame in accordance with the CCOI transport layer protocol; the first device sends the computing power optical interconnection protocol data frame to the second device through the computing power optical interconnection communication channel, so as to establish the direct interconnection link between the first device and the second device.
[0037] Preferably, in the specific technical implementation of step 2, the raw bus signals generated by the computing unit inside the first device are monitored and captured in real time. These raw bus signals involve multiple dimensions of underlying physical signals, including but not limited to memory bus extension signals, neural network processing unit inference data stream signals, and high-speed peripheral component interconnect extension signals. For the captured raw bus signals, a signal remapping mechanism is used to convert their physical level characteristics into logical representation data adapted to the optical interconnect architecture. Through this remapping process, the discretely distributed parallel bus data is transformed into a serial logic stream that meets time-division multiplexing requirements, serving as the raw input for subsequent encapsulation processing.
[0038] Preferably, for the logical representation data obtained from the remapping process, the designed consumer-grade computing optical interconnect interface transport layer protocol is invoked to perform an encapsulation operation. During encapsulation, the logical representation data is divided into fixed-length data payload blocks, and a protocol-defined message header is appended to each data payload block. The message header contains channel index information to identify the signal type, thereby ensuring that memory bus extension signals from different sources, neural network processing unit inference data stream signals, and high-speed peripheral component interconnect extension signals can achieve logical isolation and ordered multiplexing within the same optical physical link. This encapsulation process converts the original bus signals into a self-describing message structure, forming a preliminary computing optical interconnect protocol data frame.
[0039] Preferably, during the formation of the computing power optical interconnection protocol data frame, a check field is further added to the message structure. Specifically, a cyclic redundancy check (CRC) mechanism is used to calculate the check bits in the message body, and the calculated check result is filled into the end of the message, thereby constructing a complete protocol frame with error correction capabilities. Furthermore, the extended fields of the computing power optical interconnection protocol data frame also integrate metadata information such as computing power level identifier, task priority mask, and timestamp. By constructing such a highly complete computing power optical interconnection protocol data frame, technical support can be provided for data consistency maintenance and traffic priority management during subsequent high-speed optical transmission.
[0040] Preferably, the constructed computing power optical interconnection protocol data frame is injected into the previously established computing power optical interconnection communication channel. During the injection process, a miniaturized optical transceiver module integrated within the system-on-a-chip (SoC) inside the main chip package is driven to convert the electrical domain computing power optical interconnection protocol data frame into a near-infrared optical pulse signal. In this process, the signal is modulated using the single-channel transmission rate obtained in step 1 (e.g., not less than 100 gigabits per second) to efficiently load the computing power optical interconnection protocol data frame onto the optical carrier. This method of converting electrical domain protocol frames into optical pulse signals utilizes the extremely high frequency characteristics of the optical carrier, significantly improving the effective bandwidth of the original bus signal during cross-device transmission and effectively alleviating the rate bottleneck caused by signal attenuation limitations in traditional electrical interfaces.
[0041] Preferably, the first device sends a computing power optical interconnection protocol data frame, modulated as an optical pulse signal, to the second device via a computing power optical interconnection communication channel. Upon receiving the computing power optical interconnection protocol data frame, the second device performs photoelectric reconstruction and protocol decapsulation processing through its associated miniaturized optical transceiver module and sends back a computing power access confirmation message. Upon receiving the computing power access confirmation message, the first device determines that the underlying handshake between the first and second devices is complete, thereby formally establishing a direct interconnection link between the first and second devices. Because this direct interconnection link is based on direct signal mapping at the physical layer rather than traditional network protocol stack forwarding, it can achieve an access latency of less than one hundred nanoseconds. Establishing this extremely low-latency direct interconnection link provides a high-performance physical connection foundation for the subsequent dynamic registration of the second device's hardware resources as local extended resources.
[0042] Preferably, through the establishment of the aforementioned direct interconnect link, a direct interconnection at the signal level between the computing units is achieved between the first device and the second device. During the operation of the direct interconnect link, the consumer-grade computing optical interconnect interface transport layer protocol is used to maintain memory address consistency. This technical implementation allows the first device to perform physical addressing through the direct interconnect link when accessing the memory space of the second device, making the remote memory appear to the first device's operating system as a remote memory node with a non-uniform memory access architecture. By transforming complex network protocol stack operations into efficient remote bus access operations, the protocol load overhead during cross-device communication is significantly reduced, thereby improving the computing power aggregation efficiency of the neural network processing unit in multi-device collaborative inference scenarios while ensuring real-time data transmission.
[0043] Optionally, the original bus signals include memory bus extension signals, NPU inference data stream signals, and PCIe extension signals.
[0044] Preferably, in the specific technical implementation of the original bus signal, the physical level generated by the computing unit inside the first device is sensed and logically sampled in real time to extract the heterogeneous bus data stream that reflects the underlying hardware operating status. Specifically, the memory bus extension signal maps the memory controller of the second device to a remote memory node of the first device's non-uniform memory access architecture via a consumer-grade computing optical interconnect interface connector. Utilizing the low-loss characteristics provided by the optical physical layer link, memory access latency is controlled at a low level (e.g., below 100 nanoseconds), thereby enabling direct addressing and spatial aggregation of memory resources between the first and second devices. The neural network processing unit inference data stream signal carries the input tensors and output activation values of the neural network model during collaborative execution in real time. Combined with the high bandwidth carrying capacity of the computing power optical interconnect communication channel (e.g., a single-channel transmission rate of not less than 100 gigabits per second), it supports the neural network processing units between the first and second devices to execute pipelined collaborative inference according to a preset computing task distribution strategy. The high-speed peripheral component interconnect extension signal is used to transmit data packets conforming to high-speed bus protocol standards (such as the fifth-generation high-speed peripheral component interconnect extension signal), ensuring that the external computing power module can interact with the computing unit inside the first device at a communication rate close to that of the local bus. The aforementioned memory bus extension signal, neural network processing unit inference data stream signal, and high-speed peripheral component interconnect extension signal are converted into logical representation data through a signal remapping mechanism. These are then encapsulated into a self-describing message structure according to the consumer-grade computing optical interconnect interface transport layer protocol, forming a computing power optical interconnect protocol data frame that reflects the original bus characteristics. This physically solves the technical bottleneck of the inability to achieve direct connection of computing units between the first and second devices due to the physical characteristics of the electrical interface.
[0045] Optionally, step 3 includes: the first device extracting the memory bus extension signal from the computing power optical interconnection protocol data frame, mapping the memory space of the second device with which it has established the direct interconnection link to a remote memory node of a local non-unified memory access architecture (NUMA) based on the memory bus extension signal; and registering the hardware inference unit of the second device as the remote inference engine that is locally schedulable by the first device based on the remote memory node of the NUMA.
[0046] Preferably, in the specific technical implementation of step 3, the first device performs decomposition processing on the received computing power optical interconnection protocol data frame through its internally integrated protocol parsing engine. Specifically, the protocol parsing engine identifies the channel index information in the header of the computing power optical interconnection protocol data frame, and logically demultiplexes the payload content in the computing power optical interconnection protocol data frame according to the channel index information, thereby accurately extracting the memory bus extension signal. This extraction process utilizes the channel isolation mechanism defined in the consumer-grade computing optical interconnection interface transport layer protocol to ensure that the original electrical signal representation reflecting the memory state can be independently restored from the mixed data stream. By sending the extracted memory bus extension signal into the memory scheduling logic unit inside the first device, the original physical addressing information is provided for the subsequent establishment of a unified addressing space across devices between the first device and the second device.
[0047] Preferably, in one scenario, during step 3 of memory space mapping, the first device initiates a hardware-level self-discovery process for the memory space of the second device based on the extracted memory bus extension signal. The memory management unit inside the first device identifies the remote memory base address and capacity descriptor carried in the memory bus extension signal, and maps the corresponding physical address range to the memory topology within the first device, thereby defining the memory space of the second device as a remote memory node in a local non-uniform memory access architecture. Through this hardware-level address mapping method, the central processing unit or neural network processing unit of the first device can access the memory space of the second device via a direct interconnect link, just like accessing a local memory slot. This mapping process avoids the software overhead of traditional network protocol stacks, enabling the access latency of remote memory nodes in a local non-uniform memory access architecture to be maintained at a low level (e.g., less than one hundred nanoseconds).
[0048] Preferably, in step 3, the technical processing for maintaining memory address consistency utilizes the cache consistency protocol stack defined by the consumer-grade computing optical interconnect interface transport layer protocol to synchronize the memory state between the first device and the second device in real time. The first device uses consistency maintenance messages sent via the direct interconnect link to verify the validity of data copies in remote memory nodes within the local non-unified memory access architecture. When the first device needs to access a remote memory node within the local non-unified memory access architecture, it performs physical addressing via the direct interconnect link to directly operate the physical memory units of the second device, achieving cross-device memory semantic transparency. This consistency guarantee mechanism based on physical addressing ensures that in multi-device collaborative inference scenarios, the flow of tensor data between the memory of the first and second devices does not require frequent software transfers, thereby improving the real-time performance of data interaction.
[0049] Preferably, in step 3, during the specific implementation of registering the hardware inference unit, the hardware abstraction layer of the first device calls the device detection interface to evaluate the performance of the hardware inference unit of the second device associated with a remote memory node through a local non-uniform memory access architecture. The hardware abstraction layer extracts the computing power level identifier of the hardware inference unit of the second device (such as the number of computing cores, the set of supported operators, and the peak computing power), and converts these characteristic parameters into a computing resource descriptor recognizable by the first device. Subsequently, the first device loads the computing resource descriptor into its local task allocation list, thereby formally registering the hardware inference unit of the second device as a locally schedulable remote inference engine. This registration mechanism enables the task scheduler of the first device to perform nanosecond-level scheduling responses to the remote inference engine, realizing a physically separated but logically integrated computing power architecture.
[0050] Preferably, after incorporating the remote inference engine into the scheduling system, the upper-level computing power aggregation protocol layer within the first device establishes a performance prediction model for the remote inference engine. Specifically, the computing power aggregation scheduling layer within the upper-level computing power aggregation protocol layer combines the real-time communication bandwidth of the direct interconnect link with the access latency of the remote memory node in the local non-unified memory access architecture to perform a weighted calculation of the execution efficiency of the remote inference engine, generating a scheduling permission level that reflects resource availability. Based on this scheduling permission level, the first device performs cross-device splitting of the task flow that originally only ran on the local neural network processing unit of the first device, enabling a portion of the computing power load in the pending computing tasks to be distributed to the remote inference engine for execution via the computing power optical interconnect communication channel. This dynamic registration and scheduling functional collaboration ensures a high level of utilization of cross-device computing power resources while also taking into account the overall power consumption constraints of collaborative computing.
[0051] Preferably, step 3 integrates the resources of all second devices with established direct interconnect links to ultimately form a remote memory node queue. The first device performs resource marking on each node in the remote memory node queue, and after registering the corresponding hardware inference unit of the second device as a locally schedulable remote inference engine, all remote inference engines registered as locally schedulable by the first device logically form a cross-device high-speed computing power aggregation resource pool. When performing heavy inference tasks, the first device performs pipeline orchestration on the resources in the cross-device high-speed computing power aggregation resource pool, allocating different neural network model layers to different remote inference engines. This processing method fully leverages the high bandwidth characteristics of the computing power optical interconnect communication channel, and through computing power collaboration and resource aggregation between the first and second devices, significantly alleviates the computing power bottleneck and memory space limitations faced by a single portable computing device when processing large-scale neural network models.
[0052] Optionally, step 3 includes: the first device identifying all second devices with which it has established the direct interconnect link; forming a remote memory node queue based on the remote memory nodes corresponding to the local non-unified memory access architecture (NUMA) of all second devices, so that after registering the hardware inference unit of the second device as the remote inference engine that is locally schedulable by the first device, all second devices corresponding to all the remote inference engines that are registered locally by the first device form a cross-device high-speed computing power aggregation resource pool.
[0053] Preferably, in the specific technical implementation of step 3, the first device, through its internally integrated computing unit, directly connects to the protocol layer and actively polls the consumer-grade computing optical interconnect interface connectors at each physical layer. Specifically, the computing unit directly connects to the protocol layer to identify the physical access sensing results fed back by the sensing pins inside the consumer-grade computing optical interconnect interface connector, thereby determining all second devices that have established direct interconnect links with the first device. Subsequently, the first device performs device identification code extraction processing on each second device in a connected state, thereby establishing an online device list reflecting the current physical connection topology. This identification process ensures that the basis for computing power aggregation is a real, high-speed optical physical layer link, avoiding the broadcast overhead in traditional network discovery protocols, and laying a precise topological foundation for subsequent millisecond-level dynamic expansion of computing resources.
[0054] Preferably, after determining the list of online devices, the first device performs logical abstraction processing on the memory space corresponding to each second device in the list. Specifically, the first device reads the capacity information and base address offset of the memory controller of each second device through the computing power optical interconnection communication channel, and uses the driver logic of the non-uniform memory access architecture to map these remote physical spaces to remote memory nodes of the local non-uniform memory access architecture within the first device. According to the order in which these remote memory nodes of the local non-uniform memory access architecture are mapped or the communication latency parameters, the first device establishes an ordered queue of remote memory nodes in kernel mode. This queued organization can clearly reflect the distribution of cross-device memory resources, enabling the upper-level computing power aggregation protocol layer to perform more refined cross-device data scheduling based on differences in memory access efficiency.
[0055] Preferably, step 3, based on the formation of the remote memory node queue, further performs refined probing of computing power capabilities. The hardware abstraction layer of the first device sends a computing power feature query command to each second device associated in the remote memory node queue through the computing power optical interconnection communication channel to extract the computing power level identifier of its hardware inference unit. In this process, the first device extracts key parameters reflecting the processing capabilities of the hardware inference unit, such as the number of computing cores, clock frequency, supported deep learning operator set, and peak half-precision floating-point operation, and aggregates these parameters into a computing resource descriptor reflecting the computing power characteristics of the second device. In this way, the first device can grasp the performance profile of each computing node before pooling, thereby providing a decision basis for task load balancing during subsequent pipelined collaborative inference.
[0056] Preferably, after generating computing resource descriptors for each second device, the first device synchronizes these computing resource descriptors to its internal task allocation list. Specifically, the task scheduler within the first device identifies the computing power weight in each computing resource descriptor and formally registers the hardware inference units of the second devices as locally schedulable remote inference engines for the first device. During the registration process, the first device assigns a unique virtual device handle to each remote inference engine, allowing upper-layer applications to access computing resources without distinguishing whether the resources reside on the local motherboard or on a second device connected via a consumer-grade computing optical interconnect interface connector. This registration mechanism eliminates the logical boundaries of cross-device scheduling, enabling remote hardware resources to participate in real-time computing tasks with lower instruction latency.
[0057] Preferably, after all online remote inference engines have completed registration, the first device uses logical aggregation logic to uniformly group these hardware inference unit resources distributed across different second devices, thereby forming a cross-device high-speed computing power aggregation resource pool. When forming this cross-device high-speed computing power aggregation resource pool, the first device performs cumulative calculations on the total computing bandwidth and total video memory capacity of all remote inference engines within the pool through the upper-level computing power aggregation protocol layer, generating a virtual ultra-large-scale computing power entity. Through this logical resource pooling process, the first device can break down heavy-duty deep learning models that originally required computing center-level computing power into multiple task slices that can be executed in parallel within the cross-device high-speed computing power aggregation resource pool, thus overcoming the physical limitations of portable single-machine computing devices in terms of computing power scale.
[0058] Preferably, for application scenarios involving cross-device high-speed computing power aggregation resource pools, the first device performs task collaboration on each remote inference engine within the pool based on a pipelined orchestration strategy. When the first device receives a heavy inference task, the upper-level computing power aggregation protocol layer distributes different layers of the neural network model to different remote inference engines according to computational load and data throughput requirements, based on the communication latency reflected in the remote memory node queue. For example, the forward propagation layer of the neural network model is assigned to a near-end inference engine with lower latency, while the fully connected layer with a large number of parameters is assigned to a remote inference engine with larger storage space. This pipelined collaborative inference processing method based on cross-device high-speed computing power aggregation resource pools fully utilizes the high effective bandwidth provided by the computing power optical interconnection communication channel, achieving efficient aggregation and consumption of cross-device computing power resources while ensuring memory consistency.
[0059] Optionally, the method further includes: the first device calling its MDCCP protocol layer to sense the link status of the direct interconnect link in real time, so as to determine the computing power scheduling mode of the first device for the second device based on the link status, and in the computing power scheduling mode, distributing the computing tasks to be processed by the first device to the cross-device high-speed computing power aggregation resource pool for consumption through the computing power optical interconnection communication channel.
[0060] Preferably, in the specific technical implementation where the first device invokes its upper-level computing power aggregation protocol layer to sense the link status of the directly connected interconnect link in real time, the upper-level computing power aggregation protocol layer, as an upper-level protocol management entity, performs high-frequency physical heartbeat detection on the directly connected interconnect link through the interface control logic provided by the underlying computing unit direct connection protocol layer. Specifically, the upper-level computing power aggregation protocol layer injects link probe packets with timestamp information into the directly connected interconnect link and monitors the physical clock offset of the link probe packets during the round trip between the first device and the second device to extract link status characteristics reflecting the current physical connection quality. The link status characteristics include not only millisecond-level bidirectional link latency, but also link bit error rate based on error correction coding feedback and optical signal coupling fluctuation frequency. These multi-dimensional feature data together constitute an objective quantitative representation of the link status of the directly connected interconnect link. Through this real-time sensing method, the first device can obtain physical layer connection feedback superior to traditional network sensing mechanisms, thereby providing real-time physical communication boundary parameters for subsequent accurate allocation of computing power.
[0061] Preferably, after acquiring the link state characteristics, the scheduling logic module of the computing power aggregation scheduling layer in the upper-level computing power aggregation protocol layer inside the first device performs a weighted evaluation of the link state characteristics to determine the computing power scheduling mode of the first device for the second device based on the link state. During the evaluation process, the scheduling logic module compares the sensed bidirectional link latency with a preset ultra-low latency threshold (e.g., less than 100 nanoseconds), while also considering the effective remaining bandwidth of the current computing power optical interconnect communication channel. If the evaluation result shows that the link state is highly stable and has sufficient throughput margin, the computing power scheduling mode is set to high-speed optical interconnect mode; if significant jitter or a decrease in optical signal coupling strength is detected in the direct interconnect link, the computing power scheduling mode is downgraded to redundancy check transmission mode or low-speed maintenance mode through an adaptive algorithm. This dynamic mode switching based on the actual physical link performance ensures that the computing tasks to be processed can always be executed in the most suitable communication environment, thereby alleviating the problem of collaborative inference interruption caused by interface fluctuations.
[0062] Preferably, after determining the computing power scheduling mode, the first device performs feature decomposition and computing power demand matching processing on the computing task to be processed, so as to distribute the computing task to be processed by the first device to the cross-device high-speed computing power aggregation resource pool through the computing power optical interconnection communication channel under the computing power scheduling mode. Specifically, the first device divides the computing task to be processed into multiple computing slices with independent execution semantics, and calculates the matching degree score between each computing slice and each remote inference engine in the cross-device high-speed computing power aggregation resource pool according to the scheduling strategy corresponding to the computing power scheduling mode. When the computing power scheduling mode is the high-speed optical interconnection mode, the first device will prioritize mapping heavy tasks with large-scale tensor operation requirements to the remote inference engine corresponding to the second device with a large number of computing cores and low communication latency in the cross-device high-speed computing power aggregation resource pool. This deep matching process realizes the coordinated adaptation of task attributes and physical resource capabilities, and provides clear task topology guidance for subsequent efficient parallel consumption.
[0063] Preferably, for the distribution process of heavy tasks, the first device uses the consumer-grade computing optical interconnect interface transport layer protocol to encapsulate the matched heavy task into a computing power optical interconnect protocol data frame. During the encapsulation process, the first device associates a scheduling permission level reflecting the task priority with the heavy task and embeds this scheduling permission level information into the extended field of the computing power optical interconnect protocol data frame. Subsequently, the first device drives the miniaturized optical transceiver module inside its integrated system-on-a-chip to convert the computing power optical interconnect protocol data frame in the electrical domain into a high-frequency optical pulse sequence reflecting the logic of the heavy task. Through the computing power optical interconnect communication channel, the high-frequency optical pulse sequence is sent to the cross-device high-speed computing power aggregation resource pool at a single-channel transmission rate determined in step 1 (e.g., not less than 100 gigabits per second), realizing direct connection of the raw data stream at the computing unit level. This distribution method based on optical interconnect greatly compresses the time ratio of task transmission by utilizing the high-speed characteristics of optical carriers, ensuring that the heavy task can be consumed in the remote inference engine at an access speed close to that of the local bus.
[0064] Preferably, during the distribution of the computing tasks to be processed, the first device also implements dynamic distribution scale constraints based on real-time energy status feedback. Specifically, the first device calls its upper-level computing power aggregation protocol layer to construct a wireless charging channel, and obtains the wireless charging status parameters of the charging device based on the constructed wireless charging channel through the configured wireless control unit. The wireless charging status parameters involve energy input characteristics such as real-time input power, coil alignment quality, and charging area status, wherein the coil alignment quality can be further converted into a normalized alignment score. The first device performs energy balance calculation based on the wireless charging status parameters to obtain a real-time available computing power budget reflecting the current system endurance. By converting the real-time available computing power budget into a hard constraint condition for task scheduling, the first device can dynamically adjust the scheduling permission level, thereby limiting the task density distributed to the cross-device high-speed computing power aggregation resource pool, preventing device overheating or abnormal downtime due to power consumption overload during collaborative computing.
[0065] Preferably, when implementing dynamic constraints on the distribution scale of the computing tasks to be processed, the first device further refines the computing tasks to be processed into node-level micro-tasks by generating computing power reservation requests. The first device identifies the physical space covered by the charging device and divides the execution cycle of the computing tasks to be processed into multiple execution windows corresponding to the charging coverage area. Within each execution window, the first device distributes and consumes the node-level micro-tasks segment by segment based on the task throughput limit determined by the real-time available computing power budget. If the wireless charging status parameters show a decrease in energy input (such as a decrease in the normalized alignment score), the first device will automatically reduce the number of node-level micro-tasks in the current execution window. This deep coupling of computing power scheduling and wireless energy acquisition forms a complete resource, communication, and energy coordination system, enabling consumer-grade computing devices to balance high performance and device operation safety when handling heavy artificial intelligence tasks.
[0066] Optionally, during the process of mapping the memory space of the second device to a remote memory node of the Local Non-Unified Memory Access Architecture (NUMA), the first device maintains memory address consistency between the first device and the second device through the CCOI transport layer protocol. This memory address consistency enables physical addressing of the memory space of the second device when the first device accesses the remote memory node.
[0067] Preferably, in the specific technical implementation of mapping the memory space of the second device to a remote memory node of the local non-unified memory access architecture, the first device performs real-time sensing of the physical resource attributes of the memory controller of the second device through the established direct interconnect link. Specifically, the first device obtains the physical address range and memory base address offset of the second device through the computing power optical interconnect communication channel, and imports these addressing parameters reflecting the physical layout into the address translation table logic of the kernel mode of the first device. This sensing process enables the first device to obtain the original storage layout characteristics of the second device, thereby providing a clear hardware addressing basis for subsequently building a cross-device globally unified addressing view within the first device.
[0068] Preferably, the first device utilizes the consumer-grade computing optical interconnect interface transport layer protocol to initiate a hardware-level handshake process for memory address mapping between the first device and the second device. During the hardware-level handshake process, the first device logically reassembles the captured physical address range with its own local memory address space, thereby allocating a globally unique remote node identifier for the second device's memory space in the addressing logic layer that reflects the memory topology. Through this hardware-level handshake process, the first device can transform the physically discrete storage resources of the second device into logically contiguous storage units, enabling the central processing unit or neural network processing unit within the first device to directly execute instruction-level calls to the remote memory nodes of the local non-uniform memory access architecture.
[0069] Preferably, after generating the globally unique remote node identifier, the first device establishes a functional link reflecting the memory state synchronization between the first device and the second device through the cache consistency maintenance logic in the transport layer protocol of the consumer-grade computing optical interconnect interface. The first device uses the direct interconnect link to synchronize the memory access state in real time and sends memory address consistency messages to the second device to establish an address space mapping relationship at the hardware level. This processing action enables the first device to maintain memory address consistency through a hardware monitoring mechanism when performing data access on the remote memory node of the local non-unified memory access architecture, thereby avoiding computational logic anomalies caused by inconsistent data copy versions during cross-device collaborative computing.
[0070] Preferably, the first device performs low-level real-time maintenance processing for memory address consistency through the consumer-grade computing optical interconnect interface transport layer protocol. Specifically, the first device monitors memory access requests from remote memory nodes in the local non-unified memory access architecture in real time and determines whether the memory unit involved in the memory access request has undergone a state change in the local cache of the second device. If a change is determined, the first device triggers hardware-level atomic operation instructions through the computing power optical interconnect communication channel to synchronously update the state of the corresponding physical memory unit. Through this hardware synchronization method that does not rely on operating system intervention, dynamic alignment of memory address consistency between the first device and the second device is achieved, providing high execution continuity for collaborative inference tasks.
[0071] Preferably, when the first device accesses a remote memory node of the local non-unified memory access architecture, the direct connection protocol layer of the computing unit inside the first device translates the virtual memory access instruction directly into a hardware-level physical addressing processing action for the second device. When executing the hardware-level physical addressing processing action, the first device skips the traditional kernel-mode network protocol stack encapsulation and forwarding process, and instead directly sends the computing power optical interconnect protocol data frame containing the physical addressing parameters of the addressing target to the second device via the optical physical layer link. Through this direct physical mapping technique, the first device can achieve physical addressing of the memory space of the second device, keeping the end-to-end communication latency for cross-device memory access at a low level (e.g., below 100 nanoseconds).
[0072] Preferably, during the physical addressing of the memory space of the second device, the first device also dynamically adjusts the dispatch strategy corresponding to the hardware-level physical addressing processing action based on the real-time link status characteristics of the direct interconnect link. If the optical signal coupling strength of the optical physical layer link is detected to be at a high level (i.e., the normalized alignment score is high), an ultra-low latency physical pass-through mode is enabled to improve the synchronization frequency of memory address consistency. If the direct interconnect link experiences minor jitter, the flow control strategy in the consumer-grade computing optical interconnect interface transport layer protocol is used to perform ordered rearrangement and error correction reinforcement on the physical addressing requests corresponding to the hardware-level physical addressing processing action. This functional design, which deeply couples physical layer interconnect quality with logical layer address mapping, forms a robust cross-device memory sharing system, enabling the first device to achieve access performance to remote memory nodes of the local non-unified memory access architecture that is close to the operating level of the local memory bus.
[0073] Optionally, when distributing the computing tasks to be processed based on the computing power scheduling mode, if the computing power scheduling mode is a high-speed optical interconnect mode, the first device will distribute the heavy tasks in the computing tasks to be processed to the cross-device high-speed computing power aggregation resource pool to match the second device through the computing power optical interconnect communication channel, so as to build a pipelined collaborative inference cluster.
[0074] Preferably, in the logical processing of distributing computational tasks to be processed by the first device, the first device first performs a quantitative analysis of the computational overhead and data scale of the computational tasks to be processed. Specifically, the first device extracts the neural network layer topology, tensor parameter quantity, and logical operand distribution characteristics of the computational tasks to be processed, and calculates a load overhead score reflecting the complexity of the task based on these characteristics. If the load overhead score exceeds a preset heavy computation judgment threshold, the corresponding computational task to be processed is marked as a heavy task. Through this deep identification processing of the attributes of the computational tasks to be processed, the first device can selectively filter out task loads that have high requirements for both computing resources and communication bandwidth, thereby providing a decision basis for task classification hierarchy for subsequent accurate scheduling on high-speed physical channels.
[0075] Preferably, when the first device determines that the current computing power scheduling mode is high-speed optical interconnect mode, it indicates that the direct interconnect link between the first device and the second device is in a high-bandwidth and extremely low-latency ready state. Under the constraint of high-speed optical interconnect mode, the first device activates a high-priority distribution operator for heavy tasks. This high-priority distribution operator confirms that the effective remaining bandwidth of the computing power optical interconnect communication channel can carry the large-scale tensor flow required by the heavy tasks by querying the link status characteristics stored in the computing power aggregation scheduling layer in the upper-level computing power aggregation protocol layer. This processing logic transforms the advantages of physical interconnection into the execution momentum of logical scheduling, ensuring that neural network inference tasks containing massive weights can preferentially occupy the transmission resources of the optical physical layer link, thereby physically avoiding the severe congestion phenomenon caused by traditional electrical interfaces when handling large-scale data transfer.
[0076] Preferably, the first device performs specific signal dispatching actions for the identified heavy tasks via the optical interconnect communication channel. During the dispatching process, the computing unit inside the first device directly connects to the protocol layer to encapsulate the computational operators and the tensors to be inferred corresponding to the heavy tasks into the optical interconnect protocol data frames, and drives the miniaturized optical transceiver module integrated inside the main chip package of the first device to convert them into high-frequency optical pulse sequences. Since the optical interconnect communication channel has the physical characteristic of a single-channel transmission rate of not less than 100 gigabits per second as determined in step 1, the parameter matrix to be dispatched for the heavy tasks can be completely projected to the cross-device high-speed computing power aggregation resource pool within a very short time window with a transmission efficiency close to that of the main processor's internal bus. This direct-connection dispatching method based on the optical physical layer link significantly reduces the end-to-end communication ratio of task dispatching, providing a solid data synchronization foundation for subsequent real-time consumption on the remote inference engine within the cross-device high-speed computing power aggregation resource pool.
[0077] Preferably, the first device performs matching processing actions for heavy tasks based on the real-time operational profiles of each second device in the cross-device high-speed computing power aggregation resource pool. The first device evaluates the execution performance score of each node for the associated neural network operators by reading the computing power level identifiers of each remote inference engine associated in the remote memory node queue. The first device performs association mapping processing on the task slices after the heavy tasks are divided and the second devices with higher execution performance scores to realize the functional support relationship between task load and hardware resources. This matching logic based on resource profiles ensures that heavy tasks can be accurately dispatched to physical nodes with strong neural network processing capabilities, thereby achieving optimal cross-device computing power configuration at the cross-device high-speed computing power aggregation resource pool level and alleviating the execution efficiency bottleneck caused by uneven node performance.
[0078] Preferably, the first device performs functional reorganization on multiple matched second devices within a cross-device high-speed computing power aggregation resource pool to construct a pipelined collaborative inference cluster. Specifically, the first device performs hierarchical decomposition processing on the neural network model in heavy tasks according to logical sequence, and deploys the computational slices of different levels to different remote inference engines within the cross-device high-speed computing power aggregation resource pool in a cascaded manner. Each remote inference engine achieves nanosecond-level access to intermediate layer activation values through remote memory nodes with a local non-uniform memory access architecture, and transfers the computation results to the next-level remote inference engine in real time via direct interconnect links. This pipelined collaborative inference cluster construction method utilizes the parallelism of physical resources among multiple devices to transform the originally serial inference logic into a cascaded parallel computation flow, greatly improving the inference throughput when processing large-scale complex neural network models.
[0079] Preferably, during the construction of the pipelined collaborative inference cluster, the first device utilizes the upper-level computing power aggregation protocol layer to monitor the execution heartbeat and data flow rate of each node in the pipeline in real time. If it senses that the processing latency of the remote inference engine corresponding to a certain link increases due to power consumption constraints imposed by the real-time available computing power budget determined by the wireless charging status parameters, the first device will dynamically update the computing power scheduling mode and readjust the mapping relationship of task slices within the cross-device high-speed computing power aggregation resource pool. Through this dynamic orchestration processing action based on logical feedback, the pipelined collaborative inference cluster can always maintain high task consumption efficiency, enabling consumer-grade computing devices to complete complex neural network artificial intelligence inference tasks that could originally only run on server clusters in a multi-device collaborative manner, achieving efficient aggregation and secure consumption of cross-device computing power resources while ensuring memory consistency.
[0080] Optionally, the method further includes: the first device calling its MDCCP protocol layer to construct a wireless charging channel, and obtaining wireless charging status parameters of the charging device based on the constructed wireless charging channel through a configured wireless control unit, wherein the wireless charging status parameters include real-time input power, coil alignment quality, and charging area status; the first device calculating a real-time available computing power budget based on the wireless charging status parameters, establishing a scheduling permission level based on the real-time available computing power budget, and then implementing dynamic constraints on the dispatch scale of the computing tasks to be processed based on the scheduling permission level.
[0081] Preferably, in the specific technical implementation where the first device invokes its upper-level computing power aggregation protocol layer to construct a logical wireless charging channel, the upper-level computing power aggregation protocol layer activates its internal wireless charging protocol extension sublayer through a downlink control interface. This wireless charging protocol extension sublayer, through a wireless control unit connected to the hardware circuitry, initiates physical layer sensing between the first device and surrounding charging devices. Specifically, the wireless control unit, through its in-band communication channel, utilizes near-field electromagnetic coupling load modulation technology to perform a protocol handshake with the charging device, thereby establishing the logical wireless charging channel based on physical electromagnetic coupling. This method of deeply controlling the energy capture link through the upper-level computing power aggregation protocol layer allows the scheduling logic of the computing task to be processed to anticipate the possibility of energy input, laying a communication foundation for subsequently constructing a "computing power-energy" collaborative view reflecting the coupling relationship between energy and computing.
[0082] Preferably, the first device, through the configured wireless control unit, performs real-time acquisition and processing of wireless charging status parameters for the charging device based on the constructed logical wireless charging channel. During this process, the in-band protocol module within the wireless control unit demodulates the modulated signal in the coupled magnetic field, extracting multiple key indicators reflecting the current energy transmission characteristics. These indicators, serving as the wireless charging status parameters, specifically include, but are not limited to, real-time input power reflecting the capability of the primary coil transmitter (e.g., a nominal value of 15 watts or higher), coil alignment quality reflecting the spatial coupling efficiency between the transmitting and receiving coils, and the charging area status reflecting the presence of metallic foreign objects or temperature anomalies on the charging surface. Through this refined parameter extraction and logical processing, the first device can transform the electromagnetic state, originally belonging to the hardware layer, into energy characteristic data understandable by the upper-level computing power aggregation protocol layer.
[0083] Preferably, based on the acquired coil alignment quality, the first device performs coupling efficiency quantification processing using an energy efficiency evaluation algorithm. Specifically, the first device performs correlation mapping processing on the captured coil induced voltage fluctuation amplitude and a preset alignment reference spectrum, thereby generating a normalized alignment score reflecting the energy conversion efficiency coefficient. If the normalized alignment score shows that the coil offset is within a preset small range (e.g., less than 25 mm), it is determined that the current energy coupling efficiency is high; if the normalized alignment score is low, alignment compensation suggestions are sent to the user interface through the upper-level computing power aggregation protocol layer. This processing step, through objective evaluation of the physical layer coupling quality, ensures that the reference benchmark for subsequent computing power prediction has high reliability, thereby alleviating the problem of unstable computing power supply prediction caused by fluctuations in the wireless charging location.
[0084] Preferably, the first device performs a calculation of the real-time available computing power budget based on the wireless charging status parameters. Specifically, the computing power aggregation scheduling layer in the upper-level computing power aggregation protocol layer inside the first device initiates energy balance calculation logic, multiplying the real-time input power with the energy conversion efficiency coefficient determined by the normalized alignment score to obtain the effective charging power actually entering the system. Subsequently, the system deducts the reserved power for maintaining battery cycling and the basic power consumption for maintaining basic system operation (such as the static power consumption of the motherboard and display components) from the effective charging power, thereby calculating the real-time available computing power budget that can be fully allocated to the computing unit to perform tasks. This real-time budget mechanism based on dynamic energy input transforms the traditional static computing power limit into a dynamic boundary that fluctuates with the energy state, enabling consumer-grade computing devices to realistically plan the "task flow rate" based on the current "energy reserves" when processing heavy neural network artificial intelligence inference tasks.
[0085] Preferably, the first device establishes scheduling permission levels within the computing power aggregation and scheduling layer based on the real-time available computing power budget. The scheduling permission levels are defined as a series of logical tiers reflecting the degree of computing power openness, each corresponding to a different processing scale of the pending computing tasks. Specifically, the first device compares the calculated real-time available computing power budget with a preset task distribution threshold table. If the real-time available computing power budget is sufficient, a higher scheduling permission level is granted, allowing the first device to distribute all heavy tasks in the pending computing tasks to the cross-device high-speed computing power aggregation resource pool. If the real-time available computing power budget is limited (e.g., due to power limitations of the charging device or a low normalized alignment score), the scheduling permission level is automatically downgraded. By establishing this explicit level mapping relationship, the first device can perform differentiated access control on heterogeneous computing power resources at the logical level, realizing the functional support relationship between computing power consumption and energy acquisition.
[0086] Preferably, in the process of dynamically constraining the scale of the computing tasks to be processed, the first device performs logical adjustment on the task density distributed to the cross-device high-speed computing power aggregation resource pool through real-time feedback of the scheduling permission level. When the scheduling permission level switches due to changes in energy status, the high-priority distribution operator contained in the upper-level computing power aggregation protocol layer inside the first device automatically adjusts the transmission frequency of the computing power optical interconnect protocol data frames or reassembles the size of the task slices, thereby achieving dynamic constraints on the scale of task distribution without interrupting the operation of the current pipelined collaborative inference cluster. This technical solution, which deeply integrates energy awareness into the computing power scheduling process, effectively solves the risks of overheating or excessive battery discharge that are prone to occur in high-performance interconnect scenarios for consumer-grade portable computing devices, enabling the first device to execute large-scale computing tasks using the remote inference engine in a more robust manner.
[0087] Optionally, the first device and the charging device transmit a computing power access handshake command through the in-band communication channel of the wireless control unit to trigger the acquisition of the wireless charging status parameters of the charging device based on the constructed wireless charging channel.
[0088] Preferably, in the specific technical implementation of transmitting the computing power access handshake command between the first device and the charging device through the in-band communication channel of the wireless control unit, the first device initiates the data interaction process within the established physical electromagnetic coupling field through the integrated wireless control unit. Specifically, the wireless control unit utilizes near-field electromagnetic induction load modulation technology to superimpose logical transitions reflecting handshake semantics onto the wireless power transmission carrier signal, thereby constructing the in-band communication channel carrying the computing power access handshake command without adding an additional radio frequency link. This processing method, which directly achieves logical alignment between the first device and the charging device through physical electromagnetic coupling characteristics, ensures the synchronization of energy interaction and information confirmation, and provides an extremely compact physical trigger entry for subsequently obtaining the wireless charging state parameters reflecting the energy boundary.
[0089] Preferably, in the processing of transmitting the computing power access handshake instruction, the first device encapsulates the protocol data corresponding to the computing power access handshake instruction to be sent in an extended field of a wireless charging standard message frame. The extended field is defined as a binary data block with a self-describing structure, containing a message type reflecting message attributes, a computing power level reflecting the computing power capability of the first device, an estimated dwell time reflecting the estimated dwell time of the first device in the charging area, and a task priority mask used to identify the importance of the computing task to be processed. By embedding computing power metadata in the wireless charging standard message frame in this way, the first device can complete the pre-announcement of computing power access requirements during the energy handshake phase. This processing method enables the charging device and the first device to establish an access cooperation logic based on energy supply and demand balance, thereby providing forward-looking data guidance for subsequent real-time dynamic scheduling.
[0090] Preferably, when the charging device senses the computing power access handshake command, it sends back a handshake confirmation message through its internal control logic, thereby triggering the first device to obtain the wireless charging status parameters of the charging device based on the constructed logical wireless charging channel. During this triggering process, the upper-level computing power aggregation protocol layer of the first device will switch from a simple power receiving mode to an "energy-computing power collaborative sensing mode". In this mode, the wireless charging protocol extension sublayer inside the first device will continuously sample the feedback signal transmitted by the in-band communication channel and extract the modulation waveform carrying energy transmission quality information. This technical step, through a clear command triggering mechanism, transforms the static physical connection into a dynamic status monitoring link, ensuring that the energy data on which the subsequent real-time available computing power budget calculation is based has high timeliness.
[0091] Preferably, in the specific implementation of acquiring the wireless charging status parameters, the first device performs multi-dimensional feature mapping processing on the protocol signal transmitted through the demodulated in-band communication channel to extract key physical features reflecting the current energy transmission characteristics. Specifically, the first device first identifies the real-time input power reflecting the transmitting capability of the primary coil in the charging device (e.g., a nominal value of 15 watts or higher), and simultaneously acquires the coil alignment quality reflecting the spatial positional relationship between the primary coil and the receiving coil inside the first device. Furthermore, the first device also receives the charging area status reflecting the safety of the physical environment of the charging surface, determining whether there is a risk of energy loss due to foreign metal objects. Through this refined parameter extraction logic processing, the first device can obtain a set of quantitative indicators reflecting the stability of energy supply, thereby providing physical-level input for subsequently implementing dynamic constraints on the scale of computational tasks to be processed at the logical level.
[0092] Preferably, after receiving the coil alignment quality, the first device performs coupling efficiency quantification processing through a preset coupling efficiency evaluation algorithm. Specifically, the first device performs correlation mapping processing on the captured coil induced voltage fluctuation amplitude and a preset alignment reference spectrum, thereby generating a normalized alignment score reflecting the energy conversion efficiency coefficient. This normalized alignment score serves as a weighting factor in subsequent correction calculations for the real-time input power. If the coil alignment quality feedback from the handshake process is at a high level, the generated normalized alignment score is high (e.g., close to 1.0), indicating a high energy capture efficiency; otherwise, the expected energy gain is automatically lowered through the upper-level computing power aggregation protocol layer. This real-time efficiency evaluation based on handshake feedback effectively solves the problem of inaccurate real-time available computing power budget caused by the randomness of consumer-grade computing devices in wireless charging locations, enabling subsequent task distribution logic to be based on realistic energy expectations.
[0093] Preferably, through the successful transmission and feedback of the aforementioned computing power access handshake command, the computing power aggregation scheduling layer in the upper-level computing power aggregation protocol layer inside the first device will ultimately lock the wireless charging status parameters. These parameters are transmitted back in real time to the scheduling logic module inside the computing power aggregation scheduling layer via the internal bus to trigger a refresh action for the real-time available computing power power budget. Once the handshake is confirmed and the wireless charging status parameters are ready, the first device can, based on the real-time energy input intensity and combined with the current consumption capacity of the cross-device high-speed computing power aggregation resource pool, drive the high-priority distribution operator in the upper-level computing power aggregation protocol layer to dynamically constrain the task slice size distributed to the remote inference engine. This functional support relationship from handshake triggering to parameter acquisition and then to final scheduling constraints constitutes a complete cross-device computing power, communication, and energy collaborative logic, enabling the first device to robustly handle large-scale neural network artificial intelligence inference tasks in high-speed optical interconnect mode.
[0094] Optionally, the computing power access handshake instruction is encapsulated in an extended field of a wireless charging standard message frame, the extended field including message type, computing power level, expected dwell time, and task priority mask.
[0095] Preferably, in the specific technical implementation where the computing power access handshake instruction is encapsulated in the extended field of the wireless charging standard message frame, the integrated wireless control unit logic within the first device constructs a binary data block conforming to a specific binary structure. This binary data block is defined as a protocol payload with self-describing characteristics, mapped to reserved or privately defined bits in the non-energy payload of the wireless charging standard message frame. Through this encapsulation process, the first device sinks the computing power collaboration intent carried by the upper-level computing power aggregation protocol layer down to the lower-level energy interaction protocol, achieving coaxial transmission of control signaling and energy payload. This processing method enables the first device and the charging device to complete preliminary capability synchronization using physical electromagnetic coupling fields before entering a large-scale data exchange stage, providing low-overhead preliminary information support for the subsequent construction of a cross-device high-speed computing power aggregation resource pool.
[0096] Preferably, the message type included in the extended field is defined as a function identification code. During encapsulation processing, the first device sets the message type to an identifier value reflecting a "computing power access request" based on the current task distribution requirements, explicitly informing the charging device that the current session is not merely a regular power report message. When the charging device performs message parsing processing on the demodulated computing power access handshake instruction, it first identifies the message type to determine the parsing logic of subsequent protocol branches, thereby activating the corresponding resource response mechanism. This explicit message classification ensures that instructions with different service attributes can be accurately routed within the same logical wireless charging channel, avoiding semantic confusion between energy management logic and computing power scheduling logic.
[0097] Preferably, the computing power level included in the extended field is used to quantitatively characterize the static hardware characteristics of the computing units inside the first device. In a specific implementation, the first device converts characteristic parameters (i.e., computing power level identifiers) such as the number of computing cores in its integrated neural network processing unit, the completeness of its supported deep learning operator set, and the theoretical peak value of half-precision floating-point operations into a level index reflecting its hardware capabilities through a mapping algorithm. This level index, as a key component of the extended field, provides a standardized capability label for remote scheduling logic. By pre-informing the computing power level, the system can pre-assess the potential contribution of the current node when participating in the cross-device high-speed computing power aggregation resource pool, thereby providing a hardware-dimensional input reference for subsequent refined real-time available computing power power budget calculations.
[0098] Preferably, the estimated dwell time included in the extended field is generated based on a predictive analysis of the current movement state and energy demand of the first device. This field carries a parameter reflecting the estimated duration the first device will remain within the effective power coverage area of the charging device. By embedding the estimated dwell time into the computing power access handshake command, the first device can guide the upper-level computing power aggregation protocol layer to divide the computing task to be processed into multiple execution windows corresponding to the charging coverage area. This technical processing ensures that the task splitting logic (such as splitting into node-level micro-tasks) can be physically aligned with the energy acquisition window period, thereby effectively mitigating the risk of computing power coordination interruption caused by frequent entry and exit of mobile consumer computing devices from the charging area.
[0099] Preferably, the task priority mask included in the extended field is defined as a bitmap-structured scheduling filter. The first device generates a corresponding task priority mask based on the urgency of the computational tasks to be processed (e.g., high priority for real-time sensing tasks and low priority for background data analysis tasks). During the process of establishing scheduling permission levels, this task priority mask serves as input to the constraint operator, directly determining which node-level microtasks are allowed to be distributed to the remote inference engine via the computing power optical interconnect communication channel. This mask-based fine-grained admission control logic enables the system to prioritize the distribution of critical computational loads when the real-time available computing power budget is limited, thus realizing the functional support relationship between computing power resource distribution density and real-time energy supply status.
[0100] Preferably, by structurally integrating message type, computing power level, expected dwell time, and task priority mask, the computing power access handshake instruction forms a complete "energy-computing power collaboration context" within the wireless charging standard message frame. When the instruction is transmitted to the charging device via the in-band communication channel and receives feedback, the computing power aggregation scheduling layer in the upper-level computing power aggregation protocol layer within the first device immediately locks the wireless charging status parameters reflecting energy intensity. This encapsulation technology, which deeply couples computing power metadata into the energy handshake message, significantly shortens the logical path from physical access to computing power readiness. This allows the first device to trigger a refresh action for the real-time available computing power power budget with low latency, thereby enabling the high-priority distribution operator in the upper-level computing power aggregation protocol layer to dynamically and robustly constrain the task size distributed to the cross-device high-speed computing power aggregation resource pool while ensuring memory consistency.
[0101] Optionally, the method further includes: the first device generating a computing power reservation request, splitting the computing task to be processed into node-level microtasks based on the computing power reservation request, and dividing the computing task to be processed into multiple execution windows corresponding to the charging coverage area; when implementing dynamic constraints on the dispatch scale of the computing task to be processed based on the scheduling permission level, executing the node-level microtasks segment by segment through the multiple execution windows, wherein the charging coverage area consists of multiple charging devices that have established a wireless charging channel with the first device.
[0102] Preferably, in the specific technical implementation of the first device generating the computing power reservation request, the first device actively monitors its own geographical location movement trajectory and the physical distribution density of the surrounding charging devices through the upper-level computing power aggregation protocol layer to predict the energy capture potential within a preset time period. Specifically, the computing power aggregation scheduling layer in the upper-level computing power aggregation protocol layer combines the expected dwell time and preset movement speed included in the computing power access handshake instruction to construct a time-series prediction model reflecting the energy supply intensity, thereby triggering the generation of the computing power reservation request containing expected energy gain information. This processing action transforms the originally discrete wireless charging event into a logically continuous expectation of computing power resources, providing a data guide at the task planning level for the subsequent decomposition of heavy neural network artificial intelligence inference tasks into intermittently executable logical units.
[0103] Preferably, the first device performs fine-grained splitting of the computing task to be processed based on the generated computing power reservation request to obtain node-level microtasks. During the splitting process, the upper-level computing power aggregation protocol layer identifies the neural network operator dependencies in the computing task to be processed, and divides the large-scale tensor operation task into a sequence of data blocks with independent atomic execution characteristics, thereby forming the node-level microtasks. Each node-level microtask is assigned a specific execution weight and memory usage index, enabling these task slices to adapt to the fragmented scheduling needs of portable computing devices in cross-device high-speed computing power aggregation resource pools, significantly alleviating the bottleneck of being unable to complete execution within a limited execution window due to the large size of the task logic body.
[0104] Preferably, the first device identifies a charging coverage area comprised of multiple charging devices that have established a logical wireless charging channel with it, and divides the computational task to be processed into multiple execution windows corresponding to the charging coverage area. Specifically, the computing power aggregation and scheduling layer maps the movement trajectory of the first device into a series of physical intervals reflecting the energy injection intensity based on the physical location distribution of the charging devices fed back in the computing power reservation request. Each physical interval is defined as an execution window in the time dimension, used to limit the legal distribution cycle of the node-level micro-task. Through this technical processing method of aligning the task execution time with the physical charging path in a spatial dimension, an orderly distribution of the computing load on the spatial movement path is achieved.
[0105] Preferably, in the processing action of dynamically constraining the dispatch scale of the computing tasks to be processed based on the scheduling permission level, the first device uses the scheduling permission level as the flow control threshold within the execution window. The high-priority distribution operator in the upper-level computing power aggregation protocol layer reads the real-time available computing power budget mapped by the wireless charging status parameters reflecting the current energy level, and calculates the upper limit of the total number of node-level microtasks that can be dispatched within the execution window based on the expected dwell time corresponding to the current execution window. If the normalized alignment score shows that the current energy coupling efficiency is at a high level, the scheduling logic inside the first device automatically increases the task concurrency density within the execution window. This dynamic throttling mechanism based on energy input intensity ensures that the cross-device task dispatch scale is always within the safe envelope of the hardware thermal design.
[0106] Preferably, the first device guides the remote inference engine within the cross-device high-speed computing power aggregation resource pool to execute the node-level microtasks segment by segment through the multiple execution windows. During execution, the first device maintains memory address consistency using the consumer-grade computing optical interconnect interface transport layer protocol, temporarily storing the intermediate inference activation value tensor generated in the previous execution window in the remote memory node of the local non-uniform memory access architecture, as input for subsequent tasks in the next execution window. This segment-by-segment execution method allows the first device to temporarily suspend computing power distribution within the energy dead zone between the charging devices and quickly restart after entering a new execution window, thereby utilizing multiple discrete energy replenishment windows to complete the complete inference process of a heavy neural network model.
[0107] Preferably, by deeply coupling the computing power reservation request with the execution window, the first device achieves fine-grained control over the dispatch scale of the computing tasks to be processed while ensuring memory address consistency. This technical solution fully utilizes the high transmission bandwidth provided by the computing power optical interconnection communication channel and the real-time energy feedback of the wireless charging environment, solving the risk of collaborative inference crashes caused by unstable energy supply in mobile internet scenarios for consumer-grade portable computing devices. Compared with the traditional computing power distribution mode based on fixed power supply, this application discretizes the computing load and physically aligns it with the energy coverage area, enabling the device to consume heterogeneous computing power in the cross-device high-speed computing power aggregation resource pool relatively smoothly in complex mobile environments, achieving a dynamic balance between high computing performance and device operational security.
[0108] In all the above embodiments, the first device is the task dispatcher and resource aggregator, while the second device is the computing power supplier. Because this application introduces the construction of a wireless charging channel, the first device is particularly suitable for consumer electronic devices that require high-load AI computing while in motion, and which obtain energy through charging devices (such as in-vehicle wireless charging pads or desktop wireless charging stands).
[0109] Regarding the optical interconnect architecture for computing power of computing devices involved in this application, in practical application scenarios, the specific forms of the first device (usually acting as a local master device initiating scheduling) and the second device (a remote slave device providing extended computing power) can include, but are not limited to, the following combined embodiments: 1. Mobile office collaboration scenarios First device: laptop, tablet, or smartphone.
[0110] Second device: Dedicated external computing power expansion box (AI Box), high-performance desktop workstation, or another laptop with available computing power.
[0111] Scenario Description: Users connect their tablets to an external computing power unit via a fiber optic cable with a CCOI interface. The tablets register the NPU inside the computing power unit as a local inference engine, thereby enabling the smooth local execution of large language models.
[0112] 2. Intelligent New Energy Vehicle Scenarios The first device: the in-vehicle intelligent cockpit domain controller or central computing platform of an intelligent vehicle.
[0113] The second device: a smartphone or tablet equipped with a high-performance computing chip carried by the user, or a computing power enhancement module (computing blade) plugged into the CCOI interface in the vehicle.
[0114] Scenario Description: When a mobile phone is placed in the wireless charging slot in the car or connected via the CCOI interface, the vehicle system uses the optical interconnect link to call the phone's NPU to collaboratively process complex autonomous driving assistance algorithms or multimodal interaction tasks in the smart cockpit.
[0115] 3. Portable workstations and expansion modules First device: Portable ultra-thin workstation.
[0116] The second device is a stacked heterogeneous computing module (containing multiple NPU computing units) or an expansion dock with high-speed storage and computing integration.
[0117] Scenario Description: Through blind-plug connection via the CCOI interface, the first device maps the large memory capacity of the second device as a local remote NUMA node, solving the problem of insufficient video memory in portable devices when processing massive tensor data.
[0118] 4. Edge computing and industrial vision scenarios First device: Industrial-grade embedded host computer.
[0119] Second device: Vision Processing Unit (VPU) or FPGA accelerator card with a high-speed optical interface.
[0120] Scenario Description: In industrial sites with high bandwidth and severe electromagnetic interference, the main control computer utilizes the anti-interference characteristics of the optical interconnect interface to schedule remote vision units in real time for defect detection and inference.
[0121] 5. Cross-device game / entertainment rendering First device: Handheld game console or VR / AR all-in-one device.
[0122] Second device: Desktop-level high-performance gaming console or cloud-side edge computing node (local access device).
[0123] Scenario Description: The handheld device obtains support from the desktop hardware inference unit through the optical interconnect interface, enabling complex real-time ray tracing or super-resolution reconstruction algorithms and improving the rendering frame rate.
[0124] like Figure 2 As shown, an embodiment of this application discloses a computing device optical interconnection interface system, which includes: a first device and a second device, wherein a computing power optical interconnection communication channel is constructed between the first device and the second device; a direct interconnection link is established between the first device and the second device based on the computing power optical interconnection communication channel; and the first device registers the hardware inference unit of the second device as a locally schedulable remote inference engine based on the direct interconnection link.
[0125] Figure 2 An exemplary explanation can be found above. Figure 1 This will not be elaborated upon here.
Claims
1. A method for constructing a computing power optical interconnect interface for a computing device, characterized in that, Includes the following steps: Step 1: Construct a computing power optical interconnection communication channel containing link parameter information between the first device and the second device; Step 2: Establish a direct interconnection link between the first device and the second device based on the computing power optical interconnection communication channel; Step 3: Based on the direct interconnection link, the first device registers the hardware inference unit of the second device as a locally schedulable remote inference engine.
2. The method according to claim 1, characterized in that, Step 1 includes: a first device sensing the physical access of a second device through a CCOI connector; in response to sensing the physical access of the second device, triggering the establishment of an optical physical layer link between the first device and the second device; the first device obtaining the link parameter information of the second device through the optical physical layer link; and the first device establishing the computing power optical interconnection communication channel based on the link parameter information.
3. The method according to claim 2, characterized in that, The link parameter information includes single-channel rate, number of parallel channels, operating wavelength, and module power consumption level.
4. The method according to claim 1, characterized in that, Step 2 includes: the first device maps the original bus signal of its internal computing unit to the computing power optical interconnection communication channel, so as to encapsulate the original bus signal into a computing power optical interconnection protocol data frame in accordance with the CCOI transport layer protocol; the first device sends the computing power optical interconnection protocol data frame to the second device through the computing power optical interconnection communication channel, so as to establish the direct interconnection link between the first device and the second device.
5. The method according to claim 4, characterized in that, The original bus signals include memory bus extension signals, NPU inference data stream signals, and PCIe extension signals.
6. The method according to claim 2, characterized in that, Step 3 includes: the first device extracts the memory bus extension signal from the computing power optical interconnection protocol data frame, and maps the memory space of the second device with which it has established the direct interconnection link to a remote memory node of the local non-unified memory access architecture (NUMA) based on the memory bus extension signal; and registers the hardware inference unit of the second device as the remote inference engine that is locally schedulable by the first device based on the remote memory node of the local non-unified memory access architecture (NUMA).
7. The method according to claim 2, characterized in that, Step 3 includes: the first device identifying all second devices with which it has established the direct interconnect link; forming a remote memory node queue based on the remote memory nodes of the local non-unified memory access architecture (NUMA) corresponding to all second devices, so that after registering the hardware inference unit of the second device as the remote inference engine that is locally schedulable by the first device, all second devices corresponding to all the remote inference engines that are registered locally by the first device form a cross-device high-speed computing power aggregation resource pool.
8. The method according to claim 1, characterized in that, Also includes: The first device invokes its MDCCP protocol layer to sense the link status of the direct interconnect link in real time, so as to determine the computing power scheduling mode of the first device for the second device based on the link status, and in the computing power scheduling mode, distributes the computing tasks to be processed by the first device to the cross-device high-speed computing power aggregation resource pool for consumption through the computing power optical interconnection communication channel.
9. The method according to claim 6, characterized in that, During the process of mapping the memory space of the second device to a remote memory node of the Local Non-Unified Memory Access Architecture (NUMA), the first device maintains memory address consistency between the first device and the second device through the CCOI transport layer protocol. This memory address consistency enables physical addressing of the memory space of the second device when the first device accesses the remote memory node.
10. A computing power optical interconnection interface system for computing devices, characterized in that, include: The first device and the second device are connected by a computing power optical interconnection communication channel; A direct interconnection link is established between the first device and the second device based on the computing power optical interconnection communication channel; Based on the direct interconnect link, the first device registers the hardware inference unit of the second device as a locally schedulable remote inference engine.