Audio frequency integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission

The audio integrated signal processor with 40G/10G Ethernet and SRIO high-speed data transmission solves the resource waste and delay jitter problems of traditional underwater acoustic signal processors in heterogeneous networks, realizes low-overhead communication and dynamic path selection, improves transmission efficiency and GPU utilization, and supports multiple industry protocols.

CN120811989AActive Publication Date: 2025-10-17CHINA SHIP DEV & DESIGN CENT +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511277301.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-10-17
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Traditional underwater acoustic signal processors suffer from resource waste, delay jitter, low cross-board communication efficiency, and low GPU algorithm execution efficiency in multi-container communication, especially the lack of dynamic adaptation mechanism in heterogeneous networks.

Method used

The audio integrated signal processor adopts 40G/10G Ethernet and SRIO high-speed data transmission, and realizes low-overhead communication, dynamic path selection, container-level resource isolation and GPU virtualization collaborative optimization through signal classification module, heterogeneous transmission module, intelligent routing module, container resource scheduling module and heterogeneous computing offloading module.

Benefits of technology

It implements differentiated transmission strategies for different types of data streams, improves transmission efficiency and resource utilization, reduces end-to-end latency, increases GPU utilization, supports multiple industry-specific protocols, and expands application boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811989A_ABST
    Figure CN120811989A_ABST
Patent Text Reader

Abstract

The invention discloses an audio frequency integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission, and belongs to the technical field of audio frequency integrated signal processing. A signal classification module dynamically divides an audio frequency signal flow into a high-speed signal flow and a low-speed signal flow; the heterogeneous transmission module comprises a 40G Ethernet straight-through channel, a 10G Ethernet channel and an SRIO (Serial Radio Input / Output) board interconnection channel; the intelligent routing module is used for connecting the signal classification module with the heterogeneous transmission module, and the container resource scheduling module is used for carrying out GPU video memory dynamic partitioning and computing resource allocation on multi-container parallel tasks; the heterogeneous computing unloading module is used for generating a kernel-level scheduling instruction; and the closed-loop optimization module is used for acquiring network delay, jitter and GPU utilization rate data in real time and generating an optimization strategy packet containing a channel switching instruction and a resource redistribution instruction. According to the method, low-overhead communication, dynamic path selection, container-level resource isolation and GPU virtualization collaborative optimization are realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of audio integrated signal processing, and particularly relates to an audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission. BACKGROUND

[0002] The water acoustic signal processing technology is an important means to solve the underwater detection problem. The water acoustic signal processing technology completes target detection, target parameter estimation and target identification in a strong interference background by studying the sound propagation characteristics in the ocean and the acoustic characteristics of the target in reverberation and noise.

[0003] The traditional water acoustic signal processor has the problems of protocol redundancy, resource waste and delay jitter caused by single transmission path in multi-container communication; the communication efficiency between cross-boards and cross-cabin is limited by the fixed transmission mode and lacks a dynamic adaptation mechanism; the processing performance of multi-container parallel tasks is reduced due to resource competition; and the algorithm execution efficiency is low due to the conflict of video memory allocation and the insufficient scheduling strategy of GPU in the container environment.

[0004] Therefore, it is necessary to provide an audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission to solve the above problems. SUMMARY

[0005] Therefore, the purpose of the present application is to provide an audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission to solve the problem of how to realize low-overhead communication, dynamic path selection, container-level resource isolation and GPU virtualization collaborative optimization in a heterogeneous network.

[0006] To achieve the above purpose, the present application provides the following technical scheme: The present application provides an audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission, comprising: A signal classification module dynamically divides the audio signal stream into a fast signal stream and a slow signal stream through a preset data packet length threshold and a protocol whitelist; A heterogeneous transmission module, comprising: a 40G Ethernet direct channel realized through SR-IOV virtualization technology, which is configured with a hardware virtualization unit of a single I / O virtualization interface; a 10G Ethernet channel adopting a priority weighted round-robin mechanism, which is configured with at least three bandwidth isolation queues; and a SRIO inter-board interconnection channel constructed based on the RapidIO protocol, which is integrated with a DMA controller and an address mapping table; An intelligent routing module for connecting the signal classification module and the heterogeneous transmission module, comprising: a routing decision unit matching the corresponding transmission channel according to the signal stream type; and a dynamic rerouting unit triggering channel switching in response to transmission quality index changes; Container resource scheduling module, used to dynamically partition GPU memory and allocate computing resources for multi-container parallel tasks; The heterogeneous computing offloading module includes: an underwater acoustic algorithm parser that decomposes the acoustic processing algorithm into a set of microtasks with data dependencies; a GPU task scheduler that generates kernel-level scheduling instructions based on the computational density of microtasks; The closed-loop optimization module collects network latency, jitter, and GPU utilization data in real time, and generates an optimization strategy package containing channel switching instructions and resource reallocation instructions by connecting the feedback controllers of each module.

[0007] Furthermore, the dynamic classification rules of the signal classification module are: When the packet length Dynamic Threshold , or the protocol type is Sonar Raw Data Protocol; based on the DSCP marking unit, the fast signal flow is marked as EF class and mapped to the strict priority queue of the 40G Ethernet pass-through channel; When the packet length Dynamic Threshold , or the protocol type is parameter configuration instruction or log return protocol; based on the DSCP marking unit, the slow signal flow is marked as BE class and mapped to the weighted fair queue of the 10G Ethernet channel; Among them, the dynamic threshold Dynamically correct based on the average network load and GPU utilization in the previous 60 seconds: Where, is the classification threshold after dynamic adjustment, is the basic classification threshold, , is the network load weight coefficient, is the current network load, is the maximum theoretical bandwidth of the network, is the GPU utilization weight coefficient, is the current GPU utilization, is the maximum theoretical utilization of the GPU; If the same container sends more than 50 rapid signals continuously within 1 second, the traffic shaper will trigger rate limiting; the packet with conflicting markings will be forcibly reclassified and an alarm log will be generated; When the instantaneous packet loss rate of the 40G Ethernet direct channel is 5%, stop dynamic threshold correction and execute 256KB; when the packet loss rate 5% and lasts for 10 seconds, then the dynamic threshold correction is restored.

[0008] Furthermore, the method for triggering channel switching by the dynamic routing unit includes the following steps: A1: In the preset cycle Calculate the comprehensive score of the current transmission quality of all current channels : Among them, the preset cycle ; Where, is the delay weight coefficient, is the packet loss rate weight coefficient, is the bandwidth utilization weight coefficient, satisfying , is the current transmission delay, The maximum delay threshold allowed for the protocol, is the current packet loss rate, is the measured effective bandwidth, is the theoretical maximum bandwidth of the channel, is the average round trip time; A2: Select channels that meet the following conditions as candidate channels: and ; Where, Score the quality of candidate channels, Provides a real-time quality score for the channel currently being used. is the hysteresis threshold, is the absolute lower mass limit; A3: Calculate the difference in the comprehensive transmission quality score between the current channel and the candidate channels ;like Exceeding the quality difference threshold , the switching decision is triggered according to the following rules: in, Where, =0.7, is the quality difference threshold, is the slope coefficient, is the quality critical value of emergency switching, is the minimum quality improvement threshold during emergency switching. is the switching probability, The minimum probability threshold for triggering the actual switching action.

[0009] Furthermore, in step A3, during the switching process, a transition period is set. To synchronize data: And the sequence number is dynamically adjusted by using the sequence number remapping (SN Remapping) technology to compensate for the transmission delay or data packet buffering caused by switching: In the formula, is the updated sequence number, is the original sequence number before switching, is the current time, is the switching trigger time, is the number of buffered data packets per cycle, is the current round-trip time; After switching is completed, a cooling period is set to prohibit switching again, wherein ; If the candidate channel appears in the transition period , it indicates that the transmission quality is deteriorating, and it is immediately returned to the original channel, and the channel that triggers continuous switching failures is marked as a faulty channel and is shielded from switching for a duration ; When the actual number of switches exceeds the maximum allowed number of switches, the hysteresis threshold is automatically raised , wherein the maximum allowed number of switches is 5 per minute.

[0010] Further, the GPU memory dynamic partitioning step includes: B1: Establish a memory access heat evaluation unit to generate a memory block heat value by monitoring the memory access frequency of the container task within a unit time window and the single access residence duration , In the formula, is the heat value of the jth memory block, and are weighting coefficients satisfying ; B2: Construct a priority weighted heat mapping table to dynamically sort the memory blocks according to the formula and update the mapping relationship between the memory blocks and the containers according to the real-time task queue state, wherein is the container task priority weight factor, is the priority weighted heat value of the kth memory block; B3: Use a memory partitioning dynamic reconstruction mechanism to apply memory when a high-priority task is detected, The continuous GPU memory space occupied by the low-priority task in descending order is released, and a three-tuple allocation record of <task ID, start address, block size> is established.

[0011] Further, the hybrid scheduler of the computing resource allocation module operates according to the following rules: C1: Set the dynamic time slice length based on the task priority , wherein is the reference time slice; C2: Implement the preemption decision when the clock interrupt is triggered, and if there is a task request satisfying , suspend the current container context immediately and load the high-priority task pipeline; C3: Perform weighted round-robin scheduling for tasks of the same priority, and automatically adjust the time quota of the next scheduling period according to the proportion of the time slice used by each container ; , wherein is the priority weight factor of the newly requested task, is the priority weight factor of the currently executing task, is the dynamic threshold of the preemption decision, is the length of the time slice actually used by the container, is the total time slice quota allocated to the container within the scheduling period.

[0012] Further, the processing method of the heterogeneous computing offloading module includes: D1: Based on the data flow dependency relationship of the acoustic algorithm, construct a weighted directed acyclic graph DAG: , wherein is the weighted directed acyclic graph DAG, is the node set in the graph, is the edge set in the graph, is the computation density, is the computation density of the node , is the total number of computation operations of the microtask corresponding to the node, is the data volume processed by the microtask; D2: Generate a scheduling priority based on the computation density: , wherein is the dynamic weight coefficient, is the longest dependency path depth of the node in the DAG; D3: The GPU task scheduler arranges the microtasks in descending order , and generates a GPU kernel scheduling instruction when is satisfied, otherwise, it is assigned to the CPU thread pool; wherein, is a density amplification factor, is the sum of all node densities in the DAG, is the total number of nodes in the DAG.

[0013] Further, the resource reallocation instruction is used to adjust the GPU-CPU task allocation ratio: wherein, is the GPU task proportion after dynamic adjustment, is the current GPU task proportion, is the percentage of GPU utilization, is the benchmark value of GPU utilization.

[0014] The beneficial effects of the present application are: 1. Cross-layer optimization of transmission efficiency, through the dynamic threshold division of the signal classification module and the protocol awareness mechanism of intelligent routing, the differentiated transmission strategy (40G straight-through / 10G weighted polling / SRIO interconnection) of different types of data flow is realized, the end-to-end delay of real-time audio flow is reduced by 30%-50%, and the bandwidth utilization rate of file transmission is improved to more than 85% of the theoretical value; this transmission path selection mechanism based on data characteristics breaks through the traditional fixed channel allocation mode.

[0015] 2. Heterogeneous integration of resource virtualization, innovatively integrating SR-IOV virtualization technology and RapidIO hardware acceleration protocol, realizing the cooperative work of network function virtualization (NFV) and hardware acceleration unit at the single board level, the measured data shows that this architecture reduces the virtualization overhead from 15-20% of the traditional scheme to less than 5%, while maintaining the 40G line speed forwarding capability.

[0016] 3. Spatiotemporal reuse of computing resources, the container resource scheduling module realizes the dynamic partitioning of memory resources among multiple container tasks (granularity up to 256MB) through GPU memory heat perception algorithm and hybrid scheduling mechanism, combined with time slice rotation and priority preemption mechanism, the GPU utilization rate is improved from 40% in single task mode to 75% in multi-task mode.

[0017] 4. Microkernel acceleration of algorithm offloading, the heterogeneous computing offloading module adopts data flow driven microtask decomposition technology, converts traditional batch processing tasks into microkernel instruction set with parallelism up to 128 threads through the underwater acoustic algorithm parser, and cooperates with GPU three-level cache perception scheduling, so that the single frame processing time of typical acoustic processing algorithms is shortened to milliseconds.

[0018] 5、Transmission reliability dynamic guarantee, the closed loop optimization module constructs multi-dimensional QoS index system (including delay jitter, error rate, GPU display fragmentation rate, etc.), through feedback controller realizes sub-second level switching (switching time <200ms) of transmission channel, guarantees service level agreement (SLA) of key service when 40G / 10G link quality deteriorates.

[0019] 6、System architecture field adaptability, through the protocol whitelist mechanism and the reconfigurable design of DMA address mapping table, the device can be compatible with the protocol requirements of special application scenarios such as sonar array processing and underwater communication modulation, and the actual measurement supports the parallel processing capacity of not less than 8 kinds of industry special protocols, which significantly expands the application boundary of traditional audio processing equipment.

[0020] Other advantages, objects and features of the present application will be set forth in the following specification, and in part will become apparent to those skilled in the art from the present application, or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and attained by the below specification. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to make the purpose, technical scheme and beneficial effects of the present application more clear, the present application provides the following drawings for illustration: Figure 1 The framework diagram of the processing machine of the present application. DETAILED DESCRIPTION

[0022] As Figure 1 shown, the present application provides an audio integrated signal processing machine for 40G / 10G Ethernet and SRIO high-speed data transmission, comprising: A signal classification module dynamically divides the audio signal stream into a fast signal stream and a slow signal stream through a preset data packet length threshold and a protocol whitelist, wherein the fast signal stream is a data packet with a data packet length Dynamic threshold or a protocol type of real-time transmission protocol, and the slow signal stream is a data packet with a data packet length Dynamic threshold or a protocol type of file transfer protocol. A heterogeneous transmission module, comprising: a 40G Ethernet direct channel realized through SR-IOV virtualization technology, which is configured with a hardware virtualization unit of a single I / O virtualization interface; a 10G Ethernet channel adopting a priority weighted round-robin mechanism, which is configured with at least three bandwidth isolation queues; an SRIO inter-board interconnection channel constructed based on the RapidIO protocol, which is integrated with a DMA controller and an address mapping table. The intelligent routing module is used to connect the signal classification module and the heterogeneous transmission module. It includes: a routing decision unit that matches the corresponding transmission channel according to the signal flow type. The fast signal flow is bound to the 40G Ethernet pass-through channel, the slow signal flow is bound to the bandwidth isolation queue of the 10G Ethernet channel, and the communication between the on-board containers is bound to the SRIO channel; a dynamic rerouting unit that triggers channel switching in response to changes in transmission quality indicators; The container resource scheduling module is used to dynamically partition GPU memory and allocate computing resources for multi-container parallel tasks. It includes: a GPU partitioning module based on memory access popularity, which is configured with a memory mapping table that is dynamically adjusted according to the priority of container tasks; and a computing resource allocation module that uses a hybrid scheduler that combines time slicing and preemptive scheduling. The heterogeneous computing offloading module includes: an underwater acoustic algorithm parser that decomposes the acoustic processing algorithm into a set of microtasks with data dependencies; a GPU task scheduler that generates kernel-level scheduling instructions based on the computational density of microtasks; The closed-loop optimization module includes: a delay monitoring probe deployed on the network interface card, a utilization sampler configured on the GPU computing unit, and a feedback controller connected to each module, which generates an optimization strategy package containing channel switching instructions and resource reallocation instructions.

[0023] The scheme realizes the differentiated transmission strategy (40G straight-through / 10G weighted round robin / SRIO interconnection) of different types of data flow through the dynamic threshold division of the signal classification module and the protocol awareness mechanism of intelligent routing, innovatively combines the SR-IOV virtualization technology and the RapidIO hardware acceleration protocol, and realizes the cooperative work of the network function virtualization (NFV) and the hardware acceleration unit at the single board level. The container resource scheduling module realizes the dynamic partitioning (the granularity can reach 256 MB) of the GPU display memory resource among the multiple container tasks through the GPU display memory heat perception algorithm and the hybrid scheduling mechanism, combines the time slice rotation and the priority preemption mechanism, so that the GPU utilization rate is improved from 40% in the single task mode to 75% in the multi-task mode, the heterogeneous computing offloading module adopts the micro-task decomposition technology driven by the data flow, converts the traditional batch processing task into the micro-kernel instruction set with the parallelism up to 128 threads through the underwater acoustic algorithm parser, cooperates with the GPU three-level cache perception scheduling, so that the single frame processing time of the typical acoustic processing algorithm is shortened to the millisecond level; the closed-loop optimization module constructs the multi-dimensional QoS index system (including the delay jitter, the bit error rate, the GPU display memory fragmentation rate and the like), realizes the sub-second switching (the switching time is less than 200 ms) of the transmission channel through the feedback controller, and guarantees the service level agreement (SLA) of the key service when the 40G / 10G link quality is degraded; through the protocol whitelist mechanism and the reconfigurable design of the DMA address mapping table, the device can be compatible with the protocol requirements of special application scenarios such as the sonar array processing and the underwater communication modulation, and the measured parallel processing capacity of not less than 8 kinds of industry special protocols is supported, which significantly expands the application boundary of the traditional audio processing device.

[0024] In an embodiment of the present application, the dynamic classification rule of the signal classification module is: The fast signal determination condition is: the data packet length The dynamic threshold , or the protocol type is the sonar original data protocol; the fast signal flow is marked as the EF class based on the DSCP marking unit, and is mapped to the strict priority queue of the 40G Ethernet straight-through channel, The slow signal determination condition is: the data packet length The dynamic threshold , or the protocol type is the parameter configuration instruction or the log back transmission protocol; the slow signal flow is marked as the BE class based on the DSCP marking unit, and is mapped to the weighted fair queue of the 10G Ethernet channel, If more than 50 fast signal flows are continuously sent by the same container within 1 second, the flow shaper is triggered to limit the speed; the marked conflict data packet (such as DSCP=46 but the protocol type is the log back transmission) is forced to be reclassified and the alarm log is generated; The dynamic threshold The dynamic threshold is dynamically corrected according to the network average load and the GPU utilization rate within the previous 60 s: Where, is the classification threshold after dynamic adjustment, is the basic classification threshold, , is the network load weight coefficient, is the current network load, is the maximum theoretical bandwidth of the network, is the GPU utilization weight coefficient, is the current GPU utilization, is the maximum theoretical utilization of the GPU; And meet the emergency load reduction rules: When the instantaneous packet loss rate of the 40G Ethernet direct channel is 5%, stop dynamic threshold correction and execute 256KB, so that fewer data packets are judged as fast signal flow, thereby reducing the load of the 40G Ethernet pass-through channel and giving priority to high-priority traffic; when the packet loss rate 5% and lasts for 10 seconds, then the dynamic threshold correction is restored.

[0025] This solution adjusts classification thresholds based on load and GPU utilization feedback to avoid network congestion or resource waste caused by static rules. Abnormal traffic detection prevents malicious or incorrectly marked packets from occupying high-priority channels. A forced reclassification mechanism ensures the reliability of classification results, and DSCP marking is directly linked to the channel binding policy of the communication selection module. Dynamic threshold data is fed back to the performance monitoring module to form a closed-loop control.

[0026] In one embodiment of the present invention, the step of triggering channel switching by the dynamic rerouting unit includes: A1: In the preset cycle Calculate the comprehensive score of the current transmission quality of all current channels : Among them, the preset cycle , to dynamically adapt to the speed of network changes; A2: Select channels that meet the following conditions as candidate channels: and ; A3: Calculate the difference in the comprehensive transmission quality score between the current channel and the candidate channels ;like Exceeding the quality difference threshold , the decision is triggered according to the following rules: in, , =0.7, is a quality difference threshold, is a slope coefficient, , , wherein, during the switching process, a transition period is set data synchronization is performed: and a sequence number remapping (SN Remapping) technique is used to dynamically adjust the sequence number to compensate for transmission delays or data packet buffering caused by switching, ensuring that the receiving end can reassemble the data stream in the correct order: wherein, is an updated sequence number, is the original sequence number before switching, is the current time, is the switching trigger time, is the number of buffered data packets per cycle; after the switching is completed, a cooling period is set to prohibit further switching, wherein ; wherein, if the candidate channel appears within the transition period, it indicates that the transmission quality has deteriorated, and the original channel is immediately returned to, and the channel that triggers consecutive switching failures is marked as a faulty channel and is shielded from switching for a certain period of time. .

[0027] wherein, when the actual number of switches exceeds the maximum allowed number of switches, the system determines that the current switching strategy is too aggressive and automatically raises the hysteresis threshold wherein, the maximum allowed number of switches is 5 per minute, so that subsequent switching needs to meet more stringent conditions (such as a larger difference in signal quality, lower delay, etc.), thereby reducing the switching frequency.

[0028] wherein, is a delay weight coefficient, is a packet loss rate weight coefficient, is a bandwidth utilization rate weight coefficient, is the current transmission delay, is the maximum delay threshold allowed by the protocol, is the current packet loss rate, is the measured effective bandwidth, is the theoretical maximum bandwidth of the channel, wherein: ; is the real-time quality score of the channel currently being used, is the quality score of the candidate channel, is a hysteresis threshold, is an absolute quality lower limit; is a quality critical value of emergency handover, is a minimum quality improvement threshold when emergency handover, is a handover probability, is a minimum probability threshold of actual handover action triggering, is an average round trip time, is a current round trip time.

[0029] In the present scheme, by , , a hierarchical response is realized, and the transition period and the cooling period balance the handover gain and the stability loss.

[0030] In an embodiment of the present application, in step A3, during the handover process, a transition period is set to perform data synchronization: And the sequence number remapping (SN Remapping) technology is used to dynamically adjust the sequence number to compensate for the transmission delay or data packet buffering caused by handover: In the formula, is an updated sequence number, is an original sequence number before handover, is a current time, is a handover triggering time, is the number of buffered data packets per period, is a current round trip time; After the handover is completed, a cooling period is set to prohibit handover again, wherein ; If the candidate channel appears in the transition period, it indicates that the transmission quality is deteriorating, and then immediately returns to the original channel, and the channel that triggers consecutive handover failures is marked as a faulty channel and is shielded from handover for a duration ; When the actual number of handovers exceeds the maximum allowed number of handovers, the hysteresis threshold is automatically raised , wherein the maximum allowed number of handovers is 5 per minute.

[0031] In an embodiment of the present application, the GPU memory dynamic partitioning step includes: B1: Establish a memory access heat evaluation unit to monitor the memory access frequency of the container task in a unit time window and single access residence duration , generating a memory block hotness value , wherein, is the hotness value of the jth memory block, and is a weighting coefficient, satisfying ; B2: constructing a priority-weighted hotness mapping table, dynamically sorting the memory blocks according to the formula , and updating the mapping relationship between the memory blocks and the containers according to the real-time task queue state, wherein is a container task priority weight factor, is the priority-weighted hotness value of the kth memory block; B3: using a dynamic memory partitioning reconstruction mechanism, when a high-priority task is detected to apply for memory, releasing the continuous memory space occupied by low-priority tasks in descending order according to , and establishing a three-tuple allocation record of <task ID, start address, block size>.

[0032] In the present scheme, cold and hot data are separated and stored by hotness value quantification, the weighted sorting formula ensures that the memory demand of high-priority tasks is quickly responded, and the three-tuple record constrains the allocation of continuous space, thereby reducing the memory fragmentation rate.

[0033] In an embodiment of the present application, the hybrid scheduler of the computing resource allocation module operates according to the following rules: C1: setting a dynamic time slice length based on task priority , wherein is a reference time slice; C2: implementing a preemption decision when a clock interrupt is triggered, if there is a task request satisfying , then immediately suspending the current container context and loading the high-priority task pipeline; C3: performing weighted round-robin scheduling for tasks of the same priority, automatically adjusting the time quota of the next scheduling period according to the proportion of the used time slice of each container ; wherein, is the priority weight factor of the new requested task, is the priority weight factor of the task being executed, is a dynamic threshold value of the preemption decision, is the length of the time slice actually occupied by the container, is the total time slice quota allocated to the container in the scheduling period.

[0034] In one embodiment of the present invention, a method for processing a heterogeneous computing offload module includes: D1: Construct a weighted directed acyclic graph (DAG) based on the data flow dependency of the acoustic algorithm: Where, is a weighted directed acyclic graph DAG, is the set of nodes in the graph, is the edge set in the graph, To calculate the density, For nodes The calculation density of is the total number of computational operations of the microtask corresponding to the node, The amount of data processed for the microtask; D2: Generate scheduling priority based on computing density: Where, is the dynamic weight coefficient, The longest dependency path depth of the node in the DAG; D3: GPU task scheduler Arrange the execution of microtasks in descending order, when the Generate GPU kernel scheduling instructions when it is available, otherwise assign it to the CPU thread pool; Where, is the density magnification factor, is the sum of all node densities in the DAG, is the total number of nodes in the DAG.

[0035] In this program, through Accurately distinguish between compute-intensive (GPU) and memory-intensive (CPU) tasks, and introduce them into scheduling priorities. , shorten the execution delay of the DAG critical path, threshold Prevent inefficient GPU tasks from occupying streaming multiprocessor (SM) resources.

[0036] In one embodiment of the present invention, the resource reallocation instruction is used to adjust the GPU-CPU task allocation ratio: Where, To dynamically adjust the proportion of GPU tasks, is the current GPU task ratio, is the percentage of GPU utilization, It is the baseline value of GPU utilization.

[0037] The scheme guarantees that the GPU utilization rate approaches a benchmark value through an elastic resource scheduling formula, and avoids overload or idling.

[0038] Finally, it should be explained that the above preferred embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made in form and details without departing from the scope defined by the claims of the present application.

Claims

1. An audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission, characterized in that: include: The signal classification module dynamically divides the audio signal flow into fast signal flow and slow signal flow through the preset packet length threshold and protocol whitelist; Heterogeneous transport modules include: a 40G Ethernet pass-through channel implemented through SR-IOV virtualization technology, which is equipped with a hardware virtualization unit for a single I / O virtualization interface; a 10G Ethernet channel using a priority-weighted round-robin mechanism, which is equipped with at least three bandwidth isolation queues; and an SRIO inter-board interconnect channel built based on the RapidIO protocol, which integrates a DMA controller and address mapping table. The intelligent routing module is used to connect the signal classification module and the heterogeneous transmission module, including: a routing decision unit that matches the corresponding transmission channel according to the signal flow type; a dynamic rerouting unit that triggers channel switching in response to changes in transmission quality indicators; Container resource scheduling module, used to dynamically partition GPU memory and allocate computing resources for multi-container parallel tasks; The heterogeneous computing offloading module includes: an underwater acoustic algorithm parser that decomposes the acoustic processing algorithm into a set of microtasks with data dependencies; a GPU task scheduler that generates kernel-level scheduling instructions based on the computational density of microtasks; The closed-loop optimization module collects network latency, jitter, and GPU utilization data in real time, and generates an optimization strategy package containing channel switching instructions and resource reallocation instructions by connecting the feedback controllers of each module.

2. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission according to claim 1, characterized in that: The dynamic classification rules of the signal classification module are: When the packet length Dynamic Threshold , or the protocol type is Sonar Raw Data Protocol; based on the DSCP marking unit, the fast signal flow is marked as EF class and mapped to the strict priority queue of the 40G Ethernet pass-through channel; When the packet length Dynamic Threshold , or the protocol type is parameter configuration instruction or log return protocol; based on the DSCP marking unit, the slow signal flow is marked as BE class and mapped to the weighted fair queue of the 10G Ethernet channel; Among them, the dynamic threshold Dynamically correct based on the average network load and GPU utilization in the previous 60 seconds: Where, is the classification threshold after dynamic adjustment, is the basic classification threshold, , is the network load weight coefficient, is the current network load, is the maximum theoretical bandwidth of the network, is the GPU utilization weight coefficient, is the current GPU utilization, is the maximum theoretical utilization of the GPU; If the same container sends more than 50 rapid signals continuously within 1 second, the traffic shaper will trigger rate limiting; the packet with conflicting markings will be forcibly reclassified and an alarm log will be generated; When the instantaneous packet loss rate of the 40G Ethernet direct channel is 5%, stop dynamic threshold correction and execute 256KB; when the packet loss rate 5% and lasts for 10 seconds, then the dynamic threshold correction is restored.

3. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission according to claim 1, characterized in that: The method for triggering channel switching by a dynamic routing unit includes the following steps: A1: In the preset cycle Calculate the comprehensive score of the current transmission quality of all current channels : Among them, the preset cycle ; Where, is the delay weight coefficient, is the packet loss rate weight coefficient, is the bandwidth utilization weight coefficient, satisfying , is the current transmission delay, The maximum delay threshold allowed for the protocol, is the current packet loss rate, is the measured effective bandwidth, is the theoretical maximum bandwidth of the channel, is the average round trip time; A2: Select channels that meet the following conditions as candidate channels: and ; Where, Score the quality of candidate channels, Score the real-time quality of the channel currently being used. is the hysteresis threshold, is the absolute lower mass limit; A3: Calculate the difference in the comprehensive transmission quality score between the current channel and the candidate channels ;like Exceeding the quality difference threshold , the switching decision is triggered according to the following rules: in, Where, =0.7, is the quality difference threshold, is the slope coefficient, is the quality critical value of emergency switching, is the minimum quality improvement threshold during emergency switching. is the switching probability, The minimum probability threshold for triggering the actual switching action.

4. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission according to claim 3, characterized in that: In step A3, during the switching process, a transition period is set To synchronize data: Sequence number remapping technology is used to dynamically adjust sequence numbers to compensate for transmission delays or packet buffering caused by switching: Where, is the updated serial number, is the original serial number before switching, is the current time, is the switching trigger time, is the number of packets buffered per cycle, is the current round trip time; Set a cool-down period after the switch is complete To prohibit further switching, ,in ; Among them, if the candidate channel appears during the transition period , it means that the transmission quality has deteriorated, then it will immediately return to the original channel. The channel that fails to switch is marked as a faulty channel and the switching duration is blocked. ; Among them, when the actual number of switching times exceeds the maximum allowed number of switching times, the hysteresis threshold is automatically increased , where the maximum allowed switching times is 5 times per minute.

5. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission according to claim 1, characterized in that: The steps for dynamic partitioning of GPU memory include: B1: Establish a memory access heat evaluation unit to monitor the memory access frequency of container tasks within a unit time window. and single visit residence time , generate the memory block heat value , Where, is the heat value of the j-th video memory block, and is the weighting coefficient, satisfying ; B2: Construct a priority weighted heat map and map the memory blocks according to the formula Dynamic sorting, updating the mapping relationship between memory blocks and containers according to the real-time task queue status, where is the container task priority weight factor, is the priority weighted heat value of the kth video memory block; B3: Use the dynamic reconstruction mechanism of video memory partitions to detect high priority tasks. When applying for video memory, press Release the continuous video memory space occupied by low-priority tasks in descending order and create a triplet allocation record of <task ID, starting address, block size>.

6. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission according to claim 1, characterized in that: The hybrid scheduler of the computing resource allocation module operates according to the following rules: C1: Set dynamic time slice length based on task priority ,in is the base time slice; C2: Implement preemption decision when clock interrupt is triggered, if there is a condition that satisfies If a task request is received, the current container context is immediately suspended and the high-priority task pipeline is loaded; C3: Perform weighted round-robin scheduling for tasks of the same priority level, based on the proportion of time slices used by each container. Automatically adjust the time quota for the next scheduling cycle; Where, is the priority weight factor of the new request task, is the priority weight factor of the currently executing task, is the dynamic threshold for preemption decision, The actual length of the time slice occupied by the container. The total time slice quota allocated to the container during the scheduling period.

7. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission according to claim 1, characterized in that: The processing methods for heterogeneous computing offload modules include: D1: Construct a weighted directed acyclic graph (DAG) based on the data flow dependency of the acoustic algorithm: Where, is a weighted directed acyclic graph DAG, is the set of nodes in the graph, is the edge set in the graph, To calculate the density, For nodes The calculation density of is the total number of computational operations of the microtask corresponding to the node, The amount of data processed for the microtask; D2: Generate scheduling priority based on computing density: Where, is the dynamic weight coefficient, The longest dependency path depth of the node in the DAG; D3: GPU task scheduler Execute microtasks in descending order, when satisfied Generate GPU kernel scheduling instructions when it is available, otherwise assign it to the CPU thread pool; Where, is the density magnification factor, is the sum of all node densities in the DAG, is the total number of nodes in the DAG.

8. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission according to claim 1, characterized in that: Resource reallocation instructions are used to adjust the GPU-CPU task allocation ratio: Where, To dynamically adjust the proportion of GPU tasks, is the current GPU task ratio, is the percentage of GPU utilization, It is the baseline value of GPU utilization.

Citation Information

Patent Citations

  • Multi-service integration dual-redundancy network system

    CN102984057A

  • Multi-channel collaborative acceleration method for data processing and transmission of set top box

    CN119172571A

  • Self-adaptive task scheduling execution unit management method and system

    CN119376903A

  • Cross-platform communication sharing synchronization method and system

    CN119946076A

  • Marine audio frequency integrated resource scheduling configuration management system

    CN120353587A