An audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission

The integrated audio signal processor, which uses 40G/10G Ethernet and SRIO high-speed data transmission, solves the problems of resource waste and latency jitter in multi-container communication of traditional underwater acoustic signal processors, and achieves efficient data transmission and GPU resource utilization, supporting multiple industry protocols.

CN120811989BActive Publication Date: 2025-12-26CHINA SHIP DEV & DESIGN CENT +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511277301.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-12-26
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Traditional underwater acoustic signal processors suffer from resource waste, latency jitter, low efficiency in cross-board communication, and low efficiency in GPU virtualization in multi-container communication, and lack a dynamic adaptation mechanism.

Method used

The audio integrated signal processor, which adopts 40G/10G Ethernet and SRIO high-speed data transmission, achieves low-overhead communication, dynamic path selection and GPU virtualization collaborative optimization through signal classification module, heterogeneous transmission module, intelligent routing module, container resource scheduling module and heterogeneous computing offloading module.

Benefits of technology

It implements a low-latency differentiated transmission strategy, improves transmission efficiency and resource utilization, enhances GPU utilization, supports multiple industry-specific protocols, and expands application boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811989B_ABST
    Figure CN120811989B_ABST
Patent Text Reader

Abstract

This invention discloses an integrated audio signal processor for high-speed data transmission via 40G / 10G Ethernet and SRIO, belonging to the field of integrated audio signal processing technology. It includes a signal classification module that dynamically divides the audio signal stream into fast and slow signal streams; a heterogeneous transmission module comprising a 40G Ethernet pass-through channel, a 10G Ethernet channel, and an SRIO board interconnection channel; an intelligent routing module connecting the signal classification module and the heterogeneous transmission module; a container resource scheduling module for dynamically partitioning GPU memory and allocating computing resources for multi-container parallel tasks; a heterogeneous computing offloading module for generating kernel-level scheduling instructions; and a closed-loop optimization module that collects network latency, jitter, and GPU utilization data in real time and generates an optimization strategy package containing channel switching instructions and resource reallocation instructions. This invention achieves low-overhead communication, dynamic path selection, container-level resource isolation, and GPU virtualization collaborative optimization.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of audio integrated signal processing, and particularly relates to an audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission. BACKGROUND

[0002] Water acoustic signal processing technology is an important means to solve the problem of underwater detection. Through the study of the characteristics of sound propagation in the ocean and the acoustic characteristics of the target in reverberation and noise, water acoustic signal processing technology can complete target detection, target parameter estimation and target recognition in strong interference background.

[0003] The traditional water acoustic signal processor has the problems of protocol redundancy, resource waste and delay jitter caused by single transmission path in multi-container communication; the communication efficiency between cross-board cards and cross-chassis is limited by the fixed transmission mode and lacks a dynamic adaptation mechanism; the processing performance of multi-container parallel tasks is reduced due to resource competition; and the algorithm execution efficiency is low due to the conflict of video memory allocation and the insufficient scheduling strategy of GPU in the container environment.

[0004] Therefore, it is necessary to provide an audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission to solve the above problems. SUMMARY

[0005] Therefore, the present application provides an audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission to solve the problem of how to realize low-overhead communication, dynamic path selection, container-level resource isolation and GPU virtualization collaborative optimization in a heterogeneous network.

[0006] To achieve the above-mentioned purpose, the present application provides the following technical scheme:

[0007] The present application provides an audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission, comprising:

[0008] A signal classification module dynamically divides the audio signal stream into a fast signal stream and a slow signal stream through a preset data packet length threshold and a protocol whitelist;

[0009] A heterogeneous transmission module, comprising: a 40G Ethernet direct channel realized through SR-IOV virtualization technology, which is configured with a hardware virtualization unit of a single I / O virtualization interface; a 10G Ethernet channel adopting a priority weighted round-robin mechanism, which is configured with at least three bandwidth isolation queues; and a SRIO inter-board interconnection channel constructed based on the RapidIO protocol, which is integrated with a DMA controller and an address mapping table;

[0010] The intelligent routing module is used for connecting the signal classification module and the heterogeneous transmission module, and comprises: a routing decision unit, which matches a corresponding transmission channel according to a signal flow type; and a dynamic rerouting unit, which triggers channel switching in response to transmission quality index changes.

[0011] The container resource scheduling module is used for dynamically partitioning GPU display memory and allocating computing resources for a multi-container parallel task.

[0012] The heterogeneous computing offloading module comprises: an underwater acoustic algorithm parser, which decomposes an acoustic processing algorithm into a micro-task set having a data dependency relationship; and a GPU task scheduler, which generates a kernel-level scheduling instruction according to the computing density of the micro-task.

[0013] The closed-loop optimization module collects network delay, jitter and GPU utilization rate data in real time, and generates an optimization strategy package containing channel switching instructions and resource reallocation instructions through a feedback controller connected with each module.

[0014] Further, the dynamic classification rule of the signal classification module is as follows:

[0015] When the data packet length is greater than the dynamic threshold value, or the protocol type is a sonar original data protocol; Dynamic threshold value The fast signal flow is marked as an EF class based on a DSCP marking unit and is mapped to a strict priority queue of a 40G Ethernet direct channel.

[0016] When the data packet length is greater than the dynamic threshold value, or the protocol type is a sonar original data protocol; Dynamic threshold value The slow signal flow is marked as a BE class based on a DSCP marking unit and is mapped to a weighted fair queue of a 10G Ethernet channel.

[0017] The dynamic threshold value is dynamically corrected according to the network average load and the GPU utilization rate in the previous 60 s.

[0018]

[0019] In the formula, the dynamic classification threshold value is the basic classification threshold value is , , the network load weight coefficient is the current network load is the maximum theoretical bandwidth of the network is the GPU utilization rate weight coefficient is the current GPU utilization rate is the maximum theoretical utilization rate of the GPU is

[0020] ​If the same container sends more than 50 fast signals in a single second, the traffic shaper rate limit is triggered; forced reclassification is performed on the marked conflict packets and an alarm log is generated.

[0021] When the instantaneous packet loss rate of a 40G Ethernet cut-through channel At 5%, stop dynamic threshold correction and execute. 256KB; when packet loss rate After setting the threshold to 5% and holding for 10 seconds, dynamic threshold correction resumes.

[0022] Furthermore, the method for triggering channel switching by the dynamic routing unit includes the following steps:

[0023] A1: Within the preset period Internally calculates the overall current transmission quality score for all current channels. :

[0024]

[0025] Among them, the preset period ;

[0026] In the formula, For delay weighting coefficients, This is the packet loss rate weighting coefficient. The bandwidth utilization weighting coefficient satisfies , For the current transmission delay, The maximum allowed latency threshold for the protocol. The current packet loss rate, To measure the effective bandwidth, This represents the theoretical maximum bandwidth of the channel. This represents the average round-trip time.

[0027] A2: Select channels that meet the following conditions as candidate channels: and ;

[0028] In the formula, Assess the quality of candidate channels. The real-time quality score for the currently used channel. The hysteresis threshold, This is the absolute lower limit of quality.

[0029] A3: Calculate the difference in overall transmission quality scores between the current channel and the candidate channels. ;like Exceeding the quality difference threshold Then, a switching decision will be triggered according to the following rules:

[0030]

[0031] wherein,

[0032]

[0033] wherein, =0.7, is a quality difference threshold, is a slope coefficient, is a quality threshold for emergency handover, is a minimum quality improvement threshold for emergency handover, is a handover probability, is a minimum probability threshold for actual handover action triggering.

[0034] Further, in step A3, during the handover process, a transition period is set Data synchronization is performed:

[0035]

[0036] And the sequence number is dynamically adjusted by using the sequence number remapping (SN Remapping) technology to compensate for the transmission delay or data packet buffering caused by handover:

[0037]

[0038] wherein, is the updated sequence number, is the original sequence number before handover, is the current time, is the handover triggering time, is the number of buffered data packets per period, is the current round-trip time;

[0039] After the handover is completed, a cooling period is set to prohibit further handover, wherein ;

[0040] If the candidate channel appears during the transition period, it indicates that the transmission quality has deteriorated, and the original channel is immediately returned to, and the channel that triggers consecutive handover failures is marked as a faulty channel and is shielded from handover for a duration ;

[0041] When the actual number of handovers exceeds the maximum allowed number of handovers, the hysteresis threshold is automatically increased , wherein the maximum allowed number of handovers is 5 per minute.

[0042] Further, the GPU memory dynamic partitioning step includes:

[0043] B1: Establishing a video memory access heat evaluation unit to monitor the video memory access frequency of the container task within a unit time window and the residence time of single access to generate a video memory block heat value ,

[0044]

[0045] In the formula, is the heat value of the jth video memory block, and is a weighting coefficient, satisfying ;

[0046] B2: Constructing a priority weighted heat mapping table, dynamically sorting the video memory blocks according to the formula and updating the mapping relationship between the video memory blocks and the containers according to the real-time task queue state, wherein is the container task priority weight factor, is the priority weighted heat value of the kth video memory block;

[0047] B3: Using a video memory partition dynamic reconstruction mechanism, when a high-priority task is detected to apply for video memory, releasing the continuous video memory space occupied by low-priority tasks in descending order and establishing a three-tuple allocation record of <task ID, start address, block size>.

[0048] Further, the hybrid scheduler of the computing resource allocation module operates according to the following rules:

[0049] C1: Setting a dynamic time slice length based on task priority , wherein is the reference time slice;

[0050] C2: Implementing a preemption decision when a clock interrupt is triggered, if there is a task request satisfying , then immediately suspend the current container context and load the high-priority task pipeline;

[0051] C3: Performing weighted round-robin scheduling for tasks of the same priority, automatically adjusting the time quota of the next scheduling period according to the time slice proportion used by each container ;

[0052] In the formula, is the priority weight factor of the new request task, is the priority weight factor of the task being executed, is the dynamic threshold of the preemption decision, is the time slice length actually occupied by the container, is the total time slice quota allocated to the container in the scheduling period.

[0053] Further, the processing method of the heterogeneous computing offloading module comprises:

[0054] D1: Based on the data flow dependency relationship of the acoustic algorithm, a weighted directed acyclic graph DAG is constructed:

[0055]

[0056] wherein, is a weighted directed acyclic graph DAG, is a node set in the graph, is an edge set in the graph, is a calculation density, is the calculation density of the node , is the total number of calculation operations of the micro task corresponding to the node, is the data volume processed by the micro task; D2: Generate a scheduling priority based on the calculation density:

[0057]

[0058]

[0059] wherein, is a dynamic weight coefficient, is the longest dependency path depth of the node in the DAG;

[0060] D3: The GPU task scheduler arranges the micro tasks in descending order of , generates a GPU kernel scheduling instruction when is satisfied, otherwise assigns to the CPU thread pool;

[0061] wherein, is a density amplification factor, is the sum of the densities of all nodes in the DAG, is the total number of nodes in the DAG.

[0062] Further, the resource reallocation instruction is used to adjust the GPU-CPU task allocation ratio:

[0063]

[0064] wherein, is the GPU task proportion after dynamic adjustment, is the current GPU task proportion, is the percentage of GPU utilization, is the benchmark value of GPU utilization.

[0065] ​The beneficial effects of the present application are:

[0066] 1. Cross-layer optimization of transmission efficiency, through dynamic threshold division of the signal classification module and protocol-aware mechanism of intelligent routing, different types of data flow are realized Differentiated transmission strategy (40G straight through / 10G weighted round robin / SRIO interconnection), the end-to-end delay of real-time audio stream is reduced by 30%-50%, and the bandwidth utilization of file transmission is increased to more than 85% of the theoretical value; This transmission path selection mechanism based on data characteristics breaks through the traditional fixed channel allocation mode.

[0067] 2. Heterogeneous integration of resource virtualization, innovatively combining SR-IOV virtualization technology and RapidIO hardware acceleration protocol, realizing the cooperative work of network function virtualization (NFV) and hardware acceleration unit at the single board level, the measured data shows that this architecture reduces the virtualization overhead from 15-20% of the traditional scheme to less than 5%, while maintaining 40G line speed forwarding capability.

[0068] 3. Spatiotemporal reuse of computing resources, the container resource scheduling module realizes dynamic partitioning of GPU memory resources (granularity up to 256MB) among multiple container tasks through GPU memory heat sensing algorithm and hybrid scheduling mechanism, combined with time slice rotation and priority preemption mechanism, GPU utilization is increased from 40% in single task mode to 75% in multi-task mode.

[0069] 4. Microkernel acceleration of algorithm offloading, the heterogeneous computing offloading module uses data flow driven microtask decomposition technology, converts traditional batch processing tasks into microkernel instruction sets with parallelism up to 128 threads through underwater acoustic algorithm parser, and cooperates with GPU three-level cache sensing scheduling, so that the single-frame processing time of typical acoustic processing algorithms is shortened to milliseconds.

[0070] 5. Dynamic guarantee of transmission reliability, the closed-loop optimization module constructs a multi-dimensional QoS index system (including delay jitter, bit error rate, GPU memory fragmentation rate, etc.), and realizes sub-second switching (switching time <200ms) of the transmission channel through the feedback controller, and guarantees the service level agreement (SLA) of key services when the 40G / 10G link quality deteriorates.

[0071] 6. Field adaptability of system architecture, through protocol whitelist mechanism and reconfigurable design of DMA address mapping table, the device can be compatible with the protocol requirements of special application scenarios such as sonar array processing and underwater communication modulation, and can support parallel processing capability of not less than 8 kinds of industry special protocols, which significantly expands the application boundary of traditional audio processing equipment.

[0072] Additional advantages, objects, and features of the application will be set forth in part by the description that follows, and in part will become apparent to those skilled in the art upon examination of the following or can be learned from practice of the application. The objects and other advantages of the application can be realized and attained by the structure particularly pointed out in the written description and claims hereof as well as the appended drawings. BRIEF DESCRIPTION OF DRAWINGS

[0073] To make the objectives, technical solutions and beneficial effects of the present application clearer, the present application is described below with the help of the accompanying drawings:

[0074] Figure 1 A frame diagram of the processing machine of the present application. DETAILED DESCRIPTION

[0075] As shown in Figure 1 , the present application provides an audio integrated signal processor for 40G / 10G Ethernet and SRIO high-speed data transmission, comprising:

[0076] a signal classification module, which dynamically divides the audio signal stream into a fast signal stream and a slow signal stream through a preset data packet length threshold and a protocol white list, wherein the fast signal stream is a data packet with a data packet length dynamic threshold or a protocol type of real-time transmission protocol, and the slow signal stream is a data packet with a data packet length dynamic threshold or a protocol type of file transfer protocol;

[0077] a heterogeneous transmission module, comprising: a 40G Ethernet direct channel realized through SR-IOV virtualization technology, which is configured with a hardware virtualization unit of a single I / O virtualization interface; a 10G Ethernet channel adopting a priority weighted round-robin mechanism, which is configured with at least three bandwidth isolation queues; an SRIO inter-board interconnection channel constructed based on a RapidIO protocol, which is integrated with a DMA controller and an address mapping table;

[0078] an intelligent routing module, used for connecting the signal classification module and the heterogeneous transmission module, comprising: a routing decision unit, which matches a corresponding transmission channel according to a signal stream type, wherein the fast signal stream is bound to the 40G Ethernet direct channel, the slow signal stream is bound to a bandwidth isolation queue of the 10G Ethernet channel, and inter-board communication is bound to the SRIO channel; and a dynamic rerouting unit, which triggers channel switching in response to transmission quality index changes;

[0079] The container resource scheduling module is used to dynamically partition GPU memory and allocate computing resources for multi-container parallel tasks. It includes: a GPU partitioning module based on memory access popularity, which is configured with a memory mapping table that is dynamically adjusted according to the priority of container tasks; and a computing resource allocation module that uses a hybrid scheduler that combines time-slice round-robin and preemptive scheduling.

[0080] The heterogeneous computing offloading module includes: an underwater acoustic algorithm resolver, which decomposes acoustic processing algorithms into a set of microtasks with data dependencies; and a GPU task scheduler, which generates kernel-level scheduling instructions based on the computational density of the microtasks.

[0081] The closed-loop optimization module includes: a latency monitoring probe deployed on the network interface card, a utilization sampler configured on the GPU computing unit, and a feedback controller connecting each module, which generates an optimization strategy package containing channel switching instructions and resource reallocation instructions.

[0082] This solution achieves differentiated transmission strategies (40G passthrough / 10G weighted round-robin / SRIO interconnect) for different types of data streams through dynamic threshold division of the signal classification module and protocol awareness mechanism of intelligent routing. It innovatively integrates SR-IOV virtualization technology and RapidIO hardware acceleration protocol to realize the collaborative work of network function virtualization (NFV) and hardware acceleration unit at the single board level. The container resource scheduling module realizes dynamic partitioning of memory resources among multiple container tasks (granularity up to 256MB) through GPU memory heat awareness algorithm and hybrid scheduling mechanism. Combined with time slice round-robin and priority preemption mechanism, it increases GPU utilization from 40% in single-task mode to 75% in multi-task mode. The heterogeneous computing offloading module adopts data flow-driven micro-task decomposition technology. Through the underwater acoustic algorithm parser, it transforms traditional batch processing tasks into a microkernel instruction set with a parallelism of up to 128 threads. With GPU L3 cache awareness scheduling, it reduces the single-frame processing time of typical acoustic processing algorithms to the millisecond level. The closed-loop optimization module constructs a multi-dimensional QoS indicator system (including latency jitter, bit error rate, GPU memory fragmentation rate, etc.), and realizes sub-second switching of transmission channels (switching time <200ms) through the feedback controller, ensuring the service level agreement (SLA) of critical services when the quality of 40G / 10G links deteriorates. Through the protocol whitelist mechanism and the reconfigurable design of the DMA address mapping table, the device is compatible with the protocol requirements of special application scenarios such as sonar array processing and underwater communication modulation. The actual test shows that it supports the parallel processing capability of no less than 8 industry-specific protocols, which significantly expands the application boundaries of traditional audio processing equipment.

[0083] In one embodiment of the present invention, the dynamic classification rule of the signal classification module is as follows:

[0084] The fast signal determination condition is: data packet length. Dynamic threshold , or the protocol type is the sonar original data protocol; the fast signal stream is marked as the EF class by the DSCP marking unit, and is mapped to the strict priority queue of the 40G Ethernet direct channel,

[0085] The slow signal determination condition is that the data packet length Dynamic threshold , or the protocol type is the parameter configuration instruction or the log return protocol; the slow signal stream is marked as the BE class by the DSCP marking unit, and is mapped to the weighted fair queue of the 10G Ethernet channel,

[0086] If more than 50 fast signal streams are continuously sent by the same container within 1 second, the flow shaper is triggered to limit the speed; the marked conflict data packet (for example, DSCP=46 but the protocol type is the log return) is forced to be reclassified, and an alarm log is generated;

[0087] The dynamic threshold is According to the network average load and the GPU utilization rate within the previous 60 seconds, the dynamic threshold is dynamically corrected:

[0088]

[0089] In the formula, The classified threshold after dynamic adjustment is The basic classified threshold is , The network load weight coefficient is The current network load is The maximum theoretical bandwidth of the network is The GPU utilization rate weight coefficient is The current GPU utilization rate is The maximum theoretical utilization rate of the GPU is

[0090] The emergency load reduction rule is met: when the instantaneous packet loss rate of the 40G Ethernet direct channel is 5%, the dynamic threshold correction is stopped, and 256KB is executed, so that fewer data packets are determined as fast signal streams, thereby reducing the load of the 40G Ethernet direct channel and preferentially guaranteeing high-priority traffic; when the packet loss rate is 5% for 10 seconds, the dynamic threshold correction is restored.

[0091] The scheme adjusts the classified threshold through load and GPU utilization rate feedback, avoids network congestion or resource waste caused by static rules, prevents malicious or incorrectly marked data packets from occupying high-priority channels through abnormal traffic detection, guarantees the reliability of the classified result through the forced reclassification mechanism, directly associates the DSCP marking with the channel binding strategy of the communication selection module, and feeds back the dynamic threshold data to the performance monitoring module to form a closed loop control.

[0092] In one embodiment of the present invention, the step of the dynamic rerouting unit triggering channel switching includes:

[0093] A1: Within the preset period Internally calculates the overall current transmission quality score for all current channels. :

[0094]

[0095] Among them, the preset period To dynamically adapt to the speed of network changes;

[0096] A2: Select channels that meet the following conditions as candidate channels: and ;

[0097] A3: Calculate the difference in overall transmission quality scores between the current channel and the candidate channels. ;like Exceeding the quality difference threshold Then, the decision will be triggered according to the following rules:

[0098]

[0099] in, , =0.7, This is the quality difference threshold. The slope coefficient, , ,

[0100] During the switching process, a transition period is set. Perform data synchronization:

[0101]

[0102] It employs sequence number remapping (SN) technology to dynamically adjust sequence numbers to compensate for transmission delays or packet buffering caused by handover, ensuring that the receiving end can reassemble the data stream in the correct order.

[0103]

[0104] In the formula, The updated serial number. The original serial number before the switch. For the current time, To switch trigger times, This refers to the number of data packets buffered per cycle.

[0105] After the handover is completed, a cooling period is set to prohibit re-handover, wherein ;

[0106] If the candidate channel appears in the transition period , it indicates that the transmission quality is deteriorating, and the original channel is immediately returned to, and the channel that triggers continuous handover failures is marked as a faulty channel and is shielded from the handover duration .

[0107] When the actual number of handovers exceeds the maximum allowed number of handovers, the system determines that the current handover strategy is too aggressive, and automatically raises the hysteresis threshold , wherein the maximum allowed number of handovers is 5 per minute, so that subsequent handovers need to meet more stringent conditions (such as greater signal quality difference, lower delay, etc.), thereby reducing the handover frequency.

[0108] In the formula, is a delay weight coefficient, is a packet loss rate weight coefficient, is a bandwidth utilization rate weight coefficient, is the current transmission delay, is the maximum delay threshold allowed by the protocol, is the current packet loss rate, is the measured effective bandwidth, is the theoretical maximum bandwidth of the channel, wherein: ; is the real-time quality score of the channel currently being used, is the quality score of the candidate channel, is the hysteresis threshold, is the absolute quality lower limit; is the quality threshold for emergency handover, is the minimum quality improvement threshold for emergency handover, is the handover probability, is the minimum probability threshold for triggering actual handover actions, is the average round trip time, is the current round trip time.

[0109] In the present scheme, hierarchical response is achieved by , , The transition period and the cooling period balance the handover benefits and stability losses.

[0110] In one embodiment of the present application, in step A3, during the handover process, a transition period is set to perform data synchronization:

[0111]

[0112] And the sequence number remapping (SN Remapping) technology is used to dynamically adjust the sequence number to compensate for the transmission delay or data packet buffering caused by switching:

[0113]

[0114] In the formula, is the updated sequence number, is the original sequence number before switching, is the current time, is the switching trigger time, is the number of buffered data packets per cycle, is the current round-trip time;

[0115] After the switching is completed, a cooling period is set to prohibit switching again, wherein ;

[0116] If the candidate channel appears in the transition period , it indicates that the transmission quality is deteriorating, and it is immediately returned to the original channel, and the channel that triggers continuous switching failures is marked as a faulty channel and is shielded from switching for a duration ;

[0117] When the actual number of switching times exceeds the maximum allowed number of switching times, the hysteresis threshold is automatically raised , wherein the maximum allowed number of switching times is 5 times per minute.

[0118] In an embodiment of the present application, the GPU memory dynamic partitioning step comprises:

[0119] B1: Establish a memory access heat evaluation unit to generate a memory block heat value by monitoring the memory access frequency and single access residence duration of the container task in a unit time window.

[0120]

[0121] In the formula, is the heat value of the jth memory block, and are weighting coefficients, satisfying ;

[0122] B2: Construct a priority weighted heat mapping table to map the memory blocks according to the formula Dynamic sorting updates the mapping relationship between video memory blocks and containers based on the real-time task queue status. This is a weighting factor for container task priority. This is the priority-weighted heat value of the k-th memory block;

[0123] B3: Employs a dynamic memory partitioning and reconstruction mechanism; when a high-priority task is detected... When requesting video memory, press Release contiguous video memory space occupied by low-priority tasks in descending order, and create a triplet allocation record of <task ID, starting address, block size>.

[0124] In this solution, hot and cold data are separated and stored by quantifying the heat value, the weighted sorting formula ensures that the memory requirements of high-priority tasks are responded to quickly, and triplet records constrain the allocation of continuous space to reduce memory fragmentation.

[0125] In one embodiment of the present invention, the hybrid scheduler of the computing resource allocation module operates according to the following rules:

[0126] C1: Sets the dynamic time slice length based on task priority. ,in As a reference time slice;

[0127] C2: Implement preemption decision when the clock interrupt is triggered, if there exists a condition that satisfies... If a task request is received, the current container context is immediately suspended and the high-priority task pipeline is loaded.

[0128] C3: Perform weighted round-robin scheduling on tasks of the same priority, based on the proportion of time slices used by each container. Automatically adjust the time quota for the next scheduling cycle;

[0129] In the formula, This is the priority weighting factor for new request tasks. This is the priority weighting factor for the currently executing task. In order to seize the dynamic threshold for decision-making, The actual time slice length occupied by the container. This represents the total time slice quota allocated to containers within the scheduling cycle.

[0130] In one embodiment of the present invention, the processing method of the heterogeneous computing offloading module includes:

[0131] D1: Construct a weighted directed acyclic graph (DAG) based on data flow dependencies using acoustic algorithms.

[0132]

[0133] In the formula, For a weighted directed acyclic graph (DAG), Let be the set of nodes in the graph. Let be the set of edges in the graph. To calculate density, For nodes The calculated density, This represents the total number of computational operations for the microtasks corresponding to the node. The amount of data processed by a microtask;

[0134] D2: Generating scheduling priorities based on computational density:

[0135]

[0136] In the formula, For dynamic weighting coefficients, The depth of the longest dependency path of a node in the DAG;

[0137] D3: GPU Task Scheduler Microtasks are executed in descending order when the following conditions are met: GPU kernel scheduling instructions are generated in a timely manner; otherwise, they are allocated to the CPU thread pool.

[0138] In the formula, Density amplification factor It is the sum of the densities of all nodes in the DAG. The total number of nodes in the DAG.

[0139] In this plan, through Accurately distinguish between compute-intensive (GPU) and memory-intensive (CPU) tasks, and introduce [a new feature] in scheduling priority. To shorten the execution latency of the critical path in a DAG, the threshold Avoid inefficient GPU tasks consuming streaming multiprocessor (SM) resources.

[0140] In one embodiment of the present invention, the resource reallocation instruction is used to adjust the GPU-CPU task allocation ratio:

[0141]

[0142] In the formula, To dynamically adjust the GPU task ratio, This represents the current percentage of GPU tasks. This represents the percentage of GPU utilization. This is a baseline value for GPU utilization.

[0143] This solution uses an elastic resource scheduling formula to ensure that GPU utilization approaches the baseline value, avoiding overload or idleness.

[0144] Finally, it should be noted that the above preferred embodiments are merely intended to illustrate the technical solutions of the present application but not to limit the present application. Although the present application has been described in detail through the above preferred embodiments, those skilled in the art should understand that various modifications can be made in form and details without departing from the scope of the present application defined by the claims.

Claims

1. An audio integrated signal processor for 40G / 10G Ethernet and SRIO high speed data transmission, characterized in that, include: The signal classification module dynamically divides the audio signal stream into fast signal stream and slow signal stream by using a preset data packet length threshold and a protocol whitelist. The heterogeneous transmission module includes: a 40G Ethernet pass-through channel implemented through SR-IOV virtualization technology, which is equipped with a hardware virtualization unit with a single I / O virtualization interface; a 10G Ethernet channel using a priority-weighted polling mechanism, which is equipped with at least three bandwidth isolation queues; and an SRIO inter-board interconnection channel built based on the RapidIO protocol, which integrates a DMA controller and an address mapping table. The intelligent routing module connects the signal classification module and the heterogeneous transmission module, and includes: a routing decision unit that matches the corresponding transmission channel according to the signal flow type; and a dynamic rerouting unit that triggers channel switching in response to changes in transmission quality indicators. The container resource scheduling module is used to dynamically partition GPU memory and allocate computing resources for multi-container parallel tasks; The heterogeneous computing offloading module includes: an underwater acoustic algorithm resolver, which decomposes acoustic processing algorithms into a set of microtasks with data dependencies; and a GPU task scheduler, which generates kernel-level scheduling instructions based on the computational density of the microtasks. The closed-loop optimization module collects network latency, jitter, and GPU utilization data in real time, and generates an optimization strategy package containing channel switching instructions and resource reallocation instructions through the feedback controller connected to each module.

2. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high speed data transmission according to claim 1, wherein, The dynamic classification rules of the signal classification module are as follows: When the data packet length Dynamic threshold Or the protocol type is a sonar original data protocol; the fast signal stream is marked as an EF class based on the DSCP marking unit, and is mapped to a strict priority queue of a 40G Ethernet pass-through channel; When the data packet length Dynamic threshold Or the protocol type is a parameter configuration instruction or a log return protocol; based on the DSCP marking unit, the slow signal stream is marked as a BE class and mapped to a weighted fair queue of a 10G Ethernet channel; wherein the dynamic threshold According to the network average load and GPU utilization within the previous 60s, dynamic correction: In the formula, is a dynamic adjusted classification threshold value, is a basic classification threshold value, , is a network load weight coefficient, is a current network load, is a maximum theoretical bandwidth of the network, is a GPU utilization rate weight coefficient, is a current GPU utilization rate, is a maximum theoretical utilization rate of the GPU; If the same container sends more than 50 fast signals in a single second, the traffic shaper rate limit is triggered; forced reclassification is performed on the marked conflict packets and an alarm log is generated. When the instantaneous packet loss rate of a 40G Ethernet pass-through channel is 5% 5%, stop dynamic threshold correction and perform 256KB; when the packet loss rate 5% for 10 seconds, resume dynamic threshold correction.

3. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high speed data transmission of claim 1, wherein: The method for triggering channel switching by a dynamic routing unit includes the following steps: A1: Calculate a current transmission quality comprehensive score of all current channels in a preset period :​ wherein the preset period ; In the formula, is a delay weight coefficient, is a packet loss rate weight coefficient, is a bandwidth utilization rate weight coefficient, satisfying , is a current transmission delay, is a maximum delay threshold allowed by a protocol, is a current packet loss rate, is a measured effective bandwidth, is a theoretical maximum bandwidth of a channel, is an average round-trip time; A2: select a channel satisfying the following conditions as a candidate channel: and ; wherein is the quality score of the candidate channel, is the real-time quality score of the channel currently in use, is a lag threshold, is an absolute quality lower limit; A3: Calculate the difference of the transmission quality comprehensive score of the current channel and the candidate channel ; if the quality difference exceeds a quality difference threshold , then trigger the handover decision according to the following rules: in, wherein = 0.7, is a quality difference threshold, is a slope coefficient, is a quality critical value for emergency handover, is a minimum quality improvement threshold at emergency handover, is a handover probability, is a minimum probability threshold for actual handover action triggering.

4. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high speed data transmission of claim 3, wherein: In step A3, during the handover, a transition period is set Data synchronization is performed: It also employs sequence number remapping technology to dynamically adjust sequence numbers to compensate for transmission delays or data packet buffering caused by handover: In the formula, is an updated sequence number, is an original sequence number before switching, is a current time, is a switching trigger time, is a number of buffered data packets per period, is a current round trip time; After the handover is completed, a cooling period is set to prohibit a further handover, wherein ; Among them, if the candidate channel appears during the transition period If this indicates a deterioration in transmission quality, the system will immediately revert to the original channel for continuous triggering. Channels that fail to complete the handover are marked as faulty channels, and the handover duration is masked. ; wherein the hysteresis threshold is automatically raised when the actual number of handovers exceeds the maximum allowed number of handovers wherein the maximum allowed number of handovers is 5 per minute.

5. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high speed data transmission of claim 1, wherein, The steps for dynamically partitioning GPU memory include: B1: Establishing a video memory access heat evaluation unit to monitor the video memory access frequency of the container task in a unit time window and the single access residence duration to generate a video memory block heat value , In the formula, Let j be the heat value of the j-th memory block. and For the weighting coefficients, satisfying ; B2: construct a priority-weighted hotness mapping table, and map the GPU blocks to the containers according to the formula dynamically rank, update the mapping relationship between the GPU blocks and the containers according to the real-time task queue state, wherein a priority-weighted factor for the container task, a priority-weighted hotness value of the kth GPU block; B3: When a high-priority task is detected, the dynamic reconstruction mechanism of the video memory partition is adopted When the video memory is applied, the continuous video memory space occupied by the low-priority task is released in descending order, and a three-tuple allocation record of <task ID, start address, block size> is established. When the video memory is applied, the continuous video memory space occupied by the low-priority task is released in descending order, and a three-tuple allocation record of <task ID, start address, block size> is established.

6. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high speed data transmission of claim 1, wherein, The hybrid scheduler of the computing resource allocation module operates according to the following rules: C1: Set dynamic time slice length based on task priority wherein is a reference time slice, is a container task priority weight factor; C2: Implement pre-emption decision at clock interrupt trigger, if there is a task request that satisfies the task request, then immediately suspend the current container context and load the high priority task pipeline; C3: Perform a weighted round-robin scheduling for same priority tasks, by each container's used time slice proportion Automatically adjust the time quota of the next scheduling period; wherein, is a priority weight factor for a new requested task, is a priority weight factor for a currently executing task, is a dynamic threshold for preemption decision, is a time slice length that the container has actually occupied, is a total time slice quota allocated to the container within a scheduling period.

7. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high speed data transmission of claim 1, wherein, The methods for handling heterogeneous computing offloading modules include: D1: Based on the data flow dependencies using acoustic algorithms, construct a weighted directed acyclic graph (DAG). In the formula, is a weighted directed acyclic graph DAG, is a set of nodes in the graph, is a set of edges in the graph, is a calculation density, is a calculation density of a node is a calculation density of a node is the total number of calculation operations of the micro task corresponding to the node, is the data volume processed by the micro task; D2: Generating scheduling priorities based on computational density: In the formula, is a dynamic weight coefficient, is the longest dependent path depth of the node in the DAG; D3: The GPU task scheduler executes micro-tasks in descending order, generates GPU kernel scheduling instructions when the condition is met, and assigns to the CPU thread pool otherwise. ​​ wherein is a density amplification factor, is the sum of all node densities in the DAG, is the total number of nodes in the DAG.

8. The audio integrated signal processor for 40G / 10G Ethernet and SRIO high speed data transmission of claim 1, wherein, Resource reallocation commands are used to adjust the GPU-CPU task allocation ratio: In the formula, is the dynamic adjustment of the GPU task proportion, is the current GPU task proportion, is the percentage of GPU utilization, is the benchmark value of GPU utilization.

Citation Information

Patent Citations

  • Self-adaptive task scheduling execution unit management method and system

    CN119376903A

  • Apparatus and method for operating multi lane in high-rate ethernet optical link interface

    KR1020130048091A