Intelligent matrix architecture method, system, equipment and medium
The intelligent matrix architecture is built through domestic processors, and an intelligent matrix system that reduces costs and improves security is realized, the problem of dependence on foreign processors is solved, resource utilization and data transmission are optimized, and it is suitable for artificial intelligence and big data analysis.
Patent Information
- Application Number
- CN202510400813.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The dependence of existing intelligent matrix systems on foreign high-performance processors leads to high costs and supply chain security risks, the heterogeneous computing environment is complex, the lack of dynamic load scheduling mechanism, unbalanced resource utilization, and insufficient identification of performance bottlenecks.
The intelligent matrix architecture is built using domestic processors, and the point-to-point communication between processors is realized through the Ethernet interface, real-time monitoring of CPU usage and load conditions, dynamically allocate computing tasks, optimize data flow management, integrate secure encryption modules, simplify system structure and improve communication efficiency.
Reduce hardware costs, improve supply chain security and independent control, optimize resource utilization, realize efficient computing and data transmission, meet the needs of artificial intelligence and big data analysis, and have higher cost-effectiveness.
Smart Images

Figure CN120336003A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer hardware, and particularly to an intelligent matrix architecture method, system, device and medium. Background Art
[0002] With the rapid development and popular application of artificial intelligence technology, the demand for computing power has shown exponential growth. The traditional single-processor architecture faces performance bottlenecks when dealing with large-scale data sets and running complex algorithm models, and cannot meet the requirements of real-time response and high-concurrency processing. As a solution, the distributed computing architecture has emerged, especially the intelligent matrix system based on multi-processor collaborative work. Due to its efficient parallel computing ability and flexible scalability, it shows great potential and application value in the fields of scientific computing, deep learning and big data analysis.
[0003] Currently, the mainstream intelligent matrix systems on the market are mostly built with high-performance processors from foreign manufacturers. These systems are based on the principle of large-scale parallel computing, and achieve excellent computing performance by constructing a processor grid and realizing high-speed interconnection; especially in constructing artificial intelligence algorithm models and processing massive data, the existing intelligent matrix systems can decompose complex tasks into multiple subtasks for parallel processing, greatly improving the computing efficiency and resource utilization rate.
[0004] However, the existing intelligent matrix systems have many deficiencies: First, the dependence on high-performance processors brings higher hardware costs and potential supply chain security risks; second, the heterogeneous computing environment between different processors leads to high system integration and maintenance complexity; third, the lack of an intelligent scheduling mechanism for dynamic load changes results in uneven resource utilization in both peak load and light load states; finally, the existing systems have insufficient ability to identify and adaptively adjust to processor performance bottlenecks, and it is difficult to achieve optimal resource allocation and task distribution in a complex and changing computing environment.
[0005] Therefore, how to provide an intelligent matrix architecture solution that reduces costs and improves supply chain security has become an urgent problem to be solved at present. Summary of the Invention
[0006] Embodiments of the present invention provide an intelligent matrix architecture method, system, device and medium to solve the problems of high cost and large supply chain risk existing in the existing intelligent matrix systems that mostly use high-performance processors from foreign manufacturers.
[0007] To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.
[0008] According to the first aspect of the embodiments of the present invention, an intelligent matrix architecture method is provided.
[0009] In one embodiment, the intelligent matrix architecture method includes:
[0010] Construct a core computing module based on multiple domestic processors, implement point-to-point communication between processors through an Ethernet interface, and perform parallel operations using the multi-core architecture integrated in the processors;
[0011] Real-time monitor the CPU usage rate of each processor core, dynamically allocate computing tasks according to the load situation, and manage the data stream;
[0012] Establish a data transmission channel based on an interface chip, and transmit the data processed by the processor to an external device.
[0013] In one embodiment, real-time monitoring the CPU usage rate of each processor core, dynamically allocating computing tasks according to the load situation, and managing the data stream includes:
[0014] Use a performance monitoring script to read interface data, analyze the CPU usage rate, memory utilization rate, and network bandwidth utilization rate of the processor, and identify performance bottlenecks;
[0015] Through maintaining a task queue and setting task priorities, and dynamically adjusting task allocation to achieve parallel processing of computing tasks;
[0016] Based on data transmission optimization and data stream control technologies, establish a data stream management mechanism, and perform error detection and correction to achieve reliable data transmission.
[0017] In one embodiment, using a performance monitoring script to read interface data, analyze the CPU usage rate, memory utilization rate, and network bandwidth utilization rate of the processor, and identify performance bottlenecks includes:
[0018] Based on the monitoring period, regularly sample the interface data of each processor, calculate the CPU usage rate, and solve the load balancing index of the processor according to the current load and the maximum load;
[0019] Calculate the memory utilization rate and network bandwidth utilization rate of the processor, and combine the CPU usage rate of the processor to identify the performance bottleneck of the processor;
[0020] Monitor the temperature and power supply of each processor through a sensor, and issue an alarm when the temperature or power supply exceeds the safe range.
[0021] In one embodiment, the expression of the CPU usage rate is:
[0022]
[0023] Wherein, U cpu represents the CPU usage rate, T i represents the usage time of the i-th core, T total,i represents the total time of the i-th core, and n represents the number of cores;
[0024] The expression for memory utilization rate is:
[0025]
[0026] Wherein, U m represents the memory utilization rate, M used represents the used memory space, M total represents the total capacity of the DDR memory;
[0027] The expression for network bandwidth utilization rate is:
[0028]
[0029] Wherein, U n represents the network bandwidth utilization rate, R real represents the actual transmission rate, R max represents the maximum supported rate of the Ethernet interface.
[0030] In one embodiment, parallel processing of computing tasks is achieved by maintaining a task queue, setting task priorities, and dynamically adjusting task allocation, including:
[0031] Create a task queue data structure to store tasks to be processed;
[0032] Set initial priorities for tasks according to task urgency and resource requirements;
[0033] Dynamically adjust task allocation according to the load balancing index of each processor.
[0034] In one embodiment, dynamically adjusting task allocation according to the load balancing index of each processor includes:
[0035] Sort the tasks in the task queue from high to low according to the initial task priorities, and divide the tasks into high-priority tasks and low-priority tasks;
[0036] Compare the load balancing index of each processor with a preset threshold. If the load balancing index is lower than the preset threshold, mark the processor as a low-load processor;
[0037] Based on system resource fluctuations and task execution progress, when a high-priority task is detected to be inserted into the queue, the execution of the current low-priority task is paused, and a low-load processor is allocated to the newly inserted high-priority task to ensure that the response time to urgent tasks does not exceed the preset time limit.
[0038] In one embodiment, based on data transmission optimization and data flow control technologies, a data flow management mechanism is established, and error detection and correction are performed to achieve reliable data transmission, including:
[0039] Using data fragmentation technology to fragment the data packets of the data flow, and through data compression processing, reducing the data transmission volume;
[0040] Through the sliding window protocol, controlling the transmission rate of the data flow to achieve flow control;
[0041] Using cyclic redundancy check and forward error correction coding technologies to detect and correct errors during data transmission.
[0042] According to the second aspect of the embodiments of the present invention, an intelligent matrix architecture system is provided.
[0043] In one embodiment, the intelligent matrix architecture system includes:
[0044] A core computing unit, constructing a core computing module based on multiple domestic processors, realizing point-to-point communication between processors through an Ethernet interface, and performing parallel operations using the multi-core architecture integrated in the processors;
[0045] A system management unit, used to monitor the CPU usage rate of each processor core in real time, dynamically allocate computing tasks according to the load situation, and manage the data flow;
[0046] A data transmission unit, used to establish a data transmission channel based on an interface chip and transmit the data processed by the processor to external devices.
[0047] According to the third aspect of the embodiments of the present invention, a computer device is provided.
[0048] In one embodiment, the computer device includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the above method are implemented.
[0049] According to the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided.
[0050] In one embodiment, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the above method are implemented.
[0051] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0052] (1) The present invention constructs an intelligent matrix architecture using a domestic processor (RK3588), effectively reducing the system's dependence on foreign high-performance processors. This not only significantly reduces the hardware cost, but more importantly, improves the security and autonomy of the supply chain. By optimizing the PCB layout and signal transmission path, reducing signal interference and loss, and combining improvements in the implementation of the PCIe protocol stack to reduce protocol overhead, the present invention supports a data transmission rate of up to 4GB / s, which is about 50% higher than the traditional PCIe2.0x4 interface. At the same time, it improves the stability and efficiency of data transmission, providing a reliable guarantee for large-scale data processing and artificial intelligence applications.
[0053] (2) The present invention realizes a direct communication architecture between processors. Four RK3588 processors are directly interconnected through Ethernet interfaces without the need for additional switch devices, significantly simplifying the system structure, reducing the deployment cost, and improving the communication efficiency. At the same time, the system integrates dynamic voltage and frequency scaling (DVFS) technology. The system management unit monitors the load conditions of each processor in real time, reasonably allocates tasks through a load balancing algorithm to avoid resource waste, and dynamically adjusts the voltage and frequency of the processor according to the current load. For example, it reduces the frequency and voltage under low load, effectively reducing the system power consumption while ensuring performance, and realizing the efficient utilization of computing resources and higher cost performance.
[0054] (3) The present invention integrates a security encryption module in the hardware design, supporting multiple encryption algorithms such as AES and SHA to ensure the security of data during transmission and storage. The system realizes real-time encryption processing of transmitted data to prevent data from being stolen or tampered with during transmission, and supports the secure boot function to ensure that the firmware and software loaded during system startup have been strictly verified, effectively preventing malware intrusion and comprehensively improving the security performance of the system. The four RK3588 processors work together to provide powerful computing capabilities to meet the requirements of high-load application scenarios such as artificial intelligence and big data analysis, and have higher cost performance compared with foreign similar products.
[0055] (4) The present invention adopts a modular design concept. In terms of hardware, each processor module is designed to be plug-and-play and is equipped with an independent power supply and heat dissipation module. Users can seamlessly connect a new processor to the existing system through a standardized PCIe3.0x4 interface. In addition, at the software level, the extension management module is responsible for the initialization and configuration of the new processor. The system management unit can dynamically allocate tasks to the newly added processor according to the current task load and efficiently manage the data stream to ensure unblocked data transmission between processors, providing flexible expansion capabilities for different scale application scenarios and meeting the diverse needs from small-scale deployments to large-scale computing clusters.
[0056] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0058] Figure 1 is a schematic flow chart of a method for an intelligent matrix architecture shown according to an exemplary embodiment;
[0059] Figure 2 is a specific implementation diagram of an intelligent matrix architecture system shown according to an exemplary embodiment;
[0060] Figure 3 is a structural block diagram of an intelligent matrix architecture system shown according to an exemplary embodiment;
[0061] Figure 4 is a schematic structural diagram of a computer device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] The following description and the drawings fully illustrate the specific embodiments herein, enabling those skilled in the art to practice them. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. The scope of the embodiments herein includes the entire scope of the claims and all available equivalents of the claims. Herein, the terms "first", "second", etc. are only used to distinguish one element from another, and do not require or imply any actual relationship or order between these elements. In fact, the first element can also be called the second element, and vice versa. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a structure, device or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such structure, device or equipment. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the structure, device or equipment including the said element. The embodiments herein are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0063] As used herein, the terms "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. These are only for the convenience of describing the present disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation on the present invention. In the description of the present disclosure, unless otherwise specified and limited, the terms "mounted", "connected", and "coupled" shall be understood in a broad sense. For example, it may be a mechanical connection or an electrical connection, or it may be the communication inside two elements. It may be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0064] As used herein, unless otherwise specified, the term "plurality" means two or more.
[0065] As used herein, the character " / " indicates that the objects before and after are in an "or" relationship. For example, A / B means: A or B.
[0066] As used herein, the term "and / or" is an associative relationship describing an object, indicating that three relationships can exist. For example, A and / or B means: A or B, or, the three relationships of A and B.
[0067] It should be understood that although the steps in the flowchart are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in the present disclosure, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0068] Each module in the device or system of the present application can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent thereof, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0069] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0070] Figure 1An embodiment of an intelligent matrix architecture method of the present invention is shown.
[0071] In this alternative embodiment, the intelligent matrix architecture method includes:
[0072] Step S101, construct a core computing module based on multiple domestic processors, implement point-to-point communication between processors through an Ethernet interface, and perform parallel computing using the multi-core architecture integrated in the processors;
[0073] Step S102, monitor the CPU usage rate of each processor core in real time, dynamically allocate computing tasks according to the load situation, and manage the data stream;
[0074] Step S103, establish a data transmission channel based on the interface chip, and transmit the data processed by the processor to external devices.
[0075] Specifically, in this embodiment, an intelligent matrix system based on domestic high-performance processors constructed by the intelligent matrix architecture method includes: 4 RK3588 processors, as the core computing unit 201 of the system; a PCIe3.0x4 interface as the data transmission unit 203, used to transmit the processed data to external devices at high speed; a system management unit 202, used to monitor the system status, allocate computing tasks, and manage the data stream.
[0076] Specifically, the domestic processor can adopt an ARM architecture processor independently developed by Rockchip (RK3588 processor), integrated with 4 Cortex-A76 cores and 4 Cortex-A55 cores, with a maximum main frequency of up to 2.4 GHz, and equipped with an independent NPU, which can provide powerful AI computing power.
[0077] Specifically, the interface chip adopts a domestic chip (PCIe3.0x4 interface), supports a data transmission rate of up to 4 GB / s, and can meet the needs of large-scale data transmission.
[0078] Specifically, the system management unit 202 adopts a domestic MCU, runs self-developed management software, and can realize real-time monitoring of the system status, dynamic allocation of computing tasks, and efficient management of the data stream.
[0079] Specifically, the 4 RK3588 processors perform point-to-point communication through an Ethernet interface, without additional switch devices, simplifying the system structure and reducing costs.
[0080] Specifically, 1) Monitor the system status, including:
[0081] a) Initialize the monitoring module:
[0082] Start the system management unit 202 and initialize resource utilization monitoring and health status monitoring;
[0083] Set the monitoring period, for example, sample once per second;
[0084] b) Resource utilization monitoring:
[0085] Periodically sample the CPU usage, memory utilization, and network bandwidth utilization of each RK3588 processor.
[0086] Calculate the resource utilization and record it in the log file.
[0087] c) Health status monitoring:
[0088] Monitor the temperature and power supply of each RK3588 processor through sensors.
[0089] If the temperature or power supply exceeds the safe range, trigger an alarm and take corresponding measures, such as reducing the load or shutting down the system.
[0090] Specifically, 2) Allocate computing tasks, including:
[0091] a) Task queue initialization: Initialize the task queue, set the task priorities; Add new tasks to the task queue.
[0092] b) Dynamic task allocation: Periodically check the current load of each processor.
[0093] c) Dynamically allocate tasks according to the load balancing index; Adopt the round-robin scheduling algorithm and the priority scheduling algorithm to ensure fair allocation and efficient execution of tasks.
[0094] Specifically, 3) Manage the data stream, including:
[0095] a) Data transmission optimization: Fragment large data packets to reduce transmission latency; Compress the data to reduce the data volume.
[0096] b) Data stream control: Adopt the sliding window protocol to control the transmission rate of the data stream.
[0097] c) Adopt cyclic redundancy check and forward error correction coding techniques to detect and correct errors during transmission.
[0098] In this alternative embodiment, the CPU usage of each processor core is monitored in real time, computing tasks are dynamically allocated according to the load situation, and managing the data stream includes:
[0099] Step S1021, Use the performance monitoring script to read the interface data, analyze the CPU usage, memory utilization, and network bandwidth utilization of the processor, and identify the performance bottleneck;
[0100] Step S1022: Achieve parallel processing of computing tasks by maintaining a task queue, setting task priorities, and dynamically adjusting task allocation.
[0101] Step S1023: Based on data transmission optimization and data flow control technology, establish a data flow management mechanism and perform error detection and correction to achieve reliable data transmission.
[0102] In this alternative embodiment, use a performance monitoring script to read interface data, analyze the CPU usage rate, memory utilization rate, and network bandwidth utilization rate of the processor, and identify performance bottlenecks, including:
[0103] Step S10211: Based on the monitoring period, regularly sample the interface data of each processor, calculate the CPU usage rate, and solve the load balancing index of the processor according to the current load and the maximum load.
[0104] Step S10212: Calculate the memory utilization rate and network bandwidth utilization rate of the processor, and combine the CPU usage rate of the processor to identify the performance bottleneck of the processor.
[0105] Step S10213: Monitor the temperature and power supply status of each processor through sensors, and issue an alarm when the temperature or power supply exceeds the safe range.
[0106] Specifically, first, ① Monitor the system status, including:
[0107] Load monitoring: Use the built-in performance monitoring script of the system to read the relevant interface data of the operating system at a specific time interval (per second), obtain the CPU usage rate information of each core, and grasp the load status of each core in real time.
[0108] Specifically, improve the monitoring algorithm, and improve the accuracy and efficiency of system status monitoring through more efficient algorithms and parameter settings. Specifically include:
[0109] Resource utilization monitoring: Improve the calculation formula of the CPU usage rate, considering the characteristics of the multi-core architecture; the expression of the CPU usage rate in the present invention is:
[0110]
[0111] In the formula, U cpu represents the CPU usage rate, T i represents the usage time of the i-th core, T total,i represents the total time of the i-th core, and n represents the number of cores.
[0112] Specifically, health status monitoring: Improve the algorithms for temperature and power supply monitoring to improve the monitoring accuracy and response speed.
[0113] Specifically, supplement memory and network monitoring. By reintroducing memory and network monitoring metrics and providing specific calculation formulas and monitoring methods. The specific methods include:
[0114] a) Calculate the memory utilization rate. The expression for the memory utilization rate is:
[0115]
[0116] In the formula, U m represents the memory utilization rate, M used represents the used memory space, and M total represents the total DDR memory capacity;
[0117] b) Calculate the network bandwidth utilization rate. The expression for the network bandwidth utilization rate is:
[0118]
[0119] In the formula, U n represents the network bandwidth utilization rate, R real represents the actual transmission rate, and R max represents the maximum supported rate of the Ethernet interface.
[0120] Specifically, the resource utilization rate monitoring is based on the resource management mechanism of the operating system. By performing timed sampling and calculations, the usage of system resources can be grasped in real time.
[0121] Specifically, the health status monitoring is based on hardware sensors and the power management module. By monitoring the temperature and power supply in real time, the stable operation of the system can be ensured.
[0122] In this alternative embodiment, the parallel processing of computing tasks is achieved by maintaining a task queue, setting task priorities, and dynamically adjusting task allocation, including:
[0123] Step S10221: Create a task queue data structure to store tasks to be processed;
[0124] Step S10222: Set an initial priority for the tasks according to the task urgency and resource requirements;
[0125] Step S10223: Dynamically adjust the task allocation according to the load balancing index of each processor.
[0126] In this alternative embodiment, dynamically adjusting the task allocation according to the load balancing index of each processor includes:
[0127] Sort the tasks in the task queue from high to low according to the initial task priorities, and divide them into high-priority tasks and low-priority tasks;
[0128] Compare the load balancing index of each processor with a preset threshold. If the load balancing index is lower than the preset threshold, mark the processor as a low-load processor;
[0129] Based on system resource fluctuations and task execution progress, when a high-priority task is detected to be inserted into the queue, pause the execution of the current low-priority task and allocate the low-load processor to the newly inserted high-priority task to ensure that the response time to emergency tasks does not exceed the preset time limit.
[0130] In this alternative embodiment, based on data transmission optimization and data flow control technologies, establish a data flow management mechanism and perform error detection and correction to achieve reliable data transmission, including:
[0131] Step S10231: Fragment the data packets of the data flow using data fragmentation technology and reduce the data transmission volume through data compression processing;
[0132] Step S10232: Control the transmission rate of the data flow through the sliding window protocol to achieve flow control;
[0133] Step S10233: Use cyclic redundancy check and forward error correction coding technologies to detect and correct errors during data transmission.
[0134] Specifically, then, ② allocate computing tasks, and the system management unit 202 realizes the allocation of computing tasks in the following ways:
[0135] a) Task queue management:
[0136] Task queue: Maintain a task queue to store tasks to be processed.
[0137] Task priority: Set the priority of tasks according to the urgency of tasks and resource requirements.
[0138] b) Dynamic task allocation:
[0139] Load balancing: Dynamically allocate tasks according to the current load of each processor to ensure the efficient use of system resources. The expression for load balancing is: Load balancing index = maximum load / current load;
[0140] In the formula, the current load refers to the current task load of the processor, and the maximum load refers to the maximum supported load of the processor.
[0141] Specifically, in the present invention, the task scheduling algorithm adopts a combination of the round-robin scheduling algorithm and the priority scheduling algorithm to ensure the fair allocation and efficient execution of tasks.
[0142] Specifically, the task queue management is based on the design principles of the operating system. By maintaining the task queue and setting task priorities, it ensures the orderly processing of tasks.
[0143] Specifically, b) The dynamic task allocation is based on load balancing and task scheduling algorithms. By dynamically adjusting the task allocation, it improves the parallel processing ability of the system.
[0144] Specifically, for dynamic task allocation: A task scheduler is constructed, taking into account the fairness of the round-robin scheduling algorithm and the efficiency of the priority scheduling algorithm. When a new task arrives, based on the current load of each core and the task priority, an intelligent decision is made on where to allocate the task. For example, if core A has a relatively light current load and is suitable for processing high-priority tasks, the scheduler will preferentially allocate such tasks to core A to achieve the optimization of task execution.
[0145] Specifically, the expression for load balancing is: Load balancing index = Current load / Maximum load; where the current load is the total task load that the processor is currently undertaking, calculated by methods such as the weighted average of the CPU usage rates of each core; the maximum load represents the limit task load that the processor can bear under specific conditions, determined by hardware characteristics and design parameters. This formula intuitively reflects the degree of load balancing. The lower the index, the more balanced the load and the more efficient the system operation.
[0146] Specifically, the priority is dynamically adjusted according to the urgency of the task and resource requirements, ensuring the timely execution of high-priority tasks, improving the system response speed and resource utilization efficiency, and meeting the requirements of complex and changeable task processing scenarios.
[0147] Specifically, a) Task queue management: Create a task queue data structure for temporarily storing tasks to be processed. When each task enters the queue, considering the urgency of the task (including levels such as urgent, high-priority, and normal) and resource requirements (such as the estimated usage of CPU resources, memory resources, etc.), through a specific priority calculation algorithm, an initial priority is set for the task, and it is stored in the queue sorted by priority for subsequent scheduling preparation.
[0148] Specifically, for dynamic scheduling: The present invention adopts a dynamic scheduling algorithm, which pays real-time attention to changes in task priorities and system load conditions. During the task execution process, according to system resource fluctuations and task execution progress, the task priorities are adjusted in a timely manner. For example, when the system load suddenly increases, to ensure critical tasks, the priorities of non-urgent tasks are appropriately reduced; when an urgent task is inserted, the queue order is dynamically adjusted to ensure that it can quickly obtain resources for execution, optimizing the overall task execution process.
[0149] Specifically, the present invention adopts a method for intelligent identification of performance bottlenecks. By real-time monitoring the usage of system resources, it intelligently identifies performance bottlenecks and performs optimization in a timely manner. Specifically, it includes:
[0150] 1) Resource monitoring: Real-time monitoring of key metrics such as CPU usage, memory utilization, network bandwidth utilization, etc.
[0151] 2) Bottleneck identification: By analyzing these metrics, identify potential performance bottlenecks. For example, if the CPU usage persists above a certain threshold, it can be considered that there is a CPU bottleneck in the system.
[0152] 3) Optimization measures: Based on the identified bottlenecks, take corresponding optimization measures, such as increasing resource allocation, optimizing task scheduling, etc.
[0153] Specifically, finally, ③ manage the data flow, and the system management unit 202 realizes the management of the data flow in the following ways:
[0154] a) Data transmission optimization:
[0155] Data fragmentation: Fragment large data packets into multiple small data packets to reduce transmission latency.
[0156] Data compression: Compress the data to reduce the amount of data and improve transmission efficiency.
[0157] b) Data flow control:
[0158] Flow control: Control the transmission rate of the data flow through the Sliding Window Protocol to ensure the stability of data transmission.
[0159] c) Error detection and correction: Adopt technologies such as Cyclic Redundancy Check (CRC) and Forward Error Correction (FEC) to detect and correct errors during transmission.
[0160] Specifically, the data transmission optimization is based on network transmission protocols and data compression algorithms, and improves the efficiency of data transmission through fragmentation and compression.
[0161] Specifically, the data flow control is based on the Sliding Window Protocol and error detection and correction technologies, and ensures the reliability of data transmission through flow control and error handling.
[0162] Figure 2 With Figure 3 Figure [ID] shows an embodiment of an intelligent matrix architecture system of the present invention.
[0163] In this alternative embodiment, the intelligent matrix architecture system includes:
[0164] A core computing unit 201, constructing a core computing module based on multiple domestic processors, realizing point-to-point communication between processors through an Ethernet interface, and performing parallel computing using the multi-core architecture integrated in the processors;
[0165] System management unit 202, used to monitor the CPU usage of each processor core in real time, dynamically allocate computing tasks according to load conditions, and manage data flow;
[0166] The data transmission unit 203 is used to establish a data transmission channel based on the interface chip and transmit the data processed by the processor to an external device.
[0167] Specifically, the present invention proposes an intelligent matrix architecture system (an intelligent matrix system based on a domestic high-computing power processor), comprising:
[0168] The core computing unit 201 includes four RK3588 processors, which are the core computing modules of the system. Each RK3588 processor is equipped with an independent heat dissipation module and power supply module. The four RK3588 processors communicate point-to-point through the Ethernet interface to achieve high-speed data exchange and collaborative computing. Each RK3588 processor is equipped with an independent heat dissipation module and power supply module to ensure stable operation of the system.
[0169] The data transmission unit 203 includes a PCIe 3.0x4 interface for transmitting the processed data to an external device such as a GPU or FPGA at high speed to meet higher performance computing requirements;
[0170] The system management unit 202 is used to monitor the system status, allocate computing tasks and manage data flows to ensure efficient and stable operation of the system.
[0171] Specifically, the detailed description of the system management unit 202 is as follows:
[0172] The RK3588 processor in the present invention adopts a multi-core architecture of 4 Cortex-A76 cores and 4 Cortex-A55 cores. In order to make full use of hardware resources, it is necessary to design a multi-core load balancing strategy, monitor the CPU usage of each core in real time, dynamically allocate tasks according to the load situation, and combine round-robin scheduling and priority scheduling algorithms to ensure fair and efficient execution of tasks. The load balancing index is defined as the ratio of the current load to the maximum load, which is used to quantitatively evaluate the load balancing state and provide a key basis for task scheduling decisions to maximize the overall computing efficiency.
[0173] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store static information and dynamic information data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0174] Those skilled in the art can understand that Figure 4 the structure shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0175] In addition, the present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the steps in the above method embodiments.
[0176] In addition, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0177] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0178] The present invention is not limited to the structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. An intelligent matrix architecture method, characterized in that, including: Construct a core computing module based on multiple domestic processors, implement point-to-point communication between processors through Ethernet interfaces, and perform parallel computing using the multi-core architecture integrated in the processors; Real-time monitor the CPU usage rate of each processor core, dynamically allocate computing tasks according to the load situation, and manage the data stream; Establish a data transmission channel based on the interface chip to transmit the data processed by the processor to external devices.
2. The intelligent matrix architecture method according to claim 1, wherein The real-time monitoring of the CPU usage rate of each processor core, dynamically allocating computing tasks according to the load situation, and managing the data stream includes: Use a performance monitoring script to read interface data, analyze the CPU usage rate, memory utilization rate, and network bandwidth utilization rate of the processor, and identify performance bottlenecks; Through maintaining a task queue and setting task priorities, and dynamically adjusting task allocation to achieve parallel processing of computing tasks; Based on data transmission optimization and data stream control technologies, establish a data stream management mechanism, and perform error detection and correction to achieve reliable data transmission.
3. The intelligent matrix architecture method according to claim 2, wherein The use of a performance monitoring script to read interface data, analyze the CPU usage rate, memory utilization rate, and network bandwidth utilization rate of the processor, and identify performance bottlenecks includes: Based on the monitoring period, regularly sample the interface data of each processor, calculate the CPU usage rate, and solve the load balancing index of the processor according to the current load and the maximum load; Calculate the memory utilization rate and network bandwidth utilization rate of the processor, and combine the CPU usage rate of the processor to identify the performance bottleneck of the processor; Monitor the temperature and power supply of each processor through sensors, and issue an alarm when the temperature or power supply exceeds the safe range.
4. The intelligent matrix architecture method according to claim 3, wherein The expression of the CPU usage rate is: Where U cpu represents the CPU usage rate, T i represents the usage time of the i-th core, T total,i represents the total time of the i-th core, and n represents the number of cores; The expression of the memory utilization rate is: Where U m represents the memory utilization rate, M used represents the used memory space, and M total represents the total capacity of the DDR memory; The expression of the network bandwidth utilization rate is: Where, U n represents the network bandwidth utilization rate, R real represents the actual transmission rate, and R max represents the maximum supported rate of the Ethernet interface.
5. The intelligent matrix architecture method according to claim 2, wherein The implementation of parallel processing of computing tasks by maintaining a task queue, setting task priorities, and dynamically adjusting task allocation includes: Create a task queue data structure to store tasks to be processed; Set an initial priority for the task according to the task urgency and resource requirements; Dynamically adjust task allocation according to the load balancing index of each processor.
6. The intelligent matrix architecture method according to claim 5, characterized in that, The dynamically adjusting task allocation according to the load balancing index of each processor includes: Sort the tasks in the task queue from high to low according to the initial task priority, and divide them into high-priority tasks and low-priority tasks; Compare the load balancing index of each processor with the preset threshold. If the load balancing index is lower than the preset threshold, mark the processor as a low-load processor; Based on system resource fluctuations and task execution progress, when a high-priority task is detected to be inserted into the queue, pause the execution of the current low-priority task, and allocate the low-load processor to the newly inserted high-priority task to ensure that the response time for urgent tasks does not exceed the preset time limit.
7. An intelligent matrix architecture method according to claim 2, characterized in that, The establishment of a data stream management mechanism based on data transmission optimization and data stream control technologies, and performing error detection and correction to achieve reliable data transmission includes: Use data fragmentation technology to fragment the data packets of the data stream, and reduce the data transmission volume through data compression processing; Control the transmission rate of the data stream through the sliding window protocol to achieve flow control; Utilize cyclic redundancy check and forward error correction coding techniques to detect and correct errors during data transmission.
8. An intelligent matrix architecture system, characterized in that, It includes: A core computing unit that constructs a core computing module based on multiple domestic processors, realizes point-to-point communication between processors through an Ethernet interface, and performs parallel operations using the multi-core architecture integrated in the processors; A system management unit that is used to monitor the CPU usage rate of each processor core in real time, dynamically allocate computing tasks according to the load situation, and manage the data stream; A data transmission unit that is used to establish a data transmission channel based on an interface chip and transmit the data processed by the processor to an external device.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Computing platform architecture based on domestic AI processor
CN117093537A
Implementation method of novel multi-application intelligent card operating system
CN118535330A
Policy-Based Dynamic Compute Unit Adjustments
US20200341597A1