Distributed double-control server and distributed double-control system
Through the distributed dual-control server architecture, the management bus and network switching module are used to connect the processing modules, which solves the problem of low interconnection efficiency between the CPU and intelligent modules in traditional servers, realizes a high-performance and high-reliability computing environment, and supports hot plugging and expansion.
Patent Information
- Application Number
- CN202422879287.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2034-11-25
AI Technical Summary
In the traditional intelligent computing server architecture, the CPU and intelligent modules are interconnected by the PCIE bus, resulting in low computing efficiency and poor reliability. As the amount of data grows, the power consumption and cost of the entire machine increase, and the equipment complexity is high.
A distributed dual-control server architecture is adopted, and the first processing module and the second processing module are connected through the management bus and the network switching module, and interconnected with the distributed computing module to realize the switching of the main and backup modules, avoiding the master-slave relationship of the PCIE bus and improving computing performance and reliability.
It improves the performance and reliability of computing servers, reduces overall machine power consumption and costs, supports hot plugging and expansion, facilitates maintenance, and ensures business continuity.
Smart Images

Figure CN223347336U_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of server architecture, and in particular to a distributed dual-control server and a distributed dual-control system. Background Art
[0002] With the rapid rise of artificial intelligence technology, a large amount of data needs to be analyzed and processed intelligently, including data access, intelligent analysis, data storage, and data forwarding. The traditional intelligent computing server architecture is mainly centered on the CPU, and uses the PCIE bus to cascade intelligent chip modules to perform intelligent analysis and processing of data. Figure 1 As shown, the CPU pulls in external data streams (such as video data from an IPC) through the data port, then forwards it via the PCIE switch module to the various intelligent modules for intelligent analysis. The analyzed data is then transmitted back to the CPU via the PCIE switch module, where it is stored. Simultaneously, the intelligent computing server must connect to the platform and forward the analyzed data to the platform server. This places a heavy load on the CPU. The CPU must be responsible for data access, forwarding it to the intelligent modules, moving data from the intelligent modules to the platform, and storing it all at the same time.
[0003] As can be seen, in traditional server architectures, the CPU and intelligent modules are interconnected via the PCIe bus, forming a master-slave relationship. Data access, storage, forwarding, and computing tasks primarily rely on the CPU's processing power. However, as data volumes continue to grow, this architecture has gradually become inefficient and unreliable. For example, the intelligent module can perform intelligent analysis on 256 video streams, but due to CPU performance limitations, the intelligent server can only access, intelligently analyze, forward, and store 128 video streams. To match the computing power of the intelligent module, the CPU needs to be upgraded with a higher-specification, higher-performance CPU. This increases overall power consumption, overall cost, and device complexity. For another example, if a smart card experiences an error, the PCIe bus connecting the smart card to the CPU may become stuck, causing the CPU to hang and the entire device to cease operation.
[0004] Currently, no effective solution has been proposed to address the problems of low efficiency and poor reliability of computing servers in related technologies. Utility Model Content
[0005] In this embodiment, a distributed dual-control server and a distributed dual-control system are provided to solve the problems of low efficiency and poor reliability of computing servers in related technologies.
[0006] In a first aspect, a distributed dual-control server is provided in this embodiment, including: a distributed computing module, a network switching module, a first processing module and a second processing module;
[0007] The first processing module and the second processing module are connected to the distributed computing module via a management bus;
[0008] The first processing module and the second processing module are further connected to the distributed computing module via the network switching module;
[0009] The first processing module and the second processing module are connected via a keep-alive bus;
[0010] When any one of the first processing module and the second processing module is working, the other processing module is not working; the working processing module is used to control the distributed computing module to perform computing tasks.
[0011] In some embodiments, the distributed computing module includes several computing modules and expansion boards; each of the computing modules is connected to the expansion board via a connector.
[0012] In some of these embodiments, the expansion board includes a first management interface, a second management interface, and an expansion interface;
[0013] The first management interface is connected to the first processing module via a first out-of-band management bus;
[0014] The second management interface is connected to the second processing module via a second out-of-band management bus;
[0015] The expansion interface is used to connect each of the computing modules.
[0016] In some embodiments, the expansion interface includes a hot-swap interface.
[0017] In some embodiments, the computing module is an intelligent analysis module for video data.
[0018] In some of the embodiments, the processing module is further configured to schedule another computing module to continue executing the computing task when one computing module fails.
[0019] In some of the embodiments, the further comprising: a storage module;
[0020] The storage module is connected to the first processing module and the second processing module respectively.
[0021] In some embodiments, the network switching module includes a network topology module based on the Ethernet protocol.
[0022] In a second aspect, a distributed dual-control system is provided in this embodiment, the system comprising: a platform server, a data source end, and the distributed dual-control server described in any one of the first aspects;
[0023] The platform server and the data source are respectively connected to the network switching modules of the distributed dual-control servers.
[0024] In some embodiments, the data source includes a network camera.
[0025] Compared with the related art, the distributed dual-control server and distributed dual-control system provided in this embodiment include a distributed computing module, a network switching module, a first processing module and a second processing module; the first processing module and the second processing module are connected to the distributed computing module through a management bus; the first processing module and the second processing module are also connected to the distributed computing module through the network switching module; the first processing module and the second processing module are connected through a keep-alive bus; when any one of the first processing module and the second processing module is working, the other processing module is not working; the processing module when working is used to control the distributed computing module to perform computing tasks; it solves the problem that under the traditional solution, the CPU and the intelligent module are interconnected by the PCIE bus to form a master-slave relationship, resulting in low computing efficiency and poor reliability, and realizes the interconnection of the distributed computing module with the dual processing module through the network switching module, thereby improving the performance of the computing server.
[0026] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0028] Figure 1 A schematic diagram of the structure of a computing server in the prior art;
[0029] Figure 2 This is a schematic diagram of the structure of a distributed dual-control server in one embodiment;
[0030] Figure 3 This is a schematic diagram of the structure of a distributed dual-control server in another embodiment;
[0031] Figure 4 Schematic diagram of the structure of a distributed dual-control system in one embodiment;
[0032] Figure 5 FIG. 1 is a schematic diagram of the structure of a distributed dual-control server in a preferred embodiment.
[0033] Figure numerals: 100, distributed dual-control server; 110, distributed computing module; 111, computing module; 112, expansion board; 120, network switching module; 130, first processing module; 140, second processing module; 150, storage module; 200, platform server; 300, data source end. DETAILED DESCRIPTION
[0034] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0035] Unless otherwise defined, the technical terms or scientific terms involved in this application should have the general meaning understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "an", "a", "the", "these" and the like in this application do not indicate quantitative restrictions, and they can be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application refers to two or more. "And / or" describes the relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. Generally, the character " / " indicates that the related objects are in an "or" relationship. The terms "first," "second," "third," etc. used in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0036] This embodiment provides a distributed dual-control server 100. Figure 2 is a structural diagram of a distributed dual-control server 100, as shown in Figure 2 As shown, the distributed dual-control server 100 includes: a distributed computing module 110 , a network switching module 120 , a first processing module 130 and a second processing module 140 .
[0037] The first processing module 130 and the second processing module 140 are connected to the distributed computing module 110 via a management bus.
[0038] The first processing module 130 and the second processing module 140 are further connected to the distributed computing module 110 through the network switching module 120 .
[0039] The first processing module 130 and the second processing module 140 are connected via a keep-alive bus.
[0040] When any one of the first processing module 130 and the second processing module 140 is working, the other processing module is not working; the working processing module is used to control the distributed computing module 110 to perform computing tasks.
[0041] Specifically, the distributed computing module 110 includes several computing modules 111. These computing modules 111 can be intelligent computing modules suitable for intelligent analysis of video data. Computing modules 111 can also be encryption and decryption modules suitable for data encryption and decryption. In this embodiment, there is no specific limitation on computing modules 111.
[0042] Among them, the first processing module 130 acts as the main CPU management module during operation and executes the main task, which includes receiving computing tasks, such as intelligent analysis tasks; and allocating and scheduling the computing tasks to the computing module 111 in the distributed computing module. The computing tasks include information such as task type and stream pulling address, so that the computing module 111 obtains the data to be calculated, such as video data, according to the stream pulling address; controls the computing module 111 to process the data to be calculated based on the computing task, thereby completing the computing task of the computing module 111; and loads software such as the running program and the algorithm warehouse on the computing module 111.
[0043] While first processing module 130 is operating, second processing module 140 serves as a backup CPU management module, performing backup tasks. This backup task involves communicating with first processing module 130 via a keep-alive bus using a proprietary protocol to monitor the normal operation of first processing module 130. The backup CPU management module does not perform primary tasks.
[0044] When the second processing module 140 detects that the first processing module 130 is not operating normally, the second processing module 140 will replace the first processing module 130 and take over the current work of the first processing module 130 (i.e., the main task), such as docking with the platform, performing business management, monitoring computing module information, and performing business allocation, scheduling, and migration.
[0045] After the first processing module 130 recovers from an abnormality, it can continue to function as the primary CPU management module, taking over the primary tasks from the second processing module 140. Alternatively, it can function as the backup CPU management module for the second processing module 140, performing backup tasks and detecting the normal operation of the second processing module 140 via the keep-alive bus. The specific execution method after abnormality recovery can be pre-configured based on actual usage.
[0046] The network switching module 120 is a hardware device that implements data packet switching. It typically includes multiple ports, can connect to multiple devices, and supports high-speed data transmission. Its core function is to monitor data flows in real time and forward packets to the corresponding target device based on their destination MAC addresses. Network switching modules 120 can be selected, but are not limited to, Ethernet switching modules and fiber optic switching modules.
[0047] In this embodiment, the first processing module 130 and the second processing module 140 are connected to the distributed computing module through a management bus; the first processing module 130 and the second processing module 140 are also connected to the distributed computing module 110 through the network switching module 120; the first processing module 130 and the second processing module 140 are connected through a keep-alive bus; the processing module when working is used to control the distributed computing module 110 to perform computing tasks; this solves the problem that under the traditional solution, the CPU and the intelligent module are interconnected by the PCIE bus to form a master-slave relationship, resulting in low computing efficiency and poor reliability, and realizes the interconnection of the distributed computing module 110 with the dual processing modules through the network switching module 120, thereby improving the performance of the computing server.
[0048] In some of these embodiments, Figure 3 As shown, the distributed computing module 110 includes several computing modules 111 and an expansion board 112; each computing module 111 is connected to the expansion board 112 via a connector.
[0049] Specifically, each computing module 111 is hard-connected to the expansion board 112 via a connector, and the connector ensures that current or signals between the computing module 111 and the expansion board 112 can be transmitted stably and efficiently.
[0050] In some of these embodiments, Figure 3 As shown, the expansion board 112 includes a first management interface, a second management interface, and an expansion interface. The first management interface is connected to the first processing module 130 via a first out-of-band management bus; the second management interface is connected to the second processing module 140 via a second out-of-band management bus. The expansion interface is used to connect to each computing module 111.
[0051] Specifically, the first processing module 130 and the second processing module 140 control the distributed computing module 110 through different out-of-band management buses respectively without interfering with each other. When an abnormality occurs in one of the processing modules, the other processing module can take over the distributed computing module 110 to improve server reliability.
[0052] In some embodiments, the expansion interface includes a hot-swap interface.
[0053] Specifically, the expansion board 112 supports hot plugging, which enables the computing module 111 to be plugged in or out without turning off the power without affecting the normal operation of the server, thereby improving the reliability of the server and facilitating expansion, maintenance and management.
[0054] In some embodiments, the computing module 111 is an intelligent analysis module for video data.
[0055] Specifically, each intelligent analysis module is an independent subsystem that directly pulls data streams (video streams) from the front-end network camera through the network switching module 120 and performs tasks such as video decoding, intelligent analysis, and encoding. The analysis results are forwarded to the platform server 200 through the network switching module 120 and can also be forwarded to the first processing module 130 (i.e., the main CPU management module) for storage.
[0056] In some of the embodiments, the processing module is further configured to schedule another computing module 111 to continue executing computing tasks when one computing module 111 fails.
[0057] Specifically, during business operation, the working processing module, namely the main CPU management module, obtains the operating status of each computing module 111 in the distributed computing module 110 through the out-of-band management bus. When a computing module 111 malfunctions, the main CPU management module migrates the stream pulling address and computing tasks to other computing modules 111 to ensure the continuity of the entire machine business and improve product reliability.
[0058] In some of these embodiments, Figure 3 As shown, the distributed dual-control server 100 also includes a storage module 150, which is connected to the first processing module 130 and the second processing module 140. The storage medium of the storage module 150 is used to store content such as algorithm silos and program packages. When business needs arise, the distributed computing module 110 can forward some business data to the active processing modules for storage in the storage module 150.
[0059] In some embodiments, the network switching module 120 includes a network topology module based on the Ethernet protocol.
[0060] Specifically, network switching is based on the Ethernet protocol and uses network protocols such as TCP / IP for data transmission. The topology is a star, tree, or mesh topology formed by using switches or routers. It supports hot plugging, has no master-slave relationship, and all connected devices are equal, with data packets freely transmitted between devices. Traditional PCIE switching is based on the high-speed serial computer expansion bus standard (PCI Express, PCIe), uses a point-to-point high-speed serial communication protocol, has a master-slave relationship, and the master is defined as the Root Port and the slave device is defined as the End Point. Compared with traditional PCIE switching, the network switching module 120 can enhance the flexibility and scalability of distributed computing server data transmission.
[0061] In this embodiment, a distributed dual-control system is provided. Figure 4 As shown, the system includes: a platform server 200, a data source end 300, and the distributed dual-control server 100 in any of the above embodiments; the platform server 200 and the data source are respectively connected to the network switching module 120 of the distributed dual-control server 100. Preferably, the data source end 300 includes a network camera.
[0062] Specifically, the first processing module 130 in the distributed dual-control server 100 communicates with the platform server 200 through the network switching module 120, interfacing with the platform and providing services such as intelligence, compression, coding, and OSD (overlaying text on video). The distributed computing module 110 in the distributed dual-control server 100 pulls video streams from network cameras through the network switching module 120 to perform computational processing on the video.
[0063] In this preferred embodiment, computing modules 111 within distributed computing module 110 are independent subsystems that communicate via network interconnection through network switching module 120, forming a distributed system. First processing module 130 and second processing module 140 form a dual-control system. If any processing module experiences an anomaly, device services remain uninterrupted. Furthermore, if any module experiences an anomaly, services are dispatched to other modules, ensuring continuous device service.
[0064] The present embodiment is described and illustrated below through preferred embodiments.
[0065] Figure 5 This is a schematic diagram of the structure of the preferred embodiment. Figure 5, the distributed dual-control server 100 is applied to the intelligent analysis scenario, forming a distributed dual-control intelligent server all-in-one machine, which includes: a distributed computing module 110, a network switching module 120, a first processing module 130, a second processing module 140 and a storage medium. The distributed computing module 110 includes several intelligent computing modules and intelligent expansion boards. The first processing module 130 and the second processing module 140 both serve as CPU management modules, wherein the first processing module 130 is the main CPU management module and the second processing module 140 is the backup CPU management module, which can perform master-slave switching when an abnormality occurs. In this preferred embodiment, the intelligent computing module, the network switching module 120, and the CPU management module (including the main and backup) are interconnected through a network bus to form a distributed computing system. Under the traditional solution, such as Figure 1 As shown, the CPU and the intelligent module are interconnected by a PCIE bus, forming a master-slave relationship.
[0066] The CPU management module communicates with the platform server 200 through the network switch module 120, connecting to the platform and providing services such as intelligence, compression, coding, and OSD. The CPU management module is also responsible for managing the device chassis, managing the intelligent modules, assigning tasks to the intelligent modules, running programs for the intelligent computing modules, and loading software such as the algorithm warehouse.
[0067] Each intelligent computing module is hard-wired to the intelligent expansion board via a connector and is hot-swappable. Each intelligent computing module is an independent subsystem, pulling data streams (video streams) directly from the front-end network camera IPC via the network switching module 120 and performing tasks such as video decoding, intelligent analysis, and encoding. The analysis results are forwarded to the platform server 200 and can also be forwarded to the CPU management module for storage.
[0068] The control method of the distributed dual-control server 100 is as follows:
[0069] When the device is powered on, the main CPU management module is enabled by default. It first configures the network switch module 120 and simultaneously obtains information about the intelligent computing modules (such as the number of intelligent modules and slot locations) on the intelligent expansion board through the out-of-band management bus. The main CPU management module then performs port mapping between the intelligent computing modules and the network switch module 120.
[0070] The main CPU management module loads system programs and software such as the algorithm warehouse to the intelligent computing module through the network, and the intelligent computing module runs independently as a subsystem.
[0071] The main CPU management module performs business management, receives intelligent tasks sent from the platform, and dispatches tasks to various intelligent computing modules, including assigning a streaming address to each intelligent module.
[0072] The intelligent module executes the video streaming task based on the address, decodes, intelligently analyzes, and encodes the video stream, generating intelligent data streams, images, events, etc., and returns these data streams to the platform. At the same time, the intelligent module can also forward some data to the CPU for local storage if business needs require it.
[0073] During the operation of these services, the main CPU management module obtains the operating status of each intelligent computing module through the main out-of-band management bus. If an intelligent computing module experiences an abnormality, the main CPU management module migrates the stream pull address and intelligent tasks to other intelligent computing modules, ensuring the continuity of the entire intelligent service and improving product reliability.
[0074] The primary and backup CPU management modules maintain keepalive communication over the network. The backup CPU management module monitors the primary CPU management module's operation over the network. If the primary CPU management module experiences an anomaly, the backup CPU module connects to the platform and manages services. Simultaneously, the backup CPU module monitors information from each intelligent computing module via an out-of-band management bus, enabling service allocation, scheduling, and migration.
[0075] In this preferred embodiment, the active and standby CPU management modules and each intelligent computing module are divided into independent subsystems and interconnected through network switching module 120 to form a distributed system. The active and standby CPU modules form a dual-control system, so if any CPU management module experiences an error, device services are not interrupted. The intelligent computing modules form a distributed cluster, so if any module experiences an error, the service is dispatched to another module, ensuring continuous device service.
[0076] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0077] Obviously, the accompanying drawings are merely examples or embodiments of the present application. A person skilled in the art can also apply the present application to other similar situations based on these drawings without inventive effort. Furthermore, it is understandable that, although the work involved in this development process may be complex and lengthy, certain design, manufacturing, or production changes based on the technical content disclosed in this application are merely routine technical means for a person skilled in the art and should not be considered to constitute a deficiency in the disclosure of the present application.
[0078] The term "embodiment" as used in this application refers to specific features, structures, or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily mean that the embodiment is the same, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is understood, either explicitly or implicitly, by those skilled in the art that the embodiments described in this application can be combined with other embodiments when there is no conflict.
[0079] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A distributed dual-control server, characterized in that: include: A distributed computing module, a network switching module, a first processing module, and a second processing module; The first processing module and the second processing module are connected to the distributed computing module via a management bus; The first processing module and the second processing module are further connected to the distributed computing module via the network switching module; The first processing module and the second processing module are connected via a keep-alive bus; When any one of the first processing module and the second processing module is working, the other processing module is not working; the working processing module is used to control the distributed computing module to perform computing tasks.
2. The distributed dual-control server according to claim 1, characterized in that: The distributed computing module includes several computing modules and expansion boards; each computing module is connected to the expansion board via a connector.
3. The distributed dual-control server according to claim 2, characterized in that: The expansion board includes a first management interface, a second management interface and an expansion interface; The first management interface is connected to the first processing module via a first out-of-band management bus; The second management interface is connected to the second processing module via a second out-of-band management bus; The expansion interface is used to connect each of the computing modules.
4. The distributed dual-control server according to claim 3, characterized in that: The expansion interface includes a hot plug interface.
5. The distributed dual-control server according to claim 2, characterized in that: The computing module is an intelligent analysis module for video data.
6. The distributed dual-control server according to claim 2, characterized in that: The processing module when working is also used to schedule another computing module to continue to execute the computing task when one computing module fails.
7. The distributed dual-control server according to claim 1, characterized in that: Also includes: Storage module; The storage module is connected to the first processing module and the second processing module respectively.
8. The distributed dual-control server according to claim 1, characterized in that: The network switching module includes a network topology module based on the Ethernet protocol.
9. A distributed dual-control system, characterized in that: The system comprises: a platform server, a data source end, and a distributed dual-control server according to any one of claims 1 to 8; The platform server and the data source are respectively connected to the network switching modules of the distributed dual-control servers.
10. The distributed dual-control system according to claim 9, characterized in that: The data source includes a network camera.