A method and system for invocation and modeling training of data center resources

By deploying distribution assistants and control panels in the data center, and utilizing identifier verification and time window management, the problems of low resource utilization and training task stability in data center resource management are solved, achieving efficient and reliable model training.

CN122195600APending Publication Date: 2026-06-12CHONGQING HIKE NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING HIKE NETWORK TECH CO LTD
Filing Date
2026-03-04
Publication Date
2026-06-12

Smart Images

  • Figure CN122195600A_ABST
    Figure CN122195600A_ABST
Patent Text Reader

Abstract

The application discloses a kind of calling and modeling training method and system of data center resources, it is related to data center resource management technical field.The method includes the following steps: in data center deployment distribution helper, obtains training data, distribution helper is according to preset segmentation rule to training data is segmented, obtain multiple data blocks, distribution helper distributes corresponding bearing space for each data block, and bearing space stores data block;For each data block generates unique identifier, and the identifier of data block is bound with bearing space;Wherein, a plurality of connection ports are configured to bearing space;The application can avoid the data block in the same bearing space in the same time window is simultaneously called by multiple steering wheels corresponding sub-demand, effectively avoid the problem, such as insufficient memory, load is too high or data conflict, caused by multiple steering wheels competition data block in the same bearing space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data center resource management technology, specifically to a method and system for data center resource retrieval and modeling training. Background Technology

[0002] In the field of artificial intelligence, model training based on large-scale datasets is a core task, but its efficiency and flexibility have long been limited by the inherent defects of traditional data center resource management models. Existing technologies typically treat the training task as a whole, processing it centrally on a statically allocated pool of computing resources, resulting in low resource utilization, such as idle GPUs during data loading or insufficient CPU load during model computation. Furthermore, the centralized storage and management of massive amounts of training data easily creates I / O bottlenecks, slowing down the overall training process.

[0003] In addition, the traditional model training process has two major problems: First, when multiple training subtasks run in parallel, it is not easy to ensure that each subtask can accurately access its specified data block, which poses a risk of incorrect connections or data confusion. Second, when multiple training subtasks need to call the same data block within the same time period, it is easy to trigger concurrent access to the storage node (hosting space) that carries the data block, which can easily lead to insufficient memory, excessive load, or even process crash or abnormal training results, seriously reducing the stability and efficiency of the model training task. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for calling and modeling training data center resources, so as to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for accessing and modeling data center resources, comprising the following steps: A distribution assistant is deployed in the data center to acquire training data. The distribution assistant divides the training data into multiple data blocks according to preset splitting rules. The distribution assistant distributes a corresponding carrying space to each data block, and the carrying space stores the data block. A unique identifier is generated for each data block, and the identifier of the data block is bound to the carrying space. Multiple connection ports are configured for the carrying space. The distribution assistant receives the training requirements of the model and distributes a corresponding training space according to the training requirements. The distribution assistant splits the training requirements into multiple sub-requirements based on preset splitting rules. The distribution assistant distributes a control panel to each sub-requirement, and the control panel is located within the training space. Multiple carrier spaces corresponding to the sub-requirements are determined, and the same number of sub-ports are configured on the control panel corresponding to the sub-requirements based on the corresponding multiple carrier spaces. The control panel is connected one-to-one with the multiple carrier spaces corresponding to the sub-requirements through the same number of sub-ports to obtain multiple data channels. Among them, the distribution assistant copies the identifier of the carrier space to the control panel, and the control panel verifies the consistency between the copied identifier and the identifier of the carrier space. When the two are consistent, the verification is completed, and a data channel is established between the two. The model is trained using multiple control panels; multiple time windows are set, and the multiple control panels are arranged in the multiple time windows according to the priority of the sub-requirements, and the multiple sub-requirements are trained according to the arrangement order of the multiple time windows. The training requirements of the model are updated in real time. When the training requirements change, the corresponding control panel is updated.

[0006] In a preferred embodiment, the step of deploying a distribution assistant in the data center to acquire training data involves the distribution assistant segmenting the training data according to a preset segmentation rule to obtain multiple data blocks, including: Deploy the distribution assistant on the resource management node of the data center; The data center obtains training data through data interfaces or data upload channels; The preset segmentation rules allow the distribution assistant to segment the training data into multiple data blocks. The segmentation rules can be based on one of the following: data size, data type, data access frequency, or semantic logic.

[0007] In a preferred embodiment, establishing the data channel includes: The distribution assistant uses preset splitting rules to split training requirements into multiple sub-requirements; the splitting rules include splitting training requirements according to training stage or task priority. The control panel connects one-to-one with the connection port of the corresponding carrier space through the sub-port and establishes a data channel. The distribution assistant copies the identifier of the carrier space to the control panel. The control panel verifies the consistency between the copied identifier and the identifier of the carrier space. When they match, the verification is completed and a data channel is established between them.

[0008] In a preferred embodiment, the model is trained using multiple control panels; wherein multiple time windows are set, and the multiple control panels are arranged within the multiple time windows according to the priority of the sub-requirements, and the multiple sub-requirements are trained according to the arrangement order of the multiple time windows, including: Obtain the priority order of multiple sub-requirements, and arrange the control panels corresponding to the multiple sub-requirements in sequence based on the priority order to obtain the arrangement chain; Multiple time windows are set according to the training requirements of the corresponding model, and the multiple time windows are arranged in sequence to obtain the implementation chain. The multiple time windows are numbered according to the arrangement order, and each time window is bound one-to-one with the corresponding number to obtain the performance parameters of the training space. The maximum carrying capacity of the time window is set based on the performance parameters. Configure point layouts for multiple carrying spaces on a one-to-one basis, and assign multiple control panels to multiple time windows; The time windows of the multiple control panels that have been allocated are used to train the corresponding sub-requirements according to the implementation chain.

[0009] In a preferred embodiment, the step of configuring point panels on a one-to-one basis for multiple carrying spaces and allocating multiple control panels to multiple time windows includes: The multiple connection ports that connect the carrying space to the sub-ports are configured with points one-to-one. The multiple points are then connected to obtain the point layout corresponding to the carrying space. The connection ports are associated with the corresponding configured points, and the control panel is associated with the points and point layouts connected to the corresponding connection ports. Multiple control panels are allocated to corresponding time windows according to the arrangement chain and maximum carrying capacity. During the allocation process, multiple control panels carry associated point panels to the corresponding time windows, and multiple point panels are placed in the matching area within the time window according to the arrangement chain. The matching area is used to store multiple point panels.

[0010] In a preferred embodiment, training the time windows of the allocated multiple control panels to the corresponding sub-requirements according to the implementation chain includes: When multiple control panels are trained on sub-requirements within the same time window, the multiple control panels copy their associated points and return the associated points in their original positions. The control panel carries the copied points and the data channel points for consistency verification. When they match, the verification is completed, the control panel activates the data channel, and the control panel then trains the sub-requirements through the corresponding carrying space.

[0011] In a preferred embodiment, the real-time updating of the model's training requirements, when the training requirements change, involves updating the control panel corresponding to the training requirements, including: Monitor the model's training requirements in real time and determine whether the training requirements have changed; If so, mark the changes in training requirements, obtain the sub-requirements corresponding to the changes, and update the control panel corresponding to the sub-requirements based on the changes; wherein, updating the control panel includes adding or deleting sub-ports connected to the carrying space. If not, there is no need to update the control panel corresponding to the training requirements.

[0012] This invention also provides a data center resource retrieval and modeling training system, comprising: The first distribution module is used to deploy a distribution assistant in the data center to obtain training data. The distribution assistant divides the training data into multiple data blocks according to a preset segmentation rule. The distribution assistant distributes a corresponding carrying space to each data block, and the carrying space stores the data block. A unique identifier is generated for each data block, and the identifier of the data block is bound to the carrying space. Multiple connection ports are configured for the carrying space. The second distribution module is used to receive the training requirements of the model from the distribution assistant and distribute a corresponding training space according to the training requirements of the model. The distribution assistant splits the training requirements into multiple sub-requirements based on preset splitting rules. The distribution assistant distributes a control panel for each sub-requirement, and the control panel is located in the training space. Multiple carrier spaces corresponding to the sub-requirements are determined, and the same number of sub-ports are configured on the control panel corresponding to the sub-requirements based on the corresponding multiple carrier spaces. The control panel is connected one-to-one with the multiple carrier spaces corresponding to the sub-requirements through the same number of sub-ports to obtain multiple data channels. Among them, the distribution assistant copies the identifier of the carrier space to the control panel, and the control panel verifies the consistency between the copied identifier and the identifier of the carrier space. When the two are consistent, the verification is completed, and a data channel is established between the two. The training module is used to train the model using multiple control panels; it sets multiple time windows, arranges the multiple control panels in the multiple time windows according to the priority of the sub-requirements, and trains the multiple sub-requirements according to the arrangement order of the multiple time windows. The update module is used to update the model's training requirements in real time. When the training requirements change, the corresponding control panel is updated.

[0013] The technical effects and advantages provided by the present invention in the above technical solution are as follows: 1. This invention performs identifier consistency verification between sub-ports and connection ports, thereby preventing the control panel from accessing data blocks in other carrying spaces. This allows the control panel to accurately use the data blocks in the corresponding carrying spaces to train sub-requirements during subsequent model training, thus enhancing the data reliability of the entire model training process. 2. By setting the point layout, this invention can prevent data blocks in the same carrying space from being called by multiple control panels corresponding to sub-requirements at the same time window, effectively avoiding problems such as insufficient memory, excessive load or data conflict caused by multiple control panels competing for data blocks in the same carrying space. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0015] Figure 1 This is a flowchart of the method of the present invention.

[0016] Figure 2 This is a schematic diagram of the control panel structure in the method of the present invention.

[0017] Figure 3 This is a system block diagram of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Example 1, please refer to Figure 1 and Figure 2 As shown in this embodiment, a method for calling and modeling training data center resources includes the following steps: A distribution assistant is deployed in the data center to acquire training data. The distribution assistant divides the training data into multiple data blocks according to preset splitting rules. The distribution assistant distributes a corresponding carrying space to each data block, and the carrying space stores the data block. A unique identifier is generated for each data block, and the identifier of the data block is bound to the carrying space. Multiple connection ports are configured for the carrying space. The distribution assistant receives the training requirements of the model and distributes a corresponding training space according to the training requirements. The distribution assistant splits the training requirements into multiple sub-requirements based on preset splitting rules. The distribution assistant distributes a control panel to each sub-requirement, and the control panel is located within the training space. Multiple carrier spaces corresponding to the sub-requirements are determined, and the same number of sub-ports are configured on the control panel corresponding to the sub-requirements based on the corresponding multiple carrier spaces. The control panel is connected one-to-one with the multiple carrier spaces corresponding to the sub-requirements through the same number of sub-ports to obtain multiple data channels. Among them, the distribution assistant copies the identifier of the carrier space to the control panel, and the control panel verifies the consistency between the copied identifier and the identifier of the carrier space. When the two are consistent, the verification is completed, and a data channel is established between the two. The model is trained using multiple control panels; multiple time windows are set, and the multiple control panels are arranged in the multiple time windows according to the priority of the sub-requirements, and the multiple sub-requirements are trained according to the arrangement order of the multiple time windows. The training requirements of the model are updated in real time. When the training requirements change, the corresponding control panel is updated. In this embodiment, the consistency of identifiers is checked between the sub-port and the connection port, thereby preventing the control panel from accessing data blocks in other carrying spaces. This allows the control panel to accurately use the data blocks in the corresponding carrying space to train the sub-requirements during subsequent model training, thus enhancing the data reliability of the entire model training process. By setting the point layout, it is possible to prevent data blocks in the same carrying space from being called simultaneously by the sub-requirements corresponding to multiple control panels within the same time window, effectively avoiding problems such as insufficient memory, excessive load, or data conflicts caused by multiple control panels competing for data blocks in the same carrying space.

[0020] In one embodiment, the step of deploying a distribution assistant in the data center to acquire training data involves the distribution assistant segmenting the training data according to a preset segmentation rule to obtain multiple data blocks, including: Deploy the distribution assistant on the resource management node of the data center; The data center obtains training data through data interfaces or data upload channels; The distribution assistant uses preset segmentation rules to segment the training data into multiple data blocks. The segmentation rules are based on one of the following: data size, data type, data access frequency, and semantic logic. Specifically, the identifier includes a data header, a data footer, and a main content tag for each data block. The main content tag for each data block is obtained based on semantic extraction. A distribution assistant is deployed in the data center. The distribution assistant is essentially developed based on Python and integrates resource scheduling interfaces for data reading and segmentation. An index library is built for the corresponding distribution assistant in the data center. For example, a data center needs to train a ResNet50 model for industrial product defect classification. It obtains 50GB of product image data through a data interface or data upload channel. The product image data is in JPG format and includes three types of defects: cracks, deformation, and scratches. The training data can be segmented in several ways: the 50GB image data is divided into 25 2GB data blocks according to data size, named blocks 1 to 25; or blocks 1 to 20 are divided into training set data blocks according to data type. Blocks 21 to 25 are divided into validation set data blocks; high-frequency access core defect samples, such as blocks 1 to 10, are divided into 1GB blocks according to data access frequency, and low-frequency supplementary samples, such as blocks 11 to 25, are divided into 2GB blocks; based on semantic logic, they are divided by defect type: blocks 1 to 8 are crack defect data, blocks 9 to 16 are deformation defect data, and blocks 17 to 25 are scratch defect data; for example, an identifier is generated for block 1, with the identifier format being data header|source|size|timestamp|content tag|CRC32 checksum, for example, 001|defect dataset-20251128|1GB|202511281000|image classification-industrial defect-crack-training set|6A8F2D4E, this identifier is bound to the carrying space of block 1 and stored in the distribution assistant's index database; The specific process of distributing the hosting space is as follows: The distribution assistant sends a request to the data center to allocate one virtual machine as a hosting space to each of data blocks 1 to 25. The hosting space is essentially a virtual machine. Multiple connection ports are configured for each hosting space. By deploying the distribution assistant in the data center, the distribution assistant can dynamically allocate appropriate hosting spaces according to the size, type, and access requirements of the data blocks. By dividing the training data according to size, type, access frequency, etc., the data blocks are made to better fit the actual training requirements. By configuring multiple connection ports for the hosting space, it is easy for one hosting space to connect with multiple control panels, and it is easy for the data blocks in the hosting space to support the training of multiple control panels corresponding to the sub-requirements. Specifically, the distribution assistant dynamically allocates suitable hosting space based on the size, type, and access requirements of data blocks through the following steps: 1) Constructing a hosting space resource pool; 2) Classifying virtual machines in the data center according to hardware, storage, and network attributes; 3) Simultaneously establishing an evaluation benchmark table to determine resource thresholds and recommended hosting space types corresponding to different data block characteristics; 4) Extracting features such as size and type from data block identifiers, combining them with access frequency requirements, and converting them into quantitative resource requirements according to the benchmark table; 5) Screening candidate hosting spaces within the resource pool according to quantitative requirements, and determining the optimal candidate based on resource matching degree, load, and cost priority; 6) Sending a request to the resource scheduling module to complete hosting space allocation, binding data block identifiers with hosting space information and updating the index database, and verifying integrity after migrating data blocks.

[0021] In one embodiment, establishing the data channel includes: The establishment of the data channel includes: The distribution assistant uses preset splitting rules to split training requirements into multiple sub-requirements; the splitting rules include splitting training requirements according to training stage or task priority. The control panel connects one-to-one with the connection port of the corresponding carrier space through the sub-port and establishes a data channel. The distribution assistant copies the identifier of the carrier space to the control panel. The control panel verifies the consistency between the copied identifier and the identifier of the carrier space. When they match, the verification is completed and a data channel is established between them.

[0022] Specifically, the resource allocation for the training space is as follows: The distribution assistant receives the user's training requirements (training tasks) through the web interface submitted by the data center: the model type is ResNet50, the training objective is three-class classification of industrial defects, the resource requirements are 2 GPUs, 8GB of memory, and the data range is block 1 to block 25; the distribution assistant requests a virtual machine with an 8-core CPU, 10GB of memory, 2 V100 GPUs, and 50GB of storage from the data center as the training space based on the training requirements. Training requirements can be broken down in several ways: by dividing them into four sub-requirements (sub-tasks) according to the training stage, sub-requirement A corresponds to data preprocessing, sub-requirement B corresponds to feature extraction, sub-requirement C corresponds to parameter iteration, and sub-requirement D corresponds to result verification; by task priority, sub-requirement A is defined as high priority, sub-requirement B as medium to high priority, sub-requirement C as medium to low priority, and sub-requirement D as low priority. The data channel establishment process is as follows: The distribution assistant copies the identifier of the carrier space to the control panel. The control panel verifies the consistency between the copied identifier and the identifier of the carrier space. When they match, the verification is complete, and a data channel is established between them. For example, the distribution assistant generates control panels A, B, C, and D for four sub-requirements within the training space. The control panels run within the training space and are essentially virtual machines. For example, if there are 10 carrier spaces related to the training of sub-requirement A, and sub-requirement A corresponds to control panel A with 10 sub-ports, these 10 sub-ports are connected to the connection ports of the corresponding 10 carrier spaces. The distribution assistant... The identifier of the carrier space is copied to the corresponding sub-port of the corresponding control panel (e.g., the identifier is copied to the sub-port of control panel A, through which the control panel connects to the carrier space). When control panel A starts training, it initiates a connection request to the connection port of the carrier space through the sub-port and sends the identifier copied by the distribution assistant (the connection port corresponds to the sub-port). The connection port reads the identifier bound to its carrier space and compares it with the received identifier. After the two match, the verification is successful, and a dedicated data channel is established between the sub-port of control panel A and the connection port of the carrier space. At this time, the data channel is in an inactive state. By breaking down training requirements through a distribution assistant, computing resources can be precisely allocated according to sub-requirements (e.g., configuring different numbers of sub-ports for different control panels). By splitting complex training requirements into training stages and data shards, training requirements (training tasks) can be executed modularly. Establishing one-to-one data channels between the control panel's sub-ports and the connection port of the carrier space facilitates the control panel's access to data blocks within the carrier space, thereby facilitating the training of sub-requirements. Furthermore, by performing identifier consistency checks between sub-ports and connection ports, the control panel is prevented from accessing data blocks in other carrier spaces. This ensures that the control panel can accurately utilize the corresponding data blocks within the carrier space to train sub-requirements during subsequent model training, thus enhancing the data reliability of the entire model training process. It should be noted that training requirements are user-submitted model training task definitions, which include model type and objectives, required resources (such as GPU and memory), data range (such as data block numbers), and task execution strategy (a sequence of sub-requirements broken down by training stage or priority). As the starting point of the training process, it is received by the distribution assistant, which automatically allocates training space, generates control panels, and arranges time windows, ultimately forming an executable modular training pipeline. This requirement supports dynamic updates, and the system can adjust the configuration in real time through polling to achieve iteration.

[0023] In one embodiment, the model is trained using multiple control panels; wherein multiple time windows are set, and the multiple control panels are arranged within the multiple time windows according to the priority of the sub-requirements, and the multiple sub-requirements are trained according to the arrangement order of the multiple time windows, including: Obtain the priority order of multiple sub-requirements, and arrange the control panels corresponding to the multiple sub-requirements sequentially based on the priority order to obtain the arrangement chain; wherein, the priority order is arranged based on the urgency of the sub-requirements; Multiple time windows are set according to the training requirements of the corresponding model, and the multiple time windows are arranged in sequence to obtain the implementation chain. The multiple time windows are numbered according to the arrangement order, and each time window is bound one-to-one with the corresponding number to obtain the performance parameters of the training space. The maximum carrying capacity of the time window is set based on the performance parameters. Configure point layouts for multiple carrying spaces on a one-to-one basis, and assign multiple control panels to multiple time windows; The time windows of the multiple control panels that have been allocated are used to train the corresponding sub-requirements according to the implementation chain.

[0024] In one embodiment, configuring point layouts one-to-one for multiple carrying spaces and allocating multiple control panels to multiple time windows includes: The multiple connection ports that connect the carrying space to the sub-ports are configured with points one-to-one. The multiple points are then connected to obtain the point layout corresponding to the carrying space. The connection ports are associated with the corresponding configured points, and the control panel is associated with the points and point layouts connected to the corresponding connection ports. Multiple control panels are allocated to corresponding time windows according to the arrangement chain and maximum carrying capacity. During the allocation process, multiple control panels carry associated point panels to the corresponding time windows, and multiple point panels are placed in the matching area within the time window according to the arrangement chain. The matching area is used to store multiple point panels.

[0025] In one embodiment, training the time windows of the allocated multiple control panels to the corresponding sub-requirements according to the implementation chain includes: When multiple control panels are trained on sub-requirements within the same time window, the multiple control panels copy their associated points and return the associated points in their original positions. The control panel carries the copied points and the data channel points to verify consistency. When they are consistent, the verification is completed, the control panel activates the data channel, and the control panel then trains the sub-requirements through the corresponding carrying space. Specifically, the control panel is arranged sequentially according to the priority of sub-requirements, such as the urgency of sub-requirements: sub-requirement 1 > sub-requirement 2 > sub-requirement 3 > sub-requirement 4, forming a chain of control panel arrangements. This ensures that core tasks are executed first. Sequential arrangement here means arranging sub-requirements with high urgency first, and then arranging sub-requirements with low urgency, thereby ensuring that core tasks are executed first. Based on the total training duration and the performance parameters of the training space, N time windows are set. Performance parameters include the number of CPU cores, GPU computing power, and memory bandwidth. Each time window corresponds to a continuous period of time and the performance parameters of the training space. The N time windows are sequentially numbered (1 to N) to form an implementation chain. The maximum capacity of each time window is set according to the performance parameters of the training space. The method for setting N time windows based on the total training duration and the performance parameters of the training space is as follows: Based on the iterative characteristics of the training task, the total training duration is divided into N consecutive scheduling cycles of equal length, each cycle being a time window. The duration of a single time window can be set to a fixed value (e.g., 5 minutes) based on experience or experimentation, or set to an integer multiple of the time required to complete a standard training batch. The method for setting the maximum capacity of each time window based on the performance parameters of the training space is as follows: The maximum capacity refers to the highest number of concurrently executed load subtasks allowed within a time window. Its specific value depends on the type of control panel (load level) scheduled into that time window. The total resource capacity of the training space is dynamically determined. The determination principle is to ensure that the sum of the peak resource requirements (such as GPU and memory) of all concurrent control disks within the window does not exceed the total resource capacity corresponding to the training space. For example, if a high-load control disk needs to exclusively occupy 1 GPU, and there are 2 GPUs in the training space, then the maximum carrying capacity of the high-load control disk in this time window is 2; if a medium-load control disk is included, its resource requirements are halved, and the maximum carrying capacity can be increased to 4 accordingly.

[0026] The process of allocating multiple control panels to time windows is as follows: When allocating control panels, multiple control panels are sequentially assigned to time windows according to the arrangement chain. The number of control panels that can be allocated in each time window is set according to the maximum carrying capacity of the training space. If the maximum carrying capacity of the training space has been reached after arranging 3 control panels in the current time window, the remaining control panels are arranged in the next time window according to the arrangement chain, and so on, until all control panels are allocated. Finally, the sub-requirements corresponding to the control panels in each window are executed sequentially according to the implementation chain of the time windows. The next time window can only start after the training of the previous time window is completed, so as to avoid resource conflicts. Among them, the corresponding sub-requirements are executed through the control panels, and the training requirements of the model are completed based on the data blocks, thereby realizing the model modeling and training. By configuring points one-to-one on all connection ports of each carrier space, each point is essentially a trust token, such as an encrypted string, and each connection port corresponds to a unique point. By configuring points one-to-one on multiple connection ports (connection ports connected to different control panels) on each carrier space that are connected to sub-ports, multiple points are connected to obtain a point layout. The point layout is a set of trust tokens, including all trust tokens connected to all control panels on the corresponding carrier space. The control panel connects to the carrier space through the data channel between the sub-port and the corresponding connection port, and associates it with the points of the corresponding connection port and the point layout of the carrier space corresponding to the connection port. This makes it convenient for the control panel to carry the point layout in order to call the data blocks in the carrier space, and to use this to train on sub-requirements. When training the sub-requirements corresponding to the control panel within the same time window, the control panel first copies its associated points (trust tokens) and returns the associated points to the connection port of the corresponding carrying space in their original positions to confirm that the points have not been tampered with. The control panel carries the copied points through the data channel to perform consistency verification (such as hash value comparison) with the points returned in their original positions. Only when the copied points are consistent with the associated points will the control panel activate the data channel and read the data of the carrying space through the data channel to execute the training task of the corresponding sub-requirement. When the control panel completes the corresponding sub-requirement, the control panel returns the point layout in its original position (returns it to its original position). The control panel carries the point layout into the time window, while other control panels cannot carry the point layout. Therefore, by setting the point layout, it is possible to avoid data blocks in the same carrying space being called by the sub-requirements corresponding to multiple control panels at the same time window, effectively avoiding problems such as insufficient memory, excessive load, or data conflicts caused by multiple control panels competing for data blocks in the same carrying space. It should be noted that the principle behind the consistency verification between the copied points carried by the control panel and the points returned from the original location via the data channel is as follows: Essentially, it establishes a mutual exclusion lock for accessing the carrying space through a point flow rule of one-time use and reset to zero after use. Points are dynamically generated by the distribution assistant during carrying space allocation. These points are essentially encrypted strings, generated using a hash algorithm (such as SHA-256) to perform a one-way hash operation on the unique identifier of the carrying space, the connection port number, a random number, and a timestamp. The generated point structure is a 64-bit hexadecimal string, possessing three main characteristics: uniqueness, unpredictability, and tamper resistance. After generation, the points are securely stored in an index maintained by the distribution assistant and bound to a specific connection port-carrying space. When establishing a data channel, the distribution assistant copies the point corresponding to the connection port and securely sends it to the sub-port of the control panel via an encrypted channel, associating the control panel with that point. At this point, the control panel only holds a copy of the point, while the connection port of the carrying space holds the original point; the two constitute a verification pair. In-situ return is the process of temporarily migrating a point from the control panel back to the connection port of the carrier space for on-site verification. Its implementation relies on the bidirectional communication capability of the data channel. The specific process is as follows: When training starts, the control panel copies the associated point string from its local cache and sends the copied point as an authentication message to the connection port of the carrier space through the established, inactive data channel. Upon receiving the point, the connection port immediately reads the original point and performs hash value comparison using a constant-time comparison algorithm to prevent timing attacks. If the comparison matches, the corresponding data channel is activated; otherwise, verification fails, the connection port is locked, and a security alarm is triggered. The significance of in-situ return is that the point layout is returned directly to its carrier space without any intermediate nodes. Only when the data channel is activated is the control panel allowed to read data blocks within the carrier space and execute sub-requirement training tasks. This ensures that only the control panel holding the correct point can access a specific carrier space, preventing unauthorized or incorrect access. The essence of a point-to-point panel is the physical representation of access permissions, and its mutual exclusion mechanism is achieved through holding state and exclusivity. When a control panel activates the data channel between itself and the carrying space to which the point-to-point panel belongs, other control panels cannot carry the corresponding point-to-point panel, thus other control panels must wait. Among them, the matching area is a dedicated logical space defined within the time window. Its core basic function is to centrally store the point-to-point panels (i.e., the set of carrying space access trust tokens) associated with all control panels with pending sub-requirements within the current time window, according to the arrangement chain of control panels. It provides a unified storage container for the scattered point-to-point panels, avoiding the management chaos or loss problems caused by the scattered storage of point-to-point panels, and ensuring that the access credentials of each control panel can be quickly retrieved.

[0027] In one embodiment, the real-time updating of the model's training requirements, when the training requirements change, involves updating the control panel corresponding to the training requirements, including: Monitor the model's training requirements in real time and determine whether the training requirements have changed; If so, mark the changes in training requirements, obtain the sub-requirements corresponding to the changes, and update the control panel corresponding to the sub-requirements based on the changes; wherein, updating the control panel includes adding or deleting sub-ports connected to the carrying space. If not, there is no need to update the control panel corresponding to the training requirements; Specifically, the distribution assistant monitors the model's training needs (training tasks) through periodic polling once per minute. If it detects that a user has submitted changes to the training needs via the web interface (the changes being the differences between the latest and previous training needs): when the changes involve adding 1000 training data points with scratches or defects, the new training data is added to the data center. The new training data is then split into three data blocks (e.g., data block 26, data block 27, and data block 28). The distribution assistant distributes corresponding storage spaces to the three data blocks and configures multiple connection ports for each storage space. When these changes correspond to sub-needs A (i.e., initial...), If the training requirements are broken down into sub-requirements A, then the control panel A corresponding to sub-requirement A is updated based on these changes in requirements. The update process includes: adding 3 sub-ports to the control panel A, connecting the 3 sub-ports to the connection ports of the carrying spaces corresponding to data blocks 26 to 28 respectively, and establishing data channels for each; if the training requirements no longer include the training corresponding to a certain defect type, such as data block 12 corresponding to this defect type, then the carrying space where data block 12 is located is deleted, and at the same time, all sub-ports corresponding to the connection ports on this carrying space are deleted, thereby deleting the sub-ports on the control panel that are connected to the carrying space; If no change in demand is detected: the distribution assistant continues to monitor, the control panel runs according to the original configuration, and no updates are required; furthermore, the data center can respond to and process new training requests submitted by users in real time without interruption or downtime of model training requirements (training tasks), enabling the model to quickly adapt to changes in data distribution, greatly improving the agility of model iteration and training continuity.

[0028] Example 2, please refer to Figure 3 As shown in this embodiment, a data center resource retrieval and modeling training system includes: The first distribution module is used to deploy a distribution assistant in the data center to obtain training data. The distribution assistant divides the training data into multiple data blocks according to a preset segmentation rule. The distribution assistant distributes a corresponding carrying space to each data block, and the carrying space stores the data block. A unique identifier is generated for each data block, and the identifier of the data block is bound to the carrying space. Multiple connection ports are configured for the carrying space. The second distribution module is used to receive the training requirements of the model from the distribution assistant and distribute a corresponding training space according to the training requirements of the model. The distribution assistant splits the training requirements into multiple sub-requirements based on preset splitting rules. The distribution assistant distributes a control panel for each sub-requirement, and the control panel is located in the training space. Multiple carrier spaces corresponding to the sub-requirements are determined, and the same number of sub-ports are configured on the control panel corresponding to the sub-requirements based on the corresponding multiple carrier spaces. The control panel is connected one-to-one with the multiple carrier spaces corresponding to the sub-requirements through the same number of sub-ports to obtain multiple data channels. Among them, the distribution assistant copies the identifier of the carrier space to the control panel, and the control panel verifies the consistency between the copied identifier and the identifier of the carrier space. When the two are consistent, the verification is completed, and a data channel is established between the two. The training module is used to train the model using multiple control panels; it sets multiple time windows, arranges the multiple control panels in the multiple time windows according to the priority of the sub-requirements, and trains the multiple sub-requirements according to the arrangement order of the multiple time windows. The update module is used to update the model's training requirements in real time. When the training requirements change, the corresponding control panel is updated.

[0029] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for retrieving and modeling training data center resources, characterized in that, Includes the following steps: A distribution assistant is deployed in the data center to acquire training data. The distribution assistant divides the training data into multiple data blocks according to preset splitting rules. The distribution assistant distributes a corresponding carrying space to each data block, and the carrying space stores the data block. A unique identifier is generated for each data block, and the identifier of the data block is bound to the carrying space. Multiple connection ports are configured for the carrying space. The distribution assistant receives the training requirements of the model and distributes a corresponding training space according to the training requirements. The distribution assistant splits the training requirements into multiple sub-requirements based on preset splitting rules. The distribution assistant distributes a control panel to each sub-requirement, and the control panel is located within the training space. Multiple carrier spaces corresponding to the sub-requirements are determined, and the same number of sub-ports are configured on the control panel corresponding to the sub-requirements based on the corresponding multiple carrier spaces. The control panel is connected one-to-one with the multiple carrier spaces corresponding to the sub-requirements through the same number of sub-ports to obtain multiple data channels. Among them, the distribution assistant copies the identifier of the carrier space to the control panel, and the control panel verifies the consistency between the copied identifier and the identifier of the carrier space. When the two are consistent, the verification is completed, and a data channel is established between the two. The model is trained using multiple control panels; multiple time windows are set, and the multiple control panels are arranged in the multiple time windows according to the priority of the sub-requirements, and the multiple sub-requirements are trained according to the arrangement order of the multiple time windows. The training requirements of the model are updated in real time. When the training requirements change, the corresponding control panel is updated.

2. The method for data center resource retrieval and modeling training according to claim 1, characterized in that, The process involves deploying a distribution assistant in the data center to acquire training data. The distribution assistant then segments the training data according to a preset segmentation rule, resulting in multiple data blocks, including: Deploy the distribution assistant on the resource management node of the data center; The data center obtains training data through data interfaces or data upload channels; The preset segmentation rules allow the distribution assistant to segment the training data into multiple data blocks. The segmentation rules can be based on one of the following: data size, data type, data access frequency, or semantic logic.

3. The method for data center resource retrieval and modeling training according to claim 1, characterized in that, The establishment of the data channel includes: The distribution assistant uses preset splitting rules to split training requirements into multiple sub-requirements; the splitting rules include splitting training requirements according to training stage or task priority. The control panel connects one-to-one with the connection port of the corresponding carrier space through the sub-port and establishes a data channel. The distribution assistant copies the identifier of the carrier space to the control panel. The control panel verifies the consistency between the copied identifier and the identifier of the carrier space. When they match, the verification is completed and a data channel is established between them.

4. The method for data center resource retrieval and modeling training according to claim 1, characterized in that, The method involves training the model using multiple control panels; specifically, setting multiple time windows, arranging the control panels within these time windows according to the priority of the sub-requirements, and training the multiple sub-requirements according to the arrangement order of the time windows, including: Obtain the priority order of multiple sub-requirements, and arrange the control panels corresponding to the multiple sub-requirements in sequence based on the priority order to obtain the arrangement chain; Multiple time windows are set according to the training requirements of the corresponding model, and the multiple time windows are arranged in sequence to obtain the implementation chain. The multiple time windows are numbered according to the arrangement order, and each time window is bound one-to-one with the corresponding number to obtain the performance parameters of the training space. The maximum carrying capacity of the time window is set based on the performance parameters. Configure point layouts for multiple carrying spaces on a one-to-one basis, and assign multiple control panels to multiple time windows; The time windows of the multiple control panels that have been allocated are used to train the corresponding sub-requirements according to the implementation chain.

5. The method for data center resource retrieval and modeling training according to claim 4, characterized in that, The method of configuring point positions on multiple carrying spaces one-to-one and allocating multiple control panels to multiple time windows includes: The multiple connection ports that connect the carrying space to the sub-ports are configured with points one-to-one. The multiple points are then connected to obtain the point layout corresponding to the carrying space. The connection ports are associated with the corresponding configured points, and the control panel is associated with the points and point layouts connected to the corresponding connection ports. Multiple control panels are allocated to corresponding time windows according to the arrangement chain and maximum carrying capacity. During the allocation process, multiple control panels carry associated point panels to the corresponding time windows, and multiple point panels are placed in the matching area within the time window according to the arrangement chain. The matching area is used to store multiple point panels.

6. The method for data center resource retrieval and modeling training according to claim 5, characterized in that, The step of training the corresponding sub-requirements according to the implementation chain using time windows allocated to multiple control panels includes: When multiple control panels are trained on sub-requirements within the same time window, the multiple control panels copy their associated points and return the associated points in their original positions. The control panel carries the copied points and the data channel points for consistency verification. When they match, the verification is completed, the control panel activates the data channel, and the control panel then trains the sub-requirements through the corresponding carrying space.

7. The method for data center resource retrieval and modeling training according to claim 1, characterized in that, The real-time updating of the model's training requirements, when the training requirements change, updates the control panel corresponding to the training requirements, including: Monitor the model's training requirements in real time and determine whether the training requirements have changed; If so, mark the changes in training requirements, obtain the sub-requirements corresponding to the changes, and update the control panel corresponding to the sub-requirements based on the changes; wherein, updating the control panel includes adding or deleting sub-ports connected to the carrying space. If not, there is no need to update the control panel corresponding to the training requirements.

8. A data center resource retrieval and modeling training system, used to implement the data center resource retrieval and modeling training method according to any one of claims 1-7, characterized in that, include: The first distribution module is used to deploy a distribution assistant in the data center to obtain training data. The distribution assistant divides the training data into multiple data blocks according to a preset segmentation rule. The distribution assistant distributes a corresponding carrying space to each data block, and the carrying space stores the data block. A unique identifier is generated for each data block, and the identifier of the data block is bound to the carrying space. Multiple connection ports are configured for the carrying space. The second distribution module is used to receive the training requirements of the model from the distribution assistant and distribute a corresponding training space according to the training requirements of the model. The distribution assistant splits the training requirements into multiple sub-requirements based on preset splitting rules. The distribution assistant distributes a control panel for each sub-requirement, and the control panel is located in the training space. Multiple carrier spaces corresponding to the sub-requirements are determined, and the same number of sub-ports are configured on the control panel corresponding to the sub-requirements based on the corresponding multiple carrier spaces. The control panel is connected one-to-one with the multiple carrier spaces corresponding to the sub-requirements through the same number of sub-ports to obtain multiple data channels. Among them, the distribution assistant copies the identifier of the carrier space to the control panel, and the control panel verifies the consistency between the copied identifier and the identifier of the carrier space. When the two are consistent, the verification is completed, and a data channel is established between the two. The training module is used to train the model using multiple control panels; it sets multiple time windows, arranges the multiple control panels in the multiple time windows according to the priority of the sub-requirements, and trains the multiple sub-requirements according to the arrangement order of the multiple time windows. The update module is used to update the model's training requirements in real time. When the training requirements change, the corresponding control panel is updated.