Operation request processing system and method for computing power resources
By using write-before logging technology and resource locking mechanism in the computing power resource operation of the AI Intelligent Computing Center, the problem of idempotence of resource operations under concurrent operations is solved, the consistency and reliability of resource states are achieved, duplicate operations and resource waste are avoided, and the risk of business logic errors is reduced.
Patent Information
- Application Number
- CN202510447523.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-25
AI Technical Summary
In the computing power resource operation of AI Intelligent Computing Center, there are problems with idempotence of resource operations in concurrent operation scenarios, resulting in repeated operations, inconsistent data, waste of resources and inconsistent status, which may cause business logic errors.
Pre-operation records are generated using write-pre-logging technology, combining resource locking and state synchronization mechanisms, by locking resources, recording operation logs and restoring the initial state, avoiding repeated operations and state inconsistencies, and filtering repeated requests with unique request identifiers and timestamps to ensure the consistency of resource status.
It significantly reduces the operational idempotence problem in high-concurrency operation scenarios, avoids resource waste and state abnormalities, ensures the consistency of resource state, and prevents business logic errors.
Smart Images

Figure CN120371515A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a system and method for processing operation requests for computing power resources. Background Art
[0002] The computing power resources of an AI intelligent computing center refer to four categories of resources that the AI intelligent computing center can provide externally, namely CPU computing resources, GPU computing resources, storage resources, and network resources. These resources are usually managed by a Kubernetes resource scheduling center, and users can quickly and at low cost perform AI-related machine learning, large model training, and inference on the computing power resources of the AI intelligent computing center.
[0003] The operation of the computing power resources of an AI intelligent computing center refers to, including but not limited to, operations such as creating, deleting, modifying configurations, mounting, unmounting, and resizing resources on the Kubernetes scheduling system.
[0004] Although in most cases, users' operations on resources are serial and there will be no situation where the operations are not idempotent, in the actual production environment, the vast majority of scenarios are concurrent operation scenarios (for example, multiple people operate on the same computing power resources during the same time period), and the problem of operation idempotency is very prominent. Summary of the Invention
[0005] In view of the above problems, the present invention is proposed to provide a system and method for processing operation requests for computing power resources that can overcome or at least partially solve the above problems.
[0006] According to one aspect of the present invention, there is provided a system for processing operation requests for computing power resources, including: a resource locking module, including: a query sub-module, adapted to query whether the requested resource is locked by other operation requests according to the operation request sent by the user terminal; a locking sub-module, adapted to lock the requested resource when the requested resource is not locked by other operation requests; an operation log pre-writing module, adapted to generate a pre-operation record for the requested resource in the pre-operation log, and the pre-operation record at least includes the resource identifier, operation type, and initial resource state of the requested resource, so as to perform corresponding operations on the requested resource according to the resource identifier and operation type; a resource state synchronization module, including: an acquisition sub-module, adapted to acquire the operation state of the requested resource during the operation, and the operation state includes normal operation, operation failure, and operation end; a recovery sub-module, adapted to restore the resource state of the requested resource to the initial resource state according to the pre-operation log when the operation state of the requested resource is operation failure.
[0007] Optionally, in the system according to the present invention, the operation request includes a unique request identifier and a request timestamp. The unique request identifier is generated by concatenating the user terminal identifier, the request operation type, the request operation resource type, and the request operation resource identifier. Further, the system further includes: a resource operation filtering module, adapted to use the unique request identifier and the request timestamp to filter the received operation request according to the filtering rules, and return a request failure to the user terminal for which the operation request does not conform to the filtering rules.
[0008] Optionally, in the system according to the present invention, the filtering rules include: detecting whether the unique request identifier exists in the historical request records. If it exists, it indicates that the operation request does not conform to the request rules; or judging whether the operation request is a repeated request within a predetermined period according to the unique request identifier and the timestamp. If so, it indicates that the operation request does not conform to the request rules.
[0009] Optionally, in the system according to the present invention, the operation type is at least one of resource creation, deletion, configuration modification, mounting, unmounting, and resizing. The system further includes: a resource specification remaining module, adapted to judge whether the specification of the requested resource is sufficient when the operation type is the target type. If it is not sufficient, it returns a request failure.
[0010] Optionally, in the system according to the present invention, the pre-operation record further includes an operation result. The system further includes: a resource status consistent disk write module, adapted to update the operation result in the pre-operation record according to the operation execution result after the operation ends.
[0011] Optionally, in the system according to the present invention, the pre-operation record further includes a request status. The resource status consistent disk write module is further adapted to update the failure reason to the pre-operation log after the operation request fails.
[0012] Optionally, in the system according to the present invention, it further includes: a resource unlocking module, adapted to unlock the requested resource when the operation status of the requested resource is operation end.
[0013] Optionally, in the system according to the present invention, the locking sub-module is adapted to construct locking information with the resource identifier of the requested resource as the key and the resource information as the value, and use the locking information to perform a locking process on the requested resource. The resource information includes the resource type and the resource operation type of the requested resource.
[0014] Optionally, in the system according to the present invention, the query sub-module is further adapted to match the resource identifier of the requested resource with each piece of locking information. If the match is successful, it indicates that the requested resource is locked. If the match fails, it indicates that the requested resource is not locked.
[0015] Optionally, in the system according to the present invention, the acquisition sub-module acquires the operation status of the requested resource during the operation in a polling manner. The polling manner includes that if the current resource status of the requested resource is normal operation, then a status query is performed again after a predetermined time period until the resource status of the requested resource is operation failure or operation end.
[0016] Optionally, in the system according to the present invention, the target type includes the creation of resources and the modification of configurations.
[0017] Optionally, in the system according to the present invention, the acquisition sub-module determines the resource status of the requested resource according to whether the requested resource times out. If, when acquiring the resource status of the requested resource, the requested resource is not completed and does not time out, it indicates that the resource status of the requested resource is normal operation. If, when acquiring the resource status of the requested resource, the requested resource is not completed and times out, it indicates that the resource status of the requested resource is operation failure.
[0018] Optionally, in the system according to the present invention, the computing power resources include CPU computing resources, GPU computing resources, storage resources, and network resources.
[0019] According to another aspect of the present invention, there is provided a method for processing an operation request for computing power resources, which is suitable for being executed by the above-mentioned operation request processing system for computing power resources. The method includes: in response to receiving an operation request sent by a user terminal, determining whether the requested resource is locked by other operation requests; if not locked, performing a locking process on the requested resource; generating a pre-operation record for the requested resource in a pre-operation log, where the pre-operation record at least includes the resource identifier, operation type, and initial resource status of the requested resource, so as to perform corresponding operations on the requested resource according to the resource identifier and operation type; acquiring the operation status of the requested resource during the operation, where the operation status includes normal operation, operation failure, and operation end; when the operation status of the requested resource is operation failure, restoring the resource status of the requested resource to the initial resource status according to the pre-operation log.
[0020] According to another aspect of the present invention, there is provided a computing device, including: at least one processor; and a memory storing program instructions, wherein the program instructions are configured to be executed by the at least one processor, and the program instructions include instructions for executing the above method.
[0021] According to another aspect of the present invention, there is provided a computer program product, including computer programs / instructions, wherein when the computer programs / instructions are executed by a processor, the above method is implemented.
[0022] According to another aspect of the present invention, there is provided a readable storage medium storing program instructions, which when read and executed by a computing device, cause the computing device to execute the above-mentioned method.
[0023] According to the solution of the present invention, by applying the Write Ahead Log (WAL) technology in the field of database technology to the field of computing power resource scheduling and management, and combining the resource locking and unlocking technology, through locking the resources, problems such as repeated resource operations, inconsistent resource status data, and resource waste are avoided, and the idempotency problem of operations in the scenario of high-concurrency resource operations is significantly reduced or avoided. Through the generated pre-operation log, the resources with operation failures can be restored to the state before the operation, avoiding the problem of potential subsequent business logic errors caused by inconsistent states, and greatly reducing the resource status abnormal accidents.
[0024] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the embodiments of the present invention. Brief Description of the Drawings
[0025] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0026] Figure 1 A schematic diagram of scenario 100 showing an operation request for computing power resources;
[0027] Figure 2 A schematic structural diagram of an operation request processing system 200 for computing power resources according to an embodiment of the present invention is shown;
[0028] Figure 3 A block diagram of the physical components (i.e., hardware) of a computing device 300 is shown;
[0029] Figure 4 A flowchart of an operation request processing method 400 for computing power resources according to an embodiment of the present invention is shown;
[0030] Figure 5 A schematic flowchart of an operation request processing system 200 for computing power resources according to an embodiment of the present invention is shown. Detailed Embodiments
[0031] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0032] Some terms related to the present invention are explained accordingly herein.
[0033] The explanations are as follows: AI intelligent computing center: It is an infrastructure dedicated to supporting the research and development and application of AI technology. It integrates a large number of high-performance computing devices, high-speed network connections, and advanced software systems, can efficiently process large-scale AI computing tasks, provide powerful computing power support, and accelerate the training and inference processes of AI models. Through resource pooling, software definition, and efficient management, the intelligent computing center can flexibly schedule computing resources to achieve efficient processing and transfer of data.
[0034] Idempotent operation: Idempotence refers to the effect that the same operation or resource has the same effect in one or multiple requests.
[0035] The AI intelligent computing center can provide multiple types of resources such as CPU computing resources, GPU computing resources, storage resources, and network resources to the outside. Users request to operate these resources through the WEB page or API interface. For example, users request to modify the data in the resources through the WEB page. Figure 1 FIG. 100 is a schematic diagram showing a scenario of an operation request for computing power resources. As Figure 1 shown, in this scenario 100, it includes multiple user terminals, a resource scheduling center, and an AI intelligent computing center. The user terminals are configured with a WEB page or an API interface, and request computing power resources from the AI intelligent computing center through the WEB page or API interface via the resource scheduling center. The AI intelligent computing center is configured with various types of computing resources such as CPU computing resources, GPU computing resources, storage resources, and network resources. These resources are usually scheduled and managed by the resource scheduling center. The resource scheduling center can be, for example, a Kubernetes scheduling system.
[0036] User terminals can perform operations such as creating, deleting, modifying configurations, mounting, unmounting, and resizing resources on the resource scheduling center. For example, a user terminal requests to modify the configuration of a certain resource through the resource scheduling center. In this scenario, a single user is generally sequential during the request operation process, but for the same resource, multiple users can request operations simultaneously. For example, user A requests to modify resource data a to b, and at the same time user B requests to modify resource data a to c. This results in concurrent operations on the same resource, thus generating the problem of operation idempotence.
[0037] Problems with resource operation idempotence include: 1. Repeated operation problems. When a user sends the same request multiple times through a web page or API interface, the server may perform the operation multiple times, resulting in repeated creation or update of resources.
[0038] 2. Data inconsistency: Multiple operations lead to inconsistent data status, which affects the correctness of the system.
[0039] 3. Resource waste: Repeated operations may lead to resource waste, and normal users cannot use the resources.
[0040] 4. Inconsistent status may lead to later business logic errors.
[0041] In order to solve the problems existing in the above-mentioned prior art, the technical solution of the present invention is proposed. Specifically, Figure 2 and 5 As shown, Figure 2 A schematic diagram of the structure of a computing resource operation request processing system 200 according to an embodiment of the present invention is shown. Figure 5 A flow chart of a system 200 for processing operation requests for computing resources according to an embodiment of the present invention is shown.
[0042] System 200 can be deployed in Figure 1 The resource scheduling center includes a resource operation filtering module 210, a resource locking module 220, a resource specification remaining module 230, an operation log pre-writing module 240, a resource status synchronization module 250, a resource unlocking module 260 and a resource status consistent disk module 270 which are coupled to each other.
[0043] The resource operation filtering module 210 is connected to an external user terminal and is used to filter the operation requests sent by the user terminal.
[0044] It is worth noting that when a user terminal initiates an operation request, it will generate a unique request identifier (RequestID) and a 10-digit request timestamp (accurate to seconds). The RequestID consists of the following parts: user terminal identifier, requested operation type, requested operation resource type, and requested operation resource identifier.
[0045] When generating the RequestID, first convert these parts of content into lowercase characters uniformly and remove the leading and trailing spaces. For example, if the user terminal identifier (user id) is 01, the request operation type is deletion, the request operation resource type is CPU computing resource, and the request operation resource identifier (request operation resource id) is 1, then first convert "Chinese characters" to "lowercase characters". For example, convert "user id" to "user_id", "request operation type" to "request_type", "request operation resource type" to "resource_type", and "request operation resource id" to "resource_id".
[0046] Then splice the converted contents in the format of key:value\n, where the combination of the last Key and Value does not add \n.
[0047] The spliced content after the above splicing is:
[0048] user_id: 01\nrequest_type: gpu-pod\nresource_type: delete\nresource_id: 1.
[0049] After that, sort each Key in ascending order of ASCII code. Sorting in ascending order of ASCII code means arranging characters in ascending order of their ASCII code values. In the ASCII code table, the sorting rules for characters are as follows: digits: the ASCII code values of 0-9 are arranged in ascending order, capital letters: the ASCII code values of A-Z are arranged in ascending order, lowercase letters: the ASCII code values of a-z are arranged in ascending order, punctuation marks and other special characters: arranged in their order in the ASCII code table. The specific range of the ASCII code table is digits: 48-57 (0-9), capital letters: 65-90 (A-Z), lowercase letters: 97-222 (a-z), punctuation marks and control characters: 0-31 (control characters), 127 (communication special characters), 32-126 (printable characters).
[0050] Finally, the value obtained by Base64 encoding the generated string is the RequestID value for this time.
[0051] In a specific example, the following are the basic steps of Base64 encoding:
[0052] Divide the data into groups of 3 bytes (24 bits);
[0053] Convert each byte into 8-bit binary form;
[0054] Divide the 24-bit data into groups of 6 bits each, resulting in 4 groups of 6 bits;
[0055] Convert each 6-bit group to its corresponding Base64 character;
[0056] If the data is less than 3 bytes, perform padding.
[0057] Concatenate all the converted Base64 characters to form the final encoded result. The process of decoding a Base64 encoding is the reverse. Convert each Base64 character to its corresponding 6-bit binary value, and then combine these 6-bit values into the original binary data.
[0058] The 10-bit request timestamp is provided as the Unix Timestamp accurate to the second.
[0059] In a specific example, the UNIX time stamp (Unix Timestamp) is the number of seconds (or milliseconds) since January 1, 1970, 00:00:00 (UTC time). It is a marker used to record the time when an event occurred, usually represented as the date and time of a specific moment.
[0060] In a computing device, the UNIX time stamp can be calculated using the following formula:
[0061] [Timestamp = Current time - January 1, 1970 00:00:00 UTC]
[0062] In databases and programming languages, different functions can be used to obtain the current time stamp or convert a date and time to a time stamp:
[0063] MySQL: Use the unix_timestamp() function to obtain the current time stamp, and use the from_unixtime() function to convert the time stamp to the date and time format 34.
[0064] Hive: Use the unix_timestamp() function to obtain the current time stamp, and use the from_unixtime() function to convert the time stamp to the date and time format.
[0065] The resource operation filtering module 210 is adapted to use the above unique request identifier and request timestamp to filter the received operation requests according to the filtering rules, and return a request failure to the user terminal whose operation request does not conform to the filtering rules.
[0066] In some embodiments, the filtering rule includes detecting whether a unique request identifier exists in the historical request record. If it exists, it indicates that the operation request does not conform to the request rule; or judging whether the operation request is a repeated request within a predetermined period according to the unique request identifier and the timestamp. If so, it indicates that the operation request does not conform to the request rule. The predetermined period can be defined according to the actual application scenario. For example, the predetermined period is 5 seconds, and the present application does not make any limitations in this regard. In other words, the resource operation filtering module 210 can detect whether a RequestID already exists in the historical request record. If so, reject the current request; otherwise, release it. Or detect whether a RequestID requests more than once within 5 seconds. If so, reject the current request; otherwise, release it.
[0067] The resource specification remaining module 230 is adapted to judge whether the specification of the requested resource is sufficient according to the operation type in the operation request. If it is not sufficient, it returns a request failure. The operation type is at least one of resource creation, deletion, configuration modification, mounting, unmounting, and resizing. When the resource specification remaining module 230 detects that the operation type is the target type, it judges whether the specification of the requested resource is sufficient. If it is not sufficient, it returns a request failure. The target type can be resource creation, configuration modification, etc. Some operations such as creation and reconfiguration require prior confirmation of whether there are remaining resources available for operation in the specification of the resource to be operated. If the remaining resources are insufficient, it directly returns out of stock. For example, if a user wants to create a CPU computing resource but the CPU computing resources in the AI computing center are insufficient, the request will fail.
[0068] The resource locking module 220 includes a query sub-module 221 and a locking sub-module 222. Among them, the query sub-module 221 is adapted to query whether the requested resource is locked by other operation requests according to the operation request sent by the user terminal. The locking sub-module 222 is adapted to lock the requested resource when the requested resource is not locked by other operation requests.
[0069] When querying whether the requested resource is locked by other operation requests, the query sub-module 221 will query whether there is resource locking information associated with the requested resource. The resource locking information is generally stored in the cache system and is used to characterize whether a certain resource is locked. If the resource is locked, a resource locking information of the resource will be generated on the cache system accordingly. If the query sub-module 221 can find the resource locking information, it means that the resource is locked; otherwise, it is not locked.
[0070] Resource locking information is usually saved in the cache system, which contains the operating user, resource type, and resource operation type. This information is stored in the form of Key:Value. The saved Key information consists of two parts, namely [Key prefix] + [resource id], and the saved Value is a data of HSet type, in a format similar to the following
[0071]
[0072] Among them, "userId": "u-00001" indicates the user id: u-00001, where "userId" is the Key and u-00001 is the Value. Similarly, "resourceType": "gpu-pod" indicates the resource type: GPU computing resource, and "operateType": "delete" indicates the operation resource type: delete.
[0073] In some embodiments, each resource locking information corresponds to a locking identifier, which is the resource identifier of the resource. The query sub-module 221 can query whether there is resource locking information for the requested resource according to this locking identifier. If it exists, it means the requested resource is locked; if not, it means the requested resource is not locked.
[0074] When the requested resource is not locked, the locking sub-module 222 will generate a resource locking information using the resource identifier, resource type, and resource operation type, and perform a locking process on the requested resource. The lock here is an exclusive lock, and all other operation requests related to the resource will be blocked.
[0075] In some embodiments, the operation log write-ahead module 240 can generate a pre-operation record for the requested resource in the pre-operation log according to the operation request, so that the resource scheduling center can perform corresponding operations on the requested resource according to the pre-operation log.
[0076] The pre-operation record records the resource identifier, operation type, initial resource state, operation result, and request state of the requested resource.
[0077] The resource identifier indicates the resource type of the requested resource. For example, resource identifier 01 represents CPU computing resource, 02 represents GPU computing resource, 03 represents storage resource, 04 represents network resource, etc.
[0078] The operation type indicates the operation content of the requested resource, such as creating, deleting, modifying configuration, mounting, unmounting, and resizing the resource, etc.
[0079] The initial resource state indicates the current state of the requested resource. For example, the resource is running, etc.
[0080] The operation result indicates the execution result of the operation on the requested resource, and the request status indicates whether the request is successful (for example, if the resource is locked by another resource, it means the request fails). The operation result and request status in the pre-operation log can be null values at the beginning. After the actual operation on the requested resource is completed, the resource status consistency disk writing module 270 updates the actual operation result or request status to the pre-operation record.
[0081] The pre-operation log is pre-generated in the operation log pre-writing module 240. This pre-operation log refers to the WAL log method. Before performing the operation action, all information related to the operation is written as a pre-operation record into the pre-operation log. In addition to the above-mentioned resource identifier, operation type, initial resource status, operation result, and request status, this information can also include the operating user, operation time, request address, request method, request parameters, request ID, etc.
[0082] It should be noted that the writing process of the pre-operation record is carried out before operating on the requested resource, that is, "record first and then do". This mechanism ensures that even if the operation fails, the pre-operation log can be used to restore the consistency state of the requested resource.
[0083] Exemplarily, when the user terminal requests to modify data a in the requested resource to b, a pre-operation record of this request operation will be generated in the pre-operation log first. When the modification of data a to b fails during the process of the requested resource, the pre-operation log can be retrieved for rollback, so as to restore the initial resource to the state before the modification.
[0084] The resource status synchronization module 250 is used to execute the above process of retrieving the pre-operation log for rollback and restoring the status of the requested resource. The resource status synchronization module 250 includes an acquisition sub-module 251 and a restoration sub-module 252.
[0085] The acquisition sub-module 251 obtains the operation status of the requested resource during the operation in a polling manner. The operation status includes normal operation, operation failure, and operation end. This polling method includes that if the current resource status of the requested resource is normal operation, a status query is performed again after a predetermined time period until the resource status of the requested resource is operation failure or operation end. Exemplarily, if the operation of the requested resource obtained by the acquisition sub-module 251 is in an unfinished and non-timeout state, it means the operation status is normal operation, then wait for 1 second to enter the next query, and repeat this process continuously until the operation status becomes operation failure or operation end. Among them, when the operation of the requested resource is in an unfinished and timeout state, it means the operation fails and the operation status becomes operation failure. When the operation of the requested resource is in a completed state, it means the operation ends.
[0086] When the resource state of the requested resource is an operation failure, the recovery sub-module 252 restores the source state of the requested resource to the initial resource state in a rollback manner according to the above pre-operation log.
[0087] After the operation on the requested resource ends (regardless of whether the operation is successful or failed), the resource state consistent disk write module 270 updates the operation result to the above pre-operation log. If the operation is successful, the operation result is updated to "operation successful" in the pre-operation log. If the operation fails, the operation result is updated to "operation failed" in the pre-operation log.
[0088] It should be noted that although the operation result is an operation failure, the requested resource has been restored to a usable state by the recovery sub-module 252 at this time.
[0089] In addition, when the request fails (for example, when the above resource operation filtering module 210 filters the operation request and finds that the request is repeated, resulting in a request failure, or when the resource specification remaining module 230 determines that the specification of the requested resource is insufficient, resulting in a request failure, etc.), the resource state consistent disk write module 270 can also update the reason for the request failure to the pre-operation log. In other words, when the system 200 receives an operation request sent by the user terminal, the operation log pre-write module 240 records the pre-operation log of this operation request in the pre-operation log.
[0090] In addition, after the operation on the requested resource ends, the resource unlocking module 260 unlocks the requested resource to facilitate subsequent request operations.
[0091] By applying the Write Ahead Log (WAL) technology in the database technology field to the computing power resource scheduling and management field and combining the resource locking and unlocking technology, the system 200 of the present invention avoids problems such as repeated operation of resources, inconsistent resource state data, and resource waste by locking the resources, significantly reduces or avoids the operation idempotency problem in the resource high-concurrency operation scenario, and can restore the failed resources to the state before the operation through the generated pre-operation log, avoiding the problem of possible subsequent business logic errors due to inconsistent states, and greatly reducing resource state abnormal accidents.
[0092] In some embodiments, the resource scheduling center can be implemented by one or more computing devices. Figure 3A block diagram of the physical components (i.e., hardware) of computing device 300 is shown. In a basic configuration, computing device 300 includes at least one processing unit 302 and system memory 304. According to one aspect, depending on the configuration and type of the computing device, system memory 304 includes, but is not limited to, volatile storage (e.g., random access memory), non-volatile storage (e.g., read-only memory), flash memory, or any combination of such memories.
[0093] According to one aspect, system memory 304 includes operating system 305. System memory 304 also includes program modules 350. According to one aspect, operating system 305, for example, is suitable for controlling the operation of computing device 300. Additionally, the examples are practiced in conjunction with a graphics library, other operating systems, or any other application programs, and are not limited to any particular application or system. In Figure 3 this basic configuration is shown by those components within dashed line 308. According to one aspect, computing device 300 has additional features or functions. For example, according to one aspect, computing device 300 includes additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or magnetic tapes. Such additional storage is Figure 3 shown in by removable storage device 309 and non-removable storage device 310.
[0094] As stated above, according to one aspect, program modules 350 are stored in system memory 304. According to one aspect, program modules 350 may be implemented as one or more computer program products, and the type of computer program products is not limited in this application. For example, it may include: email, word processing applications, spreadsheet applications, database applications, slide show applications, painting or computer-aided applications, web browsers, etc. In an embodiment according to the present invention, the application program in program modules 350 may be the operation request processing system 200 for computing power resources, and the operation request processing system 200 for computing power resources is configured to execute the operation request processing method 400 for computing power resources of the present invention.
[0095] According to one aspect, the examples may be practiced in a circuit including discrete electronic components, a packaged or integrated electronic chip containing logic gates, a circuit utilizing a microprocessor, or on a single chip containing electronic components or a microprocessor. For example, it may be via where in Figure 3Each or many of the components shown in [the figure] can be practiced as examples by a system-on-a-chip (SOC) integrated on a single integrated circuit. According to one aspect, such an SOC device may include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all of which are integrated (or "burned") onto a chip substrate as a single integrated circuit. When operating via the SOC, the functions described herein can be operated via dedicated logic integrated with other components of the computing device 300 on a single integrated circuit (chip). Embodiments of the present invention can also be practiced using other technologies capable of performing logical operations (such as AND, OR, and NOT), including but not limited to mechanical, optical, fluidic, and quantum technologies. Additionally, embodiments of the present invention can be practiced within a general-purpose computer or in any other circuit or system.
[0096] According to one aspect, the computing device 300 may also have one or more input devices 312, such as a keyboard, mouse, pen, voice input device, touch input device, etc. An output device 314, such as a display, speaker, printer, etc., may also be included. The foregoing devices are examples and other devices may also be used. The computing device 300 may include one or more communication connections 316 that allow communication with other computing devices 318, and the other computing devices 318 may be printing devices, such as printers. Examples of suitable communication connections 316 include but are not limited to: RF transmitter, receiver, and / or transceiver circuits; Universal Serial Bus (USB), parallel, and / or serial ports.
[0097] As used herein, the term computer-readable medium includes computer storage media. Computer storage media can include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, or program modules). System memory 304, removable storage device 309, and non-removable storage device 310 are all examples of computer storage media (i.e., memory storage). Computer storage media can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other article that can be used to store information and can be accessed by the computer device 300. According to one aspect, any such computer storage media can be part of the computing device 300. Computer storage media does not include carrier waves or other propagated data signals.
[0098] According to one aspect, a communication medium is implemented by computer-readable instructions, data structures, program modules, or other data in a modulated data signal (e.g., a carrier wave or other transmission mechanism), and includes any information delivery medium. According to one aspect, the term "modulated data signal" describes a signal having one or more sets of characteristics or a signal that has been altered in a manner that encodes information in the signal. By way of example and not limitation, communication media include wired media such as a wired network or a direct wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0099] In an embodiment according to the present invention, the computing device 300 is configured to execute an operation request processing method 400 for applying computing power resources according to the present invention. The computing device 300 includes one or more processors and one or more readable storage media storing program instructions. When the program instructions are configured to be executed by the one or more processors, the computing device is caused to execute the application remote control method 300 in the embodiments of the present invention.
[0100] According to an embodiment of the present invention, an operation request processing device 500 for the computing power resources of the computing device 300 is configured to execute an operation request processing method 400 for the computing power resources according to the present invention. Among them, the operation request processing system 200 for the computing power resources contains multiple program instructions for executing the operation request processing method 400 for the computing power resources according to the present invention. These program instructions can instruct the processor to execute the operation request processing method 400 for the computing power resources according to the present invention. By executing the operation request processing method 400 for the computing power resources according to the present invention, the computing device 200 significantly reduces or avoids the problem of operation idempotency in the scenario of high-concurrency resource operations.
[0101] Figure 4 A flowchart of an operation request processing method 400 for computing power resources according to an embodiment of the present invention is shown. The method 400 is suitable for execution in a computing device (such as the aforementioned computing device 300).
[0102] As Figure 4 shown, the purpose of the method 400 is to implement a method for processing operation requests for computing power resources that, by locking resources, avoids problems such as resources being repeatedly operated, resource status data being inconsistent, and resource waste, and significantly reduces or avoids the problem of operation idempotency in the scenario of high-concurrency resource operations. Additionally, by recording the resource information of the resource before the operation, the resource with an operation failure can be restored to the state before the operation, avoiding problems that may cause subsequent business logic errors due to inconsistent states, and greatly reducing the occurrence of accidents with abnormal resource states.
[0103] Method 400 begins at step 402, in which it is determined whether the requested resource is locked by other operation requests in response to receiving an operation request sent by a user terminal.
[0104] In step 404, if the requested resource is not locked by other operation requests, a lock is applied to the requested resource.
[0105] In step 406, a pre-operation record for the requested resource is generated in the pre-operation log. The pre-operation record includes at least the resource identifier, operation type, and initial resource state of the requested resource, so that corresponding operations can be performed on the requested resource according to the resource identifier and operation type.
[0106] In step 408, the operation state of the requested resource during the operation is obtained. The operation state includes normal operation, operation failure, and operation end.
[0107] In step 410, when the operation state of the requested resource is operation failure, the resource state to which the requested resource is restored is restored to the initial resource state according to the pre-operation log.
[0108] It should be noted that the working principle and process of method 400 provided in this embodiment are similar to those of the above system 200. For the related parts, reference can be made to the description of the above system 200, which will not be elaborated here.
[0109] The various technologies described here can be implemented in combination with hardware or software, or a combination of them. Thus, the method and device of the present invention, or certain aspects or parts of the method and device of the present invention, can take the form of program code (i.e., instructions) embedded in a tangible medium, such as a removable hard disk, USB flash drive, floppy disk, CD-ROM, or any other machine-readable storage medium. When the program is loaded into a machine such as a computer and executed by the machine, the machine becomes a device for practicing the present invention.
[0110] A10. The system as described in A1, wherein the acquisition sub-module acquires the operation status of the requested resource during the operation in a polling manner, and the polling manner includes that if the current resource status of the requested resource is normal operation, a status query is performed again after a predetermined time period until the resource status of the requested resource is operation failure or operation end. A11. The system as described in A4, wherein the target type includes the creation of resources and the modification of configurations. A13. The system as described in A10, wherein the acquisition sub-module determines the resource status of the requested resource according to whether the requested resource times out. If the requested resource is not completed and does not time out when acquiring the resource status of the requested resource, it indicates that the resource status of the requested resource is normal operation. If the requested resource is not completed and times out when acquiring the resource status of the requested resource, it indicates that the resource status of the requested resource is operation failure. A14. The system as described in any one of A1 - A10, wherein the computing power resources include CPU computing resources, GPU computing resources, storage resources, and network resources.
[0111] In the case where the program code is executed on a programmable computer, the computing device generally includes a processor, a processor-readable storage medium (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device. Among them, the memory is configured to store the program code; the processor is configured to execute the method of the present invention according to the instructions in the program code stored in the memory.
[0112] By way of example and not limitation, the readable medium includes a readable storage medium and a communication medium. The readable storage medium stores information such as computer-readable instructions, data structures, program modules, or other data. The communication medium generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information delivery medium. A combination of any of the above is also included within the scope of the readable medium.
[0113] In the specification provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the examples of the present invention. Based on the above description, the structure required to construct such a system is obvious. In addition, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the description of the specific language above is to disclose the preferred embodiments of the present invention.
[0114] In the specification provided herein, a number of specific details are set forth. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail so as not to obscure an understanding of the present specification.
[0115] Those skilled in the art should understand that the modules or units or components of the devices in the examples disclosed herein may be arranged in the devices as described in the embodiments, or alternatively may be located in one or more devices different from the devices in the examples. The modules in the foregoing examples may be combined into one module or further divided into multiple sub-modules.
[0116] In addition, some of the embodiments herein are described as a combination of methods or method elements that may be implemented by a processor of a computer system or by other devices performing the functions. Therefore, a processor having the necessary instructions for implementing the methods or method elements forms a device for implementing the methods or method elements. In addition, the elements described herein in the device embodiments are examples of such devices: the devices are for implementing the functions performed by the elements for the purpose of implementing the invention.
[0117] As used herein, unless otherwise specified, the use of ordinal numbers "first", "second", "third", etc. to describe ordinary objects merely indicates different instances of similar objects and is not intended to imply that the objects so described must have a given order in time, space, ranking or in any other manner.
[0118] Although the invention has been described in terms of a limited number of embodiments, those skilled in the art within the present technical field will appreciate that other embodiments may be contemplated within the scope of the invention as thus described. In addition, it should be noted that the language used in this specification has been principally selected for readability and instructional purposes and not for the purpose of explaining or limiting the subject matter of the invention. Accordingly, many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the appended claims.
Claims
1. An operation request processing system for computing power resources, comprising: A resource locking module, comprising: A query sub-module, adapted to query whether the requested resource is locked by other operation requests according to the operation request sent by the user terminal; A locking sub-module, adapted to lock the requested resource when the requested resource is not locked by other operation requests; An operation log pre-write module, adapted to generate a pre-operation record for the requested resource in the pre-operation log, where the pre-operation record at least includes the resource identifier, operation type, and initial resource state of the requested resource, so as to perform corresponding operations on the requested resource according to the resource identifier and operation type; A resource status synchronization module, comprising: An acquisition sub-module, adapted to acquire the operation status of the requested resource during the operation, where the operation status includes normal operation, operation failure, and operation end; A recovery sub-module, adapted to restore the resource status of the requested resource to the initial resource state according to the pre-operation log when the operation status of the requested resource is operation failure.
2. The system according to claim 1, wherein, The operation request includes a unique request identifier and a request timestamp, the unique request identifier is generated by splicing the user terminal identifier, request operation type, request operation resource type, and request operation resource identifier, and the system further includes: A resource operation filtering module, adapted to filter the received operation requests according to the filtering rules using the unique request identifier and request timestamp, and return a request failure to the user terminal whose operation request does not conform to the filtering rules.
3. The system according to claim 2, wherein, The filtering rules include: Detect whether the unique request identifier exists in the historical request record, if it exists, it indicates that the operation request does not conform to the request rules; or According to the unique request identifier and timestamp, determine whether the operation request is a repeated request within a predetermined period, if so, it indicates that the operation request does not conform to the request rules.
4. The system according to claim 1, wherein The operation type is at least one of resource creation, deletion, configuration modification, mounting, unmounting, and resizing, and the system further includes: A resource specification remaining module, adapted to determine whether the specification of the requested resource is sufficient when the operation type is the target type, if not, return a request failure.
5. The system according to any one of claims 1-4, wherein, The pre-operation record further includes an operation result, and the system further includes: A resource status consistent disk write module, adapted to update the operation result in the pre-operation record according to the operation execution result after the operation ends.
6. The system according to claim 5, wherein The pre-operation record further includes a request status, and the resource status consistent disk write module is further adapted to update the failure reason to the pre-operation log after the operation request fails.
7. The system according to claim 1, wherein It further includes: A resource unlocking module, adapted to unlock the requested resource when the operation status of the requested resource is operation end.
8. The system according to claim 1, wherein The locking sub-module is adapted to construct locking information with the resource identifier of the requested resource as the key and the resource information as the value, and perform a locking process on the requested resource using the locking information, where the resource information includes the resource type and resource operation type of the requested resource.
9. The system according to claim 8, wherein, The query sub-module is also adapted to match the resource identifier of the requested resource with each of the lock information. If the match is successful, it indicates that the requested resource is locked. If the match fails, it indicates that the requested resource is not locked.
10. A method for processing an operation request for computing power resources, which is adapted to be executed by the system according to any one of claims 1-9. The method includes: In response to receiving an operation request sent by a user terminal, determining whether the requested resource is locked by other operation requests; If it is not locked, perform a locking process on the requested resource; Generate a pre-operation record for the requested resource in a pre-operation log, where the pre-operation record at least includes the resource identifier, operation type, and initial resource state of the requested resource, so as to perform corresponding operations on the requested resource according to the resource identifier and operation type; Obtain the operation state of the requested resource during the operation, where the operation state includes normal operation, operation failure, and operation end; When the operation state of the requested resource is operation failure, restore the resource state of the requested resource to the initial resource state according to the pre-operation log.