Data operation processing method and device, equipment and medium
By analyzing and executing user-submitted data operation requests in the data development environment, ensuring accurate analysis and direction of operating environment identifiers, the problems of high cost of traditional physical isolation solutions and difficulty in data sharing are solved, and efficient, accurate and secure data operation processing are achieved.
Patent Information
- Application Number
- CN202510117344.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-06
AI Technical Summary
In a data development environment, although traditional physical isolation schemes can ensure that data operations in the test environment do not affect the development environment, they are costly, difficult to share data, and are inconvenient for testing, and lack an efficient, flexible and easy-to-implement data operation processing method.
By obtaining the operation environment identifier and data operation statement in the data operation request submitted by the user, in response to the request scheduling event, the request is submitted to the data operation parser for analysis. After parsing, the data operation statement is submitted to the data operation cluster for execution. The cluster performs the corresponding data operation in the target environment database and returns the execution result to the user.
It realizes operation isolation between databases in different environments, ensures data security and integrity, avoids operation errors and data chaos caused by environmental confusion, improves the reliability and stability of data operations, and provides an efficient, accurate and secure data operation processing solution.
Smart Images

Figure CN119938646A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of e-commerce technology, and in particular to a data operation and processing method and its corresponding device, computer equipment, and computer-readable storage medium. Background Art
[0002] In a data development environment, the isolation between different environments for data operations is a crucial issue, in order to ensure that when data operations are performed in a test environment, the data in the development environment cannot be affected. In traditional technology, in a data development environment, the isolation between different environments for data operations is a crucial issue, in order to ensure that when data operations are performed in a test environment, the data in the development environment cannot be affected. In traditional technology, a physical isolation solution is adopted to achieve isolation by storing data in physically different devices or storage areas. This method usually involves two or more independent network environments, and there is no electronic communication path between these environments. Although physical isolation provides the highest level of security and is often used to handle extremely sensitive data, such as government, military or financial data, its disadvantages are high cost, difficulty in data sharing and extreme inconvenience in testing.
[0003] Therefore, exploring a more efficient, flexible and easy-to-implement data operation and processing method is of great significance to improving the security and management efficiency of the data development environment. Summary of the invention
[0004] The primary purpose of the present application is to solve at least one of the above problems and to provide a data operation and processing method and its corresponding device, computer equipment, and computer program product.
[0005] In order to meet the various objectives of this application, this application adopts the following technical solutions:
[0006] A data operation processing method provided to meet one of the purposes of this application includes the following steps:
[0007] Obtaining a data operation request submitted by a user, responding to a request scheduling event, and submitting the data operation request to a data operation parser, wherein the data operation request includes an operation environment identifier and a data operation statement;
[0008] The data operation parser performs corresponding parsing on the data operation statement according to the operation environment identifier to obtain the parsed data operation statement, so that its operation database object is designated as the target environment database pointed to by the operation environment identifier;
[0009] The parsed data operation statement is submitted to the data operation cluster for execution, and the cluster executes the corresponding data operation in the target environment database, and the execution result returned by the cluster is pushed to the user.
[0010] On the other hand, a data operation processing device provided to meet one of the purposes of the present application includes a request processing module, a statement parsing module and a statement execution module, wherein the request processing module is used to obtain a data operation request submitted by a user, respond to a request scheduling event, and submit the data operation request to a data operation parser, wherein the data operation request includes an operation environment identifier and a data operation statement; the statement parsing module is used for the data operation parser to perform corresponding parsing on the data operation statement according to the operation environment identifier, obtain the parsed data operation statement, and make its operation database object be designated as the target environment database pointed to by the operation environment identifier; the statement execution module is used to submit the parsed data operation statement to a data operation cluster for execution, and the cluster executes the corresponding data operation in the target environment database, obtains the execution result returned by the cluster and pushes it to the user.
[0011] On the other hand, a computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory, wherein the central processing unit is used to call and run a computer program stored in the memory to execute the steps of the data operation processing method described in the present application.
[0012] On the other hand, a computer program product provided to meet another purpose of the present application includes a computer program / instruction, which, when executed by a processor, implements the steps of the method described in any embodiment of the present application.
[0013] The technical solution of this application has many advantages, including but not limited to the following aspects:
[0014] This application first obtains the operating environment identifier and specific data operation statement in the data operation request submitted by the user. In response to the request scheduling event, the request is submitted to the data operation parser to parse the data operation statement according to the operating environment identifier, so that the parsed statement can clearly point to a specific target environment database, laying the foundation for subsequent precise operations. Then, the parsed data operation statement is submitted to the data operation cluster for execution, and the cluster performs the corresponding data operation in the target environment database, and returns the execution result to the user. The whole process is closely linked, which can not only efficiently perform the data operations required by the user, but also ensure the accuracy of the operation. The most important thing is that through the precise parsing and pointing of the operating environment identifier, the data operation can be accurately arranged to the corresponding target environment database for execution, realizing the isolation of operations between different environment databases without interference. This not only ensures the security and integrity of the data, but also avoids the operational errors and data confusion that may be caused by environmental confusion, greatly improves the reliability and stability of data operations, and provides users with an efficient, accurate and safe data operation processing solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0016] Figure 1 The network architecture of the e-commerce platform exemplified in this application;
[0017] Figure 2 A flowchart of a typical embodiment of the data operation processing method of the present application;
[0018] Figure 3 A principle block diagram of the data operation processing device of the present application;
[0019] Figure 4 A schematic diagram of the structure of a computer device used in this application. DETAILED DESCRIPTION
[0020] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present application.
[0021] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0022] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with those in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.
[0023] like Figure 1 In the network architecture shown, the e-commerce platform 82 is deployed on the Internet to provide corresponding services to its users. Similarly, the devices 80 of the merchant users of the e-commerce platform 82 and the devices 81 of the consumer users are also connected to the Internet to use the services provided by the e-commerce platform.
[0024] The exemplary e-commerce platform 82 provides supply and demand matching of products and / or services to the general public with the help of Internet infrastructure. In the e-commerce platform 82, products and / or services are provided as commodity information. To simplify the description, the concepts of commodity, product, etc. are used in this application to refer to the products and / or services in the e-commerce platform 82, which may specifically be physical products, digital products, tickets, service subscriptions, other offline services, etc.
[0025] In reality, all entities can access the e-commerce platform 82 as users, use various online services provided by the e-commerce platform 82, and achieve the purpose of participating in the business activities achieved by the e-commerce platform 82. These entities can be natural persons, legal persons or social organizations, etc. Corresponding to the two types of entities, merchants and consumers in business activities, the e-commerce platform 82 has two types of users, merchant users and consumer users. All entities in the product distribution chain in business activities, including manufacturers, sellers, retailers, logistics providers, etc., can use online services in the e-commerce platform 82 as merchant users, while consumers in business activities, including real or potential consumers, can use online services in the e-commerce platform 82 as their corresponding consumer users. In actual business activities, the same entity can act as both a merchant user and a consumer user, and this should be understood flexibly.
[0026] The infrastructure used to deploy the e-commerce platform 82 mainly includes the backend architecture and frontend equipment. The backend architecture runs various online services through the service cluster, including middleware or frontend services for the platform, services for consumers, services for merchants, etc., to enrich and improve its service functions; the frontend equipment mainly covers the terminal devices used by users to access the e-commerce platform 82 as clients, including but not limited to various mobile terminals, personal computers, point-of-sale devices, etc. For example, merchant users can enter product information for their online stores through their terminal devices 80, or generate their product information using the interfaces opened by the e-commerce platform; consumer users can access the webpage of the online store implemented by the e-commerce platform 82 through their terminal devices 81, trigger the shopping process through the shopping buttons provided on the webpage, and call various online services provided by the e-commerce platform 82 during the shopping process, so as to achieve the purpose of shopping orders.
[0027] In some embodiments, the e-commerce platform 82 may be implemented by a processing facility including a processor and a memory, the processing facility storing a set of instructions that, when executed, cause the e-commerce platform 82 to perform the e-commerce and support functions involved in the present application. The processing facility may be part of a server, client, network infrastructure, mobile computing platform, cloud computing platform, fixed computing platform, or other computing platform, and provides electronic components, merchant equipment, payment gateways, application developers, marketing channels, transportation providers, customer equipment, point-of-sale equipment, etc. of the e-commerce platform 82.
[0028] The e-commerce platform 82 can be implemented as online services such as cloud computing services, software as a service (SaaS), infrastructure as a service (IaaS), platform as a service (PaaS), desktop as a service (DaaS), hosted software as a service, mobile backend as a service (MBaaS), information technology management as a service (ITMaaS), etc. In some embodiments, the various functional components of the e-commerce platform 82 can be implemented to be suitable for operation on various platforms and operating systems. For example, for an online store, its administrator users enjoy the same or similar functions regardless of various embodiments such as iOS, Android, HomonyOS, or web pages.
[0029] The e-commerce platform 82 can realize the corresponding independent station for each merchant to run its corresponding online store, and provide the merchant with the corresponding business management engine instance for the merchant to establish, maintain and run one or more online stores in one or more independent stations. The business management engine instance can be used for content management, task automation and data management of one or more online stores, and various specific business processes of the online store can be configured through interfaces or built-in components to support the implementation of business activities. The independent station is the infrastructure of the e-commerce platform 82 with cross-border service functions. Merchants can maintain their online stores more centrally and autonomously based on the independent station. The independent station usually has a domain name and storage space dedicated to the merchant, and different independent stations are relatively independent. The e-commerce platform 82 can provide standardized or personalized technical support for a large number of independent stations, so that merchant users can customize their own business management engine instance and use this business management engine instance to maintain one or more online stores they own.
[0030] The online store can implement backend configuration and maintenance by having the merchant user log in to its business management engine instance as an administrator. With the support of various online services provided by the infrastructure of the e-commerce platform 82, the merchant user can configure various functions in its online store as an administrator, view various data, etc. For example, the merchant user can manage various aspects of its online store, such as viewing the online store's recent activities, updating the online store's product catalog, managing orders, recent visit activities, total order activities, etc.; the merchant user can also view more detailed information about the business and visitors to the merchant's online store by obtaining reports or metrics, such as showing a sales summary of the merchant's overall business, specific sales and participation data of active sales marketing channels, etc.
[0031] The e-commerce platform 82 may provide communication facilities and associated merchant interfaces for providing electronic communications and marketing, such as utilizing electronic message aggregation facilities to collect and analyze communication interactions between merchants, consumers, merchant devices, customer devices, point-of-sale devices, etc., aggregating and analyzing communications, such as for increasing the potential for providing product sales, etc. For example, a consumer may have a question related to a product, which may generate a dialogue between the consumer and the merchant (or an automated processor-based agent on behalf of the merchant), where the communication facility is responsible for interacting and providing the merchant with analysis on how to increase the probability of a sale.
[0032] In some embodiments, an application suitable for installation in a terminal device can be provided to serve the access needs of different users, so that various users can access the e-commerce platform 82 by running the application in the terminal device, such as the merchant backend module of the online store in the e-commerce platform 82, etc. In the process of implementing business activities through these functions, the e-commerce platform 82 can implement various functions related to supporting the implementation of business activities as middleware or online services and open corresponding interfaces, and then implant the toolkit corresponding to the interface access function into the application to implement function expansion and task implementation. The business management engine can include a series of basic functions, and expose these functions to online services and / or application calls through APIs. Online services and applications use corresponding functions by remotely calling corresponding APIs.
[0033] With the support of various components of the business management engine instance, the e-commerce platform 82 can provide online shopping functions, enabling merchants to establish connections with customers in a flexible and transparent manner, and consumer users can select items online, create product orders, provide the delivery address of the goods in the product order, and complete the payment confirmation of the product order. Then, the merchant can review and complete or cancel the order.
[0034] A data operation and processing method of the present application can be programmed as a computer program product and deployed in a client or server for execution. For example, in the exemplary application scenario of the present application, it can be deployed and implemented in the server of an e-commerce customer service platform, thereby accessing an interface opened after the computer program product is running and performing human-computer interaction with the process of the computer program product through a graphical user interface to execute the method.
[0035] See also Figure 2 The data operation processing method of the present application, in its typical embodiment, comprises the following steps:
[0036] Step S1100: obtaining a data operation request submitted by a user, responding to a request scheduling event, and submitting the data operation request to a data operation parser, wherein the data operation request includes an operation environment identifier and a data operation statement;
[0037] A data development system can be built in advance and provided to users with developer permissions to use the system to conduct data mining or data analysis on business data, product data, customer data, etc. in e-commerce platforms and / or independent site stores, and implement corresponding data operations.
[0038] The user can log in to the data development system and submit the corresponding identity information to the system for authentication. When the system confirms that the user has developer privileges, the user can enter the development mode by touching the "Development" button in the system. In this mode, in the graphical interface of the control that includes editing data operation statements, the user can edit the corresponding data operation statements according to the data operations that need to be implemented and their operating environment, and select the environment for executing the data operation statements as the test environment and / or the production environment, and then determine the operating environment identifier that represents the selected environment, encapsulate the data operation request corresponding to the data operation statement and the operating environment identifier, and submit it to the scheduling system.
[0039] The SQL statement (Structured Query Language) is a standard programming language for managing and operating relational databases, including operations such as querying, modifying, inserting and deleting data in the database. The HQL (Hive Query Language) statement is a statement edited using the query language of the Hive database. It is similar to SQL, but is specifically designed for large-scale data processing in the Hadoop ecosystem. HQL statements can be used to query, delete and modify data stored in Hive. For ease of understanding of the exemplary examples, the operating environment identifier can be 0, 1, and 2 to refer to the production environment, test environment, production environment, and test environment, respectively. Of course, those skilled in the art can flexibly set them separately.
[0040] In one embodiment, due to actual business needs, some data operations may have higher priorities than other data operations, such as urgent market analysis reports or key business decision support reports, etc. For this reason, in the data development system, the user can also edit the data operation statement and select the environment for executing the data operation statement, and edit the purpose description text corresponding to what business needs the data operation statement is used to meet, and then encapsulate the purpose description text together with the data operation statement and the operating environment identifier corresponding to the selected environment into the data operation request. After receiving each data operation request, the scheduling system uses the language model to identify the purpose description text to confirm the scheduling priority of the request, and then responds to the request scheduling event in the order of scheduling priority from high to low and submits the corresponding data operation request to the data operation parser, wherein the data operation request with the same scheduling priority is further submitted to the data operation parser in the order from first to last in which the scheduling system receives these data operation requests. The language model can be a classification mapping relationship between the semantics of the pre-modeled purpose description text and the identifier representing the scheduling priority of each level, and the corresponding model pre-trained to a convergence state. Thus, the corresponding scheduling priority can be obtained by inputting the purpose description text into the model. The structure of the model includes a text feature representation layer and a classification layer. The text feature representation layer is used to extract and vectorize the semantics of the purpose description text. The recommended selection of this layer can be a BERT model, and any other model such as Transfomer Encoder, RoBERTa, XLM-RoBERTa, MPNet, BiLSTM, etc. can also be used; the classification layer can be an MLP or a fully connected layer. The language model can also be a large language model. Before using the model, the purpose description text needs to be embedded in a preset prompt template, and the corresponding question text is obtained and then input into the large language model to obtain the corresponding scheduling priority. The prompt template includes a task prompt and the purpose description text to be embedded. The task prompt is used to guide the large language model to refer to multiple examples given to determine the scheduling priority corresponding to the embedded purpose description text. Each example demonstrates different purpose description texts and their corresponding scheduling priorities. The specific prompt template can be flexibly implemented by those skilled in the art based on the disclosure herein.
[0041] Step S1200: the data operation parser parses the data operation statement according to the operation environment identifier to obtain the parsed data operation statement, so that its operation database object is designated as the target environment database pointed to by the operation environment identifier;
[0042] The data operation parser includes a pre-processing component and a statement parsing component. The pre-processing component can be a functional component that can be directly called after being pre-packaged such as code or interface, and the statement parsing component can be QuickSQL. QuickSQL is an open source project developed by Qihoo360, which aims to provide users with a flexible, fast, and federated (3F) SQL analysis middleware for processing multiple data sources.
[0043] This application realizes the logical isolation of the production environment and the test environment. The core of this application is that there is only one physical cluster at the bottom layer, but the database of the test environment and the database of the production environment are logically divided. The data synchronization between the two databases, that is, the table structure and the data in each table of the two databases are exactly the same, and the difference is that the library names of the two databases are unique and different. In order to realize the data synchronization, it can be realized by CDC synchronization, and it can also be flexibly realized by those skilled in the art.
[0044] The pre-processing component parses the operation environment identifier in the data operation request, determines the operation database object corresponding to the data operation statement in the request, and specifies the database associated with the environment mapping for executing the statement as the target environment database. The syntax parsing layer in the QuickSQL parses, verifies, optimizes, etc. the data operation statement to generate a query plan, and the computing engine layer is responsible for interpreting the query plan into a language recognizable by a specific execution engine, thereby obtaining the parsed data operation statement.
[0045] Step S1300: submit the parsed data operation statement to a data operation cluster for execution, and the cluster executes the corresponding data operation in the target environment database, and obtains the execution result returned by the cluster and pushes it to the user.
[0046] The data operation cluster is to deploy and run Spark on a cluster composed of multiple computers. This mode is also called Spark on YARN mode. In this mode, Spark, as an application of YARN, can use the resource management and scheduling capabilities provided by YARN to optimize resource allocation and task execution in the cluster.
[0047] Submit the parsed data operation statement to the data operation cluster, distribute the to-be-executed code of the data operation corresponding to the statement to different nodes in the cluster for execution, and finally summarize the corresponding execution results to form a final execution result and return it. The returned execution result is pushed to the user to be displayed in the data development system and inform the user that the data operation has been successfully executed.
[0048] It can be known from the typical embodiments of the present application that the technical solution of the present application has many advantages, including but not limited to the following aspects:
[0049] This application first obtains the operating environment identifier and specific data operation statement in the data operation request submitted by the user. In response to the request scheduling event, the request is submitted to the data operation parser to parse the data operation statement according to the operating environment identifier, so that the parsed statement can clearly point to a specific target environment database, laying the foundation for subsequent precise operations. Then, the parsed data operation statement is submitted to the data operation cluster for execution, and the cluster performs the corresponding data operation in the target environment database, and returns the execution result to the user. The whole process is closely linked, which can not only efficiently perform the data operations required by the user, but also ensure the accuracy of the operation. The most important thing is that through the precise parsing and pointing of the operating environment identifier, the data operation can be accurately arranged to the corresponding target environment database for execution, realizing the isolation of operations between different environment databases without interference. This not only ensures the security and integrity of the data, but also avoids the operational errors and data confusion that may be caused by environmental confusion, greatly improves the reliability and stability of data operations, and provides users with an efficient, accurate and safe data operation processing solution.
[0050] In a further embodiment, step S1100, in response to a request scheduling event, submitting the data operation request to a data operation parser, comprises the following steps:
[0051] Step S1110: submit the data operation request to the scheduling system, and the scheduling system confirms whether there is a preset upstream task corresponding to the data operation request. If so, the execution progress of each upstream task is obtained. The upstream task is a task that needs to be executed before parsing the data operation statement in the request;
[0052] In the big data development scenario, some data operations need to rely on other data operations to be executed before execution, and / or, some data operations need to wait for a period of time before execution. For the sake of understanding, for the data operation of counting active users for a whole day, before the data operation, it is necessary to wait for a day, and then execute after the data operation that depends on collecting the detailed information of the user's e-commerce behavior is executed. In this regard, in the data development system, the user can also edit the tasks to be executed before executing the data operation statement, that is, the relative upstream tasks, or edit the tasks to be executed after executing the data operation statement, that is, the relative downstream tasks. The tasks include waiting or other data operations to be executed within a preset time, and the preset time can be pre-set by those skilled in the art according to business needs. For the tasks of editing other data operations to be executed, the user also edits other data operation statements and selects the environment for executing the data operation statement in the data development system, that is, the data operations that the user edits in the data development system belong to individual tasks. Furthermore, the execution order of each task edited by the user is stored in the scheduling system, so that the scheduling system controls the execution of each task in a predefined order.
[0053] After receiving the data operation request, the scheduling system queries whether there is a task that needs to be executed before executing the data operation corresponding to the request, that is, an upstream task. If there is an upstream task, the scheduling system can query the task log to confirm the execution progress of each upstream task. The task log includes the execution progress of all tasks. The execution progress includes pending execution, executing, execution failure, and execution success.
[0054] Step S1120: When all the execution progress items indicate that the execution is completed, submit the data operation request to the data operation parser;
[0055] When the execution progress of each of the upstream tasks is successful, that is, the execution is completed, the scheduling system submits the data operation request to the data operation parser, parses the data operation statement in the request, and obtains the parsed data operation statement.
[0056] Step S1130: When any execution progress indicates that the execution is not completed, control the data operation request to continue waiting until it is scheduled.
[0057] When the execution progress of any of the upstream tasks is execution failure or pending execution, which means that the execution has not been completed, the scheduling system puts the data operation request on hold, so that the execution progress of the data operation corresponding to the request continues to be set to pending execution until the execution progress of each of the upstream tasks indicates that the execution is completed and then it is scheduled.
[0058] In this embodiment, by introducing the concept of upstream tasks in the scheduling system, it is possible to effectively deal with the situation where some data operations need to rely on other data operations or need to wait for a certain period of time before they can be executed, thereby ensuring the orderliness and accuracy of data operations. After receiving the data operation request, the scheduling system will actively query whether there is a corresponding upstream task, and further check the task log to confirm the execution progress of each upstream task. This mechanism enables the system to accurately control the execution order of tasks. Only when all upstream tasks are successfully executed will the data operation request be submitted to the data operation parser for subsequent processing, avoiding data errors or operation failures caused by unfinished upstream tasks, and ensuring the integrity and reliability of data operations. At the same time, for the situation where the upstream task is not completed, the scheduling system will put the data operation request in a waiting state, continuously monitor the progress of the upstream task, and schedule it until it is fully completed, further enhancing the flexibility and stability of the system, improving the efficiency and quality of big data development, and providing strong support for complex data processing processes.
[0059] In a further embodiment, step S1200, the data operation parser performs corresponding parsing on the data operation statement according to the operation environment identifier to obtain the parsed data operation statement, includes the following steps:
[0060] Step S1210: The data operation parser confirms the target operation environment represented by the operation environment identifier, and identifies the original operation database name defined by the preset placeholder in the data operation statement;
[0061] The pre-processing component in the data operation parser parses the operation environment identifier in the data operation request, determines the environment for executing the data operation statement in the request as the target operation environment, and traverses the characters in the statement to determine the character content wrapped by the preset placeholder as the original operation database name. The preset placeholder can be a symbol such as '', "", [], [[]], etc. that can wrap the character content in the middle.
[0062] Step S1220: According to a predetermined library name mapping protocol, the preset placeholder and the original operating database name defined by it are converted into a library name of a database built for the target operating environment.
[0063] The predetermined library name mapping protocol is the character content wrapped by preset placeholders in the data operation statement, and the mapping conversion relationship between the database name of the test environment and the database name of the production environment is pre-agreed. Technical personnel in this field can flexibly set it according to the disclosure here, and can also refer to the settings of subsequent embodiments.
[0064] In one embodiment, the character content wrapped by the preset placeholder in the data operation statement is directly the database name of the production environment, and the database name of the test environment is appended with a naming suffix representing the test compared to the database name of the production environment. For the sake of easy understanding, the database name of the production environment wrapped is: [[kaixi ndou]], and the database name of the test environment is: kaixi ndou_test. In this way, when the target operation environment is the production environment, the character content wrapped by the preset placeholder according to the library name mapping protocol, that is, the original operation database name, is directly the database name of the production environment. Therefore, the preset placeholder and the original operation database name defined by it are replaced with the database name of the production environment; when the target operation environment is the test environment, the character content wrapped by the preset placeholder according to the library name mapping protocol, that is, the original operation database name, is appended with a naming suffix representing the test at the end, and the preset placeholder and the original operation database name defined by it are replaced with the original operation database name appended with the naming suffix.
[0065] In this embodiment, the original operation database name can be quickly located through the accurate confirmation of the operation environment identifier by the data operation parser and the recognition of the preset placeholder in the data operation statement. Then, according to the predetermined library name mapping protocol, the preset placeholder and the original operation database name limited by it are converted into the library name of the database built by the target operation environment. This process cleverly solves the problem of database name switching of data operation statements in different environments. This design not only avoids the tediousness and error-proneness of manually modifying the database name, but also ensures that the data operation statement can be executed accurately in different environments, so that when the user manually edits the data operation statement, he only needs to use the preset placeholder to limit the original operation database name, and then he can automatically adapt to the target operation environment and perform the corresponding conversion, thereby effectively improving the efficiency, quality and convenience of data development, fully reflecting the ingenuity and efficiency of the design, and providing strong support for the environment switching in the big data development scenario, making the entire data operation process smoother and more reliable.
[0066] In a further embodiment, before step S1100, obtaining the data operation request submitted by the user, the following steps are included:
[0067] Step S1000, responding to the temporary storage modification statement event, confirming whether the data operation statement of the previous version is in use, and if it is in use, additionally storing the modified version of the data operation statement;
[0068] It is understandable that during the data development process, considering that data development may not be completed in a short period of time, an intermediate state temporary storage function is needed so that users can save unfinished data operation statements midway and continue development when needed.
[0069] In the data development system, when editing a data operation statement, the user can touch the "temporarily save and modify" control to submit the undeveloped data operation statement and trigger the temporarily save and modify statement event. The data development system first compares whether the corresponding data operation statements before and after submission are consistent. If they are inconsistent, the submitted data operation statement is used as the modified version of the data operation statement, and the corresponding data operation statement before submission is used as the data operation statement before modification. The data development system sends a statement query request to the scheduling system in order to obtain the execution progress of the data operation request carrying the data operation statement before modification. The scheduling system receives and responds to the statement query request and queries the task log, so that the execution progress can be determined and returned to the data development system to respond to the statement query request. When the data development system receives the execution progress, it means that the data operation statement before modification is in use. At this time, in order not to interfere with its use, the data development system additionally stores the data operation statement after modification.
[0070] Step S1010: When the data operation statement of the pre-modification version is not used, the data operation statement is overwritten and stored as the data operation statement of the modified version.
[0071] When the execution progress cannot be received, it means that the data operation statement of the pre-modification version is not used. At this time, there is no problem of interference in use, so the data development system can directly overwrite the stored data operation statement with the data operation statement of the modified version to ensure the latest version.
[0072] In this embodiment, users do not need to maintain versions by themselves, and can safely use the essential development function of temporarily storing modifications provided by the data development system. The system will automatically store and update statements according to their usage status, making version management orderly. Users can focus more on the development of data operation statements without worrying about version conflicts or data loss. This effectively ensures the continuity and stability of data development and provides a more friendly and efficient working environment for data developers.
[0073] In a further embodiment, after step S1300, obtaining the execution result returned by the cluster and pushing it to the user, the following steps are included:
[0074] Step S1400, obtaining the resource consumption cost corresponding to the data operation statement executed by the data operation cluster;
[0075] In one embodiment, the resource usage of the data operation cluster during the execution of the data operation statement may be monitored and analyzed, including memory consumption and runtime, and the memory consumption is multiplied by the runtime to obtain the resource consumption cost. For ease of understanding, an exemplary example is given, where the memory consumption is 100G / minute and the runtime is 10 minutes, then the resource consumption cost is 100G / minute*10 minutes, which equals 1000G.
[0076] Step S1410: When the resource consumption cost exceeds a preset threshold, a cost excess report is generated and pushed to the user.
[0077] The preset threshold may be preset by a person skilled in the art based on prior knowledge or experimental data, as well as cost control strategies and / or budget requirements, and is used to monitor and manage resource costs to capture situations where costs are out of control.
[0078] When the resource consumption cost exceeds the preset threshold, it means that the resource consumption cost is too much at this time. The difference between the resource consumption cost and the preset threshold is calculated, and the difference and the data operation statement are embedded in the preset prompt template to obtain the corresponding prompt text input into the large language model, obtain the cost excess report generated by the large language model, and send it to the user.
[0079] The prompt template includes a task prompt, an excess cost amount to be embedded, and a data operation statement to be embedded. The task description can guide the large language model to measure whether the given excess cost amount is too high according to the given excess cost amount, and analyze the given data operation statement accordingly, give a corresponding statement optimization plan, and analyze the statement optimization plan and its corresponding optimized data operation statement and excess cost amount to form a cost excess report output. A person skilled in the art can flexibly edit the task description according to the disclosure herein. In one embodiment, the task is described as: "The following given data operation statement will cause the resource cost consumed by its execution to exceed the standard, specifically the following given cost excess. Please act as a developer of the database query statement and first measure whether the cost excess is too much. If it is too much, please analyze the operation logic in the statement and find out the part that can be optimized on a large scale, such as whether some query operations can be merged, reducing the full scan of data, etc., in order to greatly reduce the consumed resource cost through large-scale optimization measures; if it is not too much, please analyze the statement from the detail level, such as checking whether the index usage in the statement is reasonable, whether there are unnecessary data conversion steps, whether the constant expression can be calculated in advance, etc., in order to slightly reduce the consumed resource cost through small-scale optimization measures. In the end, you need to form a complete and detailed cost excess report based on your analysis of the cost excess, the statement optimization plan and the optimized statement."
[0080] This embodiment monitors the resource consumption of data operation statements, calculates the cost and compares it with the preset threshold. When the threshold is exceeded, a report containing an optimization plan is generated using a large language model and pushed to the user, thereby achieving effective cost monitoring and optimization, helping users to make optimization suggestions to reduce resource consumption costs while ensuring the quality of data operations, and improving user experience.
[0081] In a further embodiment, before step S1100, responding to the request scheduling event, the following steps are included:
[0082] Step S2210: taking the data operation statement as a target data operation statement, and using a preset statement recognition model to determine the similarity between the target data operation statement and each data operation statement in a preset statement repository;
[0083] The sentence recognition model is a dual-tower model, including two identical text feature representation layers and classification layers with shared parameters. The text feature representation layer is suitable for extracting the semantics of the input text for vector representation, and can be selected from a variety of known models, including but not limited to Bert, RNN, BiLSTM, BiGRU, RoBERTa, ALBert, ERN IE, BERT-WWM, etc. The classifier is suitable for binary classification tasks and can be MLP (feedforward neural network) or FC (fully connected layer). The sentence recognition model is pre-trained to a convergent state and acquires the ability to determine the similarity between two data operation sentences.
[0084] For each data operation statement in the statement repository, the data operation statement and the target operation statement are input into the statement recognition model, and the text feature representation layer encodes the input statement into a high-dimensional semantic vector to capture the vocabulary, grammar and semantic information in the statement. Due to parameter sharing, the two text feature representation layers respectively represent the two statements in the same way, so that the model can understand the semantics of the two statements from the same perspective.
[0085] Next, the two semantic vectors are sent to the classification layer. The classification layer is a network structure suitable for binary classification tasks that further processes the two semantic vectors and calculates the similarity between them. Specifically, the classification layer may quantify the similarity between the two statements by calculating the dot product, cosine similarity, or other distance metrics of the two vectors. Since the model has been pre-trained and has learned how to accurately judge the similarity between two data operation statements based on the characteristics of the semantic vectors, it can output a value between 0 and 1, indicating the similarity between the target data operation statement and each data operation statement in the statement repository, thereby completing the reasoning and calculation of the similarity.
[0086] Step S2220: Filter out data operation statements whose similarity exceeds a preset threshold from the statement storage library, and obtain the corresponding resource consumption costs as the estimated resource consumption costs corresponding to the target data operation statements;
[0087] The similarity between each data operation statement in the statement repository and the target data operation statement is compared with a preset threshold value. The preset threshold value is set according to actual needs and experience, and is used to determine whether two statements are similar enough so that they can be considered to consume the same resource cost when executed.
[0088] Step S2230: Push the estimated resource consumption cost to the user.
[0089] The estimated resource consumption cost of the obtained target data operation statement is pushed to the user.
[0090] In this embodiment, by evaluating the similarity between the reserved statements and the statements submitted by the user based on the statement recognition model, the resource consumption cost of the data operation statements submitted by the user is estimated in advance, so that the user can have a clear expectation of the cost before execution, helping the user to reasonably plan resources and control costs, avoid executing high-cost operations, and improve resource utilization efficiency and the scientific nature of cost management.
[0091] In a further embodiment, before step S2200, taking the data operation statement as the target data operation statement and using a preset statement recognition model to determine the similarity between the target data operation statement and each data operation statement in a preset statement repository, the following steps are included:
[0092] Step S2200, obtaining each data operation statement and its corresponding consumed resource cost, and storing them in a statement storage library;
[0093] For each data operation statement, the resource usage of the data operation cluster during the execution of the data operation statement, including memory consumption and running time, can be monitored and analyzed. The memory consumption is multiplied by the running time to obtain the resource consumption cost.
[0094] Step S2201: Obtain a training set, and train the sentence recognition model to a convergence state so that it can learn the ability to determine the similarity between two data operation sentences.
[0095] The training set contains a large number of paired data operation statement samples and the annotated similarity information between them. These sample statement pairs can be combinations of different data operation statements that have been executed in history, and the annotated similarity information is based on the experience of domain experts or predetermined by other reliable methods, and is used to indicate the semantic and functional similarity of each pair of statements.
[0096] Next, the training set is used to train the preset sentence recognition model. The sentence recognition model is a dual-tower structure, which includes two identical text feature representation layers and a classification layer with shared parameters. During the training process, the model inputs each pair of data operation sentences into two text feature representation layers, which extract the semantic vectors of the sentences. Due to parameter sharing, the model can process different sentences in a consistent manner, ensuring feature extraction from the same semantic perspective.
[0097] Then, the two semantic vectors are sent to the classification layer, whose task is to calculate the similarity between the two vectors and output it as a value. The model will adjust the parameters of each layer in the model through the optimization algorithm according to the similarity information annotated in the training set, so that the similarity value output by the model is as close as possible to the actual similarity of the annotation. This process will be repeated until the performance of the model on the training set no longer improves, that is, the training reaches a convergence state.
[0098] In this embodiment, by constructing a statement repository and fully training the statement recognition model until it converges, the model learns the ability to determine the similarity between two data operation statements, so that the similarity between new and unseen data operation statement pairs can be accurately identified. This can provide a reliable basis for determining the similarity between the target data operation statement and the data operation statement in the statement repository in the subsequent steps, so that statements with high similarity can be effectively screened out, and then their resource consumption cost can be obtained as the estimated cost, so as to achieve a fast and reliable estimation of the data operation cost.
[0099] See also Figure 3 , a data operation processing device provided to meet one of the purposes of the present application is a functional embodiment of the data operation processing method of the present application. On the other hand, the device is a data operation processing device provided to meet one of the purposes of the present application, including a request processing module 1100, a statement parsing module 1200 and a statement execution module 1300, wherein the request processing module 1100 is used to obtain a data operation request submitted by a user, respond to a request scheduling event, and submit the data operation request to a data operation parser, wherein the data operation request includes an operation environment identifier and a data operation statement; the statement parsing module 1200 is used for the data operation parser to perform corresponding parsing on the data operation statement according to the operation environment identifier, obtain the parsed data operation statement, and make its operation database object be designated as the target environment database pointed to by the operation environment identifier; the statement execution module 1300 is used to submit the parsed data operation statement to a data operation cluster for execution, and the cluster executes the corresponding data operation in the target environment database, and obtains the execution result returned by the cluster and pushes it to the user.
[0100] In a further embodiment, the request processing module 1100 includes: a system confirmation submodule, used to submit the data operation request to the scheduling system, and the scheduling system confirms whether there is a preset upstream task corresponding to the data operation request, and if so, obtains the execution progress of each upstream task, wherein the upstream task is a task that needs to be executed before parsing the data operation statement in the request; a request scheduling submodule, used to submit the data operation request to the data operation parser when all the execution progress indicates that the execution is completed; and a request waiting submodule, used to control the data operation request to continue waiting until it is scheduled when any execution progress indicates that the execution is not completed.
[0101] In a further embodiment, the statement parsing module 1200 includes: a library name identification submodule, which is used for the data operation parser to confirm the target operating environment represented by the operating environment identifier and identify the original operating database name defined by the preset placeholder in the data operation statement; and a library name conversion submodule, which is used to convert the preset placeholder and the original operating database name defined by it into a library name of a database built for the target operating environment according to a predetermined library name mapping protocol.
[0102] In a further embodiment, the request processing module 1100 includes: an additional storage submodule for responding to a temporary modification statement event, confirming whether the data operation statement of the previous version is in use, and if so, additionally storing the modified version of the data operation statement; and an overwrite storage submodule for overwriting and storing the data operation statement as the modified version of the data operation statement when the data operation statement of the previous version is not in use.
[0103] In a further embodiment, after the statement execution module 1300, it includes: a cost acquisition sub-module, used to obtain the resource consumption cost corresponding to the data operation cluster executing the data operation statement; a user alarm sub-module, used to generate a cost excess report and push it to the user when the resource consumption cost exceeds a preset threshold.
[0104] In a further embodiment, the request processing module 1100 includes: a similarity identification submodule, which is used to take the data operation statement as the target data operation statement, and use a preset statement recognition model to determine the similarity between the target data operation statement and each data operation statement in a preset statement repository; a cost estimation submodule, which is used to screen out data operation statements whose similarity exceeds a preset threshold from the statement repository, and obtain their corresponding resource consumption costs as the estimated resource consumption costs corresponding to the target data operation statement; and a cost notification submodule, which is used to push the estimated resource consumption costs to the user.
[0105] In a further embodiment, the similarity identification submodule includes: a library construction submodule, which is used to obtain the corresponding resource consumption cost associated with each data operation statement and store it in a statement reserve library; a model training submodule, which is used to obtain a training set and train the statement recognition model to a convergence state so that it can learn the ability to determine the similarity between two data operation statements.
[0106] In order to solve the above technical problems, the present application also provides a computer device. Figure 4 As shown, a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database may store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a data operation processing method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the data operation processing method of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0107] In this embodiment, the processor is used to execute Figure 3 The memory stores the program code and various data required to execute the above modules or submodules. The network interface is used to transmit data between user terminals or servers. The memory in this embodiment stores the program code and data required to execute all modules / submodules in the data operation processing device of this application, and the server can call the program code and data of the server to execute the functions of all submodules.
[0108] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the data operation processing method of any embodiment of the present application.
[0109] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments of the present application can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0110] In summary, the present application can accurately place data operations in the target environment for execution.
[0111] It will be understood by those skilled in the art that the various operations, methods, steps, measures, and schemes in the processes discussed in this application may be alternated, changed, combined, or deleted. Furthermore, other steps, measures, and schemes in the various operations, methods, and processes discussed in this application may also be alternated, changed, rearranged, decomposed, combined, or deleted. Furthermore, the steps, measures, and schemes in the various operations, methods, and processes in the prior art that are open source and in this application may also be alternated, changed, rearranged, decomposed, combined, or deleted.
[0112] The above description is only a partial implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A data operation processing method, characterized in that: The steps include: Obtaining a data operation request submitted by a user, responding to a request scheduling event, and submitting the data operation request to a data operation parser, wherein the data operation request includes an operation environment identifier and a data operation statement; The data operation parser performs corresponding parsing on the data operation statement according to the operation environment identifier to obtain the parsed data operation statement, so that its operation database object is designated as the target environment database pointed to by the operation environment identifier; The parsed data operation statement is submitted to the data operation cluster for execution, and the cluster executes the corresponding data operation in the target environment database, and the execution result returned by the cluster is pushed to the user.
2. The data operation processing method according to claim 1, characterized in that: In response to the request scheduling event, submitting the data operation request to the data operation parser includes the following steps: Submit the data operation request to the scheduling system, and the scheduling system confirms whether there is a preset upstream task corresponding to the data operation request. If so, obtain the execution progress of each upstream task, where the upstream task is a task that needs to be executed before parsing the data operation statement in the request; When all the execution progress items indicate that the execution is completed, submitting the data operation request to the data operation parser; When any execution progress indicates that the execution is not completed, the data operation request is controlled to continue waiting until being scheduled.
3. The data operation processing method according to claim 1, characterized in that: The data operation parser performs corresponding parsing on the data operation statement according to the operation environment identifier to obtain the parsed data operation statement, including the following steps: The data operation parser confirms the target operation environment represented by the operation environment identifier, and identifies the original operation database name defined by the preset placeholder in the data operation statement; According to a predetermined library name mapping protocol, the preset placeholder and the original operating database name defined by it are converted into a library name applied to a database established in the target operating environment.
4. The data operation processing method according to claim 1, characterized in that: Before obtaining the data operation request submitted by the user, the following steps are included: In response to the temporary modification statement event, confirm whether the data operation statement of the previous version is in use, and if it is in use, additionally store the modified version of the data operation statement; When the data operation statement of the pre-modification version is not used, the data operation statement is overwritten and stored as the data operation statement of the post-modification version.
5. The data operation processing method according to claim 1, characterized in that: After the execution result returned by the cluster is obtained and pushed to the user, the following steps are included: Obtaining the resource consumption cost corresponding to the data operation statement executed by the data operation cluster; When the resource consumption cost exceeds a preset threshold, a cost excess report is generated and pushed to the user.
6. The data operation processing method according to claim 1, characterized in that: Before responding to a request to dispatch an event, the following steps are included: Taking the data operation statement as a target data operation statement, and using a preset statement recognition model to determine the similarity between the target data operation statement and each data operation statement in a preset statement repository; Filtering data operation statements whose similarity exceeds a preset threshold from the statement reserve library, and obtaining the corresponding resource consumption costs as the estimated resource consumption costs corresponding to the target data operation statements; The estimated resource consumption cost is pushed to the user.
7. The data operation processing method according to claim 6, characterized in that: The data operation statement is used as a target data operation statement, and before a preset statement recognition model is used to determine the similarity between the target data operation statement and each data operation statement in a preset statement repository, the following steps are included: Obtain each data operation statement and its corresponding resource consumption cost, and store them in the statement repository; A training set is obtained, and a sentence recognition model is trained to a convergence state so that the model can learn the ability to determine the similarity between two data operation sentences.
8. A data operation processing device, characterized in that: include: A request processing module, used to obtain a data operation request submitted by a user, respond to a request scheduling event, and submit the data operation request to a data operation parser, wherein the data operation request includes an operation environment identifier and a data operation statement; A statement parsing module, configured to allow the data operation parser to parse the data operation statement accordingly according to the operation environment identifier, obtain the parsed data operation statement, and make its operation database object designated as the target environment database pointed to by the operation environment identifier; The statement execution module is used to submit the parsed data operation statement to the data operation cluster for execution, and the cluster executes the corresponding data operation in the target environment database, and the execution result returned by the cluster is pushed to the user.
9. A computer device comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.