Data query method and device, equipment, storage medium and program product
By fragmenting the data of table documents and using parallel query methods, the problem of low efficiency of table documents data query in the prior art is solved, and faster data query speed and higher efficiency are achieved.
Patent Information
- Application Number
- CN202510028120.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is less efficient when querying data in table documents, especially when the data volume is large, the query speed is slower.
By sharding the data of the table document, multiple data fragments are obtained, and parallel data query method is used to query the target data that matches the keyword characters in the target data shard.
It improves the speed and efficiency of data query, and compared with the overall scanning solution, it speeds up data query and simplifies query logic.
Smart Images

Figure CN120011404A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet technology, and in particular to a data query method, device, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] When performing data query on a table document, related technologies mostly scan the entire table document as a whole to find the queried data. However, if there is a lot of data in the table document, the data query speed will be slow, that is, the data query efficiency will be low. Summary of the invention
[0003] Embodiments of the present application provide a data query method, device, electronic device, computer-readable storage medium, and computer program product, which can improve the efficiency of data query in table documents.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present application provides a data query method, including:
[0006] Obtaining key characters, wherein the key characters are used to perform data query in a table document;
[0007] Determining a target table page for which data query is required from at least one table page included in the table document;
[0008] Acquire a plurality of data slices corresponding to the target table page, wherein the plurality of data slices are obtained by performing slice processing on the data included in the target table page;
[0009] Determine a query range for the target table page, and select at least one target data shard corresponding to the query range from the multiple data shards;
[0010] For each of the target data slices, a parallel data query method is adopted to query the target data matching the key character in the at least one target data slice.
[0011] The present application provides a data query device, including:
[0012] A first acquisition module, used to acquire key characters, wherein the key characters are used to perform data query in a table document;
[0013] A first determining module, used to determine a target table page that needs to be queried for data from at least one table page included in the table document;
[0014] A second acquisition module is used to acquire a plurality of data slices corresponding to the target table page, wherein the plurality of data slices are obtained by performing slice processing on the data included in the target table page;
[0015] A second determination module is used to determine a query range for the target table page, and select at least one target data slice corresponding to the query range from the multiple data slices;
[0016] The query module is used to query the target data matching the key character in the at least one target data slice by adopting a parallel data query method for each target data slice.
[0017] In the above scheme, the second acquisition module is also used to acquire the row data of each row in the target table page, and based on the row data, divide the data of the target table page to obtain the multiple data slices, each of the data slices includes at least one row of row data, and the row data included in different data slices are different; or, acquire the column data of each column in the target table page, and based on the column data, divide the data of the target table page to obtain the multiple data slices, each of the data slices includes at least one column of column data, and the column data included in different data slices are different.
[0018] In the above scheme, the device also includes a secondary sharding module, which is used to perform the following processing on each of the data shards to obtain multiple data sub-shards corresponding to each of the data shards: when the data shard includes multiple data types, the data shard is sharded based on the multiple data types to obtain multiple data sub-shards, and different data sub-shards correspond to different data types; the query module is also used to perform the following processing on each of the target data shards: determine the data type of the key character, and select a target data sub-shard whose data type is consistent with the data type of the key character from the multiple data sub-shards corresponding to the target data shard; in the target data sub-shard, query the target data that matches the key character.
[0019] In the above scheme, the device also includes an update module, which is used to obtain updated new data slices in response to the data slices being updated; based on the new data slices, update the data sub-slices corresponding to each of the new data slices.
[0020] In the above scheme, the device also includes a detection module, which is used to periodically detect the data volume of each of the data slices to obtain a detection result; when it is determined based on the detection result that there is a first data slice among the multiple data slices, a second data slice is selected from the multiple data slices, and the first data slice and the second data slice are merged; when it is determined based on the detection result that there is a third data slice among the multiple data slices, the third data slice is sliced; wherein the data volume of the first data slice is less than a first threshold, the data volume of the third data slice is greater than a second threshold, and the second threshold is greater than the first threshold.
[0021] In the above scheme, the detection module is also used to select at least one of the following from the multiple data slices as the second data slice: a data slice with a data volume less than the first threshold; a data slice with a data volume not less than the first threshold and a data volume less than a third threshold, wherein the third threshold is less than the second threshold; a data slice with the smallest data volume except the first data slice.
[0022] In the above scheme, the first determination module is also used to receive an operation request, the operation request includes at least one of a data query request and a data replacement request; parse the operation request to obtain a table page identifier carried by the operation request, the table page identifier is used to indicate the table page searched based on the keyword; from at least one table page included in the table document, select the table page corresponding to the table page identifier as the target table page.
[0023] In the above scheme, the second determination module is also used to, when the operation request also carries a query range for the target table page, use the query range carried in the operation request as the query range for the target table page; when the operation request does not carry a query range for the target table page, use multiple data shards corresponding to the target table page as the query range for the target table page.
[0024] In the above scheme, the table document is an online document for collaborative editing by multiple objects, and the key character is determined by the first object among the multiple objects; the device also includes a third acquisition module, and the third acquisition module is used to obtain each updated target data slice when the second object among the multiple objects performs a data update operation on the target data slice; the query module is also used to use a parallel data query method for each updated target data slice to query the target data matching the key character in the updated target data slice.
[0025] In the above scheme, the third acquisition module is also used to obtain the priority of the second object and the priority of the first object in response to the data update operation; in response to the priority of the second object being higher than the priority of the first object, each of the target data slices is updated based on the data update operation to obtain updated target data slices.
[0026] In the above scheme, the key character is determined by a received data query request; the device also includes a replacement module, which is used to record the first position of the target data in the table document in response to querying the target data matching the key character; when a data replacement request for the key character in the query range is received, the data replacement request is parsed to obtain other characters carried by the data replacement request for replacing the key character; the first position of the recorded target data is obtained; and based on the first position, the target data is replaced with the other characters.
[0027] In the above scheme, the replacement module is also used to obtain other data in the target table page except the target data, determine the second position of the other data in the table document, and cache the other data and the second position; in response to receiving a data update request for the other data, when the time interval between the receiving time of the data update request and the receiving time of the data replacement request is less than the interval threshold, obtain the cached other data and the second position, and replace the current data at the second position in the query range with the other data.
[0028] In the above scheme, the key character is carried in the data query request, and the data query request is stored in a request queue, and the request queue includes multiple requests; the device also includes a state management module, and the state management module is used to obtain the state machine corresponding to the data query request and control the state machine to be in an initial state; wherein the state machine is used to indicate the processing status of the data query request; the first determination module is also used to control the state machine to switch from the initial state to the loading state, and when the state machine is in the loading state, determine the target table page that needs to be queried for data from at least one table page included in the table document; the second determination module 4554 is also used to control the The state machine switches from the loading state to the preparation state, and when the state machine is in the preparation state, determines the query range for the target table page; the query module 4555 is used to control the state machine to switch from the preparation state to the search state, and when the state machine is in the search state, searches for target data matching the keyword in at least one target data slice, and switches the state machine from the search state to the termination state; the device also includes a response module, and the response module is used to determine the end of the response to the data query request if the state machine is in the termination state, and respond to the request in the request queue that is after the data query request.
[0029] An embodiment of the present application provides an electronic device, including:
[0030] Memory for storing computer executable instructions or computer programs;
[0031] The processor is used to implement the data query method provided in the embodiment of the present application when executing the computer executable instructions or computer programs stored in the memory.
[0032] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions or a computer program for causing a processor to execute and implement the data query method provided in the embodiment of the present application.
[0033] The embodiment of the present application provides a computer program product, which includes computer executable instructions or computer programs, and the computer executable instructions or computer programs are stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instructions or computer programs from the computer-readable storage medium, and the processor executes the computer executable instructions or computer programs, so that the electronic device executes the data query method provided in the embodiment of the present application.
[0034] The embodiments of the present application have the following beneficial effects:
[0035] After determining the key characters for data query in the table document and the target table page for data query in the table document, multiple data slices obtained by slicing the data included in the target table page are obtained, and then at least one target data slice corresponding to the query range is selected from the multiple data slices, so that for each target data slice, a parallel data query method is adopted, and the target data matching the key characters is searched in at least one target data slice. In this way, by slicing the data of the table document and then using a parallel processing method for data query for each data slice, the data query speed is accelerated and the data query efficiency is improved compared with the solution of scanning the entire table document as a whole for data query. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a schematic diagram of the architecture of a data query system 100 provided in an embodiment of the present application;
[0037] Figure 2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0038] Figure 3 It is a flowchart of a data query method provided in an embodiment of the present application;
[0039] Figure 4 is a schematic diagram of key characters provided in the embodiments of the present application;
[0040] Figure 5 is a schematic diagram of a table page provided in an embodiment of the present application;
[0041] Figure 6 is a schematic diagram of a process for determining multiple data shards provided in an embodiment of the present application;
[0042] Figure 7 This is a first schematic diagram for determining a query scope provided in an embodiment of the present application;
[0043] Figure 8 This is a second schematic diagram for determining a query scope provided in an embodiment of the present application;
[0044] Fig. 9 This is a first product interface diagram of the data query process provided by the embodiment of the present application;
[0045] Fig.10 This is a second product interface diagram of the data query process provided by the embodiment of the present application;
[0046] Fig.11 This is a third product interface diagram of the data query process provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0048] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0049] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0051] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0052] 1) In response, it is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed may be in real time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.
[0053] 2) Client, also known as the user end, refers to the program corresponding to the server that provides local services to users. Except for some applications that can only run locally, it is generally installed on the terminal and needs to cooperate with the server to run, that is, it requires corresponding servers and service programs in the network to provide corresponding services. In this way, specific communication connections need to be established on the client and server sides to ensure the normal operation of the application, such as virtual scene clients (such as game clients) and video clients.
[0054] 3) Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0055] 4) State machine, a discrete event model used to represent multiple states and the transitions between each state, can be viewed as a directed graph, where each node represents a state and each edge represents the trigger condition for transitioning from one state to another.
[0056] See also Figure 1 , Figure 1 It is a schematic diagram of the architecture of the data query system 100 provided in an embodiment of the present application, a terminal (terminal 400 is shown as an example), and the terminal 400 is connected to the server 200 through a network 300, wherein the network 300 can be a wide area network or a local area network, or a combination of the two, and data transmission is achieved using wireless or wired links.
[0057] The terminal 400 is used to receive key characters for the table document and send the key characters to the server.
[0058] The server 200 is used to obtain key characters, which are used to perform data query in a table document; determine a target table page on which data query is required from at least one table page included in the table document; obtain multiple data slices corresponding to the target table page, wherein the multiple data slices are obtained by performing slice processing on the data included in the target table page; determine a query range for the target table page, and select at least one target data slice corresponding to the query range from the multiple data slices; and for each target data slice, use a parallel data query method to query target data matching the key characters in at least one target data slice.
[0059] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Deliver Network), and big data and artificial intelligence platforms. The terminal 400 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a set-top box, an intelligent voice interaction device, a smart home appliance, a virtual reality device, a vehicle-mounted terminal, an aircraft, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device, an intelligent speaker, and a smart watch, but is not limited thereto. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0060] Next, an electronic device that implements the data query method provided in the embodiment of the present application is described. Figure 2 , Figure 2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may be a server or a terminal. Figure 1 Take the server shown in as an example, Figure 2 The electronic device shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the terminal 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .
[0061] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0062] The user interface 430 includes one or more output devices 431 that enable display of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0063] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0064] The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0065] In some embodiments, memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0066] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0067] A network communication module 452, used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB), etc.;
[0068] a presentation module 453 for enabling display of information via one or more output devices 431 (e.g., display screen, speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripherals and displaying content and information);
[0069] The input processing module 454 is used to detect one or more user inputs or interactions from one of the one or more input devices 432 and translate the detected inputs or interactions.
[0070] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 2The data query device 455 stored in the memory 450 is shown, which can be software in the form of a program and a plug-in, etc., and includes the following software modules: a first acquisition module 4551, a first determination module 4552, a second acquisition module 4553, a second determination module 4554 and a query module 4555. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.
[0071] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the data query device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the data query method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components.
[0072] In some embodiments, the terminal or server can implement the data query method provided by the embodiment of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application (APP, Application), that is, a local client, that is, a program that needs to be installed in the operating system to run, such as an instant messaging APP, a web browser APP; it can also be a small program, that is, a program that can be run only by downloading it to a browser environment; it can also be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of client, module or plug-in.
[0073] Based on the above description of the data query system and electronic device provided by the embodiment of the present application, the data query method provided by the embodiment of the present application is described below. In actual implementation, the data query method provided by the embodiment of the present application can be implemented by the terminal or the server alone, or by the terminal and the server in collaboration, so that Figure 1 The server 200 in the embodiment of the present application alone executes the data query method provided by the embodiment of the present application as an example for explanation. Figure 3 , Figure 3 is a flow chart of the data query method provided in the embodiment of the present application. Next, Figure 3The steps shown are explained.
[0074] Step 101: The server obtains key characters, which are used to perform data query in a table document.
[0075] It should be noted that the key characters are the characters to be queried, that is, the characters you want to find in the table document, which can be numbers, text, formulas, letters, etc.; and the process of the server obtaining the key characters can be an operation request for the table document sent by the receiving terminal, wherein the operation request can be a data query request or a data replacement request; thereby parsing the operation request to obtain the key characters carried by the operation request.
[0076] It should be noted that the key characters are input by the user, for example, see Figure 4 , Figure 4 is a schematic diagram of key characters provided in the embodiment of the present application, based on Figure 4 , 14 in the dotted box 401 is the key character input by the user; thus in response to the input operation, after the input key character is displayed, in response to the trigger operation of the query control indicated by 402, an operation request carrying the key character is sent to the server.
[0077] It should be noted that when the operation request is a data query request, the key characters carried in the data query request need to be searched in the table document; when the operation request is a data replacement request, the key characters carried in the data replacement request need to be searched in the table document first, and then the key characters are replaced after being found.
[0078] Step 102: determine a target table page for which data query is required from at least one table page included in the table document.
[0079] In actual implementation, the process of determining a target table page that needs to be queried for data from at least one table page included in a table document may be: receiving an operation request, the operation request including at least one of a data query request and a data replacement request; parsing the operation request to obtain a table page identifier carried by the operation request, the table page identifier being used to indicate a table page to be searched based on a keyword; and selecting a table page corresponding to the table page identifier from at least one table page included in the table document as the target table page.
[0080] It should be noted that the number of target table pages can be one or more. For example, when the number of table page identifiers carried by the operation request is one, the number of target table pages is also one. When the number of table page identifiers carried by the operation request is multiple, the number of target table pages is also multiple. At the same time, a table document includes at least one table page. Here, the table document is a workbook, and the table page is a worksheet in the workbook. For example, see Figure 5 , Figure 5 is a schematic diagram of a table page provided in an embodiment of the present application, based on Figure 5 , Figure 5 The indicated entire form document includes two form pages, that is, two worksheets, such as Worksheet 1 and Worksheet 2 indicated in the dotted frame 501 .
[0081] In actual implementation, each table page has a corresponding identifier. After obtaining the table page identifier carried by the operation request, the obtained table page identifier is matched with the identifier of each table page respectively, and based on the matching result, the table page corresponding to the identifier consistent with the obtained table page identifier is selected from multiple table pages as the target table page.
[0082] In this way, based on the table page identifier carried by the operation request, not only can the specific table page be accurately located, which can avoid confusion and errors when processing large amounts of table data, but also the corresponding table page can be quickly found by using the table page identifier as a unique identifier, thereby improving data query efficiency during the data query process; in addition, in a multi-user or multi-task environment, it can also ensure that the correct table page can be accurately accessed when responding to each operation request, which helps maintain data consistency.
[0083] It should be noted that when an operation request sent by a terminal is received, a request queue will be created. The request queue is a first-in-first-out data structure used to store and manage all operation requests to be responded to. Each request has a corresponding table page ID as a table page identifier; among them, one operation request only corresponds to one or more table page identifiers.
[0084] Step 103 : Acquire multiple data slices corresponding to the target table page, where the multiple data slices are obtained by performing slice processing on the data included in the target table page.
[0085] It should be noted that the multiple data slices corresponding to the target table page can be obtained by slicing the data included in the target table page in advance, or can be obtained by slicing the data included in the target table page in real time. Next, based on these two situations, the process of obtaining the multiple data slices corresponding to the target table page is explained.
[0086] In some embodiments, multiple data slices are obtained by slicing the data included in the target table page in real time. The process of obtaining multiple data slices corresponding to the target table page is also the process of constructing multiple data slices corresponding to the target table page. Specifically, the process of obtaining multiple data slices corresponding to the target table page can be: obtaining row data of each row in the target table page, and dividing the data of the target table page based on the row data to obtain multiple data slices, each data slice includes at least one row of row data, and different data slices include different row data; or obtaining column data of each column in the target table page, and dividing the data of the target table page based on the column data to obtain multiple data slices, each data slice includes at least one column of column data, and different data slices include different column data.
[0087] It should be noted that dividing the data of the target table page based on the row data to obtain multiple data slices means dividing the data of the target table page according to the row to obtain multiple local data areas or multiple data fragments, that is, multiple data slices; thus, different data slices include different row data, and each data slice is a collection of at least one row of data in the table document;
[0088] Correspondingly, based on the column data, the data of the target table page is divided to obtain multiple data slices, which means that the data of the target table page is divided according to the columns to obtain multiple local data areas or multiple data fragments, that is, multiple data slices; thus, different column data are included in different data slices, and each data slice is a collection of at least one column of data in the table document.
[0089] For example, see Figure 6 , Figure 6 is a schematic diagram of a process for determining multiple data fragments provided in an embodiment of the present application, based on Figure 6 , according to the row Figure 6 The data of the indicated table page is sliced, that is, the data from the 1st to 9th rows indicated in the dotted box 601 is used as one data slice, and the data from the 10th to 18th rows indicated in the dotted box 602 is used as another data slice.
[0090] In other embodiments, when multiple data slices are obtained by slicing the data included in the target table page in advance, before obtaining the multiple data slices corresponding to the target table page, it is also possible to obtain row data for each row in the target table page, and based on the row data, divide the data of the target table page to obtain multiple data slices, each data slice includes at least one row of row data, and different data slices include different row data; or, obtain column data for each column in the target table page, and based on the column data, divide the data of the target table page to obtain multiple data slices, each data slice includes at least one column of column data, and different data slices include different column data; store the multiple data slices; and the process of obtaining the multiple data slices corresponding to the target table page may be to obtain the multiple pre-stored data slices.
[0091] It should be noted that the process of dividing the data of the target table page based on row data or column data to obtain multiple data slices is as described above, and this embodiment of the present application will not be elaborated on.
[0092] In this way, the data of the table document is sharded by rows or columns. After the data is sharded, each fragment, i.e., data shard, can be processed independently. This not only improves the speed of data processing, but also for large table data, loading all the data at one time may consume a lot of memory. The sharding processing can also only load the data fragments required for processing, reducing memory usage; at the same time, for specific data query tasks, only operations need to be performed on the relevant data shards, which simplifies the query logic and improves query efficiency.
[0093] In some embodiments, after the data of the target table page is divided to obtain multiple data slices, each data slice can be further sliced. Specifically, after the data of the target table page is divided to obtain multiple data slices, the following processing can be performed on each data slice to obtain multiple data sub-slices corresponding to each data slice: when the data slice includes multiple data types, the data slice is sliced based on the multiple data types to obtain multiple data sub-slices, and different data sub-slices correspond to different data types; thus, the subsequent process of querying the target data matching the key character in at least one target data slice can be to perform the following processing on each target data slice: determine the data type of the key character, and select the target data sub-slice whose data type is consistent with the data type of the key character from the multiple data sub-slices corresponding to the target data slice; query the target data matching the key character in the target data sub-slice.
[0094] It should be noted that multiple data types include formulas, values, text, etc., so based on the multiple data types, the data slice is further sliced and processed to obtain multiple data sub-slices, and one data sub-slice corresponds to one data type; when the data slice is sliced and processed based on multiple data types to obtain multiple data sub-slices, the data type of the key character is determined, such as a formula or a value, so that the data type of the key character is matched with the data type of the multiple data sub-slices corresponding to the target data slice, and then based on the matching result, the target data sub-slice whose data type is consistent with the data type of the key character is selected from the multiple data sub-slices corresponding to the target data slice, so as to query the target data matching the key character in the target data sub-slice.
[0095] In this way, the data shards are re-sharded based on the data type, and then the target data sub-shard whose data type is consistent with the data type of the key character is selected from the multiple data sub-shards corresponding to the target data shard, so that the target data matching the key character is queried in the target data sub-shard; in this way, the data shards to be queried are preliminarily screened based on the data type, and data queries are performed only in the data shards related to the data query task, that is, the query range is dynamically adjusted based on the query content, which not only reduces the number of data shards that need to be queried and improves the speed of data processing, but also simplifies the query logic and improves the query efficiency.
[0096] In actual implementation, the data shards are sharded based on multiple data types to obtain multiple data sub-shards. If the data shards are updated, the data sub-shards under the corresponding data shards will also be updated. Specifically, based on multiple data types, the data shards are sharded to obtain multiple data sub-shards. Then, in response to updates to the data shards, updated new data shards can be obtained; based on the new data shards, the data sub-shards corresponding to each new data shard can be updated.
[0097] It should be noted that an update to a data shard refers to an addition, deletion, or change of data in the data shard, or a merge or split operation between data shards, etc., which is not limited in the embodiments of the present application; when an update to a data shard occurs, based on the updated new data shard, the data sub-shards under the corresponding new data shard are also updated in accordance with the update method of the corresponding new data shard.
[0098] In this way, when a data shard is updated, the data sub-shards under the corresponding data shard will also be updated. This ensures that the data in all data sub-shards are up-to-date, maintains data consistency, and ensures that all users or systems can obtain accurate information when accessing data, which ensures the accuracy of data queries.
[0099] In actual implementation, after the data of the target table page is divided to obtain multiple data slices, the data volume of the data slices can be detected to determine whether to merge or split the data slices. Specifically, after the data of the target table page is divided to obtain multiple data slices, the data volume of each data slice can be periodically detected to obtain a detection result; when it is determined based on the detection result that there is a first data slice among the multiple data slices, a second data slice is selected from the multiple data slices, and the first data slice is merged with the second data slice; when it is determined based on the detection result that there is a third data slice among the multiple data slices, the third data slice is sliced; wherein, the data volume of the first data slice is less than the first threshold, the data volume of the third data slice is greater than the second threshold, and the second threshold is greater than the first threshold.
[0100] It should be noted that the detection period here can be pre-set, for example, it can be one hour or half a day, etc., and the embodiments of the present application do not limit this; similarly, the first threshold and the second threshold are also pre-set, and the embodiments of the present application do not limit this; when the detection result indicates that the data volume of the corresponding data slice is less than the first threshold, the corresponding data slice is determined as the first data slice, and the first data slice is merged with the second data slice selected from the multiple data slices, that is, the first data slice and the second data slice are merged into one data slice; when the detection result indicates that the data volume of the corresponding data slice is greater than the second threshold, the corresponding data slice is determined as the third data slice, and the third data slice is split, that is, it is sliced again.
[0101] In this way, based on the amount of data in the data shards, the data shards are merged or split; in this way, when the amount of data is less than the first threshold, the corresponding data shards are merged, which reduces the total number of data shards, not only reduces the maintenance cost of the data shards, simplifies the management of the data shards, but also reduces the occupation of storage resources because the space waste between shards can be eliminated; at the same time, by merging the data shards, the number of shards that need to be accessed when performing data queries is reduced, thereby improving query performance;
[0102] When the data volume is larger than the second threshold, the corresponding data shards are split, which can reduce the load of a single data shard and avoid performance bottlenecks caused by a single data shard being too large, thereby improving query results and ensuring the stability of the data query process; at the same time, as the amount of data grows, splitting the data shards with large data volumes can provide better scalability, thereby adapting to the increase in data volume and ensuring efficient data query when the amount of data increases.
[0103] In actual implementation, the number of second data shards can be one or more. For the process of selecting the second data shard from multiple data shards, specifically, at least one of the following is selected from the multiple data shards as the second data shard: a data shard with a data volume less than a first threshold; a data shard with a data volume not less than the first threshold and a data volume less than a third threshold, wherein the third threshold is less than the second threshold; a data shard with the smallest data volume except the first data shard.
[0104] It should be noted that, in the case where the second data shard is a data shard with a data volume less than the first threshold, when there are multiple data shards with a data volume less than the first threshold, the data shards with a data volume less than the first threshold can be merged. Specifically, when the detection result indicates that there are multiple data shards with a data volume less than the first threshold, the other data shards among the multiple data shards except the first data shard are used as second data shards, thereby merging the first data shard with the second data shard, that is, merging the multiple data shards; for example, when there are two data shards with a data volume less than the first threshold, for one of the two data shards, after determining the data shard as the first data shard, the other data shard is determined as the second data shard, thereby merging the two data shards;
[0105] In the case where the second data shard is a data shard whose data volume is not less than the first threshold and whose data volume is less than the third threshold, based on the detection result, any one or more data shards whose data volume is greater than or equal to the first threshold and less than the third threshold are selected from the multiple data shards as the second data shards, thereby merging the first data shard with the second data shard;
[0106] In the case where the second data shard is the data shard with the smallest data volume except the first data shard, based on the detection result, the data shard with the smallest data volume is selected from the data shards except the first data shard as the second data shard, thereby merging the first data shard with the second data shard.
[0107] In this way, the second data shards are selected based on different methods, which enriches the types of the second data shards and ensures that the second data shards can be selected in any case, that is, the merging process of the data shards can be realized.
[0108] Step 104 : determine a query range for the target table page, and select at least one target data shard corresponding to the query range from the multiple data shards.
[0109] In actual implementation, as described above, the key characters are determined by the operation request, that is, the operation request also carries the key characters, and accordingly, the operation request can also carry the query range, so that the query range for the target table page is determined by parsing the operation request; specifically, the process of determining the query range for the target table page may be, when the operation request also carries the query range for the target table page, using the query range carried in the operation request as the query range for the target table page; when the operation request does not carry the query range for the target table page, using the multiple data shards corresponding to the target table page as the query range for the target table page.
[0110] It should be noted that the query range is used to indicate the queried data content, that is, the queried data slice, so that when the operation request carries the query range for the target table page, the query range carried in the operation request will be used as the query range for the target table page; here, before sending the operation request to the server, the terminal will receive a selection operation for the query range, and thus determine the query range in response to the selection operation for the query range, and then generate an operation request carrying key characters and the query range in response to the query instruction for the table document, and then send the operation request to the server.
[0111] The selection operation for the query range may be a selection operation for options indicating different query ranges in the table document, or may be a size adjustment operation for a range box for selecting the query range in the table document, which is not limited in the embodiments of the present application; for example, see Figure 7 , Figure 7 is a first schematic diagram of determining a query scope provided in an embodiment of the present application, based on Figure 7 ,exist Figure 7 The indicated table page displays a range box for selecting a query range as indicated by 701 , so that in response to a size adjustment operation on the range box, the range indicated by the size-adjusted range box is used as the query range;
[0112] Alternatively, see Figure 8 , Figure 8 is a second schematic diagram of determining a query range provided in an embodiment of the present application, based on Figure 8 ,exist Figure 8 The indicated table page displays multiple options for indicating different query ranges as indicated by the dotted box 801, so that in response to the selection operation of the target option among the multiple options, that is, the "Search only in formula" option, the data corresponding to the formula in the table page is used as the query range.
[0113] When the operation request does not carry the query range for the target table page, it means that the terminal has not received the selection operation for the query range, and thus uses the default query range as the query range for the target table page. Here, the default query range is the content of the entire target table page, that is, the multiple data shards corresponding to the target table page.
[0114] In this way, by clarifying the query scope, you can locate the data that needs to be queried more quickly, reduce unnecessary full table scans, and significantly reduce the amount of data that needs to be processed, thereby improving the data query speed.
[0115] In some embodiments, the table document is an online document for collaborative editing by multiple objects, and the key character is determined by the first object among the multiple objects; thus, after selecting at least one target data slice corresponding to the query range from the multiple data slices, it is also possible to obtain each updated target data slice when the second object among the multiple objects performs a data update operation on the target data slice; and in a subsequent process, for each target data slice, a parallel data query method is used to query the target data matching the key character in at least one target data slice, which can be a process of querying the target data matching the key character in the updated target data slice using a parallel data query method.
[0116] It should be noted that the object here is the user of the spreadsheet document, and the first object among the multiple objects refers to the object for determining the key characters. For example, when the key characters are determined by an operation request, the first object is the object that triggers the operation request, that is, the object that performs the operation, and the second object is any object among the multiple objects except the first object; and the data update operation refers to the addition, deletion or modification of data in the target data slice, etc., which is not limited in this embodiment of the present application. Therefore, when the second object among the multiple objects performs a data update operation on the target data slice, the updated target data slices are obtained, and data query is performed based on the updated target data slices.
[0117] In this way, when the spreadsheet document is an online document for collaborative editing by multiple users, if other users edit the online document, data queries will be performed based on the edited data, that is, each data query is based on the latest data. In this way, it is ensured that all users or systems can obtain the latest and accurate information when accessing the data, that is, the accuracy of data queries is ensured.
[0118] In actual implementation, not every second object that performs a data update operation will obtain an updated target data shard, and then perform data query based on the updated target data shard. Specifically, the process of obtaining each updated target data shard may be, in response to the data update operation, obtaining the priority of the second object and the priority of the first object; in response to the priority of the second object being higher than the priority of the first object, updating each target data shard based on the data update operation to obtain an updated target data shard.
[0119] It should be noted that the priority of an object may be the account level or authority of the object, etc.; after obtaining the priority of the second object and the priority of the first object, the priority of the second object is compared with the priority of the first object. When the priority of the second object is higher than the priority of the first object, each target data slice is updated based on the data update operation to obtain an updated target data slice; when the priority of the second object is lower than or equal to the priority of the first object, the data of each target data slice is kept unchanged;
[0120] In this way, data query will be performed based on the updated data shard only when the priority of the second object is higher than that of the first object. This reduces the possibility of conflict when multiple users edit documents at the same time. At the same time, by setting priorities, the editing permissions of different users on the table document can be controlled to ensure the accuracy and consistency of the table data.
[0121] Step 105 , for each target data slice, a parallel data query method is used to query the target data matching the key character in at least one target data slice.
[0122] In actual implementation, the target data matching the key characters refers to the data identical to the key characters. For each target data slice, a parallel data query method is adopted. After querying the target data matching the key characters in at least one target data slice, if the target data matching the key characters is found, the target data is sent to the terminal, so that the terminal renders the query result, that is, the target data, so that the user can see the search result in real time and improve the user experience; if the target data matching the key characters is not found, a query result indicating that the target data matching the key characters is not found is sent to the terminal, so that the terminal renders the query result, so that the user can determine that there is no data matching the key characters in the relevant target table page, so as to determine whether the key characters are entered incorrectly, or query the key characters in other table pages.
[0123] In some embodiments, as described above, the operation request may be a data query request or a data replacement request. When the operation request is a data query request, after the target data matching the key characters is queried, the target data may be replaced; specifically, the key characters are determined by the received data query request; in at least one target data slice, after the target data matching the key characters is queried, in response to the target data matching the key characters being queried, the first position of the target data in the table document is recorded; when a data replacement request for the key characters in the query range is received, the data replacement request is parsed to obtain other characters carried by the data replacement request for replacing the key characters; the first position of the recorded target data is obtained; and based on the first position, the target data is replaced with other characters.
[0124] It should be noted that, if the target data is the data in the entire cell, the first position refers to the position of the cell; if the target data is partial data in a cell, the first position refers to the position of the cell and the position of the partial data in the cell; for example, the target data is "Hello", and the data of the cell where the target data is located is "Hello, I am XX", then the first position of the target data includes the position of the cell, and the first and second digits of the characters included in the target data "Hello" in the cell; thus, after obtaining the other characters carried by the data replacement request for replacing the key characters, the target data is replaced with the other characters based on the first position of the record, that is, after obtaining the other characters carried by the data replacement request for replacing the key characters, the target data to be replaced is located based on the first position, and the target data is replaced with the other characters.
[0125] In this way, after searching for the target data that matches the key characters, by recording the position of the target data, when the target data needs to be replaced, the target data is directly replaced based on the recorded position. In this way, the data to be replaced can be directly accessed based on the recorded position without traversing the entire data shard, which not only improves the efficiency of data replacement but also reduces resource consumption.
[0126] In other embodiments, when the operation request is a data replacement request, the target data can be replaced after the target data matching the key characters are queried. Specifically, in response to the target data matching the key characters being queried, other characters carried in the data replacement request for replacing the key characters are obtained, and the target data is replaced with the other characters.
[0127] It should be noted that the other characters here are obtained together with the key characters when parsing the data replacement request.
[0128] In actual implementation, based on the first position, before replacing the target data with other characters, it is also possible to obtain other data in the target table page except the target data, determine the second position of the other data in the table document, and cache the other data and the second position; thus, based on the first position, after replacing the target data with other characters, it is also possible to, in response to receiving a data update request for other data, obtain the cached other data and the second position when the time interval between the receipt time of the data update request and the receipt time of the data replacement request is less than an interval threshold, and replace the current data at the second position in the query range with the other data.
[0129] It should be noted that the time of receiving the data update request for other data may be before the target data is replaced with other characters, or may be before or after the target data is replaced with other characters, and this is not limited in the embodiments of the present application; at the same time, the time interval threshold is also preset, so that when the time interval between the time of receiving the data update request and the time of receiving the data replacement request is less than the time interval threshold, the data update request is determined to be an erroneous operation, and then the other cached data and the second position are obtained, and the current data at the second position in the query range is replaced with the other data;
[0130] In addition, other data except the target data in the target table page is obtained, where the other data may be all data except the target data in the target table page, or data except the target data in the target range in the target table page, where the target range may be a query range, or a circular area with the target data as the center and the target distance as the radius, etc., wherein the target distance is pre-set, and this embodiment of the present application does not limit this;
[0131] In this way, in order to prevent data loss due to misoperation, the original data in the selected range is automatically saved to a temporary storage area, so that after the replacement is completed, the original data can be restored, thus ensuring the accuracy of the table data.
[0132] In some embodiments, as described above, the key character is carried in the data query request, the data query request is stored in the request queue, and the request queue includes multiple requests; thus, after obtaining the key character, it is also possible to obtain the state machine corresponding to the data query request, and control the state machine to be in the initial state; wherein the state machine is used to indicate the processing status of the data query request;
[0133] It should be noted that, for the process of obtaining the state machine corresponding to the data query request, when the state machine and the request are in a one-to-one correspondence, the state machine corresponding to the data query request can be constructed; when the state machine and the request are in a one-to-many relationship, that is, one state machine corresponds to multiple requests, a pre-built state machine can also be obtained; and the processing status here is used to indicate that the corresponding request, that is, the data query request, is being processed or responded to, and the processing status is also the response status.
[0134] The process of determining the target table page for which data query is required from at least one table page included in the table document may include controlling the state machine to switch from an initial state to a loading state, and when the state machine is in the loading state, determining the target table page for which data query is required from at least one table page included in the table document;
[0135] It should be noted that as the user request begins, the associated table page will be loaded into the memory for quick access. In this loading state, the terminal can also display a loading indicator to inform the user that the process is in progress. If an error occurs and the required file or resource cannot be read, the state machine will record the problem and notify the user through the feedback channel.
[0136] The process of determining the query range for the target table page may be to control the state machine to switch from the loading state to the standby state, and when the state machine is in the standby state, determine the query range for the target table page;
[0137] It should be noted that after successfully loading the necessary data, the state machine will transition to an intermediate stage, namely the preparatory state, indicating that the preparation work has been completed, and then enter the actual search process.
[0138] In at least one target data slice, the process of searching for target data matching the key character may be to control the state machine to switch from the preparation state to the search state, and when the state machine is in the search state, in at least one target data slice, search for target data matching the key character, and switch the state machine from the search state to the termination state.
[0139] It should be noted that the search state is one of the most active periods of the state machine. In the search state, the system will conduct a comprehensive and thorough search for a specific search range, that is, the target data segment. If a qualified item, that is, target data that matches the key characters, is found at a certain moment, the location information of the target data will be recorded for later use.
[0140] At the end of each search cycle or when certain important events are detected (such as the need to resolve same-name conflicts, the appearance of user intervention prompts, etc.), it is necessary to determine how to proceed next. For example, if a valid match cannot be found in a target data shard for multiple consecutive times, then a search must be performed in other target data shards.
[0141] After searching for target data matching the key character in at least one target data shard, if the state machine is in a terminated state, it is possible to determine to end the response to the data query request and respond to the request after the data query request in the request queue.
[0142] It should be noted that after the operation request response is completely completed, the state machine will enter a silent mode, that is, the termination state, and then wait for the signal to start the next cycle to arrive and then re-enter a new round of work cycle. During this period, necessary self-inspection and maintenance work can also be performed to ensure that the stable and reliable operation of the system can be maintained until it is reactivated next time.
[0143] In this way, the state machine can manage the data query process in a structured manner, which not only makes the data query process clearer and more controllable, but also ensures that the data query is executed according to the predetermined process by defining the different states of the state machine and the conversion rules between states, thereby improving the reliability, flexibility and maintainability of the data query process.
[0144] In some embodiments, if the table document is an online document for collaborative editing by multiple objects, in order to ensure the integrity and consistency of the data, during the character replacement operation, the detection and response to the external collaborative data may be temporarily stopped. Specifically, before obtaining the key character, when an operation request carrying the key character is received for the table document, the process or service for detecting and responding to the external collaborative data is closed; thereby re-analyzing the operation request to obtain the key character; in this way, conflicts and inconsistencies during multi-party editing can be effectively avoided;
[0145] Then, after replacing the target data matching the key characters obtained from the query, the process or service for detecting and responding to external collaborative data is restarted, that is, the comprehensive support and service for collaborative editing functions within the table are reactivated. This not only ensures the security of the data, but also maintains the smooth communication and information sharing experience of the entire collaborative environment.
[0146] In some embodiments, during a character replacement operation, in order to allow the user to grasp the replacement progress at any time and make corrections in a timely manner, the terminal may also display a preview window; in the preview window, a first style is used to display the target data to be replaced, and a second style is used to display other characters that replace the target data; wherein the first style is used to indicate that the target data is deleted, and the second style is used to indicate that the other characters are new data that replace the deleted data, and the first style is different from the second style.
[0147] In this way, if an error is found, the user can stop the replacement process at any time and correct the cause of the error, such as readjusting the characters used to replace the target data, or adjusting the query range.
[0148] By applying the above-mentioned embodiment of the present application, after determining the key characters for data query in the table document and the target table page for data query in the table document, multiple data slices obtained by slicing the data included in the target table page are obtained, and then at least one target data slice corresponding to the query range is selected from the multiple data slices, so that for each target data slice, a parallel data query method is adopted, and the target data matching the key characters is queried in at least one target data slice. In this way, by slicing the data of the table document, and then using a parallel processing method for data query for each data slice, the data query speed is accelerated and the data query efficiency is improved compared to the solution of scanning the entire table document as a whole to perform data query.
[0149] The following is an explanation of an exemplary application of the embodiments of the present application in a practical application scenario.
[0150] When performing data query on a table document, related technologies mostly scan the entire table document as a whole to find the queried data. However, if there is a lot of data in the table document, the data query speed will be slow, that is, the data query efficiency will be low.
[0151] Based on this, the present application provides a data query method, which performs data slicing on the table document, and then uses a parallel processing method to perform data query on each data slice. Compared with the solution of performing data query by scanning the entire table document as a whole, this method speeds up the data query speed and improves the data query efficiency.
[0152] Next, the technical solution of this application is explained from the product side.
[0153] In some embodiments, a search and replace operation may be performed on mathematical formulas, functions, regular expressions, etc. in a document, such as Fig. 9 As shown, Fig. 9This is the first product interface diagram of the data query process provided by the embodiment of the present application, based on Fig. 9 In response to the input operation for the query content (keywords), the input query content as indicated by the dotted box 901, that is, 14, is displayed; then in response to the selection operation for the "Search only in formula" option in the dotted box 902, the "Search only in formula" option in the dotted box 902 is controlled to be in a selected state; then in response to the trigger operation of the search control indicated by 903, 14 is queried in the formula of the table document.
[0154] In some embodiments, the search and replace range can be specified to achieve accurate search and replace operations in a local area, such as Fig.10 As shown, Fig.10 This is a second product interface diagram of the data query process provided by the embodiment of the present application, based on Fig.10 In response to the selection operation on the query range in the table document, the query range as indicated by 1001 is displayed; then in response to the input operation on the query content (keywords), the input query content as indicated by the dotted box 1002, that is, 115, is displayed; and then in response to the trigger operation on the search control indicated by 1003, 115 is queried in the query range indicated by 1001.
[0155] In some embodiments, search and replace operations can also be performed between different worksheets, thereby breaking the isolation restrictions between worksheets. Fig.11 As shown, Fig.11 This is a third product interface diagram of the data query process provided by the embodiment of the present application, based on Fig.11 In response to the input operation for the query content (keywords), the input query content as indicated by the dotted box 1101, that is, 115, is displayed; then in response to the selection operation for the "Search All Worksheets" option in the dotted box 1102, the "Search All Worksheets" option in the dotted box 1102 is controlled to be in a selected state; and then in response to the triggering operation of the search control indicated by 1103, 115 is queried in the two worksheets indicated by the dotted box 1104.
[0156] Next, the technical solution of the present application is explained from a technical perspective.
[0157] For the search process for table documents, specifically, first, after receiving the search behavior triggered by the user, the system will create a task queue (request queue). This queue is a first-in-first-out data structure used to store and manage all search tasks to be executed. Each task (request) has a unique subtable (table page) ID as an identifier, which facilitates the detection and management of tasks. Then, the subtable (i.e., table page or worksheet) of the table document is loaded. Specifically, when a subtable needs to be searched, the system first checks the loading status of the subtable. If the subtable has been loaded, the system can directly search. If the subtable has not been loaded, the system will pull the subtable data; then, the loaded subtable is searched in slices. Specifically, in order to improve the search performance, the system will divide the subtable data into multiple slices (data slices) for search. Each slice contains a certain number of rows or columns. The system will search each slice one by one and store the search results; then, the search results are rendered. Specifically, after each slice search is completed, the system will render the search results to the editor, so that users can see the search results in real time and improve the user experience.
[0158] It should be noted that during the entire search process, the system will continuously update the task status. For example, when a shard (data shard) search is completed, the task status will change from loading to searching, and when all shard searches are completed, the task status will change from searching to completed.
[0159] In actual implementation, the task status can be managed through the state machine, so that the execution of the task at each stage can be accurately controlled and detected through the state machine, reflecting each operation situation and the system response mode; among them, the state machine is managed in an event-driven manner. When an event occurs, the state machine will switch according to the event type and the current state. For example, when the subtable is loaded, the state machine will switch from the initial state to the loading state, and when the search starts, the state machine will switch from the loading state to the search state, including the start state, loading state, preparation state, search state, replacement state and termination state.
[0160] For the startup state, whenever a new search task is generated, the state machine will automatically be in the unstarted state, that is, the startup state, and all resources and variables are initialized, ready to accept the user's first command or operation.
[0161] For the loading state, when a user operation request for a table document, such as a query request or a replace request, is received, the associated subtable will be loaded into memory for quick access. In this state, the system will display a loading indicator to inform the user that the process is in progress. If an error occurs and the required file or resource cannot be read, the state machine will record the problem and notify the user through the feedback channel.
[0162] For the preparation state, after successfully loading the necessary data, the state machine will transition to an intermediate stage to indicate that the preparation work has been completed, and then enter the actual search process, that is, the search state.
[0163] For the search state, this is one of the most active periods of the state machine. In this state, the system will conduct a comprehensive and thorough search of the predetermined search domain (target data shard). If a qualifying item (target data) is found at a certain moment, the location information and characteristics of the item will be recorded for later use. Among them, multiple searches can also be performed in parallel to improve the overall processing efficiency.
[0164] It should be noted that at the end of each search cycle or when certain important events are detected (such as the need to resolve same-name conflicts, the appearance of user intervention prompts, etc.), the state machine will evaluate to decide what to do next. For example, if a valid match is not found in the predetermined search domain for multiple consecutive times, then a search will be performed in other search domains.
[0165] For the replacement state, in this state, if the data that meets the conditions is found, in addition to the regular text revision, other types of correction measures will also be executed, such as adjusting the background color of the cell where the numerical parameter is located to show different treatment, etc.
[0166] For the termination state, after the task is completely completed, the state machine will enter a silent mode, waiting for the signal to start the next cycle to arrive and then re-enter a new round of work cycle. During this period, necessary self-inspection and maintenance work can also be performed to ensure that the stable and reliable operation of the system can be maintained until it is reactivated next time.
[0167] In actual implementation, after finding data that meets the conditions, you can also replace the existing search results, that is, execute the replacement process; the replacement process is to replace the editable content of the cell based on the existing search results. Once the user confirms the search results and decides to perform the replacement operation, the system will enter this process. Specifically, first select the replacement range. The system defaults to replacing the cells in the entire searched range, but also provides users with options to limit a smaller replacement area. If the selected replacement area is received, the system will replace the search results in the replacement area; then back up the original data. Specifically, before performing any form of replacement, in order to prevent data loss due to misoperation, the system will automatically save the original data and format settings in the selected range to a temporary storage area, so that the original data can be restored after the replacement; then, perform the replacement action. Specifically, first, in order to ensure the integrity and consistency of the data, before performing the replacement operation During this period, the system will temporarily stop detecting and responding to external collaborative data, which can effectively avoid conflicts and inconsistencies during multi-party editing. Then, based on the replacement rules and specific content determined by the user, the system compares each item in the search result list one by one, and implements a single correction of the corresponding text or format. In addition, batch replacement operations are also supported in this link, which greatly improves work efficiency. Then, in order to allow users to grasp the replacement progress at any time and make corrections in time, the system provides a real-time preview window for viewing. If any errors are found, the user can stop at any time and make necessary adjustments before continuing to the next step. Finally, once the replacement operation is successfully completed and no problems are found, the system will restore the previously backed up original data after confirmation, and at the same time reactivate the comprehensive support and services for the collaborative editing function within this range. This not only ensures the security of the data but also maintains the smooth communication and information sharing experience of the entire team collaboration environment.
[0168] In this way, through the data query method provided by this application, in the data query process for table documents, not only the search speed is fast, the replacement accuracy is high, it is safe and reliable, and it is easy to expand and maintain, but it can also meet the user's usage needs in various scenarios and improve user experience and work efficiency.
[0169] By applying the above-mentioned embodiment of the present application, after determining the key characters for data query in the table document and the target table page for data query in the table document, multiple data slices obtained by slicing the data included in the target table page are obtained, and then at least one target data slice corresponding to the query range is selected from the multiple data slices, so that for each target data slice, a parallel data query method is adopted, and the target data matching the key characters is queried in at least one target data slice. In this way, by slicing the data of the table document, and then using a parallel processing method for data query for each data slice, the data query speed is accelerated and the data query efficiency is improved compared to the solution of scanning the entire table document as a whole to perform data query.
[0170] The following is a description of an exemplary structure of the data query device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, Figure 2 As shown, the software modules stored in the data query device 455 of the memory 450 may include:
[0171] A first acquisition module 4551 is used to acquire key characters, where the key characters are used to perform data query in a table document;
[0172] A first determining module 4552 is used to determine a target table page that needs to be queried for data from at least one table page included in the table document;
[0173] A second acquisition module 4553 is used to acquire a plurality of data slices corresponding to the target table page, wherein the plurality of data slices are obtained by performing slice processing on the data included in the target table page;
[0174] A second determination module 4554 is used to determine a query range for the target table page, and select at least one target data slice corresponding to the query range from the multiple data slices;
[0175] The query module 4555 is used to query the target data matching the key character in the at least one target data slice by adopting a parallel data query method for each target data slice.
[0176] In some embodiments, the second acquisition module 4553 is also used to acquire the row data of each row in the target table page, and based on the row data, divide the data of the target table page to obtain the multiple data slices, each of the data slices includes at least one row of row data, and the row data included in different data slices are different; or, acquire the column data of each column in the target table page, and based on the column data, divide the data of the target table page to obtain the multiple data slices, each of the data slices includes at least one column of column data, and the column data included in different data slices are different.
[0177] In some embodiments, the device also includes a secondary sharding module, which is used to perform the following processing on each of the data shards to obtain multiple data sub-shards corresponding to each of the data shards: when the data shard includes multiple data types, the data shard is sharded based on the multiple data types to obtain multiple data sub-shards, and different data sub-shards correspond to different data types; the query module 4555 is also used to perform the following processing on each of the target data shards: determine the data type of the key character, and select a target data sub-shard whose data type is consistent with the data type of the key character from the multiple data sub-shards corresponding to the target data shard; in the target data sub-shard, query the target data that matches the key character.
[0178] In some embodiments, the device further includes an update module, wherein the update module is configured to obtain updated new data shards in response to updates to the data shards; and based on the new data shards, update the data sub-shards corresponding to the new data shards.
[0179] In some embodiments, the device also includes a detection module, which is used to periodically detect the data volume of each of the data slices to obtain a detection result; when it is determined based on the detection result that there is a first data slice among the multiple data slices, a second data slice is selected from the multiple data slices, and the first data slice is merged with the second data slice; when it is determined based on the detection result that there is a third data slice among the multiple data slices, the third data slice is sliced; wherein the data volume of the first data slice is less than a first threshold, the data volume of the third data slice is greater than a second threshold, and the second threshold is greater than the first threshold.
[0180] In some embodiments, the detection module is also used to select at least one of the following from the multiple data slices as the second data slice: a data slice whose data volume is less than the first threshold; a data slice whose data volume is not less than the first threshold and whose data volume is less than a third threshold, wherein the third threshold is less than the second threshold; a data slice with the smallest data volume other than the first data slice.
[0181] In some embodiments, the first determination module 4552 is also used to receive an operation request, the operation request including at least one of a data query request and a data replacement request; parse the operation request to obtain a table page identifier carried by the operation request, the table page identifier being used to indicate the table page searched based on the keyword; and select a table page corresponding to the table page identifier from at least one table page included in the table document as the target table page.
[0182] In some embodiments, the second determination module 4554 is also used to, when the operation request also carries a query range for the target table page, use the query range carried in the operation request as the query range for the target table page; when the operation request does not carry a query range for the target table page, use multiple data shards corresponding to the target table page as the query range for the target table page.
[0183] In some embodiments, the table document is an online document for collaborative editing by multiple objects, and the key character is determined by the first object among the multiple objects; the device also includes a third acquisition module, and the third acquisition module is used to obtain each updated target data slice when the second object among the multiple objects performs a data update operation on the target data slice; the query module 4555 is also used to use a parallel data query method for each updated target data slice to query the target data matching the key character in the updated target data slice.
[0184] In some embodiments, the third acquisition module is also used to obtain the priority of the second object and the priority of the first object in response to the data update operation; in response to the priority of the second object being higher than the priority of the first object, each of the target data shards is updated based on the data update operation to obtain updated target data shards.
[0185] In some embodiments, the key character is determined by a received data query request; the device also includes a replacement module, which is used to record the first position of the target data in the table document in response to querying the target data matching the key character; when a data replacement request for the key character in the query range is received, the data replacement request is parsed to obtain other characters carried by the data replacement request for replacing the key character; the first position of the recorded target data is obtained; and based on the first position, the target data is replaced with the other characters.
[0186] In some embodiments, the replacement module is also used to obtain other data in the target table page except the target data, determine the second position of the other data in the table document, and cache the other data and the second position; in response to receiving a data update request for the other data, when the time interval between the reception time of the data update request and the reception time of the data replacement request is less than an interval threshold, obtain the cached other data and the second position, and replace the current data at the second position in the query range with the other data.
[0187] In some embodiments, the key character is carried in a data query request, and the data query request is stored in a request queue, and the request queue includes multiple requests; the device also includes a state management module, and the state management module is used to obtain a state machine corresponding to the data query request and control the state machine to be in an initial state; wherein the state machine is used to indicate a processing state for the data query request; the first determination module 4552 is also used to control the state machine to switch from an initial state to a loading state, and when the state machine is in the loading state, determine the target table page that needs to be queried for data from at least one table page included in the table document; the second determination module 4554 is also used to Control the state machine to switch from the loading state to the preparation state, and when the state machine is in the preparation state, determine the query range for the target table page; the query module 4555 is used to control the state machine to switch from the preparation state to the search state, and when the state machine is in the search state, query the target data matching the key character in at least one target data slice, and switch the state machine from the search state to the termination state; the device also includes a response module, and the response module is used to determine the end of the response to the data query request if the state machine is in the termination state, and respond to the request in the request queue that is after the data query request.
[0188] The embodiment of the present application provides a computer program product, which includes computer executable instructions or computer programs, and the computer executable instructions or computer programs are stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instructions or computer programs from the computer-readable storage medium, and the processor executes the computer executable instructions or computer programs, so that the electronic device executes the data query method described above in the embodiment of the present application.
[0189] The present application embodiment provides a computer-readable storage medium storing computer-executable instructions or computer programs. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the data query method provided by the present application embodiment, for example, Figure 3 The data query method is shown.
[0190] In some embodiments, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disk, or a CD-ROM; or it may be various devices including one or any combination of the above memories.
[0191] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
[0192] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0193] As an example, computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.
[0194] In summary, the embodiments of the present application have the following beneficial effects:
[0195] (1) By processing the data of the table document in pieces, and then using a parallel processing method to perform data query on each data piece, compared with the solution of scanning the entire table document as a whole to perform data query, the data query speed is accelerated and the data query efficiency is improved.
[0196] (2) Data in table documents is sharded by rows or columns. After data sharding, each fragment, i.e., data shard, can be processed independently. This not only improves the speed of data processing, but also solves the problem that for large table data, loading all data at once may consume a lot of memory. Sharding processing can also only load the data fragments required for processing, reducing memory usage. At the same time, for specific data query tasks, operations only need to be performed on the relevant data shards, which simplifies the query logic and improves query efficiency.
[0197] (3) The data shards are re-sharded based on the data type, and then the target data sub-shard whose data type is consistent with the data type of the key character is selected from the multiple data sub-shards corresponding to the target data shard, so as to search for the target data matching the key character in the target data sub-shard; in this way, the data shards to be queried are preliminarily screened based on the data type, and data queries are only performed in the data shards related to the data query task, that is, the query range is dynamically adjusted based on the query content, which not only reduces the number of data shards that need to be queried and improves the data processing speed, but also simplifies the query logic and improves the query efficiency.
[0198] (4) When the data volume is smaller than the first threshold, the corresponding data shards are merged, thereby reducing the total number of data shards. This not only reduces the maintenance cost of the data shards and simplifies the management of the data shards, but also reduces the occupancy of storage resources because the space waste between shards can be eliminated. At the same time, by merging the data shards, the number of shards that need to be accessed when performing data queries is reduced, thereby improving query performance.
[0199] (5) When the data volume is larger than the second threshold, the corresponding data shards are split. This can reduce the load of a single data shard and avoid performance bottlenecks caused by a single data shard being too large, thereby improving the query effect and ensuring the stability of the data query process. At the same time, as the amount of data grows, splitting the data shards with large data volumes can provide better scalability, thereby adapting to the increase in data volume and ensuring that efficient data queries can be achieved even when the amount of data increases.
[0200] (6) In order to prevent data loss due to misoperation, the original data in the selected range is automatically saved to a temporary storage area, so that the original data can be restored after the replacement is completed, thus ensuring the accuracy of the table data.
[0201] (7) The data query method provided by this application not only has fast search speed, high replacement accuracy, security and reliability, and is easy to expand and maintain during the data query process for table documents, but can also meet the user's usage needs in various scenarios and improve user experience and work efficiency.
[0202] It should be noted that in the embodiments of the present application, it involves obtaining data related to key characters, table page identifiers, and data slices. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data must comply with relevant laws, regulations, and standards of relevant countries and regions.
[0203] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A data query method, characterized in that: The method comprises: Obtaining key characters, wherein the key characters are used to perform data query in a table document; Determining a target table page for which data query is required from at least one table page included in the table document; Acquire a plurality of data slices corresponding to the target table page, wherein the plurality of data slices are obtained by performing slice processing on the data included in the target table page; Determine a query range for the target table page, and select at least one target data shard corresponding to the query range from the multiple data shards; For each of the target data slices, a parallel data query method is adopted to query the target data matching the key character in the at least one target data slice.
2. The method according to claim 1, characterized in that The obtaining of multiple data slices corresponding to the target table page includes: Acquire row data of each row in the target table page, and divide the data of the target table page based on the row data to obtain the multiple data slices, each of the data slices includes at least one row of row data, and different data slices include different row data; or, The column data of each column in the target table page is obtained, and based on the column data, the data of the target table page is divided to obtain the multiple data slices, each of which includes at least one column of column data, and different data slices include different column data.
3. The method according to claim 2, characterized in that After dividing the data of the target table page to obtain the multiple data fragments, the method further includes: For each of the data shards, perform the following processing to obtain multiple data sub-shards corresponding to each of the data shards: When the data shard includes multiple data types, the data shard is sharded based on the multiple data types to obtain multiple data sub-shards, and different data sub-shards correspond to different data types; The step of searching for target data matching the key character in the at least one target data slice includes: For each of the target data slices, perform the following processing: Determine the data type of the key character, and select a target data sub-slice whose data type is consistent with the data type of the key character from the multiple data sub-slices corresponding to the target data slice; In the target data sub-shard, the target data matching the key character is searched.
4. The method according to claim 3, characterized in that After the data sharding is performed based on the multiple data types to obtain multiple data sub-shards, the method further includes: In response to the data shard being updated, obtaining an updated new data shard; Based on the new data shards, the data sub-shards corresponding to the new data shards are updated.
5. The method according to claim 2, characterized in that: After dividing the data of the target table page to obtain the multiple data fragments, the method further includes: Periodically detecting the data volume of each of the data slices to obtain a detection result; When it is determined based on the detection result that a first data slice exists among the multiple data slices, selecting a second data slice from the multiple data slices, and merging the first data slice with the second data slice; When it is determined based on the detection result that a third data fragment exists among the multiple data fragments, fragmenting the third data fragment; The data volume of the first data slice is smaller than a first threshold, the data volume of the third data slice is larger than a second threshold, and the second threshold is larger than the first threshold.
6. The method according to claim 5, characterized in that The selecting a second data slice from the multiple data slices includes: From the multiple data shards, select at least one of the following as the second data shard: Data shards whose data volume is less than the first threshold; Data shards whose data volume is not less than the first threshold and whose data volume is less than a third threshold, wherein the third threshold is less than the second threshold; A data slice with the smallest data amount except the first data slice.
7. The method according to claim 1, characterized in that The step of determining a target table page for which data query is required from at least one table page included in the table document comprises: receiving an operation request, wherein the operation request includes at least one of a data query request and a data replacement request; Parsing the operation request to obtain a table page identifier carried in the operation request, wherein the table page identifier is used to indicate a table page to be searched based on the key character; From at least one table page included in the table document, a table page corresponding to the table page identifier is selected as the target table page.
8. The method according to claim 7, characterized in that The determining of the query scope for the target table page includes: When the operation request also carries a query range for the target table page, the query range carried in the operation request is used as the query range for the target table page; When the operation request does not carry a query range for the target table page, a plurality of data slices corresponding to the target table page are used as the query range for the target table page.
9. The method according to claim 1, characterized in that: The table document is an online document for collaborative editing by multiple objects, and the key character is determined by the first object among the multiple objects; After selecting at least one target data shard corresponding to the query range from the multiple data shards, the method further includes: When a second object among the multiple objects performs a data update operation on the target data slice, obtaining each updated target data slice; The method of using a parallel data query method for each target data slice to query target data matching the key character in at least one target data slice includes: For each of the updated target data slices, a parallel data query method is adopted to query the target data matching the key character in the updated target data slices.
10. The method according to claim 9, characterized in that The obtaining of each updated target data slice comprises: In response to the data update operation, obtaining the priority of the second object and the priority of the first object; In response to the priority of the second object being higher than the priority of the first object, each of the target data slices is updated based on the data update operation to obtain an updated target data slice.
11. The method according to claim 1, characterized in that: The key character is determined by a received data query request; after searching for target data matching the key character in the at least one target data slice, the method further comprises: In response to finding the target data matching the key character, recording the first position of the target data in the table document; When receiving a data replacement request for the key character in the query range, parsing the data replacement request to obtain other characters carried in the data replacement request and used to replace the key character; Acquire the first position of the recorded target data; Based on the first position, the target data is replaced with the other characters.
12. The method according to claim 11, characterized in that Before replacing the target data with the other characters based on the first position, the method further includes: Acquire other data in the target table page except the target data, determine a second position of the other data in the table document, and cache the other data and the second position; After replacing the target data with the other characters based on the first position, the method further comprises: In response to receiving a data update request for the other data, when the time interval between the receipt time of the data update request and the receipt time of the data replacement request is less than an interval threshold, the cached other data and the second position are obtained, and the current data at the second position in the query range is replaced with the other data.
13. The method according to claim 1, characterized in that The key character is carried in a data query request, and the data query request is stored in a request queue, wherein the request queue includes a plurality of requests; After obtaining the key characters, the method further includes: Obtaining a state machine corresponding to the data query request, and controlling the state machine to be in an initial state; wherein the state machine is used to indicate a processing status for the data query request; The step of determining a target table page for which data query is required from at least one table page included in the table document comprises: Controlling the state machine to switch from an initial state to a loading state, and when the state machine is in the loading state, determining a target table page that needs to be queried for data from at least one table page included in the table document; The determining of the query scope for the target table page includes: Controlling the state machine to switch from the loading state to the standby state, and when the state machine is in the standby state, determining a query range for the target table page; The step of searching for target data matching the key character in the at least one target data slice includes: Controlling the state machine to switch from the preparation state to the search state, and when the state machine is in the search state, searching for target data matching the key character in the at least one target data slice, and switching the state machine from the search state to the termination state; The method further comprises: If the state machine is in the termination state, it is determined to end the response to the data query request, and respond to the request in the request queue that is subsequent to the data query request.
14. A data query device, characterized in that: The device comprises: A first acquisition module, used to acquire key characters, wherein the key characters are used to perform data query in a table document; A first determining module, used to determine a target table page that needs to be queried for data from at least one table page included in the table document; A second acquisition module is used to acquire a plurality of data slices corresponding to the target table page, wherein the plurality of data slices are obtained by performing slice processing on the data included in the target table page; A second determination module is used to determine a query range for the target table page, and select at least one target data slice corresponding to the query range from the multiple data slices; The query module is used to query the target data matching the key character in the at least one target data slice by adopting a parallel data query method for each target data slice.
15. An electronic device, characterized in that: include: Memory for storing computer executable instructions or computer programs; A processor, used to implement the data query method according to any one of claims 1 to 13 when executing the computer executable instructions or computer programs stored in the memory.
16. A computer-readable storage medium, characterized in that: Computer executable instructions or computer programs are stored, which are used to cause a processor to execute and implement the data query method described in any one of claims 1 to 13.
17. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer programs are executed by a processor, the data query method according to any one of claims 1 to 13 is implemented.