NoSQL-based trajectory data deduplication method and system

Based on the NoSQL database and redis database, only one of the latest valid data is recorded for multiple operation requests of users in a single business function, which solves the problem of MySQL's read and write backlog when the number of users is large, and the system performance is improved.

CN120030004APending Publication Date: 2025-05-23当趣网络科技(杭州)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411970834.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When the number of users is large, the MySQL database has long-term high reading and writing times, resulting in data backlog and disk pressure, and the query and write pressure gradually increases. The existing technology has not yet proposed an effective solution.

Method used

The trajectory data de-repeat method based on NoSQL is adopted. By obtaining multiple write operation requests by the user on a single business function, the trajectory data is written to the redis database, the trajectory data is merged by requesting the merge plug-in, and the latest write record is written to the NoSQL database to avoid redundant data accumulation.

Benefits of technology

It effectively avoids the read and write backlog caused by MySQL due to the accumulation of invalid redundant data, reduces the read and write pressure of the database, and improves system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030004A_ABST
    Figure CN120030004A_ABST
Patent Text Reader

Abstract

The invention relates to a NoSQL (Not Only Structured Query Language)-based trajectory data deduplication method and system, and the method comprises the steps: obtaining a plurality of write operation requests of a user on a single service function; writing the corresponding track data into a redis database based on the plurality of write operation requests; merging the plurality of write operation requests into a target operation request through a request merging plug-in; and based on the target operation request, reading the newest write-in record of the redis database, and writing the newest track data of the user on the business function into the NoSQL database based on the newest write-in record. By means of the method and device, on the basis of the NoSQL database and the redis database, only one piece of latest valid data is recorded for multiple operation requests of a single user on a single service function, the situation that invalid redundant data is accumulated due to mysql is effectively avoided, and the problem of how to reduce read-write overstock of the database is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of database technology, and in particular to a NoSQL-based trajectory data deduplication method, system, device and medium. Background Art

[0002] In the current business field, user behavior record data such as user playback records and operation logs are an important basis for analyzing user interaction plans. The current common strategy in the industry is to use the MySQL database to record user behavior and perform table operations for users in the business database. This strategy is often not a problem when the number of users is relatively small. However, when the number of users is relatively large, for example, when the platform has more than 100,000 daily active users, it often receives tens of millions of data write operations every day, which will cause the MySQL read and write times to be high for a long time. In the long run, the database data backlog and disk pressure will become very large, and the pressure of data query and write will gradually increase.

[0003] Currently, no effective solution has been proposed for the problem of how to reduce the read and write backlog of the database in related technologies. Summary of the invention

[0004] The embodiments of the present application provide a NoSQL-based trajectory data deduplication method, system, device, and medium to at least solve the problem of how to reduce the read and write backlog of the database in the related art.

[0005] In a first aspect, an embodiment of the present application provides a NoSQL-based trajectory data deduplication method, the method comprising:

[0006] Get multiple write operation requests from users on a single business function;

[0007] Based on the multiple write operation requests, writing the corresponding trajectory data into the redis database;

[0008] Merging the multiple write operation requests into one target operation request through a request merging plug-in;

[0009] Based on the target operation request, the latest write record of the redis database is read, and based on the latest write record, the latest track data of the user on the business function is written to the NoSQL database.

[0010] In some embodiments, the method comprises:

[0011] A read operation request of the business function to be queried by the user is obtained, and the track data of the business function to be queried is read from the redis database or the NoSQL database based on the read operation request.

[0012] In some embodiments, obtaining a read operation request of a user for a business function to be queried, and reading the track data of the business function to be queried from the redis database or the NoSQL database based on the read operation request includes:

[0013] Obtaining a read operation request of the user's service function to be queried, and detecting whether the read operation request contains a service timestamp;

[0014] If not, reading the track data of the to-be-queried business function from the redis database based on the read operation request;

[0015] If so, the trajectory data of the to-be-queried business function is read from the NoSQL database based on the read operation request.

[0016] In some embodiments, if not, reading the track data of the to-be-queried business function from the redis database based on the read operation request; if yes, reading the track data of the to-be-queried business function from the NoSQL database based on the read operation request includes:

[0017] If not, based on the read operation request, generate a row key of the to-be-queried business function of the user through a preset naming rule, and query and read the corresponding track data from the redis database based on the row key;

[0018] If so, based on the read operation request, a row key of the to-be-queried business function of the user is generated by a preset naming rule, and corresponding track data is queried and read from the NoSQL database based on the row key.

[0019] In some embodiments, based on the read operation request, generating the row key of the to-be-queried business function of the user by a preset naming rule includes:

[0020] Generate the row key of the business function to be queried by the user through the preset naming rule, wherein the value of the generated row key is equal to [md5 (business name + user id)] 取前2位 +"_"+user id+"_"+"business id"+"_"+function id.

[0021] In some embodiments, based on the target operation request, reading the latest write record of the redis database includes:

[0022] Distribute the target operation request through the Kafka stream processing platform to obtain a current operation message after the message is distributed;

[0023] Based on the current operation message, the latest write record of the trajectory data corresponding to the multiple write operation requests is obtained from the redis database.

[0024] In some of the embodiments, the NoSQL database is a column-based HBase database, and the request merging plug-in is a collapse-executor plug-in.

[0025] In a second aspect, an embodiment of the present application provides a NoSQL-based trajectory data deduplication system, the system is used to execute the method described in the first aspect above, the system includes an acquisition module, a cache module, a merging module and a storage module;

[0026] The acquisition module is used to acquire multiple write operation requests of a user on a single business function;

[0027] The cache module is used to write the corresponding trajectory data into the redis database according to the multiple write operation requests;

[0028] The merging module is used to merge the multiple write operation requests into one target operation request through a request merging plug-in;

[0029] The storage module is used to read the latest write record of the redis database according to the target operation request, and write the latest track data of the user on the business function to the NoSQL database based on the latest write record.

[0030] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the computer program.

[0031] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect above.

[0032] Compared with the related art, the embodiments of the present application provide a NoSQL-based trajectory data deduplication method, system, device and medium, wherein the method obtains multiple write operation requests from a user on a single business function; based on the multiple write operation requests, writes the corresponding trajectory data into a redis database; merges the multiple write operation requests into a target operation request through a request merge plug-in; based on the target operation request, reads the latest write record from the redis database, and writes the latest trajectory data of the user on the business function to the NoSQL database based on the latest write record, thereby realizing that on the basis of the NoSQL database and the redis database, for multiple operation requests of a single user on a single business function, only one latest valid data is recorded, effectively avoiding the situation where MySQL will cause the accumulation of invalid redundant data, and solving the problem of how to reduce the read and write backlog of the database. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0034] Figure 1 is a flowchart of the steps of a NoSQL-based trajectory data deduplication method according to an embodiment of the present application;

[0035] Figure 2 is a timing diagram of a NoSQL-based trajectory data deduplication method according to an embodiment of the present application;

[0036] Figure 3 It is a flowchart of a NoSQL-based trajectory data deduplication method according to an embodiment of the present application;

[0037] Figure 4 It is a schematic diagram of the internal structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0039] Obviously, the drawings described below are only some examples or embodiments of the present application. For ordinary technicians in this field, the present application can also be applied to other similar scenarios based on these drawings without creative work. In addition, it can also be understood that although the efforts made in this development process may be complicated and lengthy, for ordinary technicians in this field related to the content disclosed in this application, some changes in design, manufacturing or production based on the technical content disclosed in this application are just conventional technical means, and should not be understood as insufficient content disclosed in this application.

[0040] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0041] Unless otherwise defined, the technical terms or scientific terms involved in this application should be understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the" and the like involved in this application do not indicate a quantitative limitation, and may represent the singular or plural. The terms "include", "comprise", "have" and any of their variations involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0042] The present application embodiment provides a NoSQL-based trajectory data deduplication method. Figure 1 is a flowchart of the steps of the NoSQL-based trajectory data deduplication method according to an embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps:

[0043] Step S102, obtaining multiple write operation requests of a user on a single business function;

[0044] Step S104, based on multiple write operation requests, write the corresponding trajectory data into the redis database;

[0045] Step S106, merging multiple write operation requests into one target operation request through a request merging plug-in;

[0046] Step S108, based on the target operation request, read the latest write record of the redis database, and write the latest track data of the user in the business function into the NoSQL database based on the latest write record.

[0047] Specifically, step S108 distributes the target operation request through the kafka stream processing platform to obtain the current operation message after the message distribution; based on the current operation message, the latest write record of the trajectory data corresponding to the multiple write operation requests is obtained from the redis database.

[0048] Through the above steps in the embodiment of the present application, it is achieved that on the basis of the NoSQL database and the redis database, only the latest valid data is recorded for multiple operation requests of a single user on a single business function, effectively avoiding the situation where MySQL will cause the accumulation of invalid redundant data, and solving the problem of how to reduce the read and write backlog of the database.

[0049] The specific embodiment of the present application provides a NoSQL-based trajectory data deduplication method. For the overall architecture of the method, redis is used as a read buffer layer. At the same time, the written data will be directly merged through a request merge plug-in. The request merge plug-in is preferably a collapse-executor plug-in; the merge request is written to a NoSQL database after the message is distributed. The NoSQL database is preferably a column-based HBase database, and HBase will not write back to redis.

[0050] Specifically, Figure 2 is a timing diagram of a NoSQL-based trajectory data deduplication method according to an embodiment of the present application, such as Figure 2 As shown, in the front-end service, a BFF layer is established to receive data and play the role of intermediate forwarding and anti-corrosion. This layer is used to filter some data and merge requests to forward to the next layer. In other words, the collapse-executor plug-in is used to merge multiple requests into one request and forward it to the HBase service for batch writing. Kafka is used to distribute messages to reduce the write peak.

[0051] At the database level, based on the characteristics of HBase, only one valid data is recorded for a single business function of a single user. Taking the playback record as an example, for the short video function of business A, only one record of a short video id is recorded. Multiple records use the version overwriting feature of the HBase database to allow the new version to overwrite the old version. This can avoid the situation where the use of MySQL will cause data accumulation or slow data update after the data volume increases.

[0052] It should be noted that for the collapse-executor plug-in mentioned above, Collapse Executor is a high-performance, low-latency batch merge executor designed to support high-concurrency hotspot requests. Collapse Executor can merge multiple similar requests into one to reduce thread consumption, effectively improve service resource utilization, and reduce service response time, thereby improving overall system performance. Furthermore, it supports integration with Spring Boot to provide a convenient microservice building solution. Furthermore, it supports mainstream technology stacks such as CompletableFuture, WebClient, and Servlet Async.

[0053] The present application embodiment provides a NoSQL-based trajectory data deduplication method. Figure 3 is a flow chart of a NoSQL-based trajectory data deduplication method according to an embodiment of the present application, such as Figure 3 As shown, in addition to the write operation process of step S102 to step S108 in the above embodiment, the method also includes:

[0054] Obtain the user's read operation request for the business function to be queried, and read the track data of the business function to be queried from the redis database or NoSQL database based on the read operation request.

[0055] like Figure 3 As shown, the read operation process specifically includes the following steps:

[0056] Step S31 obtains the user's read operation request for the service function to be queried, and detects whether the read operation request contains a service timestamp;

[0057] Step S32: if not, then read the track data of the business function to be queried from the redis database based on the read operation request;

[0058] Preferably, in step S32, if not, based on the read operation request, a row key of the user's business function to be queried is generated by a preset naming rule, and the corresponding track data is queried and read from the redis database based on the row key;

[0059] Step S33: If yes, then the trajectory data of the business function to be queried is read from the NoSQL database based on the read operation request.

[0060] Preferably, in step S33, if yes, based on the read operation request, a row key of the user's business function to be queried is generated by a preset naming rule, and the corresponding trajectory data is queried and read from the NoSQL database based on the row key.

[0061] It should be noted that, for step S32 and step S33, the row key of the user's to-be-queried business function is generated by a preset naming rule, wherein the value of the generated row key is equal to [md5 (business name + user id)] 取前2位 +"_"+user id+"_"+"business id"+"_"+function id. The md5 in the row key is used to eliminate hot data in a single region. Each row key records the usage track of a single user on a single business function, which can effectively increase the reading efficiency. If the business requests the usage track of a certain function of a certain user, you can directly request the row key of the user to query. However, range queries can only be performed through business timestamps. The timestamp of the last record returned by the service will become the start timestamp of the next query.

[0062] It can be seen that although the method provided in the above embodiment can solve the write bottleneck problem of the database and the problem of data hotspots, it is difficult to effectively provide conditions for filtering queries under aggregate queries (range queries can only be performed through business timestamps). All data can only be found in the memory for filtering and sorting, which will bring great pressure to the memory buffer and is not friendly to the business side. Therefore, a multi-dimensional heterogeneous index table is proposed here to perform pre-queries to reduce the reading pressure of the memory.

[0063] It should be further explained that the steps shown in the above process or the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and, although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0064] The embodiment of the present application provides a NoSQL-based trajectory data deduplication system, which includes an acquisition module, a cache module, a merging module, and a storage module;

[0065] The acquisition module is used to obtain multiple write operation requests from users on a single business function;

[0066] The cache module is used to write the corresponding trajectory data into the redis database according to multiple write operation requests;

[0067] The merge module is used to merge multiple write operation requests into one target operation request through the request merge plug-in;

[0068] The storage module is used to read the latest write record of the redis database according to the target operation request, and write the latest trajectory data of the user on the business function to the NoSQL database based on the latest write record.

[0069] Through the acquisition module, cache module, merge module and storage module in the embodiments of the present application, it is achieved that on the basis of NoSQL database and redis database, only the latest valid data is recorded for multiple operation requests of a single user on a single business function, effectively avoiding the situation where MySQL will cause the accumulation of invalid redundant data, and solving the problem of how to reduce the read and write backlog of the database.

[0070] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0071] This embodiment further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0072] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0073] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0074] In addition, in combination with the NoSQL-based trajectory data deduplication method in the above embodiment, the embodiment of the present application can provide a storage medium for implementation. The storage medium stores a computer program; when the computer program is executed by a processor, any one of the NoSQL-based trajectory data deduplication methods in the above embodiment is implemented.

[0075] In one embodiment, a computer device is provided, which may be a terminal. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for deduplicating trajectory data based on NoSQL is implemented. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball, or touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0076] In one embodiment, Figure 4 is a schematic diagram of the internal structure of an electronic device according to an embodiment of the present application, such as Figure 4 As shown, an electronic device is provided, which may be a server, and its internal structure diagram may be as shown in Figure 4 As shown. The electronic device includes a processor, a network interface, an internal memory and a non-volatile memory connected through an internal bus, wherein the non-volatile memory stores an operating system, a computer program and a database. The processor is used to provide computing and control capabilities, the network interface is used to communicate with an external terminal through a network connection, the internal memory is used to provide an environment for the operation of the operating system and the computer program, the computer program is executed by the processor to implement a NoSQL-based trajectory data deduplication method, and the database is used to store data.

[0077] Those skilled in the art will understand that Figure 4 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.

[0078] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0079] Those skilled in the art should understand that the technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0080] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A NoSQL-based trajectory data deduplication method, characterized in that: The method comprises: Get multiple write operation requests from users on a single business function; Based on the multiple write operation requests, writing the corresponding trajectory data into the redis database; Merging the multiple write operation requests into one target operation request through a request merging plug-in; Based on the target operation request, the latest write record of the redis database is read, and based on the latest write record, the latest track data of the user on the business function is written to the NoSQL database.

2. The method according to claim 1, characterized in that The method comprises: A read operation request of the business function to be queried by the user is obtained, and the track data of the business function to be queried is read from the redis database or the NoSQL database based on the read operation request.

3. The method according to claim 2, characterized in that Obtaining a read operation request of a user for a business function to be queried, and reading the track data of the business function to be queried from the redis database or the NoSQL database based on the read operation request includes: Obtaining a read operation request of the user's service function to be queried, and detecting whether the read operation request contains a service timestamp; If not, reading the track data of the to-be-queried business function from the redis database based on the read operation request; If so, the trajectory data of the to-be-queried business function is read from the NoSQL database based on the read operation request.

4. The method according to claim 3, characterized in that: If not, reading the track data of the to-be-queried business function from the redis database based on the read operation request; If so, reading the trajectory data of the to-be-queried business function from the NoSQL database based on the read operation request includes: If not, based on the read operation request, generate a row key of the to-be-queried business function of the user through a preset naming rule, and query and read the corresponding track data from the redis database based on the row key; If so, based on the read operation request, a row key of the to-be-queried business function of the user is generated by a preset naming rule, and corresponding track data is queried and read from the NoSQL database based on the row key.

5. The method according to claim 3, characterized in that: Based on the read operation request, generating a row key of the to-be-queried business function of the user by using a preset naming rule includes: Generate the row key of the business function to be queried by the user through the preset naming rule, wherein the value of the generated row key is equal to [md5 (business name + user id)] 取前2位 +"_"+user id+"_"+"business id"+"_"+function id.

6. The method according to claim 1, characterized in that Based on the target operation request, reading the latest write record of the redis database includes: Distribute the target operation request through the Kafka stream processing platform to obtain a current operation message after the message is distributed; Based on the current operation message, the latest write record of the trajectory data corresponding to the multiple write operation requests is obtained from the redis database.

7. The method according to claim 1, characterized in that The NoSQL database is a column-based HBase database, and the request merging plug-in is a collapse-executor plug-in.

8. A NoSQL-based trajectory data deduplication system, characterized in that: The system is used to execute the method according to any one of claims 1 to 7, and the system comprises an acquisition module, a cache module, a merging module and a storage module; The acquisition module is used to acquire multiple write operation requests of a user on a single business function; The cache module is used to write the corresponding trajectory data into the redis database according to the multiple write operation requests; The merging module is used to merge the multiple write operation requests into one target operation request through a request merging plug-in; The storage module is used to read the latest write record of the redis database according to the target operation request, and write the latest track data of the user on the business function to the NoSQL database based on the latest write record.

9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.