Data processing method and device, electronic equipment and computer readable medium
By introducing a cache database and message queue mechanism in Elasticsearch, the change messages of user information are merged and processed and cached in the user dimension, which solves the problem of excessive CPU caused by frequent writes of Elasticsearch, and reduces data volume and saves cluster resources, ensuring the stability of the cluster.
Patent Information
- Application Number
- CN202311498608.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-13
AI Technical Summary
In scenarios where user data is large and high in dimensions, frequent writes of Elasticsearch lead to excessive CPU, inability to implement fast writes, and put a lot of pressure on the cluster and affect stability.
When receiving the information change message of the database, the attribute change messages of each user in the same business line are first merged into the user dimension and cached in the cache database. Then, the cached data is processed regularly and stored in the search engine cluster.
It reduces the amount of data required by the search engine cluster, saves the cluster's processing resources and storage resources, ensures the stability of the cluster, and reduces the repeated consumption of messages through message idempotence level prevention.
Smart Images

Figure CN119988437A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of big data processing technology, and in particular to a data processing method, device, electronic device, and computer-readable medium. Background Art
[0002] In scenarios where the amount of user data is large and the data dimension is high, in order to realize the search function of user information, user information is usually stored in the Elasticsearch search engine. Among them, Elasticsearch is generally a Lucene-based search server, which provides a distributed multi-user full-text search engine. Elasticsearch is developed in Java and released as open source under the Apache license terms. It is a popular enterprise-level search engine.
[0003] However, the inventors found that the underlying layer of Elasticsearch is generally implemented in Java (a programming language). Due to language limitations, frequent writing will cause the CPU (Central Processing Unit) to be too high, making it impossible to write quickly. In addition, it will put a lot of pressure on the Elasticsearch cluster, making it impossible to provide stable services.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention
[0005] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.
[0006] Some embodiments of the present disclosure propose a data processing method, a data processing device, an electronic device, a computer-readable medium, and a computer program product to solve one or more of the technical problems mentioned in the above background technology section.
[0007] In a first aspect, some embodiments of the present disclosure provide a data processing method, including: in response to receiving an information change message from a database, determining whether a message matching the identifier of the information change message is received within a specified time period, wherein message marking is performed when the information change message is generated; in response to determining that a matching message is not received, determining the key-value group of the target business line to which the information change message belongs in a cache database based on the business line identifier in the information change message; based on the object identifier in the information change message, caching the data in the information change message in the key-value group of the target business line on the key corresponding to the object identifier; in response to reaching a set time length, generating message data representing attribute information changes based on the key-value data of each business line in the cache database; processing the attribute information in the message data, storing the processed attribute information in a search engine cluster, and deleting the data related to the message data in the cache database.
[0008] In some embodiments, the method further includes: in response to determining that a message matching the identifier of the information change message is received within a specified time period, determining that the information change message is a retry message, and filtering out the information change message without processing it.
[0009] In some embodiments, based on the business line identifier in the information change message, the key value group of the target business line to which the information change message belongs in the cache database is determined, including: in response to determining that the business line indicated by the business line identifier in the information change message exists among the business lines divided in the cache database, the business line is determined as the target business line to which the information change message belongs; in response to determining that the business line indicated by the business line identifier in the information change message does not exist among the business lines divided in the cache database, the information change message is intercepted.
[0010] In some embodiments, based on the key-value data of each business line in the cache database, message data representing changes in attribute information are generated, including: for the key-value data in each business line in the cache database, according to the information of the business line and the number of keys in the business line, the name of each key representing the change in attribute information in the business line is generated as message data, and multiple message data for the business line are obtained; and the multiple message data for each business line obtained are sent to a distributed message queue.
[0011] In some embodiments, attribute information in message data is processed and the processed attribute information is stored in a search engine cluster, including: in response to a consumer determining that message data exists and successfully acquiring a distributed lock, obtaining a set amount of message data from a distributed message queue, and obtaining attribute information indicated by the set amount of message data from a cache database; determining whether the attribute information contains information other than target attribute information, wherein the target attribute information is attribute information stored in the search engine cluster; in response to determining that information other than the target attribute information exists, filtering out the information, and storing the attribute information belonging to the target attribute information in the search engine cluster.
[0012] In some embodiments, the attribute information belonging to the target attribute information is stored in the search engine cluster, including: for the attribute information belonging to the target attribute information, each attribute information of the same object is deduplicated according to the event dimension to obtain at least one event data of the same object; and the event data of each object is stored in the search engine cluster.
[0013] In some embodiments, storing the event data of each object in the search engine cluster includes: storing the event data of each object in the search engine cluster, and using the object identifier of the object as a routing key, wherein the object identifier is also included in the request for querying data in the search engine cluster.
[0014] In some embodiments, the method also includes: regularly obtaining the data volume of each cluster in the search engine cluster; in response to determining that the data volume deviation between two clusters is greater than a set value, marking the two clusters; exchanging the business line data in each marked cluster to achieve data balance between clusters; and updating the correspondence between relevant clusters and business lines.
[0015] In some embodiments, the method further includes: in response to detecting that only data is stored in the cache database and no data is read from the cache database within a certain time period, determining whether the number of keys in the key-value group of each business line in the cache database reaches a set threshold; in response to determining that the set threshold is reached, stopping storing data in the business line that reaches the set threshold.
[0016] In some embodiments, the method further includes: in response to determining that the creation time of a key in a business line in a cache database has reached the effective time, clearing the data stored on the key, wherein, for the key-value group under the business line, when generating a key corresponding to an object identifier, the effective time of the key is set.
[0017] In a second aspect, some embodiments of the present disclosure provide a data processing device, including: a message determination unit, configured to determine whether a message matching the identifier of the information change message is received within a specified time period in response to receiving an information change message from a database, wherein message marking is performed when the information change message is generated; a business line determination unit, configured to determine, in response to determining that a matching message has not been received, the key value group of the target business line to which the information change message belongs in a cache database according to the business line identifier in the information change message; a data cache unit, configured to cache the data in the information change message in the key value group of the target business line, on the key corresponding to the object identifier, according to the object identifier in the information change message; a message generation unit, configured to generate message data representing the change of attribute information based on the key value data of each business line in the cache database in response to reaching a set time length; a data storage unit, configured to process the attribute information in the message data, store the processed attribute information in a search engine cluster, and delete data related to the message data in the cache database.
[0018] In some embodiments, the data processing device also includes a data filtering unit, which is configured to determine that the information change message is a retry message in response to determining that a message matching the identifier of the information change message is received within a specified time period, and filter out the information change message without processing it.
[0019] In some embodiments, the business line determination unit is further configured to, in response to determining that among the business lines divided in the cache database, there is a business line indicated by the business line identifier in the information change message, determine the business line as the target business line to which the information change message belongs; in response to determining that among the business lines divided in the cache database, there is no business line indicated by the business line identifier in the information change message, intercept the information change message.
[0020] In some embodiments, the message generation unit is further configured to generate the names of each key representing the change of attribute information in each business line in the cache database as message data based on the information of the business line and the number of keys in the business line, and obtain multiple message data for the business line; and send the obtained multiple message data of each business line to the distributed message queue.
[0021] In some embodiments, the data storage unit includes: a data acquisition subunit, configured to obtain a set amount of message data from a distributed message queue and obtain attribute information indicated by the set amount of message data from a cache database in response to a consumer determining the existence of message data and successfully acquiring a distributed lock; an information determination subunit, configured to determine whether there is information other than target attribute information in the attribute information, wherein the target attribute information is attribute information stored in a search engine cluster; an information storage subunit, configured to filter out information other than the target attribute information in response to determining the existence of the information, and store the attribute information belonging to the target attribute information in the search engine cluster.
[0022] In some embodiments, the information storage subunit is further configured to deduplicate attribute information belonging to the target attribute information according to the event dimension, and obtain at least one event data of the same object; and store the event data of each object in the search engine cluster.
[0023] In some embodiments, the information storage subunit is further configured to store the event data of each object in the search engine cluster, and use the object identifier of the object as a routing key, wherein the object identifier is also included in the request for querying data in the search engine cluster.
[0024] In some embodiments, the data processing device also includes a data balancing unit, which is configured to periodically obtain the data volume of each cluster in the search engine cluster; mark the two clusters in response to determining that the data volume deviation between two clusters is greater than a set value; exchange the business line data in each marked cluster to achieve data balance between clusters; and update the correspondence between relevant clusters and business lines.
[0025] In some embodiments, the data processing device also includes a data cache detection unit, which is configured to determine whether the number of keys in the key-value group of each business line in the cache database reaches a set threshold in response to detecting that only data is stored in the cache database and no data is read from the cache database within a certain time period; in response to determining that the set threshold is reached, stop storing data in the business line that reaches the set threshold.
[0026] In some embodiments, the data processing device also includes a data clearing unit, which is configured to clear the data stored on a key in a business line in response to determining that the creation time of the key in the cache database has reached the effective time, wherein for the key-value group under the business line, when generating a key corresponding to the object identifier, the effective time of the key is set.
[0027] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the data processing method described in any implementation method in the above-mentioned first aspect.
[0028] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the data processing method described in any one of the implementations in the first aspect above is implemented.
[0029] In a fifth aspect, some embodiments of the present disclosure provide a computer program product, including a computer program, which, when executed by a processor, implements the data processing method described in any implementation manner in the above-mentioned first aspect.
[0030] The above-mentioned embodiments of the present disclosure have the following beneficial effects: the data processing methods of some embodiments of the present disclosure can greatly reduce the amount of data that the search engine cluster needs to process, thereby saving the processing resources and storage resources of the cluster and ensuring the stability of the cluster. Specifically, the existing user information storage method is to store the generated user information directly in the search engine cluster. In a scenario where there are many dimensions of user information and the amount of data in each dimension is huge, the realization of quasi-real-time and frequent writing of data will inevitably lead to a higher CPU of the cluster, which in turn leads to the problem of instability of the cluster and query service.
[0031] Based on this, the data processing method in some embodiments of the present disclosure, when receiving an information change message, can first merge the attribute change messages of each user in the same business line in the user dimension and cache them in the cache database. That is, all the information of a user is allocated and cached on the same key. Then, the attribute information in the cache database is stored in the search engine cluster after processing. Compared with the method of directly storing data in the cluster, the method disclosed in the present disclosure can greatly reduce the amount of data by merging information in the user dimension. This can not only reduce the CPU pressure, but also save cluster resources, thereby ensuring the stability of the cluster. In addition, before processing the information change message, message power level anti-duplicate can also be performed. That is, determine whether multiple matching messages are received in a short time. That is, determine whether the received message is a retry message. This can avoid or reduce the repeated consumption of messages and improve data processing efficiency. It can also avoid writing retry messages into the cache database, reducing the storage space occupied. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0033] Figure 1 are flow charts of some embodiments of the data processing method disclosed herein;
[0034] Figure 2 It is a schematic diagram of application scenarios of some embodiments of the data processing method disclosed in the present invention;
[0035] Figure 3 are flow charts of other embodiments of the data processing method disclosed herein;
[0036] Figure 4 is a data flow diagram of some embodiments of the data processing method disclosed herein;
[0037] Figure 5 is a schematic diagram of the structure of some embodiments of the data processing device disclosed in the present invention;
[0038] Figure 6 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0039] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0040] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0041] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0042] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0043] With regard to the collection, storage, and use of user information involved in this disclosure, such as basic user information, user relationship information, and user task information, before performing the corresponding operations, the relevant organizations or individuals shall fulfill their obligations, including conducting personal information security impact assessments, fulfilling the obligation to inform the personal information subject, and obtaining the authorization and consent of the personal information subject in advance.
[0044] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0045] Figure 1 A process 100 of some embodiments of a data processing method according to the present disclosure is shown. The data processing method comprises the following steps:
[0046] Step 101, in response to receiving an information change message from a database, determining whether a message matching an identifier of the information change message is received within a specified time period.
[0047] It should be noted that various applications can be installed on the terminal devices used by users, such as shopping applications, instant messaging tools, browsers, etc. Users can operate on these applications through terminal devices. Therefore, various user information will be generated on these applications, such as user basic information (such as user account, name, etc.), user relationship information (such as user type, group, etc.), user feedback information, user reservation information, etc. These large amounts of user relationships and multi-dimensional user information are complex and cannot be simply supported by relational databases. In order to realize the search function of user information, the above information is usually stored in the Elasticsearch search engine.
[0048] Generally, any operation performed by a user on an application can be characterized as a change in the database data of the backend server. It is understandable that log information (binlog) is generally generated when database data is changed. The log information usually contains complete information about the changed data.
[0049] Based on this, in some embodiments, when receiving a database information change message, the execution subject (e.g., a server) of the data processing method can first determine whether a message matching the identifier of the information change message is received within a specified time period. Here, when data changes occur in the database, the above log information can be directly sent to the execution subject.
[0050] For example, Figure 2As shown, when the database data changes, the database information change message can be sent to the distributed message queue (JMQ, message middleware system) through DTS (Data Transmission Service). Among them, DTS usually distributes data according to the table id (identity) and hash value, so it can be guaranteed that there is an order between multiple concurrent changes of a piece of data. For example, if the id of a table is 1, name = a, if the name of this data is changed to b and c successively, then the received message must also be in the order of name = b, name = c, which can ensure that the name information will not be changed incorrectly. In this case, the execution subject can monitor the message. When a message is monitored in the message queue, the information change message can be obtained from the message queue.
[0051] It should be noted that during the message consumption process, repeated retry messages may be generated due to consumption errors and other reasons. For example, if 10 messages are pulled down and the 9th message is consumed with an error, all 10 messages will be retried for consumption. At this time, the first 9 messages are repeated consumption, which will waste system resources and bring additional data storage pressure. Therefore, here, in order to avoid repeated processing of messages, message marking can be performed when the above-mentioned information change message is generated. In this way, when the execution subject receives the information change message, it can first determine whether a message matching (such as the same) the identifier of the information change message is received within a specified time period (such as 5 minutes). In other words, first determine whether multiple information change messages with the same identifier are received in a short period of time. If no message matching the identifier is received, the execution subject can continue to execute step 102.
[0052] Optionally, if a message matching the identifier of the information change message is received within a specified time period, the information change message can be determined to be a retry message. At this time, the execution subject can filter out the information change message and not process it. In other words, if multiple messages contain the same identifier, they can be identified as the same message and directly filtered out. This can avoid subsequent repeated processing of multiple identical messages and save system resources. And through filtering processing, these retry messages can be avoided from being cached in the cache database, thereby reducing the amount of data and storage space occupied by the cache database.
[0053] Step 102, in response to determining that no matching message is received, determining the key value group of the target business line to which the information change message belongs in the cache database according to the business line identifier in the information change message.
[0054] It is understandable that, normally, a large number of information change messages about the same customer are received every second. In the disclosed embodiment, after receiving the change message, it is not written directly to Elasticsearch, but is first placed in the cache database redis. In the cache database, users can be grouped and multiple messages from the same user can be merged into one piece of data. Then write one piece of data to Elasticsearch to reduce the tps (Transaction PerSecond, the number of messages processed per second) of Elasticsearch. That is, data is stored in the user dimension. Specifically, in the cache database, the maximum number of rediskeys can be configured for each business line to prevent a business line from having too much data, causing hot key problems and affecting throughput. The business line here can be understood as the smallest business operation unit, such as a gold bar business line, an IOU business line, etc. Normally, data between different business lines is generally isolated, so you only need to consider how to handle user information under the same business line.
[0055] In some embodiments, if it is determined that no matching message has been received, the execution entity may process the received information change message. That is, according to the business line identifier in the information change message, determine the key value group of the target business line to which the information change message belongs in the cache database (redis). As an example, if it is determined that among the business lines divided in the cache database, there is a business line indicated by the business line identifier in the information change message, then the business line can be determined as the target business line to which the information change message belongs. Among them, the business line identifier is not limited here, and can be at least one of numbers, letters, text, symbols, etc.
[0056] For another example, if it is determined that the business line indicated by the business line identifier in the information change message does not exist in the business lines divided in the cache database, it means that the message is unnecessary data. Then the information change message can be intercepted. For example, 1,000 data change messages are received. Among them, 100 messages come from non-business concern tables, so these 100 messages can be intercepted. This can further reduce the amount of data that needs to be processed.
[0057] Step 103: based on the object identifier in the information change message, cache the data in the information change message in the key-value group of the target business line, on the key corresponding to the object identifier.
[0058] In some embodiments, when the target business line is determined, the execution subject can further determine the key under the target business line based on the object (user) identifier in the information change message. The object identifier here is also not limited. Specifically, the execution subject can determine the key corresponding to the object identifier in the key value group of the target business line. Thereby, the data in the information change message is cached on the key. That is, the data in the information change message is used as the value of the key. In other words, after receiving the data, the data can be assigned to a fixed key according to the business line. And it can be based on the user dimension (such as the hash value of the user number) to ensure that the information of the same user can be assigned to the same key. In this way, data stream merging can be achieved, the amount of data can be reduced, and the frequency of subsequent data writing to Elasticsearch can be reduced.
[0059] Step 104 , in response to reaching a set duration, based on the key-value data of each business line in the cache database, generates message data representing a change in attribute information.
[0060] In some embodiments, in order to avoid frequently writing data to Elasticsearch, the execution entity can set a scheduling task to regularly write the data in the cache database to the search engine cluster. Here, in response to determining that the set duration has been reached, the execution entity can generate message data representing the change of attribute information based on the key-value data of each business line in the cache database. As an example, a sliding time window (such as 1 second) can be set. When the execution frequency of the schedule reaches 1 second, the execution entity can obtain all business line data in the cache database that is within the sliding time window, such as the data written to the cache database within the 1 second. At this time, the execution entity can use each pair of key-value data under each business line as message data representing the change of attribute information.
[0061] Optionally, for the key-value data in each business line in the cache database, the execution subject can generate the name of each key in the business line that represents the change of attribute information according to the information of the business line and the number of keys in the business line. And the name of each key can be used as a message data respectively, so as to obtain multiple message data of the business line. Then, the obtained multiple message data of each business line can be sent to the distributed message queue for consumption processing by consumers.
[0062] As an example, the execution entity can obtain all business lines and the maximum number of keys corresponding to the business lines in the cache database every 1 second. All keys are generated based on the business line information and the maximum number of keys. For example, the maximum number of keys for the business line 100 is 2. Then the following two keys will be generated: binlog_redis_100_0 and binlog_redis_100_1. Among them, 100 is the business line ID; 0 and 1 are hash key suffixes used to prevent large keys and hot keys. And the generated keys are sent to the distributed message queue.
[0063] Step 105: Process the attribute information in the message data, store the processed attribute information in the search engine cluster, and delete the data related to the message data in the cache database.
[0064] It is understandable that, in order to avoid repeated message processing, the execution subject can also determine whether a message matching the identifier of the message data is received within a specified time period before processing the message data. If there is a message with a matching identifier, the message data can be filtered out and not processed. This can avoid storing retry messages in the search engine cluster, reduce the amount of data and pressure on the cluster, and improve overall performance and stability.
[0065] In some embodiments, if it is determined that there is no matching message, the execution subject may process the attribute information in the message data. And the processed attribute information may be stored in the search engine cluster. At the same time, the data related to the message data in the cache database may be deleted. The processing method here may be set according to the actual situation. For example, the clue data may be merged and deduplicated.
[0066] It should be noted that the following phenomenon usually exists: one user browsing one product corresponds to one event. When the user refreshes the page, multiple events (assuming N) will be repeated in a short period of time. And one event generally corresponds to X clues. At this time, when it is hit in MySQL (relational database management system), there will be X pieces of data, but the search terms corresponding to these X pieces of data are exactly the same (number of refreshes (number of repeated events) N*number of single event clues X). So for the Elasticsearch cluster used for search, these N*X pieces of data are completely duplicate and avoidable data.
[0067] To this end, X clues of a user can be deduplicated in the event dimension, that is, N*X message data can be reduced to 1 message. This can further reduce the number of change events that the Elasticsearch cluster needs to process. In other words, for the attribute information in each message data, the execution subject can deduplicate the attribute information of the same object according to the event dimension, thereby obtaining at least one event data of the same object. Then, the event data of each object can be stored in the search engine cluster.
[0068] Optionally, the execution entity may also Figure 3 The process shown in the figure is to process the message data, which will not be described in detail here.
[0069] It can be seen from the above description that the data processing method in some embodiments of the present disclosure, when receiving an information change message, can first merge the attribute change messages of each user in the same business line in the user dimension and cache them in the cache database. That is, all the information of a user is allocated and cached on the same key. Then, the attribute information in the cache database is stored in the search engine cluster after processing. Compared with the method of directly storing data in the cluster, the method disclosed in the present disclosure can greatly reduce the amount of data by merging information in the user dimension. This can not only reduce the CPU pressure, but also save cluster resources, thereby ensuring the stability of the cluster. In addition, before processing the information change message, message power level anti-duplicate can also be performed. That is, determine whether multiple matching messages are received in a short time. That is, determine whether the received message is a retry message. This can avoid or reduce the repeated consumption of messages and improve data processing efficiency. It can also avoid writing retry messages into the cache database, reducing the storage space occupied.
[0070] In some embodiments, for Figure 2 In the application scenario shown, when the Elasticsearch cluster fails or network jitter causes temporary unavailability, if the received data is directly written (pushed) to redis infinitely, it often causes redis data to only enter the queue but not exit the queue, which in turn causes the redis storage space to be full, triggering the redis elimination strategy. This leads to a series of problems such as the inability to ensure the final consistency of MySQL and Elasticsearch data, and the unavailability of redis reading and writing. For this reason, the embodiment of the disclosure can also set a redis overload protection mechanism.
[0071] As an example, if the execution subject detects that only data is stored in the cache database within a certain period of time, and no data is read from the cache database, it can be determined whether the number of keys in the key-value group of each business line in the cache database has reached the set threshold. If it is determined that the set threshold has been reached, it can stop storing data in the business line that has reached the set threshold. In other words, when the number of key elements of a business line in redis exceeds the set threshold, it can be prohibited to push data into redis, thereby avoiding the problem of too much data filling up redis and ensuring its high availability.
[0072] In addition, it should be noted that the Redis key is designed based on the effective business line. Only when the business line is effective, data will be written to the business line key. Only when the business line is effective, data will be read from the business line key. If the business line fails when the business line finishes writing data, the data written before will never have a chance to be dequeued and will remain in Redis. This will cause certain pressure on Redis performance and storage space.
[0073] In order to improve the storage space utilization of the cache database, in some application scenarios, the execution entity can also identify useless redis keys. Here, if it is determined that the creation time of a key in a business line in the cache database has reached the effective time, the data stored on the key can be cleared. The creation time is usually the time from the creation time to the current time. Among them, for the key-value group under the business line, when generating the key corresponding to the object identifier, the effective time of the key can be set. In other words, the hash-based validity period can be set for the key under each business line. In this way, even if the business line is deactivated, the data will be automatically released. Thereby reducing the storage space occupied by invalid data.
[0074] Continue to refer Figure 3 , which shows a process 300 of some other embodiments of the data processing method according to the present disclosure. The data processing method may also include the following steps:
[0075] Step 301, in response to the consumer determining that message data exists and successfully acquiring a distributed lock, a set amount of message data is acquired from a distributed message queue, and attribute information indicated by the set amount of message data is acquired from a cache database.
[0076] In some embodiments, the execution subject of the data processing method can use the redis distributed lock to try to lock when the consumer determines that there is message data, for example, through message subscription. And when the distributed lock is successfully acquired, a set number of message data can be obtained from the distributed message queue, and the attribute information indicated by these message data can be obtained from the cache database. That is, the data of the value in the key-value data corresponding to the name of the key. The set number here is also not limited.
[0077] That is to say, if the lock is successfully acquired, consumption can continue. If it fails, it means that the last message has not been processed yet, so you can exit and wait for the next distribution. After successfully acquiring the lock, consider the redis performance and blocking situation, set a maximum of 1,000 data to be taken out of the message queue. Then process the 1,000 data (see the following steps for details).
[0078] Step 302: Determine whether the attribute information contains information other than the target attribute information.
[0079] In some embodiments, the execution subject may also identify invalid data in the attribute information, that is, determine whether the attribute information contains information other than the target attribute information, wherein the target attribute information is usually the attribute information required to be stored in the search engine cluster.
[0080] For example, if 1,000 attribute information change messages are received, and the attribute information contains five fields, a, b, c, d, and e, but the Elasticsearch cluster only needs four fields, a, b, c, and d. If 400 of the 1,000 change messages change the e field information, then for the Elasticsearch cluster, these 400 changes are invalid and can be ignored. This can reduce the pressure on Elasticsearch writing.
[0081] Step 303: In response to determining that information other than the target attribute information exists, the information is filtered out, and the attribute information belonging to the target attribute information is stored in the search engine cluster.
[0082] In some embodiments, when it is determined that there is information other than the target attribute information, the execution subject can filter out the invalid information, thereby storing the attribute information belonging to the target attribute information in the search engine cluster. After processing, the relevant data in the cache database can be deleted and the lock can be released.
[0083] Optionally, the execution subject may also merge and de-duplicate the attribute information belonging to the target attribute information. That is, for the attribute information belonging to the target attribute information, de-duplicate the attribute information of the same object according to the event dimension, thereby obtaining at least one event data of the same object. Then, the event data of each object is stored in the search engine cluster.
[0084] In some application scenarios, in order to further reduce the overall pressure of the search engine cluster and ensure the stability of the cluster service, the execution entity can use the object identifier of each object as the routing key when storing the event data of each object in the search engine cluster. The object identifier can also be included in the request to query the data in the search engine cluster.
[0085] Normally, most read scenarios will use user IDs (such as cusNo user code). Therefore, when writing data to the Elasticsearch cluster, the user ID can be used as the routing key (_routing). And it is required that scenarios involving data writing (insert) and updating (update) must have user IDs. That is, the read-write routing key design. When the user ID is used for query search, the request will be sent to the node of the specified Elasticsearch cluster. In this way, the Elasticsearch cluster coordination node does not need to perform operations such as aggregation and paging across machines, which can greatly reduce the pressure on the cluster CPU and improve the stability and performance of the Elasticsearch cluster.
[0086] It should be noted that if _routing is not specified, assuming that the Elasticsearch cluster has 30 nodes, then a request will be made to all 30 nodes at the same time. Each node needs to perform query and aggregation, and report the results to the coordination node for final paging. Therefore, the overall pressure on the cluster is relatively large.
[0087] Furthermore, the disclosed embodiment can also realize automatic calibration of cluster data offset. In actual application, if a business line has a surge or a sharp drop in data due to business fluctuations, it will inevitably lead to uneven data distribution among clusters. This will cause high CPU pressure in some clusters, low cluster stability and performance, and CPU idleness regardless of cluster. For this reason, the data self-balancing logic between Elasticsearch clusters is added.
[0088] Specifically, the execution entity can periodically obtain the data volume of each cluster in the search engine cluster. If it is determined that the data volume deviation between two clusters is greater than a set value, the two clusters can be marked. In addition, the business line data in each marked cluster can be exchanged to achieve data balance between clusters. And the corresponding relationship between the relevant clusters and business lines is updated.
[0089] That is to say, in the daily data writing process, the total amount of data in each cluster can be captured regularly. If the cluster data is seriously uneven, such as when the deviation of the total amount of data is greater than 0.5 times the low data volume (or high data volume), the cluster will be marked as a cluster with data offset problems. In addition, the data between clusters will be automatically balanced during low business peaks.
[0090] Specifically, first calculate the data proportion of each business line in the cluster with data offset problem. Then calculate the business line to be replaced. For example, in cluster one, the gold bar business line has 5 million data and the white bar business line has 4 million data; in cluster two, the Jing card has 1 million data and the insurance business line has 1 million data. Then, based on the statistical results of the business line data, the business line that needs to be exchanged can be identified. For example, for the above example, the gold bar business line in cluster one can be swapped with the Jing card business line in cluster two. After the swap, there are 5 million data in cluster one and 6 million data in cluster two. The data volume of the two clusters is relatively balanced, and the cluster resource utilization and stability are maximized. Finally, the mapping relationship between the business line and the cluster is changed to the matching correspondence after the business line is replaced. When new data is written subsequently, the latest matching correspondence shall prevail, so that the pressure of each cluster is evenly distributed, avoiding resource waste and performance stability problems.
[0091] The data processing method of the disclosed embodiment further enriches and improves the process of message data processing and storage in the search engine cluster. Through a series of processes such as invalid data filtering and user clue data merging and deduplication, the amount of data that the search engine cluster needs to process can be further reduced. This can effectively alleviate cluster pressure and save cluster resources. In the scenario of massive data and frequent changes, the stability of the Elasticsearch cluster and services can be guaranteed, and the user experience can be improved.
[0092] Through the above optimization means and the data processing method of the disclosed embodiment, the frequently changed data of the same customer will be merged and deduplicated in the user dimension and written to Elasticsearch. The detailed data flow process can be as follows: Figure 4 As shown in the figure, this can greatly reduce the pressure on Elasticsearch writing and improve the stability of the elasticsearch cluster and the corresponding query service. It can also help reduce data delays during data synchronization and improve the overall stability and performance of the system.
[0093] It should be noted that the execution subject in the embodiment of the present disclosure may be Figure 2The system server shown may also be a server of each subsystem in the system, such as an application server, a search engine server, etc. The system may control the subsystem server to execute the above processing by sending instructions to each subsystem.
[0094] Further references Figure 5 , as a response to the above Figures 1 to 4 In order to realize the processing method shown in the figure, the present disclosure provides some embodiments of a data processing device. These processing device embodiments are Figures 1 to 4 The data processing device can be specifically applied to various electronic devices.
[0095] like Figure 5 As shown, the data processing device 500 of some embodiments may include: a message determination unit 501, configured to determine whether a message matching the identifier of the information change message is received within a specified time period in response to receiving an information change message from a database, wherein message marking is performed when the information change message is generated; a business line determination unit 502, configured to determine the key value group of the target business line to which the information change message belongs in the cache database according to the business line identifier in the information change message in response to determining that no matching message has been received; a data cache unit 503, configured to cache the data in the information change message in the key value group of the target business line, on the key corresponding to the object identifier, according to the object identifier in the information change message; a message generation unit 504, configured to generate message data representing the change of attribute information based on the key value data of each business line in the cache database in response to reaching a set time length; a data storage unit 505, configured to process the attribute information in the message data, store the processed attribute information in the search engine cluster, and delete the data related to the message data in the cache database.
[0096] In some embodiments, the data processing device 500 may also include a data filtering unit (not shown in the figure), which is configured to, in response to determining that a message matching the identifier of the information change message is received within a specified time period, determine that the information change message is a retry message, and filter out the information change message without processing it.
[0097] In some embodiments, the business line determination unit 502 can be further configured to, in response to determining that among the business lines divided in the cache database, there is a business line indicated by the business line identifier in the information change message, determine the business line as the target business line to which the information change message belongs; in response to determining that among the business lines divided in the cache database, there is no business line indicated by the business line identifier in the information change message, intercept the information change message.
[0098] In some embodiments, the message generation unit 504 can be further configured to generate the names of each key representing the change of attribute information in each business line in the cache database as message data based on the information of the business line and the number of keys in the business line, and obtain multiple message data for the business line; and send the obtained multiple message data of each business line to the distributed message queue.
[0099] In some embodiments, the data storage unit 505 may include: a data acquisition subunit (not shown in the figure), configured to obtain a set amount of message data from the distributed message queue and obtain attribute information indicated by the set amount of message data from the cache database in response to the consumer determining the existence of message data and successfully acquiring a distributed lock; an information determination subunit (not shown in the figure), configured to determine whether there is information other than target attribute information in the attribute information, wherein the target attribute information is attribute information stored in the search engine cluster; an information storage subunit (not shown in the figure), configured to filter out the information in response to determining the existence of information other than the target attribute information, and store the attribute information belonging to the target attribute information in the search engine cluster.
[0100] In some embodiments, the information storage subunit can be further configured to deduplicate the attribute information belonging to the target attribute information according to the event dimension, and obtain at least one event data of the same object; and store the event data of each object in the search engine cluster.
[0101] In some embodiments, the information storage subunit may be further configured to store the event data of each object in the search engine cluster and use the object identifier of the object as a routing key, wherein the object identifier is also included in the request for querying data in the search engine cluster.
[0102] In some embodiments, the data processing device 500 may also include a data balancing unit (not shown in the figure), which is configured to periodically obtain the data volume of each cluster in the search engine cluster; in response to determining that the data volume deviation between two clusters is greater than a set value, mark the two clusters; exchange the business line data in each marked cluster to achieve data balance between clusters; and update the correspondence between related clusters and business lines.
[0103] In some embodiments, the data processing device 500 may also include a data cache detection unit (not shown in the figure), which is configured to determine whether the number of keys in the key-value group of each business line in the cache database reaches a set threshold in response to detecting that only data is stored in the cache database and no data is read from the cache database within a certain time period; in response to determining that the set threshold is reached, stop storing data in the business line that reaches the set threshold.
[0104] In some embodiments, the data processing device 500 may also include a data clearing unit (not shown in the figure), which is configured to clear the data stored on a key in a business line in response to determining that the creation time of the key in the cache database has reached the effective time, wherein for the key-value group under the business line, when generating a key corresponding to the object identifier, the effective time of the key is set.
[0105] It is understood that the units described in the data processing device 500 are similar to those described in the reference Figures 1 to 4 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the data processing device 500 and the units included therein, and will not be described in detail here.
[0106] Reference below Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing some embodiments of the present disclosure. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0107] like Figure 6 As shown, the electronic device 600 may include a processing device 601 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0108] Typically, the following devices may be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 607 including, for example, a speaker, a vibrator, etc.; storage devices 608 including, for example, a disk, a hard disk, etc.; and communication devices 609. The communication devices 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 6 Each block shown in the figure may represent one device, or may represent multiple devices as required.
[0109] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.
[0110] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0111] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0112] The computer-readable medium may be included in the electronic device; or it may exist independently without being assembled into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: in response to receiving an information change message from a database, determines whether a message matching the identifier of the information change message is received within a specified time period, wherein message marking is performed when the information change message is generated; in response to determining that a matching message is not received, determines the key value group of the target business line to which the information change message belongs in the cache database according to the business line identifier in the information change message; in response to determining that no matching message is received, determines the key value group of the target business line to which the information change message belongs in the cache database according to the object identifier in the information change message; in response to reaching a set time length, generates message data representing attribute information changes based on the key value data of each business line in the cache database; processes the attribute information in the message data, stores the processed attribute information in the search engine cluster, and deletes the data related to the message data in the cache database.
[0113] In addition, computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0114] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0115] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor, for example, may be described as: a processor including a message determination unit, a business line determination unit, a data cache unit, a message generation unit, and a data storage unit. Among them, the names of these units do not constitute a limitation on the unit itself in certain circumstances, for example, the message determination unit may also be described as "a unit for determining whether a message matching the identifier of an information change message is received within a specified time period".
[0116] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0117] Some embodiments of the present disclosure further provide a computer program product, including a computer program, which implements any of the above-mentioned data processing methods when executed by a processor.
[0118] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.
Claims
1. A data processing method, comprising: In response to receiving an information change message from a database, determining whether a message matching the identifier of the information change message is received within a specified time period, wherein message marking is performed when the information change message is generated; In response to determining that no matching message has been received, determining, according to the business line identifier in the information change message, a key value group of the target business line to which the information change message belongs in the cache database; According to the object identifier in the information change message, cache the data in the information change message in the key-value group of the target business line, on the key corresponding to the object identifier; In response to reaching a set time, generating message data representing a change in attribute information based on the key-value data of each business line in the cache database; The attribute information in the message data is processed, the processed attribute information is stored in a search engine cluster, and the data related to the message data in the cache database is deleted.
2. The data processing method according to claim 1, wherein: The method further comprises: In response to determining that a message matching the identifier of the information change message is received within a specified time period, determining that the information change message is a retry message, and filtering out the information change message without processing it.
3. The data processing method according to claim 1, wherein: The determining, according to the business line identifier in the information change message, a key value group of a target business line to which the information change message belongs in a cache database includes: In response to determining that among the business lines divided in the cache database, there is a business line indicated by the business line identifier in the information change message, determining the business line as the target business line to which the information change message belongs; In response to determining that the business line indicated by the business line identifier in the information change message does not exist among the business lines divided in the cache database, the information change message is intercepted.
4. The data processing method according to claim 1, wherein: The generating of message data representing attribute information change based on the key value data of each business line in the cache database includes: For the key-value data in each business line in the cache database, based on the information of the business line and the number of keys in the business line, the names of the keys representing the changes in the attribute information in the business line are generated as message data, thereby obtaining multiple message data of the business line; The obtained multiple message data of each business line are sent to the distributed message queue.
5. The data processing method according to claim 4, wherein: The processing of the attribute information in the message data and storing the processed attribute information in the search engine cluster includes: In response to the consumer determining that the message data exists and successfully acquiring the distributed lock, acquiring a set amount of message data from the distributed message queue, and acquiring attribute information indicated by the set amount of message data from the cache database; Determining whether there is information other than target attribute information in the attribute information, wherein the target attribute information is attribute information stored in the search engine cluster; In response to determining that information other than the target attribute information exists, the information is filtered out, and attribute information belonging to the target attribute information is stored in the search engine cluster.
6. The data processing method according to claim 5, wherein: The storing the attribute information belonging to the target attribute information in the search engine cluster includes: For the attribute information belonging to the target attribute information, each attribute information of the same object is deduplicated according to the event dimension to obtain at least one event data of the same object; The event data of each object is stored in the search engine cluster.
7. The data processing method according to claim 6, wherein: The storing the event data of each object in the search engine cluster includes: The event data of each object is stored in the search engine cluster, and the object identifier of the object is used as a routing key, wherein the object identifier is also included in a request for querying data in the search engine cluster.
8. The data processing method according to claim 1, wherein: The method further comprises: Regularly obtaining the data volume of each cluster in the search engine cluster; In response to determining that the data volume deviation between the two clusters is greater than a set value, marking the two clusters; Exchanging the business line data in each marked cluster to achieve data balance between clusters; and Update the correspondence between related clusters and business lines.
9. The data processing method according to claim 1, wherein: The method further comprises: In response to detecting that data is only stored in the cache database and no data is read from the cache database within a certain time period, determining whether the number of keys in the key-value group of each business line in the cache database reaches a set threshold; In response to determining that a set threshold is reached, storage of data to the business line that has reached the set threshold is stopped.
10. The data processing method according to any one of claims 1 to 9, wherein: The method further comprises: In response to determining that the creation time of a key in a business line in the cache database has reached the effective time, the data stored on the key is cleared, wherein, for the key-value group under the business line, when generating a key corresponding to the object identifier, the effective time of the key is set.
11. A data processing device, comprising: a message determination unit configured to determine, in response to receiving an information change message from a database, whether a message matching the identifier of the information change message is received within a specified time period, wherein message marking is performed when the information change message is generated; A business line determination unit, configured to determine, in response to determining that no matching message has been received, a key value group of a target business line to which the information change message belongs in a cache database according to the business line identifier in the information change message; a data caching unit configured to cache the data in the information change message in the key-value group of the target business line, on a key corresponding to the object identifier, according to the object identifier in the information change message; A message generating unit, configured to generate message data representing a change in attribute information based on the key value data of each business line in the cache database in response to reaching a set time length; The data storage unit is configured to process the attribute information in the message data, store the processed attribute information in the search engine cluster, and delete the data related to the message data in the cache database.
12. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method according to any one of claims 1 to 10.
13. A computer readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 10 is implemented.
14. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the data processing method according to any one of claims 1 to 10.